mirror of https://github.com/meilisearch/MeiliSearch synced 2025-06-14 20:11:38 +02:00

History

Jakub Jirutka 13f1277637 Allow to disable specialized tokenizations (again)

In PR #2773, I added the `chinese`, `hebrew`, `japanese` and `thai`
feature flags to allow melisearch to be built without huge specialed
tokenizations that took up 90% of the melisearch binary size.
Unfortunately, due to some recent changes, this doesn't work anymore.
The problem lies in excessive use of the `default` feature flag, which
infects the dependency graph.

Instead of adding `default-features = false` here and there, it's easier
and more future-proof to not declare `default` in `milli` and
`meilisearch-types`. I've renamed it to `all-tokenizers`, which also
makes it a bit clearer what it's about.

2023-05-04 15:45:40 +02:00

benches

Remove formating benchmark because they can't be isoloated easily anymore

2023-04-06 15:10:16 +02:00

scripts

Get rid of the threshold when comparing benchmarks

2022-04-19 15:39:58 +02:00

src

move the benchmarks to another crate so we can download the datasets automatically without adding overhead to the build of milli

2021-06-02 11:11:50 +02:00

.gitignore

add a gitignore to avoid pushing the autogenerated file

2021-06-02 17:03:30 +02:00

build.rs

retry downloading the benchmarks datasets

2022-08-17 19:25:05 +02:00

Cargo.toml

Allow to disable specialized tokenizations (again)

2023-05-04 15:45:40 +02:00

README.md

Update links of the docs

2023-05-03 19:14:57 +02:00

README.md

Benchmarks

Run the benchmarks

On our private server

The Meili team has self-hosted his own GitHub runner to run benchmarks on our dedicated bare metal server.

To trigger the benchmark workflow:

Go to the Actions tab of this repository.
Select the Benchmarks workflow on the left.
Click on Run workflow in the blue banner.
Select the branch on which you want to run the benchmarks and select the dataset you want (default: songs).
Finally, click on Run workflow.

This GitHub workflow will run the benchmarks and push the critcmp report to a DigitalOcean Space (= S3).

The name of the uploaded file is displayed in the workflow.

More about critcmp.

💡 To compare the just-uploaded benchmark with another one, check out the next section.

On your machine

To run all the benchmarks (~5h):

cargo bench

To run only the search_songs (~1h), search_wiki (~3h), search_geo (~20m) or indexing (~2h) benchmark:

cargo bench --bench <dataset name>

By default, the benchmarks will be downloaded and uncompressed automatically in the target directory.
If you don't want to download the datasets every time you update something on the code, you can specify a custom directory with the environment variable MILLI_BENCH_DATASETS_PATH:

mkdir ~/datasets
MILLI_BENCH_DATASETS_PATH=~/datasets cargo bench --bench search_songs # the four datasets are downloaded
touch build.rs
MILLI_BENCH_DATASETS_PATH=~/datasets cargo bench --bench songs # the code is compiled again but the datasets are not downloaded

Comparison between benchmarks

The benchmark reports we push are generated with critcmp. Thus, we use critcmp to show the result of a benchmark, or compare results between multiple benchmarks.

We provide a script to download and display the comparison report.

Requirements:

grep
curl
critcmp

List the available file in the DO Space:

./benchmarks/script/list.sh

songs_main_09a4321.json
songs_geosearch_24ec456.json
search_songs_main_cb45a10b.json

Run the comparison script:

# we get the result of ONE benchmark, this give you an idea of how much time an operation took
./benchmarks/scripts/compare.sh son songs_geosearch_24ec456.json
# we compare two benchmarks
./benchmarks/scripts/compare.sh songs_main_09a4321.json songs_geosearch_24ec456.json
# we compare three benchmarks
./benchmarks/scripts/compare.sh songs_main_09a4321.json songs_geosearch_24ec456.json search_songs_main_cb45a10b.json

Datasets

The benchmarks uses the following datasets:

smol-songs
smol-wiki
movies
smol-all-countries

Songs

smol-songs is a subset of the songs.csv dataset.

It was generated with this command:

xsv sample --seed 42 1000000 songs.csv -o smol-songs.csv

Download the generated smol-songs dataset.

Wiki

smol-wiki is a subset of the wikipedia-articles.csv dataset.

It was generated with the following command:

xsv sample --seed 42 500000 wiki-articles.csv -o smol-wiki-articles.csv

Download the smol-wiki dataset.

Movies

movies is a really small dataset we uses as our example in the getting started

Download the movies dataset.

All Countries

smol-all-countries is a subset of the all-countries.csv dataset It has been converted to jsonlines and then edited so it matches our format for the _geo field.

It was generated with the following command:

bat all-countries.csv.gz | gunzip | xsv sample --seed 42 1000000 | csv2json-lite | sd '"latitude":"(.*?)","longitude":"(.*?)"' '"_geo": { "lat": $1, "lng": $2 }' | sd '\[|\]|,$' '' | gzip > smol-all-countries.jsonl.gz

Download the smol-all-countries dataset.

README.md

Benchmarks

TOC

Run the benchmarks

On our private server

On your machine

Comparison between benchmarks

Datasets

Songs

Wiki

Movies

All Countries