MeiliSearch

mirror of https://github.com/meilisearch/MeiliSearch synced 2025-07-04 20:37:15 +02:00

Author	SHA1	Message	Date
meili-bors[bot]	9a756cf2c5	Merge #4888 4888: bring back v1.10.0 into main r=Kerollmops a=ManyTheFish Co-authored-by: Louis Dureuil <louis@meilisearch.com> Co-authored-by: meili-bors[bot] <89034592+meili-bors[bot]@users.noreply.github.com> Co-authored-by: Tamo <tamo@meilisearch.com> Co-authored-by: ManyTheFish <many@meilisearch.com>	2024-08-27 14:02:08 +00:00
meili-bors[bot]	36d8684dc8	Merge #4881 4881: Infer locales from index settings r=curquiza a=ManyTheFish # Pull Request ## Related issue Fixes #4828 Fixes #4816 ## What does this PR do? - Add some test using `AttributesToSearchOn` - Make the search infer the language based on the index settings when the `locales` filed is not precise CI is now working: https://github.com/meilisearch/meilisearch/actions/runs/10490050545/job/29055955667 Co-authored-by: ManyTheFish <many@meilisearch.com>	2024-08-21 14:18:16 +00:00
ManyTheFish	b12e997c8a	Add pinyin flag	2024-08-21 14:38:04 +02:00
ManyTheFish	8bf89ec394	Infer locales from index settings	2024-08-21 10:47:40 +02:00
meili-bors[bot]	ee62d9ce30	Merge #4845 4845: Fix perf regression facet strings r=ManyTheFish a=dureuill Benchmarks between v1.9 and v1.10 show a performance regression of about x2 (+3dB regression) for most indexing workloads (+44s for hackernews). [Benchmark interpretation in the engine weekly meeting](https://www.notion.so/meilisearch/Engine-weekly-4d49560d374c4a87b4e3d126a261d4a0?pvs=4#98a709683276450295fcfe1f8ea5cef3). - Initial investigation pointed to #4819 as the origin of the regression. - Further investigation points towards the hypernormalization of each facet value in `extract_facet_string_docids` - Most of the slowdown is in `normalize_facet_strings`, and precisely in `detection.language()`. This PR improves the situation (-10s compared with `main` for hackernews, so only +34s regression compared with `v1.9`) by skipping normalization when it can be skipped. I'm not sure how to fix the root cause though. Should we skip facet locale normalization for now? Cc `@ManyTheFish` --- Tentative resolution options: 1. remove locale normalization from facet. I'm not sure why this is required, I believe we weren't doing this before, so maybe we can stop doing that again. 2. don't do language detection when it can be helped: won't help with the regressions in benchmark, but maybe we can skip language detection when the locales contain only one language? 3. use a faster language detection library: `@Kerollmops` told me about https://github.com/quickwit-oss/whichlang which bolsters x10 to x100 throughput compared with whatlang. Should we consider replacing whatlang with whichlang? Now I understand whichlang supports fewer languages than whatlang, so I also suggest: 4. use whichlang when the list of locales is empty (autodetection), or when it only contains locales that whichlang can detect. If the list of locales contains locales that whichlang cannot detect, then use whatlang instead. --- > [!CAUTION] > this PR contains a commit that adds detailed spans, that were used to detect which part of `extract_facet_string_docids` was taking too much time. As this commit adds spans that are called too often and adds 7s overhead, it should be removed before landing. Co-authored-by: Louis Dureuil <louis@meilisearch.com> Co-authored-by: ManyTheFish <many@meilisearch.com>	2024-08-19 06:29:48 +00:00
ManyTheFish	0f965d3574	Remove hotloop's spans	2024-08-14 14:33:36 +02:00
ManyTheFish	ade54493ab	Only detect language for a facet if several locales have been specified by the user in the settings	2024-08-14 12:03:52 +02:00
meili-bors[bot]	07c8ed0459	Merge #4864 4864: Don't remove facet value when multiple original values map to the same normalized value r=ManyTheFish a=dureuill # Pull Request ## Related issue Fixes #4860 > [!WARNING] > This PR contains a fix to the immediate issue, but it looks like the underlying data model is faulty: there is only one possible "original" value for each normalized value in a facet of a document, while because of array values (or manually written nested fields, if you're evil), it is technically possible to have multiple, distinct original values mapping to the same normalized value. Co-authored-by: Louis Dureuil <louis@meilisearch.com>	2024-08-13 14:04:17 +00:00
Louis Dureuil	c3cdc407ec	Avoid unnecessary clone()	2024-08-08 14:57:02 +02:00
Louis Dureuil	2f10273d14	Group by normalized values, make sure you don't remove a value where there remains at still one value that normalizes towards it	2024-08-08 14:02:53 +02:00
meili-bors[bot]	321639364f	Merge #4861 4861: Make sure the index scheduler never stops running r=irevoire a=irevoire # Pull Request ## Related issue Fixes https://github.com/meilisearch/meilisearch/issues/4748 ## What does this PR do? - Whatever happens, we always try to process tasks once every minute (if no tasks are enqueued that's practically free) Co-authored-by: Tamo <tamo@meilisearch.com>	2024-08-07 16:21:54 +00:00
Tamo	442d06dce7	ensure the run function doesn't panic even if the tick function does	2024-08-07 17:50:32 +02:00
Tamo	8f6a98df07	make sure the index scheduler never stops running	2024-08-07 17:06:43 +02:00
meili-bors[bot]	b44e17c4c3	Merge #4858 4858: also intersect the universe for searchOnAttributes r=irevoire a=dureuill # Pull Request ## Related issue Fixes #4857 ## What does this PR do? - intersect with the universe (which does not contain the filtered out ids) when looking up documents for words, even when using `searchOnAttributes` Co-authored-by: Louis Dureuil <louis@meilisearch.com>	2024-08-07 13:15:26 +00:00
Louis Dureuil	e3ef0ae19e	also intersect the universe for searchOnAttributes	2024-08-06 14:06:56 +02:00
meili-bors[bot]	57f7af77c7	Merge #4846 4846: Add OpenAI tests r=dureuill a=dureuill # Pull Request ## Related issue Part of fixing #4757 ## What does this PR do? - OpenAI embedder: don't pass apiKey when it is empty (slightly improves error messages) - rest embedder and rest-based embedders: specialize the authorization denied error message depending on the configuration source - fix existing tests - Adds assets containing prerecorded texts to embed and the embeddings obtained from OpenAI - Adds an asset containing a tokenized long document and the embedding obtained from OpenAI for this token - Uses the wiremock crate to mock the OpenAI API: parse the openai request, lookup the response in assets, craft an openai response Co-authored-by: Louis Dureuil <louis@meilisearch.com>	2024-08-05 10:49:28 +00:00
meili-bors[bot]	2d16d0aea1	Merge #4839 4839: In prometheus metrics return the route pattern instead of the real route when returning the HTTP requests total r=irevoire a=irevoire # Pull Request ## Related issue Fixes https://github.com/meilisearch/meilisearch/issues/4825 ## What does this PR do? - return the route pattern instead of the real route when returning the HTTP requests total Co-authored-by: Tamo <tamo@meilisearch.com>	2024-08-05 10:14:51 +00:00
meili-bors[bot]	c817718e07	Merge #4853 4853: Fix rhai deletion r=irevoire a=dureuill # Pull Request ## Related issue Fixes #4849 ## What does this PR do? - insert inside of the bitmap instead of pushing into it. Co-authored-by: Louis Dureuil <louis@meilisearch.com>	2024-08-01 16:34:31 +00:00
Louis Dureuil	e64d0e0ca8	use insert instead of push for bitmaps	2024-08-01 18:32:45 +02:00
Louis Dureuil	21aa430b5e	Fix openai tests	2024-07-31 17:57:55 +02:00
Louis Dureuil	8535dc0be2	Fix existing tests	2024-07-31 17:57:32 +02:00
Louis Dureuil	72b9005344	Redact uid for Value	2024-07-31 17:57:13 +02:00
meili-bors[bot]	420c33132c	Merge #4850 4850: Use a fixed date format regardless of features r=irevoire a=dureuill # Pull Request ## Related issue Fixes #4844 ## What does this PR do? Given the following script: ``` cargo run -- --db-path meili.ms sleep 3 curl -s -X POST http://127.0.0.1:7700/indexes -H 'Content-Type: application/json' --data-binary '{"uid": "movies", "primaryKey": "id"}' sleep 3 cargo run -p meilisearch --db-path meili.ms sleep 3 curl -s -X POST http://127.0.0.1:7700/indexes/movies/search -H 'Content-Type: application/json' --data-binary '{}' ``` - Before this PR, the final search returns a decoding error. - After this PR, the search completes successfully ### Technical standpoint This PR fixes two locations where the formatting of dates were dependent on the feature set of the `time` crate. 1. The `IndexStats` had two fields without the serialization format specified 2. More subtly, the index dates (`createdAt,` `updatedAt`) were using value remapping in the main DB to `SerdeJson<OffsetDateTime>`, which was using whatever default format was available. This was fixed by creating a local `OffsetDateTime` wrapper that would specify the serialization format Co-authored-by: Louis Dureuil <louis@meilisearch.com>	2024-07-31 15:32:26 +00:00
Louis Dureuil	9ef710cad4	Use wrapper that forces the desired date format	2024-07-31 17:12:19 +02:00
Louis Dureuil	48f7329a83	Specify index_mapper on `IndexStats`	2024-07-31 17:11:28 +02:00
Louis Dureuil	ab1ec9ca21	Add tokenized test	2024-07-31 15:03:45 +02:00
Louis Dureuil	9d6efd92d2	new assets for tokenized test	2024-07-31 15:03:45 +02:00
Louis Dureuil	abdb337fd6	Add openai tests	2024-07-31 15:03:45 +02:00
Louis Dureuil	1c755c8899	Add openai responses	2024-07-31 15:03:45 +02:00
Louis Dureuil	3a42c3134e	update tests after changing authorized error message	2024-07-31 15:03:45 +02:00
Louis Dureuil	5aa6cb3600	Specialize authorized error message depending on config source	2024-07-31 15:03:44 +02:00
Louis Dureuil	9b7764575b	openai: don't pass apiKey when it is empty	2024-07-31 15:03:44 +02:00
Louis Dureuil	0e68718027	Add detailed spans	2024-07-31 13:05:47 +02:00
Louis Dureuil	7c3fc8c655	Split settings and document facet string extractions	2024-07-31 10:57:46 +02:00
Louis Dureuil	8acd3f50bb	skip normalization when the locales and values are the same	2024-07-31 09:53:00 +02:00
meili-bors[bot]	25791e3f46	Merge #4836 4836: Attach declared localized-attributes subroutes r=dureuill a=dureuill RC.0 unexpectedly doesn't contain the `GET /indexes/{indexUid}/localized-attributes` and `PUT /indexes/{indexUid}/localized-attributes` subroute. This PR makes them available. Co-authored-by: Louis Dureuil <louis@meilisearch.com> Co-authored-by: Tamo <tamo@meilisearch.com>	2024-07-30 19:01:54 +00:00
meili-bors[bot]	866922ecc3	Merge #4808 4808: Make the tests run faster r=irevoire a=irevoire ## Index-Scheduler ### Only check the consistency of the index-scheduler on snapshots when running in release mode This saves 12s on the tests, and since the tests run in release mode in the CI, we don't lose any information. From 28s to 16s ### We were snapshotting the index for no reason in `advance_till`, I removed this call This saved an additional 8s on the tests, going from 16s to 8s. ---- After these two optimizations, the test suite as a whole executes 14% quicker ## Meilisearch integration tests While profiling this test suite, nothing stands out. The only noticeable thing is that we're losing most of our time creating and dropping threads. I made the theory that by sharing a single common instance between all integrations tests I would gain some time again. In https://github.com/meilisearch/meilisearch/pull/4808/commits/355a7acd1c85a1f728f31fb7f4e9c17233ea1a95 I saved another 15s by only testing this theory on the module that tests the error messages. But we can do it on many more tests. We must take care of not making any test flaky, though. ## Use two indexing threads By moving from one to two indexing threads, we gain an additional 30% in performance. # Conclusion ## Before The execution of the test suite was taking around: - 4m40s on my computer - 15 minutes on the debug CI with cache - 29 minutes on the Windows CI with cache ## After The execution of the test suite is taking around: - 2m20 on my computer - 8 minutes on the debug CI with cache - 29 minutes on the Windows CI with cache ## This means the test suite should now run ~50% faster on your computer; the CI may report errors twice faster, but we'll still wait for ~the same amount of time to merge a PR Co-authored-by: Tamo <tamo@meilisearch.com>	2024-07-30 15:11:30 +00:00
Tamo	f05ea04879	In prometheus metrics return the route pattern instead of the real route when returning the HTTP requests total	2024-07-30 16:24:49 +02:00
Tamo	b1b3a1a98b	add a get, set and put test for the localized attributes setting	2024-07-30 15:51:02 +02:00
meili-bors[bot]	143d6cde10	Merge #4835 4835: Log error from main using tracing r=irevoire a=dureuill Engine follow-up to https://github.com/meilisearch/meilisearch-support/issues/252#issuecomment-2251288276 (private link) > `@meilisearch/engine-team` we need to open a PR to tracing::error! when an error occurs in the Meilisearch main. It would be nice to have it included in the second RC <img width="1349" alt="Error logged when launching Meilisearch to import dump on path where the dump doesn't exist" src="https://github.com/user-attachments/assets/e5d2ae6e-f810-4029-9787-3b6ea9d47cfd"> --- <img width="1349" alt="Error logges when launching Meilisearch with a db path that is not writeable" src="https://github.com/user-attachments/assets/f672d78d-04b0-4d02-9402-259eaa6e2b62"> Co-authored-by: Louis Dureuil <louis@meilisearch.com>	2024-07-30 13:43:50 +00:00
Tamo	c457069367	ensure a test is 100% not flaky	2024-07-30 15:41:51 +02:00
Tamo	bb1283222e	make clippy happy	2024-07-30 15:10:56 +02:00
Tamo	7a5a38f870	fix a sync issue on empty indexes	2024-07-30 15:09:12 +02:00
Tamo	ded3cd0dd6	an additionnal 30% of perf for the tests	2024-07-30 15:03:54 +02:00
Tamo	68f885f1c4	fix two snapshots	2024-07-30 14:45:59 +02:00
Tamo	9372c34dab	prepare the tests to share indexes with api key	2024-07-30 14:34:11 +02:00
Tamo	6666c57880	reduce the number of thread spawned by milli	2024-07-30 14:34:10 +02:00
Tamo	b53a019b07	fix the initialization problem over the shared indexes with documents	2024-07-30 14:24:57 +02:00
Tamo	d262b1df32	craft an API over the Shared Server and Shared index to avoid hard to debug mistakes	2024-07-30 14:24:57 +02:00
Tamo	ed795bc837	fmt	2024-07-30 14:24:57 +02:00

1 2 3 4 5 ...

9832 commits