MeiliSearch

mirror of https://github.com/meilisearch/MeiliSearch synced 2025-02-26 20:21:32 +01:00

Author	SHA1	Message	Date
Louis Dureuil	ed19b7c3c3	Only reindex if the size increased	2024-09-03 12:07:59 +02:00
Louis Dureuil	1ac008926b	Add maxBytes parameter	2024-09-03 12:07:15 +02:00
Louis Dureuil	c49d892c82	Changes to prompt	2024-09-03 12:07:10 +02:00
Louis Dureuil	de962a26f3	New error type when maxBytes is null	2024-09-03 12:01:04 +02:00
Louis Dureuil	21296190a3	Reindex embedders	2024-09-02 13:00:53 +02:00
Louis Dureuil	4464d319af	Change default template to use the new facility	2024-09-02 11:30:59 +02:00
Louis Dureuil	580ea2f450	Pass the fields <-> ids map with metadata to render	2024-09-02 11:30:10 +02:00
Louis Dureuil	915cf4bae5	Add field.is_searchable property to fields	2024-09-02 11:28:53 +02:00
meili-bors[bot]	9a756cf2c5	Merge #4888 4888: bring back v1.10.0 into main r=Kerollmops a=ManyTheFish Co-authored-by: Louis Dureuil <louis@meilisearch.com> Co-authored-by: meili-bors[bot] <89034592+meili-bors[bot]@users.noreply.github.com> Co-authored-by: Tamo <tamo@meilisearch.com> Co-authored-by: ManyTheFish <many@meilisearch.com>	2024-08-27 14:02:08 +00:00
ManyTheFish	b12e997c8a	Add pinyin flag	2024-08-21 14:38:04 +02:00
ManyTheFish	8bf89ec394	Infer locales from index settings	2024-08-21 10:47:40 +02:00
meili-bors[bot]	ee62d9ce30	Merge #4845 4845: Fix perf regression facet strings r=ManyTheFish a=dureuill Benchmarks between v1.9 and v1.10 show a performance regression of about x2 (+3dB regression) for most indexing workloads (+44s for hackernews). [Benchmark interpretation in the engine weekly meeting](https://www.notion.so/meilisearch/Engine-weekly-4d49560d374c4a87b4e3d126a261d4a0?pvs=4#98a709683276450295fcfe1f8ea5cef3). - Initial investigation pointed to #4819 as the origin of the regression. - Further investigation points towards the hypernormalization of each facet value in `extract_facet_string_docids` - Most of the slowdown is in `normalize_facet_strings`, and precisely in `detection.language()`. This PR improves the situation (-10s compared with `main` for hackernews, so only +34s regression compared with `v1.9`) by skipping normalization when it can be skipped. I'm not sure how to fix the root cause though. Should we skip facet locale normalization for now? Cc `@ManyTheFish` --- Tentative resolution options: 1. remove locale normalization from facet. I'm not sure why this is required, I believe we weren't doing this before, so maybe we can stop doing that again. 2. don't do language detection when it can be helped: won't help with the regressions in benchmark, but maybe we can skip language detection when the locales contain only one language? 3. use a faster language detection library: `@Kerollmops` told me about https://github.com/quickwit-oss/whichlang which bolsters x10 to x100 throughput compared with whatlang. Should we consider replacing whatlang with whichlang? Now I understand whichlang supports fewer languages than whatlang, so I also suggest: 4. use whichlang when the list of locales is empty (autodetection), or when it only contains locales that whichlang can detect. If the list of locales contains locales that whichlang cannot detect, then use whatlang instead. --- > [!CAUTION] > this PR contains a commit that adds detailed spans, that were used to detect which part of `extract_facet_string_docids` was taking too much time. As this commit adds spans that are called too often and adds 7s overhead, it should be removed before landing. Co-authored-by: Louis Dureuil <louis@meilisearch.com> Co-authored-by: ManyTheFish <many@meilisearch.com>	2024-08-19 06:29:48 +00:00
ManyTheFish	0f965d3574	Remove hotloop's spans	2024-08-14 14:33:36 +02:00
ManyTheFish	ade54493ab	Only detect language for a facet if several locales have been specified by the user in the settings	2024-08-14 12:03:52 +02:00
Louis Dureuil	c3cdc407ec	Avoid unnecessary clone()	2024-08-08 14:57:02 +02:00
Louis Dureuil	2f10273d14	Group by normalized values, make sure you don't remove a value where there remains at still one value that normalizes towards it	2024-08-08 14:02:53 +02:00
Louis Dureuil	e3ef0ae19e	also intersect the universe for searchOnAttributes	2024-08-06 14:06:56 +02:00
meili-bors[bot]	57f7af77c7	Merge #4846 4846: Add OpenAI tests r=dureuill a=dureuill # Pull Request ## Related issue Part of fixing #4757 ## What does this PR do? - OpenAI embedder: don't pass apiKey when it is empty (slightly improves error messages) - rest embedder and rest-based embedders: specialize the authorization denied error message depending on the configuration source - fix existing tests - Adds assets containing prerecorded texts to embed and the embeddings obtained from OpenAI - Adds an asset containing a tokenized long document and the embedding obtained from OpenAI for this token - Uses the wiremock crate to mock the OpenAI API: parse the openai request, lookup the response in assets, craft an openai response Co-authored-by: Louis Dureuil <louis@meilisearch.com>	2024-08-05 10:49:28 +00:00
Louis Dureuil	e64d0e0ca8	use insert instead of push for bitmaps	2024-08-01 18:32:45 +02:00
Louis Dureuil	9ef710cad4	Use wrapper that forces the desired date format	2024-07-31 17:12:19 +02:00
Louis Dureuil	5aa6cb3600	Specialize authorized error message depending on config source	2024-07-31 15:03:44 +02:00
Louis Dureuil	9b7764575b	openai: don't pass apiKey when it is empty	2024-07-31 15:03:44 +02:00
Louis Dureuil	0e68718027	Add detailed spans	2024-07-31 13:05:47 +02:00
Louis Dureuil	7c3fc8c655	Split settings and document facet string extractions	2024-07-31 10:57:46 +02:00
Louis Dureuil	8acd3f50bb	skip normalization when the locales and values are the same	2024-07-31 09:53:00 +02:00
Tamo	d262b1df32	craft an API over the Shared Server and Shared index to avoid hard to debug mistakes	2024-07-30 14:24:57 +02:00
meili-bors[bot]	c2c1ba39ee	Merge #4826 4826: Update Charabia v0.9.0 r=dureuill a=ManyTheFish # Pull Request ## Related Changelog https://github.com/meilisearch/charabia/releases/tag/v0.9.0 ## Notable Change for Meilisearch Adds all math symbols from https://www.compart.com/en/unicode/category/Sm to the default separator list. Co-authored-by: ManyTheFish <many@meilisearch.com>	2024-07-25 14:08:38 +00:00
ManyTheFish	35567b2137	Update Charabia v0.9.0	2024-07-25 16:02:14 +02:00
Louis Dureuil	d4ea7cc2a9	fix clippy 👉👈	2024-07-25 12:10:32 +02:00
Louis Dureuil	2413592bbf	Display docid when there are documents without manual embeddings for a manual embedder	2024-07-25 12:10:32 +02:00
Louis Dureuil	553440632e	Introduce Setting::some_or_not_set	2024-07-25 12:01:52 +02:00
Louis Dureuil	7a347966da	Allow explicit `dimensions` for ollama	2024-07-25 12:01:51 +02:00
Louis Dureuil	4654d51e05	Add custom headers for REST embedder	2024-07-25 12:01:51 +02:00
ManyTheFish	a918561ac1	Fix PR comments	2024-07-25 10:52:56 +02:00
ManyTheFish	70d71581ee	fix clippy	2024-07-25 10:52:56 +02:00
ManyTheFish	04fa44e7eb	Implement localized attributes settings	2024-07-25 10:51:27 +02:00
ManyTheFish	90c0a6db7d	Implement localized search	2024-07-25 10:51:27 +02:00
ManyTheFish	cc02920f2b	Update charabia	2024-07-25 10:51:27 +02:00
Tamo	988552e178	add tests on the rest embedder	2024-07-24 14:34:17 +02:00
Louis Dureuil	0d8199f3b7	Change parameters in milli settings	2024-07-24 14:34:17 +02:00
Louis Dureuil	4b74803dae	Change parameters in vector settings	2024-07-24 14:34:17 +02:00
Louis Dureuil	d731fa661b	ollama and openai use new EmbedderOptions	2024-07-24 14:34:17 +02:00
Louis Dureuil	a1beddd5d9	rest embedder: use json_template	2024-07-24 14:34:17 +02:00
Louis Dureuil	4109182ca4	Add json_template module	2024-07-24 14:34:12 +02:00
Louis Dureuil	1a297c048e	Error changes	2024-07-24 14:34:12 +02:00
Louis Dureuil	303e601b87	HuggingFace: Clearer error message when a model is not supported	2024-07-23 15:13:22 +02:00
meili-bors[bot]	ea73615abf	Merge #4804 4804: Implements the experimental contains filter operator r=irevoire a=irevoire # Pull Request Related PRD: (private link) https://www.notion.so/meilisearch/Contains-Like-Filter-Operator-0d8ad53c6761466f913432eb1d843f1e Public usage page: https://meilisearch.notion.site/Contains-filter-operator-usage-3e7421b0aacf45f48ab09abe259a1de6 ## Related issue Fixes https://github.com/meilisearch/meilisearch/issues/3613 ## What does this PR do? - Extract the contains operator from this PR: https://github.com/meilisearch/meilisearch/pull/3751 - Gate it behind a feature flag - Add tests Co-authored-by: Tamo <tamo@meilisearch.com>	2024-07-17 15:47:11 +00:00
Tamo	02c61eabfa	fix the range reported when the experimental feature has not been set	2024-07-17 16:54:33 +02:00
Tamo	2af9481804	Implements the experimental contains filter operator«	2024-07-17 11:13:37 +02:00
Louis Dureuil	24240934f9	Improve errors when indexing documents with a user provided embedder	2024-07-16 13:39:01 +02:00

1 2 3 4 5 ...

2462 Commits