Hello,
I ran a few benchmarks and asked myself a question… #SOS-5166
My understanding: We activated the chunked search of GroupDocs.Search and I think, it will only speed up the search if there are multiple segments. In your documentation, you talk about tens or hundreds of thousands of files.
Your documentation: Search by chunks | GroupDocs
But what is still slow, are very short words (and because it’s probably 1 large segment in my benchmark index of 10k files, chunking brings no speed up). Benchmark:
| Scenario | Headers | First hit | All hits | (median) | (worst) | Hits |
|---|---|---|---|---|---|---|
| SingleLetter [chunked] | 40,907s | 40,907s | 40,996s | 42,999s | 45,519s | 9.691 |
| SingleLetter [one-shot] | 41,809s | 41,809s | 41,956s | 43,974s | 44,723s | 9.691 |
| ShortWord [chunked] | 2,467s | 2,467s | 2,623s | 2,769s | 3,111s | 9.049 |
| ShortWord [one-shot] | 2,301s | 2,301s | 2,394s | 2,631s | 2,869s | 9.049 |
| MediumWord [chunked] | 2,074s | 2,074s | 2,121s | 2,352s | 2,458s | 5.262 |
| MediumWord [one-shot] | 2,381s | 2,381s | 2,417s | 2,504s | 2,607s | 5.262 |
| LongWord [chunked] | 2,236s | 2,237s | 2,241s | 2,435s | 2,590s | 505 |
| LongWord [one-shot] | 1,837s | 1,838s | 1,841s | 2,115s | 2,483s | 505 |
| VeryLongWord [chunked] | 1,915s | 1,915s | 1,925s | 2,437s | 2,680s | 1.662 |
| VeryLongWord [one-shot] | 1,855s | 1,856s | 1,863s | 1,970s | 2,061s | 1.662 |
| Phrase [chunked] | 1,998s | 1,998s | 1,999s | 2,162s | 2,379s | 137 |
| Phrase [one-shot] | 1,991s | 1,991s | 1,992s | 2,145s | 2,214s | 137 |
When I see the “SingleLetter” case, I’m asking myself:
When I search for “a” or for all words containing “a”, couldn’t it be possible to do this on your side:
For segment 1, dot his:
- look for “a” in the term list and find the first 100 documents containing it.
- get the first 100 documents and return them
…
… - get the last up to 100 documents for “a” and return them in the search
- find the next
Then, for segment 2, do these steps again…
Then for segment 3, then 4, …
Goal: For us, it’s about how long it takes to get the first search result. It’d be good to use the following code we already use to get a stream of search results immediately instead of waiting 40 seconds. (Getting 10 results immediately feels so much faster than waiting 40 seconds):
while (result.NextChunkSearchToken != null)
{
result = index.SearchNext(result.NextChunkSearchToken);
Console.WriteLine("Document count: " + result.DocumentCount);
Console.WriteLine("Occurrence count: " + result.OccurrenceCount);
}
BTW: .Net offers IAsyncEnumerable in the newer versions, which is usually a good way to implement something like that in .Net (just want to mention it in case it might be useful. The “while (result.NextChunkSearchToken != null)” way is fine, too.)
Best regards,
jam