Speedup of first search results?

Hello,

I ran a few benchmarks and asked myself a question… #SOS-5166

My understanding: We activated the chunked search of GroupDocs.Search and I think, it will only speed up the search if there are multiple segments. In your documentation, you talk about tens or hundreds of thousands of files.

Your documentation: Search by chunks | GroupDocs

But what is still slow, are very short words (and because it’s probably 1 large segment in my benchmark index of 10k files, chunking brings no speed up). Benchmark:

Scenario Headers First hit All hits (median) (worst) Hits
SingleLetter [chunked] 40,907s 40,907s 40,996s 42,999s 45,519s 9.691
SingleLetter [one-shot] 41,809s 41,809s 41,956s 43,974s 44,723s 9.691
ShortWord [chunked] 2,467s 2,467s 2,623s 2,769s 3,111s 9.049
ShortWord [one-shot] 2,301s 2,301s 2,394s 2,631s 2,869s 9.049
MediumWord [chunked] 2,074s 2,074s 2,121s 2,352s 2,458s 5.262
MediumWord [one-shot] 2,381s 2,381s 2,417s 2,504s 2,607s 5.262
LongWord [chunked] 2,236s 2,237s 2,241s 2,435s 2,590s 505
LongWord [one-shot] 1,837s 1,838s 1,841s 2,115s 2,483s 505
VeryLongWord [chunked] 1,915s 1,915s 1,925s 2,437s 2,680s 1.662
VeryLongWord [one-shot] 1,855s 1,856s 1,863s 1,970s 2,061s 1.662
Phrase [chunked] 1,998s 1,998s 1,999s 2,162s 2,379s 137
Phrase [one-shot] 1,991s 1,991s 1,992s 2,145s 2,214s 137

When I see the “SingleLetter” case, I’m asking myself:
When I search for “a” or for all words containing “a”, couldn’t it be possible to do this on your side:

For segment 1, dot his:

  1. look for “a” in the term list and find the first 100 documents containing it.
  2. get the first 100 documents and return them
    …
    …
  3. get the last up to 100 documents for “a” and return them in the search
  4. find the next

Then, for segment 2, do these steps again…

Then for segment 3, then 4, …


Goal: For us, it’s about how long it takes to get the first search result. It’d be good to use the following code we already use to get a stream of search results immediately instead of waiting 40 seconds. (Getting 10 results immediately feels so much faster than waiting 40 seconds):

while (result.NextChunkSearchToken != null)
{
    result = index.SearchNext(result.NextChunkSearchToken);
    Console.WriteLine("Document count: " + result.DocumentCount);
    Console.WriteLine("Occurrence count: " + result.OccurrenceCount);
}

BTW: .Net offers IAsyncEnumerable in the newer versions, which is usually a good way to implement something like that in .Net (just want to mention it in case it might be useful. The “while (result.NextChunkSearchToken != null)” way is fine, too.)

Best regards,
jam

Hello @jamsharp

Search speed depends heavily on the type of search.
An exact word search offers the fastest performance.
Fuzzy search yields results somewhat more slowly.
Regex search is the slowest, as it involves exhaustively scanning the words in the index to match the regex pattern.
If you need to use wildcards, it is better to use wildcard search rather than regex:

Wildcard search is optimized for fast index lookups. Moreover, the fewer wildcard characters there are at the beginning of the pattern, the faster the search executes.
Tell us what types of search you use and what kind of queries you write.

Best regards,
Andrey Golubkov

Hello,

We understand that some search types are faster than others.

Our question is: If one of our users wants to search for all words containing “en” for example and when chunking search is activated, couldn’t the first results come faster than after like 30 seconds?

In our case, we translate a search " a word that contains ‘en’ " to this:

content:?(0~255)en?(0~255)

This type of query - where a large number of unknown characters precede known ones - takes the longest to execute. Conversely, if the known characters appear at the beginning of the pattern, the query executes very quickly.
Unfortunately, this is currently a system limitation due to the algorithms and data structures being used.
However, if you know, for example, that there are no more than three characters before ‘en’ in the word, the query will execute significantly faster.

Ok, so this would be a bigger change?

However, if you know, for example, that there are no more than three characters before ‘en’ in the word, the query will execute significantly faster.

What we offer the users in this case, is a LIKE search (a Contains-Search), so we don’t know the exact number of letters before and after. We figured out that 255 chars are the maximum and that

content:?(0~255)en?(0~255)

is the best way to represent such a “Contains” search with GroupDocs.

Hello @jamsharp
The upcoming release will also include optimizations for wildcard search.
Please wait for the release and check the search speed again.
The speed may well be satisfactory after the optimization.