java 在 lucene 中使用命中荧光笔

声明:本页面是StackOverFlow热门问题的中英对照翻译,遵循CC BY-SA 4.0协议,如果您需要使用它,必须同样遵循CC BY-SA许可,注明原文地址和作者信息,同时你必须将它归于原作者(不是我):StackOverFlow 原文地址: http://stackoverflow.com/questions/2409870/
Warning: these are provided under cc-by-sa 4.0 license. You are free to use/share it, But you must attribute it to the original authors (not me): StackOverFlow

提示:将鼠标放在中文语句上可以显示对应的英文。显示中英文
时间:2020-10-29 21:00:44  来源:igfitidea点击:

using hit highlighter in lucene

javalucenehit-highlighting

提问by Rohit Banga

I have two questions regarding hit highlighter provided with apache lucene:

我有两个关于 apache lucene 提供的命中荧光笔的问题:

  1. see thisfunction could you explain the use of token stream parameter.

  2. I have several large lucene document containing many fields and each field has some strings in it. Now I have found the most relevant document for a particular query. Now this document was found because several words in the query might have matched with the words in the document. I want to find out what words in the query caused this. So for this I plan to use Lucene Hit Highlighter. Example: if the query is "skin doctor delhi" and the document titled "dermatologist" contains the words "skin" and "doctor" then after hit highlighting i should be able to separate out "skin" and "doctor" from the query. I have been trying to write the code for this for several weeks now. Not able to get what i want. Could you help me please?

  1. 看到这个函数,你能解释一下令牌流参数的使用吗?

  2. 我有几个包含许多字段的大型 lucene 文档,每个字段中都有一些字符串。现在我找到了与特定查询最相关的文档。现在找到了这个文档,因为查询中的几个词可能与文档中的词匹配。我想找出查询中的哪些词导致了这种情况。因此,为此我计划使用 Lucene Hit Highlighter。示例:如果查询是“skin doctor delhi”并且标题为“dermatologist”的文档包含单词“skin”和“doctor”,那么在点击突出显示后,我应该能够从查询中分离出“skin”和“doctor”。几个星期以来,我一直在尝试为此编写代码。无法得到我想要的。请问你能帮帮我吗?

Thanks in advance.

提前致谢。

Update:

更新:

Current Approach: I create a query containing all the words in the document.

当前方法:我创建了一个包含文档中所有单词的查询。

Field[] field = doc.getFields("description");
String desc = "";
for (int j = 0; j < field.length; ++j) {
     desc += field[j].stringValue() + " ";
}

Query q = qp.parse(desc);
QueryScorer scorer = new QueryScorer(q, reader, "description");
Highlighter highlighter = new Highlighter(scorer);

String fragment = highlighter.getBestFragment(analyzer, "description", text);

It works for small documents but does not work for large documents. The following stacktrace is obtained.

它适用于小文件,但不适用于大文件。获得以下堆栈跟踪。

    org.apache.lucene.search.BooleanQuery$TooManyClauses: maxClauseCount is set to 1024
    at org.apache.lucene.search.BooleanQuery.add(BooleanQuery.java:152)
    at org.apache.lucene.queryParser.QueryParser.getBooleanQuery(QueryParser.java:891)
    at org.apache.lucene.queryParser.QueryParser.getBooleanQuery(QueryParser.java:866)
    at org.apache.lucene.queryParser.QueryParser.Query(QueryParser.java:1213)
    at org.apache.lucene.queryParser.QueryParser.TopLevelQuery(QueryParser.java:1167)
    at org.apache.lucene.queryParser.QueryParser.parse(QueryParser.java:182)

It is obvious that the approach is unreasonable for large documents. What should be done to correct this?

很明显,这种方法对于大文档是不合理的。应该怎么做才能纠正这个问题?

BTW I am using FuzzyQuery matching.

顺便说一句,我正在使用 FuzzyQuery 匹配。

采纳答案by Yuval F

EDIT: added some details about explain().

编辑:添加了一些关于解释()的细节。

Some general introduction: The Lucene Highlighter is meant to find text snippets from a hit document, and to highlight tokens matching the query.

一些一般性介绍:Lucene Highlighter 旨在从命中文档中查找文本片段,并突出显示与查询匹配的标记。

  1. Therefore, The TokenStream parameter is used to break the hit text into tokens. The highlighter's scorer then scores each token, in order to score fragments and choose snippets and tokens to be highlighted.
  2. I believe you are doing it wrong. If all you want to do is understand which query terms were matched in the document, you should use the explain()method. Basically, after you have instantiated a searcher, use:
  1. 因此,TokenStream 参数用于将命中文本分解为标记。高亮器的记分员然后对每个标记进行评分,以便对片段进行评分并选择要突出显示的片段和标记。
  2. 我相信你做错了。如果您只想了解文档中匹配了哪些查询词,则应使用解释()方法。基本上,在您实例化搜索器后,请使用:

Explanation expl = searcher.explain(query, docId);

Explanation expl = searcher.explain(query, docId);

String asText = expl.toString();

String asText = expl.toString();

String asHtml = expl.toHtml();

String asHtml = expl.toHtml();

docId is the raw document id from the search results.

docId 是来自搜索结果的原始文档 ID。

Only if you do need the snippets and/or highlights, you should use the Highlighter. If you still want to use the highlighter, follow Nicholas Hrychan's advice. Beware, though, as he describes the Lucene 2.4.1 API - If you use a more advanced version, you should use "QueryScorer" where he says "SpanScorer" .

只有当您确实需要片段和/或突出显示时,才应该使用荧光笔。如果您仍想使用荧光笔,请遵循Nicholas Hrychan 的建议。不过要小心,因为他描述了 Lucene 2.4.1 API - 如果你使用更高级的版本,你应该使用 "QueryScorer" 他说 "SpanScorer" 。