Statement query method and device based on semantic aggregation degree and topic preposition degree

By introducing quantitative methods of semantic aggregation degree and topic preposition degree in text retrieval technology, the problem of bottleneck in text retrieval accuracy in the prior art is solved, and a higher accuracy rate and wider scope of application are achieved in the multi-alternative segment screening task.

CN119938841APending Publication Date: 2025-05-06CHENGDU HANLAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510015033.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

There are bottlenecks in improving the accuracy rate of existing text retrieval technologies, especially in the fields of medicine, law, etc., and it is difficult for the main search algorithm to significantly improve the accuracy rate in the multi-alternative segment screening task.

Method used

By introducing two new dimensions of semantic aggregation and topic preposition degree, this information is quantified using a naive method and fused into the main retrieval algorithm to assist in selecting segments that are more likely to contain the sentences to be queried from multiple alternative segments.

Benefits of technology

It effectively improves the accuracy of text retrieval, especially in the multi-alternative segment screening task, and expands the scope of application of auxiliary retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938841A_ABST
    Figure CN119938841A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data query processing, in particular to a statement query method and device based on a semantic aggregation degree and a topic preposition degree, and the statement query method comprises the following steps: obtaining a to-be-queried statement and an alternative text segment expressed by a natural language; performing word segmentation on the to-be-queried statement to obtain at least two target segmented words; searching all target segmented words in the alternative text segment, and determining the position of each target segmented word appearing in the alternative text segment for the first time; calculating a semantic aggregation degree and a topic preposition degree of the to-be-queried statement in the alternative text segment, and performing quantitative fusion on information of the semantic aggregation degree and the topic preposition degree; and based on the semantic aggregation degree and the topic preposition degree, assisting a main retrieval algorithm to select a text segment which more possibly contains the to-be-queried statement from a plurality of alternative text segments. According to the method, information, which is easy to neglect in text retrieval, in two dimensions of semantic aggregation degree and topic preposition degree is quantified in a naive mode, so that the accuracy of a main query algorithm in a multi-alternative text segment screening task is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data query processing, and in particular to a sentence query method and device based on semantic aggregation and subject fronting degree. Background Art

[0002] In some fields that have high requirements for the accuracy of text retrieval tasks (such as medicine, law and other fields with high decision-making risks), after the accuracy is raised to a certain level (such as 80%) by using various auxiliary retrieval technologies, a bottleneck in the accuracy will appear. That is, no matter how the various technologies or parameters are adjusted, the accuracy will not be significantly improved. From the perspective of information theory, this phenomenon is essentially that the information contained in the text and used by various technologies to assist retrieval has been fully used. At this time, if other available dimensions of information that can assist retrieval cannot be found from this information, it is almost impossible to significantly improve the accuracy.

[0003] The current text retrieval technologies usually consider information in dimensions such as the semantics of the text, the subject content of the text, and the syntactic structure. From traditional statistical-based retrieval algorithms to various emerging language models, the related technologies are becoming more and more complex, but the effects are showing a trend of diminishing marginal utility. The design of various new technologies is often just a superposition of existing technologies, or a combination of existing technologies. The fundamental reason for the lack of obvious effects is that new information dimensions are not sought from a more basic perspective. Some complex technologies may improve the accuracy rate within a certain range by improving the fit to the content, but they lose the ability to generalize, that is, the improved accuracy rate cannot play a role in text retrieval tasks for different content.

[0004] Therefore, it is particularly important to find a method that can assist the main retrieval algorithm to further improve the accuracy and has a wide range of applicability. Summary of the invention

[0005] In view of this, the purpose of the present invention is to provide a sentence query method and device based on semantic aggregation and topic preposition, which quantifies the information of two dimensions, semantic aggregation and topic preposition, which are easily overlooked in text retrieval, in a simple way, so as to improve the accuracy of the main query algorithm in the task of screening multiple alternative text segments.

[0006] The present invention solves the above technical problems by the following technical means:

[0007] In a first aspect, an embodiment of the present invention provides a sentence query method based on semantic aggregation and topic front degree, comprising the following steps:

[0008] Obtaining query sentences and alternative text segments expressed in natural language;

[0009] Segmenting the query sentence to obtain at least two target segmented words;

[0010] Searching for all target participles in the candidate text segment, and determining the first occurrence position of each target participle in the candidate text segment;

[0011] Calculating the semantic aggregation degree and topic preposition degree of the query sentence in the candidate text segment, and quantitatively fusing the information of the semantic aggregation degree and topic preposition degree;

[0012] Based on the semantic aggregation and topic front degree, the main search algorithm is assisted to select the text segment that is more likely to contain the query sentence from multiple candidate text segments.

[0013] In some implementations, the semantic convergence is characterized by the variance or standard deviation of the positions where all target words first appear in the candidate text segments.

[0014] In some implementations, calculating the semantic aggregation and topic preposition of the query sentence in the candidate text segment, and quantitatively fusing the information of the semantic aggregation and topic preposition, includes:

[0015] Calculating the semantic aggregation degree of the query sentence in the candidate text segment;

[0016] Calculate the topic front degree of the query sentence in the candidate text segment;

[0017] Merging the length information of the candidate text segments to obtain a fusion result;

[0018] Adjust the fusion result.

[0019] In some implementations, calculating the semantic aggregation of the query statement in the candidate text segment includes:

[0020] Obtain the number n of target word segments in the query sentence;

[0021] Calculate the average position P of all target words appearing for the first time in the candidate text segment avg ;

[0022] The semantic aggregation degree is calculated using the following formula:

[0023]

[0024] Among them, P i Indicates the position where each target word first appears in the candidate text segment.

[0025] In some implementations, the formula for calculating the topic fronting degree of the query sentence in the candidate text segment is as follows:

[0026]

[0027] In some implementations, the fusing the length information of the candidate text segments to obtain a fusion result includes:

[0028] Get the length len of the candidate text segment;

[0029] The length information of the candidate paragraphs and the degree of topic preposition are combined, and the formula involved is as follows:

[0030]

[0031] In some implementations, the fusion result is adjusted according to the following formula:

[0032]

[0033] In some implementations, based on the semantic aggregation, the standard deviation of the first occurrence position of all target words in the candidate text segment is used to represent the final fusion result, and the calculation formula is as follows:

[0034]

[0035] in, It is a method for calculating the standard deviation of the first occurrence position of all target words in the word sequence of the query sentence in the candidate text segment.

[0036] The present invention breaks away from the thinking framework of various current auxiliary retrieval technologies, takes the habit of people in "getting straight to the point" in the process of formal document writing as the degree of subject preposition, and borrows the idea of ​​variance and standard deviation in statistics as a tool for amplifying and measuring the degree of data dispersion, and uses a simple and easy method to calculate the semantic aggregation degree and subject preposition degree of the word sequence of the query sentence in the candidate text segment, and finally adds the dilution effect of the length of the candidate text segment on the subject preposition degree, and comprehensively calculates a quantitative score value, that is, the overall score value of the candidate text segment containing all target word segments in the word sequence of the query sentence, and uses this score value to assist the main retrieval algorithm, so that the text segment that is more likely to contain the query sentence can be selected from multiple candidate text segments, which can effectively improve the accuracy of text retrieval.

[0037] In the second aspect, an embodiment of the present invention also provides a statement query device based on semantic aggregation and topic front degree, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the steps of the statement query method described in the first aspect above are implemented.

[0038] In a third aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the statement query method described in the first aspect above are implemented.

[0039] It can be understood that the beneficial effects of the second and third aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a flow chart of a sentence query method based on semantic aggregation and topic front degree;

[0041] Figure 2 This is a diagram of the general situation in which the subject content is placed at the beginning of the paragraph in formal document writing;

[0042] Figure 3 This is a schematic diagram of 5 words obtained from the word segmentation of a query statement;

[0043] Figure 4 It is a zero-spacing aggregation distribution diagram of the query statement word sequence;

[0044] Figure 5 It is a completely average distribution diagram of the query statement word sequence;

[0045] Figure 6 It is a schematic diagram of the dilution of the topic preposition degree by text length. DETAILED DESCRIPTION

[0046] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0047] The terms "first" and "second" in the specification and claims herein are used to distinguish different objects rather than to describe a specific order of objects. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "multiple" refers to two or more, for example, multiple processing units refer to two or more processing units, etc., multiple elements refer to two or more elements, etc.

[0048] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.

[0049] The following is a brief description of the technical field of the technical solution of this application:

[0050] Data processing, including data collection, storage, retrieval, processing, change and transmission, aims to extract and derive valuable and meaningful data for certain specific users from large amounts of data that may be disorganized and difficult to understand.

[0051] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making. Artificial intelligence hardware technology generally includes computer vision technology, speech recognition technology, natural language processing technology, as well as its learning / deep learning, big data processing technology, knowledge graph and other aspects.

[0052] Natural language processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Natural language processing involves natural language, that is, the language people use in daily life, which is closely related to linguistic research; it also involves computer science and mathematics. The pre-trained model, an important technology for model training in the field of artificial intelligence, is developed from the Large Language Model (LLM) in the field of NLP. After fine-tuning, the large language model can be widely used in downstream tasks.

[0053] At present, various text retrieval technologies usually consider information in dimensions such as text semantics, text subject content, and syntactic structure. From traditional statistical-based retrieval algorithms to various emerging language models, related technologies are becoming more and more complex, but the effects are showing a trend of diminishing marginal utility. The design of various new retrieval technologies is often just a superposition of existing technologies, or a combination of existing technologies. The fundamental reason for the lack of obvious effects is that new information dimensions are not found from a more basic perspective.

[0054] The sentence query method based on semantic aggregation and topic preposition of the present invention breaks away from the thinking framework of various current auxiliary retrieval technologies, takes the habit of people in "getting straight to the point" in the process of formal document writing as a new information dimension (i.e., "topic preposition"), and borrows the idea of ​​"variance" and "standard deviation" as tools in statistics to amplify and measure the degree of data dispersion, uses a simple method to calculate the semantic aggregation of the query sentence word sequence in the alternative text segment, and finally adds the dilution effect of the alternative text segment length on the topic preposition, and designs an algorithm to comprehensively calculate the semantic aggregation and topic preposition into a quantitative score value, which is used to assist the main retrieval algorithm to further improve the accuracy.

[0055] For details, see Figure 1 :

[0056] The sentence query method based on semantic aggregation degree and topic front degree of the present invention comprises the following steps:

[0057] Step 100, obtaining a query sentence and candidate text segments expressed in natural language;

[0058] Step 200, segmenting the query statement to obtain at least two target segmented words;

[0059] Step 300, searching for all target participles in the candidate text segment, and determining the first occurrence position of each target participle in the candidate text segment;

[0060] Step 400, calculating the semantic aggregation and topic preposition of the query sentence in the candidate text segment, and quantitatively integrating the information of the semantic aggregation and topic preposition, wherein the semantic aggregation is represented by the variance or standard deviation of the first appearance position of all target words in the candidate text segment;

[0061] Step 500 , based on the semantic aggregation degree and the topic preposition degree, the main search algorithm is assisted to select a text segment that is more likely to contain the query sentence from multiple candidate text segments.

[0062] In the above technical solution, based on the common phenomenon that in formal document writing, authors are accustomed to directly describe the content theme at the beginning of the entire document, a single chapter, a single paragraph or even a single sentence, and the characterization effect of the discrete degree of the query sentence word sequence on the semantic aggregation degree, combined with the statistical characteristics of the mean and variance in statistics, the information of these two dimensions is quantitatively integrated. At the same time, based on the consideration of the dilution effect of adding the length of the text segment on the advantage of the topic pre-degree, and measures are implemented to avoid the calculation value returning to zero and causing the algorithm to fail, the information of the two dimensions of semantic aggregation and topic pre-degree that are neglected in text retrieval is quantified in a simple way to improve the accuracy of the main query algorithm in the multi-alternative text segment screening task.

[0063] In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0064] In step 100, the natural language in the query sentence expressed in natural language is the language used by people in daily life, for example, the query sentence expressed in natural language is: a path planning method for a drone.

[0065] It should be noted that in the present application, users can ask questions in a variety of ways such as inputting text information, inputting language information or selecting text information. After the user asks a question, the text information can be used directly as a query statement, or the voice information can be converted into text information and the converted text information can be used as the query statement.

[0066] The alternative paragraphs in this embodiment generally refer to more formal documents. The alternative paragraphs usually have the characteristics of the degree of subject preposition. The degree of subject preposition means that in the document, for the content to be expressed, the author will habitually directly put forward the subject content at the beginning, and this feature exists in the document structure in a recursive form. For example, at the beginning of the entire document, the author will usually clarify the theme of the entire document. At the beginning of each chapter in the document, the author will explain the main content of this chapter. This tendency will continue to exist in each sub-chapter, each paragraph, and even each sentence. And the more formal and standardized the document is, the more obvious the characteristics of the degree of subject preposition will be at each level of the document structure. Please refer to Figure 2 , Figure 2 The dark blocks represent the position of the topic in different paragraphs. The dark blocks appear in the front a lot in formal documents, which represents a common phenomenon in formal document writing.

[0067] The convention or habit of putting the topic first is not a mandatory writing standard, nor is it related to the content or semantics of the document. This feature originates from the purpose of the existence of this type of text. For formal or non-fictional texts, the purpose of writing is to convey information as efficiently and directly as possible, so it is undoubtedly the most suitable practice to put forward the topic as early as possible in the different levels of the text structure. This is different from fictional (literary) works, which need to create suspense and emotional ups and downs through various techniques such as flashbacks and interspersions, so there is no unified tendency for the placement of topics in this type of text.

[0068] In step 200, the query statement can be parsed according to the concepts, relationships, attributes, etc. of the words, phrases or short phrases in the natural language. For example, the user query statement (query statement) can be segmented according to the concepts, relationships, attributes, etc. of the words, phrases or short phrases to obtain at least two target segmentations. For another example, the query statement "drone path planning method" is segmented as follows:

[0069] "Drone", "of", "path", "planning", "method".

[0070] In step 300, all target words are searched in the candidate text segment, and the first occurrence position of each target word in the candidate text segment is determined, so as to calculate the variance or standard deviation of the first occurrence position of all target words in the candidate text segment in the subsequent step.

[0071] In step 400 , the semantic aggregation degree in the present application may be represented by the variance or standard deviation of the positions where all target words first appear in the candidate text segments.

[0072] Calculate the semantic aggregation and topic preposition of the query sentence in the candidate text segment, and quantitatively integrate the information of semantic aggregation and topic preposition, including:

[0073] Step 410: Calculate the semantic aggregation degree of the query statement in the candidate text segments.

[0074] Step 420, calculating the topic fronting degree of the query sentence in the candidate text segment.

[0075] Step 430: fuse the length information of the candidate text segments to obtain a fusion result.

[0076] Step 440, adjusting the fusion result.

[0077] Generally speaking, when a query statement has multiple candidate text segments as answers, and every word in the query statement word sequence exists in each candidate text segment, if the query statement word sequence exists in the candidate text segment in a relatively compact manner, then the probability of this text segment being the correct choice is higher, and if the query statement word sequence exists in the candidate text segment in a relatively dispersed manner, then the probability of this text segment being the correct choice is lower than that of the above-mentioned compact matching text segment. In the embodiment of the present application, the degree of compactness (or discreteness) of the distribution of the query statement word sequence in the candidate text segment is referred to as semantic aggregation. The feature represented by this semantic aggregation can be understood by two extreme cases of this feature. Please refer to Figure 3 If you use Figure 3 The five blocks in the figure represent five words obtained by segmenting a query sentence. These five words have two special distribution modes in the text, as follows:

[0078] One is the zero-spacing aggregate distribution method, see Figure 4 , when each word in the word segmentation sequence of the query sentence exists in the text segment without spacing, that is, the entire query sentence (without word segmentation and spacing) exists in the form of a complete sentence, the text segment has a very high probability of being the correct choice.

[0079] The other is a completely even distribution method, see Figure 5 , when the first and last words in the word sequence of the query sentence are exactly at the beginning and end of a paragraph, and the other words are distributed in the paragraph with an average distance, then compared with the zero-spacing aggregate distribution of the word sequence of the query sentence, its probability of being the correct choice will be much lower.

[0080] Based on the above content, the semantic aggregation degree of the query statement in the candidate text segment is calculated as follows:

[0081] In the statistical variance calculation, the square value of the difference between the value of each data point and the average value is summed and finally divided by the number of data points. Corresponding to the relationship between the word segmentation sequence of the sentence to be queried and the alternative text segment involved in this application, the square of the difference between the first appearance position of each target word segment of the sentence to be queried in the alternative text segment and the average position of their first appearance in all alternative text segments is added and then divided by the length of the word sequence. This method of calculating the complete translation variance can more clearly reflect the degree of discreteness of the word segmentation sequence of the sentence to be queried in the alternative text segment, and thus can be used to quantify the semantic aggregation degree. The semantic aggregation degree of the word segmentation sequence of the sentence to be queried in the alternative text segment can be calculated using the following formula:

[0082]

[0083] In the above formula, n represents the number of target words in the query sentence, P i represents the first occurrence position of each target word in the candidate text, P avg Represents the average of the first occurrence positions of all target words in the candidate texts.

[0084] The above calculation method can reflect the semantic aggregation degree, but it cannot reflect the semantic preposition degree, because no matter what the average first appearance position of all target words in the word sequence of the query sentence is, as long as the degree of discreteness within the word sequence is the same, the semantic aggregation degree calculated by this calculation method will not change. Therefore, it is also necessary to consider the quantification of the semantic preposition degree, as follows:

[0085] In order to add quantification of the degree of semantic preposition, it is necessary to use the average position P of the first appearance of all target words in the word sequence of the query sentence again. avg In statistics, the essence of the average is to use a single number to represent all sample data. In the case involved in this application, P avg The average position of the first appearance of all target words represents the position of the entire sequence. The smaller this value is, the earlier the entire sequence appears in the candidate text segment. The larger the calculated value of the semantic aggregation degree is, the greater the discreteness between words is, and the smaller the semantic aggregation degree is. The smaller the calculated value of the semantic aggregation degree is, the lower the discreteness is, and the greater the semantic aggregation degree is. In order to make the two play a synergistic role, the embodiment of the present application adopts P avg As a multiplication factor in the denominator, the method of using P avg The representation of the degree of semantic preposition, therefore, the calculation of the degree of semantic preposition is as follows:

[0086]

[0087] When P avg After adding the denominator, under the same semantic aggregation degree, P avg The smaller it is, the higher the final score is, so the information of semantic aggregation and semantic preposition degree is fused and used.

[0088] In actual application, we may also encounter the situation where the length of the text affects the degree of topic fronting. Figure 3 Take the word sequence shown in the figure as an example, when a query sentence has two candidate text segments with different text lengths and: the aggregation degree of the word sequence in the two text segments is consistent; the relative position of the word sequence in the two text segments is consistent. For example, see Figure 6, the semantic aggregation degree of the word sequence in the two paragraphs is the same, and the relative position in the two paragraphs is about half of the paragraph. If we want to consider the topic preposition degree of the word sequence in the two paragraphs, we can only measure it from the absolute length of the text. Compared with the longer paragraph, the shorter paragraph has a higher topic preposition degree (fewer words from the beginning to the topic position), but another aspect that needs to be considered is: regardless of whether the expanded description of the topic is before or after the word sequence, the longer paragraph has longer content to describe the content related to the topic. Therefore, when calculating the topic preposition length of different paragraphs, we should moderately consider the dilution effect of text length on the advantage of topic preposition degree.

[0089] For the calculation methods of semantic aggregation and semantic preposition, if the semantic aggregation is the same, a longer paragraph length will result in a longer P avg The length of the candidate text segment increases, which leads to a decrease in the overall score. In order to neutralize this effect, it is necessary to consider the length information of the fused candidate text segment. The method used in the embodiment of the present application is to take the logarithm of the text segment length as base 10, and then multiply it by the calculated semantic preposition degree score. If len is used to represent the length of the candidate text segment, the overall score calculation method becomes:

[0090]

[0091] If we simply use the absolute value of the paragraph length (number of words) as a multiplier, the paragraph length will have too much impact on the score, and the semantic aggregation and topic preposition information contained in the score will be lost. Therefore, we use the logarithm method to compress the difference between different length values. For example, for paragraphs with a length of 100 words and 1000 words, although the length of the two differs by 10 times, after taking the logarithm, this difference is reduced to 2 times.

[0092] Even if the length information of the candidate text is integrated, there are still some special cases in the formula used to calculate the overall score. These special cases may cause the calculation to fail. For example, if the local value in the formula is zero, the overall calculation value will be zero. It is necessary to fine-tune the calculation method to avoid the risks brought by such special cases. avg -P i The value of is 0, which will eventually lead to the result of the entire formula being 0. Therefore, the embodiment of the present application adopts a method of adding 1 to both the numerator and the denominator to avoid this risk without significantly affecting the original score. The modified calculation formula is:

[0093]

[0094] The above calculation formula is based on a word segmentation sequence of a query sentence, and scores an alternative text segment containing all target word segments in the word segmentation sequence. The higher the final calculated value, the greater the probability that the query sentence is included in the alternative text segment.

[0095] The above algorithm in this embodiment can quantify information reflecting the semantic aggregation and topic preposition of the word sequence of the query sentence in the alternative text segments, which is used to assist the main query algorithm in selecting a text segment that is more likely to contain the correct answer from multiple alternative text segments.

[0096] The above-mentioned quantitative method for the probability of the query statement appearing in the candidate text segment is mainly characterized by the variance of the position of the first appearance of the target word in the candidate text segment in the word sequence of the query statement. In other implementations, the calculation method of the key dimension information can also be changed to weaken or strengthen the role of this part of the information, that is, the standard deviation calculation method can be used. For example, in the calculation of the semantic aggregation part, if it is necessary to reduce the sensitivity of the algorithm to the discrete degree of the target word in the word sequence, the calculation formula is designed as follows:

[0097]

[0098] In the above calculation formula, It is a method for calculating the standard deviation of the first occurrence position of all target words in the word sequence of the query sentence in the candidate text segment.

[0099] In addition, if you need to increase or decrease the effect of paragraph length on the score increase in actual application, you can change It is implemented in a partial base way. If you need to reduce the effect of the length on the final calculated value, you can increase the base, for example, to 20 or even more. On the contrary, when you want to increase the effect of the length on the final calculated value, you can reduce the base to achieve it.

[0100] In step 500, after steps 100-400, the information such as sentence aggregation, topic preposition and text length are integrated to calculate the overall score of the candidate text segment containing all target words in the word sequence of the sentence to be queried. This score is used to assist the main search algorithm, so that the text segment that is more likely to contain the sentence to be queried can be selected from multiple candidate text segments, thereby improving the accuracy of text retrieval.

[0101] Another embodiment of the present invention provides a statement query device based on semantic aggregation and topic preposition, including: a processor, a memory, and a computer program in the memory that can be run on the processor, such as a statement query method program based on semantic aggregation and topic preposition. When the processor executes the computer program, the steps in the above-mentioned statement query method embodiments based on semantic aggregation and topic preposition are implemented, such as Figure 1 steps.

[0102] Exemplarily, the above-mentioned computer program can be divided into one or more modules / units, one or more modules / units are stored in a memory and executed by a processor to complete the present invention. One or more modules / units can be a series of computer program instruction segments that can complete specific functions. The instruction segments are used to describe the execution process of the computer program in a sentence query device based on semantic aggregation and topic preposition. For example, the computer program can be divided into a sentence text acquisition module, a word segmentation module, a word segmentation search module, a calculation fusion module, and an auxiliary retrieval module. The specific functions of each module are as follows:

[0103] The sentence and text acquisition module is used to obtain the sentences to be queried and the candidate texts expressed in natural language.

[0104] The word segmentation module is used to segment the query sentence to obtain at least two target word segments.

[0105] The word segmentation search module is used to search for all target word segments in the candidate text segment and determine the position where each target word appears for the first time in the candidate text segment.

[0106] The calculation fusion module is used to calculate the semantic aggregation and topic preposition degree of the query sentence in the candidate text segment, and quantitatively fuse the information of the semantic aggregation and topic preposition degree.

[0107] The auxiliary retrieval module is used to assist the main retrieval algorithm to select a text segment that is more likely to contain the query sentence from multiple candidate text segments based on semantic aggregation and topic front degree.

[0108] The sentence query device based on semantic aggregation and subject preposition can be a computing device such as a desktop computer, a notebook, a PDA, a cloud server, etc. The sentence query device based on semantic aggregation and subject preposition can include, but is not limited to, a processor and a memory, for example, it can also include an output device, a network access device, a bus, an audio access device, etc.

[0109] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the statement query device based on semantic aggregation and subject preposition, and uses various interfaces and lines to connect the various parts of the statement query device based on semantic aggregation and subject preposition.

[0110] The memory can be used to store computer programs and / or modules. The processor implements various functions of the statement query device based on semantic aggregation and subject front degree by running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory.

[0111] If the module / unit integrated in the statement query device based on semantic aggregation and subject preposition is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of each embodiment of the above-mentioned statement query method based on semantic aggregation and subject preposition can be implemented.

[0112] The computer program includes computer program code, which may be in source code form, object code form, executable file or some intermediate form, etc. Computer readable media may include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.

[0113] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention is described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions of the present invention, which should be included in the scope of the claims of the present invention. The techniques, shapes, and structural parts not described in detail in the present invention are all known technologies.

Claims

1. A sentence query method based on semantic aggregation and topic fronting degree, characterized in that: The following steps are involved: Obtaining query sentences and alternative text segments expressed in natural language; Segmenting the query sentence to obtain at least two target segmented words; Searching for all target participles in the candidate text segment, and determining the first occurrence position of each target participle in the candidate text segment; Calculating the semantic aggregation degree and topic preposition degree of the query sentence in the candidate text segment, and quantitatively fusing the information of the semantic aggregation degree and topic preposition degree; Based on the semantic aggregation and topic front degree, the main search algorithm is assisted to select the text segment that is more likely to contain the query sentence from multiple candidate text segments.

2. The sentence query method based on semantic aggregation and topic front degree according to claim 1 is characterized in that: The semantic aggregation degree is characterized by the variance or standard deviation of the positions where all target words appear for the first time in the candidate text segment.

3. The sentence query method based on semantic aggregation and topic front degree according to claim 2 is characterized in that: The calculating of the semantic aggregation degree and the topic preposition degree of the query sentence in the candidate text segment, and quantitatively fusing the information of the semantic aggregation degree and the topic preposition degree, includes: Calculating the semantic aggregation degree of the query sentence in the candidate text segment; Calculate the topic front degree of the query sentence in the candidate text segment; Merging the length information of the candidate text segments to obtain a fusion result; Adjust the fusion result.

4. The sentence query method based on semantic aggregation and topic front degree according to claim 3 is characterized in that: The calculating the semantic aggregation degree of the query statement in the candidate text segment includes: Obtain the number n of target word segments in the query sentence; Calculate the average position P of all target words appearing for the first time in the candidate text segment avg ; The semantic aggregation degree is calculated using the following formula: Among them, P i Indicates the position where each target word first appears in the candidate text segment.

5. The sentence query method based on semantic aggregation and topic front degree according to claim 4 is characterized in that: The formula for calculating the topic preposition degree of the query sentence in the candidate text segment is as follows:

6. The sentence query method based on semantic aggregation and topic front degree according to claim 5 is characterized in that: The step of fusing the length information of the candidate text segments to obtain a fusion result includes: Get the length len of the candidate text segment; The length information of the candidate paragraphs and the degree of topic preposition are combined, and the formula involved is as follows:

7. The sentence query method based on semantic aggregation and topic front degree according to claim 6 is characterized in that: The adjustment fusion result is adjusted according to the following formula:

8. The sentence query method based on semantic aggregation and topic front degree according to claim 2 is characterized in that: Based on the semantic aggregation, the standard deviation of the first appearance position of all target words in the candidate text segment is used to represent the final fusion result, and the calculation formula is as follows: in, It is a method for calculating the standard deviation of the first occurrence position of all target words in the word sequence of the query sentence in the candidate text segment.

9. A sentence query device based on semantic aggregation and topic front degree, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the statement query method according to any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the statement query method according to any one of claims 1 to 8 are implemented.