Search method, device, equipment and storage medium
By combining fuzzy search and semantic search methods, the text to be searched in the double-recorded video is processed, which solves the problem of inaccurate search results in the prior art, and achieves more efficient and accurate standard speech search.
Patent Information
- Application Number
- CN202510173564.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-17
AI Technical Summary
When searching for standard speech in dual-recorded videos, the prior art cannot effectively deal with typos, symbol errors and changes in expression methods in the text to be searched, resulting in inaccurate search results.
By obtaining the text to be searched, a fuzzy search is first performed to obtain the initial target fragment. If its similarity to the standard speech is insufficient, the search text is segmented, multiple text fragments are generated, and semantic search is performed using feature vectors until the search result that satisfies the similarity threshold is determined.
While taking into account the search efficiency, this method effectively reduces the influence of noise such as typos, wildcards, word order inversion, semantic conversion in the text, and improves the accuracy of search results.
Smart Images

Figure CN119646118B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a search method, device, equipment and storage medium. Background Art
[0002] In the fields of finance, law, and government services, in order to ensure that the business process is compliant and legal, it is usually necessary for both parties to complete the business process according to fixed standard scripts and record the entire business process in audio and video (referred to as double recording). The double-recorded video can be checked later. One of the inspection contents involves searching for standard scripts to determine whether both parties have completed the business process in accordance with regulations.
[0003] The existing method for searching for standard words in dual-recorded videos is: formulate a standard words text template; convert the audio content in the dual-recorded video into the text to be searched; match the standard words text template with the transcribed text to be searched in sequence, and output the search results.
[0004] However, in the existing dual-recording video inspection process, since the text to be searched is the text converted from audio, there are inevitably typos, missing symbols and other problems. The current string exact matching or regular matching cannot obtain accurate search results when there are typos in the text to be searched. In addition, in the actual business process, there may be situations where the business personnel mention the content in the standard script, but the words and sentences used are different from the standard script. The existing text search method cannot obtain accurate search results when there are certain changes in the expression. Summary of the invention
[0005] The present disclosure provides a search method, device, equipment and storage medium to at least solve the above technical problems existing in the prior art.
[0006] According to a first aspect of the present disclosure, a search method is provided, including: acquiring a text to be searched; performing a fuzzy search on the text to be searched based on a standard wording established in advance to obtain a first target segment; in response to a text similarity between the first target segment and the standard wording being less than a first threshold, segmenting the text to be searched to obtain a plurality of text segments; performing a semantic search on the plurality of text segments based on a feature vector of the standard wording and a feature vector of the text segment to obtain a second target segment; in response to a similarity between a feature vector of the second target segment and a feature vector of the standard wording being greater than a second threshold, determining the second target segment as the search result corresponding to the text to be searched.
[0007] In one possible implementation, the method of performing a fuzzy search on the text to be searched based on a pre-established standard phraseology to obtain a first target fragment includes: performing a fuzzy search on the text to be searched based on the pre-established standard phraseology to obtain a matching local fragment; determining a first number of characters in front of the local fragment and a second number of characters behind the local fragment in the standard phraseology; extending the local fragment forward by the first number of characters in the text to be searched to obtain a first extended fragment, and extending the local fragment backward by the second number of characters to obtain a second extended fragment; determining the first target fragment based on the first extended fragment, the second extended fragment and the standard phraseology.
[0008] In one possible implementation, the first target segment is determined based on the first extended segment, the second extended segment and the standard speech, including: determining the extended segment with the highest text similarity between the first extended segment and the second extended segment and the standard speech as the target extended segment corresponding to the local segment; and determining the target extended segment with the highest text similarity between the standard speech and the target extended segment corresponding to all the local segments as the first target segment.
[0009] In one possible implementation, the segmenting of the text to be searched to obtain a plurality of text fragments includes: segmenting the text to be searched into a plurality of first fragments based on punctuation in the text to be searched; determining a cosine distance between feature vectors of two adjacent first fragments in the text to be searched; in response to the cosine distance being less than a third threshold, splicing the two adjacent first fragments to obtain a spliced fragment, until the cosine distances between the feature vectors of all two adjacent fragments are not less than the third threshold, and determining all the obtained fragments as the text fragments.
[0010] In one possible implementation, the semantic search of the multiple text fragments to obtain the second target fragment includes: determining the text fragment whose feature vector among the multiple text fragments has the greatest similarity with the feature vector of the standard speech as the best fragment; in the text to be searched, splicing the best fragment with the preceding or succeeding text fragment to obtain a third extended fragment; the length of the third extended fragment is greater than or equal to the standard speech; and determining the second target fragment based on the feature vectors of the third extended fragment and the standard speech.
[0011] In one possible implementation, the determining of the second target segment based on the feature vectors of the third extended segment and the standard speech includes: in response to a first similarity between the feature vectors of the third extended segment and the standard speech being less than or equal to a second similarity between the feature vectors of the best segment and the standard speech, and the second similarity being greater than a fourth threshold, determining the best segment as the second target segment; in response to the first similarity being greater than the second similarity, concatenating the third extended segment with a preceding or succeeding text segment to obtain a fourth extended segment, until a third similarity between the feature vectors of the fourth extended segment and the standard speech is less than or equal to the first similarity, and the first similarity is greater than a fourth threshold, determining the third extended segment as the second target segment.
[0012] In one possible implementation, the method of determining the text segment having the greatest similarity between a feature vector of the multiple text segments and a feature vector of the standard speech as the best segment includes: segmenting the standard speech to obtain a plurality of standard speech segments; determining the text segment having the greatest similarity between a feature vector of the multiple text segments and a feature vector of the standard speech segment as the best sub-segment; in response to the number of text segments between two adjacent best sub-segments being less than a fifth threshold, splicing the two adjacent best sub-segments and the text segment between the two adjacent best sub-segments to obtain the best segment.
[0013] According to a second aspect of the present disclosure, a search device is provided, the device comprising: an acquisition module for acquiring a text to be searched; a first search module for performing a fuzzy search on the text to be searched based on a pre-established standard language to obtain a first target segment; a segmentation module for segmenting the text to be searched to obtain a plurality of text segments in response to a text similarity between the first target segment and the standard language being less than a first threshold value; a second search module for performing a semantic search on the plurality of text segments based on a feature vector of the standard language and a feature vector of the text segment to obtain a second target segment; and a determination module for determining the second target segment as a search result corresponding to the text to be searched in response to a similarity between a feature vector of the second target segment and a feature vector of the standard language being greater than a second threshold value.
[0014] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0015] at least one processor; and
[0016] a memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the present disclosure.
[0018] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method described in the present disclosure.
[0019] A search method, device, equipment and storage medium disclosed in the present invention combine fuzzy search and semantic search, and can effectively reduce the impact of various types of noise such as typos, wildcards, word order inversion, semantic conversion, etc. in the text while taking into account the search efficiency, thereby avoiding the problem of low search result accuracy due to typos and different expressions.
[0020] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner, in which:
[0022] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts.
[0023] Figure 1 A schematic diagram showing a search method according to an embodiment of the present disclosure Figure 1 ;
[0024] Figure 2 A schematic diagram showing a search method according to an embodiment of the present disclosure Figure 2 ;
[0025] Figure 3 A schematic diagram showing a search method according to an embodiment of the present disclosure Figure 3 ;
[0026] Figure 4 A schematic diagram showing a search method according to an embodiment of the present disclosure Figure 4 ;
[0027] Figure 5 A schematic diagram showing a search method according to an embodiment of the present disclosure Figure 5 ;
[0028] Figure 6 A schematic diagram of the structure of a search device according to an embodiment of the present disclosure is shown;
[0029] Figure 7 A schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0030] In order to make the purpose, features, and advantages of the present disclosure more obvious and easy to understand, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present disclosure.
[0031] Figure 1 A schematic diagram showing a search method according to an embodiment of the present disclosure Figure 1 ,like Figure 1 As shown, a search method includes:
[0032] Step S101, obtaining the text to be searched.
[0033] In this embodiment, the text to be searched may be text obtained by performing voice recognition on the audio content of a dual-recorded video. The dual-recorded video may be a video obtained by recording and videotaping business processes in the fields of finance, law, and government services. In the business process, both parties usually need to complete the business process according to fixed standard scripts.
[0034] Step S102: Based on the standard words established in advance, a fuzzy search is performed on the search text to obtain a first target segment.
[0035] In this embodiment, fuzzy search is used to perform local string matching between standard words and the text to be searched, and can search for accurate results when there are erroneous or missing words in the text to be searched. It is more robust and practical in text processing in the speech-to-text scenario. Standard words are words that need to be involved in the business process specified in advance. For example, in the financial field, there may be the following standard words:
[0036] Standard Script 1: Dear Mr. / Ms. {Customer Name}, Hello;
[0037] Standard speech 2: I am the bank’s account manager, my name is {account manager’s name}, and I am happy to have the opportunity to introduce our bank’s services and products to you.
[0038] In this embodiment, a fuzzy search may be performed on the text to be searched based on the standard terms. If the text similarity between a certain segment in the text to be searched and the standard terms is greater than a certain threshold, the segment is determined as the first target segment.
[0039] Step S103 , in response to the text similarity between the first target segment and the standard speech being less than a first threshold, segmenting the text to be searched to obtain a plurality of text segments.
[0040] In this embodiment, if the text similarity between the first target segment and the standard speech is less than the first threshold, it is considered that the results retrieved by the fuzzy search do not meet the search requirements, and a semantic search needs to be performed on the search text. First, the search text needs to be segmented to obtain multiple text segments. In one example, the search text can be segmented based on a tokenizer, where a tokenizer is a tool that converts a text string into smaller units (called "tokens" or "tags"). The tokenizer can be a rule-based tokenizer and a deep learning-based tokenizer. The rule-based tokenizer can be segmented based on spaces, punctuation, or regular expressions, and the deep learning-based tokenizer can be segmented based on the vector of the text to be searched.
[0041] Step S104, based on the feature vector of the standard speech and the feature vector of the text segment, a semantic search is performed on the multiple text segments to obtain a second target segment.
[0042] In this embodiment, for a certain standard speech, the similarity between the feature vector of the standard speech and the feature vectors of multiple text segments can be calculated in sequence. If the similarity is greater than a certain threshold, the corresponding text segment is determined as the second target segment.
[0043] Step S105 , in response to the similarity between the feature vector of the second target segment and the feature vector of the standard speech being greater than a second threshold, the second target segment is determined as the search result corresponding to the text to be searched.
[0044] In this embodiment, if the similarity between the feature vector of the second target segment and the feature vector of the standard speech is greater than a second threshold, it is considered that the results retrieved by the semantic search meet the search requirements, and thus the second target segment can be determined as the search result corresponding to the text to be searched.
[0045] Figure 2 A schematic diagram showing a search method according to an embodiment of the present disclosure Figure 2 ,like Figure 2As shown, the complete process of fuzzy search in this embodiment is as follows: input standard words, perform fuzzy search on the search text based on the standard words, and obtain a first target segment. If the text similarity between the first target segment and the standard words is not less than a first threshold, it is considered that the first target segment meets the search requirements, and the first target segment is directly output as the search result; if the text similarity between the first target segment and the standard words is less than the first threshold, it is considered that the first target segment does not meet the search requirements, so it is necessary to first vectorize the standard words to obtain a feature vector of the standard words, and segment the search text to obtain multiple text segments, and then perform semantic search on the multiple text segments based on the feature vector of the standard words and the feature vector of the text segments to obtain a second target segment. If the similarity between the feature vector of the second target segment and the feature vector of the standard words is greater than a second threshold, it is considered that the second target segment meets the search requirements, and the second target segment is output as the search result; if the similarity between the feature vector of the second target segment and the feature vector of the standard words is not greater than the second threshold, it is considered that there is no segment in the search text that matches the input standard words, and therefore the search failure can be output.
[0046] It should be emphasized that there may be multiple standard phrases, and a search method in the present disclosure may be used to search the search text based on each standard phrase, and the final search result may be a collection of search results corresponding to all standard phrases.
[0047] In the present disclosure, fuzzy search can search for accurate results even when there are typos or missing characters in the text to be searched, and the search efficiency is higher. Semantic search can search for accurate results even when there are sentence sequences swapped or certain changes in the expression in the text to be searched. Combining fuzzy search and semantic search can effectively reduce the impact of various types of noise in the text, such as typos, wildcards, word order inversion, semantic conversion, etc., while taking into account the search efficiency, and avoid the problem of low accuracy of search results due to typos and different expressions.
[0048] Figure 3 A schematic diagram showing a search method according to an embodiment of the present disclosure Figure 3 ,like Figure 3 As shown, step S102 "based on the pre-established standard words, fuzzy search the search text to obtain the first target segment" includes:
[0049] Step S201, based on the pre-defined standard words, a fuzzy search is performed on the search text to obtain matching local segments.
[0050] In this embodiment, during the fuzzy search of the text to be searched based on the standard terms, if the text similarity between a certain segment in the text to be searched and the standard terms is greater than a certain threshold, the segment is determined as a local segment matching the standard terms.
[0051] Step S202: Determine the first number of characters before the partial segment and the second number of characters after the partial segment in the standard speech.
[0052] Step S203: In the text to be searched, the local segment is extended forward by a first number of characters to obtain a first extended segment, and the local segment is extended backward by a second number of characters to obtain a second extended segment.
[0053] In this embodiment, the matched local segment may not cover the standard speech, and the local segment needs to be extended in the text to be searched until the number of characters in the local segment is the same as the number of characters in the standard speech, so that the extended segment can cover the standard speech. First, in the standard speech, it is necessary to determine the first number of characters in front of the local segment and the second number of characters behind the local segment, so as to determine how many characters the local segment needs to be extended before and after in the text to be searched. Then, in the text to be searched, the local segment is extended forward by the first number of characters to obtain a first extended segment, and the local segment is extended backward by the second number of characters to obtain a second extended segment.
[0054] In one example, if the text to be searched is: aaabbbccc, and the standard phraseology is abbc, a fuzzy search is performed on the text to be searched based on the standard phraseology to obtain matching local fragments abb and c; for the local fragment abb, the number of the first character in front of it in the standard phraseology is 0, and the number of the second character in the back of it in the standard phraseology is 1, so there is no need to extend the local fragment abb forward in the text to be searched, that is, there is no first extended fragment, but the local fragment abb needs to be extended backward by one character in the text to be searched to obtain abbb, that is, the second extended fragment is abbb; for the local fragment c, the number of the first character in front of it in the standard phraseology is 3, and the number of the first character in the back of it in the standard phraseology is 0, so the local fragment c is extended forward by 3 characters in the text to be searched to obtain bbbc, that is, the first extended fragment is bbbc, and there is no second extended fragment.
[0055] Step S204, determining a first target segment based on the first extended segment, the second extended segment and the standard speech.
[0056] In this embodiment, the first target segment can be determined based on the first extended segment, the second extended segment and the standard speech in the following manner: the extended segment with the highest text similarity between the first extended segment and the second extended segment and the standard speech is determined as the target extended segment corresponding to the local segment, and then the target extended segment with the highest text similarity between the target extended segments corresponding to all local segments and the standard speech is determined as the first target segment.
[0057] In one example, for the above-mentioned text to be searched aaabbbccc and the standard speech abbc, the matching local segments are abb and c. For the local segment abb, there is no first extended segment, and the second extended segment is abbb. The second extended segment abbb can be directly used as the target extended segment; for the local segment c, the first extended segment is bbbc, and there is no second extended segment. The first extended segment bbbc can be directly used as the target extended segment. If the target extended segment corresponding to the local segment abb is the second extended segment abbb, and the target extended segment corresponding to the local segment c is the first extended segment bbbc, then the target extended segment with the highest text similarity between abbb and bbbc and the standard speech abbc can be determined as the first target segment. If the text similarity between abbb and abbc is greater than the text similarity between bbbc and abbc, then the first target segment is abbb.
[0058] In one example, the text similarity between two segments is determined by the following formula:
[0059] Formula 1
[0060] Among them, S is the text similarity, is the edit distance between two fragment strings, The lengths of the strings for the two segments respectively.
[0061] In the present disclosure, the local segment obtained by the fuzzy search is extended forward and backward in the text to be searched, respectively, until the length of the extended segment is the same as the length of the standard speech, and the first target segment corresponding to the fuzzy search is determined based on the text similarity between the extended segment and the standard speech. In this way, it is possible to ensure that the most complete first target segment is obtained, thereby improving the accuracy of the first target segment.
[0062] In another embodiment, the step S103 of “segmenting the text to be searched to obtain multiple text segments” includes:
[0063] Based on punctuation marks in the text to be searched, the text to be searched is divided into multiple first fragments; the cosine distance between the feature vectors of two adjacent first fragments in the text to be searched is determined; in response to the cosine distance being less than a third threshold, the two adjacent first fragments are spliced to obtain spliced fragments, until the cosine distances between the feature vectors of all two adjacent fragments are not less than the third threshold, and all the obtained fragments are determined as text fragments.
[0064] In this embodiment, the text to be searched can be first divided into multiple first segments based on the punctuation in the text to be searched, and the first segments can be vectorized to obtain the feature vectors of the first segments, and then the cosine distance between the feature vectors of two adjacent first segments is calculated in sequence. If the cosine distance is less than a third threshold value, it proves that the two adjacent first segments are semantically similar, and then the two adjacent first segments are spliced to obtain the spliced segments, until the cosine distance between the feature vectors of all two adjacent segments is not less than the third threshold value, and all the obtained segments are determined as text segments; if the cosine distance is not less than the third threshold value, it proves that the two adjacent first segments are semantically dissimilar, and the two adjacent first segments need to be segmented. In this way, the text to be searched can be accurately segmented, and the accuracy of the text segments can be improved.
[0065] In one example, you can also create an MNLI (Multi-Genre Natural Language Inference) dataset for relevant scenarios based on standard speech templates, fine-tune the vectorized Embedding model, and improve the accuracy of segmentation. The MNLI dataset is mainly used for sentence relevance judgment.
[0066] Figure 4 A schematic diagram showing a search method according to an embodiment of the present disclosure Figure 4 ,like Figure 4 As shown, the "performing semantic search on multiple text segments to obtain a second target segment" in step S104 includes:
[0067] Step S301 : determining the text segment whose feature vector among multiple text segments has the greatest similarity with the feature vector of the standard speech as the best segment.
[0068] In this embodiment, for a certain standard speech, the similarities between the feature vector of the standard speech and the feature vectors of multiple text segments of the text to be searched can be calculated respectively, and the text segment with the greatest similarity can be determined as the best segment.
[0069] Step S302: In the text to be searched, the best segment is concatenated with the preceding or following text segment to obtain a third extended segment.
[0070] In this embodiment, the standard speech may be relatively long, and the best segment may not cover the standard speech. It is necessary to extend the best segment forward or backward in the text to be searched until the length of the best segment is greater than or equal to the length of the standard speech to obtain the third extended segment.
[0071] In one example, if multiple text fragments of the text to be searched are: text fragment 1, text fragment 2, text fragment 3, text fragment 4, text fragment 5, ..., text fragment n, among which text fragment 3 is the best fragment, the best fragment can be extended forward and backward to obtain two third extended fragments, namely (text fragment 2, text fragment 3) and (text fragment 3, text fragment 4), where () represents splicing.
[0072] Step S303: Determine a second target segment based on the feature vector of the third extended segment and the standard speech.
[0073] In this embodiment, the second target segment can be determined based on the feature vectors of the third extended segment and the standard speech in the following manner: in response to the first similarity between the feature vectors of the third extended segment and the standard speech being less than or equal to the second similarity between the best segment and the feature vectors of the standard speech, and the second similarity being greater than a fourth threshold, the best segment is determined as the second target segment; in response to the first similarity being greater than the second similarity, the third extended segment is spliced with a preceding or succeeding text segment to obtain a fourth extended segment, until the third similarity between the feature vectors of the fourth extended segment and the standard speech is less than or equal to the first similarity, and the first similarity is greater than the fourth threshold, the third extended segment is determined as the second target segment.
[0074] Figure 5 A schematic diagram showing a search method according to an embodiment of the present disclosure Figure 5 ,like Figure 5As shown, the complete process of semantic search in the present disclosure is as follows: determine the feature vector of the standard speech, determine the text segment whose feature vector in multiple text segments has the largest similarity with the feature vector of the standard speech as the best segment, and then, in the text to be searched, splice the best segment with the previous or next text segment to obtain a third extended segment. If the similarity between the feature vector of the third extended segment and the standard speech is not improved relative to the similarity between the feature vector of the best segment and the standard speech, it is considered that there is no need to continue extending the third extended segment, and the second similarity between the feature vector of the best segment and the standard speech is greater than the fourth threshold, the best segment can be directly determined as the second target segment. If the similarity between the feature vector of the best segment and the standard speech is greater than the fourth threshold, the best segment can be directly determined as the second target segment. If the second similarity between the vectors is not greater than the fourth threshold, it is considered that there is no segment corresponding to the standard speech in the text to be searched, that is, the search fails; if the similarity between the feature vector of the third extended segment and the standard speech is improved relative to the similarity between the feature vector of the best segment and the standard speech, it is also necessary to determine the fourth extended segment after the third extended segment is further spliced with a text segment in front or behind, if the similarity between the feature vector of the fourth extended segment and the standard speech is not improved relative to the similarity between the feature vector of the third extended segment and the standard speech, and the similarity between the feature vector of the third extended segment and the standard speech is greater than the fourth threshold, the third extended segment is determined as the second target segment.
[0075] In the present disclosure, in the process of semantic search of the text to be searched, the text segment whose feature vector has the greatest similarity with the feature vector of the standard speech among multiple text segments is determined as the best segment, and then in the text to be searched, the best segment is spliced with the preceding or following text segment to obtain a third extended segment, the length of the third extended segment is greater than or equal to the standard speech, and finally the second target segment is determined based on the feature vector of the third extended segment and the standard speech. In this way, it is possible to ensure that the most complete second target segment is obtained, and the accuracy of the second target segment is improved.
[0076] In another embodiment, step S301 of “determining the text segment having the greatest similarity between a feature vector of the plurality of text segments and a feature vector of the standard speech as the best segment” includes:
[0077] The standard speech is segmented to obtain multiple standard speech segments; the text segment whose feature vector in the multiple text segments has the greatest similarity with the feature vector of the standard speech segment is determined as the best sub-segment; in response to the number of text segments between two adjacent best sub-segments being less than a fifth threshold, two adjacent best sub-segments and the text segment between two adjacent best sub-segments are spliced to obtain the best segment.
[0078] In this embodiment, since the Embedding model has a limit on the length of the input token, and when the input token is too long, on the one hand, the sentence covers too much information, and on the other hand, the model's ability to express the characteristics of the sentence will also decrease. Therefore, in this embodiment, it is necessary to segment the standard speech to obtain multiple standard speech segments. In one example, punctuation marks can be used as segmentation points to segment the standard speech text to ensure that the segmented standard speech segments do not exceed 128 words. Among them, since the use of text and punctuation in standard speech is highly standardized, punctuation marks that have the meaning of completely ending a sentence can be selected as separators, such as periods, question marks, exclamation marks, line breaks, etc.
[0079] In this embodiment, for each standard speech segment, it is necessary to determine the text segment whose feature vector in multiple text segments has the greatest similarity with the feature vector of the standard speech segment as the best sub-segment. If the number of text segments between two adjacent best sub-segments is less than the fifth threshold, the two adjacent best sub-segments and the text segment between the two adjacent best sub-segments are spliced to obtain the best segment. In one example, under normal circumstances, each best sub-segment should be continuous in the text to be searched, but due to the presence of typos, disorder and other noise in the text, the actual result is not satisfactory. Therefore, it is necessary to formulate a splicing rule: the first best sub-segment is the starting point. If the interval between the starting point of the subsequent best sub-segment and the end point of the previous best sub-segment does not exceed N sentences, the previous best sub-segment, the interval segment, and the next best sub-segment are spliced together to form a complete and continuous best segment. Otherwise, the splicing is terminated and the spliced result is output. In one example, the value of N can be 3. In this way, a complete, continuous and accurate best segment can be obtained.
[0080] In another embodiment, the standard words may also include wildcards to represent information such as names and product numbers. When searching the search text, the wildcards need to be specially processed to avoid the wildcard characters interfering with the calculation results of text similarity. When performing a fuzzy search, a wildcard can be used to replace a string of any length, where when the wildcard is located at both ends of a sentence, it is represented as any word; when performing a semantic search, the wildcard can be replaced with "[MASK][MASK][MASK]", where [MASK] is a special token. In one example, the number of [MASK] can be 3.
[0081] Figure 6 A schematic diagram of the structure of a search device according to an embodiment of the present disclosure is shown. Figure 6 As shown, a search device comprises:
[0082] An acquisition module 10 is used to acquire the text to be searched; a first search module 11 is used to perform a fuzzy search on the text to be searched based on a pre-established standard language to obtain a first target segment; a segmentation module 12 is used to segment the text to be searched in response to the text similarity between the first target segment and the standard language being less than a first threshold value to obtain multiple text segments; a second search module 13 is used to perform a semantic search on multiple text segments based on a feature vector of the standard language and a feature vector of the text segment to obtain a second target segment; a determination module 14 is used to determine the second target segment as the search result corresponding to the text to be searched in response to the similarity between the feature vector of the second target segment and the feature vector of the standard language being greater than a second threshold value.
[0083] In one possible implementation, the first search module 11 is also used to: perform a fuzzy search on the text to be searched based on a pre-established standard phraseology to obtain a matching local segment; in the standard phraseology, determine a first number of characters in front of the local segment and a second number of characters behind the local segment; in the text to be searched, extend the local segment forward by a first number of characters to obtain a first extended segment, and extend the local segment backward by a second number of characters to obtain a second extended segment; determine a first target segment based on the first extended segment, the second extended segment and the standard phraseology.
[0084] In one possible implementation, the first search module 11 is also used to: determine the extended segment with the highest text similarity between the first extended segment and the second extended segment and the standard speech as the target extended segment corresponding to the local segment; determine the target extended segment with the highest text similarity between the target extended segments corresponding to all local segments and the standard speech as the first target segment.
[0085] In one embodiment, the segmentation module 12 is also used to: segment the text to be searched into multiple first segments based on punctuation in the text to be searched; determine the cosine distance between the feature vectors of two adjacent first segments in the text to be searched; in response to the cosine distance being less than a third threshold, splice the two adjacent first segments to obtain spliced segments, until the cosine distances between the feature vectors of all two adjacent segments are not less than the third threshold, and determine all the obtained segments as text segments.
[0086] In one possible implementation, the second search module 13 is also used to: determine the text segment whose feature vector in multiple text segments has the greatest similarity with the feature vector of the standard speech as the best segment; in the text to be searched, splice the best segment with the previous or next text segment to obtain a third extended segment; the length of the third extended segment is greater than or equal to the standard speech; based on the feature vectors of the third extended segment and the standard speech, determine the second target segment.
[0087] In one embodiment, the second search module 13 is also used to: in response to the first similarity between the feature vector of the third extended segment and the standard speech being less than or equal to the second similarity between the feature vector of the best segment and the standard speech, and the second similarity being greater than a fourth threshold, determine the best segment as the target segment; in response to the first similarity being greater than the second similarity, splice the third extended segment with a preceding or succeeding text segment to obtain a fourth extended segment, until the third similarity between the feature vector of the fourth extended segment and the standard speech is less than or equal to the first similarity, and the first similarity is greater than the fourth threshold, determine the third extended segment as the target segment.
[0088] In one possible implementation, the segmentation module 12 is further used to: segment the standard speech to obtain multiple standard speech segments; the second search module 13 is further used to: determine the text segment whose feature vector in the multiple text segments has the greatest similarity with the feature vector of the standard speech segment as the best sub-segment; in response to the number of text segments between two adjacent best sub-segments being less than a fifth threshold, splicing the two adjacent best sub-segments and the text segment between the two adjacent best sub-segments to obtain the best segment.
[0089] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.
[0090] Figure 7 A schematic block diagram of an example electronic device 800 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0091] like Figure 7 As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 to a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0092] A number of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0093] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as a search method. For example, in some embodiments, a search method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of a search method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform a search method in any other appropriate manner (e.g., by means of firmware).
[0094] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0095] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0096] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0097] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0098] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0099] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0100] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0101] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of the present disclosure, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0102] The above is only a specific embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art who is familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present disclosure, which should be included in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be based on the protection scope of the claims.
Claims
1. A search method, characterized in that: The method comprises: Get the text to be searched; Based on the standard words set in advance, a fuzzy search is performed on the text to be searched to obtain a first target segment; In response to the text similarity between the first target segment and the standard speech being less than a first threshold, segmenting the text to be searched to obtain a plurality of text segments; Based on the feature vector of the standard speech and the feature vector of the text segment, a semantic search is performed on the plurality of text segments to obtain a second target segment; In response to the similarity between the feature vector of the second target segment and the feature vector of the standard speech being greater than a second threshold, determining the second target segment as a search result corresponding to the text to be searched; The step of performing a fuzzy search on the text to be searched based on a pre-defined standard phrase to obtain a first target segment includes: Based on the standard words formulated in advance, a fuzzy search is performed on the text to be searched to obtain matching local segments; In the standard speech, determining a first number of characters before the partial segment and a second number of characters after the partial segment; In the text to be searched, the partial segment is extended forward by the first number of characters to obtain a first extended segment, and the partial segment is extended backward by the second number of characters to obtain a second extended segment; Determine the first target segment based on the first extended segment, the second extended segment and the standard speech; The step of determining the first target segment based on the first extended segment, the second extended segment, and the standard speech includes: Determine the extended segment having the highest text similarity with the standard speech in the first extended segment and the second extended segment as the target extended segment corresponding to the local segment; The target extended segment corresponding to all the local segments and having the highest text similarity with the standard speech is determined as the first target segment.
2. The method according to claim 1, characterized in that: The step of segmenting the text to be searched to obtain a plurality of text segments includes: Based on punctuations in the text to be searched, segment the text to be searched into a plurality of first segments; Determine the cosine distance between the feature vectors of two adjacent first segments in the text to be searched; In response to the cosine distance being less than a third threshold, the two adjacent first segments are spliced to obtain a spliced segment, until the cosine distances between the feature vectors of all two adjacent segments are not less than the third threshold, and all the obtained segments are determined as the text segments.
3. The method according to claim 1, characterized in that: The performing semantic search on the plurality of text segments to obtain a second target segment includes: Determine the text segment whose feature vector among the plurality of text segments has the greatest similarity with the feature vector of the standard speech as the best segment; In the text to be searched, the best segment is spliced with the preceding or following text segment to obtain a third extended segment; the length of the third extended segment is greater than or equal to the standard speech; The second target segment is determined based on the feature vector of the third extended segment and the standard speech.
4. The method according to claim 3, characterized in that The determining the second target segment based on the feature vector of the third extended segment and the standard speech includes: In response to a first similarity between the feature vector of the third extended segment and the standard speech being less than or equal to a second similarity between the feature vector of the best segment and the standard speech, and the second similarity being greater than a fourth threshold, determining the best segment as the second target segment; In response to the first similarity being greater than the second similarity, the third extended segment is spliced with a preceding or succeeding text segment to obtain a fourth extended segment, until the third similarity between the fourth extended segment and the feature vector of the standard speech is less than or equal to the first similarity, and the first similarity is greater than a fourth threshold, and the third extended segment is determined as the second target segment.
5. The method according to claim 3, characterized in that: The step of determining the text segment having the greatest similarity between the feature vectors of the plurality of text segments and the feature vector of the standard speech as the best segment comprises: Segmenting the standard speech to obtain multiple standard speech segments; Determine the text segment whose feature vector in the plurality of text segments has the greatest similarity to the feature vector of the standard speech segment as the best sub-segment; In response to the number of text segments between two adjacent best sub-segments being less than a fifth threshold, the two adjacent best sub-segments and the text segment between the two adjacent best sub-segments are concatenated to obtain the best segment.
6. A search device, characterized in that: The device comprises: The acquisition module is used to obtain the text to be searched; A first search module, configured to perform a fuzzy search on the text to be searched based on a pre-defined standard phrase to obtain a first target segment; a segmentation module, configured to segment the text to be searched to obtain a plurality of text segments in response to the text similarity between the first target segment and the standard speech being less than a first threshold; A second search module performs semantic search on the plurality of text segments based on the feature vector of the standard speech and the feature vector of the text segment to obtain a second target segment; a determination module, configured to determine the second target segment as a search result corresponding to the text to be searched in response to a similarity between a feature vector of the second target segment and a feature vector of the standard speech being greater than a second threshold; The step of performing a fuzzy search on the text to be searched based on a pre-defined standard phrase to obtain a first target segment includes: Based on the standard words formulated in advance, a fuzzy search is performed on the text to be searched to obtain matching local segments; In the standard speech, determining a first number of characters before the partial segment and a second number of characters after the partial segment; In the text to be searched, the partial segment is extended forward by the first number of characters to obtain a first extended segment, and the partial segment is extended backward by the second number of characters to obtain a second extended segment; Determine the first target segment based on the first extended segment, the second extended segment and the standard speech; The step of determining the first target segment based on the first extended segment, the second extended segment, and the standard speech includes: Determine the extended segment having the highest text similarity with the standard speech in the first extended segment and the second extended segment as the target extended segment corresponding to the local segment; The target extended segment corresponding to all the local segments and having the highest text similarity with the standard speech is determined as the first target segment.
7. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to make a computer execute the method according to any one of claims 1-5.
Citation Information
Patent Citations
Text positioning method and system and storage medium
CN117743510A
Data retrieval method and device and deep learning model training method and device
CN118312544A
Intelligent search method based on retrieval enhancement generation
CN118733712A