Error correction method and device for search terms, equipment and storage medium

By building a historical search data index file in search engines, the problem of inaccurate search terms recommendations when users enter errors is solved, fast and accurate search terms recommendations are achieved, and user search experience is improved.

CN120144847APending Publication Date: 2025-06-13BEIJING IQIYI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510145079.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing search engines cannot provide search terms quickly and accurately when users enter errors, which increases the search cost and time for users.

Method used

By constructing historical search data, the index file includes error correction words corresponding to the input characters. When the user enters the target character, the matching error correction words are found from the index file and displayed as recommended search terms.

Benefits of technology

It realizes the rapid and accurate provision of search terms when users enter errors, reduces the search cost and time of users and improves search efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144847A_ABST
    Figure CN120144847A_ABST
Patent Text Reader

Abstract

The invention relates to an error correction method and device for search terms, equipment and a storage medium. The method comprises the following steps: acquiring a target character input by a user, and searching an index file for a target error correction word matched with the target character; wherein the index file comprises error correction words which are determined through historical search data and correspond to the input characters; and displaying the target error correction word as a recommended search word. It can be seen that the index file is pre-established through the historical search data, and the index file comprises the error correction word corresponding to the input character; therefore, in the application, when the user inputs the target character to search, the corresponding target error correction word can be searched through the index file and is displayed as the recommended search word, so that the recommended search word is quickly and accurately provided, and the search efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of search term error correction, and in particular, to a method, device, equipment and storage medium for correcting search terms. Background Art

[0002] A search dropdown box refers to a function in which, when a user uses a search engine, the search engine automatically provides a list of search term candidates for the user to select based on the user's current input. Currently, when a user conducts a search, the process of the search dropdown box providing some search terms is called Suggest (recommendation), and it is implemented through a Suggest module, thereby improving the user's search efficiency. However, when providing search terms through the Suggest module, if the corresponding search term cannot be accurately provided to the user due to an input error, the user needs to further search and input to use the search engine for searching, increasing the user's search cost.

[0003] Therefore, how to quickly and accurately provide search terms and improve search efficiency is a problem that needs to be solved by those skilled in the art. Summary of the Invention

[0004] This application provides a method, device, equipment and storage medium for correcting search terms to quickly and accurately provide search terms and improve search efficiency.

[0005] In a first aspect, this application provides a method for correcting search terms, including:

[0006] Obtain a target character input by a user;

[0007] Search for a target error correction term that matches the target character in an index file; the index file includes: error correction terms corresponding to the input characters determined through historical search data;

[0008] Display the target error correction term as a recommended search term.

[0009] Optionally, before obtaining the target character input by the user, it further includes:

[0010] Determine a first historical search term according to historical search data;

[0011] Convert the first historical search term into a historical nine-grid code;

[0012] Construct a first index file with the historical nine-grid code as the key and the first historical search term corresponding to the historical nine-grid code as the value.

[0013] Optionally, the searching for a target error correction term that matches the target character in the index file includes:

[0014] Convert the target character into a target nine - grid code;

[0015] Search for a target historical nine - grid code that matches the target nine - grid code in the first index file;

[0016] Use the first historical search term corresponding to the target historical nine - grid code as the target error - correction term corresponding to the target character.

[0017] Optionally, before obtaining the target character input by the user, it further includes:

[0018] Determine an input candidate word and a second historical search term according to historical search data; wherein, the input candidate word is a candidate word with a click - through rate lower than a first predetermined threshold input by the user;

[0019] Calculate the similarity value between each input candidate word and each second historical search term;

[0020] Filter out error - correction word pairs with similarity values higher than a second predetermined threshold; the error - correction word pairs include: the input candidate word and the second historical search term with a similarity value greater than the second predetermined threshold;

[0021] Construct a second index file with the input candidate word in each error - correction word pair as the key and the second historical search term as the value.

[0022] Optionally, the calculating the similarity value between each input candidate word and each second historical search term includes:

[0023] Convert each character in each input candidate word into a four - corner code, and convert each character in each second historical search term into a four - corner code;

[0024] Calculate the glyph similarity of different character pairs between the input candidate word and the second historical search term according to the four - corner codes of each character in the input candidate word and the corresponding characters in the second historical search term;

[0025] Calculate the similarity value between the input candidate word and the second historical search term according to the glyph similarity of different character pairs between the input candidate word and the second historical search term.

[0026] Optionally, the calculating the glyph similarity of different character pairs between the input candidate word and the second historical search term according to the four - corner codes of each character in the input candidate word and the corresponding characters in the second historical search term includes:

[0027] Determine each character pair between the input candidate word and the second historical search term;

[0028] Determine, for each character pair, the number of the same codes and the total number of codes in the four - corner code of the corresponding character in the input candidate word and the four - corner code of the corresponding character in the second historical search term;

[0029] Calculate the glyph similarity of each character pair according to the same number of encodings of each character pair and the total number of encodings.

[0030] Optionally, the finding of the target error correction word matching the target character from the index file includes:

[0031] Find a target input candidate word that matches the target character from a second index file;

[0032] Use the second historical search word corresponding to the target input candidate word as the target error correction word corresponding to the target character.

[0033] In a second aspect, the present application provides an error correction device for search terms, including:

[0034] An acquisition module, configured to acquire a target character input by a user;

[0035] A search module, configured to search for a target error correction word that matches the target character from an index file; the index file includes: error correction words corresponding to the input characters determined through historical search data;

[0036] A display module, configured to display the target error correction word as a recommended search term.

[0037] In a third aspect, the present application provides an electronic device, including:

[0038] A processor, a memory, and a computer program stored on the memory and executable on the processor, and the processor executes the steps of the above error correction method of the present application through the computer program.

[0039] In a fourth aspect, the present application further provides a computer storage medium, where the computer storage medium stores computer executable instructions, and the computer executable instructions are used to execute the steps of the above error correction method of the present application.

[0040] The above technical solutions provided by the embodiments of the present application have the following advantages compared with the prior art: The present application discloses an error correction solution for search terms. In this solution, a target character input by a user is acquired, and a target error correction word that matches the target character is searched from an index file; wherein, the index file includes: error correction words corresponding to the input characters determined through historical search data; the target error correction word is displayed as a recommended search term. It can be seen that the present application pre-establishes an index file through historical search data, and the index file includes error correction words corresponding to the input characters; therefore, in the present application, when a user inputs a target character for searching, the corresponding target error correction word can be searched through the index file and displayed as a recommended search term, so as to quickly and accurately provide recommended search terms, thereby improving the search efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention and, together with the specification, are used to explain the principles of the present invention.

[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other accompanying drawings can also be obtained based on these drawings without creative efforts.

[0043] One or more embodiments are exemplarily illustrated by the pictures in the corresponding accompanying drawings. These exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the accompanying drawings represent similar elements, unless otherwise stated, and the drawings in the accompanying drawings do not constitute a proportional limitation.

[0044] Figure 1 Schematic flowchart of a method for correcting search terms provided by an embodiment of the present application;

[0045] Figure 2 Schematic flowchart of another method for correcting search terms provided by an embodiment of the present application;

[0046] Figure 3a Schematic flowchart of an offline processing based on nine - grid coding provided by an embodiment of the present application;

[0047] Figure 3b Schematic flowchart of an online processing based on nine - grid coding provided by an embodiment of the present application;

[0048] Figure 4 Schematic flowchart of another method for correcting search terms provided by an embodiment of the present application;

[0049] Figure 5a Schematic flowchart of an offline processing based on similar - character correction provided by an embodiment of the present application;

[0050] Figure 5b Schematic flowchart of an online processing based on similar - character correction provided by an embodiment of the present application;

[0051] Figure 6 Corresponding diagram of the four - corner coding of Chinese characters provided by an embodiment of the present application;

[0052] Figure 7 Schematic structural diagram of a search - term correction device provided by an embodiment of the present application;

[0053] Figure 8 Schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0054] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Apparently, the described embodiments are some but not all of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without making creative efforts shall fall within the scope of protection of this application.

[0055] The following disclosure provides many different embodiments or examples for implementing different structures of the present invention. To simplify the disclosure of the present invention, components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present invention. In addition, the present invention may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplification and clarity and does not itself indicate the relationship between the various embodiments and / or settings discussed.

[0056] This application discloses a method, apparatus, device, and storage medium for correcting search terms. By constructing a word recall channel for complex error correction scenarios, users can still find the content they want to search for when they enter incorrect terms, improving search efficiency.

[0057] Refer to Figure 1 , which is a schematic flowchart of a method for correcting search terms provided by an embodiment of this application. The method specifically includes the following steps:

[0058] S101. Obtain a target character input by a user;

[0059] In this application, the target character input by a user is the character entered by the user in the search box of the browser when searching through the browser. In this embodiment, the currently input character by the user is referred to as the target character. The target character may be in the form of Chinese characters or letters. For example, the target character is "kuangchan" entered by the user, or it may be "Ding Baozhen" entered by the user. It is not specifically limited here.

[0060] When this application obtains the target character input by the user, the acquisition condition can be preset in advance. If it is detected that the target character currently input by the user meets the acquisition condition, the target character input by the user will be obtained. For example: set the acquisition condition as: if it is detected that the number of letters input by the user is greater than or equal to 3, then obtain the target character input by the user; if the user is about to input "kuangchan", when the user inputs "k", the acquisition operation will not be triggered, when the user inputs "ku", the acquisition operation will not be triggered either, when the user inputs "kua", the acquisition operation will be triggered, and the target character obtained at this time is "kua", when the user inputs "kuan", the acquisition operation will be triggered, and the target character obtained at this time is "kuan", and so on.

[0061] That is to say, when the target character currently input by the user meets the acquisition condition, the error correction word will be matched through the error correction method of this application and displayed as a recommended search term in the search dropdown box. For example: when the user inputs "kua", the target character "kua" is obtained, and after matching the error correction word using the index file, it is displayed as a recommended search term. If the currently displayed recommended search term is the complete content that the user wants to input, the corresponding recommended search term can be directly selected for searching. If it is not the complete content that the user wants to input, at this time the user will continue to input the next character. For example: the user inputs "kuan", and at this time it is necessary to continue to obtain the target character "kuan" and continue to match the error correction word using the index file, and so on, until the recommended search term includes the content that the user wants to input.

[0062] S102. Search for the target error correction word that matches the target character from the index file; the index file includes: the error correction word corresponding to the input character determined through historical search data;

[0063] In this application, the index file is a file pre-constructed by the user, and the index file stores: the error correction word corresponding to the character input by the user, and the index file is determined through historical search data. For example: by analyzing historical search data, it can be known that the content that the user wants to input is "Ding Baozhen", but due to input errors, the content actually input by the user is "Ding Baozhen", and at this time the error correction word "Ding Baozhen" corresponding to the character "Ding Baozhen" input by the user can be saved in the index file.

[0064] It should be noted that when the present application saves the error correction words corresponding to the characters input by the user in the index file, it can be saved in various ways, and no specific limitation is imposed here. In the present application, the relationship between the characters input by the user and the error correction words can be directly saved in the index file. If the target character currently input by the user can match the corresponding character in the retrieval file, the corresponding error correction word is obtained. For example, the corresponding relationship between "Ding Baozhen" and "Ding Baozheng" is saved in the index file. If the target character currently input by the user is "Ding Baozheng", the matching error correction word is "Ding Baozheng"; it is also possible to save the relationship between the error correction word and a certain type of encoding thereof in the index file. If a certain type of encoding of the target character currently input by the user can match the corresponding encoding in the index file, the corresponding error correction word is obtained. For example, "kuangbiao" and the corresponding encoding are saved in the index file. If the target character currently input by the user is "kuangchan", since the similarity between "kuangchan" and "kuangbiao" is high, the similarity of their encodings is also high. Therefore, when the user inputs "kuangchan", the encoding of "kuangbiao" can be matched in the index file, and then the error correction word "kuangbiao" is matched.

[0065] S103. Display the target error correction word as the recommended search term.

[0066] It should be noted that the solution described in the present application is implemented on the basis of the Suggest module providing search terms, that is: this solution can be implemented through the Suggest module or through a separate error correction module. However, no matter which module is used for implementation, while the Suggest module provides the original search term, the present application can match the error correction word through the index file and display the error correction word as the recommended search term together with the original search term in the search drop-down box. Moreover, when the present application displays the search terms, the display position of the recommended search term can be set according to the search results of the Suggest module. If the Suggest module directly provides a matching search term, the display position of the recommended search term is set at a later position. If the Suggest module does not provide a matching search term, the display position of the recommended search term is set at a forward position. If no matching error correction word is found through this solution, only the search terms provided by the Suggest module are displayed in the search drop-down box. By this means, when the user uses the search engine to search, every time a character is input, the search Suggest module / error correction module will provide a batch of search terms (Query) for the user, aiming to reduce the search cost for the user and improve the search efficiency of the user.

[0067] For example: if the target character input by the user is "kuangbiao", the Suggest module directly provides the matching search term "狂飙", in which case the "狂飙" of the Suggest module can be displayed in the front position, while other recommended search terms obtained through this scheme can be displayed in the back position; if the target character input by the user is "kuangchan", the Suggest module does not directly provide a matching search term, and the corrected recommended search term obtained through this scheme is "狂飙", in which case the "狂飙" obtained by this scheme can be displayed in the front position, while other search terms can be displayed in the back position, thereby improving user search efficiency.

[0068] From the above, it can be seen that the present application pre-establishes an index file through historical search data, and the index file includes correction words corresponding to the input characters; therefore, in the present application, when the user enters the target characters for search, the corresponding target correction words can be found through the index file, and displayed as recommended search words, thereby quickly and accurately providing recommended search words, thereby improving search efficiency.

[0069] See also Figure 2 , is a flow chart of another search word error correction method provided in an embodiment of the present application, the method specifically comprising the following steps:

[0070] S201, determining a first historical search term according to historical search data;

[0071] S202, converting the first historical search term into a historical nine-square grid code;

[0072] S203, constructing a first index file using the historical nine-square grid code as a key and the first historical search term corresponding to the historical nine-square grid code as a value;

[0073] S204, obtaining a target character input by the user;

[0074] S205, converting the target character into a target nine-square grid code;

[0075] S206, searching the first index file for a target historical nine-square grid code that matches the target nine-square grid code;

[0076] S207, using the first historical search term corresponding to the target historical nine-square grid code as the target error correction term corresponding to the target character;

[0077] S208: Display the target error correction word as a recommended search word.

[0078] In this application, the historical search data includes the user's search logs, historical search records, etc., as long as it can contain the historical search information that the user may search. Since the search contents of different users are different, the historical search data obtained when generating the index file can include the historical search data of different users. Different index files are generated for the historical search data of different users. When looking up correction words through the index file, the identity information of the current user can be recognized, the corresponding index file can be found, and the correction words can be matched through the corresponding index file.

[0079] According to the differences in the index file construction process, this application can divide the index file into a first index file and a second index file. Among them, the construction process of the first index file is as follows: determine the first historical search term according to the historical search data, convert the first historical search term into a historical nine-square grid code, use the historical nine-square grid code as the key, and use the first historical search term corresponding to the historical nine-square grid code as the value to construct the first index file.

[0080] In this construction process, the first historical search term determined through the historical search data is specifically the Query (search term) that the user may search. After this application obtains each first historical search term, it needs to be counted and a Query list is generated. Each first historical search term is saved in this Query list. In order to more quickly and effectively match correction words, this application needs to use the nine-square grid coding method to convert each first historical search term into a coding format. This application calls the coding of the first historical search term the historical nine-square grid code. After obtaining each first historical search term and its corresponding historical nine-square grid code, it is necessary to construct an inverted index with the historical nine-square grid code as the key and the first historical search term as the value to generate the first index file.

[0081] The present application uses the nine-square grid encoding method to encode the first historical search term because when the user uses the nine-square grid input, due to the different algorithms of the input method, the input method has different understandings of the input content, and the default recognized pinyin sequence may be far from the pinyin sequence the user wants to enter, causing the user to enter the wrong pinyin sequence. For example: when the user wants to search for Kuangbiao, what needs to be entered is "kuangbiao", but the pinyin sequence of "kuangchan" is actually entered through the nine-square grid input method, which increases the user's input cost. To solve this problem, the present application converts pinyin into a nine-square grid input code, that is, the letters in the same nine-square grid are mapped to the digital value of the current position, such as: abc are mapped to 2, def are mapped to 3, ghi are mapped to 4, and so on. If the first historical search term determined by historical search data is "狂飙", and the pinyin sequence of "狂飙" is "kuangbiao", then each pinyin in "kuangbiao" is converted into a nine-square grid code, that is: map "k" to "5", map "u" to "8", and so on, and the historical nine-square grid code for the first historical search term "狂飙" is "582642426".

[0082] In this embodiment, the first index file can be updated in time according to the addition of historical search data. After the updated first index file is generated, it needs to be loaded so that the user can use the updated first index file to find the error correction word when searching.

[0083] After obtaining the target character currently input by the user, the present application also needs to use the nine-square grid encoding method to convert the target character into a nine-square grid code. The present application refers to the nine-square grid code converted from the target character as the target nine-square grid code; then, a table is looked up in the index according to the target nine-square grid code to find the inverted zipper corresponding to the target nine-square grid code, wherein the inverted zipper is a specific implementation form of the inverted index, including the historical nine-square grid code and the first historical search term, and the inverted zipper corresponding to the target nine-square grid code is: the historical nine-square grid code in the inverted zipper matches the target nine-square grid code, and then the first historical search term in the inverted zipper is recalled for sorting, and the sorting result is fed back to the user.

[0084] For example, if the target character currently input by the user is "kuangchan", the target character "kuangchan" can be converted into the target nine-square grid code using the nine-square grid encoding method:

[0085] "582642426"; After searching, the historical nine - grid code of the first historical search term "Kuáng Biāo" in the first index file is "582642426", and these two codes are exactly the same. At this time, the historical nine - grid code of the first historical search term "Kuáng Biāo" can be used as the target historical nine - grid code. Correspondingly, the first historical search term "Kuáng Biāo" is the target error - correction term. Through this method of the present application, when the user inputs "kuangchan", "Kuáng Biāo" can be recalled through the nine - grid code, improving the user's search efficiency.

[0086] See Figure 3a , which is an offline processing flow chart based on nine - grid coding provided by an embodiment of the present application. In this process, taking "Kuáng Biāo" in the Corpus data (text data set) as an example, the nine - grid code of "Kuáng Biāo" is calculated as: "582642426", and an inverted index is constructed: "582642426" - "Kuáng Biāo"; See Figure 3b , which is an online processing flow chart based on nine - grid coding provided by an embodiment of the present application. In this process, the user inputs the target character "kuangchan", calculates the nine - grid code "582642426" of the target character, queries the inverted index with "582642426" to obtain the error - correction term "Kuáng Biāo", recommends the Query "Kuáng Biāo", and after the user clicks "Kuáng Biāo", enters the search result page.

[0087] It should be noted that since when the present application obtains the target character input by the user, the acquisition condition can be preset in advance. If it is detected that the target character currently input by the user meets the acquisition condition, the target character input by the user is acquired. Therefore, in the present application, the target character is not necessarily a complete pinyin. For example: set the acquisition condition as: if it is detected that the number of letters input by the user is greater than or equal to 3, acquire the target character input by the user; then when the user is about to input "kuangchan" and the character input by the user is "kua", the acquisition operation will be triggered. At this time, the acquired target character is "kua". Therefore, in this case, when searching for the target historical nine - grid code that matches the target nine - grid code from the first index file through the target nine - grid code of the target character, it is also to search for the historical nine - grid code that is the same as the target nine - grid code. For example: the target character is "kua" and the target nine - grid code is "582", then in the first index file, search for the historical nine - grid code whose first three codes are "582". If the historical nine - grid code of the first historical search term "Kuáng Biāo" is "582642426", which is the same as "582", then the historical nine - grid code of "Kuáng Biāo" is the target historical nine - grid code. Correspondingly, the first historical search term "Kuáng Biāo" is the target error - correction term.

[0088] In this application, the first historical search terms in the Query list can be sorted in descending order according to the number of searches, that is: the first historical search terms with more searches are stored in the front positions in the Query list. If the number of matching target correction terms is greater than one, the sorting of the target correction terms is determined according to the sorting of the first historical search terms in the Query list, and then the target correction terms are displayed to the user according to the sorting of the target correction terms, so that the user can see the target correction terms with more searches first.

[0089] In summary, when establishing the index file in this application, the first historical search terms can be found through historical search data, and the historical search terms can be converted into historical nine-grid codes by means of nine-grid coding. The first index file is constructed according to the historical nine-grid codes and the first historical search terms, so as to quickly find the corresponding correction terms through this index file. When finding correction terms through the first index file in this application, the target characters can be converted into the form of nine-grid codes for searching, so as to quickly and accurately find the target correction terms of the target characters.

[0090] See Figure 4 , which is a schematic flowchart of another method for correcting search terms provided by an embodiment of this application. The method specifically includes the following steps:

[0091] S301. Determine the input candidate terms and the second historical search terms according to the historical search data; among them, the input candidate terms are candidate terms with a click-through rate lower than the first predetermined threshold input by the user;

[0092] S302. Calculate the similarity values between each input candidate term and each second historical search term;

[0093] S303. Screen the correction term pairs with similarity values higher than the second predetermined threshold; the correction term pairs include: the input candidate terms and the second historical search terms with similarity values greater than the second predetermined threshold;

[0094] S304. Construct a second index file with the input candidate terms in each correction term pair as keys and the second historical search terms as values;

[0095] S305. Obtain the target characters input by the user;

[0096] S306. Search for the target input candidate terms that match the target characters from the second index file;

[0097] S307. Use the second historical search term corresponding to the target input candidate term as the target correction term corresponding to the target character;

[0098] S308. Display the target correction term as a recommended search term.

[0099] According to the differences in the index file construction process, this application can divide the index file into a first index file and a second index file. Among them, the construction process of the second index file is as follows: determine the input candidate words and the second historical search words according to the historical search data, calculate the similarity values between each input candidate word and each second historical search word, and screen out the error correction word pairs with similarity values higher than the second predetermined threshold; the error correction word pairs include: the input candidate words with similarity values greater than the second predetermined threshold and the second historical search words, and construct the second index file with the input candidate words in each error correction word pair as the keys and the second historical search words as the values.

[0100] In this construction process, the historical search words can be obtained by analyzing the historical search data. The historical search words are specifically the Queries that the user may search. In order to distinguish, this application calls the historical search words obtained in this embodiment the second historical search words. After obtaining each second historical search word, it is necessary to count and generate a Query list, and each second historical search word is saved in the Query list. By analyzing the historical search data, this application can also obtain the input candidate words, which are the candidate words with a click-through rate lower than the first predetermined threshold input by the user, and perform statistics to obtain the input content list; the first predetermined threshold is a pre-set click-through rate threshold, and only when the click-through rate of the candidate word input by the user is lower than this first predetermined threshold, will the candidate word be used as the input candidate word. Among them, the input candidate word is a string that the user often enters in the search box and has a poor click-through rate when searching. For example: when the user wants to search for "Ding Baozhen", but due to input errors, enters "Ding Baozhen" in the search box. At this time, the user will not click "Ding Baozhen" to search, but re-enter the correct "Ding Baozhen" to search. Since the click-through rate of "Ding Baozhen" is low, this application takes "Ding Baozhen" as the input candidate word.

[0101] After this application obtains each second historical search word in the Query list and each input candidate word in the input content list, it is necessary to cross-calculate the similarity relationship between different input candidate words and different second historical search words in the Query list and the input content list, and screen out the error correction word pairs pair with similarity values higher than the second predetermined threshold. The error correction word pairs include: the input candidate words with similarity values greater than the second predetermined threshold and the second historical search words. Finally, construct an inverted index with the input candidate words as the keys and the second historical search words as the values according to the error correction pairs to generate the second index file. Among them, the second predetermined threshold is a pre-set similarity threshold, and only when the similarity value between the input candidate word and the second historical search word is greater than this second predetermined threshold, will an error correction word pair be generated. The second predetermined threshold can be set according to actual needs and is not specifically limited here. In this embodiment, the second predetermined threshold can be set to 0.75.

[0102] When calculating the similarity values between each input candidate word and each second historical search word in this application, various methods can be used for calculation, which are not specifically limited here. For example, the similarity value can be calculated by splitting characters. The calculation process can be as follows: The two characters for which the similarity value is calculated are split into radicals, including top-bottom splitting or left-right splitting, and the probability of each part being the same after splitting is calculated, thereby obtaining the similarity value of the two characters. It is also possible to perform image recognition through a trained model to obtain the similarity value between two characters.

[0103] In this embodiment, the second index file can be updated in a timely manner according to the addition of historical search data. After generating the updated second index file, it is necessary to load this second index file so that when the user searches, the updated second index file can be used to find correction words.

[0104] After obtaining the target character currently input by the user in this application, directly use this target character to look up in the second index file to find the inverted index zipper corresponding to this target character. This inverted index zipper is a specific implementation form of the inverted index, including the input candidate word and the second historical search word. The inverted index zipper corresponding to this target character is: the input candidate word in this inverted index zipper matches the target character; then recall the second historical search words in this inverted index zipper for sorting, and feedback the sorting result to the user.

[0105] See Figure 5a , which is an offline processing flowchart for similarity-based character error correction provided by an embodiment of this application. In this process, taking "Ding Baozhen" in the Corpus data (text data set) as the second historical search word and "Ding Baozhen" as the input candidate word as an example for illustration, in the offline stage, the similarity value between "Ding Baozhen" and "Ding Baozhen" is calculated as: 0.83, which is greater than the second similarity threshold of 0.75, then construct the inverted index: "Ding Baozhen" - "Ding Baozhen"; See Figure 5b , which is an online processing flowchart for similarity-based character error correction provided by an embodiment of this application. In this process, the user inputs the target character "Ding Baozhen", queries the inverted index in the second index file with "Ding Baozhen" as the key, finds Query "Ding Baozhen", recommends Query "Ding Baozhen", and after the user clicks on "Ding Baozhen", enters the search result page.

[0106] It should be noted that when obtaining the target character input by the user in this application, the acquisition condition can be preset in advance. If it is detected that the target character currently input by the user meets the acquisition condition, the target character input by the user is acquired. Therefore, in this application, the target character is not necessarily a complete word. For example: set the acquisition condition as: if it is detected that the number of Chinese characters input by the user is greater than or equal to 2, then acquire the target character input by the user; then when the user is about to input "Ding Baozhen", it is detected that the character input by the user is "Ding Bao", and the acquisition operation will be triggered. At this time, the acquired target character is "Ding Bao". Therefore, in this case, through the target character "Ding Bao", the matching target input candidate word can be directly found from the second index file. Since the first two characters of "Ding Bao" are the same as those of "Ding Baozhen", it is determined that "Ding Baozhen" is the matching target input candidate word, and the second historical search word "Ding Baozhen" corresponding to "Ding Baozhen" is directly used as the target correction word.

[0107] In this application, when sorting the correction word pairs in the second index file, they can be sorted in descending order according to the similarity value between the input candidate word and the second historical search word. If the similarity values of two correction words are the same, they are sorted in descending order according to the search times of the second historical search word. Therefore, if the number of matching target correction words is greater than one, the sorting of each target correction word is determined according to the sorting of each correction pair, and then each target correction word is displayed to the user according to the sorting of each target correction word, so that the user can see the target correction words with high similarity and more search times first.

[0108] In summary, when establishing the index file in this application, the input candidate word and the second historical search word can be found through the historical search data, and the second index file can be constructed according to the similarity between the input candidate word and the second historical search word, so as to quickly find the corresponding correction word through this index file. When finding the correction word through the second index file in this application, the target character input by the user can be directly matched with each input candidate word in the second index file. After finding the matching target input candidate word, the corresponding second historical search word can be used as the target correction word, so as to quickly and accurately find the target correction word of the target character.

[0109] In another embodiment of this application, the process of calculating the similarity value between each input candidate word and each second historical search word specifically includes the following content: convert each character in each input candidate word into a four-corner code, and convert each character in each second historical search word into a four-corner code; calculate the glyph similarity of different character pairs in the input candidate word and the second historical search word according to the four-corner codes of each character in the input candidate word and the corresponding characters in the second historical search word; calculate the similarity value between the input candidate word and the second historical search word according to the glyph similarity of different character pairs in the input candidate word and the second historical search word.

[0110] In this application, through analysis, it is found that some users may have the problem of frequent input errors when inputting, possibly due to similar glyphs. For example, "Ding Baozhen" is misinput as "Ding Baozheng". Therefore, in this application, a four-corner coding method is proposed to solve this problem and improve the input efficiency of users. Specifically, the similarity value between each input candidate word and each second historical search word can be calculated through the four-corner coding method of Chinese characters. The four-corner coding method of Chinese characters is to set the shapes of the four corners of a Chinese character as a four-corner code respectively. By using the four-corner coding method of Chinese characters, the four-corner codes of each character in the input candidate word and the second historical search word are determined. Then, by comparing the four-corner codes of the two Chinese characters in the word pair, the glyph similarity between the two Chinese characters in the word pair can be determined, and further the similarity value between each input candidate word and each second historical search word can be calculated.

[0111] For example: The input candidate word is "Ding Baozheng", and the second historical search word is "Ding Baozhen". Each character in "Ding Baozhen" is converted into the corresponding four-corner code, and each character in "Ding Baozheng" is converted into the corresponding four-corner code. And in the input candidate word: "Ding Baozheng" and the second historical search word: "Ding Baozhen", there are three word pairs: "Ding - Ding", "Bao - Bao", "Zheng - Zhen". This word pair is formed by combining the Chinese characters in the input candidate word and the Chinese characters in the second historical search word according to the character order.

[0112] See Figure 6 , which is the corresponding diagram of the four-corner coding of Chinese characters provided by the embodiment of this application. In this diagram, the corresponding relationship between different glyphs and coding values of Chinese characters is set. According to this corresponding relationship, the character is then converted into a coding value in the order of upper left, upper right, lower left, and lower right. For example: Using the four-corner coding of Chinese characters, "Zhen" in "Ding Baozhen" can be converted into "4498", and "Zhen" in "Ding Baozheng" can be converted into "3428".

[0113] In another embodiment of this application, according to the four-corner codes of each character in the input candidate word and the four-corner codes of the corresponding characters in the second historical search word, the glyph similarity of different word pairs in the input candidate word and the second historical search word is calculated. The specific content is as follows: Determine each word pair in the input candidate word and the second historical search word; Determine the number of the same codes and the total number of codes in the four-corner codes of the corresponding characters of the input candidate word and the four-corner codes of the corresponding characters of the second historical search word in each word pair; Calculate the glyph similarity of each word pair according to the number of the same codes and the total number of codes of each word pair.

[0114] In this embodiment, among the three pairs of characters "Ding-Ding", "Bao-Bao", and "Zhen-Zhen", the two Chinese characters in the pair "Ding-Ding" are exactly the same, and the glyph similarity of this pair of characters can be directly set to 1. The two Chinese characters in the pair "Bao-Bao" are also exactly the same, and the glyph similarity of this pair of characters can be directly set to 1. In the pair "Zhen-Zhen", the four-corner code of "Zhen" is "4498", and the four-corner code of "Zhen" is "3428". The number of identical codes in "4498" and "3428" is 2, and the total number of codes is 4. Then the glyph similarity of this pair of characters is 2 / 4 = 0.5.

[0115] After determining the glyph similarity of each pair of characters in this application, the similarity value between the input candidate word and the second historical search word can be calculated according to the glyph similarity of different pairs of characters in the input candidate word and the second historical search word. In this embodiment, the average similarity of all pairs of characters is calculated as the similarity value between the input candidate word and the second historical search word. For example: the input candidate word is "Ding Baozhen", and the second historical search word is "Ding Baozhen". The glyph similarities of the three pairs of characters here: "Ding-Ding", "Bao-Bao", and "Zhen-Zhen" are 1, 1, and 0.5 respectively. Then the similarity between "Ding Baozhen" and "Ding Baozhen" is (1 + 1 + 0.5) / 3 ≈ 0.83. It can be seen that the similarity between "Ding Baozhen" and "Ding Baozhen" is very high. Therefore, after adding it to the second index file, when the user enters "Ding Baozhen", "Ding Baozhen" can be recalled for the user to achieve error correction.

[0116] When calculating the similarity value between the input candidate word and the second historical search word in this application, the glyph similarity of each pair of characters can be calculated through the four-corner code method. Through the glyph similarity, the similarity value between the input candidate word and the second historical search word can be accurately and effectively calculated. When determining the glyph similarity in this application, specifically, it is based on the number of identical codes and the total number of codes of the four-corner codes of different characters in the pair of characters, so as to accurately calculate the glyph similarity.

[0117] In summary, this application can recall error correction words through the nine-square grid coding method and correct the input content through the four-corner code method. Through this method, the recall error correction in the Suggest scenario can be realized, the recall problem when the user enters wrongly can be solved, the search cost of the user can be saved, and the search efficiency can be improved.

[0118] See Figure 7 , Figure 7 which is a schematic structural diagram of an error correction device for search terms provided by an embodiment of this application. The device specifically includes:

[0119] An acquisition module 11, configured to acquire the target character input by the user;

[0120] A search module 12, configured to search for a target error correction word that matches the target character from an index file; the index file includes: error correction words corresponding to the input characters determined from historical search data;

[0121] A display module 13, configured to display the target error correction word as a recommended search term.

[0122] As an optional embodiment, the error correction device further includes:

[0123] A first determination module, configured to determine a first historical search term according to historical search data;

[0124] A conversion module, configured to convert the first historical search term into a historical nine-grid code;

[0125] A first construction module, configured to construct a first index file with the historical nine-grid code as the key and the first historical search term corresponding to the historical nine-grid code as the value.

[0126] As an optional embodiment, the search module includes:

[0127] A first conversion unit, configured to convert the target character into a target nine-grid code;

[0128] A first search unit, configured to search in the first index file for a target historical nine-grid code that matches the target nine-grid code, and use the first historical search term corresponding to the target historical nine-grid code as the target error correction word corresponding to the target character.

[0129] As an optional embodiment, the error correction device further includes:

[0130] A second determination module, configured to determine an input candidate word and a second historical search term according to historical search data; wherein, the input candidate word is a candidate word input by the user with a click-through rate lower than a first predetermined threshold;

[0131] A calculation module, configured to calculate a similarity value between each input candidate word and each second historical search term;

[0132] A screening module, configured to screen error correction word pairs with a similarity value higher than a second predetermined threshold; the error correction word pairs include: an input candidate word and a second historical search term with a similarity value greater than the second predetermined threshold;

[0133] A second construction module, configured to construct a second index file with the input candidate word in each error correction word pair as the key and the second historical search term as the value.

[0134] As an optional embodiment, the calculation module includes:

[0135] A second conversion unit, configured to convert each character in each input candidate word into a four-corner code, and convert each character in each second historical search term into a four-corner code;

[0136] A first calculation unit, configured to calculate the glyph similarity of different character pairs between the input candidate word and the second historical search term according to the four-corner codes of each character in the input candidate word and the four-corner codes of the corresponding characters in the second historical search term;

[0137] A second calculation unit, configured to calculate the similarity value between the input candidate word and the second historical search term according to the glyph similarity of different character pairs between the input candidate word and the second historical search term.

[0138] As an optional embodiment, the first calculation unit includes:

[0139] A first determination subunit, configured to determine each character pair between the input candidate word and the second historical search term;

[0140] A second determination subunit, configured to determine the number of identical codes and the total number of codes in the four-corner codes of the corresponding characters of the input candidate word and the four-corner codes of the corresponding characters in the second historical search term for each character pair;

[0141] A calculation subunit, configured to calculate the glyph similarity of each character pair according to the number of identical codes and the total number of codes of each character pair.

[0142] As an optional embodiment, the search module includes:

[0143] A second search unit, configured to search for a target input candidate word that matches the target character from a second index file, and use the second historical search term corresponding to the target input candidate word as the target correction word corresponding to the target character.

[0144] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.

[0145] See Figure 8 , Figure 8 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device specifically includes:

[0146] A processor 21, a memory 22, and a computer program stored on the memory 22 and executable on the processor 21. The processor 21 executes the steps of the error correction method described in any of the above method embodiments through the computer program.

[0147] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 21 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0148] The memory 22 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 22 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory 22 is at least used to store the following computer program 221. After the computer program is loaded and executed by the processor 21, it can implement the relevant steps in the error correction method disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 22 may further include an operating system 222 and data 223, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 222 may include Windows, Unix, Linux, etc.

[0149] In some embodiments, the electronic device may further include a display screen 23, an input / output interface 24, a communication interface 25, a sensor 26, a power supply 27, and a communication bus 28.

[0150] Of course, Figure 8 The structure of the shown electronic device does not constitute a limitation on the electronic device in the embodiments of the present application. In actual applications, the electronic device may include more or fewer components than Figure 8 shown, or combine certain components.

[0151] In another exemplary embodiment, a computer storage medium is further provided. When the program instructions are executed by a processor, the steps of the error correction method described in any of the above method embodiments are implemented. Among them, the storage medium may include: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.

[0152] Optionally, specific examples in this embodiment may refer to the examples described in the above embodiments, and will not be elaborated here.

[0153] It should be understood that the terms used herein are only for the purpose of describing specific example embodiments and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" as used herein may also include the plural forms. The terms "comprising", "including", "containing", and "having" are inclusive and thus specify the presence of the stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be performed in the particular order described or illustrated, unless the order of performance is explicitly stated. It should also be understood that alternative or additional steps may be used.

[0154] The above description is only the specific implementation manners of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for correcting search terms, characterized in that: include: Get the target character input by the user; Searching for a target error correction word matching the target character from an index file; The index file includes: error correction words corresponding to the input characters determined by historical search data; The target error correction word is displayed as a recommended search word.

2. The error correction method according to claim 1, characterized in that: Before obtaining the target character input by the user, the method further includes: Determine a first historical search term according to the historical search data; Convert the first historical search term into a historical nine-square grid code; A first index file is constructed by taking the historical nine-square grid code as a key and the first historical search term corresponding to the historical nine-square grid code as a value.

3. The error correction method according to claim 2, characterized in that: The step of searching the index file for a target error correction word that matches the target character comprises: Convert the target character into a target nine-square grid code; Searching the first index file for a target historical nine-square grid code that matches the target nine-square grid code; The first historical search word corresponding to the target historical nine-square grid code is used as the target error correction word corresponding to the target character.

4. The error correction method according to claim 1, characterized in that: Before obtaining the target character input by the user, the method further includes: Determining an input candidate word and a second historical search word according to the historical search data; wherein the input candidate word is a candidate word input by a user whose click rate is lower than a first predetermined threshold; Calculate the similarity value between each input candidate word and each second historical search word; Screening error correction word pairs whose similarity values ​​are higher than a second predetermined threshold; the error correction word pairs include: an input candidate word whose similarity value is greater than the second predetermined threshold and a second historical search word; A second index file is constructed using the input candidate words in each error correction word pair as a key and the second historical search word as a value.

5. The error correction method according to claim 4, characterized in that: The calculating of the similarity value between each input candidate word and each second historical search word includes: Convert each character in each input candidate word into a four-corner code, and convert each character in each second historical search word into a four-corner code; Calculate the glyph similarity between different character pairs in the input candidate word and the second historical search word based on the four-corner code of each character in the input candidate word and the four-corner code of the corresponding character in the second historical search word; The similarity value between the input candidate word and the second historical search word is calculated according to the similarity of the glyphs of different character pairs in the input candidate word and the second historical search word.

6. The error correction method according to claim 5, characterized in that: The step of calculating the glyph similarity between different character pairs in the input candidate word and the second historical search word based on the four-corner code of each character in the input candidate word and the four-corner code of the corresponding character in the second historical search word includes: Determine each word pair between the input candidate word and the second historical search word; Determine the number of identical codes and the total number of codes in the four-corner codes of the corresponding characters of the input candidate word and the four-corner codes of the corresponding characters of the second historical search word in each character pair; The glyph similarity of each character pair is calculated based on the number of identical codes and the total number of codes of each character pair.

7. The error correction method according to any one of claims 4 to 6, characterized in that: The step of searching the index file for a target error correction word that matches the target character comprises: searching the target input candidate word matching the target character from the second index file; The second historical search word corresponding to the target input candidate word is used as the target error correction word corresponding to the target character.

8. A search word error correction device, characterized in that: include: An acquisition module is used to acquire the target characters input by the user; A search module, used for searching a target error correction word matching the target character from an index file; The index file includes: error correction words corresponding to the input characters determined by historical search data; A display module is used to display the target error correction word as a recommended search word.

9. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the steps of the error correction method described in any one of claims 1 to 7 of the present application through the computer program.

10. A computer storage medium, characterized in that: The computer storage medium stores computer executable instructions, and the computer executable instructions are used to execute the steps of the error correction method described in any one of claims 1 to 7 of the present application.