Character sequence searching method and related equipment

By compressing the character sequence and query index, the compressed character sequence and query index are constructed, the problem of character sequence search consumes resources in the prior art is solved and the search efficiency is improved.

CN120104769APending Publication Date: 2025-06-06TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311641174.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-01
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing character sequence similarity search methods consume huge storage space and computing resources, affecting search efficiency.

Method used

By compressing the character sequence to be searched and the character sequence in the query index, the compressed character sequence and query index are constructed, and the compressed character sequence is used for searching.

Benefits of technology

It reduces the storage space and computing resources required for search, improves search efficiency, and realizes cost reduction and efficiency improvement of character sequence search.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104769A_ABST
    Figure CN120104769A_ABST
Patent Text Reader

Abstract

The invention discloses a character sequence searching method and related equipment, and the method comprises the steps: obtaining a to-be-searched first character sequence, and carrying out the compression processing of the first character sequence through employing a target compression mode, and obtaining a first compressed character sequence corresponding to the first character sequence; a query index is obtained, the query index is constructed based on second compressed character sequences corresponding to second character sequences included in the character sequence set, and each second character sequence is compressed in a target compression mode to obtain the corresponding second compressed character sequence, the query index defines an index structure used during character sequence search, and the query index is used for indicating a character sequence search rule; and searching a second character sequence matched with the first character sequence in the character sequence set based on the first compressed character sequence and the at least one second compressed character sequence according to the indication of the query index. The search efficiency can be improved based on the compressed character sequence and the query index.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a character sequence search method, a character sequence search device, a computer device, a computer-readable storage medium, and a computer program product. Background Art

[0002] With the development of computer technology, string (also called character sequence, the characters it contains are text characters) similarity search as a basic operation in data processing has become a core function in many applications, including but not limited to: data cleaning, near-duplicate object detection and data integration. String similarity search makes it more and more convenient to find information in various scenarios. For example, a user can input a character sequence or an object containing a character sequence (such as a picture) through a computer device. The computer device performs a similarity search based on the character sequence, and can find other character sequences that are most similar to the character sequence to obtain search results. However, it has been found in practice that the current search based on character sequence similarity still consumes huge storage space and computing resources, which affects the search efficiency. Summary of the invention

[0003] The embodiments of the present application provide a character sequence search method and related devices, which can improve search efficiency based on a compressed character sequence and a query index.

[0004] On the one hand, an embodiment of the present application provides a character sequence search method, the method comprising:

[0005] Acquire a first character sequence to be searched, and compress the first character sequence using a target compression method to obtain a first compressed character sequence corresponding to the first character sequence;

[0006] Obtaining a query index, where the query index is constructed based on second compressed character sequences corresponding to respective second character sequences included in the character sequence set; wherein each second character sequence is compressed using a target compression method to obtain a corresponding second compressed character sequence; the query index defines an index structure used when searching for a character sequence, and the query index is used to indicate a character sequence search rule;

[0007] According to the indication of the query index, based on the first compressed character sequence and at least one second compressed character sequence, a second character sequence matching the first character sequence is searched in the character sequence set.

[0008] On the one hand, an embodiment of the present application provides a character sequence search device, the device comprising:

[0009] An acquisition unit, used for acquiring a first character sequence to be searched;

[0010] A processing unit, configured to compress the first character sequence using a target compression method to obtain a first compressed character sequence corresponding to the first character sequence;

[0011] The acquisition unit is further used to acquire a query index, where the query index is constructed based on the second compressed character sequences corresponding to the second character sequences included in the character sequence set; wherein each second character sequence is compressed using a target compression method to obtain a corresponding second compressed character sequence; the query index defines an index structure used when searching for a character sequence, and the query index is used to indicate a character sequence search rule;

[0012] The processing unit is further configured to search the character sequence set for a second character sequence matching the first character sequence based on the first compressed character sequence and at least one second compressed character sequence according to the indication of the query index.

[0013] In one aspect, an embodiment of the present application provides a computer device, the computer device comprising:

[0014] a processor suitable for executing a computer program;

[0015] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned character sequence search method is implemented.

[0016] Accordingly, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. The computer program is loaded by a processor and executes the above-mentioned character sequence search method.

[0017] Accordingly, an embodiment of the present application provides a computer program product, which includes a computer program or computer instructions, and when the computer program or computer instructions are executed by a processor, the above-mentioned character sequence search method is implemented.

[0018] In an embodiment of the present application, a first character sequence to be searched can be obtained, and the first character sequence can be compressed using a target compression method to obtain a first compressed character sequence corresponding to the first character sequence; thereafter, a query index can be obtained, the query index being constructed based on the second compressed character sequences corresponding to each second character sequence included in the character sequence set, and each second character sequence is compressed using the target compression method to obtain a corresponding second compressed character sequence; it can be seen that the first compressed character sequence and the second compressed character sequence use the same compression method and are obtained based on the same compression method, which enables the extraction of each character sequence to use a unified compression standard, so as to facilitate the comparison of each compressed character sequence in the character sequence search process. The query index defines the index structure used when searching for a character sequence, and the query index is used to indicate a character sequence search rule; according to the indication of the query index, based on the first compressed character sequence and at least one second compressed character sequence, a second character sequence matching the first character sequence is searched in the character sequence set. Since the query index constructed based on each second compressed character sequence is a simpler and lighter index structure, the character sequence search rule indicated by it can quickly find the second character sequence that matches the first character sequence in the character sequence set, reducing the space occupied by the search and the time spent on the search, thereby improving the search efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0020] Figure 1 is an architectural diagram of a character sequence search system provided by an exemplary embodiment of the present application;

[0021] Figure 2 is a flowchart of a character sequence search method provided by an exemplary embodiment of the present application;

[0022] Figure 3 is a schematic diagram of a character sequence search scenario provided by an exemplary embodiment of the present application;

[0023] Figure 4 is a flowchart of another character sequence search method provided by an exemplary embodiment of the present application;

[0024] Figure 5a is a schematic diagram of a compression process provided by an exemplary embodiment of the present application;

[0025] Figure 5b is a schematic diagram of comparing a compressed character sequence provided by an exemplary embodiment of the present application;

[0026] Figure 6a is a schematic diagram of a query tree structure provided by an exemplary embodiment of the present application;

[0027] Figure 6b It is a schematic diagram of the structure of a multi-level inverted index provided by an exemplary embodiment of the present application;

[0028] Figure 7a is a schematic diagram of searching a character sequence based on a query tree provided by an exemplary embodiment of the present application;

[0029] Figure 7b is a schematic diagram of searching a character sequence based on a multi-level inverted index provided by an exemplary embodiment of the present application;

[0030] Figure 8 is a structural schematic diagram of a character sequence search device provided by an exemplary embodiment of the present application;

[0031] Fig. 9 It is a structural diagram of a computer device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0032] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0033] The present application proposes a character sequence search method, which can use a target compression method to compress the first character sequence to be searched into a first compressed character sequence for a first character sequence to be searched; and can also use a target compression method to compress multiple second character sequences included in the character sequence set to obtain corresponding second compressed character sequences for a character sequence set that provides a search range. Whether it is the first character sequence or the second character sequence, the length of the character sequence can be shortened by compression processing, and the compressed character sequence (such as the first compressed character sequence / the second compressed character sequence) can be regarded as a short representation of the character sequence (the first character sequence / the second character sequence). Since the length of the character sequence becomes shorter after compression, performing a character similarity search based on the compressed character sequence can reduce the amount of character processing and improve search efficiency. In addition, a query index can be constructed based on each second compressed character sequence. Since the compressed character sequence includes a small number of characters, the constructed query index is a simpler and lighter index structure. Moreover, the character sequence search rule indicated by the query index is also based on the search rule of the compressed character sequence. Based on the index structure and search rules provided by the query index, searches can be performed with low space occupancy, thereby reducing the space cost and time cost of the search, further improving the search efficiency, and achieving cost reduction and efficiency improvement of character sequence searches.

[0034] In this application, the terms "first", "second", etc. are used to distinguish between identical or similar items with substantially the same effects and functions. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor is there a limitation on the quantity and execution order. In this application, the term "at least one" means one or more, and the meaning of "multiple" means two or more; for example: multiple second character sequences means two or more second character sequences.

[0035] The character sequence mentioned above refers to a sequence of multiple characters arranged in order, which can also be called a string. Among them, characters can be letters, numbers, symbols, etc.; in this application, a character sequence can represent a paragraph or a text. For example: the character sequence is "weather temperature", and for another example, the character sequence is "@#%¥1244", and for another example, the character sequence is "abode". The character sequence has a length, which refers to the number of characters included in the character sequence. For example, the length of the above-mentioned character sequence "weather temperature" is 4. For another example: the length of the character sequence "abode" is 5. In a feasible implementation, the character sequence can be divided according to the length of the character sequence, and the character sequence may include a long character sequence and a short character sequence. The long character sequence refers to a character sequence whose length is greater than or equal to the preset length threshold, and the short character sequence refers to a character sequence whose length is less than the preset length threshold. Exemplarily, the preset length threshold is 128. If the number of characters included in the character sequence s1 is 200, then the character sequence s1 is a long character sequence; if the number of characters included in the character sequence s2 is 50, then the character sequence s1 is a short character sequence. The length of a character sequence may change before and after compression: the length of a character sequence after compression is smaller than the length before compression, for example, the length of the first compressed character sequence is smaller than the length of the first character sequence, and the length of the second compressed character sequence is smaller than the length of the second character sequence.

[0036] The first character sequence refers to the character sequence to be searched, which can be a character sequence of any length. Since the first character sequence is the query standard referenced when searching for a character sequence, the first character sequence in this application may also be referred to as a query character sequence (denoted as q). The character sequence set is used to provide a search range for a character sequence. The character sequence set refers to a set of multiple second character sequences, which may also be referred to as a data set in this application, denoted as S = {s 1 ,s 2 , ..., s N}, where s i represents the i-th (i∈[1, N]) second character sequence, and N represents the number of second character sequences included in the character sequence set. For example, the character sequence set includes 5 second character sequences, and the character sequence set can be represented as S={s 1 ,s 2 , ..., s 5 Each second character sequence in the character sequence set is used to compare with the first character sequence. In a specific search process, a second character sequence whose similarity with the first character sequence is greater than a certain similarity threshold can be determined as a second character sequence that matches the first character sequence. A second character sequence that matches the first character sequence (for example, is completely the same or has a high similarity) may be searched out from the character sequence set, or a second character sequence that matches the first character sequence may not be searched out.

[0037] The query index defines an index structure used for searching a character sequence, and the index structure is a data structure that arranges and organizes each second compressed character sequence. For example, the index structure defined by the query index may be a tree structure, and each node in the tree structure is used to store a character in each second compressed character sequence; for another example, the index structure defined by the query index may be a table structure, and each second compressed character sequence may be stored in the table structure according to a certain rule. The query index is used to indicate a character sequence search rule, and the character sequence search rule may be adapted to the index structure defined by the query index, and different index structures may have different character sequence search rules.

[0038] The character sequence search method provided in this application can be applied to various business scenarios, including but not limited to: Internet scenarios, local search scenarios and other scenarios. The Internet scenario refers to a scenario where business processing is performed based on the Internet, including but not limited to: online search scenarios, content browsing scenarios, shopping scenarios, advertising recommendation scenarios, and live broadcast scenarios, etc. The local search scenario refers to a search performed locally on a computer device, and the search scope is limited to data stored in the local space. Other scenarios include data cleaning scenarios, reread object detection scenarios, and other scenarios. Corresponding Internet products can be provided in Internet scenarios, and the character sequence search method provided in this application can help improve the search accuracy, efficiency, and user experience of Internet products; based on different types of Internet scenarios, Internet products include but are not limited to: search engines, online media platforms, e-commerce platforms, various service products, and the like. Among them, in search engines, the character sequence search method can help search engines more accurately match users' search keywords and web page content, thereby improving the quality and accuracy of search results; in online media platforms, it can help online media platforms more accurately match users' search keywords and article content, thereby improving user experience and platform security; in e-commerce platforms, it can more accurately match users' search keywords and product information, thereby improving the accuracy and efficiency of product search and recommendation; in the field of electronic resource services, it can more accurately match users' search keywords and product information that supports electronic resource replacement, thereby improving product recommendation effects.

[0039] Based on the above definition, the principle of the character sequence search method proposed in the embodiment of the present application is explained below. Specifically, the general principle of the method is as follows: obtain the first character sequence to be searched, and compress the first character sequence using the target compression method to obtain the first compressed character sequence corresponding to the first character sequence; obtain the query index, and the query index is constructed based on the second compressed character sequences corresponding to each second character sequence included in the character sequence set, wherein each second character sequence is compressed using the target compression method to obtain the corresponding second compressed character sequence; it can be seen that the first compressed character sequence and the second compressed character sequence use the same compression method, and based on the same compression method, the extraction of each character sequence can adopt a unified compression standard, so as to facilitate the comparison of each compressed character sequence in the character sequence search process. The query index defines the index structure used when searching for a character sequence, and the query index is used to indicate the character sequence search rule; according to the indication of the query index, based on the first compressed character sequence and at least one second compressed character sequence, search the character sequence set for a second character sequence that matches the first character sequence. In this process, the query index constructed based on each second compressed character sequence is a simpler and lighter index structure. The character sequence search rule indicated by it can quickly find the second compressed character sequence that matches the first compressed character sequence, so that based on the found second compressed character sequence, the second character sequence that matches the first character sequence can be quickly searched in the character sequence set, thereby reducing the space occupied by the search and the time spent on the search, thereby improving the search efficiency.

[0040] The character sequence search method of the present application is a method for searching based on string similarity. The method can achieve similarity retrieval of character sequences by matching a first compressed character sequence with a second compressed character sequence, and quickly search for one or more second character sequences similar to the first character sequence. In many applications, character sequence similarity search is an important function, including data cleaning, near-duplicate object detection, data integration, etc. By giving a character sequence set and a query character sequence q, we can find the ones that satisfy the query condition under the similarity metric. In a specific implementation, in order to more quickly find a character sequence that matches the query character sequence, the character sequence can be first compressed into a short character sequence for representation. On the one hand, the search can be based on a compressed character sequence of shorter length, which can reduce the resources and time spent on calculation. On the other hand, since the compressed character sequence focuses on key characters that are helpful in determining similarity, the correctness of the character sequence that matches the query character sequence based on the similarity between the compressed character sequences can be guaranteed.

[0041] The method provided in the embodiment of the present application may involve cloud technology, specifically cloud computing and cloud storage in the cloud basic technology category. Cloud computing is a computing model that distributes computing tasks on a resource pool composed of a large number of computers, so that various application systems can obtain computing power, storage space and information services as needed. The network that provides resources is called a "cloud". The resources in the "cloud" are infinitely expandable in the eyes of users, and can be obtained at any time, used on demand, expanded at any time, and paid for by use. For example, the search for character sequences in this application can be understood as a kind of data calculation, which can be implemented based on cloud computing. Cloud storage is a new concept extended and developed from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as a storage system) refers to a storage system that uses cluster applications, grid technology, and distributed storage file systems to bring together a large number of different types of storage devices (storage devices are also called storage nodes) in the network through application software or application interfaces to work together and provide data storage and business access functions to the outside world. For example, all character sequences in the embodiment of the present application can be stored in a distributed storage file system to provide a character sequence search function through a distributed storage file system to improve search efficiency.

[0042] In one implementation, the character sequence search method mentioned above can be executed by a computer device, which can be a terminal or a server. Figure 1 As shown, when a user has a search requirement, a character string can be input through the terminal as the first character sequence q. The terminal sends the first character sequence to the server, and the server can obtain the first character sequence and compress the first character sequence to obtain a first compressed character sequence q′; the server can also obtain a query index from the database, and the query index is constructed based on the second compressed character sequences corresponding to each second character sequence included in the character sequence set, wherein each second character sequence is compressed using the target compression method to obtain the corresponding second compressed character sequence. Before performing a character sequence search, the character sequence set (i.e. ), and construct a query index based on the second compressed character sequences corresponding to each second character sequence in the character sequence set; then, the server can perform a character sequence search based on the first compressed character sequence and each second compressed character sequence according to the instruction of the query index, thereby searching the character sequence set for a second character sequence s matching the first character sequence q. i. In one embodiment, after the query index is constructed for the first time, the query index can be stored in a database, and the query index can be directly obtained from the database for use in subsequent searches without repeated construction, thereby improving search efficiency. Optionally, if the character sequence set is updated, for example, the character sequence set currently obtained has a new second character sequence or has a reduction in the second character sequence compared to the character sequence set obtained last time, then the server or other servers can also update the query index based on the updated character sequence set, whereby the query index after the update is organized into the second compressed character sequences corresponding to the respective second character sequences included in the latest character sequence set, and the query index matches the updated character sequence set. The above-mentioned method can also be executed by multiple computer devices, and the computer device can be a terminal or a server. For example, it can be executed jointly by a terminal and a server. For example: the terminal can obtain a first character sequence and compress the first character sequence, and then send the compressed first compressed character sequence to the server. The server can pre-acquire a character sequence set to build a query index, and then after obtaining the first compressed character sequence, directly search for a second character sequence matching the first character sequence in the character sequence set based on the character sequence search rule indicated by the query index and based on the first compressed character sequence and the second compressed character sequence.

[0043] It is understandable that the above-mentioned terminals include but are not limited to: smart phones, tablet computers, smart wearable devices, smart voice interaction devices, smart home appliances, personal computers, vehicle-mounted terminals, smart cameras and virtual reality devices, etc., and this application does not limit this. This application does not limit the number of terminals. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, but is not limited to this. This application does not limit the number of servers.

[0044] Based on the above description, an exemplary embodiment of the present application proposes a character sequence search method. The character sequence search method can be executed by a computer device (terminal or server) mentioned above, or by a terminal and a server together; for ease of explanation, the following description will be based on an example of a computer device executing the character sequence search method. Figure 2 , the character sequence searching method may include the following steps S201-S203.

[0045] S201, obtaining a first character sequence to be searched, and compressing the first character sequence using a target compression method to obtain a first compressed character sequence corresponding to the first character sequence.

[0046] In a specific implementation, the first character sequence may be a character sequence input by a user, for example, a text sequence input by a user in a service interface of an application (such as a web page of a browser) may be used as the first character sequence. The first character sequence may also be a character sequence extracted from other types of data, for example, if an image input by a user includes text, then the image may be subjected to character extraction processing to obtain a text sequence in the image and use it as the first character sequence. For another example, if a user inputs audio for search, then the audio may be converted to obtain a text sequence corresponding to the audio and use it as the first character sequence. For another example, the first character sequence may be obtained by analyzing the browsing data of the user in the application.

[0047] The first compressed character sequence corresponding to the first character sequence is also the compressed first character sequence. Each character in the first compressed character sequence is a key character extracted from the first character sequence, and the character features it possesses are called key character features. The key characters are used to generate a compressed character sequence with a relatively similar degree of similarity. The character features may include the content features expressed by the characters and the position features of the characters in the original character sequence (here, the first character sequence); the target compression method in this application may be a method of independently selecting key characters from multiple position intervals of the character sequence according to the importance of the character features to achieve compression. Therefore, the general logic of compressing the first character sequence may be: first select the characters in the corresponding position interval from the first character sequence, and then determine the importance of the character features corresponding to the characters, for example, apply an independent random hash function (such as minhash) to determine whether a certain character expresses relatively key information in the first character sequence, that is, judge the importance of the character features. After that, select the characters with higher importance as the target characters, and organize multiple target characters (i.e., key characters) according to the character positions of the target characters in the first character sequence to obtain the first compressed character sequence.

[0048] The above-mentioned target compression method is a character sequence overview extraction, and the first compressed character sequence obtained based on this method can be called a character sequence overview representation (or an overview of the first character sequence, or an overview representation). The character sequence overview extraction is a technology that compresses a character sequence into a short representation. In a specific implementation, by capturing key characters of a character sequence, an overview representation of a character sequence can be constructed based on multiple captured key characters, and in a character sequence similarity search, the overview representation method can be used to quickly find candidate character sequences similar to the query character sequence, because if two character sequences are highly similar, then the matching degree between their overview representations is also high.

[0049] It can be seen that after the compression process, the length of the first compressed character sequence is less than the length of the first character sequence, the first compressed character sequence includes multiple key characters extracted from the first character sequence, and the first compressed character sequence retains the relative position relationship of the multiple key characters in the first character sequence. Exemplarily, the first character sequence is "above", and after the compression process is performed by the target compression method, the first compressed character sequence obtained is "abv". It can be seen that the relative position relationship between the characters a, b and v has not changed, so that the position characteristics of the original character sequence can be maintained, which is conducive to ensuring the accuracy of the character sequence search.

[0050] S202, obtaining a query index, where the query index is constructed based on second compressed character sequences corresponding to each second character sequence in the character sequence set; wherein each second character sequence is compressed using a target compression method to obtain a corresponding second compressed character sequence; the query index defines an index structure used when searching for a character sequence, and the query index is used to indicate a character sequence search rule.

[0051] In a specific implementation, the query index can be pre-built and stored in a database, so that when there is a need to use the query index during a character sequence search, the query index can be obtained from the database. In one implementation, the query index can be built by following the steps (1)-(3).

[0052] (1) Obtain a character sequence set, where the character sequence set includes a plurality of second character sequences.

[0053] (2) Each second character sequence in the character sequence set is compressed using a target compression method to obtain a second compressed character sequence corresponding to each second character sequence. That is, each second character sequence is compressed using a target compression method to obtain a corresponding second compressed character sequence. The second compressed character sequence is a compressed second character sequence, and one second compressed character sequence corresponds to one second character sequence, and different second compressed character sequences correspond to different second character sequences. Exemplarily, a second character sequence in the character sequence set is "abode", and the second compressed character sequence obtained after compression is "abd", that is, the second compressed character sequence is the character sequence after the second character sequence "abode" is compressed. In one embodiment, the computer device can compress each second character sequence in the character sequence set before obtaining the first character sequence, or when the character sequence set is first obtained, so as to directly use the second compressed character sequence corresponding to the second character sequence to construct a query index, thereby speeding up the search process.

[0054] As can be seen from the above, the first character sequence and the second character sequence are compressed using the same compression method (i.e., the target compression method). In one embodiment, the character sequence can be compressed into a compressed character sequence of fixed length by the target compression method, so that the length of each second compressed character sequence is the same, and the length of the first compressed character sequence is also the same as that of the second compressed character sequence. The target compression method can implicitly encode each character sequence, so that the lengths of different compressed character sequences are equal, so that the original character sequences of different lengths are aligned by the compressed character sequences. When performing a character sequence similarity search by aligning the compressed character sequences, the following benefits are achieved: ① Improve comparison efficiency: Aligned character sequences can be compared faster. After alignment, similar character sequences will have similar characters or character blocks (i.e., multiple continuous characters) at the same or similar positions, which makes the comparison operation more direct and efficient. ② Reduce the error rate: Alignment helps to reduce the error rate when calculating the similarity of character sequences. It ensures that when comparing two character sequences, the corresponding characters or character blocks are correctly matched, thereby reducing the possibility of misjudging similar or dissimilar character sequences. ③ Optimize storage and processing: Aligned character sequences are more efficient in storage and processing. They can make the index structure more compact while reducing the computing resources required when processing character sequences. ④ Adapting to long character sequence processing: For long character sequences, alignment can significantly improve processing efficiency. It allows the system to ignore irrelevant parts and focus on those key character sequences that help determine similarity. Therefore, aligning character sequences while compressing them can effectively improve the efficiency and accuracy of character sequence similarity searches, especially when processing long character sequences. This method has important application value in data processing, information retrieval and other fields.

[0055] (3) Constructing a query index based on the second compressed character sequences corresponding to each second character sequence in the character sequence set.

[0056] As can be seen from the foregoing, the query index can be used to arrange and organize character sequences (here, the second compressed character sequences), so the query index can include the second compressed character sequences corresponding to each second character sequence in the character sequence set. However, based on the different index structures defined by the query index, there are also differences in the organization of the second compressed character sequences, so that any second compressed character sequence can appear once or multiple times in the query index. Exemplarily, the index structure defined by the query index is a tree-structured index, and the tree structure is organized by the common prefix characters of each second compressed character sequence, so each second compressed character sequence can appear once in the query index. For another example, the index structure defined by the query index is a table-structured index, in which each second compressed character sequence is organized by character and character position, and the multiple characters included in a second compressed character sequence can cause a second compressed character sequence to appear multiple times in the query index.

[0057] The character sequence search rule indicated by the query index matches the defined index structure. Based on the definition of the query index, the character sequence search rule indicated is a rule for searching the second compressed character sequence. Different index structures correspond to different character sequence search rules. In a tree-structured index structure, each second compressed character sequence corresponds to a search path in the tree structure, so the character search rule may include a search path; and in a table-structured index structure, each second compressed character sequence is stored according to the character classification of each position, so the character search rule may include a search order.

[0058] S203: Search, in accordance with the indication of the query index, in the character sequence set for a second character sequence matching the first character sequence based on the first compressed character sequence and at least one second compressed character sequence.

[0059] In a specific implementation, at least one second compressed character sequence can be traversed based on the first compressed character sequence according to the indication of the query index, and the traversal can be terminated after traversing a preset number of second compressed character sequences that match the first compressed character sequence. Exemplarily, the traversal can be terminated after traversing two second compressed character sequences that match the first compressed character sequence. At this time, the two second compressed character sequences may not be obtained after all the second compressed character sequences are traversed, or the two second compressed character sequences may be obtained after all the first compressed character sequences are traversed. The second character sequences corresponding to the preset number of second compressed character sequences that can be traversed are then determined to be second character sequences that match the first character sequence that are searched from the character sequence set, and the search for the character sequence is terminated.

[0060] In another specific implementation, at least one candidate compressed character sequence can be screened out from each second compressed character sequence based on the first compressed character sequence according to the indication of the query index, each candidate compressed character sequence is a second compressed character sequence matching the first compressed character sequence, and then the second character sequence matching the first character sequence is searched in the character sequence set based on these candidate compressed character sequences. In this way, each second compressed character sequence can be traversed, and a second compressed character sequence matching the first compressed character sequence can be screened out from each second compressed character sequence.

[0061] When searching in a character sequence set, the similarity between the second character sequence corresponding to each candidate compressed character sequence and the first character sequence can be calculated, and then based on the similarity, it is determined whether the second character sequence corresponding to each candidate compressed character sequence can be used as a second character sequence that matches the first character sequence. In this implementation logic, based on the comparison between the first compressed character sequence and the second compressed character sequence, some second character sequences that are already very different from the first character sequence can be filtered out based on the compressed character sequence, thereby reducing the amount of data for similarity calculation; then, to ensure the accuracy of the match, the second character sequence corresponding to the second compressed character sequence is verified based on the similarity, so that the selected second character sequence is a character sequence that meets the query condition and has a high similarity with the first character sequence.

[0062] In a feasible implementation, when screening candidate compressed character sequences, the screening can be based on character difference information between the first compressed character sequence and each second compressed character sequence. The character difference information may include the number of different characters in the same character position between the first compressed character sequence and each second compressed character sequence. If this number is greater than a preset number threshold, then the corresponding second compressed character sequence can be used as a candidate compressed character sequence.

[0063] It is understandable that the character sequence search performed by the method of the present application may not find one or more second character sequences that match the first character sequence. This is because if the query index does not include a second compressed character sequence that matches the first compressed character sequence, then no character sequence that matches the first character sequence can be searched from the character sequence set, and the search result obtained is empty. In this case, a search result prompt can be output, and the search result prompt is used to prompt that the second character sequence that matches the first character sequence has not been searched. Exemplarily, the first character sequence is "house", and the multiple second character sequences included in the character sequence set are "abode", "about" and "abound", respectively. No second character sequence that matches the first character sequence can be searched from the character sequence set.

[0064] The character sequence search performed by the method of the present application may also search for one or more second character sequences in the character sequence set that match the first character sequence. For example, the first character sequence is "above", and the multiple second character sequences included in the character sequence set are "abode", "about" and "abound", respectively. The second character sequence "abode" can be obtained by searching the character sequence set, and the second character sequence matches the first character sequence "above".

[0065] Based on different application scenarios, the computer device may display the search returned data in different forms. For example, by inputting a text, the product carrying the text can be searched for; by inputting an image including a first character sequence, the image including the same character sequence as the first character sequence can be searched for and returned.

[0066] In one embodiment, the method provided in the embodiment of the present application can be applied to an Internet scenario. If one or more second character sequences matching the first character sequence are searched, the computer device can also output the searched second character sequence in the form of data adapted to the Internet scenario. That is, output target business data adapted to the Internet scenario, and the target business data contains a second character sequence matching the first character sequence. Based on different scenario types, in various Internet scenarios, the data type of the target business data also has corresponding differences. Specifically, it may include but is not limited to the contents shown in the following ①-④.

[0067] ① If the Internet scenario includes an advertising scenario, the target business data refers to the advertising data in the advertising scenario. The advertising data here includes but is not limited to a combination of one or more of the following forms: images, text videos, and audio, etc. In the advertising scenario, the determination of the first character sequence can be a search term determined from multiple search terms according to the user's search frequency, or a category label determined based on the number of views of a certain content. Then, after the character search process is performed through the above process, advertisements matching the first character sequence can be searched, thereby recommending advertisements that may be of interest to the user in the application. For example, if a user browses a large number of game videos within a period of time, then through analysis, the first character sequence can be extracted as the tag word: game; thus, game-type advertisements can be recommended to the user.

[0068] ② If the Internet scenario includes a search scenario, the target business data refers to the search data in the search engine; the search data includes at least one of the following search types of data: text, image, video and audio. For example, if a search term is entered in the browser, articles, videos and images related to the search term can be displayed on the browser's web page; further, the type of search data can be selected by the user, for example, if only data of a certain search type (such as articles) is searched under a certain search term, the search data under the selected search type can be displayed.

[0069] ③ If the Internet scenario includes a shopping scenario, the target business data refers to the item data in the corresponding shopping platform. The shopping platform refers to the interactive platform provided by the shopping application, and the item data includes but is not limited to: item details (including item name, item function, etc.), item sales volume, item popularity, etc. For example, Figure 3 A scene diagram of a character sequence search is shown. If a search term, such as "down jacket", is entered in a shopping application, graphic data of various items with item tags including "down jacket" can be displayed in the service interface provided by the shopping application.

[0070] ④ If the Internet scenario includes an electronic resource transaction scenario, the target business data refers to product data that supports replacement through electronic resources. The so-called electronic resources refer to resources that are stored and exchanged in electronic form. In one embodiment, digital assets can be replaced by electronic resources. The digital assets refer to assets that exist in a digitized form. These digital assets can be protected and traded through encryption algorithms and can usually be traded and stored on a blockchain network. Based on this, product data can include the type of data assets and the replacement relationship between digital assets and electronic resources.

[0071] The character sequence search method provided by the present application can compress the first character sequence to be searched into a first compressed character sequence by using a target compression method; for a character sequence set providing a search range, a target compression method can also be used to compress multiple second character sequences in the character sequence set to obtain corresponding second compressed character sequences. Whether it is the first character sequence or the second character sequence, the length of the character sequence can be shortened by compression processing, and the compressed character sequence (such as the first compressed character sequence / the second compressed character sequence) can be regarded as a short representation of the character sequence (the first character sequence / the second character sequence). Since the length of the character sequence becomes shorter after compression, character similarity search based on the compressed character sequence can reduce the number of character processing and improve search efficiency. In addition, a query index can be constructed based on each second compressed character sequence. Since the number of characters included in the compressed character sequence is small, the constructed query index is a simpler and lighter index structure, and the character sequence search rule indicated by the query index is also based on the search rule of the compressed character sequence. Based on the index structure and search rule provided by the query index, it is possible to search with low space occupancy, reduce the space cost and time cost of the search, further improve the search efficiency, and achieve cost reduction and efficiency improvement of character sequence search.

[0072] See also Figure 4 , is another character sequence search method provided by an example embodiment of the present application. The character sequence search method can be executed by the computer device (terminal or server) mentioned above, or can be executed jointly by the terminal and the server; for the convenience of explanation, the following description will be given by taking the computer device executing the character sequence search method as an example. This embodiment mainly introduces the specific implementation principle of character sequence compression, and the character sequence search method may include the following steps S401-S406. Since the following contents will be introduced in this embodiment: compression of character sequences, generation of query indexes, and specific implementation processes of character sequence search, and some definitions may be involved in these specific implementation processes, for the convenience of explanation and understanding, the symbols shown in Table 1 below are defined below.

[0073]

[0074]

[0075] S401, obtaining a first character sequence to be processed.

[0076] S402, obtaining an initialization character sequence, where the initialization character sequence is used to store L characters.

[0077] In a specific implementation, L is determined based on a preset number of times l of recursive processing of the first character sequence, where L is an integer greater than 1 and l is a positive integer. The computer device may initialize a character sequence y of size L. ′ , that is, the initialization character sequence. When the initialization character sequence is obtained, no characters have been stored yet. The length of the initialization character sequence is set to L. Based on the recursive processing logic of the first character sequence, L=2 l -1, l represents the number of times the first character sequence is recursively processed. Based on the setting of the number of recursions, after l recursions, the character sequence of length n can be compressed to a character sequence of length 2. l A compressed character sequence of -1.

[0078] S403: Select a target character from the first character sequence according to the character features of each character in the first character sequence, and store the target character in the initialization character sequence to update the initialization character sequence.

[0079] The target character is a key character in the first character sequence, and the key character refers to a character with key character features. Since the characters in the character sequence are arranged in sequence, each character has a character position in the first character sequence, and the character position can be used to indicate the order of the character in the first character sequence. Exemplarily, the first character sequence q is: "above", the character "a" is the first character, and the character position can be recorded as 1, "b" is the second character, and the character position can be recorded as 2, and so on. q[i] can represent the character at the i-th position in the first character sequence, i represents the character position, and i∈[1, L]. Then for each character in the first character sequence, the character position of the character in the first character sequence can be regarded as a character feature. In addition, characters play an important role in expressing the semantic information of the character sequence. Therefore, the semantic information expressed by the character can also be regarded as a character feature. In summary, the key character feature can include at least one of the following: the position information of the key character in the first character sequence and the semantic information of the key character.

[0080] In one implementation, the characteristics of the first character sequence include the length of the first character sequence. When selecting the target character from the first character sequence according to the character characteristics of each character in the first character sequence, the selection can be made specifically according to the length of the first character sequence and the character position of each character in the first character sequence. The specific implementation logic can include the following steps (1.1)-(1.3).

[0081] (1.1) Obtain an interval length parameter, and determine a target position interval of the first character sequence based on the interval length parameter and the length of the first character sequence.

[0082] Specifically, the interval length parameter ε is used to determine the target position interval of the first character sequence. The interval length parameter may be a preset parameter value, which specifically affects the interval length of the target position interval. Optionally, the interval length parameter may be determined based on one or more of the following: ① Characteristics of the character sequence: Since the key information of character sequences of different types and lengths may be distributed in different areas, selecting a suitable length can help better capture the key features of the string. Therefore, the interval length parameter can be determined based on the length and type of the first character sequence. ② Balancing compression and accuracy: A larger interval length parameter can contain more characters, thereby possibly obtaining a more accurate compression representation, but this may also increase the complexity of processing. A smaller interval length parameter may make the compression process more efficient, but may sacrifice some accuracy. Therefore, the interval length parameter can be determined based on the compression length requirement and the compression accuracy requirement. ③ Performance optimization: In actual applications, the selection of the interval length parameter can also be adjusted according to specific application scenarios and performance requirements. For example, in scenarios with high performance requirements, a smaller length may be selected to increase processing speed. Therefore, the interval length parameter can be determined based on scenario requirements. ④ Experience and experiment: In many cases, the selection of the interval length is based on experience or determined through experiments. Therefore, a series of experiments can be conducted according to actual data sets and application requirements to determine the most suitable value as the interval length parameter.

[0083] The length n of the first character sequence refers to the number of characters included in the first character sequence, that is, n is a positive integer greater than 1. When determining the target position interval based on the interval length parameter ε and the length n of the first character sequence, it can be implemented according to the following expression, specifically: [(1 / 2-ε)n:(1 / 2+ε)n]. Among them, ε>0, the target position interval is a position range composed of multiple consecutive character positions in the middle of the first character sequence, and the interval length of the target position interval is 2nε; when ε≤0.5, the interval length of the target position interval can be made less than or equal to the length of the first character sequence.

[0084] (1.2) A hash function is used to perform hash calculation on characters in the first character sequence whose character positions are within the target position interval to obtain a hash value of at least one character within the target position interval.

[0085] After determining the target position interval, the characters whose character positions in the first character sequence are in the target position interval include one or more. The computer device can apply an independent hash function to each character in the target position interval to perform hash calculations to obtain a hash value for each character in the target position interval. Among them, a hash function is a function used to map any information to a shorter and fixed-length value, and the mapped value is called a hash value (also known as a hash value). The hash function includes but is not limited to: minhash (used to compare the similarity of sets), LSH (Locality Sensitive Hashing) and other hash functions. Based on the advantages of minhash, that is, it is easy to implement and has a higher search accuracy, the hash function in this application can adopt a minhash function.

[0086] Exemplarily, the length of the first character sequence q is 12, and the determined target position interval is [4, 6]. The characters in the target position interval include: q[4]=b, q[5]=0, and q[6]=v. Minhash is used to perform hash calculation on these characters to obtain three hash values, one for each hash value.

[0087] (1.3) According to the hash value of at least one character in the target position interval, a character with the smallest hash value is selected from the at least one character in the target position interval, and the selected character is determined as the target character.

[0088] After the above processing, each character in the target position interval has a corresponding hash value, and the hash values ​​of the characters in the target position interval can be compared to select the character with the smallest hash value as the target character. Furthermore, after each target character is obtained, the target character can be pushed into the initialization character sequence to update the initialization character sequence, and the updated initialization character sequence includes the extracted target character.

[0089] In addition to being used to update the initialization character sequence, the target character can also be used to divide the first character sequence, and perform similar processing on each sub-character sequence obtained by the division, so as to recursively obtain multiple target characters. In this way, the hash values ​​of the target characters extracted from the first character sequence are all minimum hash values, so that the key information of the first character sequence can be retained in the compressed first compressed character sequence, which is beneficial to the subsequent similarity search.

[0090] In general, the target character selection method shown in (1.1)-(1.3) above can recursively obtain the target character from the first character sequence according to the length of the first character sequence and the character position of each character in the first character sequence. In the selection process, by adopting a hash function that retains similarity, the first compressed character sequence can retain the key information in the original first character sequence, so that the second character sequence that matches the first character sequence can be accurately found in the subsequent character sequence search process.

[0091] In another specific implementation, different compression parameters under the target compression method can be used for the first character sequence to obtain multiple first compressed character sequences corresponding to the first character sequence. For example, different hash functions can be used to perform hash calculations in the process of selecting target characters, so that each character in the target position interval corresponds to multiple character hash values, and different target characters can be obtained under different hash functions, thereby obtaining different compressed character sequences, so that the features contained in the first character sequence can be captured from multiple angles, and based on the multiple first compressed character sequences corresponding to the first character sequence, a more comprehensive search can be performed from different angles, thereby obtaining more accurate search results.

[0092] S404: recursively process the first character sequence according to the target character to obtain a first compressed character sequence corresponding to the first character sequence based on the updated initialization character sequence.

[0093] In a specific implementation, the first character sequence can be divided based on the target character, that is, the first character sequence is divided into two sub-character sequences by the target character, and then the sub-character sequence is recursively processed as the next input. This process is repeated l times recursively to obtain a first compressed character sequence, and the first compressed character sequence includes multiple target characters extracted by l recursions. With respect to the above logic, the specific implementation can include the following steps (2.1)-(2.3).

[0094] (2.1) The first character sequence is divided according to the target character to obtain two sub-character sequences.

[0095] The target character selected from the first character sequence may be a key character in the middle position of the first character sequence. Therefore, the first character sequence can be divided into two sub-character sequences by using the target character. And the sub-character sequence does not contain the target character. Exemplarily, the first character sequence is stkil_tdwcqkovgradap, and the target character determined by the above (1.1)-(1.3) is character c. Therefore, the first character sequence can be divided into two sub-character sequences by character c, and the two sub-character sequences are stkil_tdw and qkovgradap.

[0096] (2.2) determining each sub-character sequence as a first character sequence, and repeatedly executing the target step, wherein the target step is: selecting a target character from the first character sequence according to the character features of each character in the first character sequence, and storing the target character in the initialization character sequence, so as to update the initialization character sequence;

[0097] By determining each sub-character sequence as the first character sequence, for each sub-character sequence obtained by division, the target character can be selected in the same way as the first character sequence (such as the process described in (1.1)-(1.3) above), and the selected target character can be stored in the initialization character sequence. Therefore, for the two sub-character sequences obtained by division, the target character can be selected from each sub-character sequence, and the two target characters can be stored in the initialization character sequence. So far, the updated initialization character sequence includes 3 target characters. Furthermore, the sub-character sequence can be divided using the target characters selected from the sub-character sequence and the above process can be repeated to obtain more target characters and construct a first compressed character sequence.

[0098] (2.3) When the number of target characters stored in the updated initialization character sequence reaches the length L of the initialization character sequence, the updated initialization character sequence is determined as the first compressed character sequence corresponding to the first character sequence.

[0099] In a specific implementation, recursively performing the first character sequence includes dividing the first character sequence based on the target character and storing the target character. In this process, based on the number of sub-character sequences obtained by each division, the target step can be executed multiple times, and the sub-character sequences processed by the target step are different when they are executed at different times. Exemplarily, in the first recursion, the first character sequence is divided into two sub-character sequences, and the target character needs to be selected for each sub-character sequence. Therefore, the target step can be executed twice in this recursion, and in the second recursion, since each sub-character sequence can be further divided to obtain 4 sub-character sequences, the target step can be executed four times. Each time the target character is stored in the initialization character sequence, it means an update of the initialization character sequence, and the initialization character sequence fixes the number of characters stored. Therefore, when the number of target characters stored in the updated initialization character sequence is equal to L, it means that the number of times the original first character sequence is recursively processed reaches the preset number of recursions l, then after selecting the target character from each sub-character sequence for the lth time and storing each target character in the initialization character sequence, the updated initialization character sequence can be obtained. Therefore, the updated initialization character sequence can be determined as the first compressed character sequence corresponding to the first character sequence.

[0100] It is worth noting that: ① when storing the selected target characters, the storage order of the target characters can be determined according to the character position of the target characters in the first character sequence. In this way, the position information of the target characters in the original character sequence is taken into account when the target characters are pushed into the initial character sequence, so the relative position relationship between the target characters in the first compressed character sequence finally obtained remains unchanged with the relative position relationship in the first character sequence. That is to say, in the recursive processing process, the target characters selected each time are stored in the initialization character sequence according to their position order in the original character sequence. Exemplarily, if the key character found in the first recursion is a, the key character found in the second recursion is b, and a appears before b in the original character sequence, then in the compressed character sequence y', a will also be in front of b. In this way, not only the character sequence is compressed, but also the relative position information of the characters in the original character sequence is retained in the compressed character sequence, which will play a relatively important role in the subsequent string similarity search. This method achieves effective compression while maintaining the characteristics of the character sequence, making the character sequence similarity search more efficient and accurate. ② In the first recursion, after selecting the target character from each sub-character sequence and storing the target character in the initialization character sequence, the updated initialization character sequence includes multiple target characters. After selecting the target character from the sub-character sequence, each sub-character sequence can be further divided to obtain 4 sub-character sequences, and these sub-character sequences are all used as the first character sequence for the second recursion. And so on, until the lth recursion stops.

[0101] The recursive processing shown in (2.1)-(2.3) above can divide the first character sequence based on the target character, and continuously divide the new sub-character sequence during the recursive process, so as to locate different position intervals, and perform feature analysis on the first character sequence at the granularity of the corresponding position interval, thereby ensuring the accuracy of the compression processing.

[0102] In the above process, based on the preset number of recursions, a character sequence of any length can be compressed into a compressed character sequence of a fixed length. The target compression method can be called MHCompact, which is a method for extracting an overview representation of a character sequence by compressing the character sequence through min-hash. Based on the inspiration of embedding the character sequence into the Hamming space (a space based on binary coding), unlike embedding long and sparse character sequences, MHCompact can compress the character sequence into a short and compact compressed character sequence, thereby retaining the key information of the character sequence while improving the efficiency of similarity calculation in the search process. However, some challenges may be encountered in practice. For some character sequences whose length cannot be divided evenly, it may be impossible to compress the character sequence into a compressed character sequence of a fixed length. In order to solve this problem, when compressing the first character sequence, the computer device can also adopt the following methods shown in ①-③ to ensure the preset length L of the compressed character sequence finally compressed.

[0103] ① Padding: For character sequences whose length cannot be completely recursively extended to the first layer, specific padding characters can be added to expand the character sequence so that its length meets the recursive requirement. This method can ensure that all character sequences can be compressed to a length of 2 l -1. Therefore, before selecting the target character, the computer device can first obtain the length of the first character sequence. If the length cannot meet the recursive requirement, the first character sequence is padded to obtain the padded first character sequence. The padded first character sequence can meet the recursive requirement, so that subsequent processing can be performed based on the padded first character sequence. Conversely, if the length of the first character sequence meets the recursive requirement, the target character can be directly selected based on the first character sequence. The length of the first character sequence meets the requirement that the length can be divisible.

[0104] ② Special processing of short character sequences: For particularly short character sequences, different processing methods can be used, such as using different compression strategies or applying specific hash functions to them, to ensure that they can also be compressed into compressed character sequences of fixed length. Therefore, in the process of selecting the target character from the first character sequence, the matching hash function can be determined based on the length of the first character sequence, and then the determined hash function can be used for hash calculation.

[0105] ③ Hash technology: In min-hash technology, a special hash function can be designed to ensure that character sequences of different lengths have the same length of overview representation after compression (i.e. compressed character sequence). This method does not depend on the actual length of the character sequence, but on the characteristics of the hash function.

[0106] It can be seen that a fixed-length overview representation can be obtained through the above method, that is, even if the lengths of the original character sequences are different, the designed overview representation method can ensure that all compressed character sequences have the same length. This is achieved through the internal design of the algorithm, ensuring that no matter what the length of the input character sequence is, the output is a compressed character sequence of fixed length. That is, MHCompact can ensure that no matter what the length of the original character sequence is, a compressed character sequence of consistent length can be generated in the above method. In this way, the originally unaligned character sequences can be aligned after compression processing, which can play an important role in the subsequent indexing and search process and help improve efficiency and accuracy.

[0107] To facilitate understanding of the compression process shown in S402-S404 described above, Figure 5a The following example content is provided: Assume that the first character sequence q is: stkil_tdwcqkovgradap, the length of the first character sequence is recorded as |q|=19, and the number of recursions is 2. First, the characters in the target position interval of the first character sequence q can be determined, that is, the middle 6 characters "dwcqko", and then minhash is applied to obtain a key character c from the middle 6 character intervals, and then the first character sequence is divided into two sub-character sequences by the key character c, namely stkil_tdw and qkovgradap, and for each sub-character sequence, a new key character is obtained in the same way. Specifically, the key character k and the key character a can be selected from the two sub-character sequences respectively. At this time, the number of recursions is 2, and the first compressed character sequence q corresponding to the first character sequence can be obtained. ′ =cka. In this way, the first character sequence is compressed into a shorter overview character sequence.

[0108] For the above compression process flow, the algorithm logic shown in the following Algorithm 1 can also be provided.

[0109]

[0110] It can be seen from the above algorithm that in order to construct an overview representation (i.e., a compressed character sequence), an independent random hash function (e.g., a minhash function) can be first applied to an interval in the middle of the first character sequence (i.e., the target position interval, which is the middle interval of the character sequence) to find the key character with the minimum hash value, so as to extract the key character (i.e., the target character), and the first character sequence can be divided into two sub-character sequences based on the key character, and then these sub-character sequences can be recursively obtained to obtain more key characters. In the recursive process, the key character can be stored in the overview character sequence y', and the process of storing it in the overview character sequence y' takes into account the position of the character in the original character sequence, so the position order of multiple target characters can be kept unchanged, thereby maintaining the relative position relationship between the selected target characters in the original character sequence. Among them, by using an independent minhash family to compress the long character sequence into a short overview character sequence, the character sequence can be implicitly aligned. Experiments have shown that MHCompact can achieve approximate matching of queries with an accuracy of 0.99. The present application can adopt a string similarity calculation method based on the minimum hash value, which can be used to construct an overview representation of a character sequence (i.e., a compressed character sequence). Specifically, the overview representation is constructed by obtaining key characters in the character sequence and recursively obtaining key characters (i.e., fulcrum characters). Based on the overview representation, a character sequence similarity search can be performed.

[0111] S405, obtaining a query index, which is constructed based on the second compressed character sequences corresponding to each second character sequence in the character sequence set; wherein each second character sequence is compressed using a target compression method to obtain a corresponding second compressed character sequence; the query index defines an index structure used when searching for a character sequence, and the query index is used to indicate a character sequence search rule.

[0112] In a specific implementation, the query index is constructed as follows: a character sequence set is obtained, wherein the character sequence set includes multiple second character sequences; then, each second character sequence in the character sequence set is compressed using a target compression method to obtain a second compressed character sequence corresponding to each second character sequence; and then a query index is constructed based on the second compressed character sequences corresponding to each second character sequence in the character sequence set.

[0113] For the compression processing of each second character sequence, since the same compression method is used as the first character sequence, the compression processing flow of the first character sequence introduced in S402-S404 can be referred to. That is, the first character sequence is replaced with the second character sequence, and the same process is executed to obtain the second compressed character sequence corresponding to the compressed second character sequence. The same processing flow as above is adopted for the second compressed character sequence. Based on the target compression method, characters of any length can be compressed into a compressed character sequence of fixed length. The lengths of the second compressed character sequences are the same, and the lengths of the second compressed character sequences are the same as the first compressed character sequences, so as to ensure the alignment between the character sequences. The search efficiency and accuracy can be improved through alignment.

[0114] This is because, after preliminary experiments based on existing data sets, it was observed that the distribution of characters to be edited in the second character sequence s is close to uniform distribution with a high probability, especially when the character sequence is very long (such as the distribution of spelling errors in an article). If it is assumed that the characters to be edited in the character sequence are uniformly distributed, then the probability that a randomly selected character needs to be edited is k / n, where k is the number of characters to be edited and n is the length of the character sequence. Correspondingly, the probability that the character does not need to be edited is 1-k / n. The probability is high when k / n is small, that is, when the edit distance between the second character sequence s and the first character sequence q is small. When characters are randomly selected from a certain segment of the character sequence, the probability remains unchanged. Therefore, if characters are independently randomly selected from multiple intervals to construct an overview representation of the character sequence, the overviews of similar character sequences are likely to be similar. In other words, if the candidate overview (i.e., the second compressed character sequence) is similar to the query overview (i.e., the first compressed character sequence), then the candidate character sequence (i.e., the second character sequence) is likely to be similar to the query character sequence (i.e., the first character sequence), and the results found by the overview character sequence have high accuracy. Therefore, the present application provides a new approximation method, which can use a series of key characters to build an overview representation, and calculate the number of different key characters in the overview representation between the candidate string and the query string, so as to quickly find the candidate string that needs to be edited. In the process of building the overview representation, an independent random hash function (such as minhash) can be first applied to an interval in the middle of the string to extract the key characters, and the string is divided into two substrings by the key characters. These substrings are then processed recursively to obtain more key characters. Based on the previous assumption, at each recursion, the probability that a string and a query string generate the same key characters is 1-k / n, and the probability of generating different key characters each time is k / n. The above compression processing method is a novel overview representation method, which implicitly encodes and aligns different strings, and guarantees the similarity between the overview representations of similar strings with a high probability.

[0115] Assuming that the characters to be edited in the character sequence are uniformly distributed, then at each recursion, the probability that a second character sequence and the first character sequence are compressed to produce the same key character is 1-k / n; and the probability of generating different key characters each time is k / n. If the two character sequences are the same, their overview representations are likely to be the same, or there are only a few different key characters; conversely, if the two character sequences are not similar, most of the key characters between their overview representations are different. Obviously, if the two character sequences are similar, their overviews are likely to be the same or there are only a few different key characters. On the contrary, if the two character sequences are not similar, most of the key characters between the overview representations are different. Corresponding to the present application, if a second character sequence is similar to the first character sequence, then the second compressed character sequence corresponding to the second character sequence and the first compressed character sequence corresponding to the first character sequence may be exactly the same, or there are only a few different key characters.

[0116] For the above process, you can combine Figure 5b The following example content is provided: Assume that the first character sequence q is: stkil_tdwcqkovgradap, and the second character sequence s is stkilatdwcqkovgradbp; wherein the first length of the first character sequence is recorded as |q|=19, and the second length of the second character sequence is recorded as |s|=20. The underline shown in q represents the position offset between q and s, the edit distance between q and s is 2, and the probability of recursively obtaining the same key character each time is close to 0.9. First, applying minhash can obtain a key character from the middle 6 character intervals of s and q, and the key characters are the same, because s and q have the same interval "dwcqko", and "c" is captured from both s and q. Then s and q are divided into two sub-character sequences, and the key characters are recursively obtained from the center of the 6 character intervals of the sub-character sequences. Finally, s and q are compressed into shorter overview character sequences s′ and q′, which are likely to be the same, or only one character is different. For example, (1) s′=q′=″cka″, or (2) s′=″caa″, q′=″cta″.

[0117] In one embodiment, the specific implementation of the computer device constructing the query index based on the second compressed character sequences respectively corresponding to each second character sequence in the character sequence set may include the following steps 3.1 to 3.3.

[0118] Step 3.1 performs character scanning on the second compressed character sequences corresponding to each second character sequence in the character sequence set to obtain at least one character located at the same position.

[0119] The computer device can perform character scanning on each second compressed character sequence, and obtain the characters included in each second compressed character sequence in the order of character positions. In different second compressed character sequences, the characters at the same position may be the same or different. Exemplarily, each second compressed character sequence includes "abt", "abr" and "abs", and these character sequences have the same prefix ab, but the characters at the third character position are all different; therefore, the character at the first position scanned is character a, the character at the second position is character b, but the characters at the third position involve three, including character t, character r and character s.

[0120] Step 3.2: If the scanned character is not included in the query tree, insert the scanned character into the query tree.

[0121] The scanned characters here refer to at least one character located at the same position. If the query tree is constructed for the first time, the query tree includes a root node but does not include a child node. The root node does not store characters, so the query tree does not include the scanned characters. In the specific implementation process of inserting the scanned characters into the query tree, a corresponding number of child nodes can be created according to the number of scanned characters, and the scanned characters can be stored in the corresponding child nodes, that is, these child nodes in the query tree are used to store the scanned characters; then, according to the position of the scanned characters in the corresponding second compressed character sequence, the depth of the child node is set, and a dependency relationship between existing nodes and child nodes is established; wherein, the existing node can be a root node or a child node that has stored the corresponding character; since the depth of the child node in the query tree is determined based on the position of the stored character in the corresponding second compressed character sequence, the dependency relationship between nodes can be used to indicate the relative position relationship between the characters stored between nodes of different depths.

[0122] To facilitate understanding of the above process, an example is given below. Assume that the compressed character sequence set S ′={aba, abb, aca, baa, bab, bba, bdb}, then the query tree is first constructed based on the compressed character sequence set, and each second compressed character sequence is scanned to obtain the characters at the first position, including: a and b, so that these two characters can be inserted into the query tree. Specifically, two child nodes can be created to store character a and character b respectively. The depth of these two child nodes in the query tree is d=1, and they are connected to the root node. After that, the scanning can be continued to obtain the characters at the second position, including a, b, c and d. These characters can be inserted into the query tree according to the original position relationship. Specifically, since the first position includes different characters, the query tree includes two branches. In the branch where character a is located, two child nodes can be created. These two child nodes are used to store character b and character c respectively, and are connected to the child node storing character a; in the branch where character b is located, three child nodes can be created and used to store characters a, b, d respectively. These child nodes can be connected to the child node storing character b. By analogy, the following can be obtained: Figure 6a In one embodiment, each leaf node in the query tree may be linked to a record list, which includes a second compressed character sequence consisting of characters stored in each node in the path from the root node to the leaf node. A leaf node refers to a child node with the largest depth in the query tree and is also the node farthest from the leaf node, such as Figure 6a The child nodes with depth d=3 in the query tree shown are leaf nodes, and the record lists connected to some leaf nodes are not shown. Based on the second character sequences shown in the above diagram, they can be implicitly aligned after compression, that is, the lengths of the second compressed character sequences are the same, so the maximum depth of the query tree is equal to the length of any second compressed character sequence. Based on the length of the second compressed character sequence being 3, the maximum depth of the query tree is 3, and the depths of the leaf nodes in the query tree are the same.

[0123] Based on the above method, a query tree is constructed. If multiple characters starting from the first character in different compressed character sequences are the same, then the child nodes in the query tree are in the same branch. If there is a difference starting from a certain character, then there is a child node on the corresponding branch in the query tree. The query tree corresponds to at least one search path, each search path refers to a path formed from the root node to the leaf node, and each search path corresponds to a second compressed character sequence. Figure 6a The search path node r-node a-node b-node a in the query tree shown represents a search path, and the corresponding second compressed character sequence is "aba".

[0124] Step 3.3: If the query tree includes the scanned characters, continue scanning until all compressed character sequences are inserted into the query tree.

[0125] In a specific implementation, if at least one of the scanned characters is included in the query tree, it means that the currently scanned character has been previously stored in a node at the correct position in the query tree, because some other compressed character sequence has the same prefix as the second compressed character sequence. If there is a character among the scanned multiple characters that is not included in the query tree, it can be inserted into the query tree in the manner shown in step 3.3 above, and then the characters at other positions in each second compressed character sequence can be continuously scanned, and the characters not stored in the query tree can be inserted into the query tree.

[0126] In one embodiment, before each character sequence search is performed based on the query tree, an initialized tag value (e.g., 0) may be added to each child node in the query tree. Thereafter, based on the first compressed character sequence, when the child nodes in the query tree are traversed to perform a character sequence search, the tag value of the traversed child node may be updated to determine whether the compressed character sequence corresponding to the search path where the corresponding child node is located is a second compressed character sequence that matches the first compressed character sequence based on the updated tag value, thereby determining the search result. Based on the construction of the query tree, the query index can be used to indicate the search path. If different second compressed character sequences have at least one continuous character in common starting from the first character position, then the search paths in the query tree overlap, that is, they pass through at least one common node.

[0127] It can be seen that the process of constructing the query index shown in the above steps 3.1-3.3 can construct a query tree (Trie tree, also known as a dictionary tree) based on each second compressed character sequence, and specifically each second compressed character sequence can be inserted into the tree to construct it. The query tree can be used as a query index to organize multiple second compressed character sequences, and specifically, characters with a common prefix can be organized in the same branch. By using the common prefixes of different second compressed character sequences in the query tree, the query time overhead can be reduced by exchanging space for time, and the unnecessary comparison of character sequences can be minimized to achieve the purpose of improving query efficiency. If a string search is performed based on the query tree, the time complexity of the search is O(LN), where L is the length of the second compressed character sequence and N is the number of second compressed character sequences. Compared with the hash tree, the efficiency of character sequence search based on the query tree is higher.

[0128] In another embodiment, the specific implementation of the computer device constructing the query index based on the second compressed character sequences respectively corresponding to each second character sequence in the character sequence set may include the following steps 4.1 to 4.5.

[0129] Step 4.1 obtains a character set corresponding to the character sequence set, where the character set includes multiple characters.

[0130] In a feasible implementation, the computer device can construct a character set based on the character sequence set, which is denoted as ∑ in this application. The character set construction includes the following two methods: obtaining based on statistics of the uncompressed second character sequence, and obtaining based on statistics of the compressed second character sequence.

[0131] I. Obtained based on statistics of the uncompressed second character sequence.

[0132] In a specific implementation, character statistics may be performed on the second character sequences included in the character sequence set. Specifically, characters in each second character sequence may be extracted, and then a character set is constructed based on the extracted different characters. Then the character set ∑={a, b, c, d, e, i, o, n, r, u, y} obtained after counting each second character sequence. It can be seen that this method counts the characters of each second character sequence (an original character sequence) in the entire character sequence set (a data set) before compressing the character sequence. Its advantage is that it can fully capture the character diversity in the data set and ensure that important character information is not lost during the compression process. In this way, based on the character diversity in the original data set reflected by the character set, it can be ensured that the compressed second character sequence can still effectively represent the original second character sequence.

[0133] II. Obtained based on statistics of the compressed second character sequence.

[0134] In a specific implementation, character statistics may be performed on the second compressed character sequence corresponding to the second character sequence included in the character sequence set, thereby obtaining the character set. The character sequences included therein are all second compressed character sequences, and the character set ∑={a, b, c, d} is obtained based on the statistics of each second compressed character sequence. Since the length of the compressed character sequence is usually shorter, compared with method I, the character set obtained in this method includes fewer characters, the character set will be smaller, and the statistical process will be faster. Based on the reduction in the number of characters in the character set, the search speed can be accelerated.

[0135] In practical applications, the method to choose to construct the character set may depend on the specific application scenario and data characteristics. If the compression algorithm changes the frequency of occurrence of characters to some extent or ignores certain characters, this may result in the final character set being insufficient to represent the diversity of the entire data set. In some cases, in order to ensure the integrity and diversity of the data, statistics based on the uncompressed second character sequence are more appropriate; in other cases, if the compression algorithm can effectively retain character information, statistics based on the compressed second character sequence is a more efficient choice.

[0136] Step 4.2 traverses the characters included in the character set, determines the traversed character as the current character, and selects at least one second compressed character sequence including the current character from the second compressed character sequences corresponding to each second character sequence in the character sequence set according to the current character.

[0137] In a specific implementation, the characters in the character set can be arranged in sequence. For example, each character in the character set is a letter, and the character set can be called an alphabet, so the characters in the character set can be arranged in alphabetical order. The computer device can traverse the characters in the character set in order, and the traversed characters can be used as the current character; regardless of whether the characters in the character set are derived from the second character sequence or from the second compressed character sequence, based on the correspondence between the second compressed character sequence and the second character sequence, there is at least one second compressed character sequence in each second compressed character sequence that contains the current character. Therefore, at least one second compressed character sequence can be selected from each second compressed character sequence according to the current character, and each second compressed character sequence includes the current character. Exemplarily, a compressed character sequence set for storing second compressed character sequences For the character set Σ={a,b,c,d}, when the traversed character is character a, 4 second compressed character sequences can be selected, namely {aca, baa, bab, bba}.

[0138] The second compressed character sequence is selected to construct a more lightweight query index. Compared with directly using the second character sequence to construct the query index, on the one hand, the second compressed character sequence can effectively represent the second character sequence, and the search for the second compressed character sequence can indirectly represent the search for the second character sequence; on the other hand, the size of the second compressed character sequence is smaller, and the number of characters that need to be processed in the process of searching based on the query index is smaller, which can speed up the search of the character sequence.

[0139] Step 4.3 obtains the reference character position of the current character in each selected second compressed character sequence, and adds each selected second compressed character sequence to the record list corresponding to the current character in the corresponding reference character position.

[0140] In a specific implementation, the current character corresponds to one or more character positions in each selected second compressed character sequence, which are referred to as reference character positions in this application. The number of times the current character appears in the same second compressed character sequence is equal to the number of reference character positions, that is, if a second compressed character sequence includes multiple current characters, then the reference character position includes multiple character positions. Exemplarily, using the above example, the current character is character a, and the character a appears twice in the second compressed character sequence aca, then the reference character position of character a in the second compressed character sequence aca includes position 1 and position 3, and the reference character position of character a in the second compressed character sequence baa includes position 2 and position 3, the reference character position of character a in the second compressed character sequence bab includes position 2, and the reference character position of character a in the second compressed character sequence bab includes position 3.

[0141] The current character has a corresponding record list at the reference character position, and the record list can be used to record the character at the reference character position as the second compressed character sequence of the current character. Based on the number of reference character positions, the same second compressed character sequence can be added to different record lists. For example, the second compressed character sequences selected above are: {aca, baa, bab, bba}, then the current character (i.e., character a) can record "aca" in the record list at the first position, the current character a can record "baa" and "bab" in the record list at the second position, and the current character a can record "bba" in the record list at the third position.

[0142] Step 4.4: After all characters in the character set are traversed, a list of records corresponding to each character in the character set at each character position is obtained.

[0143] For each character in the character set, the corresponding second compressed character sequence can be selected according to the above process, and the second compressed character sequence can be recorded in the corresponding record list. In this way, a record list corresponding to each character in the character set at each character position can be obtained. In one embodiment, the length of the compressed character sequence is L, so the number of character positions can be L. Since the first compressed character sequence and the second compressed character sequence are obtained by the same compression method, the first compressed character sequence and the second compressed character sequence are aligned, so setting the number of character positions to L can help more efficiently find the record list in which the character at each character position is a character in the first compressed character sequence.

[0144] It is understandable that, since the position of a character in the character set in the corresponding second compressed character sequence does not necessarily cover all character positions, a certain character does not necessarily have the second compressed character sequence recorded in the record list corresponding to each character position. For example, for character c, only the second compressed character sequence aca includes character c, and character c has the second compressed character sequence aca recorded in the record list of the second character position, but character c does not have any second compressed character sequence recorded in the record lists of the first character position and the third character position.

[0145] In one embodiment, each character in the character set or the second compressed character sequence itself can be used as an index to implement the search for the character sequence. In another embodiment, a sequence number can be assigned to each second compressed character sequence, and a character number can be assigned to each character in the character set. Then, when recording the second compressed character sequence to the record list, the sequence number of the second compressed character sequence can be specifically stored in the record list, so that when searching for the character sequence, the record list of the corresponding character at each character position can be searched according to the character number of each character, and the corresponding second compressed character sequence can be found based on the sequence number recorded in the record list. In this embodiment, the sequence number, character number, etc. can be used as an index to find the character sequence.

[0146] Step 4.5 determines each character in the character set and the record list corresponding to each character at the same character position as the inverted index of the corresponding character position, and integrates the inverted indexes of each character position to obtain a query index.

[0147] The computer device can determine all the characters in the character set and the record list of each character at a certain character position as an inverted index of a certain character position (also known as a first-level inverted index). It is called an inverted index because compared to the forward index (forward index) that obtains the characters included therein through a character sequence, the present application can find a character sequence containing the character or related to the character through a character during the search process. Therefore, this index is called an inverted index (inverted index, or reverse index), that is, an index that finds the character sequence through the reverse order of the characters. Inverted indexes of L character positions can be obtained in the above manner, and the inverted index of each character position is a first-level inverted index, and then L-level inverted indexes can be obtained (since L is greater than 1, it can also be called a multi-level inverted index), and the query index can be composed of these L-level inverted indexes.

[0148] Corresponding to the length L of the compressed character sequence, for each position j of the compressed character sequence, there is a level inverted index, each level inverted index consists of a record list of characters c∈∑, and each record list of character c contains the second compressed character sequence of character c at position j. For the query index constructed in the above manner, the following can be provided Figure 6b The structure diagram of the query index is shown in FIG. 1 , where the character set ∑ = {a, b, c...x, y, z}, and the character position includes L positions, and each character c at position j can be recorded as c j , j∈[1, L]. For each character position j, there is a corresponding record list of each character, and the record list of any character is used to record the second compressed character sequence of the corresponding character at the character position. 1 When the character at position is character a, the record list of character a is used to record the second compressed character sequence whose first character position is character a. It should be noted that Figure 6b Only the record lists of a and b at various character positions are shown, while the record lists of other characters at various character positions are not shown.

[0149] For the above construction process, the algorithm logic shown in the following Algorithm 2 can be provided.

[0150]

[0151]

[0152] As shown in the algorithm logic above, the index can be initialized using level L. Each second character sequence is converted into the second compressed character sequence s′ i , the target compression method MHCompact is used in the conversion process. For the second compressed character sequence s′ i Each character c in position j can be the second compressed character sequence s′ i Added to the j-level record list of c in the inverted index, represented as After the second character sequence is processed, an index can be set, for example, a sequence number is set for each second compressed character sequence, and finally a query index (such as Figure 6b shown).

[0153] In one embodiment, the first compressed character sequence and the second compressed character sequence are aligned, which ensures that in the character sequence similarity search, both the second character sequence and the first character sequence in the character sequence set can be processed in an efficient and unified manner. For example, the query index constructed here can be applied to any first character sequence without considering the length of the first character sequence.

[0154] It can be seen that the process shown in the above steps 4.1 to 4.5 can construct a multi-level inverted index (also called a multi-level reverse index / multi-layer reverse index, minIndex) as a query index based on the second compressed character sequence and each character. Since the character sequence set contains a large number of second character sequences (i.e., the data set size is large), the space consumption of the Trie tree-based index is still not negligible. Compared with the Trie tree-based index, the multi-level inverted index does not have an overly complex spatial structure for the organization of the second compressed character sequence. It is a simpler and smaller index structure. Therefore, searching for character sequences based on the multi-level inverted index can further reduce space consumption and improve search efficiency.

[0155] In addition, unlike the existing indexes that store a large number of redundant sub-character sequences, the multi-level inverted index retains sketch character sequences with less redundancy and does not use additional structures compared to the index based on the Trie tree. By searching the compressed character sequence through the multi-level inverted index, lower space consumption can be achieved and a better balance can be achieved in time and space. Among them, the characters included in the sketch character sequence are called sketch characters, which refer to representative characters selected in a specific character sequence summary or compressed representation. In this application, these sketch characters are selected from the original character sequence to represent or summarize the key features of the original character sequence. The selection of sketch characters is usually based on a certain algorithm or rule, such as the selection of key characters in the minimum hash (MinHash) algorithm. Through these representative sketch characters, the resources required for processing and storage can be reduced while retaining the key features of the original character sequence. This method plays an important role in applications such as data compression, index construction, and fast search.

[0156] S406: Search the character sequence set for a second character sequence matching the first character sequence based on the first compressed character sequence and at least one second compressed character sequence according to the indication of the query index.

[0157] In one embodiment, the query index includes a query tree, and the character sequence search rule includes a search path; the query tree is linked to multiple record lists, each of which is used to record a second compressed character sequence corresponding to a search path. A search path refers to a path formed by a root node in the query tree to reach a leaf node, and each child node included in the path stores characters. A search path corresponds to a second compressed character sequence; a record list only records one second compressed character sequence, and the number of record lists is equal to the number of second compressed character sequences. Based on this, when the computer device executes the above S406, it can specifically execute according to the following steps 5.1-5.3.

[0158] Step 5.1 searches the query tree along the search path indicated by the query index based on the character difference information between the first compressed character sequence and the second compressed character sequence corresponding to the search path to obtain a target record list.

[0159] In a specific implementation, based on the characters stored in the nodes in the search path and the characters in the corresponding positions in the first compressed character sequence, the character difference information between the first compressed character sequence and the second compressed character sequence corresponding to the search path can be determined, and the character difference information includes the number of different characters in the same position between the first compressed character sequence and the corresponding compressed character sequence; if the characters in each position of the second compressed character sequence and the first compressed character sequence are exactly the same, then the number can be 0; if the characters in at least one position of the second compressed character sequence and the first compressed character sequence are different, then the number can be non-zero. Based on the character difference information, the query tree is searched to determine which search paths in the query tree correspond to the second compressed character sequence and the first compressed character sequence, so that at least one record list in each record list linked to the query tree can be determined as the target record list, that is, the target record list includes at least one record list linked to the query tree. The character difference information between the second compressed character sequence and the first compressed character sequence recorded in these record lists is within the difference allowable range. In other words, the second compressed character sequence recorded in the target record list matches the first compressed character sequence.

[0160] In one achievable manner, based on the aforementioned query tree construction process, the query tree may include a root node and multiple child nodes, the root node does not store characters, and each child node is used to store a character in the corresponding second compressed character sequence, and the characters stored in each child node with the same parent node are different. The query tree includes multiple child nodes on the same search path, and the child node with the largest number of nodes separated from the root node is the child node, that is, the node with the greatest depth in the query tree, and the record list linked to the query tree is linked to the leaf node, that is, one leaf node is linked to one record list. And the record list linked to each leaf node records the second compressed character sequence represented by the search path from the root node to the leaf node. Exemplarily, as mentioned above Figure 6a The leaf node a shown is linked to a record list, and the record list records the second compressed character sequence "aba". Based on this, the computer device can specifically implement the following steps A to D when executing the above step 5.1.

[0161] Step A: Starting from the root node of the query tree, the query tree is traversed along the search path indicated by the query index, and the traversed child node is used as the current child node.

[0162] In a specific implementation, the computer device may start from the root node and traverse each child node in the query tree along the search path in sequence. The following is an exemplary description using a traversed child node as the current node. Each traversed child node may be processed according to the same process.

[0163] Step B: determining the target character position corresponding to the character stored in the current child node in the corresponding second compressed character sequence according to the depth of the current child node in the query tree, and comparing the character stored in the current child node with the character at the target character position in the first compressed character sequence.

[0164] In a specific implementation, the depth of a child node in the query tree is used to indicate the character position of the character stored in the corresponding child node in the corresponding second compressed character sequence. Therefore, the computer device can first determine the target character position corresponding to the character stored in the current child node in at least one second compressed character sequence according to the depth of the current child node. Figure 6a Taking the query tree shown as an example, the current child node is a node storing character a among all nodes with a depth of 1, then it can be determined that the character stored in the current child node is at the first character position. Furthermore, the character at the target character position can be determined from the first compressed character sequence. For example, if the target character position is the first character position, it can be determined that the character at the target character position in the first compressed character sequence is also a. Thus, the character stored in the current child node is compared with the character at the target character position in the first character sequence, that is, the characters at the same character position in the first compressed character sequence and the second compressed sequence can be compared to see if they are the same, so as to obtain the character difference information between the two compressed character sequences.

[0165] Step C: if the character stored in the current child node is different from the character at the target character position in the first compressed character sequence, the tag value of the current child node is updated, and when the updated tag value is greater than the preset tag value, the branch where the current child node is located is pruned.

[0166] In a specific implementation, if they are different, it means that the second compressed character sequence corresponding to one or more search paths where the current child node is located has at least one different character from the first character sequence. The tag value Used to record: the number of difference characters accumulated when the first compressed character sequence and the corresponding second compressed character sequence reach the position indicated by the current child node. In one embodiment, before searching, the tag value of each child node is initialized to 0, and when the current child node is reached during the search process, it can be updated based on the accumulated number of difference characters. Exemplarily, the character stored in the current child node is the character at the first character position, and is different from the character at the first character position in the first compressed character sequence, then the tag value of the current child node can be updated from 0 to 1 to indicate that the number of difference characters accumulated at the first character position between the first compressed character sequence and the corresponding second compressed character sequence is 1. The updated tag value of the current child node can also be passed to other child nodes that have not been traversed in one or more search paths where the current child node is located, so as to facilitate updating based on the latest tag value when the node is traversed later. The transfer here can also be understood as updating the tag values ​​of other child nodes that have not been traversed in one or more search paths where the current child node is located.

[0167] For example, the character stored in the current child node is the character b at the first character position in the second compressed character sequence, and the character at the second character position in the first compressed character sequence is the character a, so the tag value of the current child node can be updated to 1, and the tag value is passed to the child node used to store the character a. If the child node used to store the character a is subsequently traversed, since the character at the same position in the first compressed character sequence is different, the tag value 1 can be further updated to 2. This can also indicate that there are a total of two different characters at the same position between the second compressed character sequence corresponding to the search path and the first compressed character sequence.

[0168] Further, based on the updated tag value of the current child node The size relationship between the preset mark value α can determine whether to prune the branch where the current child node is located. The preset mark value α is the upper limit of the cumulative number of different characters between the first compressed character sequence and the second compressed character sequence. For example, α=1, which means that at most one character position is allowed to be different between the first compressed character sequence and the second compressed character sequence. If the preset mark value is exceeded, the corresponding second compressed character sequence can be discarded. Among them, the branch where the current child node is located may involve one or more search paths. If the branch is pruned, one or more second compressed character sequences can be excluded, and then the corresponding second character sequence can be excluded. By pruning the corresponding branches, the leaf nodes in the pruned branches will not be traversed when traversing the query tree.

[0169] Step D: If the comparison shows that the character stored in the current child node is the same as the character at the target character position in the first compressed character sequence, or the updated tag value of the current child node is less than or equal to the preset tag value, continue to traverse the query tree until the leaf node of the query tree is traversed, and determine the record list linked to the traversed leaf node as the target record list.

[0170] In one specific implementation, if they are the same, it means that the characters at the same character position between the first compressed character sequence and the second compressed character sequence are also the same, so the tag value of the current child node can be kept unchanged and the query tree can be traversed continuously. In another specific implementation, if they are different, since the tag value of the current child node will be updated to record the number of different characters, and the updated tag value and the preset tag value will be distinguished, under the condition that the updated tag value is less than or equal to the preset tag value, it means that the number of different characters between the first compressed character sequence and the second compressed character sequence is allowed, and the query tree can be traversed continuously.

[0171] It is understandable that, in the process of continuing to traverse the query tree, a new child node can be traversed as the current child node, and the above steps B to D are recursively executed, and finally a leaf node can be traversed. If a leaf node is traversed, it means that the second compressed character sequence represented from the root node to the leaf node matches the first compressed character sequence, and then the record list linked to the traversed leaf node can be determined as the target record list, thereby obtaining the second compressed character sequence recorded in the target record list.

[0172] For the process shown in the above steps A to D, Figure 6a Based on the query tree shown, Figure 7a The character sequence search process based on the query tree is exemplarily described. Assume that the first compressed character sequence q′=″aba″, the preset tag value α=1, and traverse all nodes in the query tree from the root node r to the leaf node. When node a is encountered in a node with a depth of 1, since node a stores the character a, which is equal to q′[0]=″a″, and node b stores the character b, which is not equal to q′[1]=″b″, the tag value of node a remains 0, while the tag value of node b is updated to 1 and passed to their child nodes. When node α is encountered in a node with a depth of 2, since it is not equal to q′[1]=″b″, its tag value becomes 2, and the tag value is greater than the preset tag value α, so its successor nodes can be pruned (including nodes a and b with a depth of 3, whose parent nodes are both node a with a depth of 2). Similarly, the successor nodes of node d with a depth of 2 can be pruned. For details, please refer to Figure 7aThe nodes in the dashed box are the nodes to be pruned. Once a leaf node is encountered and its label does not exceed α, the record list linked to it is the target record list. The record list can be further filtered and merged into the result.

[0173] Step 5.2 performs filtering processing on the target record list to obtain a filtered target record list. Based on the target record list including at least one record list, the filtered target record list includes the at least one filtered record list.

[0174] In one embodiment, the filtering process here may include at least one of the following: length filtering and position filtering. This is because, although each character sequence can be compressed into a compressed character sequence of the same length to achieve character sequence alignment, the lengths of the original character sequences may be different, resulting in a large length difference between the second character sequence finally searched and the first character sequence, so the second character sequence with a large length difference can be filtered out by length filtering; in addition, character sequences with low similarity may also produce similar compressed character sequences due to coincidence, so it is also necessary to compare the positions of the same characters in the compressed character sequence at the same position in the original character sequence, so as to filter out the second character sequence with a large position difference. Therefore, after determining the second compressed character sequence that matches the first compressed character sequence, it is necessary to further determine the difference information between the first character sequence and the corresponding second character sequence, and through the difference between the original character sequences, the second character sequence that matches the first character sequence can be more accurately screened out.

[0175] The specific implementation of length filtering and position filtering is introduced in detail below.

[0176] (i) Length filtering. In order to implement length filtering, the length of the second character sequence corresponding to the second compressed character sequence can be recorded in the corresponding record list, and the target record list also records the length of the second character sequence corresponding to the corresponding second compressed character sequence. The target record list includes at least one record list, and different record lists record different second compressed character sequences. In addition to recording the second compressed character, each record list also records the length of the second character sequence corresponding to the corresponding second compressed character sequence. Exemplarily, the second compressed character sequence recorded in the record list list1 is "aba", and the second compressed character sequence is obtained by compressing the second character sequence "abstract", so the length recorded in the record list is the length 8 of the second character sequence abstract. Based on this, the specific process of length filtering includes the following contents (i)-(iii).

[0177] (i) Get the length of the first character sequence.

[0178] The length of the first character sequence refers to the number of characters included in the first character sequence. For example, the first character sequence q is: stkil_tdwcqkovgradap, and the length of the first character sequence is |q|=19. The computer device can count the number of characters included in the first character sequence to obtain the length of the first character sequence.

[0179] (ii) determining a length difference between the first character sequence and a corresponding second character sequence based on the length of the first character sequence and the length of a record in the target record list.

[0180] The target record list includes at least one record list, and the computer device can subtract the length of the first character sequence from the length recorded in each record list included in the target record list to obtain the length difference between the two character sequences, and the absolute value of the length difference can be used as the length difference between the first character sequence and the corresponding second character sequence. Exemplarily, the length of the first character sequence is 19, the target record list includes a record list, and the length of the second character sequence recorded in the record list is 20, so that the length difference between the first character sequence and the second character sequence recorded in the record list is 1. If the target record list includes two record lists, the length of the records in one record list is 20, and the length of the records in the other record list is 24, then two length differences can be obtained, which are 1 and 5 respectively, and the two length differences correspond to different second character sequences respectively.

[0181] (iii) If the length difference is greater than the first difference threshold, the second compressed character sequence recorded in the corresponding record list included in the target record list is removed to obtain a filtered target record list.

[0182] The first difference threshold sets an upper limit of the length difference between the second character sequence and the first character sequence, and the first difference threshold can be defined by the user or set based on an empirical value. For example, the first difference threshold is 1, indicating that the length difference between the two character sequences must be less than or equal to 1 in order to retain the second compressed character sequence recorded in the record list, so as to retain the second character sequence corresponding to the second compressed character sequence.

[0183] Each length difference determined based on the length of the first character sequence and the length of each record in the record list shall be compared with the first difference threshold. If a certain length difference exceeds the first difference threshold, it means that the length difference between the corresponding second character sequence and the first character sequence is large, then the second compressed character sequence recorded in the corresponding record list can be eliminated to exclude the second character sequence to which the second compressed character sequence should be in the final search results. If a certain length difference is less than or equal to the first difference threshold, then the second compressed character sequence recorded in the corresponding record list can be retained for further verification based on similarity. Exemplarily, if the first difference threshold is 1, the determined length difference includes the length difference of 1 between the first character sequence and the second character sequence corresponding to the second compressed character sequence in the record list list1, and the length difference of 5 between the first character sequence and the second character sequence corresponding to the second compressed character sequence in the record list list2. Since the length difference of 1 is equal to the first difference threshold of 1, the second compressed character sequence in the record list list1 can be retained. If the length difference of 5 is greater than the first difference threshold of 1, then the second compressed character sequence in the record list list2 can be deleted. According to the above method, the filtered target record list includes at least one filtered record list, and a certain filtered record list may remain unchanged from the record list before filtering due to the judgment of length difference, or may not include the second compressed character sequence originally recorded. For example, in the above example, the filtered record list list2 does not include the original second compressed character sequence because the second compressed character sequence is removed, while the filtered record list list1 still retains the second compressed character sequence originally recorded, that is, the filtered record list list1 and the original record list list1 remain unchanged.

[0184] It can be seen that the computer device performs filtering by appending the length of the second character sequence in each record list. When searching for a record list, the length of the record in the record list can be compared with the length of the first character sequence. If the length difference is greater than a given first difference threshold, the record list can be pruned. This is because the second character sequence with a too large length difference does not match the first character sequence, so that at least one record list can be screened out from multiple record lists as a target record list.

[0185] In a specific implementation, the computer device can call a traditional length filter to perform the above-mentioned length filtering process; the basic idea is to filter out candidate character sequences with legal lengths within a length range according to the length of the query character sequence and a set threshold, and then further process these candidate character sequences.

[0186] In another specific implementation, the computer device can call a character sequence processing model to perform the above-mentioned length filtering process. The character sequence processing model is a model trained based on machine learning and used to implement length filtering. By training the model, the second character sequence with a length within a specified range can be retrieved more efficiently and the query efficiency can be improved. The model is an application embodiment of a learning index, which can replace the traditional length filter to achieve more efficient and accurate filtering processing, thereby further improving the accuracy of search results. Specifically, during the model training process, a compressed character sequence and a specified length can be input, and after the model is processed, a character sequence of the specified length can be output. Furthermore, during the model application process, only the second compressed character sequence recorded in the target record list and the length of the first character sequence need to be input into the model, and a second character sequence with the same length or a smaller difference as the first character sequence can be obtained. In this way, artificial intelligence technology is involved, specifically natural language processing (NLP) technology and machine learning (ML) technology. Among them, natural language processing is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can realize effective communication between people and computers using natural language. Natural language processing involves natural language, which is the language people use in daily life. It is closely related to linguistic research. It also involves computer science and mathematics. It is an important technology for model training in the field of artificial intelligence. The pre-trained model is developed from the large language model in the field of NLP. After fine-tuning, the large language model can be widely used in downstream tasks. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies. Machine learning is a multi-disciplinary cross-disciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all fields of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning by teaching. The pre-trained model is the latest development of deep learning, which integrates the above technologies. The character sequence processing model mentioned above is a combination of machine learning technology and natural language processing technology.

[0187] (ii) Position filtering. In order to perform position filtering, the character positions of each character in the second compressed character sequence in the corresponding second character sequence can also be recorded in the corresponding record list, and then the target record list records the second character positions of each character in the corresponding second compressed character sequence in the corresponding second character sequence. Exemplarily, a certain record list records the second compressed character sequence as "abb", and the record list also records the second character position of the character a in the corresponding second character sequence (such as the first character position), and the second character position of the character b in the corresponding second character sequence; since the compressed character sequence includes characters b at different positions, which are selected from different position intervals in the corresponding second character sequence, the characters b at different positions all have corresponding second character positions, which are the fourth character position and the eighth character position, respectively. Based on this, the specific process of position filtering includes the following contents (I)-(III).

[0188] (I) If there is at least one identical reference character between the second compressed character sequence recorded in the target record list and the first compressed character sequence, the first character position of each reference character in the first character sequence is obtained.

[0189] For ease of explanation, the target record list includes a record list as an example. The reference character refers to the character at the same position between the second compressed character sequence and the first compressed character sequence recorded in the target record list. Exemplarily, the first compressed character sequence is "aba" and the second compressed character sequence is "abb", then at least one reference character includes the character a at the first character position and the character b at the second character position. Since the characters in the compressed character sequence are extracted from the original character sequence, the reference character is derived from the original character sequence and also has a corresponding character position in the original character sequence. The computer device can obtain the character position of each reference character in the first character sequence, which is referred to as the first character position in this application for ease of description. Exemplarily, the character a obtained in the above example is the second character in the first character sequence, so the character position of the character in the first character sequence is the second character position. The character a is the sixth character in the first character sequence, so the character position of the character in the first character sequence is the sixth character position.

[0190] (II) performing difference calculations based on the first character position and the second character position of each reference character recorded in the target record list to obtain a position difference of each reference character.

[0191] When the computer device performs a difference calculation based on the character position of the reference character, the first character position and the second character position can be subtracted, and the absolute value of the difference obtained by the subtraction can be used as the position difference of the reference character, and the position difference can represent the offset of the same character in the second character sequence relative to the first character sequence. Exemplarily, the reference character is character a, and the first character position of character a is the first character position, and the second character position of character a is the third character position, then it can be determined that the position difference of the character a in the two character sequences is 2, that is, the same character is offset by 2 character positions in the second character sequence compared to the first character sequence. For each reference character, the corresponding position difference can be calculated in the above manner, and then the position difference of each reference character can be compared with the second difference threshold to determine whether the position offset of the same character in different character sequences is too large.

[0192] (III) If the position difference of the corresponding reference characters is greater than the second difference threshold, the second compressed character sequence where the corresponding reference characters are recorded in the target record list is deleted to obtain a filtered target record list.

[0193] If the position difference of a reference character is greater than the second difference threshold, it means that the position offset of the reference character in the first character sequence and the second character sequence is large, and the first character sequence and the second character sequence are not effectively aligned based on the reference character, then the similarity between the second character sequence and the first character sequence is low, so the second compressed character sequence recorded in the target record list can be deleted, and the obtained filtered target record list does not include the originally recorded second compressed character sequence.

[0194] It should be noted that if the position difference of the corresponding reference characters is less than or equal to the second difference threshold, it means that the position offset of the reference character in the first character sequence and the second character sequence is within the allowable range, and the first character sequence and the second character sequence can be effectively aligned based on the reference character. Then the second compressed character sequence where the corresponding reference character recorded in the target record list is located can be retained, so that the filtered target record list remains unchanged from the original target record list.

[0195] It can be seen that in the above method, by attaching the position of the original character sequence to each character in the second compressed character sequence recorded in the record list, if the first compressed character sequence recorded in the record list and the first compressed character sequence queried contain the same characters, their positions in their respective original character sequences can be compared. If the position difference is greater than the threshold, it means that it is not a feasible alignment, that is, the two identical characters are actually characters at different positions in the original character sequence, so the second compressed character sequence recorded in the record list can be deleted. The above method can reduce the number of candidates with overlapping characters but not appropriate in the compressed character sequence, thereby improving query efficiency and search accuracy.

[0196] It is understandable that if the filtering process performed on the target record list includes length filtering and position filtering, then the length filtering and position filtering of the target record list can be performed successively or concurrently, and the present application does not limit the execution order of these two types of filtering processes. Exemplarily, the target record list can be first subjected to length filtering to obtain a target record list after length filtering, and then position filtering can be performed based on the target record list after length filtering to finally obtain a filtered target record list.

[0197] Step 5.3: Obtain a second character sequence matching the first character sequence according to the filtered target record list.

[0198] The filtered target record list includes at least one filtered record list. A filtered record list (e.g., list1) in the at least one filtered record list is used as an example for explanation. If the corresponding filtered record list (e.g., list1) includes a second compressed character sequence, in one embodiment, a second character sequence corresponding to the second compressed character sequence included in the filtered record list can be determined as a second character sequence matching the first character sequence. Since the record list of the query tree link only includes one second compressed character sequence, after the filtering process, the filtered record list may not include the second compressed character sequence previously recorded, and the second character sequence corresponding to the filtered second compressed character sequence does not match the first character sequence. Therefore, the number of second character sequences matching the first character sequence obtained according to the filtered target record list is less than or equal to the number of filtered record lists. Exemplarily, the number of filtered target record lists includes three filtered record lists, but one of the filtered record lists does not include the second compressed character sequence, so that two second compressed character sequences can be obtained based on the three filtered record lists, and the second character sequences corresponding to the two second compressed character sequences match the first character sequence.

[0199] In another embodiment, the above filtering process can be used to screen out a set of candidate character sequences in the character sequence similarity search. Specifically, the second character sequence corresponding to the second compressed character sequence recorded in the filtered target record list can be determined as the candidate character sequence. Further, in order to verify whether the candidate character sequence truly matches the first character sequence, the similarity of the first character sequence and the second character sequence can also be calculated. For example, a threshold-based distance similarity calculation can be used, that is, the edit distance between the first character sequence and the candidate character sequence is calculated. When the edit distance is less than the distance threshold, the candidate character sequence can be determined as the second character sequence that matches the first character sequence.

[0200] For the process of searching for a character sequence based on a query tree as shown in the above steps 5.1 to 5.3, the algorithm logic shown in the following Algorithm 3 can be provided, which shows the recursive search process on the Trie tree.

[0201]

[0202]

[0203] For ease of description, the serial number (such as 1, 2, 3, etc.) before each line in Algorithm 2 indicates the line number. According to the above content, we start from the root node of the Trie tree and traverse all the child nodes of the root node. Check whether the mark on the child node is not greater than α. If the mark If the value exceeds α, the successor of the child node can be pruned (lines 6-7). If the mark is within the range of α, check the character represented by the node. If the character is equal to the current character of q′, recursively process the child nodes of the node (line 10), otherwise recursively process the child nodes with the mark increased by 1 (line 12). In each iteration, if the input node is a leaf node, after length filtering and position filtering, the linked record list is merged into the result It can be seen that the above-mentioned Trie tree-based search is a tree-structured character sequence matching algorithm, which can realize fast character sequence search and prefix search, specifically by dividing each compressed character sequence into individual character nodes and storing them in a tree structure to realize fast search.

[0204] The following is a cost analysis of the process of searching for a character sequence based on a query tree. The space cost of a Trie tree index is smaller than O(LN) because common prefixes reduce storage. However, the Trie-based index needs to express various relationships between characters, so its implementation is more complicated, which incurs additional space costs. Assuming that the average number of branches of a node is σ (σ ≤ |∑|, where ∑ is the alphabet of all strings), the time cost of searching a Trie is at least O(σ α), since all nodes within depth α need to be traversed before any pruning. In the worst case, the time cost is O(σ L ). To this end, in this application, character sequence search can also be performed based on a multi-level inverted index. Since the multi-level inverted index (minIndex) is a simple and lightweight index, it can be searched with low space consumption, thereby reducing space costs. Experiments on data sets in various practical scenarios have proved that minIndex has excellent performance, high efficiency and low storage space consumption.

[0205] In another embodiment, the query index is a multi-level inverted index, and the character sequence search rule includes a search order, and optionally, the search order includes a position order and an arrangement order of characters in a character set. Based on this, when the computer device executes the above S406, it can specifically execute the following steps 6.1 to 6.4.

[0206] Step 6.1: In accordance with the search order indicated by the query index and based on the character position information of the first compressed character sequence, at least one second compressed character sequence is selected from the second compressed character sequences corresponding to each second character sequence in the character sequence set.

[0207] The character position information includes the character position of each character in the first compressed character sequence. According to the search order and the characters at each position in the first compressed character sequence, a second compressed character sequence with the same characters at the same character position can be searched based on a multi-level inverted index, thereby obtaining at least one second compressed character sequence. In other words, in each of the screened second compressed character sequences, at least one character position has the same character as the character at the same character position in the first compressed character sequence.

[0208] In one implementation, based on the aforementioned multi-level inverted index construction process, it can be known that each level of the inverted index includes each character in the character set and a record list of each character; a level one inverted index corresponds to a character position, for example, if the length of the first compressed character sequence is L, then the query index can be an L-level inverted index. A character record list is used to record at least one second compressed character sequence where the corresponding character is at the character position corresponding to the corresponding level of the inverted index; illustratively, as described above Figure 6bIn the multi-level inverted index shown, the inverted index corresponding to the first character position includes all the characters in the character set and a record list of each character (part of the record list is not shown). For example, the record list corresponding to character a can be used to record at least one second compressed character sequence of character a at the first character position (for example, "aba", "acc", "abc"). The character set is determined based on the character sequence set. For the character set, please refer to the relevant introduction when constructing the query index, which will not be repeated here. Based on this, when the computer device executes the above step 6.1, it can specifically follow the steps 6.1.1 to 6.1.3 as shown below.

[0209] Step 6.1.1 searches the record list included in the multi-level inverted index according to the search order indicated by the query index and the character positions of the characters in the first compressed character sequence to obtain a target record list.

[0210] As can be seen from the above, the record list included in the multi-level inverted index is specifically: a record list corresponding to each character at each character position. In the specific search process, the record list can be searched in each level of the inverted index according to the characters at each position in the first compressed character. Figure 6b Based on the structure of the multi-level inverted index shown in Figure 7b The selection of the target record list is illustrated as shown. Assuming that the first compressed character sequence is "aba", then according to the search order indicated by the multi-level inverted index, the record list of character a can be searched in the inverted index corresponding to the first character position, the record list of character b can be searched in the inverted index corresponding to the second character position, and the record list of character a can be searched in the inverted index corresponding to the third character position. These record lists can all be determined as target record lists. Since the first compressed character sequence includes multiple characters, the search is performed according to each character and character position, and the obtained target record list includes multiple record lists, each record list corresponding to a character at a character position in the first compressed character sequence.

[0211] In a specific implementation, the target record list is obtained according to the following steps ① to ③.

[0212] Step ① traverses the characters in the first compressed character sequence, and determines the traversed character as the current character.

[0213] Since the characters in the first compressed character sequence have a relative position relationship and are arranged in sequence, the computer device can traverse the characters in the first compressed character sequence in the order of character positions, and the traversed character is the current character. For example, if the first compressed character sequence is "aba", the character a at the first character position traversed is the current character.

[0214] Step ② scans the j-th level inverted index in the query index in the search order according to the current character and the character position j of the current character to obtain a record list of the current character.

[0215] Specifically, the character position of the current character is character position j in the first compressed character sequence, and based on the length of the first compressed character sequence being L, j∈[1,L]. For example, the character position j=1 of the character a traversed above. According to the character position of the current character, the inverted index corresponding to the character position can be determined from the multi-level inverted index, that is, the j-th level inverted index. Each level of the inverted index includes each character in the character set and a list of records corresponding to each character. Therefore, each character in the inverted index of that level can be scanned in the search order. When the current character is scanned, the record list of the current character can be obtained. The character at character position j in the second compressed character sequence recorded in the record list of the current character is the current character. For example: the current character is character a at the first character position, that is, q′[1]=a, then the first level inverted index can be scanned. Get the list of records of character a in the inverted index at this level And the first character in each second compressed character sequence recorded in the record list is character a.

[0216] Optionally, the character at each position in the first compressed character sequence can be called an axis for the corresponding inverted index. The axis refers to a reference point or element used to help quickly split a data set or accelerate the search process. The selection and use of the axis can effectively reduce the amount of data that needs to be compared, thereby accelerating the search process. Specifically, when the axis is scanned during the search process, the scan can be stopped, thereby reducing the number of scans.

[0217] Step ③: When all characters in the first compressed character sequence are traversed, a record list of each character in the first compressed character sequence is obtained, and the record list of each character is determined as a target record list.

[0218] After all characters in the first compressed character sequence are traversed, all levels of inverted indexes are scanned, and record lists of all characters in the first compressed character sequence can be obtained, and these record lists can all be used as target record lists. Based on the length of the first compressed character sequence, the number of record lists included in the target record list is the same as the length. For example, if the length of the first compressed character sequence is 3, then 3 record lists can be obtained as target record lists (such as Figure 7b shown).

[0219] The method of obtaining the target record list shown in the above steps ① to ③ can determine the stop point of character scanning in each level of the inverted index by traversing the characters in the first character sequence. When the same character as the traversed character is scanned in the inverted index, the record list of the character can be directly obtained and the scanning can be stopped, thereby improving the speed of obtaining the record list. Since the second compressed character sequence recorded in the record list obtained in this way has overlapping characters with the first character sequence (that is, the characters in the same character position are the same), the second compressed character sequence obtained in the above manner is related to the first compressed character sequence, so that the second character sequence matching the first character sequence can be more accurately determined based on these second compressed character sequences.

[0220] Step 6.1.2 filters the target record list to obtain a filtered target record list.

[0221] Based on the length L of the first compressed character sequence (L is an integer greater than 1), the target record list may include L record lists. Each record list may be filtered to obtain a filtered record list, so the filtered target record list includes multiple filtered record lists. The filtering process here may include at least one of the following: length filtering and position filtering. For the specific implementation of length filtering and position filtering, please refer to the above-mentioned relevant introduction, and this application will not be repeated here.

[0222] It should be noted that since each record list included in the target record list selected from the multi-level inverted index may record more than one second compressed character sequence, each second compressed character sequence needs to be processed during the filtering process. For example, when performing length filtering, it is necessary to calculate the length difference between the second character sequence corresponding to each second compressed character sequence in the record list and the first character sequence. When performing position filtering, for each second compressed character sequence in the record list, it is also necessary to first detect whether the same characters are in the same character position as the first compressed character sequence, and then, when the same characters exist in the same character position, detect the position difference between the positions of the characters in their respective original character sequences. Through filtering processing, it is possible to further filter out second compressed character sequences that match the first compressed character sequence from the target record list, thereby improving search accuracy.

[0223] Step 6.1.3 determines the second compressed character sequence included in the filtered target record list as the screened second compressed character sequence.

[0224] After the above filtering, some second compressed character sequences originally recorded in some record lists may be filtered out. Therefore, the total number of second compressed character sequences included in the filtered target record list may be reduced. In addition, since the second compressed character sequences included in different record lists may be repeated, when determining the second compressed character sequence, different second compressed character sequences in the filtered target record list may be counted, and then these second compressed character sequences may be determined as the filtered second compressed character sequences. The second compressed character sequences included in the filtered target record list are all character sequences with a high degree of matching with the first compressed character sequence.

[0225] The process shown in the above steps 6.1.1 to 6.1.3 can specifically filter the record list in the multi-level inverted index according to the characters at each position in the first compressed character sequence, so as to quickly obtain a second compressed character sequence containing the same characters at the same position. In addition, further filtering the filtered record list through length filtering and / or position filtering can improve the search accuracy and reduce the number of character sequences involved in subsequent similarity calculations to improve the search efficiency.

[0226] Step 6.2 determines the number of different characters between at least one of the selected second compressed character sequences and the first compressed character sequence.

[0227] Although the first character sequence and the second character sequence use the same compression method to obtain the corresponding compressed character sequence, the compressed character sequences may have different characters in the same position due to the difference between the original character sequences. Therefore, filtering the compressed character sequences in the target record list may not be sufficient to determine that the filtered second character sequence matches the first character sequence. In order to further improve the accuracy of the search results, the computer device can determine the number of difference characters between each filtered second compressed character sequence and the first compressed character sequence, so that each filtered second compressed character sequence has a difference character number with the first compressed character, and the difference character number refers to the total number of different characters in the same character position. When the difference character number is 0, it means that the second compressed character sequence is exactly the same as the first compressed character sequence and there is no difference. When the difference character number is greater than 0, it means that there is at least one position where the characters are different between the second compressed character sequence and the first compressed character sequence. Based on the length L of the first compressed character sequence, the difference character number ≤ L.

[0228] In a specific implementation, when determining the number of different characters, it can be implemented according to the process shown in ①-② below.

[0229] ① Count the number of occurrences of each second compressed character sequence recorded in the filtered target record list to obtain the occurrence frequency of each second compressed character sequence.

[0230] If the second compressed character sequence and the first compressed character sequence have multiple identical characters in the same character position, the second compressed character sequence is recorded in multiple record lists after filtering, and a second compressed character sequence may have appeared in multiple record lists of the multi-level inverted index. Therefore, for any second compressed character sequence, it may appear once or multiple times in the filtered target record list. The computer device can count the number of occurrences of each second compressed character sequence, and can determine the number of occurrences as the occurrence frequency f of the corresponding second compressed character sequence, then the number of occurrences of a second compressed character sequence is f∈[1, L]. Exemplarily, the filtered target record list includes 3 filtered record lists, and the second compressed character sequence aba is recorded in these record lists, so it can be determined that the occurrence frequency of the second compressed character sequence aba is 3, and the occurrence frequencies of other second compressed character sequences are all 1.

[0231] In a feasible implementation, a statistical tool can be called to count the number of occurrences of the second compressed character sequence in the filtered target record list. For example, a statistical tool Map can be used to specifically store and manage different second compressed character sequences in the record list and their occurrence frequencies, thereby providing support for optimization of the search algorithm and data analysis. In an implementation, for each second compressed character sequence in the filtered target record list, the second compressed character sequence s′ can be detected. i Is it recorded in Map? <s′ i , f>, if not recorded in Map <s′ i , f>, then it can be inserted into the Map, which contains at least one second compressed character sequence and the frequency of occurrence of each second compressed character sequence; if it has been recorded in the Map <s′ i , f>, then the occurrence frequency of the second compressed character sequence recorded in the Map can be increased by 1. After processing the second compressed character sequence in the filtered target record list, a Map containing the intersection of the corresponding second compressed character sequence and the first compressed character sequence can be obtained. In another implementation, the second character sequence s can also be detected. i Is it recorded in Map? i ​, f>, and then add a new second character sequence to the Map or update the frequency of occurrence of an existing compressed character sequence according to the same logic. It can be seen that the role of Map is to store records (such as the second compressed character sequence / or the second character sequence) and the frequency of occurrence of the second compressed character sequence. The frequency of occurrence here refers to the number of times or frequency that a specific record (such as the second compressed character sequence) appears in the filtered list of multiple records. Based on the above example, the role of Map can be summarized as follows:

[0232] 1. Store records and their frequencies: The key in the Map can be the second compressed character sequence s′ i or the original second character sequence s i , the value is the frequency f of the second compressed character sequence in the multi-level inverted index. This data structure makes it possible to quickly query the frequency of occurrence of any specific compressed character sequence.

[0233] 2. Improve retrieval efficiency: By monitoring the frequency of occurrence of different records (i.e., the second compressed character sequence), performance can be optimized when performing searches or queries. For example, in some algorithms, records with higher frequencies may be given priority because they are more likely to be relevant to the query results.

[0234] 3. Data analysis and optimization: Knowing the frequency of different records is also useful for data analysis. It can help identify common patterns, hot spots, or data items that need special treatment.

[0235] ② Calculate the difference between the occurrence frequency of each second compressed character sequence and the length of the first compressed character sequence to obtain the number of different characters between each second compressed character sequence and the first compressed character sequence.

[0236] Specifically, each second compressed character sequence recorded in the filtered target record list has an occurrence frequency. Based on the structural characteristics of the multi-level inverted index, for the second compressed character sequence, the corresponding occurrence frequency can represent the number of identical characters in the same position between it and the first compressed character sequence. Exemplarily, the length of the first compressed character sequence is 3, and the occurrence frequency of the second compressed character sequence is 3, which means that the second compressed character sequence is exactly the same as the first compressed character sequence, that is, the characters in each position are the same, and the number of difference characters is 0 at this time. Based on the occurrence frequency f of the second compressed character sequence ≤ the length L of the first compressed character sequence, therefore, when performing the difference calculation, the occurrence frequency of the second compressed character sequence can be directly subtracted from the length of the first compressed character sequence. Specifically, the length of the first compressed character sequence can be subtracted from the occurrence frequency of the second compressed character sequence, and the difference can be directly determined as the number of difference characters. Further, based on the number of difference characters, it can be determined whether the corresponding second character sequence can be determined as a candidate character sequence.

[0237] In the methods shown in ①-② above, based on the search characteristics of the multi-level inverted index, by counting the frequency of occurrence of the second compressed character sequence in the filtered record list, the number of difference characters between the second compressed character sequence and the first compressed character sequence can be accurately obtained, which is conducive to further screening of character sequences based on the number of difference characters.

[0238] Step 6.3: If the number of different characters is less than the preset number threshold, the second character sequence corresponding to the corresponding second compressed character sequence is determined as a candidate character sequence.

[0239] Specifically, the preset number threshold is α, which represents the maximum number of different characters between the first compressed character sequence and the second compressed character sequence. The number of different characters between a second compressed character sequence and the first character sequence can be expressed as Lf, where L is the length of the first compressed character sequence and f is the frequency of occurrence of the second compressed character sequence. If Lf≤α, that is, the frequency of occurrence f satisfies the following condition: f≥L-α, the second character sequence corresponding to the second compressed character sequence involved in the frequency of occurrence f can be determined as a candidate character sequence. In other words, for a second compressed character sequence, if its frequency of occurrence meets the above conditions, then it can be determined that the difference between the second compressed character sequence and the first compressed character sequence is within the allowable range, so that the second character sequence can be determined as a candidate character sequence. On the contrary, if the number of different characters is less than the preset number threshold, that is, L-f<α, the second character sequence corresponding to the second compressed character sequence can be determined as a non-candidate character sequence, and these non-candidate character sequences are second character sequences that do not match the first character sequence.

[0240] It is understandable that, for the multiple second compressed character sequences included in the filtered target record list, it is possible to determine whether the corresponding second character sequence is a candidate character sequence in the above manner. After the above processing, at least one candidate character sequence can be obtained, and the number of the candidate character sequences is less than or equal to the number of the screened second compressed character sequences.

[0241] In one embodiment, the size setting of the preset quantity threshold α will affect the discrimination of the difference between the second compressed character sequence and the first compressed character sequence. The smaller the preset quantity threshold is set, the smaller the difference between the second compressed character sequence and the first compressed character sequence is required, so that a more similar second compressed character sequence can be matched to the first compressed character sequence to obtain a more accurate search result. Therefore, the preset quantity threshold α can be adjusted at any time to achieve a higher accuracy (>0.99). The preset quantity threshold α is independent, and the accuracy depends on the threshold factor t=k / n and the parameter l. However, the preset quantity threshold α may not be appropriately selected, that is, the setting of the parameter α does not meet the expected accuracy standard. The accuracy here refers to the degree to which the compressed character sequence correctly reflects the characteristics of the original character sequence. Whether a parameter α is appropriately selected can be determined based on the following conditions:

[0242] 1. The k and n in the threshold factor represent some key character sequence features. If the selected alpha cannot make the threshold factor t reach a reasonable range, that is, the satisfaction of the threshold factor t = k / n does not meet the requirements, this may mean that alpha is not appropriately selected.

[0243] 2. Accuracy evaluation: The accuracy of a specific α value can usually be evaluated through experiments or simulations. If the accuracy obtained is significantly lower than expected (for example, lower than 0.99), it means that α may not be appropriately selected.

[0244] 3. Actual effect: In actual applications, if it is found that the similarity between the generated compressed character sequence and the original character sequence is low, or the search effect is not good, this may be a sign that the α value is set improperly.

[0245] 4. Comparison of the number of errors to the compressed character sequence: If the number of errors due to an incorrect α setting exceeds the number of characters in the compressed character sequence, this is usually a clear sign that α needs to be readjusted.

[0246] In general, the appropriate selection of the α parameter is a process of trial and adjustment, which needs to be determined based on the threshold factor, accuracy assessment, and actual application effect. If the accuracy in the actual application does not meet expectations, or the number of errors is large, the value of α may need to be readjusted.

[0247] Step 6.4: Obtain a second character sequence matching the first character sequence based on at least one candidate character sequence.

[0248] In one implementation, the computer device may directly determine each candidate character sequence as a second character sequence that matches the first character sequence. In another implementation, in order to ensure the accuracy of the search results, similarity calculation may be performed to verify whether the candidate character sequence truly matches the first character sequence. Specifically, any candidate character sequence in at least one candidate character sequence is represented as a candidate character sequence s i The following is an example of calculating the similarity between a candidate character sequence and a first character sequence. The computer device may execute the following contents (1)-(2).

[0249] (1) For the candidate character sequence s i Calculate the similarity with the first character sequence to obtain the candidate character sequence s i The similarity between the first character sequence and the

[0250] In a specific implementation, in the process of searching for a character sequence, the similarity measurement methods between the first character sequence and the candidate character sequence include but are not limited to: cosine similarity, Jaccard similarity (a method for comparing similarities and differences between finite sample sets, the larger the Jaccard value, the higher the similarity), overlapping similarity, edit distance, etc. Among these similarity metrics, edit distance has an important advantage, which preserves the order of characters and captures a better match between two character sequences, and is often used in business scenarios such as spelling check, plagiarism check, and speech recognition in applications. Therefore, in the present application, the computer device can calculate the candidate character sequence s i The edit distance between the candidate character sequence s and the first character sequence specifically includes the following steps: first, i The minimum number of editing operations required to convert to the first character sequence is counted to obtain the first character sequence and the candidate character sequence s i The edit distance between the first character sequence and the candidate character sequence s i The edit distance between the first character sequence and the candidate character sequence s i The similarity between .

[0251] For a given first character sequence q and candidate character sequence s i , the edit distance between the two can be expressed as ED(s i ,q), the edit distance refers to the candidate character sequence s iThe minimum number of editing operations required to convert to the first character sequence, where the editing operation includes: replacing, inserting or deleting a single character, so the first character sequence and the candidate character sequence s i The similarity between them can be the edit distance. For each candidate character sequence, the edit distance between it and the first character sequence can be calculated, so as to determine whether the corresponding candidate character sequence actually matches the first character sequence based on the comparison between the edit distance and the distance threshold. It should be understood that the specific implementation of other similarity measurement methods will not be described in detail here.

[0252] (2) The candidate character sequence s i The corresponding similarity is compared with the similarity threshold; if the similarity is greater than or equal to the similarity threshold, the candidate character sequence s i A second character sequence is determined to match the first character sequence.

[0253] In a specific implementation, when the similarity is the edit distance, the similarity threshold is determined according to the distance threshold k, k∈[0,1]. The similarity threshold can be 1-k, and then when the edit distance is greater than or equal to the similarity threshold, the candidate character sequence s i The second character sequence is determined to match the first character sequence. If the similarity threshold is directly taken as the distance threshold k, then the similarity greater than or equal to the similarity threshold means that the edit distance is less than or equal to the distance threshold. That is, when the edit distance is less than or equal to k, the candidate character sequence s i A second character sequence is determined to match the first character sequence.

[0254] In the above method, the candidate character sequence s can be calculated i The edit distance between the candidate character sequence and the first character sequence, if the edit distance is less than the distance threshold, then it means that the candidate character sequence meets the search requirements and is a character sequence with a high similarity to the first character sequence, so that the character sequence can be determined as the second character sequence matching the first character sequence. For each candidate character sequence, the similarity calculation can be performed according to the contents shown in (1)-(2) above, so as to obtain the similarity between each candidate character sequence and the first character sequence, and each similarity can be compared with the similarity threshold to determine whether the corresponding candidate character sequence can be used as the second character sequence matching the first character sequence.

[0255] Thus, for a given character sequence set, a distance threshold k and a query character sequence q, a set of all second character sequences whose edit distance to q is not greater than k can be found. Exemplarily, the first character sequence q = "above", the length is 5, k = 1, and each candidate character sequence includes the content shown in Table 2 below. The threshold similarity search can return "abode", because the edit distance between "abode" and the query "above" is 1≤k.

[0256] Table 2

[0257]

[0258]

[0259] Corresponding to the specific content of the embodiment of the present application, a set of candidate character sequences, a first character sequence q and a distance threshold k can be given, and an edit distance similarity search based on the threshold can obtain all sequences that satisfy ED(s i ,q)≤k second character sequence Collection Compared to directly calculating the edit distance for all character sequences in a character sequence set, the present application first filters the character sequence set by querying the index and then performs the edit distance calculation, which can greatly reduce the computing resources required to calculate the edit distance and can achieve fast search with low space consumption.

[0260] Specifically, for the character sequence similarity search problem with a given edit distance threshold, the problem requires that the distance between the result and the query character sequence is within the given threshold. It is well known that the time complexity of the edit distance calculation is O(n 2 ), where n is the length of the character sequence. Therefore, a large number of second character sequences need to be verified, the query efficiency is low, and the performance of long character sequences is even worse. The present application compresses each character sequence, organizes the compressed character sequence through a query index, and provides a search logic for the compressed character sequence through the query index, thereby quickly searching for candidate character sequences. This is because the length of the compressed character sequence is shorter, and the index structure provided by the query index is simple and efficient, so that some candidate character sequences that already have a high degree of similarity can be quickly screened out, so that when calculating the edit distance, there is no need to calculate all the second character sequences in the character sequence set, but only part of the second character sequences.

[0261] For the character sequence search process shown in the above steps 6.1 to 6.3, an exemplary algorithm logic shown in the following Algorithm 4 can be provided.

[0262]

[0263]

[0264] The above algorithm is an algorithm based on threshold similarity search on a multi-level inverted index. Given a character sequence set A query character sequence q and a similarity threshold k, which is a threshold used to measure the edit distance. The algorithm returns all the characters that satisfy ED(s i ,q)≤k second character sequence Different thresholds can be used in this query method, which has different query time and accuracy. In the specific implementation process, the compression processing of all input character sequences (including the first character sequence and the second character sequence) can be implemented by using the MHcompact method introduced in the above algorithm 1. Then, the construction logic shown in the above algorithm 2 can be used to construct a multi-level inverted index. Afterwards, for each character q′[i] in the first compressed sequence q′, the multi-level inverted index can be scanned. The record list is filtered by length (LengthFilter) and position (PositionFilter) for the second compressed character sequence in the record list, and then the frequency of occurrence f of the remaining second compressed character sequence in the L record lists after filtering is counted and recorded in the Map. Based on the frequency of occurrence recorded in the Map, it is determined whether the second compressed character sequence matches the first compressed character sequence (i.e., Lf≤α). If they match, the edit distance between the original character sequences corresponding to the two compressed character sequences can be further calculated, i.e., ED(s i , q), and when the edit distance is less than or equal to a given threshold, the second character sequence is added to the search results In the search result, the search result is used to record a second character sequence that matches the first character sequence.

[0265] The following is a cost analysis of the character sequence search process shown in step 6.1-step 6.4. This character sequence search is a search based on a multi-level inverted index, which does not use additional structures compared to the Trie tree-based index. Although the space cost is still O(LN), the cost in practice is smaller than that of the Trie tree-based index. Specifically, the time cost of the search algorithm based on the multi-level inverted index is O(LN / |∑|), where N is the total number of second character sequences and |∑| is the number of different characters in the character set.

[0266] In a multi-level inverted index, each record list contains all occurrences of a particular sketch character. If the sketch character is 'a', then the corresponding record list will contain all compressed character sequences that contain 'a' as the sketch character. Considering that each sketch character may theoretically be evenly distributed in all character sequences, the frequency of occurrence of each character is roughly the same. Therefore, the average number of times a given sketch character appears in all N character sequences is N / |∑|. The search algorithm needs to scan L different record lists, each record list corresponds to a sketch character, and the average length of each record list is N / |∑|. Therefore, the total time cost of scanning all L record lists is roughly O(L*(N / |∑|)).

[0267] It should be noted that this time cost estimate is based on the assumption that each sketch character is evenly distributed across all strings. In practice this distribution may not be completely uniform, but this expression provides a rough estimate of the time complexity. In this way, it can be seen that multi-level inverted indexes may be more efficient than Trie-based indexes when processing a large number of strings. The time cost of the search method does not include the time cost of the verification phase O(LN / |∑|).

[0268] The above steps 6.1 to 6.4 can filter out at least one second compressed character sequence from the multi-level inverted index by searching the record list through the search order indicated by the query index. Furthermore, by comparing the difference characters between the first compressed character sequence and the second compressed character sequence, specifically by counting the number of different characters at the same position, the second compressed character sequence matching the first compressed character sequence can be quickly filtered out, thereby filtering out candidate character sequences from the second character sequence in the character sequence set based on the correspondence between these second compressed character sequences and the second character sequence, and obtaining a second character sequence matching the first character sequence based on the candidate character sequence. In this way, the candidate character sequence can be quickly filtered out through the compressed character sequence, and the accuracy of the character sequence can be further guaranteed based on the similarity calculation between the original character sequences, which can reduce the time cost of the similarity calculation between the original character sequences and improve the search efficiency.

[0269] Based on the above Figure 4The introduction of the illustrated embodiment, in this application, the compressed character sequence can be combined with two streamlined query indexes to obtain two search methods respectively. Specifically, if the above algorithm 1 and algorithm 3 are combined, a character overview representation + Trie tree search method can be provided. Among them, the overview representation is used to implicitly encode the string, and align each character sequence, and use the Trie tree (a search tree) to organize and index the compressed string, which is particularly suitable for quickly retrieving a large amount of string data. If the above algorithm 1, algorithm 2 and algorithm 4 are combined, a character overview representation + multi-level inverted index (minIndex) search method can be provided. In the character sequence similarity search process, the overview representation is used to implicitly encode the string, and the multi-level inverted index is used to efficiently organize and search these strings, and the multi-level inverted index provides a low space consumption and high efficiency way to process string search. This method can be used to quickly find candidate strings similar to the query string. Specifically, minhash can be used to capture key characters and construct an overview representation of the character sequence, and then a multi-level inverted index is used to search under low space occupancy. Low space consumption is achieved by searching the overview through a multi-level inverted index. In the above search method, due to the use of overview representation (i.e. compressed character sequence), the space cost of the index is reduced to O(LN), where N is the cardinality of the data set and L is the length of the overview representation. Since the space cost is independent of the string length, this method has better performance on long strings. This reduces both space cost and time cost.

[0270] Furthermore, in order to improve the query efficiency, the present application also uses the learning index technology to replace the length filter, and quickly locates the candidate position in the inverted list through the learning index technology, and the experimental results show that the space cost and query efficiency of the proposed method are better than the existing methods. Specifically: 1. Use of learning index: The learning index technology is applied to replace the length filter to improve the query efficiency. This means that when processing string similarity search, the traditional length filtering method is replaced by a method based on machine learning. 2. Improve query efficiency: Through the learning index technology, the candidate position in the inverted list can be located faster, so that when performing string search, the strings that do not meet the conditions can be quickly filtered out, thereby improving the overall query efficiency. 3. Comparison between learning index and traditional index: Learning index is an index method based on machine learning. Compared with the traditional index structure, it can handle complex query conditions more intelligently and efficiently, especially in a big data environment. The application of learning index technology is mainly reflected in optimizing the query process through machine learning methods, especially in replacing the traditional length filter, so as to achieve the purpose of improving query efficiency.

[0271] In another embodiment, the number of first compressed character sequences includes M, and the number of second compressed character sequences corresponding to each second character sequence includes M; M is an integer greater than 1; the M first compressed character sequences and the M second compressed character sequences are obtained by using different compression functions in the compression process using the target compression method. That is, the first character sequence and each second character sequence in the character sequence set can be compressed using the same compression method but different compression parameters to obtain corresponding multiple compressed character sequences. Among them, the compression parameter can be a compression function or parameter set for selecting target characters. For example, different minhash functions can be used to generate, such as after the first character sequence q is compressed by different minhash functions, 2 first compressed character sequences can be obtained; after the second character sequence s is compressed by different minhash functions, 2 second compressed character sequences can be obtained. By generating multiple compressed character sequences for each character sequence, the search accuracy can be improved or potential hash conflicts can be dealt with. When generating multiple compressed character sequences, each compressed character sequence is generated using a different minhash function or a different parameter set, which can capture the characteristics of the character sequence from different angles, thereby improving the comprehensiveness and accuracy of the search.

[0272] It should be noted that after the first character sequence and the second character sequence are compressed using the same compression parameter, an association relationship can be set between the compressed character sequences obtained. In this way, in the character sequence search process in the above-mentioned embodiment, based on the association relationship, a character sequence search can be performed based on each second compressed character sequence that uses the same compression parameter as the first compressed character sequence. Each second compressed character sequence obtained using the same compression parameter can be used to construct a query index. For example, the first compressed character sequence q′ obtained after the first character sequence q is compressed by the minhash function F1 1 , the first compressed character sequence q′ obtained after compression by minhash function F2 2 ; Each second character sequence is compressed by the minhash function F1 to obtain a second compressed character sequence s′ 1 , after compression by minhash function F2, the second compressed character sequence s′ is obtained 2 ; Further, based on the compression process, a second compressed character sequence s′ is obtained 1 A query index can be constructed, and according to the instruction of the query index, based on the first compressed character sequence q′ 1 and each second compressed character sequence s′ 1 Perform character sequence search. After compression processing, obtain the second compressed character sequence s′ 2A query index can be constructed, and according to the instruction of the query index, based on the first compressed character sequence q′ 2 and each second compressed character sequence s′ 2 Perform a character sequence search. In one embodiment, each query index may also be searched in parallel or serially. For example, multiple compressed character sequences corresponding to the first character sequence may be processed in parallel, thereby searching multiple query indexes at the same time to speed up the search. Alternatively, multiple compressed character sequences corresponding to the first character sequence may be processed serially, and each query index may be checked in turn. This application does not limit this.

[0273] It can be seen that generating multiple compressed character sequences for each character sequence will increase the corresponding search range. This is because using multiple compressed character sequences means that multiple different indexes or record lists will be checked during the search, which can increase the probability of finding relevant results when the original character sequence has a certain complexity or diversity.

[0274] Based on the multiple compressed character sequences corresponding to each character sequence, the following can be used: Figure 2 or Figure 4 The process of the embodiment shown is executed to obtain multiple result sets. The computer device may execute the following steps 1 to 3.

[0275] Step 1: Obtain M result sets, and merge the M result sets to obtain a merged result.

[0276] Step 2: Sort the second character sequence included in the merging result according to the similarity between the first character sequence and the second character sequence to obtain a sorting result.

[0277] Step 3: Perform service processing on the second character sequence matching the first character sequence according to the sorting result.

[0278] Each result set is used to record at least one second character sequence that matches the first character sequence; and each result set is obtained based on searching a first compressed character sequence and a second compressed character sequence corresponding to each second character sequence in the character sequence set. To obtain a result set, the first compressed character sequence and the second compressed character sequence used use the same compression parameter under the same compression method. The merged result obtained by merging the result sets includes the second character sequence that matches the first character sequence, but there may be repeated second character sequences in different result sets. Therefore, when merging the result sets, different second character sequences can be selected to form the merged result, so that the merged result includes different second character sequences that match the first character sequence.

[0279] In a specific implementation, the computer device may obtain the similarity between each second character sequence included in the merged result and the first character sequence, and sort the second character sequences in the merged result in descending order, and the obtained sorting result includes a plurality of second character sequences arranged in sequence. Afterwards, business processing may be performed based on the sorting result, and the business processing mentioned here may include but is not limited to at least one of the following: ① displaying the second character sequence matching the first character sequence according to the sorting result; ② outputting business data adapted to the corresponding business scenario according to the sorting result, etc.

[0280] By searching using multiple compressed character sequences corresponding to the first character sequence, multiple result sets can be obtained. These result sets can be merged and ranked according to relevance or other criteria, and then corresponding business processing can be performed based on the ranking results.

[0281] It should be noted that the use of multiple compressed character sequences corresponding to the first character sequence for searching will increase the corresponding computational cost, because generating and processing multiple compressed character sequences will increase the computational cost and time overhead. Therefore, in practical applications, it is necessary to find a balance between the accuracy and efficiency of the search. In addition, the balance between redundancy and accuracy must also be considered. Although the use of multiple compressed character sequences can improve accuracy, it may also introduce redundant data. Therefore, the reasonable selection of minhash functions and parameters can maximize accuracy while controlling redundancy, which is also a factor that needs to be considered.

[0282] In a feasible implementation, searching through multiple compressed character sequences corresponding to each character sequence can be adopted in some special cases, such as when the preset number threshold α for determining the number of difference characters is not appropriately selected. This method can achieve high accuracy by repeating MHCompact using different minhash families. In this way, multiple compressed character sequences are generated for each character sequence, and the multi-level inverted index can be scanned simultaneously without modification. In actual application, this method will result in a larger index size, and a single MHCompact is good enough. In summary, generating multiple compressed character sequences for a character sequence is a strategy to improve search accuracy, but the increased computational cost and complexity must also be considered. Therefore, in actual applications, the use of this method should be determined based on specific needs and resource constraints.

[0283] Next, the character sequence search device provided in the embodiment of the present application is described.

[0284] See also Figure 8 , Figure 8: is a schematic diagram of a character sequence search device provided by an exemplary embodiment of the present application. The above-mentioned character sequence search device may be a computer program (including program code) running in a computer device, for example, the character sequence search device is an application software; the character sequence search device may be used to execute the corresponding steps of the method provided by the embodiment of the present application. Figure 8 As shown, the character sequence search device may include: an acquisition unit 801, a processing unit 802 and an output unit 803.

[0285] An acquisition unit 801 is used to acquire a first character sequence to be searched;

[0286] The processing unit 802 is configured to compress the first character sequence using a target compression method to obtain a first compressed character sequence corresponding to the first character sequence;

[0287] The acquisition unit 801 is further used to acquire a query index, where the query index is constructed based on the second compressed character sequences corresponding to the second character sequences included in the character sequence set; wherein each second character sequence is compressed using a target compression method to obtain a corresponding second compressed character sequence; the query index defines an index structure used when searching for a character sequence, and the query index is used to indicate a character sequence search rule;

[0288] The processing unit 802 is further configured to search the character sequence set for a second character sequence matching the first character sequence based on the first compressed character sequence and at least one second compressed character sequence according to the indication of the query index.

[0289] In one embodiment, the acquisition unit 801 is also used to acquire a character sequence set, which includes multiple second character sequences; the processing unit 802 is also used to compress each second character sequence in the character sequence set using a target compression method to obtain a second compressed character sequence corresponding to each second character sequence; and a query index is constructed based on the second compressed character sequences corresponding to each second character sequence in the character sequence set.

[0290] In one embodiment, when the processing unit 802 compresses the first character sequence in a target compression manner to obtain a first compressed character sequence corresponding to the first character sequence, it is specifically configured to:

[0291] Get an initialization character sequence, the initialization character sequence is used to store L characters, where L is an integer greater than 1;

[0292] According to the character features of each character in the first character sequence, a target character is selected from the first character sequence, and the target character is stored in the initialization character sequence to update the initialization character sequence; wherein the character feature of the target character is a key character feature in the first character sequence;

[0293] The first character sequence is recursively processed according to the target character to obtain a first compressed character sequence corresponding to the first character sequence based on the updated initialization character sequence.

[0294] In one embodiment, when the processing unit 802 selects the target character from the first character sequence according to the character features of each character in the first character sequence, it is specifically configured to:

[0295] Obtaining an interval length parameter, and determining a target position interval of the first character sequence according to the interval length parameter and the length of the first character sequence;

[0296] Using a hash function to perform hash calculation on characters in the first character sequence whose character positions are in the target position interval, to obtain a hash value of at least one character in the target position interval;

[0297] According to the hash value of at least one character in the target position interval, a character with the smallest hash value is selected from the at least one character in the target position interval, and the selected character is determined as the target character.

[0298] In one embodiment, L is determined based on a preset number of times l of recursively processing the first character sequence, where l is a positive integer; when the processing unit 802 recursively processes the first character sequence according to the target character to obtain a first compressed character sequence corresponding to the first character sequence based on the updated initialization character sequence, it is specifically configured to:

[0299] Divide the first character sequence according to the target character to obtain two sub-character sequences;

[0300] Determine each sub-character sequence as a first character sequence, and repeatedly execute the target step, which is: select a target character from the first character sequence according to the character features of each character in the first character sequence, and store the target character in the initialization character sequence to update the initialization character sequence;

[0301] When the number of target characters stored in the updated initialization character sequence reaches the length L of the initialization character sequence, the updated initialization character sequence is determined as the first compressed character sequence corresponding to the first character sequence.

[0302] In one embodiment, when the processing unit 802 constructs the query index based on the second compressed character sequences respectively corresponding to the second character sequences in the character sequence set, it is specifically configured to:

[0303] Get the character set corresponding to the character sequence set, the character set includes multiple characters;

[0304] Traversing the characters included in the character set, determining the traversed character as the current character, and selecting at least one second compressed character sequence including the current character from the second compressed character sequences corresponding to each second character sequence in the character sequence set according to the current character;

[0305] Obtaining the reference character position of the current character in each selected second compressed character sequence, and adding each selected second compressed character sequence to the record list corresponding to the current character at the corresponding reference character position;

[0306] After all characters in the character set have been traversed, a list of records corresponding to each character in the character set at each character position is obtained;

[0307] Each character in the character set and a record list corresponding to each character at the same character position are determined as an inverted index of the corresponding character position, and the inverted indexes of each character position are integrated to obtain a query index.

[0308] In one embodiment, the character sequence search rule includes a search order, and the processing unit 802, when searching for a second character sequence matching the first character sequence in the character sequence set based on the first compressed character sequence and at least one second compressed character sequence according to the instruction of the query index, is specifically configured to:

[0309] According to the search order indicated by the query index, based on the character position information of the first compressed character sequence, at least one second compressed character sequence is selected from the second compressed character sequences corresponding to each second character sequence in the character sequence set;

[0310] Determine the number of different characters between the at least one second compressed character sequence selected and the first compressed character sequence;

[0311] If the number of different characters is less than the preset number, determining the second character sequence corresponding to the corresponding second compressed character sequence as the candidate character sequence;

[0312] A second character sequence matching the first character sequence is obtained according to at least one candidate character sequence.

[0313] In one embodiment, the query index includes a multi-level inverted index, each level of the inverted index includes each character in a character set and a record list of each character; the character set is determined based on a character sequence set; a record list of a character is used to record at least one second compressed character sequence in which the corresponding character is located at a character position corresponding to the inverted index of a corresponding level;

[0314] When the processing unit 802 selects at least one second compressed character sequence from the second compressed character sequences corresponding to each second character sequence in the character sequence set according to the search order indicated by the query index and based on the character position information of the first compressed character sequence, the processing unit 802 is specifically configured to:

[0315] According to the search order indicated by the query index, and according to each character and the character position of each character in the first compressed character sequence, the record list included in the multi-level inverted index is searched to obtain a target record list; the target record list includes multiple record lists, each record list corresponds to a character at a character position in the first compressed character sequence;

[0316] Filtering the target record list to obtain a filtered target record list; the filtered target record list includes: multiple filtered record lists;

[0317] The second compressed character sequence included in the filtered target record list is determined as the screened second compressed character sequence.

[0318] In one embodiment, the character sequence search rule includes a search order; when the processing unit 802 searches the record list included in the multi-level inverted index according to the search order indicated by the query index and the character positions of the characters in the first compressed character sequence to obtain the target record list, it is specifically used to:

[0319] Traversing the characters in the first compressed character sequence, and determining the traversed character as the current character;

[0320] According to the current character and the character position j of the current character, the j-th level inverted index in the query index is scanned in the search order to obtain a record list of the current character, and the character at the character position j in the second compressed character sequence recorded in the record list of the current character is the current character; wherein the length of the first compressed character sequence is L, L is an integer greater than 1 and j∈[1, L];

[0321] When all characters in the first compressed character sequence are traversed, a record list of each character in the first compressed character sequence is obtained, and the record list of each character is determined as a target record list.

[0322] In one embodiment, the selected second compressed character sequence is recorded in the filtered target record list; when determining the number of different characters between the at least one selected second compressed character sequence and the first compressed character sequence, the processing unit 802 is specifically configured to:

[0323] Counting the number of occurrences of each second compressed character sequence recorded in the filtered target record list to obtain the occurrence frequency of each second compressed character sequence;

[0324] The occurrence frequency of each second compressed character sequence is calculated separately from the length of the first compressed character sequence to obtain the number of different characters between each second compressed character sequence and the first compressed character sequence.

[0325] In one embodiment, any candidate character sequence in at least one candidate character sequence is represented as candidate character sequence s i When the processing unit 802 obtains a second character sequence matching the first character sequence according to at least one candidate character sequence, it is specifically configured to:

[0326] For the candidate character sequence s i Calculate the similarity with the first character sequence to obtain the candidate character sequence s i Similarity with the first character sequence;

[0327] The candidate character sequence s i The corresponding similarity is compared with the similarity threshold;

[0328] If the similarity is greater than or equal to the similarity threshold, the candidate character sequence s i A second character sequence is determined to match the first character sequence.

[0329] In one embodiment, the query index includes a query tree, and the character sequence search rule includes a search path; the query tree is linked to a plurality of record lists, each of which is used to record a second compressed character sequence corresponding to a search path;

[0330] When the processing unit 802 searches for a second character sequence matching the first character sequence in the character sequence set based on the first compressed character sequence and at least one second compressed character sequence according to the instruction of the query index, it is specifically configured to:

[0331] Searching the query tree along the search path indicated by the query index based on character difference information between the first compressed character sequence and the second compressed character sequence corresponding to the search path to obtain a target record list; the target record list includes at least one record list linked to the query tree;

[0332] Filtering the target record list to obtain a filtered target record list; the filtered target record list includes at least one filtered record list;

[0333] According to the filtered target record list, a second character sequence matching the first character sequence is obtained.

[0334] In one embodiment, the filtering process includes: length filtering process, in which the target record list records the length of the second character sequence corresponding to the corresponding second compressed character sequence; when the processing unit 802 performs filtering process on the target record list to obtain the filtered target record list, it is specifically used to:

[0335] Get the length of the first character sequence;

[0336] Determining a length difference between the first character sequence and a corresponding second character sequence based on the length of the first character sequence and the length of the record in the target record list;

[0337] If the length difference is greater than the first difference threshold, the second compressed character sequence recorded in the corresponding record list included in the target record list is removed to obtain a filtered target record list.

[0338] In one embodiment, the filtering process includes: position filtering process, in which the target record list records the second character position of each character in the corresponding second compressed character sequence in the corresponding second character sequence; when the processing unit 802 performs filtering process on the target record list to obtain the filtered target record list, it is specifically used to:

[0339] If there is at least one identical reference character between the second compressed character sequence and the first compressed character sequence recorded in the target record list, obtaining the first character position of each reference character in the first character sequence;

[0340] Performing difference calculations based on the first character position and the second character position of each reference character recorded in the target record list to obtain a position difference of each reference character;

[0341] If the position difference of the corresponding reference characters is greater than the second difference threshold, the second compressed character sequence where the corresponding reference characters are recorded in the target record list is deleted to obtain a filtered target record list.

[0342] In one embodiment, the query tree includes a root node and a plurality of child nodes; each child node is used to store a character in the corresponding second compressed character sequence; the child node with the most nodes between it and the root node on the same search path in the query tree is a leaf node, and one leaf node is linked to one record list;

[0343] When the processing unit 802 searches the query tree along the search path indicated by the query index based on the character difference information between the first compressed character sequence and the second compressed character sequence corresponding to the search path to obtain the target record list, it is specifically configured to:

[0344] Starting from the root node of the query tree, traverse the query tree along the search path indicated by the query index, and use the traversed child node as the current child node;

[0345] Determine a target character position corresponding to the character stored in the current child node in the corresponding second compressed character sequence according to the depth of the current child node in the query tree, and compare the character stored in the current child node with the character at the target character position in the first compressed character sequence;

[0346] If the comparison shows that the character stored in the current child node is different from the character at the target character position in the first compressed character sequence, the tag value of the current child node is updated, and when the updated tag value is greater than the preset tag value, the branch where the current child node is located is pruned; wherein the tag value is used to record: the number of difference characters accumulated between the first compressed character sequence and the corresponding second compressed character sequence when reaching the position indicated by the current child node;

[0347] If the comparison shows that the character stored in the current child node is the same as the character at the target character position in the first compressed character sequence, or the updated tag value of the current child node is less than or equal to the preset tag value, the query tree continues to be traversed until the leaf node of the query tree is traversed, and the record list linked to the traversed leaf node is determined as the target record list.

[0348] In one embodiment, the number of first compressed character sequences includes M, and the number of second compressed character sequences corresponding to each second character sequence includes M; M is an integer greater than 1; the M first compressed character sequences and the M second compressed character sequences are obtained by using different compression parameters during compression using the target compression method;

[0349] The acquisition unit 801 is further used to: acquire M result sets;

[0350] The processing unit 802 is further configured to merge the M result sets to obtain a merged result; wherein each result set is used to record at least one second character sequence matching the first character sequence; and each result set is obtained by searching a first compressed character sequence and a second compressed character sequence corresponding to each second character sequence in the character sequence set;

[0351] The processing unit 802 is further configured to sort the second character sequence included in the merging result according to the similarity between the first character sequence and the second character sequence to obtain a sorting result; and perform service processing on the second character sequence matching the first character sequence according to the sorting result.

[0352] In one embodiment, the method provided by the embodiment of the present application is applied to an Internet scenario, and the output unit 803 is used to: output target service data adapted to the Internet scenario, where the target service data includes a second character sequence matching the first character sequence;

[0353] Among them, if the Internet scenario includes an advertising scenario, the target business data refers to the advertising data in the advertising scenario; if the Internet scenario includes a search scenario, the target business data refers to the search data in the search engine; the search data includes at least one of the following search types of data: text, image, video and audio; if the Internet scenario includes a shopping scenario, the target business data refers to the item data in the corresponding shopping platform; if the Internet scenario includes an electronic resource trading scenario, the target business data refers to product data that supports replacement through electronic resources.

[0354] The embodiment of the present application can reduce the character length of the character sequence originally searched by compressing the character sequence, thereby making the comparison in the character search process faster; in addition, the query index is a data structure for character sequence search constructed based on the compressed character sequence, which can provide a simple and lightweight index, and the character search rule indicated by the query index is a relatively efficient and simple rule, which can reduce the space occupied by the search and the search time during the character sequence search, thereby improving the search efficiency.

[0355] Next, the computer device provided in the embodiment of the present application is described.

[0356] See also Fig. 9 , Fig. 9 It is a schematic diagram of the structure of a computer device (i.e., the first consensus node mentioned above) provided in an embodiment of the present application. The computer device may include an independent device (e.g., one or more of a server, a node, a terminal, etc.), or may include components inside an independent device (e.g., a chip, a software module, or a hardware module, etc.). The computer device may include at least one processor 901 and a communication interface 902. Further optionally, the computer device may also include at least one memory 903 and a bus 904. Among them, the processor 901, the communication interface 902, and the memory 903 are connected via a bus 904.

[0357] Among them, the processor 901 is a module that performs arithmetic operations and / or logical operations, and can specifically be a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a complex programmable logic device (CPLD), a coprocessor (assisting the central processing unit to complete corresponding processing and applications), a microcontroller unit (MCU) and other processing modules, or a combination of multiple of them.

[0358] The communication interface 902 may be used to provide information input or output for at least one processor. And / or, the communication interface 902 may be used to receive data sent externally and / or send data externally, and may be a wired link interface including an Ethernet cable, etc., or a wireless link (Wi-Fi, Bluetooth, general wireless transmission, vehicle-mounted short-range communication technology, and other short-range wireless communication technologies, etc.) interface. The communication interface 902 may be used as a network interface.

[0359] The memory 903 is used to provide a storage space, and the storage space can store data such as an operating system and a computer program. The memory 903 can be a random access memory (RAM), a read-only memory (ROM), an erasable programmable read only memory (EPROM), or a portable read only memory (CD-ROM) or a combination of multiple thereof.

[0360] In one possible implementation, the processor 901 in the computer device is used to call a computer program stored in at least one memory 903 to perform the following operations: obtain a first character sequence to be searched; compress the first character sequence using a target compression method to obtain a first compressed character sequence corresponding to the first character sequence; obtain a query index, where the query index is constructed based on second compressed character sequences corresponding to each second character sequence included in the character sequence set; wherein each second character sequence is compressed using the target compression method to obtain a corresponding second compressed character sequence; the query index defines an index structure used when searching for a character sequence, and the query index is used to indicate a character sequence search rule; and according to the indication of the query index, based on the first compressed character sequence and at least one second compressed character sequence, search the character sequence set for a second character sequence that matches the first character sequence.

[0361] In one embodiment, the processor 901 is further used to: obtain a character sequence set, the character sequence set including multiple second character sequences; compress each second character sequence in the character sequence set using a target compression method to obtain a second compressed character sequence corresponding to each second character sequence; and construct a query index based on the second compressed character sequences corresponding to each second character sequence in the character sequence set.

[0362] In one embodiment, when the processor 901 compresses the first character sequence in a target compression manner to obtain a first compressed character sequence corresponding to the first character sequence, the processor 901 is specifically configured to:

[0363] Get an initialization character sequence, the initialization character sequence is used to store L characters, where L is an integer greater than 1;

[0364] According to the character features of each character in the first character sequence, a target character is selected from the first character sequence, and the target character is stored in the initialization character sequence to update the initialization character sequence; wherein the character feature of the target character is a key character feature in the first character sequence;

[0365] The first character sequence is recursively processed according to the target character to obtain a first compressed character sequence corresponding to the first character sequence based on the updated initialization character sequence.

[0366] In one embodiment, when the processor 901 selects the target character from the first character sequence according to the character features of each character in the first character sequence, it is specifically configured to:

[0367] Obtaining an interval length parameter, and determining a target position interval of the first character sequence according to the interval length parameter and the length of the first character sequence;

[0368] Using a hash function to perform hash calculation on characters in the first character sequence whose character positions are in the target position interval, to obtain a hash value of at least one character in the target position interval;

[0369] According to the hash value of at least one character in the target position interval, a character with the smallest hash value is selected from the at least one character in the target position interval, and the selected character is determined as the target character.

[0370] In one embodiment, L is determined based on a preset number of times l of recursively processing the first character sequence, where l is a positive integer; when the processor 901 recursively processes the first character sequence according to the target character to obtain a first compressed character sequence corresponding to the first character sequence based on the updated initialization character sequence, it is specifically configured to:

[0371] Divide the first character sequence according to the target character to obtain two sub-character sequences;

[0372] Determine each sub-character sequence as a first character sequence, and repeatedly execute the target step, which is: select a target character from the first character sequence according to the character features of each character in the first character sequence, and store the target character in the initialization character sequence to update the initialization character sequence;

[0373] When the number of target characters stored in the updated initialization character sequence reaches the length L of the initialization character sequence, the updated initialization character sequence is determined as the first compressed character sequence corresponding to the first character sequence.

[0374] In one embodiment, when the processor 901 constructs the query index based on the second compressed character sequences respectively corresponding to the second character sequences in the character sequence set, it is specifically configured to:

[0375] Get the character set corresponding to the character sequence set, the character set includes multiple characters;

[0376] Traversing the characters included in the character set, determining the traversed character as the current character, and selecting at least one second compressed character sequence including the current character from the second compressed character sequences corresponding to each second character sequence in the character sequence set according to the current character;

[0377] Obtaining the reference character position of the current character in each selected second compressed character sequence, and adding each selected second compressed character sequence to the record list corresponding to the current character at the corresponding reference character position;

[0378] After all characters in the character set have been traversed, a list of records corresponding to each character in the character set at each character position is obtained;

[0379] Each character in the character set and a record list corresponding to each character at the same character position are determined as an inverted index of the corresponding character position, and the inverted indexes of each character position are integrated to obtain a query index.

[0380] In one embodiment, the character sequence search rule includes a search order, and the processor 901, when searching for a second character sequence matching the first character sequence in the character sequence set based on the first compressed character sequence and at least one second compressed character sequence according to the instruction of the query index, is specifically configured to:

[0381] According to the search order indicated by the query index, based on the character position information of the first compressed character sequence, at least one second compressed character sequence is selected from the second compressed character sequences corresponding to each second character sequence in the character sequence set;

[0382] Determine the number of different characters between the at least one second compressed character sequence selected and the first compressed character sequence;

[0383] If the number of different characters is less than the preset number, determining the second character sequence corresponding to the corresponding second compressed character sequence as the candidate character sequence;

[0384] A second character sequence matching the first character sequence is obtained according to at least one candidate character sequence.

[0385] In one embodiment, the query index includes a multi-level inverted index, each level of the inverted index includes each character in a character set and a record list of each character; the character set is determined based on a character sequence set; a record list of a character is used to record at least one second compressed character sequence in which the corresponding character is located at a character position corresponding to the inverted index of a corresponding level;

[0386] When the processor 901 selects at least one second compressed character sequence from the second compressed character sequences corresponding to each second character sequence in the character sequence set according to the search order indicated by the query index and based on the character position information of the first compressed character sequence, the processor 901 is specifically configured to:

[0387] According to the search order indicated by the query index, and according to each character and the character position of each character in the first compressed character sequence, the record list included in the multi-level inverted index is searched to obtain a target record list; the target record list includes multiple record lists, each record list corresponds to a character at a character position in the first compressed character sequence;

[0388] Filtering the target record list to obtain a filtered target record list; the filtered target record list includes: multiple filtered record lists;

[0389] The second compressed character sequence included in the filtered target record list is determined as the screened second compressed character sequence.

[0390] In one embodiment, the character sequence search rule includes a search order; when the processor 901 searches the record list included in the multi-level inverted index according to the search order indicated by the query index and according to each character and the character position of each character in the first compressed character sequence to obtain the target record list, it is specifically used to:

[0391] Traversing the characters in the first compressed character sequence, and determining the traversed character as the current character;

[0392] According to the current character and the character position j of the current character, the j-th level inverted index in the query index is scanned in the search order to obtain a record list of the current character, and the character at the character position j in the second compressed character sequence recorded in the record list of the current character is the current character; wherein the length of the first compressed character sequence is L, L is an integer greater than 1 and j∈[1, L];

[0393] When all characters in the first compressed character sequence are traversed, a record list of each character in the first compressed character sequence is obtained, and the record list of each character is determined as a target record list.

[0394] In one embodiment, the selected second compressed character sequence is recorded in the filtered target record list; when determining the number of different characters between at least one selected second compressed character sequence and the first compressed character sequence, the processor 901 is specifically configured to:

[0395] Counting the number of occurrences of each second compressed character sequence recorded in the filtered target record list to obtain the occurrence frequency of each second compressed character sequence;

[0396] The occurrence frequency of each second compressed character sequence is calculated separately from the length of the first compressed character sequence to obtain the number of different characters between each second compressed character sequence and the first compressed character sequence.

[0397] In one embodiment, any candidate character sequence in at least one candidate character sequence is represented as candidate character sequence s i When the processor 901 obtains a second character sequence matching the first character sequence according to at least one candidate character sequence, it is specifically configured to:

[0398] For the candidate character sequence s i Calculate the similarity with the first character sequence to obtain the candidate character sequence s i Similarity with the first character sequence;

[0399] The candidate character sequence s i The corresponding similarity is compared with the similarity threshold;

[0400] If the similarity is greater than or equal to the similarity threshold, the candidate character sequence s i A second character sequence is determined to match the first character sequence.

[0401] In one embodiment, the query index includes a query tree, and the character sequence search rule includes a search path; the query tree is linked to a plurality of record lists, each of which is used to record a second compressed character sequence corresponding to a search path;

[0402] When the processor 901 searches for a second character sequence matching the first character sequence in the character sequence set based on the first compressed character sequence and at least one second compressed character sequence according to the instruction of the query index, it is specifically configured to:

[0403] Searching the query tree along the search path indicated by the query index based on character difference information between the first compressed character sequence and the second compressed character sequence corresponding to the search path to obtain a target record list; the target record list includes at least one record list linked to the query tree;

[0404] Filtering the target record list to obtain a filtered target record list; the filtered target record list includes at least one filtered record list;

[0405] According to the filtered target record list, a second character sequence matching the first character sequence is obtained.

[0406] In one embodiment, the filtering process includes: length filtering process, in which the target record list records the length of the second character sequence corresponding to the corresponding second compressed character sequence; when the processor 901 performs filtering process on the target record list to obtain the filtered target record list, it is specifically used to:

[0407] Get the length of the first character sequence;

[0408] Determining a length difference between the first character sequence and a corresponding second character sequence based on the length of the first character sequence and the length of the record in the target record list;

[0409] If the length difference is greater than the first difference threshold, the second compressed character sequence recorded in the corresponding record list included in the target record list is removed to obtain a filtered target record list.

[0410] In one embodiment, the filtering process includes: position filtering process, in which the target record list records the second character position of each character in the corresponding second compressed character sequence in the corresponding second character sequence; when the processor 901 performs filtering process on the target record list to obtain the filtered target record list, it is specifically used to:

[0411] If there is at least one identical reference character between the second compressed character sequence and the first compressed character sequence recorded in the target record list, obtaining the first character position of each reference character in the first character sequence;

[0412] Performing difference calculations based on the first character position and the second character position of each reference character recorded in the target record list to obtain a position difference of each reference character;

[0413] If the position difference of the corresponding reference characters is greater than the second difference threshold, the second compressed character sequence where the corresponding reference characters are recorded in the target record list is deleted to obtain a filtered target record list.

[0414] In one embodiment, the query tree includes a root node and a plurality of child nodes; each child node is used to store a character in the corresponding second compressed character sequence; the child node with the most nodes between it and the root node on the same search path in the query tree is a leaf node, and one leaf node is linked to one record list;

[0415] When the processor 901 searches the query tree along the search path indicated by the query index based on the character difference information between the first compressed character sequence and the second compressed character sequence corresponding to the search path to obtain the target record list, it is specifically configured to:

[0416] Starting from the root node of the query tree, traverse the query tree along the search path indicated by the query index, and use the traversed child node as the current child node;

[0417] Determine a target character position corresponding to the character stored in the current child node in the corresponding second compressed character sequence according to the depth of the current child node in the query tree, and compare the character stored in the current child node with the character at the target character position in the first compressed character sequence;

[0418] If the comparison shows that the character stored in the current child node is different from the character at the target character position in the first compressed character sequence, the tag value of the current child node is updated, and when the updated tag value is greater than the preset tag value, the branch where the current child node is located is pruned; wherein the tag value is used to record: the number of difference characters accumulated between the first compressed character sequence and the corresponding second compressed character sequence when reaching the position indicated by the current child node;

[0419] If the comparison shows that the character stored in the current child node is the same as the character at the target character position in the first compressed character sequence, or the updated tag value of the current child node is less than or equal to the preset tag value, the query tree continues to be traversed until the leaf node of the query tree is traversed, and the record list linked to the traversed leaf node is determined as the target record list.

[0420] In one embodiment, the number of first compressed character sequences includes M, and the number of second compressed character sequences corresponding to each second character sequence includes M; M is an integer greater than 1; the M first compressed character sequences and the M second compressed character sequences are all obtained using different compression parameters during compression using a target compression method; the processor 901 is also used to perform the following operations: obtain M result sets, and merge the M result sets to obtain a merged result; wherein each result set is used to record at least one second character sequence matching the first character sequence; and each result set is obtained by searching for a second compressed character sequence corresponding to a first compressed character sequence and each second character sequence in the character sequence set; according to the similarity between the first character sequence and the second character sequence, the second character sequences included in the merged result are sorted to obtain a sorted result; and the second character sequence matching the first character sequence is processed according to the sorted result.

[0421] In one embodiment, the method provided in the embodiment of the present application is applied to an Internet scenario, and the processor 901 is further used to call the communication interface 902: output target service data adapted to the Internet scenario, the target service data including a second character sequence matching the first character sequence;

[0422] Among them, if the Internet scenario includes an advertising scenario, the target business data refers to the advertising data in the advertising scenario; if the Internet scenario includes a search scenario, the target business data refers to the search data in the search engine; the search data includes at least one of the following search types of data: text, image, video and audio; if the Internet scenario includes a shopping scenario, the target business data refers to the item data in the corresponding shopping platform; if the Internet scenario includes an electronic resource trading scenario, the target business data refers to product data that supports replacement through electronic resources.

[0423] The embodiment of the present application can reduce the character length of the character sequence originally searched by compressing the character sequence, thereby making the comparison in the character search process faster; in addition, the query index is a data structure for character sequence search constructed based on the compressed character sequence, which can provide a simple and lightweight index, and the character search rule indicated by the query index is a relatively efficient and simple rule, which can reduce the space occupied by the search and the search time during the character sequence search, thereby improving the search efficiency.

[0424] It should be noted that in the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0425] In addition, it should be noted that the embodiment of the present application also provides a computer-readable storage medium, in which a computer program of the aforementioned character sequence search method is stored. When the computer program is executed by a processor, the description of the character sequence search method in the embodiment of the present application is executed. That is, when one or more processors load and execute the computer program, the description of the character sequence search method in the embodiment can be implemented, which will not be repeated here, and the description of the beneficial effects of using the same method will not be repeated here.

[0426] The computer-readable storage medium may be the character sequence search device provided in any of the aforementioned embodiments or the internal storage unit of the computer device, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the computer device. Further, the computer-readable storage medium may also include both the internal storage unit of the computer device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.

[0427] In one aspect of the present application, a computer program product or computer program is provided, the computer program product includes a computer program, the computer program is stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device performs the method for searching a character sequence provided in one aspect of the embodiments of the present application.

[0428] In one aspect of the present application, another computer program product or computer program is provided. The computer program product includes a computer program. When the computer program is executed by a processor, the steps of the character sequence search method provided in the embodiment of the present application are implemented.

[0429] The steps in the method of the embodiment of the present application can be adjusted in order, combined and deleted according to actual needs.

[0430] The units in the device of the embodiment of the present application can be merged, divided and deleted according to actual needs.

[0431] The above disclosure is only part of the embodiments of the present application, which cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made in the claims of the present application still fall within the scope covered by the present application.

Claims

1. A character sequence search method, It is characterized in that The method comprises: Acquire a first character sequence to be searched, and compress the first character sequence using a target compression method to obtain a first compressed character sequence corresponding to the first character sequence; Obtaining a query index, wherein the query index is constructed based on second compressed character sequences corresponding to respective second character sequences included in the character sequence set; wherein each second character sequence is compressed using the target compression method to obtain a corresponding second compressed character sequence; the query index defines an index structure used when searching for a character sequence, and the query index is used to indicate a character sequence search rule; According to the instruction of the query index, based on the first compressed character sequence and at least one second compressed character sequence, the character sequence set is searched for a second character sequence matching the first character sequence.

2. The method according to claim 1, It is characterized in that The step of compressing the first character sequence in a target compression manner to obtain a first compressed character sequence corresponding to the first character sequence includes: Obtain an initialization character sequence, where the initialization character sequence is used to store L characters, where L is an integer greater than 1; According to the character features of each character in the first character sequence, a target character is selected from the first character sequence, and the target character is stored in the initialization character sequence to update the initialization character sequence; wherein the character feature of the target character is a key character feature in the first character sequence; The first character sequence is recursively processed according to the target character to obtain a first compressed character sequence corresponding to the first character sequence based on the updated initialization character sequence.

3. The method according to claim 2, It is characterized in that The step of selecting a target character from the first character sequence according to the character features of each character in the first character sequence includes: Acquire an interval length parameter, and determine a target position interval of the first character sequence according to the interval length parameter and the length of the first character sequence; Using a hash function to perform hash calculation on characters in the first character sequence whose character positions are in the target position interval, to obtain a hash value of at least one character in the target position interval; According to the hash value of at least one character in the target position interval, a character with the smallest hash value is selected from the at least one character in the target position interval, and the selected character is determined as the target character.

4. The method according to claim 2, It is characterized in that The L is determined based on a preset number of times l of recursively processing the first character sequence, where l is a positive integer; the recursively processing the first character sequence according to the target character to obtain a first compressed character sequence corresponding to the first character sequence based on the updated initialization character sequence includes: Dividing the first character sequence according to the target character to obtain two sub-character sequences; Determine each of the sub-character sequences as the first character sequence, and repeatedly execute the target step, wherein the target step refers to the step of selecting a target character from the first character sequence according to the character features of each character in the first character sequence, and storing the target character in the initialization character sequence to update the initialization character sequence; When the number of target characters stored in the updated initialization character sequence reaches the length L of the initialization character sequence, the updated initialization character sequence is determined as a first compressed character sequence corresponding to the first character sequence.

5. The method according to claim 1, It is characterized in that The method further comprises: Obtain a character set corresponding to the character sequence set, the character set including a plurality of characters; Traversing the characters included in the character set, determining the traversed character as the current character, and selecting at least one second compressed character sequence including the current character from the second compressed character sequences corresponding to each second character sequence in the character sequence set according to the current character; Obtaining the reference character position of the current character in each of the selected second compressed character sequences, and adding each of the selected second compressed character sequences to a record list corresponding to the current character at the corresponding reference character position; After all characters in the character set are traversed, a record list corresponding to each character in the character set at each character position is obtained; Each character in the character set and a record list corresponding to each character at the same character position are determined as an inverted index of the corresponding character position, and the inverted indexes of each character position are integrated to obtain a query index.

6. The method according to any one of claims 1 to 5, It is characterized in that The character sequence search rule includes a search order, and the searching, in accordance with the instruction of the query index, for a second character sequence matching the first character sequence in the character sequence set based on the first compressed character sequence and at least one second compressed character sequence, includes: According to the search order indicated by the query index, based on the character position information of the first compressed character sequence, at least one second compressed character sequence is selected from the second compressed character sequences corresponding to each second character sequence in the character sequence set; Determine the number of different characters between at least one of the second compressed character sequences selected and the first compressed character sequence; If the number of the difference characters is less than a preset number, determining the second character sequence corresponding to the second compressed character sequence as a candidate character sequence; A second character sequence matching the first character sequence is obtained according to at least one of the candidate character sequences.

7. The method according to claim 6, It is characterized in that The query index includes a multi-level inverted index, each level of the inverted index includes each character in a character set and a record list of each character; the character set is determined based on the character sequence set; a record list of a character is used to record at least one second compressed character sequence in which the corresponding character is located at a character position corresponding to the inverted index of a corresponding level; The step of selecting at least one second compressed character sequence from second compressed character sequences corresponding to each second character sequence in the character sequence set according to the search order indicated by the query index and based on character position information of the first compressed character sequence includes: According to the search order indicated by the query index, and according to each character and the character position of each character in the first compressed character sequence, the record list included in the multi-level inverted index is searched to obtain a target record list; the target record list includes a plurality of record lists, each record list corresponds to a character at a character position in the first compressed character sequence; Filtering the target record list to obtain a filtered target record list; the filtered target record list includes: a plurality of filtered record lists; The second compressed character sequence included in the filtered target record list is determined as the screened second compressed character sequence.

8. The method according to claim 7, It is characterized in that The character sequence search rule includes a search order; the search order indicated by the query index is followed by searching the record list included in the multi-level inverted index according to each character and the character position of each character in the first compressed character sequence to obtain a target record list, including: Traversing the characters in the first compressed character sequence, and determining the traversed character as the current character; According to the current character and the character position j of the current character, the j-th level inverted index in the query index is scanned in the search order to obtain a record list of the current character, wherein the character at the character position j in the second compressed character sequence recorded in the record list of the current character is the current character; wherein the length of the first compressed character sequence is L, L is an integer greater than 1, and j∈[1, L]; When all characters in the first compressed character sequence are traversed, a record list of each character in the first compressed character sequence is obtained, and the record list of each character is determined as a target record list.

9. The method according to claim 6, It is characterized in that The screened second compressed character sequence is recorded in the filtered target record list; The determining the number of different characters between at least one of the second compressed character sequences selected and the first compressed character sequence comprises: Counting the number of occurrences of each second compressed character sequence recorded in the filtered target record list in the filtered target record list to obtain the occurrence frequency of each second compressed character sequence; The occurrence frequency of each second compressed character sequence is calculated separately from the length of the first compressed character sequence to obtain the number of different characters between each second compressed character sequence and the first compressed character sequence.

10. The method according to claim 6, It is characterized in that Any one of the at least one candidate character sequence is represented as a candidate character sequence s i ; The step of obtaining a second character sequence matching the first character sequence based on at least one of the candidate character sequences comprises: For the candidate character sequence s i Calculate the similarity with the first character sequence to obtain the candidate character sequence s i The similarity between the first character sequence and the first character sequence; The candidate character sequence s i The corresponding similarity is compared with the similarity threshold; If the similarity is greater than or equal to the similarity threshold, the candidate character sequence s i A second character sequence is determined to match the first character sequence.

11. The method according to claim 1, It is characterized in that The query index includes a query tree, and the character sequence search rule includes a search path; the query tree is linked to a plurality of record lists, each of which is used to record a second compressed character sequence corresponding to a search path; The step of searching, in accordance with the instruction of the query index, for a second character sequence matching the first character sequence in the character sequence set based on the first compressed character sequence and at least one second compressed character sequence, comprises: Searching the query tree along the search path indicated by the query index based on character difference information between the first compressed character sequence and a second compressed character sequence corresponding to the search path to obtain a target record list; the target record list includes at least one record list linked to the query tree; Filtering the target record list to obtain a filtered target record list; the filtered target record list includes at least one filtered record list; A second character sequence matching the first character sequence is obtained according to the filtered target record list.

12. The method according to claim 7 or 11, It is characterized in that The filtering process includes: length filtering process, wherein the target record list records the length of the second character sequence corresponding to the corresponding second compressed character sequence; the filtering process on the target record list to obtain the filtered target record list includes: Obtaining the length of the first character sequence; Determining a length difference between the first character sequence and a corresponding second character sequence according to the length of the first character sequence and the length of a record in the target record list; If the length difference is greater than a first difference threshold, the second compressed character sequence recorded in the corresponding record list included in the target record list is removed to obtain a filtered target record list.

13. The method according to claim 7 or 11, It is characterized in that The filtering process includes: position filtering process, wherein the target record list records the second character position of each character in the corresponding second compressed character sequence in the corresponding second character sequence; the filtering process on the target record list to obtain the filtered target record list includes: If there is at least one identical reference character between the second compressed character sequence recorded in the target record list and the first compressed character sequence, obtaining the first character position of each reference character in the first character sequence; Based on the first character position and the second character position of each reference character recorded in the target record list, respectively, a difference calculation is performed to obtain a position difference of each reference character; If the position difference of the corresponding reference character is greater than a second difference threshold, the second compressed character sequence where the corresponding reference character is recorded in the target record list is deleted to obtain a filtered target record list.

14. The method according to claim 11, It is characterized in that The query tree includes a root node and a plurality of child nodes; each child node is used to store a character in the corresponding second compressed character sequence; the child node in the query tree that is separated from the root node by the most nodes on the same search path is a leaf node, and one leaf node is linked to one record list; The step of searching the query tree along the search path indicated by the query index based on character difference information between the first compressed character sequence and a second compressed character sequence corresponding to the search path to obtain a target record list includes: Starting from the root node of the query tree, traverse the query tree along the search path indicated by the query index, and use the traversed child node as the current child node; Determine a target character position corresponding to the character stored in the current child node in the corresponding second compressed character sequence according to the depth of the current child node in the query tree, and compare the character stored in the current child node with the character at the target character position in the first compressed character sequence; If the comparison shows that the character stored in the current child node is different from the character at the target character position in the first compressed character sequence, the tag value of the current child node is updated, and when the updated tag value is greater than a preset tag value, the branch where the current child node is located is pruned; wherein the tag value is used to record: the number of difference characters accumulated between the first compressed character sequence and the corresponding second compressed character sequence when reaching the position indicated by the current child node; If comparison shows that the character stored in the current child node is the same as the character at the target character position in the first compressed character sequence, or the updated tag value of the current child node is less than or equal to the preset tag value, continue traversing the query tree until the leaf node of the query tree is traversed, and determine the record list linked to the traversed leaf node as the target record list.

15. The method of claim 1, It is characterized in that The number of the first compressed character sequences includes M, and the number of second compressed character sequences corresponding to each second character sequence includes M; M is an integer greater than 1; the M first compressed character sequences and the M second compressed character sequences are obtained by using different compression parameters during compression using the target compression method; the method further includes: Obtaining M result sets, and merging the M result sets to obtain a merged result; wherein each result set is used to record at least one second character sequence matching the first character sequence; and each result set is obtained by searching based on one of the first compressed character sequences and a second compressed character sequence corresponding to each second character sequence in the character sequence set; sorting the second character sequence included in the merging result according to the similarity between the first character sequence and the second character sequence to obtain a sorting result; Business processing is performed on a second character sequence matching the first character sequence according to the sorting result.

16. The method of claim 1, It is characterized in that The method is applied to an Internet scenario, and the method further includes: Outputting target service data adapted to the Internet scenario, wherein the target service data includes a second character sequence matching the first character sequence; Wherein, if the Internet scenario includes an advertising scenario, the target service data refers to the advertising data in the advertising scenario; If the Internet scenario includes a search scenario, the target business data refers to search data in a search engine; the search data includes at least one of the following search types of data: text, image, video, and audio; If the Internet scenario includes a shopping scenario, the target business data refers to item data in the corresponding shopping platform; If the Internet scenario includes an electronic resource transaction scenario, the target business data refers to product data that supports replacement through electronic resources.

17. A character sequence search device, It is characterized in that The device comprises: An acquisition unit, used for acquiring a first character sequence to be searched; a processing unit, configured to compress the first character sequence in a target compression manner to obtain a first compressed character sequence corresponding to the first character sequence; The acquisition unit is further used to acquire a query index, wherein the query index is constructed based on second compressed character sequences corresponding to respective second character sequences included in the character sequence set; wherein each second character sequence is compressed using the target compression method to obtain a corresponding second compressed character sequence; the query index defines an index structure used when searching for a character sequence, and the query index is used to indicate a character sequence search rule; The processing unit is further configured to search the character sequence set for a second character sequence matching the first character sequence based on the first compressed character sequence and at least one second compressed character sequence according to an instruction of the query index.

18. A computer device, It is characterized in that include: a processor suitable for executing a computer program; A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the character sequence search method according to any one of claims 1 to 16 is executed.

19. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the character sequence search method according to any one of claims 1 to 16 is executed.

20. A computer program product, It is characterized in that The computer program product comprises a computer program or computer instructions, and the computer program or computer instructions are executed by a processor to implement the character sequence searching method as claimed in any one of claims 1 to 16.