A search method, a backend server and a search system

By introducing filters into the backend server and using the feature information of the search information to perform the first search, the problem of high performance pressure on the search server is solved and more efficient search services are achieved.

CN118551095BActive Publication Date: 2025-05-23HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410559132.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-19
Publication Date
2025-05-23
Estimated Expiration
2043-12-19

AI Technical Summary

Technical Problem

In the prior art, the search server has high performance pressure due to processing massive application data, resulting in slowing down the search speed and affecting the user experience.

Method used

By introducing filters into the background server, the first search is performed using the characteristic information of the search information, and whether the content exists. If it exists, invalid requests will be intercepted and the load on the search server will be reduced.

Benefits of technology

Effectively intercepting most invalid requests reduces the pressure on search servers, improves search efficiency, and provides users with more efficient search services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118551095B_ABST
    Figure CN118551095B_ABST
Patent Text Reader

Abstract

The present application relates to the field of data processing, and in particular to a search method, a backend server, and a search system. The method is applied to a backend server, and a filter in the backend server is used to perform a first search on search information. If the content corresponding to the search information exists, the search server performs a second search; if the content corresponding to the search information does not exist, a preset information indicating that the content does not exist is returned to the client. In the above method, most invalid requests can be intercepted by the filter in the backend server, thereby effectively reducing the pressure on the search server and providing users with more efficient search services.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the Chinese patent application submitted to the State Intellectual Property Office on December 19, 2023, with application number 202311744805.2 and application name "A search method, backend server and search system". Technical Field

[0002] The present application relates to the field of data processing, and in particular to a search method, a backend server and a search system. Background Art

[0003] Application search refers to the technology of searching for applications (or application software, apps, and applications) based on demand. Application search services usually need to process massive amounts of application data, which are generally stored in search servers. As time goes by, application data continues to increase, and more and more data is stored in search servers, the performance pressure of search servers increases, and search becomes slower and slower, affecting user experience. Summary of the invention

[0004] The present application provides a search method, a backend server and a search system, which solve the problem of high search pressure on the search server in the prior art.

[0005] In order to achieve the above objectives, this application adopts the following technical solutions:

[0006] In a first aspect, a search method is provided, which is applied to a backend server, and the method comprises:

[0007] After receiving the search information sent by the client, performing a first search according to the search information to obtain a first result;

[0008] If the first result indicates that the content corresponding to the search information does not exist, returning preset information to the client;

[0009] If the first result indicates that the content corresponding to the search information exists, sending the search information to a search server, so that the search server performs a second search according to the search information and returns a second result obtained by the second search;

[0010] After receiving the second result returned by the search server, the second result is sent to the client.

[0011] In the embodiment of the present application, based on the principle that existence may exist and non-existence must not exist, when the backend server determines that the search information does not exist, there is no need for the search server to search again. Accordingly, the filter in the backend server sends preset information to the client. Through the above method, most invalid requests can be intercepted by the backend server, thereby effectively reducing the pressure on the search server and providing users with more efficient search services.

[0012] In an implementation of the first aspect, performing a first search according to the search information to obtain a first result includes:

[0013] Extracting first characteristic information from the search information;

[0014] Acquire a first storage location of the first feature information in the storage space;

[0015] If the first characteristic information is stored in the first storage location, the first result indicates that the content corresponding to the search information exists;

[0016] If the first characteristic information is not stored in the first storage location, the first result indicates that the content corresponding to the search information does not exist.

[0017] In the embodiment of the present application, since the data volume of the first characteristic information is usually smaller than the data volume of the search information, the method of searching by the filter according to the first characteristic information is more efficient than the method of searching by the search server according to the search information. The first search of the search information using the filter can intercept most invalid requests, which can not only improve the search efficiency, but also help reduce the search pressure of the search server, and provide users with more efficient search services.

[0018] In an implementation manner of the first aspect, extracting the first feature information from the search information includes:

[0019] Performing hash processing according to the character position of each character in the search information in the character string corresponding to the search information to obtain a first character hash value;

[0020] The first feature information is extracted according to the first character hash value.

[0021] Since the order of characters in a sentence is closely related to the semantics of the sentence, the hash processing in the embodiment of the present application is equivalent to extracting the feature information of the search information through the character positions of the characters in the search information, which can reflect the semantic features of the search information to a certain extent, thereby facilitating improving the search accuracy of subsequent information searches.

[0022] In an implementation of the first aspect, performing hash processing according to the character position of each character in the search information in the character string corresponding to the search information to obtain a first character hash value includes:

[0023] For the first character in the search information, calculate the hash value of the character position of the first character in the character string corresponding to the search information to obtain the first character hash value of the first character; update the initial memory value according to the first character hash value of the first character; update the initial shift bit number according to the first preset value to obtain the shift bit number corresponding to the first character;

[0024] For the i-th character in the search information, calculate the hash value of the character position of the i-th character in the character string corresponding to the search information to obtain the first character hash value of the i-th character;

[0025] Shift the first character hash value of the i-th character according to the number of shift bits corresponding to the i-1-th character to obtain the first character hash value of the i-th character after the shift, where i is an integer greater than 1;

[0026] Update the first memory value according to the first character hash value of the i-th character after the shift to obtain a second memory value, wherein the first memory value is used to represent the first character hash value of the first i-1 characters after the shift, and the second memory value is used to represent the first character hash value of the first i characters after the shift;

[0027] The number of shift bits corresponding to the i-1th character is updated according to a first preset value to obtain the number of shift bits corresponding to the i-th character, wherein the first preset value is determined according to the data interval to which the first character hash value of the i-th character before shifting belongs.

[0028] In the embodiment of the present application, as the character position increases, the first preset value also increases, which can effectively reduce the conflict between the character hash values ​​of each character in the memory value.

[0029] In an implementation of the first aspect, the method further includes:

[0030] After obtaining the number of shift bits corresponding to the i-th character, if the number of shift bits corresponding to the i-th character is greater than a shift threshold, obtaining data from the second memory value according to the shift threshold to obtain a third memory value;

[0031] The number of shift bits corresponding to the i-th character is reduced according to the shift threshold to obtain an updated number of shift bits corresponding to the i-th character.

[0032] In the embodiment of the present application, when the number of shift bits exceeds the shift threshold, the number of shift bits is reduced, which can effectively reduce the occurrence of excessive data bits after the shift processing due to excessive shift bits; in addition, obtaining part of the data from the current memory value can also effectively control the length of the memory value. In the above manner, the length of the memory value can be effectively controlled to ensure that the memory value does not exceed the memory space.

[0033] In an implementation manner of the first aspect, acquiring data from the second memory value according to the shift threshold to obtain a third memory value includes:

[0034] According to the order of the number of digits from high to low, the values ​​of M digits in the second memory value are obtained to obtain the third memory value, wherein M is determined according to the shift threshold.

[0035] In the embodiment of the present application, the high-order numerical value is equivalent to the hash value of the character position of the later character in the string, and the low-order numerical value is equivalent to the hash value of the character position of the earlier character in the string. In the above method, the high-order numerical value is retained and the low-order numerical value is filtered out, which is equivalent to retaining the hash value of the character position of the later character in the string as much as possible. In some application scenarios, the characters at the later position in the string are mostly searched keywords. Therefore, the memory value obtained by the above method can retain more characteristic information of the keywords in the string.

[0036] In an implementation of the first aspect, extracting the first feature information according to the first character hash value includes:

[0037] Perform magnification processing on the fourth memory value to obtain the processed fourth memory value, wherein the fourth memory value is used to represent the first character hash value of the N characters after the shift, and N is the total number of characters in the search information;

[0038] The first feature information is calculated according to the fourth memory value.

[0039] In the embodiment of the present application, by amplifying the memory value, the number of bits of the processed memory value is increased, which can reduce data conflicts in subsequent calculation processes.

[0040] In an implementation manner of the first aspect, calculating the first feature information according to the fourth memory value includes:

[0041] Perform an XOR operation on the second preset value and the fourth memory value to obtain an XOR value;

[0042] The first feature information is calculated according to the XOR value.

[0043] In the XOR operation, if the two values ​​are the same, the XOR result is 0; if the two values ​​are different, the XOR result is 1. The XOR operation is equivalent to detecting whether the fourth memory value is the same as the second preset value. If the fourth memory value is the same as the second preset value, the XOR value is 0. When two sets of application data are the same, and their corresponding fourth memory values ​​are the same as the second preset value, the XOR value can be used to determine that the two sets of application data are the same. Correspondingly, the hash values ​​corresponding to the two sets of application data are the same and the feature information is also the same. When the XOR value is 0, the filter does not process the search information corresponding to the current first feature information, that is, it does not filter, and it is handed over to the search server for processing. In this way, the probability of missed detection and false detection can be reduced, and the search accuracy can be improved.

[0044] Since the first character hash value of each character in the search information can represent the feature information of the character position in the character string, the fourth memory value is equivalent to containing the feature information of the character positions of all characters in the search information. Therefore, in the embodiment of the present application, calculating the first feature information based on the XOR value is equivalent to extracting feature information based on the character positions of all characters in the search information. The feature information extracted in this way can represent the positional relationship between the characters in the search information, thereby representing the semantics corresponding to the search information.

[0045] In an implementation of the first aspect, the method further includes:

[0046] After obtaining the number of shift bits corresponding to the i-th character, if the number of shift bits corresponding to the i-th character is greater than the shift threshold, amplifying the second preset value to obtain the processed second preset value;

[0047] The performing an XOR operation on the second preset value and the fourth memory value to obtain an XOR value includes:

[0048] If the number of shift bits corresponding to the i-th character is greater than the shift threshold, an XOR operation is performed based on the processed second preset value and the fourth memory value to obtain an XOR value.

[0049] In the embodiment of the present application, when the number of shift bits is greater than the shift threshold, it means that the number of bits in the current memory value is relatively large. In this case, the second preset value is magnified, which is equivalent to increasing the number of bits of the second preset value and increasing the complexity of the second preset value. In this way, data conflicts in subsequent calculations can be reduced.

[0050] In an implementation manner of the first aspect, obtaining a first storage position of the first feature information in a storage space includes:

[0051] Calculating a first Hash value of the first feature information;

[0052] The first storage location is calculated according to the first Hash value and a storage capacity of the storage space.

[0053] The hash algorithm is a secure hash algorithm. In the embodiment of the present application, the storage location is determined by the hash algorithm, which can effectively reduce the probability of collision of feature information and help improve the accuracy of subsequent information searches.

[0054] In an implementation of the first aspect, calculating the first storage location according to the first hash value and a storage capacity of the storage space includes:

[0055] A modulo operation is performed according to the first Hash value and the storage capacity of the storage space to obtain the first storage position.

[0056] The modulo operation method can quickly and simply determine the storage location and can ensure that the calculated storage location falls within the range of the storage space.

[0057] In an implementation of the first aspect, the method further includes:

[0058] Receive application data uploaded by the developer;

[0059] extracting second characteristic information from the application data;

[0060] Calculating a second storage position of the second characteristic information in the storage space;

[0061] The second feature information is stored in the second storage location.

[0062] Since the second characteristic information can reflect the characteristics of the application data and the data volume is relatively small, storing the second characteristic information is conducive to saving storage space compared to storing the application data.

[0063] In an implementation manner of the first aspect, storing the second feature information in the second storage location includes:

[0064] If the second storage location is not empty, the second characteristic information replaces the third characteristic information, wherein the third characteristic information is the data stored in the second storage location.

[0065] Through the above method, it is possible to ensure that the second characteristic information of the latest application data is saved.

[0066] In an implementation of the first aspect, the second storage location includes a main location and a plurality of sub-locations; and the method further includes:

[0067] If the main location in the second storage location is not empty, determining whether the sub-location in the second storage location is empty;

[0068] If each sub-position in the second storage position is not empty, determining that the second storage position is not empty;

[0069] If any sub-location in the second storage location is empty, it is determined that the second storage location is empty.

[0070] Through the above storage method, each storage location in the storage space is expanded with multiple sub-locations, which can accommodate more feature information; when the feature information of different application data is the same, the feature information corresponding to each of the different application data can be saved. Accordingly, in the subsequent search process, as long as the feature information corresponding to the search information exists in the storage space, it means that the search information exists, which can ensure the filtering accuracy of the filter. In addition, since the feature information corresponding to multiple different application data is stored in the same storage location, when a certain application data is deleted later, only one set of feature information on the storage location can be deleted accordingly, and other feature information stored in the storage location can also be used to query other search information. While ensuring the query accuracy, it can also realize the flexible deletion of data in the storage space, which is conducive to the subsequent maintenance of the storage space.

[0071] In an implementation of the first aspect, the method further includes:

[0072] After the second characteristic information replaces the third characteristic information, the number of replacements is increased by one;

[0073] If the current number of replacements has reached a preset number, the storage space is expanded to obtain the expanded storage space;

[0074] A third storage position of the feature information stored in the storage space before the expansion in the storage space after the expansion is calculated.

[0075] In the embodiment of the present application, by expanding the storage space, more feature information can be stored and the same feature information previously stored can be retained as much as possible, which is beneficial to improving the subsequent search accuracy and facilitating the subsequent maintenance of the storage space.

[0076] In a second aspect, a backend server is provided, comprising:

[0077] A filter, configured to, upon receiving search information sent by a client, perform a first search according to the search information to obtain a first result; if the first result indicates that the content corresponding to the search information does not exist, return preset information to the client; wherein the preset information is used to indicate that the content corresponding to the search information does not exist;

[0078] A service module is used to send the search information to a search server when the first result indicates that the content corresponding to the search information exists, so that the search server performs a second search based on the search information and returns a second result obtained by the second search; after receiving the second result returned by the search server, the second result is sent to the client.

[0079] In a third aspect, a search system is provided, comprising:

[0080] A search server and a background server as described in the second aspect.

[0081] According to a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by one or more processors, the method according to any one of the first aspects is implemented.

[0082] According to a fifth aspect, a computer program product is provided. When the computer program product is run on a terminal device, the terminal device can implement the method described in any one of the first aspects. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] Figure 1 is a schematic diagram of a scenario of application search provided in an embodiment of the present application;

[0084] Figure 2 It is a schematic diagram of the basic architecture of the ES server provided in the embodiment of the present application;

[0085] Figure 3 This is a schematic diagram of data interaction based on the ES search architecture provided in an embodiment of the present application;

[0086] Figure 4 is a schematic diagram of a search system provided in an embodiment of the present application;

[0087] Figure 5 It is a schematic diagram of the interactive process of the search method provided in the embodiment of the present application;

[0088] Figure 6 It is a flowchart of a method for storing data in a filter provided in an embodiment of the present application;

[0089] Figure 7 is a schematic diagram of a storage space provided in an embodiment of the present application;

[0090] Figure 8 is a schematic diagram of storage space before and after capacity expansion provided in an embodiment of the present application;

[0091] Fig. 9 is a schematic diagram of the first search process provided by an embodiment of the present application;

[0092] Fig.10 It is a schematic diagram of the interactive flow of the search process provided in an embodiment of the present application. DETAILED DESCRIPTION

[0093] In the following description, specific details such as specific system structures and technologies are provided for illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application can also be implemented in other embodiments without these specific details.

[0094] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.

[0095] It should also be understood that in the embodiments of the present application, "one or more" refers to one, two or more than two; "and / or" describes the association relationship of associated objects, indicating that three relationships may exist; for example, A and / or B may represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B may be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.

[0096] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", "fourth", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0097] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0098] Application search refers to the technology of searching for applications (or application software, apps, and applications) based on demand. In some application scenarios, application search is aimed at mobile terminals. Taking a mobile phone as an example, an application with an application search function is installed in the mobile phone, and the user searches and downloads the required application software through the application with the application search function.

[0099] See also Figure 1, is a schematic diagram of a scenario of application search provided by an embodiment of the present application. Figure 1 As shown in (a) of FIG. 1 , an interface 10 of the "Application Center" installed on the mobile phone (for ease of description, in the embodiment of the present application, the application with the application search function is referred to as the "Application Center"). The interface 10 includes a search box 101. When the user does not enter search information, some recommended applications may be displayed in the interface of the "Application Center". As shown in the interface 10, the priority recommended applications are "Application 1" and "Application 2", and the second priority recommended applications are "Application 3", "Application 4" and "Application 5".

[0100] When a user enters the name of the application software or related information they want to search in the "App Center", the App Center can search based on the search information entered by the user and display the search results to the user. Figure 1 Enter "Application 1" in the search box 101 of the interface 10 shown in (a), and the mobile phone displays the following Figure 1 Interface 11 shown in (b) of FIG. The icon of "Application 1" and its "Install" button in interface 11. The application center can also search for some applications related to or similar to the search information and display them to the user. For example, "Application 6", "Application 7" and "Application 8" shown in interface 11 are applications related to "Application 1".

[0101] It should be noted that the search method provided in the embodiment of the present application can also be applied to other search scenarios, such as searching for information in a search website, etc. The embodiment of the present application does not specifically limit the application scenario of the search method.

[0102] Application search services usually need to process massive amounts of application data, which are generally stored in search servers. For example, the search server can use Elasticsearch (ES for short). ES provides a distributed, highly scalable, and highly real-time search and data analysis engine. The ES server can be regarded as a cluster consisting of multiple nodes, and each node can be a server in the cluster. Each node stores the status information of the cluster, such as all node information, all index information, and shard routing information.

[0103] Application data is usually stored in the indexes of the ES server. Each index can be regarded as a collection of documents with similar characteristics. An index can store a large amount of data that exceeds the hardware limitations of a single node. ES can divide a complete index into multiple shards. In this way, a large index can be split into multiple shards and distributed to different nodes to form a distributed search. Shards can be divided into primary shards and replica shards. A primary shard can correspond to multiple replica shards. The primary shard is responsible for processing write requests and storing data, and the replica shard is responsible for storing data. In order to ensure high availability, in some application scenarios, the primary shard and the corresponding replica shard can also be distributed on different nodes.

[0104] For example, see Figure 2 , is a schematic diagram of the basic architecture of the ES server provided in the embodiment of the present application. As an example and not a limitation, Figure 2 As shown in the figure, the ES cluster includes 3 nodes, and the index includes 3 primary shards and 3 replica shards. Among them, primary shard 1 and replica shard 1 are located on node 1, primary shard 2 and replica shard 2 are located on node 2, and primary shard 3 and replica shard 3 are located on node 3, forming a distributed architecture.

[0105] The primary shard can synchronize data with the replica shards on different nodes. Each primary shard corresponds to one replica shard. Figure 2 As shown, primary shard 1 corresponds to replica shard 2, and primary shard 1 synchronizes data with replica shard 2; primary shard 2 corresponds to replica shard 3, and primary shard 2 synchronizes data with replica shard 3; primary shard 3 corresponds to replica shard 1, and primary shard 3 synchronizes data with replica shard 1. In this way, replica shards can improve the fault tolerance of the system. When a shard of a node is damaged or lost, it can be recovered from the replica shard.

[0106] It should be noted that Figure 2 Only the case where the index includes three primary shards and each primary shard corresponds to one replica shard is shown. In other examples, the index may include more or fewer primary shards, and each primary shard may correspond to multiple replica shards. In the embodiment of the present application, the number of shards of the index is not specifically limited.

[0107] It should be noted that Figure 2 Only one example of synchronizing data between the primary shard and the replica shard is shown. In other examples, the primary shard of each node may also synchronize data with each replica shard of a different node. This is not specifically limited in the embodiments of the present application.

[0108] As an example of ES data storage, Node 1 in ES receives a write request from a client, finds the corresponding primary shard 2 according to the routing algorithm, and forwards the write request to Node 2 corresponding to the primary shard for processing, so that Node 2 writes the application data to the primary shard 2; after the primary shard 2 is written, Node 2 sends the write request in parallel to Node 3 where the replica shard 3 corresponding to the primary shard 2 is located, so that Node 3 writes the application data to the replica shard 3. After the replica shard 3 is written, Node 3 feeds back the write result to Node 2, Node 2 feeds back the write result to Node 1, and Node 1 feeds back the write result to the client, completing a complete write process. During the search process, a node receives a query request from the client and sends the query request to the primary and replica shards corresponding to the query request.

[0109] In the ES basic architecture, nodes can be divided into hot nodes and cold nodes. The hardware configuration corresponding to hot nodes is higher than that corresponding to cold nodes. Among them, hot nodes are used to store application data with high search frequency, and cold nodes are used to store application data with low search frequency.

[0110] See also Figure 3 , is a schematic diagram of data interaction based on the ES search architecture provided in the embodiment of the present application. As an example and not a limitation, Figure 3 As shown, the search architecture includes an ES server 31 , and the ES server 31 includes a hot node 311 and a cold node 312 .

[0111] Developers upload application data to the ES server 31 through the development end or remove application data from the ES server 31. After receiving the application data uploaded by the development end, the ES server 31 preferentially stores the application data in the hot node 311; if the application data is not queried for a certain period of time, the hot node 311 migrates the application data to the cold node 312.

[0112] During the search process, the user inputs the search information in the client, and the client sends the search information to the ES server 31. The ES server 31 first searches the hot node 311 for application data related to the search information, and if relevant application data is found, the search result is returned to the client; if relevant application data is not found, the cold node 312 is searched for application data related to the search information, and if relevant application data is found, the search result is returned to the client.

[0113] As time goes by, the number of applications continues to increase, and the data stored in the hot nodes and cold nodes of the ES server is increasing. The performance pressure of the ES server increases accordingly, and the search becomes slower and slower, affecting the user experience.

[0114] Based on this, the embodiment of the present application provides a search architecture. Through the search architecture provided by the embodiment of the present application, the user's search information is first filtered through a filter, and most invalid requests can be intercepted by the filter, thereby effectively reducing the pressure on the ES server and providing users with more efficient search services.

[0115] For some examples, see Figure 4 , is a schematic diagram of a search system provided in an embodiment of the present application. Figure 4 As shown, the search system may include a backend server 41 and an ES server 42 (also referred to as a search server in the present embodiment). For ease of description, in the present embodiment, the server corresponding to the application software (search application) / website (search website) with search function is referred to as a backend server. Figure 1 In the embodiment, the "application center" on the mobile phone is a search software, and its corresponding server can be called a backend server. For another example, in some application scenarios, users use search websites to search for content, and correspondingly, the server corresponding to the search website can be called a backend server.

[0116] like Figure 4 As shown, the backend server 41 may include a filter 411 and an ES service module 412 (also referred to as a service module in the embodiment of the present application). The filter 411 communicates with the client to receive search information sent by the client or to send search results to the client. The ES service module 412 communicates with the ES server 42 to call the search service of the ES server 42. It should be noted that the communication between the ES service module 412 and the ES server 42 may refer to the communication between the ES service module 412 and the nodes in the ES server cluster.

[0117] During the search process, the user enters search information in the search application of the client, and the client sends the search information to the filter 411 in the background server of the search application; the filter 411 performs a first search based on the search information; if the first result of the first search indicates that the search information does not exist, the filter 411 feeds back the first result of the first search to the client; if the first result of the first search indicates that the search information exists, the search information is sent to the ES server 42 via the ES service module 412; the ES server 42 performs a second search based on the search information, and feeds back the second result of the second search (the searched application data) to the ES service module 412, and the ES service module 42 feeds back to the client.

[0118] It should be noted that, in some other examples, the background server 41 may not include a separate ES service module 412, and the filter 411 interacts with the ES server 42. In some other examples, the ES server may also be deployed in the background server.

[0119] For application data maintenance, the filter is synchronized with the application data in the ES server. Figure 4 As shown, the developer uploads application data to the filter through the development end, and at the same time, the developer uploads the application data to the ES server through the development end; the developer deletes application data from the filter through the development end, and at the same time, the developer deletes the application data from the ES server through the development end.

[0120] and Figure 3 Compared to the search architecture shown, Figure 4 In the search system shown, a filter is added before the ES server. Through such a search system, the user's search information is filtered by the filter first. Since 80% of the application data searched by the user does not exist, only 20% of the application data actually exists. Therefore, based on the principle that what the filter determines to exist may exist, and what the filter determines to not exist must not exist, most invalid requests can be intercepted through the filter, thereby effectively reducing the pressure on the ES server and providing users with more efficient search services.

[0121] based on Figure 4 The search system shown in the figure, the embodiment of the present application provides a search method. Figure 5 , is a schematic diagram of the interactive flow of the search method provided in the embodiment of the present application. As an example and not a limitation, Figure 5 As shown, the search method may include the following steps:

[0122] S501, the client sends search information to the backend server.

[0123] Correspondingly, the filter in the background server receives the search information sent by the client.

[0124] like Figure 1 In the application search scenario shown, the search information may be the name, type and / or other descriptive information of the application software. For example, the user wants to search for application AA, which is named AA, belongs to the game category, and is a 3D game. The user can enter "AA", "game", "3D game", etc. in the search box of the search application.

[0125] In the scenario of content search, users can search for content of interest in the search website. The search information can be information such as keywords of the content. For example, if a user wants to search for the latest news reports, the user can enter "today's news" in the search box of the search website.

[0126] S502: The backend server performs a first search based on the search information to obtain a first result.

[0127] In the embodiment of the present application, S502 is executed by the filter in the background server. The specific method for the filter to perform the first search according to the search information can be found in the following Fig. 9 Description in the Examples.

[0128] S503: If the first result indicates that the content corresponding to the search information exists, the background server sends the search information to the search server, so that the search server performs a search according to the search information.

[0129] If the first search result indicates that the content corresponding to the search information does not exist, the backend server sends preset information to the client.

[0130] The preset information is used to indicate that the content corresponding to the search information does not exist. In some examples, the preset information may be text information, a picture, or a logo.

[0131] The search server can be Figure 4 The ES server shown.

[0132] like Figure 4 As shown, when the backend server includes an ES service module, S503 can be implemented by the filter in the backend server sending the search information to the ES service module, and then the ES service module sends the search information to the ES server. When the backend server does not include an ES service module, S503 can be implemented by the filter in the backend server sending the search information to the ES server.

[0133] S504: After receiving the search information sent by the backend server, the search server performs a second search according to the search information, obtains a second result, and returns the second result to the backend server.

[0134] like Figure 4 As shown, when the backend server includes an ES service module, the search server can send the second result to the ES service module in the backend server. When the backend server does not include an ES service module, the search server can send the second result to the filter in the backend server.

[0135] It should be noted that the method by which the search server performs the second search based on the search information may adopt the search method of the search server in the related art, and the embodiment of the present application does not specifically limit the search method of the search server.

[0136] In the embodiment of the present application, the second result is the search result obtained by the search server according to the search information. If the search server does not search for the content corresponding to the search information, the second result may include preset information, and the preset information indicates that the content corresponding to the search information does not exist. The preset information may be text information, a picture or a logo, etc.; if the search server retrieves the content corresponding to the search information, the second result may include the content corresponding to the search information.

[0137] Continuing with the example in S501, in the scenario of application search, the second result may include all applications related to the search information and relevant information of each application (such as application size, number of downloads, detailed description, etc.). In the scenario of content search, the second result may include all content related to the search information. For example, after the user enters "today's news" in the search box of the search website, the third search result may include all news content updated today (such as news summary, news details, etc.).

[0138] S505: After receiving the second result returned by the search server, the backend server sends the second result to the client.

[0139] like Figure 4 As shown, when the backend server includes an ES service module, S505 can be executed by the ES service module in the backend server; or the ES service module forwards the second result to the filter in the backend server, and the filter executes S505. When the backend server does not include an ES server module, S505 can be executed by the filter in the backend server.

[0140] Of course, the backend server may also include a communication module for data interaction with the client. Specifically, the client sends the search information to the communication module, which is forwarded to the filter; when the filter detects that the search information does not exist, the filter sends the first result to the communication module, which is then forwarded to the client; when the filter detects that the search information exists, it sends the search information to the search server, and the search server sends the second result to the communication module, which is then forwarded to the client.

[0141] In the embodiment of the present application, based on the principle that what exists may exist and what does not exist must not exist, when the filter in the background server determines that the search information does not exist, there is no need for the ES server to search again. Accordingly, the filter in the background server sends preset information to the client. Through the above method, most invalid requests can be intercepted by the background server, thereby effectively reducing the pressure on the search server and providing users with more efficient search services.

[0142] For ease of explanation, the process of storing data by the filter in the background server is first introduced.

[0143] See also Figure 6 , is a flowchart of a method for storing data in a filter provided in an embodiment of the present application. As an example and not a limitation, Figure 6 As shown, the method for storing filter data may include the following steps:

[0144] S601, the filter receives application data uploaded by the development end.

[0145] like Figure 4 As shown, the developer uploads application data to the filter in the background server through the development end. Correspondingly, the filter receives the application data uploaded by the development end, and then stores the application data according to steps S602-S604.

[0146] S602: The filter extracts second characteristic information from the application data.

[0147] The data volume of the second characteristic information is smaller than the data volume of the application data.

[0148] The method for extracting the second characteristic information may refer to the description in the following step I-II embodiment.

[0149] S603: The filter calculates a second storage position of the second feature information in the storage space.

[0150] In some storage methods, the second eigenvalues ​​can be randomly stored in the storage space, or the second eigenvalues ​​can be stored in the storage space in order. However, in the above methods, it is difficult to form a mapping relationship between the second eigenvalues ​​and the storage locations, which is not conducive to subsequent information search.

[0151] To solve the above problem, in some embodiments, the method of calculating the second storage location in step S603 may include: calculating a second Hash value of the second feature information; and calculating the second storage location according to the second Hash value and the storage capacity of the storage space.

[0152] In some implementations of calculating the second hash value, a secure hash algorithm (SHA) can be used to calculate the second hash value. SHA includes algorithms such as SHA-1, SHA-224, SHA-256, SHA-384, and SHA-512. Among them, SHA-256 is a safe and reliable encryption algorithm that can provide high collision resistance, that is, the probability of different input data generating the same hash value is low. In an embodiment of the present application, SHA-256 can be used to calculate the second hash value of the second feature information, thereby effectively reducing the collision probability of the feature information, which is conducive to improving the accuracy of subsequent information search.

[0153] In the embodiment of the present application, the storage method may be: performing a modulo operation according to the second hash value and the storage capacity of the storage space to obtain the second storage location.

[0154] In the embodiments of the present application, the storage capacity of the storage space may refer to the number of storage locations in the storage space. Of course, the storage capacity of the storage space may also be understood as the amount of data that can be stored in the storage space. Under this understanding, the number of storage locations in the storage space may be calculated based on the unit data amount of data that can be stored in each storage location and the storage capacity of the storage space.

[0155] Among them, the remainder operation is also called the modulus operation, which refers to calculating the remainder after dividing two numbers. In the hash algorithm, the remainder operation refers to dividing the hash value by the storage capacity of the storage space, and the remainder obtained is used as the storage location. For example, assuming that the second hash value is 1025, the number of storage locations in the storage space is 1024, and the remainder of the second hash value in the storage capacity is 1, then the first storage location of the storage space is used to store the second feature information corresponding to the second hash value 1025.

[0156] The modulo operation method can quickly and simply determine the storage location and can ensure that the calculated storage location falls within the range of the storage space.

[0157] In some implementations, the storage capacity Q of the storage space may be set to an exponent of 2. In this manner, a modulo operation is performed based on the second hash value and the storage capacity of the storage space, which is equivalent to taking the last Q bits of the second hash value. This calculation method is simpler and helps improve processing efficiency.

[0158] In an embodiment of the present application, the storage location of the second characteristic information is determined based on the second hash value of the second characteristic information and the storage capacity of the storage space, so that a clear mapping relationship can be formed between the characteristic information of the application data and the storage location of the storage space, which is beneficial to subsequent information search.

[0159] S604: The filter stores the second feature information in a second storage location.

[0160] Since the second characteristic information can reflect the characteristics of the application data and the data volume is relatively small, storing the second characteristic information is conducive to saving storage space compared to storing the application data.

[0161] In some application scenarios, there may be a situation where the feature information of different application data is the same. In this case, the storage locations corresponding to the feature information of different application data are the same. For example, the second feature information of application data XX extracted according to step S602 is the same as the second feature information of application data YY extracted, and the second hash value calculated based on the same second feature information is also the same, therefore, the storage location determined based on the same second hash value is also the same.

[0162] One solution is that if data is already stored in the storage location calculated based on the second characteristic information, the current second characteristic information will no longer be stored. In this way, it is equivalent to giving up storing the characteristic information of the latest application data. This method is prone to the following problems: when searching for subsequent information, it is impossible to determine whether the characteristic information at the storage location is the characteristic information corresponding to the current search information, which will reduce the filtering accuracy of the filter; in addition, since the characteristic information at a certain storage location may correspond to multiple sets of application data, it is not convenient to delete the characteristic information in the storage location later, which is not conducive to the maintenance of storage space.

[0163] To solve the above problem, in some embodiments, step S604 may include:

[0164] If the second storage location is empty, storing the second characteristic information in the second storage location;

[0165] If the second storage location is not empty, the second characteristic information replaces the third characteristic information, wherein the third characteristic information is the data stored in the second storage location.

[0166] Through the above method, it is possible to ensure that the second characteristic information of the latest application data is saved.

[0167] In order to store as much identical feature information as possible, an embodiment of the present application provides a structure of a storage space. In the embodiment of the present application, the storage space may include multiple main positions and at least one sub-position under each main position. Each main position is the same as the storage position corresponding to the sub-position under the main position. Through this structure of the storage space, it is equivalent to expanding each storage position in the storage space to facilitate the storage of multiple groups of identical feature information.

[0168] See also Figure 7 , is a schematic diagram of a storage space provided in an embodiment of the present application. Figure 7 As shown, the storage space includes a plurality of main locations 71 and four sub-locations 72 under each main location. For ease of explanation, Figure 7 Only the situation of 4 sub-positions under one main position is shown.

[0169] In some implementations, the data in the storage space may be stored in the form of a matrix array. Exemplarily, the main position may correspond to a row in the matrix array, and the sub-position may correspond to a column in the matrix array.

[0170] It should be noted that in the embodiment of the present application, the number of main positions and the number of sub-positions under each main position are not specifically limited. However, it is understandable that the more sub-positions under each main position, the more identical feature information can be carried, but the higher the capacity requirement for the storage space. Therefore, the number of sub-positions under each main position can be set according to actual needs.

[0171] based on Figure 7 The structure of the storage space shown in FIG. 1 , an implementation of step S604 may include:

[0172] If the main location in the second storage location is empty, storing the second feature information in the main location of the second storage location;

[0173] If the main location in the second storage location is not empty, determining whether the sub-location in the second storage location is empty;

[0174] If any sub-position in the second storage position is empty, the second feature information is stored in the free sub-position in the second storage position;

[0175] If each sub-location in the second storage location is not empty, it is determined that the second storage location is not empty.

[0176] The process of judging whether a sub-position is empty may be judged in sequence according to the arrangement order of the sub-positions, or may be judged in parallel without following the arrangement order.

[0177] by Figure 7 Taking the storage space shown as an example, assuming that the calculated second storage position is P, first determine whether the main position P0 on the second storage position P is empty; if P0 is empty, store the second feature information in P0; if P0 is not empty, determine whether the first sub-position P1 of the second storage position P is empty; if P1 is empty, store the second feature information in P1; if P1 is not empty, then determine whether the second sub-position P2 of the second storage position P is empty; if P2 is empty, store the second feature information in P2; if P2 is not empty, then determine whether the third sub-position P3 of the second storage position P is empty; if P3 is empty, store the second feature information in P3; if P3 is not empty, then determine whether the fourth sub-position P4 of the second storage position P is empty; if P4 is empty, store the second feature information in P4; if P4 is not empty, determine that the second storage position P is not empty.

[0178] In the embodiment of the present application, if it is determined that the second storage location is not empty, the second characteristic information replaces the third characteristic information as described in the above steps. In some implementations, the second characteristic information can be randomly replaced with any set of third characteristic information on the second storage location, such as the third characteristic information on the main location or the third characteristic information on any sub-location. In other implementations, the second characteristic information can also be replaced in the order of first the main location and then the sub-location.

[0179] Through the above storage method, each storage location in the storage space is expanded with multiple sub-locations, which can accommodate more feature information; when the feature information of different application data is the same, the feature information corresponding to each of the different application data can be saved. Accordingly, in the subsequent search process, as long as the feature information corresponding to the search information exists in the storage space, it means that the search information exists, which can ensure the filtering accuracy of the filter. In addition, since the feature information corresponding to multiple different application data is stored in the same storage location, when a certain application data is deleted later, only one set of feature information on the storage location can be deleted accordingly, and other feature information stored in the storage location can also be used to query other search information. While ensuring the query accuracy, it can also realize the flexible deletion of data in the storage space, which is conducive to the subsequent maintenance of the storage space.

[0180] In order to ensure that previously stored feature information (such as the third feature information replaced by the second feature information) is not deleted, in some embodiments, the method may further include:

[0181] After the second characteristic information replaces the third characteristic information, the number of replacements is increased by one;

[0182] If the current number of replacements does not reach the preset number, the third characteristic information is used to replace the fourth characteristic information, and the number of replacements is increased by one, wherein the fourth characteristic information is the data stored in the second storage location except the third characteristic information;

[0183] If the current number of replacements has reached a preset number, the storage space is expanded to obtain an expanded storage space; and a third storage position of the feature information stored in the storage space before the expansion in the storage space after the expansion is calculated.

[0184] For example, continue with the above Figure 7For example, assume that the preset number is 5. When the second storage position P is not empty, first replace the characteristic information (third characteristic information) at the main position P0 of the second storage position P with the second characteristic information, and the number of replacements is 1 at this time, which does not reach the preset number; replace the characteristic information (third characteristic information) at the main position P0 with the characteristic information (fourth characteristic information) at the sub-position P1 of the second storage position P, and the number of replacements is 2 at this time, which does not reach the preset number; replace the characteristic information (fourth characteristic information) at the sub-position P2 of the second storage position P with the characteristic information at the sub-position P2, and the number of replacements is 3 at this time, which does not reach the preset number; replace the characteristic information at the sub-position P3 of the second storage position P with the characteristic information at the sub-position P2, and the number of replacements is 4 at this time, which does not reach the preset number; replace the characteristic information at the sub-position P4 of the second storage position P with the characteristic information at the sub-position P3, and the number of replacements is 5 at this time, which reaches the preset number; expand the storage space to obtain the expanded storage space; calculate the third storage position of the characteristic information at the sub-position P4 in the expanded storage space, and store the characteristic information at the sub-position P4 of the original second storage position P in the new third storage position.

[0185] The preset number of times can be preset according to actual needs. If the preset number of times is small, the probability of the feature information being saved may be reduced; if the preset number of times is large, although the probability of the feature information being saved can be increased, it will also increase the processing time of the algorithm.

[0186] In the embodiment of the present application, the preset number of times can be set according to the number of main positions and sub-positions corresponding to each storage position in the storage space. Figure 7 For example, each storage location includes 1 main location and 4 sub-locations, and the preset number can be set to 5 (the sum of the number of main locations and sub-locations). That is, when the replacement reaches 5 times, it means that the current storage location is full, and its main location and sub-location have stored data.

[0187] As described in the above step S603 embodiment, the storage location corresponding to the characteristic information is calculated by performing a modulo operation on the hash value of the characteristic information and the storage capacity of the storage space to obtain the storage location corresponding to the characteristic information. Since the storage space is expanded, the storage capacity of the expanded storage space changes, and accordingly, the storage location corresponding to the characteristic information in the expanded storage space also changes accordingly. Therefore, it is necessary to recalculate the storage location corresponding to each of the characteristic information stored in the storage space in the expanded storage space.

[0188] For some examples, see Figure 8 , is a schematic diagram of the storage space before and after the expansion provided by the embodiment of the present application. Figure 8As shown in (a) in FIG, it is the storage space before expansion, and the number of storage locations it contains is 1024. Figure 8 As shown in (b) of FIG. 1 , the storage space after expansion contains 2048 storage locations. Assuming that the hash value of the characteristic information ZZ is 1025, according to the calculation method of the storage location in S603, the storage location corresponding to 1025 in the storage space before expansion is 1, and the storage location corresponding to 1025 in the storage space after expansion is 1025. Figure 8 As shown, the storage position corresponding to the characteristic information ZZ in the storage space before the expansion is different from the storage position corresponding to the characteristic information ZZ in the storage space after the expansion.

[0189] In the embodiments of the present application, expanding the storage space refers to increasing the number of storage locations in the storage space. In some implementations, the number of main locations can be increased, but the number of sub-locations under each main location remains unchanged. In other implementations, the number of sub-locations under each main location can be increased, but the number of main locations remains unchanged. In other implementations, the number of main locations and the number of sub-locations under each main location can be increased at the same time. The storage space can be expanded according to actual needs.

[0190] In the embodiment of the present application, by expanding the storage space, more feature information can be stored and the same feature information previously stored can be retained as much as possible, which is beneficial to improving the subsequent search accuracy and facilitating the subsequent maintenance of the storage space.

[0191] The method of extracting the second feature information in step S602 is described below.

[0192] In some embodiments, the method of extracting the second feature information in step S602 may include:

[0193] I. Perform hash processing according to the character position of each character in the application data in the character string corresponding to the application data to obtain a second character hash value.

[0194] II. Extract the second feature information according to the second character hash value.

[0195] Since the order of characters in a sentence is closely related to the semantics of the sentence, the hash processing in the embodiment of the present application is equivalent to extracting the feature information of the application data through the character positions of the characters in the application data, which can reflect the semantic features of the application data to a certain extent, thereby facilitating improving the search accuracy of subsequent information searches.

[0196] In some embodiments, the shifting process in step I may include:

[0197] For the first character in the application data, calculate the hash value of the character position of the first character in the string corresponding to the application data to obtain the second character hash value of the first character; update the initial memory value according to the second character hash value of the first character; update the initial shift bit number according to the first preset value to obtain the shift bit number corresponding to the first character.

[0198] For the jth character in the application data, calculate the hash value of the character position of the ith character in the string corresponding to the application data to obtain the second character hash value of the ith character;

[0199] Shift the second character hash value of the j-th character according to the number of shift bits corresponding to the j-1-th character to obtain the second character hash value of the j-th character after the shift, where j is an integer greater than 1;

[0200] Update the fifth memory value according to the second character hash value of the j-th character after the shift to obtain a sixth memory value, wherein the fifth memory value is used to represent the second character hash value of the first j-1 characters after the shift, and the sixth memory value is used to represent the second character hash value of the first j characters after the shift;

[0201] The number of shift bits corresponding to the j-1th character is updated according to the first preset value to obtain the number of shift bits corresponding to the jth character.

[0202] In some examples, the initial shift number can be set to 0. The initial memory value can be 0000.

[0203] In an embodiment of the present application, the first preset value is determined according to the data interval to which the second character hash value of the j-th character before the shift belongs. Exemplarily, three data intervals are set, wherein interval 1 indicates that the second character hash value is less than 128, interval 2 indicates that the second character hash value is greater than 128 and less than 2048, and interval 3 indicates that the second character hash value is greater than 2048. Each data interval corresponds to a first preset value, for example, the first preset value corresponding to interval 1 is 8, the first preset value corresponding to interval 2 is 16, and the first preset value corresponding to interval 3 is 24. As the character position increases, the first preset value also increases, which can effectively reduce the conflict between the hash values ​​of each character in the memory value.

[0204] In some implementations of updating the fifth memory value, the method may include performing an OR operation on the second character hash value of the j-th character after the shift and the fifth memory value, and obtaining the operation result as the sixth memory value.

[0205] It should be noted that in the embodiment of the present application, the fifth memory value refers to the memory value before the update, and the sixth memory value refers to the memory value after the update. For example, when j=2, the fifth memory value refers to the memory value obtained after the initial memory value is updated according to the second character hash value of the first character; when j=3, the fifth memory value refers to the memory value obtained after the memory value is updated according to the second character hash value of the second character after the shift (i.e., the sixth memory value when j=2). And so on.

[0206] Exemplarily, taking the data interval and the first preset value of the above example, the second character hash value of the first character is 0001, and the second character hash value of the character does not need to be shifted; the second character hash value 0001 of the first character is ORed with the initial memory value 0000, and the resulting memory value is 0001. The calculated number of shift bits corresponding to the first character is 8; the second character hash value 0010 of the second character is shifted left by 8 bits to obtain the second character hash value 0010 0000 0000 of the shifted second character; the second character hash value of the shifted second character is ORed with the memory value 0001 (the fifth memory value) to obtain the updated memory value 0010 0000 0001 (the sixth memory value). The calculated number of shift bits corresponding to the second character is 16. And so on.

[0207] It should be noted that the above is only an example of updating the memory value, and does not specifically limit the calculation method of the hash value and the obtained hash value.

[0208] In some implementations, after obtaining the number of shift bits corresponding to the j-th character, the method may further include:

[0209] If the number of shift bits corresponding to the j-th character is less than or equal to the shift threshold, continue to shift the second character hash value of the j+1-th character according to the number of shift bits corresponding to the j-th character;

[0210] If the number of shift bits corresponding to the jth character is greater than the shift threshold, data is obtained from the sixth memory value according to the shift threshold to obtain the seventh memory value; the number of shift bits corresponding to the jth character is reduced according to the shift threshold to obtain the updated number of shift bits corresponding to the jth character.

[0211] In the embodiment of the present application, when the number of shift bits exceeds the shift threshold, the number of shift bits is reduced, which can effectively reduce the occurrence of excessive data bits after the shift processing due to excessive shift bits; in addition, obtaining part of the data from the current memory value can also effectively control the length of the memory value. In the above manner, the length of the memory value can be effectively controlled to ensure that the memory value does not exceed the memory space.

[0212] Optionally, one way to obtain data from the sixth memory value according to the shift threshold is to obtain the values ​​of M digits in the sixth memory value in order from high to low, so as to obtain the seventh memory value.

[0213] Wherein, the M is determined according to the shift threshold. In some implementations, M may be set equal to the shift threshold. For example, assuming that the shift threshold is 32, M=32, and the current sixth memory value is 0101 0000 0100 0000 0011 0000 001000000001, then in the order of the number of digits from high to low, the values ​​of the 32 high digits are taken, and the seventh memory value obtained is 0101 0000 0100 0000 0011 0000 0010 0000.

[0214] It can be seen from the above embodiments that the high-order numerical value is equivalent to the hash value of the character position of the later character in the string, and the low-order numerical value is equivalent to the hash value of the character position of the earlier character in the string. In the above method, the high-order numerical value is retained and the low-order numerical value is filtered out, which is equivalent to retaining the hash value of the character position of the later character in the string as much as possible. In some application scenarios, the characters at the later position in the string are mostly searched keywords. Therefore, the memory value obtained by the above method can retain more characteristic information of the keywords in the string.

[0215] In some embodiments, extracting the second feature information according to the second character hash value in step II may include:

[0216] The eighth memory value is magnified to obtain a processed eighth memory value, wherein the eighth memory value is used to represent a second character hash value of N characters after the shift, where N is the total number of characters in the application data; and the second feature information is calculated according to the eighth memory value.

[0217] It should be noted that if the shift threshold is not exceeded during the above shift processing, that is, the low-order values ​​in the memory value are not filtered out, then the eighth memory value includes the hash values ​​of the character positions corresponding to the N characters in the application data. If the shift threshold is exceeded during the above shift processing, that is, the low-order values ​​in the memory value are filtered, then the eighth memory value includes the hash values ​​of the character positions corresponding to the last n characters in the application data, but the hash values ​​of the character positions of the n characters can be used to represent the characteristic information of the character positions of the N characters in the application data after the shift. Wherein, n is a positive integer less than N.

[0218] In some implementations, performing magnification processing on the eighth memory value may include: calculating the complement of the eighth memory value, shifting the complement of the eighth memory value left by L bits, and obtaining a processed eighth memory value. Wherein L is a preset value. For example, assuming that the eighth memory value is 0010 0000, L=8, then the processed eighth memory value is 0010 0000 0000 0000.

[0219] In the embodiment of the present application, by amplifying the memory value, the number of bits of the processed memory value is increased, which can reduce data conflicts in subsequent calculation processes.

[0220] In some implementations, the second characteristic information is calculated according to the eighth memory value as follows:

[0221] An XOR operation is performed on the second preset value and the eighth memory value to obtain an XOR value; and the second feature information is calculated based on the XOR value.

[0222] The second preset value may be preset by a developer. In some implementations, the second preset value may be set to a prime number. Since a prime number cannot be decomposed any further, the reliability of subsequent calculations is improved.

[0223] In the XOR operation, if the two values ​​are the same, the XOR result is 0; if the two values ​​are different, the XOR result is 1. Exemplarily, assuming that the second preset value is 0000 0111 and the eighth memory value is 0010 0000, the XOR value obtained by the XOR operation of the two is 0010 0111. Assuming that the second preset value is 0000 0111 and the eighth memory value is 0000 0111, the XOR value obtained by the XOR operation of the two is 0000 0000.

[0224] Through the XOR operation, it is equivalent to detecting whether the eighth memory value is the same as the second preset value. If the eighth memory value is the same as the second preset value, the XOR value is 0. When two sets of application data are the same, and their corresponding eighth memory values ​​are the same as the second preset value, the XOR value can be used to determine that the two sets of application data are the same. Correspondingly, the hash values ​​corresponding to the two sets of application data are the same and the feature information is also the same. In this case, one of the two sets of application data can be selected for storage. In this way, the probability of duplicate storage can be reduced.

[0225] Since the second character hash value of each character in the application data can represent the characteristic information of the character position in the character string, the eighth memory value is equivalent to containing the characteristic information of the character positions of all characters in the application data. Therefore, in the embodiment of the present application, calculating the second characteristic information based on the XOR value is equivalent to extracting characteristic information based on the character positions of all characters in the application data. The characteristic information extracted in this way can represent the positional relationship between the characters in the application data, thereby representing the semantics corresponding to the application data.

[0226] In some implementations, after obtaining the number of shift bits corresponding to the j-th character, if the number of shift bits corresponding to the j-th character is greater than a shift threshold, amplification processing is performed on the second preset value to obtain the processed second preset value;

[0227] Correspondingly, the XOR operation performed on the second preset value and the eighth memory value is:

[0228] If the number of shift bits corresponding to the j-th character is greater than the shift threshold, an XOR operation is performed on the processed second preset value and the eighth memory value to obtain an XOR value;

[0229] If the number of shift bits corresponding to the j-th character is less than or equal to the shift threshold, there is no need to perform amplification processing on the second preset value. Accordingly, an XOR operation is performed on the second preset value (i.e., the second preset value without amplification processing) and the eighth memory value to obtain an XOR value.

[0230] In the embodiment of the present application, when the number of shift bits is greater than the shift threshold, it means that the number of bits in the current memory value is relatively large. In this case, the second preset value is magnified, which is equivalent to increasing the number of bits of the second preset value and increasing the complexity of the second preset value. In this way, data conflicts in subsequent calculations can be reduced.

[0231] As an example of the method for extracting the second feature information in S602, the code implementation logic is as follows:

[0232]

[0233]

[0234] Based on the above Figure 6 The process of storing data in the filter is described below. The first search method in S502 is described below. Fig. 9 , is a schematic diagram of the first search process provided in the embodiment of the present application. As an example and not a limitation, Fig. 9 As shown, the first search may include the following steps:

[0235] S901, extracting first characteristic information from the search information.

[0236] In some embodiments, step S901 may include:

[0237] A shift process is performed according to the character position of each character in the search information in the character string corresponding to the search information to obtain a processed character position; and first feature information is extracted according to the processed character position.

[0238] As an implementation method of shift processing, it includes:

[0239] For the i-th character in the search information, calculate the hash value of the character position of the first character in the character string corresponding to the search information to obtain the first character hash value of the first character;

[0240] Shift the first character hash value of the i-th character according to the number of shift bits corresponding to the i-1-th character to obtain the first character hash value of the i-th character after the shift, where i is an integer greater than 1;

[0241] Update a first memory value according to the first character value of the i-th character after the shift to obtain a second memory value, wherein the first memory value is used to represent the first character hash value of the first i-1 characters after the shift, and the second memory value is used to represent the first character hash value of the first i characters after the shift;

[0242] The number of shift bits corresponding to the i-1th character is updated according to a first preset value to obtain the number of shift bits corresponding to the i-th character, wherein the first preset value is determined according to the data interval to which the first character hash value of the i-th character before shifting belongs.

[0243] In some embodiments, the method further comprises:

[0244] After obtaining the number of shift bits corresponding to the i-th character, if the number of shift bits corresponding to the i-th character is greater than the shift threshold, obtaining data from the second memory value according to the shift threshold to obtain a third memory value;

[0245] The number of shift bits corresponding to the i-th character is reduced according to the shift threshold to obtain the updated number of shift bits corresponding to the i-th character.

[0246] In one implementation of obtaining the third memory value, the values ​​of M digits in the second memory value are obtained in order of the number of digits from high to low to obtain the third memory value, wherein M is determined according to a shift threshold.

[0247] In some implementations, extracting the first feature information according to the first character position may include:

[0248] The fourth memory value is magnified to obtain a processed fourth memory value, wherein the fourth memory value is used to represent the first character hash value of N characters after the shift, where N is the total number of characters in the search information; and the first feature information is calculated according to the fourth memory value.

[0249] In some implementations, calculating the first characteristic information according to the fourth memory value may include: performing an XOR operation on the second preset value and the fourth memory value to obtain an XOR value; and calculating the first characteristic information according to the XOR value.

[0250] In some implementations, the method further includes:

[0251] After obtaining the number of shift bits corresponding to the i-th character, if the number of shift bits corresponding to the i-th character is greater than the shift threshold, the second preset value is magnified to obtain the processed second preset value. Accordingly, performing an XOR operation on the second preset value and the fourth memory value to obtain an XOR value includes: performing an XOR operation on the current second preset value and the fourth memory value to obtain an XOR value.

[0252] The method for extracting the first characteristic information in the search information in step S901 is similar to the method described above. Figure 6 The method of extracting the second characteristic information in the application data in step S602 in the embodiment is the same. For details, please refer to the description in the embodiment of S602, which will not be repeated here.

[0253] It should be noted that the method used by the filter to extract the second characteristic information of the application data in the process of storing data is the same as the method used to extract the first characteristic information of the search information in the process of searching, and the parameters involved in the method are also the same. For example, the parameters include: the first preset value, the shift threshold, the M value involved in the shift processing in the method of extracting the characteristic information, and the second preset value involved in calculating the characteristic information.

[0254] S902: Obtain a first storage location of first feature information in the storage space.

[0255] In some embodiments, S902 may include:

[0256] Calculate a first Hash value of the first feature information; and calculate the first storage location according to the first Hash value and the storage capacity of the storage space.

[0257] In one implementation, a modulo operation is performed based on the first hash value and the storage capacity of the storage space to obtain the first storage location.

[0258] The method of obtaining the first storage location of the first feature information in the storage space in step S902 is the same as the above Figure 6 The method of calculating the second storage position of the second characteristic information in the storage space in step S603 in the embodiment is the same. For details, please refer to the description in the embodiment of S603, which will not be repeated here.

[0259] S903: If the first characteristic information is stored in the first storage location, the first result indicates that the content corresponding to the search information exists.

[0260] S904: If the first characteristic information is not stored in the first storage location, the first result indicates that the content corresponding to the search information does not exist.

[0261] For example, see Fig.10 , is a schematic diagram of the interactive flow of the search process provided by the embodiment of the present application. Fig.10 As shown, the search process may include the following steps:

[0262] S1001, the client sends search information to the backend server.

[0263] S1002: The filter in the background server receives the search information and extracts the first characteristic information in the search information.

[0264] S1003: The filter in the background server obtains a first storage position of the first feature information in the storage space.

[0265] Steps S1002-S1003 are the same as the above steps S901-S902, and for details, please refer to the description in the above embodiment of steps S901-S902.

[0266] S1004: The filter in the background server determines whether the first storage stores the first characteristic information.

[0267] S1005: If the first storage location does not store the first characteristic information, the filter in the background server returns preset information to the client.

[0268] The preset information is used to indicate that the content corresponding to the search information does not exist.

[0269] S1006: If the first storage location stores the first characteristic information, the filter in the background server sends the search information to the search server through the service module in the background server.

[0270] S1007: After receiving the search information, the search server performs a second search.

[0271] S1008: The search server returns the second result of the second search to the client through the service module in the background server.

[0272] In the embodiment of the present application, since the data volume of the first characteristic information is usually smaller than the data volume of the search information, the method of searching by the filter according to the first characteristic information is more efficient than the method of searching by the ES server according to the search information. Using the filter to perform the first search on the search information can intercept most invalid requests, which can not only improve the search efficiency, but also help reduce the search pressure of the ES server and provide users with more efficient search services.

[0273] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0274] The embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.

[0275] The embodiment of the present application also provides a computer program product. When the computer program product is executed on a terminal device, the terminal device can implement the steps in the above-mentioned method embodiments.

[0276] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device capable of carrying the computer program code to the first device, a recording medium, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electric carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.

[0277] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0278] Those of ordinary skill in the art will appreciate that the units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application. Finally, it should be noted that the above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto, and any changes or substitutions within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of the present application shall be based on the scope of protection of the claims.

Claims

1. A search method, It is characterized in that Applied to a backend server, the method comprises: Get the first application name sent by the client; Detecting a plurality of first storage locations corresponding to the first application name in a storage space corresponding to the background server, the storage space comprising a plurality of storage locations, the storage locations being used to store characteristic information corresponding to the application name, the plurality of storage locations comprising the plurality of first storage locations; wherein the storage locations in the storage space are determined according to the characteristic information corresponding to the application name and the storage capacity of the storage space; If feature information corresponding to the first application name exists in any of the first storage locations, returning preset information to the client, the preset information indicating that the background server includes an application corresponding to the first application name; If the characteristic information corresponding to the first application name does not exist in each of the first storage locations, the first application name is sent to the search server, so that the search server searches for the first application name in the application data stored in the search server and returns the search results to the client; wherein the data stored in the storage space corresponding to the background server is consistent with the application data stored in the search server.

2. The method according to claim 1, It is characterized in that The detecting a plurality of first storage locations corresponding to the first application name in the storage space corresponding to the background server includes: Extracting first characteristic information from the first application name; A plurality of first storage locations are detected in the storage space according to the first characteristic information.

3. The method according to claim 2, It is characterized in that The extracting the first characteristic information from the first application name includes: Performing hash processing according to the character position of each character in the first application name in the character string corresponding to the first application name to obtain a first character hash value; The first feature information is extracted according to the first character hash value.

4. The method according to claim 3, It is characterized in that The performing hash processing according to the character position of each character in the first application name in the character string corresponding to the first application name to obtain a first character hash value includes: For the i-th character in the first application name, calculate the hash value of the character position of the i-th character in the character string corresponding to the first application name to obtain the first character hash value of the i-th character; Shift the first character hash value of the i-th character according to the number of shift bits corresponding to the i-1-th character to obtain the first character hash value of the i-th character after the shift, where i is an integer greater than 1; Update the first memory value according to the first character hash value of the i-th character after the shift to obtain a second memory value, wherein the first memory value is used to represent the first character hash value of the first i-1 characters after the shift, and the second memory value is used to represent the first character hash value of the first i characters after the shift; The number of shift bits corresponding to the i-1th character is updated according to a first preset value to obtain the number of shift bits corresponding to the i-th character, wherein the first preset value is determined according to the data interval to which the first character hash value of the i-th character before shifting belongs.

5. The method according to claim 4, It is characterized in that The method further comprises: After obtaining the number of shift bits corresponding to the i-th character, if the number of shift bits corresponding to the i-th character is greater than a shift threshold, obtaining data from the second memory value according to the shift threshold to obtain a third memory value; The number of shift bits corresponding to the i-th character is reduced according to the shift threshold to obtain an updated number of shift bits corresponding to the i-th character.

6. The method according to claim 5, It is characterized in that The acquiring data from the second memory value according to the shift threshold to obtain a third memory value includes: According to the order of the number of digits from high to low, the values ​​of M digits in the second memory value are obtained to obtain the third memory value, wherein M is determined according to the shift threshold.

7. The method according to claim 4, It is characterized in that The extracting the first feature information according to the first character hash value includes: Perform magnification processing on the fourth memory value to obtain the processed fourth memory value, wherein the fourth memory value is used to represent the first character hash value of the N characters after the shift, and N is the total number of characters in the first application name; The first feature information is calculated according to the fourth memory value.

8. The method according to claim 7, It is characterized in that The calculating the first feature information according to the fourth memory value includes: Perform an XOR operation on the second preset value and the fourth memory value to obtain an XOR value; The first feature information is calculated according to the XOR value.

9. The method according to claim 8, It is characterized in that The method further comprises: After obtaining the number of shift bits corresponding to the i-th character, if the number of shift bits corresponding to the i-th character is greater than the shift threshold, amplifying the second preset value to obtain the processed second preset value; The performing an XOR operation on the second preset value and the fourth memory value to obtain an XOR value includes: If the number of shift bits corresponding to the i-th character is greater than the shift threshold, an XOR operation is performed based on the processed second preset value and the fourth memory value to obtain an XOR value.

10. The method according to claim 2, It is characterized in that The plurality of first storage locations include a main location and at least one sub-location corresponding to the main location; wherein the storage location corresponding to the main location and the sub-location corresponding to the main location are the same; The detecting a plurality of the first storage locations in the storage space according to the first characteristic information includes: Calculating a first Hash value of the first feature information; Performing a modulo operation based on the first hash value and the storage capacity of the storage space to obtain a main position among the plurality of first storage positions; A sub-location among the plurality of first storage locations is obtained according to a main location among the plurality of first storage locations.

11. The method according to any one of claims 1 to 10, It is characterized in that The method further comprises: Receive the second application name sent by the development end; Extracting second characteristic information from the second application name; Calculate a plurality of second storage locations of the second application name in the storage space corresponding to the background server according to the second feature information; The second feature information is stored according to a plurality of the second storage locations.

12. The method according to claim 11, It is characterized in that The storing the second feature information according to the plurality of second storage locations comprises: If there is an empty second storage location, storing the second feature information in any empty second storage location; If each of the second storage locations is not empty, the second characteristic information replaces the third characteristic information, wherein the third characteristic information is data stored in any one of the second storage locations.

13. The method according to claim 12, It is characterized in that The plurality of second storage locations include a main location and at least one sub-location corresponding to the main location; The method further comprises: If the main location among the plurality of second storage locations is not empty, determining whether the sub-location among the plurality of second storage locations is empty; If each sub-location in the plurality of second storage locations is not empty, it is determined that each of the second storage locations is not empty.

14. The method according to claim 12, It is characterized in that The method further comprises: After the second characteristic information replaces the third characteristic information, the number of replacements is increased by one; If the current number of replacements has reached a preset number, the storage space is expanded to obtain the expanded storage space; A third storage position of the feature information stored in the storage space before the expansion in the storage space after the expansion is calculated.

15. The method according to any one of claims 1 to 14, It is characterized in that The method further comprises: Obtaining a deletion instruction sent by the development end, wherein the deletion instruction includes a name of the third application to be deleted; Detecting a plurality of fourth storage locations corresponding to the third application name in the storage space corresponding to the background server; If there is a fourth storage location that meets the preset condition, the data in any fourth storage location that meets the preset condition is deleted; wherein the fourth storage location that meets the preset condition includes feature information corresponding to the third application name.

16. A backend server, It is characterized in that The backend server includes a filter and service module; The filter obtains the first application name sent by the client; The filter detects a plurality of first storage locations corresponding to the first application name in the storage space corresponding to the background server; the storage space includes a plurality of storage locations, the storage locations are used to store characteristic information of the application name, and the plurality of storage locations include a plurality of the first storage locations; wherein the storage locations in the storage space are determined according to the characteristic information corresponding to the application name and the storage capacity of the storage space; If data exists in any of the first storage locations, the filter returns preset information to the client, where the preset information indicates that the background server includes an application corresponding to the first application name; If no data exists in each of the first storage locations, the server module sends the first application name to the search server, so that the search server searches for the first application name in the application data stored in the search server and returns the search results to the client; wherein the data stored in the storage space corresponding to the background server is consistent with the application data stored in the search server.

17. A search system, It is characterized in that include: A search server and a backend server as claimed in claim 16.

18. A search method, It is characterized in that Applied to a system including a client, a backend server and a search server, the method includes: The client displays a first interface of the application center, wherein the first interface includes a search box; The client receives a search instruction input by a user in the search box, where the search instruction includes a first application name; The client sends the first application name to the backend server; In response to receiving the first application name sent by the client, the backend server detects a plurality of first storage locations corresponding to the first application name in a storage space corresponding to the backend server, the storage space comprising a plurality of storage locations, the storage locations being used to store feature information corresponding to the application name, the plurality of storage locations comprising the plurality of first storage locations; If feature information corresponding to the first application name exists in any of the first storage locations, the background server returns preset information to the client, where the preset information indicates that the background server includes an application corresponding to the first application name; If the feature information corresponding to the first application name does not exist in each of the first storage locations, the background server sends the first application name to the search server; In response to receiving the first application name sent by the background server, the search server searches for the first application name in the application data stored in the search server and sends the search results to the client; wherein the data stored in the storage space corresponding to the background server is consistent with the application data stored in the search server.

19. The search method according to claim 18, It is characterized in that The method further comprises: In response to receiving the preset information sent by the backend server, the client displays a second interface of the application center, where the second interface includes an installation control corresponding to the first application name; In response to a user operating an installation control corresponding to the first application name on the second interface, the client sends a download request to the backend server, where the download request includes the first application name; In response to receiving the download request sent by the client, the backend server sends the application corresponding to the first application name to the client; In response to receiving the application corresponding to the first application name sent by the background server, the client installs the application corresponding to the first application name.

Citation Information

Patent Citations

  • Searching method and device, electronic equipment and storage medium

    CN108090153A

  • Resource search result display method and device and computer readable storage medium

    CN110134886A