Method of searching data and related device
The cloud management platform optimizes search quality and efficiency by using an index table with preset terms and media data to handle diverse data types, addressing inefficiencies in traditional inverted indexes and improving relevance in search results.
Patent Information
- Application Number
- PCT/RU2023/000417
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-03
AI Technical Summary
Current text search engines, particularly those using inverted indexes, face challenges with high storage costs and lower search efficiency due to the inclusion of numerous tokens with low usefulness and the inability to combine different data types, leading to suboptimal search results.
A cloud management platform that receives search requests, splits them into terms, and uses an index table with preset terms and media data to provide search results with different data types, including text, image, and audio, while optimizing the index table based on historical search requests to reduce storage costs and improve efficiency.
The solution enhances search quality and efficiency by providing relevant results with various data types, reducing storage costs, and improving the alignment of search terms with actual user intent through a tokenizer model trained on historical data.
Smart Images

Figure RU2023000417_03072025_PF_FP_ABST
Abstract
Description
METHOD OF SEARCHING DATA AND RELATED DEVICETECHNICAL FIELD
[0001] Embodiments of the present application relate to the field of cloud computing, and more particularly, to a method of searching data, and related devices for searching data.BACKGROUND
[0002] The domain of text search engines has a long history that dates back to the early days of computing. The first text search engine was used mainly for retrieving relevant information from academic archives. The advent of the internet led to a major shift in the use of text search engines. With the explosion of online content, there was a need for more sophisticated search engines that could handle complex queries and return relevant results from a vast array of online resources.
[0003] Local private datasets become large as well but current market of local search engines offer solutions for traditional text search based on inverted index. The purpose of inverted index is to optimize the speed of the query operations, by maintaining a sorted list of data records, called posting lists (postings). Each token has a list of the documents which contain it. Due to the fact that the tokens in the inverted index table are based on the documents, the inverted index table includes a large number of tokens with lower usefulness during the process of searching data, resulting in higher storage costs and lower search efficiency. Meanwhile, the inverted index table doesn’t allow combination of several modalities in the same index, and the search results that include different data types cannot be obtained through the inverted index table.
[0004] Therefore, how to improve the search quality is a challenge.SUMMARY
[0005] Embodiments of the present application provide a method of searching data and irelated devices for searching data. The technical solution may improve the search quality.
[0006] According to a first aspect, an embodiment of this application provides a method of searching data. This method is applied to cloud management platforms. The cloud management platform manages the infrastructure used to provide cloud services. This method includes: receiving a search request of a tenant; obtaining a set of search terms by splitting the search request; obtaining a set of search results according to the set of search terms and an index table, wherein the set of search results include search results with different data types, the index table includes a set of preset terms and identification information of a set of media data corresponding to each preset term in the set of preset terms, and the set of media data corresponding to each preset term has different data types; and providing the set of search results to the tenant.
[0007] According to the method above, the cloud management platform provides the set of search results with different data types to a tenant according to the search request of the tenant, thereby improving the search quality.
[0008] In a possible design, the data types include: text, image, video, or audio.
[0009] According to the method above, the cloud management platform provides various types of search results, including text, image, video, or audio, thereby improving the search quality.
[0010] In a possible design, a frequency of each search term in the set of search terms appearing in a preset database is lower than a first preset threshold, and the preset database includes at least one document and a set of historical search requests of tenant.
[0011] According to the method above, due to the term appearing in the preset database more frequently has lower usefulness, the cloud management platform uses the search term appearing in the preset database less frequently based on the search request for search, thereby providing the search results with higher quality.
[0012] In a possible design, the obtaining a set of search results according to the set of search terms and an index table includes: searching the index table based on a first search term in the set of search terms; when there is a first preset term in the index table that matches the first search term, obtaining the set of search results based on a set of media data corresponding to the first preset term; or, when there is not a preset term in the index table that matches the first search term, obtaining the set of search results according to a media database and the firstsearch term, wherein the media database includes multiple media data with different data types, and each search result in the set of search results belongs to the media database and corresponds to the first search term.
[0013] According to the method above, the cloud management platform quickly obtains the search results with different data types by searching a predetermined index table and matching the search term with the preset term in the index table, thereby improving the search efficiency.
[0014] In a possible design, the index table also includes relevance between each preset term and the media data corresponding to each preset term, and the identification information of multiple media data corresponding to each preset term is sorted in descending order of relevance in the index table.
[0015] According to the method above, the cloud management platform sorts media data based on the relevance between media data and the preset term when obtaining the index table. Therefore, the cloud management platform directly provides search results to tenants in the order of the media data in the index table, thereby improving the search quality.
[0016] In a possible design, the method further includes: obtaining the set of preset terms by splitting a set of historical search requests; and generating the index table according to the set of preset terms and the media database.
[0017] According to the method above, the cloud management platform generates the index table according to the preset term by splitting the historical search request. Due to the different style of the text (such as news articles, academic papers, and web pages) and the search request, the index table established based on search requests includes fewer invalid terms, which are less useful during search, thereby reducing the storage cost of the index table and improving the search efficiency.
[0018] In a possible design, the generating the index table according to the set of preset terms and the media database includes: obtaining a set of media data in the media database corresponding to second preset term according to the second preset term, wherein the second preset term belongs to the set of preset terms, the media data corresponding to the second preset term includes the second preset term, or semantic information in the media data corresponding to the second preset term matches the second preset term; calculating relevance between the second preset term and the media data corresponding to the second preset term; and generatingthe index table according to the second preset term, the media data corresponding to the second preset term and the relevance between the second preset term and the media data corresponding to the second preset term.
[0019] According to the method above, the cloud management platform generates the index table based on the preset terms and the media data corresponding to the preset terms, and uses this index table for searching to improve the search quality and efficiency.
[0020] In a possible design, when there is not a preset term in the index table that matches the first search term, the method further includes: adding the first search term and the identification information of the media data corresponding to the first search term to the index table.
[0021] According to the method above, the cloud management platform updates the index table based on the search term which does not exist in the index table and the media data corresponding to this search term. Therefore, the cloud management platform provides search results according to the updated index table when a tenant performs searching with this search term, thereby improving the search quality and efficiency.
[0022] In a possible design, the obtaining a set of search terms by splitting the search request includes: obtaining the set of search terms by splitting the search request based on a tokenizer model, wherein the tokenizer model is trained based on a set of historical search requests.
[0023] According to the method above, due to the fact that the tokenizer model is trained based on a set of historical search requests, the tokenizer model obtains terms that better represent the actual search intention of tenant based on the language style of the search request, thereby improving the search quality.
[0024] In a possible design, the method further includes: obtaining the tokenizer model by training the initial tokenizer model based on the set of historical search requests, wherein the frequency of a term obtained by the tokenizer model appearing in a preset database is lower than a first preset threshold, and the preset database includes at least one document and the set of historical search requests.
[0025] According to the method above, due to the term appearing in the preset database more frequently has lower usefulness, the cloud management platform trains the initialtokenizer model based on the set of historical search requests to obtain the tokenizer model. Therefore the terms obtained by the tokenizer model have higher usefulness during the search, thereby improving the search quality.
[0026] According to a second aspect, an embodiment of this application provides a method of searching data. This method is applied to cloud management platforms. The cloud management platform manages the infrastructure used to provide cloud services. This method includes: receiving a set of historical search requests of a tenant; obtaining a set of preset terms by splitting the set of historical search requests; and generating an index table according to the set of preset terms and a media database, wherein the index table includes the set of preset terms and identification information of media data in the media database corresponding to each preset term in the set of preset terms, and the index table is used to obtain a set of search results based on a search request.
[0027] According to the method above, the cloud management platform generates the index table according to the preset term by splitting the historical search request. Due to the different style of the text (such as news articles, academic papers, and web pages) and the search request, the index table established based on search requests includes fewer invalid terms, which are less useful during the search, thereby reducing the storage cost of the index table and improving the search efficiency. And the preset term in the index table better represents the actual search intention of tenant, therefore the search quality can be improved by using the preset term.
[0028] In a possible design, a frequency of each preset term in the set of preset terms appearing in a preset database is lower than a first preset threshold, and the first preset database includes at least one document and the set of historical search requests.
[0029] According to the method above, due to the term appearing in the preset database more frequently has lower usefulness, the cloud management platform uses the preset term appearing in the preset database less frequently based on the historical search request to generate the index table, in order to improve the search quality based on the index table.
[0030] In a possible design, the index table also includes relevance between each preset term and the media data corresponding to each preset term, and the identification information of multiple media data corresponding to each preset term is sorted in descending order of relevance in the index table.
[0031] According to the method above, the cloud management platform sorts media data based on the relevance between media data and the preset term when obtaining the index table. Therefore, the cloud management platform directly provides search results to tenants in the order of the media data in the index table, thereby improving the search quality.
[0032] In a possible design, the generating an index table according to the set of preset terms and a media database includes: obtaining a set of media data in the media database corresponding to second preset term according to the second preset term, wherein the second preset term belongs to the set of preset terms, the media data corresponding to the second preset term includes the second preset term, or semantic information in the media data corresponding to the second preset term matches the second preset term; and calculating the relevance between the second preset term and the media data corresponding to the second preset term; generating the index table according to the second preset term, the media data corresponding to the second preset term and the relevance between the second preset term and the media data corresponding to the second preset term.
[0033] According to the method above, the cloud management platform generates the index table based on the preset terms and the media data corresponding to the preset terms, and uses this index table for search to improve the search quality and efficiency.
[0034] In a possible design, the obtaining a set of preset terms by splitting the set of historical search requests includes: obtaining the set of preset terms by splitting the set of historical search requests based on a tokenizer model, wherein the tokenizer model is trained based on multiple historical search requests in the set of historical search requests.
[0035] According to the method above, due to the fact that the tokenizer model is trained based on a set of historical search requests, the tokenizer model obtains terms that better represent the actual search intention of tenant based on the language style of the search requests, thereby improving the search quality.
[0036] In a possible design, the method further includes: obtaining the tokenizer model by training the initial tokenizer model based on multiple historical search requests in the set of historical search requests, wherein a frequency of a term obtained by the tokenizer model appearing in a preset database is lower than a first preset threshold, and the preset database includes at least one document and the multiple historical search requests in the set of historicalsearch requests.
[0037] According to the method above, due to the term appearing in the preset database more frequently has lower usefulness, the cloud management platform trains the initial tokenizer model based on the set of historical search requests to obtain the tokenizer model. Therefore the terms obtained by the tokenizer model have higher usefulness during the search, thereby improving the search quality.
[0038] In a possible design, the media database includes multiple media data with different data types, and a set of media data in the index table corresponding to each preset term has different data types.
[0039] In a possible design, the data types include: text, image, video, or audio.
[0040] According to the method above, the media data corresponding to the preset term in the index table has different data types, therefore the cloud management platform provides various types of search results during search by the tenant, thereby improving the search quality.
[0041] According to a third aspect, an embodiment of this application provides a device of searching data. This device is applied to cloud management platforms. The cloud management platform manages the infrastructure used to provide cloud services. This device includes a transceiver unit and a processing unit. The transceiver unit is configured to receive a search request of a tenant. The processing unit is configured to obtain a set of search terms by splitting the search request. The processing unit is also configured to obtain a set of search results according to the set of search terms and an index table, wherein the set of search results includes search results with different data types, the index table includes a set of preset terms and identification information of a set of media data corresponding to each preset term in the set of preset terms, and the set of media data corresponding to each preset term has different data types. The transceiver unit is also configured to provide the set of search results to the tenant.
[0042] In a possible design, the data types include: text, image, video, or audio.
[0043] In a possible design, a frequency of each search term in the set of search terms appearing in a preset database is lower than a first preset threshold, and the preset database includes at least one document and a set of historical search requests of tenant.
[0044] In a possible design, the processing unit is configured to search the index table based on a first search term in the set of search terms; when there is a first preset term in the indextable that matches the first search term, the processing unit is also configured to obtain the set of search results based on a set of media data corresponding to the first preset term; or, when there is not a preset term in the index table that matches the first search term, the processing unit is also configured to obtain the set of search results according to a media database and the first search term, wherein the media database includes multiple media data with different data types, and each search result in the set of search results belongs to the media database and corresponds to the first search term.
[0045] In a possible design, the index table also includes relevance between each preset term and the media data corresponding to each preset term, and the identification information of multiple media data corresponding to each preset term is sorted in descending order of relevance in the index table.
[0046] In a possible design, the processing unit is also configured to obtain the set of preset terms by splitting a set of historical search requests; and the processing unit is also configured to generate the index table according to the set of preset terms and the media database.
[0047] In a possible design, the processing unit is configured to obtain a set of media data in the media database corresponding to second preset term according to the second preset term, wherein the second preset term belongs to the set of preset terms, the media data corresponding to the second preset term includes the second preset term, or semantic information in the media data corresponding to the second preset term matches the second preset term; the processing unit is also configured to calculate relevance between the second preset term and the media data corresponding to the second preset term; and the processing unit is also configured to generate the index table according to the second preset term, the media data corresponding to the second preset term and the relevance between the second preset term and the media data corresponding to the second preset term.
[0048] In a possible design, when there is not a preset term in the index table that matches the first search term, the processing unit is also configured to add the first search term and the identification information of the media data corresponding to the first search term to the index table.
[0049] In a possible design, the processing unit is configured to obtain the set of search terms by splitting the search request based on a tokenizer model, wherein the tokenizer modelis trained based on a set of historical search requests.
[0050] In a possible design, the processing unit is configured to obtain the tokenizer model by training the initial tokenizer model based on the set of historical search requests, wherein a frequency of a term obtained by the tokenizer model appearing in a preset database is lower than a first preset threshold, and the preset database includes at least one document and the set of historical search requests.
[0051] According to a fourth aspect, an embodiment of this application provides a device of searching data. This device is applied to cloud management platforms. The cloud management platform manages the infrastructure used to provide cloud services. This device includes a transceiver unit and a processing unit. The transceiver unit is configured to receive a set of historical search requests of a tenant; the processing unit is configured to obtain a set of preset terms by splitting the set of historical search requests; and the processing unit is also configured to generate an index table according to the set of preset terms and a media database, wherein the index table includes the set of preset terms and identification information of media data in the media database corresponding to each preset term in the set of preset terms, and the index table is used to obtain a set of search results based on a search request.
[0052] In a possible design, a frequency of each preset term in the set of preset terms appearing in a preset database is lower than a first preset threshold, and the first preset database includes at least one document and the set of historical search requests.
[0053] In a possible design, the index table also includes relevance between each preset term and the media data corresponding to each preset term, and the identification information of multiple media data corresponding to each preset term is sorted in descending order of relevance in the index table.
[0054] In a possible design, the processing unit is configured to obtain a set of media data in the media database corresponding to second preset term according to the second preset term, wherein the second preset term belongs to the set of preset terms, the media data corresponding to the second preset term includes the second preset term, or semantic information in the media data corresponding to the second preset term matches the second preset term; the processing unit is also configured to calculate relevance between the second preset term and the media data corresponding to the second preset term; and the processing unit is also configured to generatethe index table according to the second preset term, the media data corresponding to the second preset term and the relevance between the second preset term and the media data corresponding to the second preset term.
[0055] In a possible design, the processing unit is configured to obtain the set of preset terms by splitting the set of historical search requests based on a tokenizer model, wherein the tokenizer model is trained based on multiple historical search requests in the set of historical search requests.
[0056] In a possible design, the processing unit is also configured to obtain the tokenizer model by training the initial tokenizer model based on multiple historical search requests in the set of historical search requests, wherein a frequency of a term obtained by the tokenizer model appearing in a preset database is lower than a first preset threshold, and the preset database includes at least one document and the multiple historical search requests in the set of historical search requests.
[0057] In a possible design, the media database includes multiple media data with different data types, and a set of media data in the index table corresponding to each preset term has different data types.
[0058] In a possible design, the data types include: text, image, video, or audio.
[0059] According to a fifth aspect, an embodiment of this application provides a computing device cluster, including at least one computing device, where the computing device includes a processor and a memory coupled with the processor, where the memory is configured to store instructions, and the processor is configured to invoke and run the computer program stored in the memory, so that the computing device executes the method in any of the first aspect, the second aspect or any possible design of the first aspect or the second aspect.
[0060] According to a sixth aspect, an embodiment of this application provides a computer readable storage medium including instructions. When the instructions are run on a computer device cluster, the computer device cluster is enabled to perform the method in any of the first aspect, the second aspect or any possible design of the first aspect or the second aspect.
[0061] According to a seventh aspect, a chip system is provided, where the chip system includes a memory and a processor, the memory is configured to store a computer program, and the processor is configured to invoke the computer program from the memory and run thecomputer program, so that a server on which a chip is disposed performs the method in any of the first aspect, the second aspect or any possible design of the first aspect or the second aspect.
[0062] According to an eighth aspect, a computer program product including instructions is provided, where when the instructions are run on a computer device cluster, the computer device cluster is enabled to perform the method in any of the first aspect, the second aspect or any possible design of the first aspect or the second aspect.DESCRIPTION OF DRAWINGS
[0063] FIG. 1 is a schematic block diagram of a cloud management platform according to an embodiment of this application.
[0064] FIG. 2 shows an example of an inverted index structure.
[0065] FIG. 3 is a schematic diagram of a method of searching data according to an embodiment of this application.
[0066] FIG. 4 is a schematic diagram of a method of searching data according to an embodiment of this application.
[0067] FIG. 5 is a schematic diagram of a method of searching data according to an embodiment of this application.
[0068] FIG. 6 is a schematic diagram of a device of searching data according to an embodiment of this application.
[0069] FIG. 7 is a schematic block diagram of a computing device according to an embodiment of this application.
[0070] FIG. 8 is a schematic block diagram of a computing device cluster according to an embodiment of this application.
[0071] FIG. 9 is a schematic block diagram of a computing device 700 A and a computing device 700B connected by a network according to an embodiment of this application.DESCRIPTION OF EMBODIMENTS
[0072] The following describes the technical solutions in the present application with reference to the accompanying drawings. Obviously, the described embodiments are part of theembodiments of the present application, but not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without making creative labor shall fall within the scope of protection of the present application.
[0073] The present application will present aspects, embodiments, or features around systems that include multiple devices, components, modules, etc. It should be understood and appreciated that the individual systems may include additional devices, components, modules, etc., and / or may not include all of the devices, components, modules, etc. discussed in connection with the accompanying drawings. In addition, combinations of these options may be used.
[0074] In addition, in the embodiments of the present application, the word “exemplarily” and the phrase “as an example” are used to indicate for example, illustration or description. Any embodiment or design solution described as “exemplarily” in this application should not be construed as being superior to or more advantageous than other embodiments or design solutions. Rather, the use of the word “example” is intended to present the concept in a specific manner.
[0075] The phrases “in some possible embodiments”, “in some possible application scenarios”, etc., appearing in various places in this description, do not necessarily refer to the same embodiments, but rather mean “one or more, but not all, embodiments” unless otherwise specifically emphasized. Unless otherwise specifically emphasized, the terms “including”, “comprising”, “having”, and variations thereof all mean “including but not limited to”.
[0076] In the present application, “at least one” refers to one or more, and “multiple” refers to two or more, “and / or”, describing the association of the associated objects, indicates that three relationships can exist. For example, A and / or B can mean A alone, both A and B, and B alone, where A and B can be singular or plural. The character “ / ” generally indicates that the preceding and following associated objects are in an “or” relationship.
[0077] The application scenarios described in the present application embodiments are intended to illustrate the technical solutions of the present application embodiments more clearly and do not constitute a limitation to the technical solutions provided by the present application embodiments. It is known to those of ordinary skill in the art that the technicalsolutions provided by the present application embodiments are equally applicable to similar technical problems as the system architecture evolves and new application scenarios emerge.
[0078] In order to better describe the solutions of embodiments in the present application, concepts and terms that may be involved in the present application will be described below.
[0079] (1) Approximate nearest neighbors (ANN)
[0080] ANN is a search algorithm that allows for the identification of points in a dataset that are closest to a given point. The term “approximate” is used because the algorithm does not always guarantee to find the exact nearest neighbor. Instead, it finds points that are nearly as close to the given point as the true nearest neighbor.
[0081] (2) Posting lists
[0082] Mapping from token to positions where it occurs in text. There could be more than one appearance of same token in text.
[0083] (3) Multi-modal data
[0084] The data have different types of representation (text, image, video, audio, etc.).
[0085] (4) Tokenizer model
[0086] A component that takes text as input, breaks it up into individual tokens (e.g. individual words), and outputs collection of tokens. For instance, a whitespace tokenizer breaks text into tokens whenever it sees any whitespace.
[0087] (5) Tokenization
[0088] The process of splitting text to separate tokens.
[0089] (6) Token
[0090] A token is an atomic part of text. Usually it is a single word, subword (sequence of characters), or sequence of words.
[0091] (7) Term
[0092] A term is a synonym for ‘token’.
[0093] The technical solution in the embodiments of the present application is applied to computing devices, such as servers, hosts, personal computers, laptops, desktop computers, etc. The server is a local server or a cloud server. This technical solution is applied to a cloud management platform when the server is a cloud server.
[0094] FIG. 1 is a schematic block diagram of a cloud management platform 100 accordingto an embodiment of this application. In FIG. 1, the cloud management platform 100 manages the infrastructure used to provide cloud services. This infrastructure is deployed in one or more data centers. This infrastructure includes a computing device cluster and / or a storage device cluster. The computing device cluster includes at least one computing device. When computing devices are included in the computing device cluster, they can be directly connected or connected with each other through a network. This network is, for example, a local area network or a wide area network. Each computing device in the computing device cluster can independently execute the method in the embodiments of the present application, or computing devices in the computing device cluster can jointly execute the method in the embodiments of the present application. The storage device cluster includes at least one storage device. When storage devices are included in the storage device cluster, they can be directly connected or connected with each other through a network. This network is, for example, a local area network or a wide area network. The storage device in the storage device cluster is either a centralized storage device or a distributed storage device, and embodiments of the present application is not limited to this.
[0095] The cloud management platform 100 includes a search engine 110 and a storage module 120. In some embodiments, the search engine 110 and the storage module 120 are deployed on the same or different computing devices, and embodiments of the present application is not limited to this.
[0096] The search engine 110 receives a search request of a tenant. The tenant is a public cloud tenant who has registered a public cloud account and purchased public cloud resources or services. The search engine 110 also obtains a set of search terms by splitting the search request, and obtains a set of search results according to the set of search terms and an index table. The set of search results include search results with different data types. The index table includes a set of preset terms and identification information of a set of media data corresponding to each preset term in the set of preset terms. The set of media data corresponding to each preset term has different data types. The search engine 110 also provides the set of search results to the tenant.
[0097] In some embodiments, the data types include: text, image, video, audio, etc.
[0098] In some embodiments, the search engine 110 searches the index table based on afirst search term in the set of search terms. When there is a first preset term in the index table that matches the first search term, the search engine 110 obtains the set of search results based on a set of media data corresponding to the first preset term. When there is not a preset term in the index table that matches the first search term, the search engine 110 obtains the set of search results according to a media database and the first search term. The media database includes multiple media data with different data types. Each search result in the set of search results belongs to the media database and corresponds to the first search term.
[0099] In some embodiments, the media database is either a tenant configured media database or a preset media database.
[0100] When there is not a preset term in the index table that matches the first search term, the search engine 110 adds the first search term and the identification information of the media data corresponding to the first search term to the index table.
[0101] In some embodiments, the search engine 110 obtains the set of search terms by splitting the search request based on a tokenizer model. The tokenizer model is trained based on historical search requests. A frequency of a term obtained by the tokenizer model appearing in a preset database is lower than a first preset threshold. The preset database includes at least one document and multiple historical search requests in the set of historical search requests.
[0102] In some embodiments, the search engine 110 obtains a set of preset terms by splitting a set of historical search requests. The search engine 110 also generates the index table according to the set of preset terms and the media database.
[0103] In some embodiments, the search engine 110 obtains a set of historical search requests of a tenant, and obtains a set of preset terms by splitting the set of historical search requests. A frequency of each preset term in the set of preset terms appearing in a preset database is lower than the first preset threshold. The first preset database includes at least one document and the set of historical search requests.
[0104] In some embodiments, the search engine 110 receives the set of historical search requests of a tenant, or the search engine 110 obtains the set of historical search requests through other module which is connected with the search engine 110.
[0105] In some embodiments, the search engine 110 obtains a set of media data in the media database corresponding to a second preset term according to the second preset term. The secondpreset term belongs to the set of preset terms. The media data corresponding to the second preset term includes the second preset term, or semantic information in the media data corresponding to the second preset term matches the second preset term. The search engine 110 also calculates relevance between the second preset term and the media data corresponding to the second preset term. And the search engine 110 generates the index table according to the second preset term, the media data corresponding to the second preset term and the relevance between the second preset term and the media data corresponding to the second preset term.
[0106] In some embodiments, the index table also includes relevance between each preset term and the media data corresponding to each preset term, and the identification information of multiple media data corresponding to each preset term is sorted in descending order of relevance in the index table.
[0107] The storage module 120 is used to store intermediate data or results generated by the search engine 110 during the execution of the method in the embodiments of the present application. For example, the storage module 120 is used to store the index table, media database, historical search requests, tokenizer model, and so on. The search engine 110 obtains the required data from the storage module 120 to execute the method in the embodiments of the present application.
[0108] The cloud management platform 100 in FIG. 1 obtains the media data with different data types corresponding to the search request according to the index table, which includes a set of preset terms and the media data corresponding to each term in the set of preset terms. Therefore, the cloud management platform provides the set of search results with different data types to a tenant, thereby improving the search quality.
[0109] FIG. 2 shows an example of an inverted index structure. A set of texts are included in FIG. 2, which are namely textl , text2, and text3. Each text in FIG. 2 includes different content. After inputting the set of text from FIG. 2 into an index builder 210, a list 220 as shown in FIG. 2 can be generated. The list 220 includes a set of terms and identification information corresponding to each term in the set of terms. As shown in list 220, the term “class” is included in the text with identification information of 1, 5, 10, etc. Similarly, the term “children” is included in text with identification information of 1, 5, 20, etc. From list 220, it can be seen that splitting each text to generate inverted indexes easily leads to a large amount of data in list 220.Moreover, due to the difference in language style between the text and the search request used by tenant during search, there are many words in list 220 that are rarely retrieved by tenants, which increases storage costs and reduces search efficiency. Meanwhile, the text corresponding to each term in list 220 is sorted according to the identification information of the text, which does not allow to store and retrieve text in order of their usefulness, thereby reducing the search efficiency.
[0110] FIG. 3 is a schematic diagram of a method of searching data according to an embodiment of this application. The method in FIG. 3 is applied to the cloud management platform 100 in FIG.1. The method in FIG. 3 includes the following steps.
[0111] Step 310: receiving a search request of a tenant.
[0112] The cloud management platform receives a search request of a tenant. The search request includes at least one character, which includes at least one of the following: letter, number, symbol, etc.
[0113] Optionally, the cloud management platform provides the tenant with a first graphical interface, allowing to upload or select operations on the first graphical interface, thereby obtaining the search request of the tenant. The embodiments of the present application do not limit the specific manifestation of the first graphical interface.
[0114] Step 320: obtaining a set of search terms by splitting the search request.
[0115] The cloud management platform splits the search request to obtain a set of search terms. The set of search terms includes at least one search term. Each search term belongs to the search request.
[0116] For example, assuming the search request is “a person wearing a hat”, the cloud management platform splits the search request to obtain the following search terms: “wearing”, “hat”, and “person”.
[0117] Optionally, the cloud management platform obtains the set of search terms by splitting the search request based on a tokenizer model as shown in FIG. 4. FIG. 4 is a schematic diagram of a method of searching data according to an embodiment of this application. In FIG. 4, the search request is input into a tokenizer model 410 to obtain n search terms: Ti, Ta, ..., Tn. The n is a positive integer.
[0118] In some embodiments, the tokenizer model is trained based on a set of historicalsearch requests. The set of historical search requests is from at least one tenant. The set of historical search requests includes at least one historical search request.
[0119] In some embodiments, a frequency of a search term obtained by the tokenizer model appearing in a preset database is lower than a first preset threshold, and the preset database includes at least one preset document and the set of historical search requests. The embodiments of the present application do not limit the size of the first preset threshold.
[0120] Before the step 320, the cloud management platform obtains the tokenizer model. Alternatively, the cloud management platform obtains a set of historical search requests, and obtains the tokenizer model by training the initial tokenizer model based on historical search requests in the set of historical search requests.
[0121] Step 330: obtaining a set of search results according to the set of search terms and an index table.
[0122] The cloud management platform searches the index table according to each search term in the set of search terms to obtain the set of search results. The index table includes a set of preset terms and identification information of a set of media data corresponding to each preset term in the set of preset terms. The set of media data corresponding to each preset term has different data types. Therefore, the set of search results includes search results with different data types.
[0123] In some embodiments, the data types include: text, image, video, audio, etc. In other words, the set of search results includes at least two types of text, image, audio, video, etc.
[0124] In some embodiments, the identification information of each media data is used to identify the media data. For example, the identification information of the media data is the number of the media data, access path of the media data, storage address of the media data, etc.
[0125] When there is a first preset term in the index table that matches the first search term, the cloud management platform obtains the set of search results based on a set of media data corresponding to the first search term in the index table. The first preset term belongs to the set of preset terms in the index table. The first search term belongs to the set of search terms. The set of media data corresponding to the first search term includes at least two media data.
[0126] For example, in FIG. 4, the cloud management platform searches the index table 420 based on each search term of Ti, T2, ..., Tnto obtain the preset term corresponding to eachsearch term. The index table in FIG. 4 includes n preset terms: Ti, T2, ..., Tn. The n is a positive integer. The preset term “Ti” corresponds to multiple media data: Dl_l, Dl_2, etc. with different data types. Similarly, the preset term “T2” corresponds to multiple media data: D2_l, D2_2, etc. with different data types. The preset term “Tn” corresponds to multiple media data: Dn_l, Dn_2, etc. with different data types. The cloud management platform obtains the preset term Ti corresponding to the search term Ti, where the i=l, ..., n. The preset term T; corresponds to the search term Ti, which means that the preset term Ti is the same as the search term Ti. The cloud management platform obtains multiple media data according to the identification information of the media data corresponding to the preset term Ti in the index table, and uses these multiple media data as the set of search results.
[0127] It should be understood that FIG. 4 only shows one possible form of an index table, which is also presented in other forms. The naming methods for search terms, identification information of media data, etc. in FIG. 4 are only examples, and other naming methods can be used. The embodiments of the present application are not limited to this. For example, the search term or the identification information of media data includes at least one character, which includes at least one of the following: letter, number, symbol, etc.
[0128] When there is not a preset term in the index table that matches the first search term, the cloud management platform obtains the set of search results according to a media database and the first search term. Each search result in the set of search results belongs to the media database and corresponds to the first search term. The media database includes multiple media data with different data types.
[0129] When there is not a preset term in the index table that matches the first search term, the cloud management platform obtains multiple media data corresponding to the first search term in the media database, and uses these multiple media data as the set of search results. The media data corresponding to the first search term includes the first search term, or semantic information in the media data corresponding to the first search term matches the first search term. The specific implementation method can be found in step 530 of FIG. 5.
[0130] In some embodiments, the cloud management platform updates the index table by adding the first search term and the identification information of the media data corresponding to the first search term to the index table.
[0131] In some embodiments, before the step 330, the cloud management platform obtains the index table. Alternatively, the cloud management platform obtains a set of preset terms and a media database, and generates the index table based on the set of preset terms and the media database. The set of preset terms includes at least one preset term.
[0132] In some embodiments, the set of preset terms is obtained based on a set of historical search requests. Before the step 330, the cloud management platform obtains the set of preset terms. Or the cloud management platform obtains a set of historical search requests and obtains the set of preset terms by splitting the set of historical search requests.
[0133] Step 340: providing the set of search results to the tenant.
[0134] The cloud management platform provides the set of search results to the tenant. For example, the cloud management platform provides the tenant with a second graphical interface, allowing to upload or select operations on the second graphical interface, thereby providing the set of search results to the tenant. The embodiments of the present application do not limit the specific manifestation of the second graphical interface.
[0135] In some embodiments, the cloud management platform determines the sorting order of search results based on at least one of the following: the relevance between each search result and the set of search terms, the relevance between each search result and a set of extension terms, and the number of visits to each search result. The set of extension terms includes at least one search term and / or at least one generative term. The meaning of each generative term is the same or similar to that of a search term.
[0136] For example, the higher the relevance between a search result and the set of search terms is, the higher the ranking of the search result is. Alternatively, the higher the relevance between a search result and the set of extension terms is, the higher the ranking of the search result is. Alternatively, the more the visits to a search result are, the higher the ranking of the search result is. Alternatively, the higher the relevance between a search result and the set of search terms is, and the more the visits to a search result are, the higher the ranking of the search result is. Alternatively, the higher the relevance between a search result and the set of extension terms is, and the more the visits to a search result are, the higher the ranking of the search result is.
[0137] For example, the more search terms are included in a search result, the higher therelevance between the search result and the set of search terms is. The less search terms are included in a search result, the lower the relevance between the search result and the set of search terms is. Alternatively, the relevance between a search result and the set of search terms is determined based on the relevance between the search result and each search term included in the search result. For example, When the number of search terms included in two search results is same, the higher the relevance between a search result and each search term included in the search result is, the higher the relevance between the search result and the set of search terms is. Alternatively, the higher the weighted sum of the relevance between a search result and each search term included in the search result is, the higher the relevance between the search result and the set of search terms is.
[0138] Similarly, the more extension terms are included in a search result, the higher the relevance between the search result and the set of extension terms is. The less extension terms are included in a search result, the lower the relevance between the search result and the set of extension terms is. Alternatively, the relevance between a search result and the set of extension terms is determined based on the relevance between the search result and each extension term included in the search result. For example, when the number of extension terms included in two search results is same, the higher the relevance between a search result and each extension term included in the search result is, the higher the relevance between the search result and the set of extension terms is. Alternatively, the higher the weighted sum of the relevance between a search result and each extension term included in the search result is, the higher the relevance between the search result and the set of extension terms is.
[0139] For example, as shown in FIG. 4, the cloud management platform sorts multiple search results based on at least one method of ranking: method of ranking 1, method of ranking 2, or method of ranking 3, in order to provide the tenant with multiple search results after sorting. The method of ranking 1 performs sorting based on the relevance between each search result and the set of search terms. The method of ranking 2 performs sorting based on the relevance between each search result and the set of extension terms. The method of ranking 3 performs sorting based on the number of visits to each search result. In FIG. 4, the search results are obtained based on the method of ranking 1 and the method of ranking 3 as shown in search results page 430. In the search results page 430, there are m search results, where m is a positiveinteger greater than 1. The m search results are Data_ 131, Data_ 5, Data_ 168. The m search results include different data types. In the search results page 430, the relevance between Data_ 131 and the set of search terms Ti-Tnis higher than that between Data_ 5 and the set of search terms Ti -Tn, and / or, Data_ 131 has more visits than Data_ 5.
[0140] Optionally, the cloud management platform obtains the method of ranking to sort search results based on the tenant’s configuration. For example, the cloud management platform provides the tenant with a third graphical interface, allowing for uploading or selecting operations on the third graphical interface, thereby obtaining the method of ranking based on the tenant’s configuration. The embodiments of the present application do not limit the specific manifestation of the third graphical interface.
[0141] The method in FIG. 3 obtains the media data with different data types corresponding to the search request according to the index table, which includes a set of preset terms and the media data corresponding to each term in the set of preset terms. Therefore, the cloud management platform provides the set of search results with different data types to the tenant, thereby improving the search quality.
[0142] FIG. 5 is a schematic diagram of a method of searching data according to an embodiment of this application. The method in FIG. 5 is applied to the cloud management platform 100 in FIG. 1. The method in FIG. 5 includes the following steps.
[0143] Step 510: obtaining a set of historical search requests and a tokenizer model.
[0144] The cloud management platform directly obtains a set of historical search requests from at least one tenant. Alternatively, the cloud management platform receives a set of historical search requests uploaded by a tenant. For example, the cloud management platform provides the tenant with a fourth graphical interface, allowing for uploading or selecting operations on the fourth graphical interface, thereby obtaining the set of historical search requests of the tenant. The embodiments of the present application do not limit the specific manifestation of the fourth graphical interface.
[0145] The cloud management platform directly obtains a tokenizer model. Alternatively, the cloud management platform obtains the tokenizer model based on the set of historical search requests.
[0146] In some embodiments, the cloud management platform obtains the tokenizer modelby training the initial tokenizer model based on historical search requests in the set of historical search requests. The frequency of the term obtained by the tokenizer model appearing in the preset database is lower than the first preset threshold, and the preset database includes at least one preset document and the set of historical search requests. The embodiments of the present application do not limit the size of the first preset threshold.
[0147] In some embodiments, the cloud management platform obtains a set of first training terms by splitting a first historical search request based on the initial tokenizer model. The first historical search request belongs to the set of historical search requests. The set of first training terms includes at least one first training term. The cloud management platform also obtains the tokenizer model by adjusting the initial tokenizer model based on the frequency of each first training term appearing in the preset database and the first preset threshold.
[0148] For example, when the frequency of at least one first training term appearing in the preset database is higher than the first preset threshold, the cloud management platform adjusts at least one parameter in the initial tokenizer model to obtain a set of second training terms. The set of second training terms is obtained by splitting the first historical search request based on the adjusted initial tokenizer model. The set of second training terms includes at least one second training term. The frequency of each second training term in the set of second training terms appearing in the preset database is lower than the first preset threshold. The cloud management platform also uses the adjusted initial tokenizer model as the tokenizer model.
[0149] Step 520: obtaining a set of preset terms by splitting the set of historical search requests based on the tokenizer model.
[0150] The cloud management platform obtains the set of preset terms by splitting the set of historical search requests based on the tokenizer model. The frequency of each preset term in the set of preset terms appearing in the preset database is lower than the first preset threshold.
[0151] In some embodiments, after obtaining the set of preset terms by splitting the set of historical search requests, the cloud management platform removes a duplicate preset term in the set of preset terms.
[0152] Step 530: obtaining a set of media data in the media database corresponding to a second preset term according to the second preset term.
[0153] The cloud management platform obtains a set of media data corresponding to eachpreset term by searching the media database.
[0154] In some embodiments, the cloud management platform obtains a set of media data corresponding to the second preset term according to the second preset term. The second preset term belongs to the set of preset terms. The set of media data corresponding to the second preset term includes at least one media data. The set of media data corresponding to the second preset term includes the second preset term. Alternatively, semantic information in the media data corresponding to the second preset term matches the second preset term.
[0155] For example, when the first media data is text and the first media data includes the second preset term, the first media data corresponds to the second preset term.
[0156] For example, when the second media data is any of image, video, audio, etc. and the semantic information in the second media data matches the second preset term, the second media data corresponds to the second preset term. The semantic information in the second media data includes at least one semantic term. Alternatively, the semantic information in the second media data includes at least one semantic term and relevance between each semantic term and the second media data.
[0157] In some embodiments, matching semantic information with the second preset term means: a semantic term in the semantic information is the same as the second preset term. Alternatively, matching semantic information with the second preset term means: a semantic term in the semantic information has similar meanings to the second preset term. In other words, matching semantic information with the second preset term means: the relevance between a semantic term in the semantic information and the second preset term is higher than a second preset threshold. The embodiments of the present application do not limit the size of the second preset threshold.
[0158] In some embodiments, before the step 530, the cloud management platform obtains the semantic information in each image, video, audio, etc. in the media database. Alternatively, the cloud management platform obtains a semantic information extraction model, which is used to obtain the semantic information in each image, video, audio, etc. in the media database.
[0159] For example, the semantic information extraction model is a mapping relationship between image, video, audio and the semantic information.
[0160] For example, the semantic information extraction model is trained through machinelearning based on the first training dataset. The first training dataset includes any one of image, video, audio, semantic information, as well as the mapping relationship between any one of image, video, audio and semantic information.
[0161] In some embodiments, before the step 530, the cloud management platform obtains the semantic information extraction model. Alternatively, the cloud management platform obtains the first training dataset and trains a model based on the first training dataset to obtain the semantic information extraction model.
[0162] Optionally, after obtaining the media data corresponding to each preset term, the cloud management platform performs the step 540.
[0163] Step 540: calculating the relevance between the second preset term and the media data corresponding to the second preset term.
[0164] Optionally, the cloud management platform calculates the relevance between the second preset term and the media data corresponding to the second preset term based on a relevance calculation model.
[0165] For example, the relevance calculation model is a mapping relationship between the media data, the preset term, and the relevance between the preset term and the media data.
[0166] For example, the relevance calculation model is trained through machine learning based on the second training dataset. The second training dataset includes the media data, the preset term, the relevance between the preset term and the media data, as well as the mapping relationship between the media data, the preset term, and the relevance.
[0167] In some embodiments, before the step 540, the cloud management platform obtains the relevance calculation model. Alternatively, the cloud management platform obtains the second training dataset and trains a model based on the second training dataset to obtain the relevance calculation model.
[0168] In some embodiments, the relevance calculation model calculates the relevance between the second preset term and the first text based on the position and / or frequency of the second preset term in the first text.
[0169] For example, the cloud management platform calculates the first relevance between the second preset term and the first text according to the position of the second preset term in the first text. Alternatively, the cloud management platform calculates the second relevancebetween the second preset term and the first text according to the frequency of the second preset term in the first text. Alternatively, the cloud management platform calculates the third relevance between the second preset term and the first text according to the first relevance and the second relevance. The cloud management platform uses any one of the first relevance, second relevance and third relevance as the relevance between the second preset term and the first text.
[0170] As shown in FIG. 4, the cloud management platform calculates the first relevance S 1 between the preset term T i and the text D 1 1 according to the position of the preset term T i in the text Dl_l. The cloud management platform also calculates the second relevance S2 between the preset term Ti and the text Dl_l according to the frequency of the preset term Ti in the text Dl_l.
[0171] It should be understood that the relevance between the preset term T i and the text Dl_l in FIG. 4 is only an illustrative example. In some embodiments, there is only one relevance between the preset term Ti and the text Dl_l in FIG. 4, and the relevance between the preset term Ti and the text Dl_l is calculated based on the first relevance SI and second relevance S2.
[0172] In some embodiments, the relevance calculation model calculates the relevance between the second preset term and the first image based on the second preset term and the semantic information in the first image.
[0173] For example, the cloud management platform calculates the relevance between the second preset term and the first image according to the relevance between the second preset term and at least one semantic term in the semantic information of the first image. Alternatively, the cloud management platform calculates the relevance between the second preset term and the first image according to the relevance between the first image and each semantic term in the semantic information of the first image and the relevance between the second preset term and at least one semantic term in the semantic information of the first image.
[0174] As shown in FIG. 4, the cloud management platform calculates the relevance S3 between the preset term Ti and the image Dl_2 according to the relevance between the preset term Ti and at least one semantic term in the semantic information of the image Dl_2.
[0175] In some embodiments, the relevance calculation model calculates the relevancebetween the second preset term and the first video based on the second preset term and the semantic information in the first video. For example, the cloud management platform calculates the relevance between the second preset term and the first video according to the relevance between the second preset term and at least one semantic term in the semantic information of the first video. Alternatively, the cloud management platform calculates the relevance between the second preset term and the first video according to the relevance between the first video and each semantic term in the semantic information of the first video and the relevance between the second preset term and at least one semantic term in the semantic information of the first video.
[0176] In some embodiments, the relevance calculation model calculates the relevance between the second preset term and the first audio based on the second preset term and the semantic information in the first audio. For example, the cloud management platform calculates the relevance between the second preset term and the first audio according to the relevance between the second preset term and at least one semantic term in the semantic information of the first audio. Alternatively, the cloud management platform calculates the relevance between the second preset term and the first audio according to the relevance between the first audio and each semantic term in the semantic information of the first audio and the relevance between the second preset term and at least one semantic term in the semantic information of the first audio.
[0177] In some embodiments, the cloud management platform calculates a set of relevance between the preset term and the media data corresponding to the preset term according to the P relevance calculation models. P is a positive integer. Each relevance in the set of relevance is calculated by a relevance calculation model. When P is greater than 0, the P relevance calculation models are different. Specifically, the cloud management platform calculates the p- th relevance between the second preset term and each media data corresponding to the second preset term based on the p-th relevance calculation model, p=l, ..., P. The cloud management platform adds the P relevance between the media data and the preset term in the index table. Alternatively, the cloud management platform calculates final relevance based on the P relevance between the media data and the preset term, and adds this final relevance in the index table.
[0178] It should be understood that the step 530 and step 540 are executed separately. Alternatively, the step 530 and step 540 are merged for execution. In other words, while or afterobtaining the media data corresponding to the second preset term, the cloud management platform calculates the relevance between the second preset term and the media data. Alternatively, after calculating the relevance between the second preset term and the media data, the cloud management platform obtains the media data corresponding to the second preset term.
[0179] Step 550: generating the index table according to the second preset term and the media data corresponding to the second preset term.
[0180] The cloud management platform generates the index table according to the preset term, the media data corresponding to the preset term, and the relevance between the preset term and the media data. The index table includes a set of preset terms and identification information of a set of media data corresponding to each preset term in the set of preset terms. The index table is used to obtain a set of search results based on a search request.
[0181] In some embodiments, the identification information of each media data is used to identify the media data. For example, the identification information of the media data is the number of the media data, access path of the media data, storage address of the media data, and etc.
[0182] In some embodiments, the identification information of multiple media data corresponding to the second preset term is sorted in descending order of relevance to the second preset term in the index table.
[0183] For example, when the index table includes P relevance between the second preset term and the media data corresponding to the second preset term, the cloud management platform calculates final relevance based on the P relevance between the media data and the second preset term. The final relevance is the weighted average of the P relevance. Alternatively, the final relevance is one of the P relevance. The cloud management platform decides the order of the media data based on the final relevance.
[0184] As shown in index table 420 of FIG. 4, the final relevance (calculated based on the relevance SI and S2) between the media data Dl_l and the preset term Ti is greater than the relevance S3 between the media data Dl_2 and the preset term Ti.
[0185] In some embodiments, the index table also includes the relevance between each preset term and the media data corresponding to each preset term. The index table also includes the number of media data corresponding to each preset term. The index table also includes theposition and / or frequency of each preset term in the media data.
[0186] The embodiments of the present application do not limit the specific representation of the index table, for example, the index table is represented as any one of a table (as shown in FIG. 4), list, array, matrix, etc.
[0187] In some embodiments, when the number of media data corresponding to a third preset term is greater than a third threshold, it indicates a high probability that the third preset term is invalid. The invalid term is less useful during search. In other words, the usefulness of the third preset term is lower than a fourth preset threshold. Alternatively, the usefulness of the third preset term is determined manually. The embodiments of the present application do not limit the size of the third preset threshold or the fourth preset threshold. The cloud management platform removes the third preset term from the index table. Alternatively, the cloud management platform requests engineers or tenants to confirm whether the third preset term is invalid. When the engineer or tenant confirms that the third preset term is invalid, the cloud management platform removes the third preset term from the index table. When the engineer or tenant confirms that the third preset term is not invalid, the cloud management platform keeps the third preset term in the index table.
[0188] Optionally, when the third preset term is invalid, the cloud management platform adjusts the tokenizer model so that the third preset term will be not obtained based on the adjusted tokenizer model.
[0189] The method in FIG. 5 generates the index table based on the media database and the set of preset terms, and implements search based on the index table. Due to the different style of the text and the search request, the preset terms obtained based on search requests include fewer invalid terms, the fewer invalid terms are less useful during search, thereby reducing the storage cost of the index table and improving the search efficiency. Meanwhile, even if media data is added to the media database, there is no need to update the preset terms in the index table. Instead, only the identification information of the media data corresponding to each preset term needs to be updated, thereby making it easier to maintain and update the index table. And, the identification information of the media data corresponding to each preset term in the index table is sorted in descending order of the relevance between each preset term and the media data, so it can quickly provide the tenant with search results sorted by relevance, therebyimproving the search quality. At the same time, the index table includes multiple media data with different data types corresponding to the preset term, so it can directly provide search results with different data types, thereby further improving the search quality.
[0190] FIG. 6 is a schematic diagram of a device of searching data 600 according to an embodiment of this application. As shown in FIG. 6, the device of searching data 600 includes: a transceiver unit 610 and a processing unit 620. The device of searching data 600 implements the method as shown in FIG. 3 -FIG. 5. The device of searching data 600 is applied to cloud management platforms.
[0191] The transceiver unit 610 is configured to receive a search request of a tenant and provide a set of search results to the tenant. The transceiver unit 610 performs the step 310 and 340 in FIG. 3.
[0192] In some embodiments, the transceiver unit 610 is configured to receive a set of historical search requests of a tenant and performs the step 510 in FIG. 5.
[0193] The processing unit 620 is configured to obtain a set of search terms by splitting the search request. The processing unit 620 is also configured to obtain a set of search results according to the set of search terms and an index table. The processing unit 620 performs the step 320-330 in FIG. 3 and the step 520-550 in FIG. 5.
[0194] In some embodiments, the processing unit 620 is configured to obtain a set of preset terms by splitting the set of historical search requests, and generate an index table according to the set of preset terms and a media database.
[0195] Among them, both the transceiver unit 610 and the processing unit 620 are implemented through software or hardware. The processing unit 620 is taken as an example to introduce the implementation of the processing unit 620. Similarly, the implementation of the transceiver unit 610 refers to the implementation of the processing unit 620.
[0196] When the unit is an example of a software functional unit, the processing unit 620 includes code running on a computational instance. Among them, the computational instance includes at least one of physical hosts (computing devices), virtual machines, and containers. Furthermore, the above computational instance can be one or more. For example, the processing unit 620 includes code running on multiple hosts / virtual machines / containers. It should be noted that multiple hosts / virtual machines / containers used to run the code are distributed in the sameregion or in different regions. Furthermore, multiple hosts / virtual machines / containers used to run the code are distributed within the same availability zone (AZ) or across different AZs, each of which includes a data center or multiple geographically close data centers. Typically, a region includes multiple AZs.
[0197] Similarly, multiple hosts / virtual machines / containers used to run the code are distributed within the same virtual private cloud (VPC) or across multiple VPCs. Among them, usually one VPC is set within a region, and cross regional communication between two VPCs within the same region, as well as VPCs from different regions, requires a communication gateway to be set up within each VPC to achieve interconnection between VPCs.
[0198] When the unit is an example of a hardware functional unit, the processing unit 620 includes at least one computing device, such as a server. Alternatively, the processing unit 620 also is a device implemented using application specific integrated circuits (ASIC) or programmable logic devices (PLD). Among them, the above-mentioned PLD is a complex programmable logic device (CPLD), field programmable gate array (FPGA), general array logic (GAL), or any combination thereof.
[0199] The multiple computing devices included in the processing unit 620 are distributed in the same region or in different regions. The multiple computing devices included in the processing unit 620 are distributed within the same AZ or across different AZs. Similarly, the multiple computing devices included in the processing unit 620 are distributed within the same VPC or across multiple VPCs. Among them, the multiple computing devices are any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0200] Therefore, the units of each example described in the embodiments of the present application are implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians may use different methods to achieve the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0201] It should be noted that the device provided in the above embodiments only provides examples of the division of various functional units when executing the above methods. Inpractical applications, the above functions can be assigned to different functional units according to needs, that is, the internal structure of the device can be divided into different functional units to complete all or part of the functions described above. For example, the transceiver unit 610 can be used to perform any step in the above method, and the processing unit 620 can be used to perform any step in the above method. The steps responsible for implementing the transceiver unit 610 and the processing unit 620 can be specified as needed. The transceiver unit 610 and the processing unit 620 respectively implement different steps in the above methods to achieve all the functions of the above device.
[0202] In addition, the device and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments mentioned above, which will not be repeated here.
[0203] The method provided in the embodiments of the present application may be executed by a computing device, which may also be referred to as a computer system. This includes a hardware layer, an operating system layer running on top of the hardware layer, and an application layer running on top of the operating system layer. This hardware layer includes hardware such as processing units, memory, and memory control units, followed by a detailed explanation of the hardware’s functions and structure. This operating system is any one or more computer operating systems that implement business processing through processes, such as Linux operating system, Unix operating system, Android operating system, iOS operating system, or Windows operating system. This application layer includes applications such as browsers, contacts, word processing software, instant messaging software, etc. And, alternatively, the computer system can be a handheld device such as a smartphone, or a terminal device such as a personal computer, which is not specifically limited by the present application, as long as it can be implemented through the methods provided in the embodiments of the present application. The execution subject of the method provided in the embodiments of the present application can be a computing device, or a functional module in the computing device that can call and execute the program.
[0204] FIG. 7 is a schematic block diagram of a computing device according to an embodiment of this application. The computing device 700 can be a server, a computer, or other device with computing power. The computing device 700 shown in FIG. 7 includes at least oneprocessor 710 and a memory 720.
[0205] It should be understood that embodiments of the present application do not limit the number of processors and memory in the computing device 700.
[0206] The processor 710 executes instructions in the memory 720 to enable the computing device 700 to implement the method provided in the embodiments of the present application. Alternatively, the processor 710 executes instructions in the memory 720 to enable the computing device 700 to implement the various functional modules provided in the embodiments of the present application, thereby implementing the methods provided in the embodiments of the present application.
[0207] Optionally, the computing device 700 also includes a communication interface 730. The communication interface 730 uses transceiver modules such as but not limited to network interface cards and transceivers to achieve communication between the computing device 700 and other devices or communication networks.
[0208] Optionally, the computing device 700 also includes a system bus 740, where the processor 710, the memory 720, and the communication interface 730 are respectively connected to the system bus 740. The processor 710 can access the memory 720 through the system bus 740, for example, the processor 710 can read and write data or execute code in the memory 720 through the system bus 740. The system bus 740 is either a peripheral component interconnect express (PCI) bus or an extended industry standard architecture (EISA) bus. The system bus 740 is divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in FIG. 7, but it does not mean that there is only one bus or one type of bus.
[0209] One possible implementation is that the function of the processor 710 is mainly to interpret instructions (or code) of computer programs and process data in computer software. Among them, the instructions of the computer program and the data in the computer software can be stored in the cache of the memory 720 or the processor 710.
[0210] Optionally, the processor 710 may be an integrated circuit chip with signal processing capabilities. As an example rather than a limitation, the processor 710 is a general- purpose processor, digital signal processor (DSP), ASIC, FPGA or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. Among them,general-purpose processors are microprocessors, etc. For example, the processor 710 is a central processing unit (CPU).
[0211] The memory 720 can provide running space for processes in the computing device 700, for example, storing computer programs (specifically, program code) used to generate processes in the memory 720. After the computer program is run by the processor and generates a process, the processor allocates corresponding storage space for the process in the memory 720. Furthermore, the above storage space further includes text segments, initialization data segments, bit initialization data segments, stack segments, heap segments, and so on. The memory 720 stores data generated during the operation of the process, such as intermediate data or process data, in the storage space corresponding to the above process.
[0212] Optionally, the memory is used to temporarily store operational data in the processor 710 and data exchanged with external memory such as hard drives. As long as the computer is running, the processor 710 will transfer the data that needs to be processed to memory for processing, and then transmit the result after the operation is completed.
[0213] As an example rather than a limitation, the memory 720 may be either volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Among them, non-volatile memory is read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), flash, or electrically EPROM (EEPROM). Volatile memory is a random access memory (RAM) used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus DRAM (DRDRAM). It should be noted that the memory 720 of the system and method described in this application is intended to include but not limited to these and any other suitable types of memory.
[0214] The structure of the computing device 700 listed above is only for illustrative purposes, and this application is not limited to it. The computing device 700 in the embodiments of this application includes various hardware in computer systems in prior art. For example, the computing device 700 also includes other storage devices besides memory 720, such as disk storage, etc. Technicians in this field should understand that the computing device 700 may alsoinclude other devices necessary for normal operation. Meanwhile, according to specific needs, technical personnel in this field should understand that the above-mentioned computing device 700 may also include hardware devices for implementing other additional functions. In addition, those skilled in the art should understand that the above-mentioned computing device 700 may only include the devices necessary to implement the embodiments of the present application, without necessarily including all the devices shown in FIG. 7.
[0215] The embodiment of this application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server. In some embodiments, the computing device may also be a terminal device such as a desktop computer, laptop, or smartphone.
[0216] As shown in FIG. 8, the computing device cluster includes at least one computing device 700. The memory 720 in one or more computing devices 700 in a computing device cluster may contain the same instructions for executing the above method.
[0217] In some possible implementations, the memory 720 in one or more computing devices 700 within the computing device cluster may also hold partial instructions for executing the above method. In other words, a combination of one or more computing devices 700 can jointly execute the instructions of the above method.
[0218] It should be noted that the memory 720 in different computing devices 700 in the computing device cluster can store different instructions, which are used to perform some of the functions of the above-mentioned devices. That is to say, the instructions stored in the memory 720 of different computing devices 700 can realize the function of one or more modules in the above-mentioned device.
[0219] In some possible implementations, one or more computing devices in a computing device cluster can be connected through a network. Among them, the network can be a wide area network, a local area network, or the like. FIG. 9 illustrates a possible implementation approach. As shown in FIG. 9, two computing devices 700A and 700B are connected through a network. Specifically, the computing devices 700A and 700B are connected to the network through their communication interfaces.
[0220] It should be understood that the functionality of the computing device 700A shown in FIG. 9 can also be accomplished by multiple computing devices 700. Similarly, thefunctionality of computing device 700B can also be accomplished by multiple computing devices 700.
[0221] An embodiment of this application provides a computer program product including instructions, which can run on a computing devices cluster or be stored in any available medium. When it is run by a computing device cluster, the computing device cluster is made to execute the methods provided above, or the computing device cluster is made to implement the functions of the devices provided above.
[0222] An embodiment of this application provides a computer readable storage medium including instructions. The computer readable storage medium is any available medium that computing devices can store, or a data storage device such as a data center containing one or more available media. The available media can be magnetic media (such as floppy disks, hard drives, magnetic tapes), optical media (such as digital video disc (DVD)), or semiconductor media (such as solid-state drives), etc. The computer readable storage medium includes instructions. When the instructions are run on a computer device cluster, the computer device cluster executes the methods provided above.
[0223] An embodiment of this application provides a chip system, where the chip system includes a memory and a processor, the memory is configured to store a computer program, and the processor is configured to invoke the computer program from the memory and run the computer program, so that a server on which a chip is disposed performs the methods provided above.
[0224] A person of ordinary skill in the art may be aware that, in combination with the examples described in the embodiments disclosed in this specification, units and algorithm steps can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed by hardware or software depends on particular applications and design constraints of the technical solutions. A person skilled in the art may use different methods to implement the described functions for each particular application, but it should not be considered that the implementation goes beyond the scope of this application.
[0225] It may be clearly understood by a person skilled in the art that, for the purpose of convenient and brief description, for a detailed working process of the foregoing system,apparatus, and unit, refer to a corresponding process in the foregoing method embodiment. Details are not described herein again.
[0226] In the several embodiments provided in this application, it should be understood that the disclosed system, apparatus, and method may be implemented in other manners. For example, the described apparatus embodiment is merely an example. For example, the unit division is merely logical function division and may be other division in actual implementation. For example, a plurality of units or components may be combined or integrated into another system, or some features may be ignored or not performed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections may be implemented through some interfaces. The indirect couplings or communication connections between the apparatuses or units may be implemented in electronic, mechanical, or other forms.
[0227] The units described as separate parts may be or may not be physically separate, and parts displayed as units may be or may not be physical units, may be located in one position, or may be distributed on a plurality of network units. Some or all of the units may be selected based on actual requirements to achieve the objectives of the solutions of the embodiments.
[0228] In addition, functional units in the embodiments of this application may be integrated into one processing unit, or each of the units may exist alone physically, or two or more units are integrated into one unit.
[0229] When the functions are implemented in a form of a software functional unit and sold or used as an independent product, the functions may be stored in a computer readable storage medium. Based on such an understanding, the technical solutions in this application essentially, or the part contributing to the prior art, or some of the technical solutions may be implemented in a form of a software product. The computer software product is stored in a storage medium, and includes several instructions for instructing a computer device (which may be a personal computer, a server, a network device, or the like) to perform all or some of the steps of the methods described in the embodiments of this application. The foregoing storage medium includes: any medium that can store program code, such as a USB flash drive, a removable hard disk, a read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), a magnetic disk, or an optical disc.
[0230] The foregoing descriptions are merely specific implementations of this application,but are not intended to limit the protection scope of this application. Any variation or replacement readily figured out by a person skilled in the art within the technical scope disclosed in this application shall fall within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.
Claims
CLAIMSWhat is claimed is:
1. A method of searching data, wherein the method is applied to a cloud management platform, the cloud management platform manages the infrastructure used to provide cloud services, and the method comprises: receiving a search request of a tenant; obtaining a set of search terms by splitting the search request; obtaining a set of search results according to the set of search terms and an index table, wherein the set of search results comprises search results with different data types, the index table comprises a set of preset terms and identification information of a set of media data corresponding to each preset term in the set of preset terms, and the set of media data corresponding to each preset term has different data types; and providing the set of search results to the tenant.
2. The method according to claim 1, wherein the data types comprise: text, image, video, or audio.
3. The method according to claim 1 or 2, wherein a frequency of each search term in the set of search terms appearing in a preset database is lower than a first preset threshold, and the preset database comprises at least one document and a set of historical search requests of tenant.
4. The method according to any one of claims 1-3, wherein the obtaining a set of search results according to the set of search terms and an index table comprises: searching the index table based on a first search term in the set of search terms; when there is a first preset term in the index table that matches the first search term, obtaining the set of search results based on a set of media data corresponding to the first preset term; or when there is not a preset term in the index table that matches the first search term, obtaining the set of search results according to a media database and the first search term, wherein the media database comprises multiple media data with different data types, and each search result in the set of search results belongs to the media database and corresponds to thefirst search term.
5. The method according to claim 4, wherein the index table also comprises relevance between each preset term and the media data corresponding to each preset term, and the identification information of multiple media data corresponding to each preset term is sorted in descending order of relevance in the index table.
6. The method according to claim 4 or 5, wherein the method further comprises: obtaining the set of preset terms by splitting a set of historical search requests; and generating the index table according to the set of preset terms and the media database.
7. The method according to claim 6, wherein the generating the index table according to the set of preset terms and the media database comprises: obtaining a set of media data in the media database corresponding to second preset term according to the second preset term, wherein the second preset term belongs to the set of preset terms, the media data corresponding to the second preset term comprises the second preset term, or semantic information in the media data corresponding to the second preset term matches the second preset term; calculating relevance between the second preset term and the media data corresponding to the second preset term; and generating the index table according to the second preset term, the media data corresponding to the second preset term and the relevance between the second preset term and the media data corresponding to the second preset term.
8. The method according to any one of claims 4-7, wherein when there is not a preset term in the index table that matches the first search term, the method further comprises: adding the first search term and the identification information of the media data corresponding to the first search term to the index table.
9. The method according to any one of claims 1-8, wherein the obtaining a set of search terms by splitting the search request comprises: obtaining the set of search terms by splitting the search request based on a tokenizer model, wherein the tokenizer model is trained based on a set of historical search requests.
10. The method according to claim 9, wherein the method further comprises: obtaining the tokenizer model by training the initial tokenizer model based on the set ofhistorical search requests, wherein a frequency of a term obtained by the tokenizer model appearing in a preset database is lower than a first preset threshold, and the preset database comprises at least one document and the set of historical search requests.
11. A method of searching data, wherein the method is applied to a cloud management platform, the cloud management platform manages the infrastructure used to provide cloud services, and the method comprises: receiving a set of historical search requests of a tenant; obtaining a set of preset terms by splitting the set of historical search requests; and generating an index table according to the set of preset terms and a media database, wherein the index table comprises the set of preset terms and identification information of media data in the media database corresponding to each preset term in the set of preset terms, and the index table is used to obtain a set of search results based on a search request.
12. The method according to claim 11, wherein a frequency of each preset term in the set of preset terms appearing in a preset database is lower than a first preset threshold, and the first preset database comprises at least one document and the set of historical search requests.
13. The method according to claim 11 or 12, wherein the index table also comprises relevance between each preset term and the media data corresponding to each preset term, and the identification information of multiple media data corresponding to each preset term is sorted in descending order of relevance in the index table.
14. The method according to any one of claims 11-13, wherein the generating an index table according to the set of preset terms and a media database comprises: obtaining a set of media data in the media database corresponding to second preset term according to the second preset term, wherein the second preset term belongs to the set of preset terms, the media data corresponding to the second preset term comprises the second preset term, or semantic information in the media data corresponding to the second preset term matches the second preset term; calculating relevance between the second preset term and the media data corresponding to the second preset term; and generating the index table according to the second preset term, the media data corresponding to the second preset term and the relevance between the second preset term andthe media data corresponding to the second preset term.
15. The method according to any one of claims 11-14, wherein the obtaining a set of preset terms by splitting the set of historical search requests comprises: obtaining the set of preset terms by splitting the set of historical search requests based on a tokenizer model, wherein the tokenizer model is trained based on multiple historical search requests in the set of historical search requests.
16. The method according to claim 15, wherein the method further comprises: obtaining the tokenizer model by training the initial tokenizer model based on multiple historical search requests in the set of historical search requests, wherein a frequency of a term obtained by the tokenizer model appearing in a preset database is lower than a first preset threshold, and the preset database comprises at least one document and the multiple historical search requests in the set of historical search requests.
17. The method according to any one of claims 11-16, wherein the media database comprises multiple media data with different data types, and a set of media data in the index table corresponding to each preset term has different data types.
18. The method according to claim 17, wherein the data types comprise: text, image, video, or audio.
19. A device of searching data, comprising units to perform the method according to any one of claims 1-18.
20. A computing device cluster, comprising at least one computing device, wherein the computing device comprises a processor and a memory coupled with the processor, wherein the memory is configured to store a computer program, and the processor is configured to invoke and run the computer program stored in the memory, so that the computing device executes the method according to any one of claims 1-18.
21. A computer program product, wherein when the computer program product is run on a server, the server is enabled to perform the method according to any one of claims 1-18.
22. A computer readable storage medium storing instructions that, when run on a server, enable the server to perform the method according to any one of claims 1-18.
Citation Information
Patent Citations
Search request processing method and device
CN113139113A