A vertical search method, device, electronic device and storage medium

By extending the training data in the service knowledge graph and training the semantic matching model, the problem of poor training effect of vertical search model is solved, and the correlation between search results and search keywords is improved.

CN113821711BActive Publication Date: 2025-05-13TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110640134.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-09
Publication Date
2025-05-13
Estimated Expiration
2041-06-09

AI Technical Summary

Technical Problem

In the prior art, the vertical search model training effect is poor, resulting in low correlation between search results and search keywords. Especially in service search, the model training effect is worse due to the large and complex domain knowledge.

Method used

By extending the initial training data using entity associations in the service knowledge graph, a diverse extended data set is generated and a semantic matching model is trained based on this data set.

Benefits of technology

It improves data differences, enhances model training effect and accuracy, and improves the correlation between search results and search keywords.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113821711B_ABST
    Figure CN113821711B_ABST
Patent Text Reader

Abstract

The present application relates to the field of computer technology, and provides a vertical search method, device, electronic device and storage medium for improving the model training effect, thereby improving the vertical search relevance, wherein the method includes: after receiving a vertical search request, inputting the target search keyword contained in the vertical search request into a semantic matching model trained based on an extended data set, obtaining search results of a set search type, and then returning a search result page containing the search results. In this way, the initial training data is expanded through the entity association relationship contained in the service knowledge graph, which increases the data diversity, thereby improving the model training effect and improving the relevance of the search results and the search keywords.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] With the development of search engine technology, vertical search has been widely used to meet the diversified search needs. Vertical search is a professional search service proposed for a specific field, such as public account search, mini program search, service search, etc. Among them, service search includes services in various fields such as education, government affairs, and housekeeping. For example, when a user enters the search keyword "nanny", a list of services that can provide nanny services can be displayed in the search results.

[0003] Since the amount of training data is limited, the training data needs to be expanded before training the model used for vertical search. In related technologies, data expansion is performed through historical behavior data. For example, search results with high click-through rates are used as positive samples, and search results with low click-through rates are used as negative samples. However, when using the above data expansion scheme, it is difficult to learn the correlation between search keywords and search results, resulting in poor model training results and low model accuracy. Especially in service search, because it involves multiple fields with a lot of complex domain knowledge, the model training effect is even worse. Summary of the invention

[0004] The embodiments of the present application provide a service search method, device, electronic device and storage medium to solve the problems in the related art of poor model training effect and low correlation between search results and search keywords.

[0005] In a first aspect, an embodiment of the present application provides a vertical search method, including:

[0006] receiving a vertical search request, wherein the vertical search request includes at least one target search keyword;

[0007] Inputting the at least one target search keyword into a trained semantic matching model to obtain a search result set of a set search type; wherein the semantic matching model is trained based on an extended data set, and the extended data set is obtained by expanding the initial training data based on the entity association relationship contained in a preset service knowledge graph;

[0008] A search results page containing the search result set is returned.

[0009] In a second aspect, an embodiment of the present application provides a vertical search device, including:

[0010] A receiving unit, configured to receive a vertical search request, wherein the vertical search request includes at least one target search keyword;

[0011] A search unit, configured to input the at least one target search keyword into a trained semantic matching model to obtain a search result set of a set search type; wherein the semantic matching model is trained based on an extended data set, and the extended data set is obtained by expanding initial training data based on entity associations contained in a preset service knowledge graph;

[0012] The sending unit is used to return a search result page including the search result set.

[0013] Optionally, after combining a historical search result included in a piece of training data with each second adjacent entity in a set of second adjacent entities associated with the historical search result to obtain a third set of extended data, and before obtaining the extended data set based on the at least one set of extended data, the extension unit is further configured to:

[0014] If the entity association relationship between the historical search result and a second adjacent entity indicates that the historical search result and the second adjacent entity are in a hierarchical relationship, then the label of the extended data obtained by combining the historical search result and the second adjacent entity is set to be first-level related;

[0015] If the entity association relationship between the historical search result and a second adjacent entity indicates that the historical search result and the second adjacent entity are in a synonym relationship, then the label of the extended data obtained by combining the historical search result and the second adjacent entity is set as a secondary correlation;

[0016] Each tag is used to represent the real relevance between the corresponding historical search keyword and the corresponding historical search result, and the relevance of the first-level relevance representation is lower than the relevance of the second-level relevance representation.

[0017] Optionally, when obtaining the extended data set based on the at least one set of extended data, the extension unit is specifically configured to:

[0018] Extracting a fifth set of training data from the first set of extended data and the second set of extended data according to the set extraction ratio;

[0019] The extended data set is obtained based on the third set of extended data, the fourth set of extended data and the fifth set of extended data.

[0020] Optionally, when inputting the at least one target search keyword into a trained semantic matching model to obtain a search result set of a set search type, the search unit is further configured to:

[0021] Obtaining the relevance of each search result contained in the search result set and the at least one target search keyword, each relevance being used to characterize the relevance between the corresponding search result and the at least one target search keyword;

[0022] When returning the search result page including the search result set, the sending unit is specifically used for:

[0023] Based on the obtained relevances, the search results are ranked, and a search result page containing the ranked search results is returned.

[0024] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above-mentioned vertical search method.

[0025] An embodiment of the present application provides a computer-readable storage medium, which includes a computer program. When the computer program is executed on an electronic device, the computer program is used to enable the electronic device to execute the steps of the vertical search method.

[0026] In the embodiment of the present application, after receiving a vertical search request, the target search keyword contained in the vertical search request is input into the semantic matching model trained based on the extended data set to obtain the search results of the set search type, and then the search results page containing the search results is returned. In this way, the initial training data is expanded through the entity association relationship contained in the service knowledge graph, and diversified extended data can be obtained, long-tail data is increased, thereby improving data differentiation. When the semantic matching model is subsequently trained based on the extended data set, the model training effect is improved, the model accuracy is improved, and the relevance of the search results to the search keywords is improved.

[0027] Other features and advantages of the present application will be described in the following description, and partly become apparent from the description, or be understood by practicing the present application. The purpose and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0029] Figure 1A A schematic diagram of a possible application scenario provided in an embodiment of the present application;

[0030] Figure 1B An optional schematic diagram of a blockchain structure provided in an embodiment of the present application;

[0031] Figure 1C A flowchart of a block generation method provided in an embodiment of the present application;

[0032] Figure 2 A schematic diagram of a vertical search method provided in an embodiment of the present application;

[0033] Figure 3 A schematic diagram of an operation interface provided in an embodiment of the present application;

[0034] Figure 4A A schematic diagram of entity attributes provided in an embodiment of the present application;

[0035] Figure 4B A schematic diagram of a service knowledge graph provided in an embodiment of the present application;

[0036] Figure 5 A schematic diagram of a search result provided in the implementation of this application;

[0037] Figure 6 A schematic diagram of a flow chart of a method for obtaining an extended data set provided in an embodiment of the present application;

[0038] Figure 7 A logical diagram of extended data provided in an embodiment of the present application;

[0039] Figure 8 A schematic diagram of the structure of a vertical search device provided in an embodiment of the present application;

[0040] Fig. 9 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application;

[0041] Fig.10 A schematic diagram of the hardware composition structure of a terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the technical solution of the present application, rather than all of the embodiments. Based on the embodiments recorded in the application documents, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the technical solution of the present application.

[0043] The following is an introduction to some concepts involved in the embodiments of the present application.

[0044] 1. Vertical search is a professional search engine for a specific field, which is a subdivision and extension of the search engine. It integrates information in a specific field in the database, extracts data by field, processes the data and returns it to the terminal device. Taking WeChat's search function as an example, public account search, mini program search, etc. are all vertical searches.

[0045] 2. Service search is used to search for services related to the search keyword, which is a vertical search. Taking the search keyword "nanny" as an example, service search can directly provide a nanny service menu.

[0046] 3. Semantic matching model, which is used to obtain the correlation between search keywords (query) and search results (doc). Correlation is used to describe the degree of mutual connection between two things.

[0047] For example, referring to Table 1, when the search keyword is "Human Papilloma Virus (HPV) vaccine appointment", if the search results include "Cervical cancer vaccine appointment" and "Children's vaccine appointment", since children's vaccine appointment is equivalent to cervical cancer vaccine, which is an adult vaccine and not a children's vaccine, the search keyword "HPV vaccine appointment" is related to the search result "Cervical cancer vaccine appointment", and the search keyword "HPV vaccine appointment" is not related to the search result "Children's vaccine appointment".

[0048] For another example, still referring to Table 1, when the search keyword is "medical insurance payment inquiry", if the search results include "medical insurance information inquiry" and "medical insurance designated point inquiry", since payment inquiry is a type of information inquiry, the search keyword "medical insurance payment inquiry" is related to the search result "medical insurance information inquiry", and the search keyword "medical insurance payment inquiry" is not related to the search result "medical insurance designated point inquiry".

[0049] For another example, still referring to Table 1, when the search keyword is "personal tax inquiry", if the search results include "tax processing" and "general taxpayer inquiry", since general taxpayers refer to corporate taxes rather than individuals, the search keyword "personal tax inquiry" is related to the search result "tax processing", and the search keyword "personal tax inquiry" is not related to the search result "general taxpayer inquiry".

[0050] Table 1 Correlation between search keywords and search results

[0051]

[0052]

[0053] In an embodiment of the present application, the semantic matching model may adopt but is not limited to a convolutional kernel-based neural ranking model (Convolutional Kernel-based Neural Ranking Model, Conv-KNRM), a bidirectional encoder representation from transformers (Bidirectional Encoder Representations from Transformers, BERT), a deep semantic model (Deep Structured Semantic Models, DSSM), a convolutional latent semantic model (Convolutional latent semantic model, CDSSM or CLSM), a match-pyramid, a deep relevance matching model (Deep Relevance Matching Model, DRMM), etc.

[0054] The preferred embodiments of the present application are described below in conjunction with the drawings in the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. In addition, the embodiments and features in the embodiments of the present application may be combined with each other if there is no conflict.

[0055] Vertical search is a professional search service for a specific field, such as public account search, mini-program search, service search, etc. Taking WeChat search as an example, when the user enters the search keyword "nanny", a list of services that can provide nanny services can be displayed in the search results.

[0056] Vertical search is usually implemented based on semantic matching models, but due to limited training data, it is necessary to expand the training data to increase the training data. In related technologies, data expansion is performed through historical behavior data. For example, search results with high click-through rates are used as positive samples, and search results with low click-through rates and non-service types are used as negative samples.

[0057] However, since vertical search is designed for a specific field, if the above data expansion solution is used for fields containing a large amount of complex and non-universal corpus, the long-tail data obtained will be small, resulting in model overfitting and poor model training effect, which in turn leads to low correlation between search results and search keywords. Especially in service search, since service search involves multiple fields with large differences in data distribution and the domain knowledge is large and complex, the model training effect is even worse.

[0058] For example, referring to Table 2, in the related art, due to the data augmentation and expansion scheme based on historical behavior, the training data obtained cannot reflect that the correlation between the search keywords and the search results is positively correlated, and it is difficult to learn the precise correlation between the search keywords and the search results. Therefore, the real correlation between the search keyword "check violation" and the search result "traffic violation query" is higher than the real correlation between the search keyword "traffic road condition query" and the search result "traffic violation query". However, the predicted correlation between the search keyword "check violation" and the search result "traffic violation query" is lower than the predicted correlation between the search keyword "traffic road condition query" and the search result "traffic violation query". Obviously, the model accuracy in the related art is low.

[0059] Table 2 Prediction correlation between search keywords and search results in related technologies

[0060] Search Keywords Search Results Predicting relevance True correlation Check for violations Traffic violation inquiry 0.90 powerful Traffic Conditions Query Traffic violation inquiry 0.92 none Check Violations Traffic violation inquiry 0.95 powerful Driver's license renewal Driver's license score check 0.95 weak

[0061] In the embodiment of the present application, after receiving the vertical search request, the target search keyword contained in the vertical search request is input into the semantic matching model trained based on the extended data set to obtain the search results of the set search type, and then the search results page containing the search results is returned. In this way, the initial training data is expanded through the entity association relationship contained in the service knowledge graph, and diversified extended data can be obtained, which increases the long-tail data and improves the data difference. When the semantic matching model is subsequently trained based on the extended data set, the model training effect is improved, the model accuracy is improved, and the relevance of the search results and the search keywords is improved.

[0062] The embodiments of the present application relate to artificial intelligence (AI) and machine learning technology, and are designed based on voice technology and machine learning (ML) in artificial intelligence.

[0063] Artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0064] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0065] Machine learning is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.

[0066] The embodiment of the present application adopts a semantic matching model of machine learning when obtaining search results based on target search keywords. The method for training a semantic matching model proposed in the embodiment of the present application can be divided into two parts, including a training part and an application part; wherein, the training part involves the technical field of machine learning. In the training part, the semantic matching model is trained by the technology of machine learning, and the extended data set given in the embodiment of the present application is used as a training data set to train the semantic matching model. After the extended data in the extended data set is input into the semantic matching model, the output result of the semantic matching model is obtained, and the model parameters are continuously adjusted through the optimization algorithm in combination with the output result; the application part is used to detect the target search keywords using the semantic matching model obtained by training in the training part, and obtain the search results of the target search keywords. In addition, it should be noted that in the embodiment of the present application, the target search keywords can be trained online or offline, which is not specifically limited here.

[0067] The preferred embodiments of the present application are described below in conjunction with the drawings in the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. In addition, the embodiments of the present application and the features in the embodiments may be combined with each other if there is no conflict.

[0068] See also Figure 1A As shown, it is a schematic diagram of a possible application scenario in an embodiment of the present application. The application scenario includes a terminal device 110, a server 120 and a data sharing system 130. The terminal device 110, the server 120 and the data sharing system 130 communicate with each other through a communication network.

[0069] In a possible implementation, the communication network is a wired network or a wireless network. The terminal device 110, the server 120, and the data sharing system 130 may be directly or indirectly connected via wired or wireless communication, which is not limited in this application.

[0070] The user logs in to the application operation interface through the terminal device 110. The terminal device 110 sends a vertical search request to the server 120 in response to the operation triggered by the user in the application operation interface, so that the server 120 performs a vertical search based on at least one target search keyword contained in the vertical search request. For example, based on the vertical search request containing at least one target search keyword, the server inputs at least one target search keyword into the trained semantic matching model to obtain the search results of the set search type. Exemplarily, after responding to the user operation, the terminal device 110 can also receive and present the search result page containing the search results returned by the server 120.

[0071] In the embodiment of the present application, the application can be a social software, such as an instant messaging software, a short video software, or a small program, a web page, etc., which is not specifically limited here. Among them, the terminal device 110 is installed with an application, and the server 120 is a server corresponding to the software or web page, small program, etc.

[0072] In the embodiment of the present application, the terminal device 110 is an electronic device used by the user, which may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. Each terminal device 110 is connected to the server 120 via a wireless network. The server 120 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. The terminal and the server may be directly or indirectly connected via wired or wireless communication, which is not limited in this application.

[0073] The data sharing system 130 communicates with the server 120 via a communication network. Exemplarily, the data sharing system 130 is used to store a preset service knowledge graph, an initial training data set, and an extended data set.

[0074] The data sharing system 130 refers to a system for sharing data between nodes, and the data sharing system may include multiple nodes 131, and the multiple nodes 131 may refer to each client in the data sharing system. Each node 131 can receive input information when performing normal work, and maintain the shared data in the data sharing system based on the received input information. In order to ensure the information intercommunication within the data sharing system, there may be an information connection between each node in the data sharing system, and information can be transmitted between nodes through the above information connection. For example, when any node in the data sharing system receives input information, other nodes in the data sharing system obtain the input information according to the consensus algorithm, and store the input information as data in the shared data, so that the data stored on all nodes in the data sharing system are consistent.

[0075] Each node in the data sharing system has a node identifier corresponding to it, and each node in the data sharing system can store the node identifiers of other nodes in the data sharing system, so that the generated blocks can be broadcast to other nodes in the data sharing system according to the node identifiers of other nodes. A node identifier list as shown in the following table can be maintained in each node, and the node name and node identifier are stored in the node identifier list accordingly. Among them, the node identifier can be an Internet Protocol (IP) address and any other information that can be used to identify the node. Table 3 only takes the IP address as an example for illustration.

[0076] Table 3 Node identification list

[0077] Node Name Node ID Node 1 117.114.151.174 Node 2 117.116.189.145 …… …… Node N 119.123.789.258

[0078] Each node in the data sharing system stores the same blockchain. The blockchain consists of multiple blocks, see Figure 1B The blockchain consists of multiple blocks. The genesis block includes a block header and a block body. The block header stores the input information characteristic value, version number, timestamp and difficulty value, and the block body stores the input information; the next block of the genesis block takes the genesis block as the parent block, and the next block also includes a block header and a block body. The block header stores the input information characteristic value of the current block, the block header characteristic value, version number, timestamp and difficulty value of the parent block, and so on, so that the block data stored in each block in the blockchain is associated with the block data stored in the parent block, ensuring the security of the input information in the block.

[0079] When generating each block in the blockchain, see Figure 1CWhen the node where the blockchain is located receives the input information, it verifies the input information. After the verification is completed, the input information is stored in the memory pool and the hash tree used to record the input information is updated. After that, the update timestamp is updated to the time when the input information is received, and different random numbers are tried to calculate the eigenvalue multiple times so that the calculated eigenvalue can satisfy the following formula:

[0080] SHA256(SHA256(version+prev_hash+merkle_root+ntime+nbits+x))

[0081] <TARGET

[0082] Among them, SHA256 is the eigenvalue algorithm used to calculate the eigenvalue; the version number (version) is the version information of the relevant block protocol in the blockchain; prev_hash is the block header eigenvalue of the parent block of the current block; merkle_root is the eigenvalue of the input information; ntime is the update time of the update timestamp; nbits is the current difficulty, which is a fixed value within a period of time and is determined again after exceeding the fixed time period; x is a random number; TARGET is the eigenvalue threshold, which can be determined based on nbits.

[0083] In this way, when the random number that satisfies the above formula is calculated, the information can be stored accordingly, the block header and block body can be generated, and the current block can be obtained. Subsequently, the node where the blockchain is located sends the newly generated block to other nodes in the data sharing system according to the node identification of other nodes in the data sharing system. Other nodes verify the newly generated block and add the newly generated block to the blockchain stored in them after the verification is completed.

[0084] See also Figure 2 As shown, it is an implementation flow chart of a vertical search method provided in an embodiment of the present application, which is applied to a vertical search device. The vertical search device can be a server 120 or a device deployed on the server 120. The specific implementation process of the method is as follows:

[0085] S201. A vertical search device receives a vertical search request, where the vertical search request includes at least one target search keyword.

[0086] In the embodiment of the present application, the vertical search request may be sent by the terminal device to the vertical search apparatus in response to an input operation triggered in the operation interface. The input operation includes but is not limited to a touch operation, a mouse operation, a keyboard operation, and the like.

[0087] See also Figure 3As shown, it is a possible operation interface provided in an embodiment of the present application, in which a search box, a "friend circle" button, an "article" button, a "public account" button, a "music" button, a "mini program" button, a "music" button, and an "expression" button are included for receiving target search keywords input by the user. Among them, the "friend circle" button is used to provide a friend circle search service, the "article" button is used to provide an article search service, and the "public account" button is used to provide a public account search service. Similarly, the functions of the "mini program" button, the "music" button, and the "expression" button are not repeated. After the user enters the target search keyword, search results of search types such as friend circle, article, and public account can be obtained.

[0088] For example, suppose that the user inputs the target search keyword "babysitter" in the operation interface, and the terminal device responds to the input operation triggered in the operation interface and sends a vertical search request 1 to the vertical search device. The vertical search device receives the vertical search request 1, wherein the vertical search request includes the target search keyword "babysitter".

[0089] For example, suppose that the user enters the target search keyword "air conditioner cleaning" in the operation interface, and the terminal device responds to the input operation triggered in the operation interface and sends a vertical search request 2 to the vertical search device. The vertical search device receives the vertical search request 2, wherein the vertical search request contains the target search keyword "air conditioner cleaning".

[0090] Considering that if the user inputs the target search keyword in Chinese, the vertical search device needs to segment the text to be processed input by the user to obtain at least one target search keyword. Specifically, the vertical search request carries the text to be processed. After receiving the vertical search request, the vertical search device uses a preset segmentation tool to segment the text to obtain at least one target search keyword. Among them, the preset segmentation tool can be used but not limited to Jieba segmentation, Language Technology Platform (LTP) segmentation, etc.

[0091] For example, assuming that the text to be processed in the vertical search request is "air conditioning cleaning", after the vertical search device receives the vertical search request, it uses jieba word segmentation to obtain the target search keywords "air conditioning" and "cleaning".

[0092] S202. The vertical search device inputs at least one target search keyword into a trained semantic matching model to obtain a search result set of a set search type; wherein the semantic matching model is trained based on an extended data set, and the extended data set is obtained by expanding the initial training data based on the entity association relationship contained in a preset service knowledge graph.

[0093] The set search type includes but is not limited to one or more of public accounts, services, articles, music, friend circles, emoticons, etc.

[0094] The service knowledge graph refers to a knowledge graph specifically built for services in scenarios where services are provided to users. The service knowledge graph may contain one or more category systems, such as housekeeping categories, education categories, transportation categories, etc. In the following, the service knowledge graph is described using only the housekeeping category as an example.

[0095] Each category system contains various entities, and the attributes of each entity include but are not limited to one of the following: service entity word, service behavior word, service status word, service brand, etc. Exemplarily, there is an entity association relationship between each entity, and the entity association relationship may include but is not limited to a hyponym relationship, a synonym relationship, an extended word relationship, etc.

[0096] See also Figure 4A As shown, the service knowledge graph includes service entity words: "home appliances" and "air conditioners", among which "home appliances" and "air conditioners" are hierarchical words, that is, the entity association relationship between "home appliances" and "air conditioners" is a hierarchical relationship, and "home appliances" is the superordinate entity of "air conditioners"; it includes service action words: "cleaning" and "cleaning", among which "cleaning" and "cleaning" are synonyms, that is, the entity association relationship between "cleaning" and "cleaning" is a synonym relationship; it includes service status words: "home"; and it includes service brands: "home appliance brand A".

[0097] See also Figure 4B As shown, it is a possible schematic diagram of the service knowledge graph provided in the embodiment of the present application. Figure 4B In the service knowledge graph, the following entities are included: "housekeeping category", "nanny", "cleaning", "cleaning", "home appliances", "washing machine", "air conditioner", "home appliance brand A", etc. Among them, the entity association relationship between "housekeeping category" and "nanny" is a hierarchical relationship, the entity association relationship between "housekeeping category" and "cleaning" is a hierarchical relationship, the entity association relationship between "cleaning" and "cleaning" is a synonym relationship, the entity association relationship between "home appliances" and "washing machine" is a hierarchical relationship, and the entity association relationship between "home appliances" and "air conditioner" is a hierarchical relationship.

[0098] For example, suppose that the vertical search device inputs the target search keyword "nanny" into the trained semantic matching model and obtains a set of search results of service type and public account type, wherein the search results of service type include "A nanny service", "B nanny service", "C nanny service", etc., and the search results of public account type include "A public account", "B public account", etc.

[0099] S203: The vertical search device returns a search result page including a search result set.

[0100] For example, see Figure 5 As shown, the vertical search device returns a search result page containing a search result set, and the search result page contains at least search results of service types: "A nanny service", "B nanny service", "C nanny service", and search results of public account types: "A public account".

[0101] To improve search efficiency, in some embodiments, the vertical search device inputs at least one target search keyword into a trained semantic matching model, and when obtaining search results of a set search type, also obtains the relevance of each search result contained in the search result set to the at least one target search keyword, and then the vertical search device returns a search result page containing the search result set, including:

[0102] Based on the obtained relevance levels, the search results are ranked, and a search result page containing the ranked search results is returned.

[0103] Each relevance is used to represent the relevance between the corresponding search results and at least one target search keyword. The relevance can be expressed in numerical values ​​or in levels. In the following, only the level is used as an example for explanation.

[0104] For example, assuming that the relevance representation of the search result "A nanny service" and the target search keyword "nanny" is that the search result "A nanny service" is strongly correlated with the target search keyword "nanny", and the relevance representation of the search result "A public account" and the target search keyword "nanny" is that the search result "A public account" is weakly correlated with the target search keyword "nanny", then based on the obtained relevances, the search results "A nanny service" and the search result "A public account" are sorted, and the sorted search results are: "A nanny service", "A public account", and then the search result page containing the sorted search results is returned.

[0105] To avoid model overfitting, in some embodiments, see Figure 6 As shown, the extended dataset can be obtained by, but not limited to, the following methods:

[0106] S601. The vertical search device obtains an initial training data set, in which each piece of training data includes at least one historical search keyword and a corresponding historical search result.

[0107] The following description only takes the training data x as an example, where the training data x is any one of the training data.

[0108] For example, the training data x includes historical search keywords "air conditioner" and "cleaning" and the corresponding historical search results "air conditioner cleaning service".

[0109] It should be noted that in the embodiment of the present application, considering that the user inputs historical search keywords in Chinese, the vertical search device can also use the word segmentation tool mentioned above to obtain at least one historical search keyword contained in each training data. Since the process of obtaining the historical search keywords is similar to the process of obtaining the target search keywords, it will not be repeated here.

[0110] S602. The vertical search device determines from the service knowledge graph, based on the entity association relationships contained in the preset service knowledge graph, the first adjacent entity sets associated with each historical search keyword and the second adjacent entity sets associated with each historical search result contained in the initial training data set.

[0111] In order to improve the efficiency of model training, in some embodiments, the vertical search device can perform the following operations for each historical search keyword included in the initial training data set to obtain the first adjacent entity set associated with each historical search keyword. Specifically, taking the historical search keyword x as an example, the historical search keyword x is any one of the historical search keywords, and the following operations are performed for the historical search keyword x:

[0112] The vertical search device obtains a set of candidate adjacent entities associated with the historical search keyword x based on the entity association relationship contained in the preset service knowledge graph, wherein the set of candidate adjacent entities associated with the historical search keyword x includes each adjacent entity of the historical search keyword x in the service knowledge graph;

[0113] The vertical search device deletes a target adjacent entity from each candidate adjacent entity included in the obtained candidate adjacent entity set, wherein the target adjacent entity is: an entity having a subordinate entity in the candidate adjacent entity set based on the service knowledge graph;

[0114] The vertical search device uses the candidate adjacent entity set from which the target adjacent entity is deleted as the first adjacent entity set associated with the historical search keyword x.

[0115] Taking the historical search keyword x as "home appliances" as an example, based on the entity association relationship contained in the service knowledge graph, the candidate adjacent entity set associated with "home appliances" is obtained as {"washing machine", "air conditioner", "cleaning"}. Then, the vertical search device deletes the target adjacent entity "cleaning" from "washing machine", "air conditioner", and "cleaning". Then, the vertical search device uses the candidate adjacent entity set from which the target adjacent entity has been deleted as the first adjacent entity set associated with "home appliances": {"washing machine", "air conditioner"}.

[0116] It should be noted that, in the embodiment of the present application, the vertical search device can adopt but is not limited to the N-gram model (N-Gram model) to determine the entity corresponding to each target search keyword in the service knowledge graph. For the sake of description, in the embodiment of the present application, the entity corresponding to the target search keyword in the service knowledge graph is referred to as the target search keyword.

[0117] Since the process of obtaining the second adjacent entity set associated with each historical search result is similar to the process of obtaining the first adjacent entity set associated with each historical search keyword, it will not be described in detail here.

[0118] S603: The vertical search device composes at least one set of extended data based on each of the first adjacent entity sets and each of the second adjacent entity sets obtained; wherein the at least one set of extended data is different from each of the training data.

[0119] In order to further increase the diversity of training data, improve data differences, and enhance model training effects, in some embodiments, when executing S603, the vertical search device performs at least one of the following operations for each piece of training data included in the initial training data set:

[0120] The following only takes the training data x and the training data y as examples for explanation.

[0121] The training data x contains a historical search keyword (A), and the training data x can be expressed as<A,B> , where A is used to represent the historical search keywords contained in the training data x, and B is used to represent the historical search results contained in the training data x. The first adjacent entity set associated with A is {A1, A2, A3, ..., AN A}, the first adjacent entity set associated with B is {B1, B2, B3, ..., BN B}, where N A 、N B All are positive integers.

[0122] The training data y contains two historical search keywords (C and D). The training data y can be expressed as<C、D,E> , where C and D are used to represent the two historical search keywords contained in the training data y, and E is used to represent the historical search results contained in the training data y. The first adjacent entity set associated with C is {C1, C2, C3, ..., CN C}, the first adjacent entity set associated with D is {D1, D2, D3, ..., DN D}, the first adjacent entity set associated with E is {E1, E2, E3, ..., EN E}, where N C、N D 、N E All are positive integers.

[0123] Operation 1: The vertical search device combines at least one historical search keyword included in the training data x with each second adjacent entity in the second adjacent entity set associated with the corresponding historical search result to obtain a first set of extended data.

[0124] Take the training data x as an example, see Figure 7 As shown, the vertical search device combines A with {B1, B2, B3, ..., BN B} to obtain the first set of extended data:<A,B1> ,<A,B2> ,……, <A,BN B >.

[0125] It should be noted that, in the embodiment of the present application, if the historical search results include multiple keywords, at least one historical search keyword is combined with each second adjacent entity in the second adjacent entity set associated with each of the multiple keywords.

[0126] Taking the training data x as <"cleaning", "home appliance cleaning"> as an example, "home appliance cleaning" contains "home appliances" and "cleaning", among which the second adjacent entity set associated with "home appliances" is {"air conditioning", "washing machine"}, and the second adjacent entity set associated with "cleaning" is {"cleaning", "clearing"}. The vertical search device combines "cleaning" with {"air conditioning", "washing machine"} {"cleaning", "clearing"} respectively to obtain the first set of extended data: <"cleaning", "air conditioning">, <"cleaning", "washing machine">, <"cleaning", "cleaning">, <"cleaning", "clearing">.

[0127] Taking the training data y as an example, the vertical search device compares C with {E1, E2, E3, ..., EN E}, and D with {E1, E2, E3, ..., EN E} to obtain the first set of extended data:<C,E1> ,<C,E2> ,……, <C,EN E >,<D,E1> ,<D,E2> ,……, <D,EN E >.

[0128] Taking the training data y as <"air conditioning", "cleaning", "home appliances"> as an example, the second adjacent entity set associated with "home appliances" is {"air conditioning", "washing machine"}. The vertical search device combines "cleaning" with {"air conditioning", "washing machine"}, and combines "air conditioning" with {"air conditioning", "washing machine"} to obtain the first set of extended data: <"cleaning", "air conditioning">, <"cleaning", "washing machine">, <"air conditioning", "air conditioning">, <"air conditioning", "washing machine">.

[0129] It should be noted that, in the first set of extended data obtained by operation 1, the labels of the extended data are the same as the labels of the corresponding training data. For example, the first set of extended data obtained by expanding the training data x is:<A,B1> ,<A,B2> ,……, <A,BN B >, which is the same as the label of the training data x.

[0130] Operation 2: The vertical search device combines the historical search results contained in the training data x with each first adjacent entity in the first adjacent entity set associated with the corresponding historical search keyword to obtain a second set of extended data.

[0131] Take the training data x as an example, see Figure 7 As shown, the vertical search device connects B with {A1, A2, A3, ..., AN A} to obtain the second set of extended data:<A1,B> ,<A2,B> ,……, <AN A , B>.

[0132] Still taking the training data x as <"cleaning", "home appliance cleaning"> as an example, the first adjacent entity set associated with "cleaning" is {"cleaning", "clearing"}, and the vertical search device combines "home appliance cleaning" with {"cleaning", "clearing"} to obtain the second set of extended data: <"cleaning", "home appliance cleaning">, <"clearing", "home appliance cleaning">.

[0133] Taking the training data y as an example, the vertical search device compares E with {C1, C2, C3, ..., CN C}, and E is combined with {D1, D2, D3, ..., DN D} to obtain the second set of extended data:<C1,E> ,<C2,E> ,……, <CN C , E>,<D1,E> ,<D2,E> ,……, <DN D , E>.

[0134] Still taking the training data y as <"air conditioning", "cleaning", "home appliances"> as an example, the first adjacent entity set associated with "air conditioning" is {"refrigeration equipment"}, and the first adjacent entity set associated with "cleaning" is {"cleaning", "clearing"}. The vertical search device combines "home appliance cleaning" with {"refrigeration equipment"}, and combines "home appliances" with {"cleaning", "clearing"} to obtain the second set of extended data: <"refrigeration equipment", "home appliances">, <"cleaning", "home appliances">, <"clearing", "home appliances">.

[0135] It should be noted that the labels of the respective extended data included in the second set of extended data obtained by operation 2 are the same as the labels of the corresponding training data. For example, the second set of extended data obtained by expanding the training data x is:<A1,B> ,<A2,B> ,……, <AN A , B>, the same as the label of the training data x.

[0136] Operation three: the vertical search device combines the historical search results contained in the training data x with each second adjacent entity in the second adjacent entity set associated with the historical search results to obtain a third set of extended data.

[0137] Take the training data x as an example, see Figure 7 As shown, the vertical search device combines B with {B1, B2, B3, ..., BN B} to obtain the third set of extended data:<B,B1> ,<B,B2> ,……, <B,BN B >.

[0138] Still taking the training data x as <"cleaning", "home appliance cleaning"> as an example, "home appliance cleaning" contains "home appliances" and "cleaning", where the second adjacent entity set associated with "home appliances" is {"air conditioning", "washing machine"}, and the second adjacent entity set associated with "cleaning" is {"cleaning", "clearing"}. The vertical search device combines "home appliance cleaning" with {"air conditioning", "washing machine"}, and combines "home appliance cleaning" with {"cleaning", "clearing"} to obtain the third set of extended data: <"home appliance cleaning", "air conditioning">, <"home appliance cleaning", "washing machine">, <"home appliance cleaning", "cleaning">, <"home appliance cleaning", "clearing">.

[0139] In order to improve the efficiency of model training, in some embodiments, the vertical search device may set a label for the third group of extended data according to the entity association relationship between different entities. Specifically, if the entity association relationship between the historical search result and a second adjacent entity is characterized as: the historical search result and the second adjacent entity are in a hierarchical relationship, then the vertical search device sets the label of the extended data obtained by combining the historical search result and the second adjacent entity as first-level correlation;

[0140] If the entity association relationship between the historical search result and a second adjacent entity indicates that the historical search result and the second adjacent entity are in a synonym relationship, the vertical search device sets the label of the extended data obtained by combining the historical search result and the second adjacent entity as a secondary correlation;

[0141] Among them, each label is used to represent the real relevance between the corresponding historical search keyword and the corresponding historical search result, and the relevance of the first-level relevance representation is lower than the relevance of the second-level relevance representation.

[0142] It should be noted that in the embodiments of the present application, the correlation can be expressed by numerical values ​​or by levels, and the present application does not impose any restrictions on this. In the following, only the example of expressing the correlation by levels is used for explanation.

[0143] Taking the extended data <"home appliances", "air conditioner"> as an example, the entity association relationship between "home appliances" and "air conditioner" is a hierarchical relationship. Therefore, the vertical search device sets the labels of <"home appliances", "air conditioner"> to first-level correlation, where first-level correlation indicates moderate correlation.

[0144] Taking the extended data of <"clean", "clean"> as an example, the entity association relationship between "clean" and "clean" is a synonym relationship. Therefore, the vertical search device sets the label of <"clean", "clean"> to a secondary correlation, where the secondary correlation indicates a strong correlation.

[0145] Operation 4: The vertical search device combines at least one historical search keyword included in the training data x with part of the second adjacent entities in the second adjacent entity set associated with the corresponding historical search results to obtain a fourth set of extended data.

[0146] In order to increase negative samples in the training data to improve the model training effect, in some embodiments, the fourth set of extended data can be obtained by adopting but not limited to at least one of the following methods:

[0147] Method 4-1: The vertical search device determines the target domain type to which the historical search results contained in the training data x belong, and filters out some second adjacent entities whose domain types are different from the target domain type from the set of second adjacent entities associated with the historical search results, and combines at least one historical search keyword with some of the second adjacent entities to obtain a fourth set of extended data.

[0148] It should be noted that in the embodiment of the present application, the category system can also be called a domain type.

[0149] Take the training data x as an example, see Figure 7As shown, the vertical search device combines A with {B1, B2, B3, ..., BN B} in the second adjacent entity {B1, B2, ..., BM B} to obtain the first set of extended data:<A,B1> ,<A,B2> ,……, <A,BM B >, where B1, B2, ..., BM B Different from the target domain type of B, M B is a positive integer, M B The value is less than or equal to N B .

[0150] Taking the training data x as <"cleaning", "cleaning"> as an example, assuming that the second adjacent entity set associated with "cleaning" is {"cleaning", "clearing", "maintenance"}, where the target domain types of "cleaning" and "clearing" are both: housekeeping category, and the target domain type of "maintenance" is: auto repair category. The vertical search device determines that the target domain type of "home appliance cleaning" is: housekeeping category, and from the second adjacent entity set associated with "home appliance cleaning", some second adjacent entities of non-housekeeping categories are screened out: {"cleaning", "clearing"}, and "cleaning" is combined with {"cleaning", "clearing"} respectively to obtain the fourth group of extended data: <"cleaning", "cleaning">, <"cleaning", "clearing">.

[0151] In order to better learn the differences between different category systems, in the embodiment of the present application, as an example, B1, B2, ..., BM B The target domain type is different from that of B, and B1, B2, ..., BM B The text similarity with A exceeds the preset similarity threshold. As another example, B1, B2, ..., BM B The target domain type is different from that of B, and B1, B2, ..., BM B The text similarity with B exceeds a preset similarity threshold. The text similarity can be calculated using, but not limited to, Euclidean distance, Manhattan distance, and cosine similarity.

[0152] Method 4-2: The vertical search device determines the target upper entity associated with the historical search results contained in the training data x, and filters out some second adjacent entities whose associated upper entities are the same as the target upper entity from the set of second adjacent entities associated with the historical search results, and combines the historical search keywords with some second adjacent entities to obtain a fourth set of extended data.

[0153] Take the training data x as an example, see Figure 7 As shown, the vertical search device combines A with {B1, B2, B3, ..., BN B} in the second adjacent entity {B1, B2, ..., BM B} to obtain the first set of extended data:<A,B1> ,<A,B2> ,……, <A,BM B >, where B1, B2, ..., BM B The superordinate entities associated with B are all the same, M B is a positive integer, M B The value is less than or equal to N B .

[0154] Taking the training data x as <"cleaning", "air conditioning"> as an example, assuming that the second adjacent entity set associated with "air conditioning" is {"refrigeration equipment"}, where the upper entity of "air conditioning" is "home appliances" and the upper entity of "refrigeration equipment" is "home appliances", the vertical search device determines that "commercial appliances" and "air conditioning" have the same upper entity, and from the second adjacent entity set associated with "air conditioning", screen out some second adjacent entities {"refrigeration equipment"}, and combine "cleaning" with {"refrigeration equipment"} to obtain the fourth set of extended data: <"cleaning", "refrigeration equipment">.

[0155] Operation 5: The vertical search device combines the historical search results contained in the training data x with some of the first adjacent entities in the first adjacent entity set associated with the corresponding historical search keyword to obtain a fifth set of extended data. It should be noted that since operation 5 is similar to operation 4, it will not be described in detail here.

[0156] S604: The vertical search device obtains an extended data set based on at least one set of extended data.

[0157] Specifically, the vertical search device can extract the fifth set of training data from the first set of extended data and the second set of extended data according to the set extraction ratio; and obtain the extended data set based on the third set of extended data, the fourth set of extended data and the fifth set of extended data.

[0158] For example, assuming that the extraction ratio is 1:1, the data from the first group is expanded<A,B1> ,<A,B2> ,……, <A,BN B > and the second set of extended data:<A1,B> ,<A2,B> ,……, <AN A , B>, extract the fifth set of training data:<A,B1> ,<A2,B> ,……, <A,BN B > Then, the vertical search means based on the third set of extended data:<B,B1> ,<B,B2> ,……, <B,BN B >, the fourth set of extended data:<A,B1> ,<A,B2> ,……, <A,BM B>And the fifth set of extended data:<A1,B> ,<A2,B> ,……, <AM A , B>, M A is a positive integer, M A The value is less than or equal to N B , and get the extended data set.

[0159] Taking into account that there is duplicate extended data in at least one set of extended data obtained through the above operations one to five, in order to improve the efficiency of model training, in an embodiment of the present application, the vertical search device can deduplicate the extended data set, and subsequently train the semantic matching model based on the deduplicated extended data set.

[0160] Based on the same inventive concept, the present application embodiment provides a vertical search device. Figure 8 As shown, it is a schematic diagram of the structure of a vertical search device 800. The vertical search device 800 may include a receiving unit 801, a searching unit 802 and a sending unit 803, wherein:

[0161] A receiving unit 801 is configured to receive a vertical search request, wherein the vertical search request includes at least one target search keyword;

[0162] A search unit 802 is used to input the at least one target search keyword into a trained semantic matching model to obtain a search result set of a set search type; wherein the semantic matching model is trained based on an extended data set, and the extended data set is obtained by expanding the initial training data based on the entity association relationship contained in a preset service knowledge graph;

[0163] The sending unit 803 is used to return a search result page including the search result set.

[0164] Optionally, the vertical search device 800 further includes an expansion unit 804, wherein the expansion unit 804 is configured to:

[0165] Acquire an initial training data set, wherein each piece of training data in the initial training data set includes at least one historical search keyword and a corresponding historical search result;

[0166] Based on the entity association relationship contained in the preset service knowledge graph, determine from the service knowledge graph the first adjacent entity set associated with each historical search keyword and the second adjacent entity set associated with each historical search result contained in the initial training data set;

[0167] Based on each of the obtained first adjacent entity sets and each of the obtained second adjacent entity sets, at least one set of extended data is formed; wherein the at least one set of extended data is different from each of the training data;

[0168] The extended data set is obtained based on the at least one set of extended data.

[0169] Optionally, when determining, from the service knowledge graph, first adjacent entity sets associated with each historical search keyword contained in the initial training data set based on the entity association relationship contained in the preset service knowledge graph, the expansion unit 804 is specifically used to:

[0170] For each historical search keyword included in the initial training data set, the following operations are performed respectively:

[0171] Based on the entity association relationship contained in the preset service knowledge graph, a candidate adjacent entity set associated with a historical search keyword is obtained, wherein the candidate adjacent entity set associated with the historical search keyword includes each adjacent entity of the historical search keyword in the service knowledge graph;

[0172] Deleting a target adjacent entity from each candidate adjacent entity included in the obtained candidate adjacent entity set, wherein the target adjacent entity is: an entity having a subordinate entity in the candidate adjacent entity set based on the service knowledge graph;

[0173] The candidate adjacent entity set from which the target adjacent entity is deleted is used as the first adjacent entity set associated with the one historical search keyword.

[0174] Optionally, when at least one set of extended data is obtained based on each of the obtained first adjacent entity sets and each of the obtained second adjacent entity sets, the extending unit 804 is configured to perform at least one of the following operations:

[0175] For each piece of training data included in the initial training data set, the following operations are performed respectively:

[0176] Combine at least one historical search keyword contained in a piece of training data with each second adjacent entity in a second adjacent entity set associated with the corresponding historical search result to obtain a first set of extended data;

[0177] Combine each first adjacent entity in the first adjacent entity set associated with each of the historical search results contained in a piece of training data and the corresponding at least one historical search keyword to obtain a second set of extended data;

[0178] Combine a historical search result included in a piece of training data with each second adjacent entity in a second adjacent entity set associated with the historical search result to obtain a third set of extended data;

[0179] Combining at least one historical search keyword contained in a piece of training data with some of the second adjacent entities in the second adjacent entity set associated with the corresponding historical search result to obtain a fourth set of extended data;

[0180] A historical search result included in a piece of training data is combined with some of the first adjacent entities in the first adjacent entity set associated with the corresponding at least one historical search keyword to obtain a fifth set of extended data.

[0181] Optionally, when combining at least one historical search keyword contained in a piece of training data with some of the second adjacent entities in the second adjacent entity set associated with the corresponding historical search result to obtain a fourth set of extended data, the extension unit 804 is used to perform at least one of the following operations:

[0182] Determine the target domain type to which the historical search results contained in the piece of training data belong, and filter out some second adjacent entities whose domain types are different from the target domain type from the set of second adjacent entities associated with the historical search results, and combine the at least one historical search keyword with the some second adjacent entities, respectively, to obtain the fourth set of extended data;

[0183] Determine the target upper entity associated with the historical search results contained in the training data, and filter out some second adjacent entities whose associated upper entities are the same as the target upper entity from the set of second adjacent entities associated with the historical search results, and combine the at least one historical search keyword with the some second adjacent entities respectively to obtain the fourth set of extended data.

[0184] Optionally, after combining a historical search result included in a piece of training data with each second adjacent entity in a second adjacent entity set associated with the historical search result to obtain a third set of extended data, and before obtaining the extended data set based on the at least one set of extended data, the extending unit 804 is further configured to:

[0185] If the entity association relationship between the historical search result and a second adjacent entity indicates that the historical search result and the second adjacent entity are in a hierarchical relationship, then the label of the extended data obtained by combining the historical search result and the second adjacent entity is set to be first-level related;

[0186] If the entity association relationship between the historical search result and a second adjacent entity indicates that the historical search result and the second adjacent entity are in a synonym relationship, then the label of the extended data obtained by combining the historical search result and the second adjacent entity is set as a secondary correlation;

[0187] Each tag is used to represent the real relevance between the corresponding historical search keyword and the corresponding historical search result, and the relevance of the first-level relevance representation is lower than the relevance of the second-level relevance representation.

[0188] Optionally, when obtaining the extended data set based on the at least one set of extended data, the extending unit 804 is specifically configured to:

[0189] Extracting a fifth set of training data from the first set of extended data and the second set of extended data according to the set extraction ratio;

[0190] The extended data set is obtained based on the third set of extended data, the fourth set of extended data and the fifth set of extended data.

[0191] Optionally, when inputting the at least one target search keyword into the trained semantic matching model to obtain a search result set of a set search type, the search unit 802 is further configured to:

[0192] Obtaining the relevance of each search result contained in the search result set and the at least one target search keyword, each relevance being used to characterize the relevance between the corresponding search result and the at least one target search keyword;

[0193] When returning the search result page including the search result set, the sending unit is specifically used for:

[0194] Based on the obtained relevances, the search results are ranked, and a search result page containing the ranked search results is returned.

[0195] For the convenience of description, the above parts are divided into modules (or units) according to their functions and described separately. Of course, when implementing this application, the functions of each module (or unit) can be implemented in the same or multiple software or hardware.

[0196] Regarding the device in the above embodiment, the specific manner in which each unit executes the request has been described in detail in the embodiment of the method, and will not be elaborated here.

[0197] Those skilled in the art will appreciate that various aspects of the present application may be implemented as a system, method or program product. Therefore, various aspects of the present application may be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software, which may be collectively referred to as "circuit", "module" or "system" herein.

[0198] After introducing the method and device for generating an audio library according to an exemplary embodiment of the present application, next, an electronic device according to another exemplary embodiment of the present application is introduced.

[0199] Fig. 9 is a block diagram of an electronic device 900 according to an exemplary embodiment, the device comprising:

[0200] Processor 910;

[0201] A memory 920 for storing instructions executable by the processor 910;

[0202] The processor 910 is configured to execute instructions to implement the vertical search method in the embodiment of the present disclosure, for example Figure 2 or Figure 6 Follow the steps shown in .

[0203] In an exemplary embodiment, a storage medium including operations is also provided, such as a memory 920 including operations, and the operations can be executed by the processor 910 of the electronic device 900 to complete the above method. Optionally, the storage medium can be a non-temporary computer-readable storage medium, for example, the non-temporary computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a portable compact disk read only memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.

[0204] Based on the same inventive concept, see Fig.10 As shown, the embodiment of the present application further provides a terminal device 1000, which can be an electronic device such as a smart phone, a tablet computer, a laptop or a PC.

[0205] The terminal device 1000 includes a display unit 1040, a processor 1080, and a memory 1020, wherein the display unit 1040 includes a display panel 1041, which is used to display information input by a user or information provided to a user and various operation interfaces of the terminal device 1000, and in the embodiment of the present application, is mainly used to display the operation interface, shortcut window, etc. of the application program installed in the terminal device 1000. Optionally, the display panel 1041 can be configured in the form of LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0206] The processor 1080 is used to read the computer program and then execute the method defined by the computer program. For example, the processor 1080 reads the application, thereby running the application on the terminal device 1000 and displaying the operation interface on the display unit 1040. The processor 1080 may include one or more general-purpose processors and may also include one or more DSPs (Digital Signal Processors) to perform related operations to implement the technical solutions provided in the embodiments of the present application.

[0207] The memory 1020 generally includes internal memory and external memory, and the internal memory may be RAM, ROM, and cache (CACHE), etc. The external memory may be a hard disk, an optical disk, a USB disk, a floppy disk, or a tape drive, etc. The memory 1020 is used to store computer programs and other data, and the computer programs include applications, etc. The other data may include data generated after the operating system or the application is run, and the data includes system data (such as configuration parameters of the operating system) and user data. In the embodiment of the present application, program instructions are stored in the memory 1020, and the processor 1080 executes the program instructions in the memory 1020 to implement the vertical search method discussed above.

[0208] In addition, the terminal device 1000 may also include a display unit 1040 for receiving input digital information, character information or contact touch operation / contactless gesture, and generating signal input related to user settings and function control of the terminal device 1000. Specifically, in the embodiment of the present application, the display unit 1040 may include a display panel 1041. The display panel 1041, for example, a touch screen, can collect the user's touch operation on or near it (such as the user's operation on the display panel 1041 or on the display panel 1041 using any suitable object or accessory such as a finger, stylus, etc.), and drive the corresponding connection device according to a pre-set program. Optionally, the display panel 1041 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact point coordinates, and then sends it to the processor 1080, and can receive and execute commands sent by the processor 1080. In an embodiment of the present application, if a user selects a control in the operation interface, the touch detection device in the display panel 1041 detects the touch operation, and sends a signal corresponding to the detected touch operation to the touch controller. The touch controller converts the signal into touch point coordinates and sends them to the processor 1080. The processor 1080 determines the control selected by the user based on the received touch point coordinates.

[0209] The display panel 1041 may be implemented in various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the display unit 1040, the terminal device 1000 may further include an input unit 1030, which may include but is not limited to one or more of a physical keyboard, a function key (such as a volume control key, a switch key, etc.), a trackball, a mouse, and a joystick. Fig.10 In the figure, the input unit 1030 includes an image input device 1031 and other input devices 1032 as an example.

[0210] In addition to the above, the terminal device 1000 may also include a power supply 1090 for supplying power to other modules, an audio circuit 1060, a near field communication module 1070, and a radio frequency (RF) circuit 1010. The terminal device 1000 may also include one or more sensors 1050, such as an accelerometer, a light sensor, a pressure sensor, etc. The audio circuit 1060 specifically includes a speaker 1061 and a microphone 1062, etc. For example, the user can use voice control, the terminal device 1000 can collect the user's voice through the microphone 1062, can be controlled by the user's voice, and when the user needs to be prompted, the corresponding prompt tone is played through the speaker 1061.

[0211] Based on the same inventive concept, the present application also provides a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the vertical search method provided in various optional implementations in the above embodiments.

[0212] In some possible implementations, various aspects of the vertical search method provided by the present application may also be implemented in the form of a program product, which includes a computer program. When the program product is run on a computer device, the computer program is used to enable the computer device to execute the steps of the vertical search method according to various exemplary implementations of the present application described above in this specification. For example, the computer device may execute the following steps: Figure 2 or Figure 6 Follow the steps shown in .

[0213] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a RAM, a ROM, an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a CD-ROM, an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0214] The program product of the embodiment of the present application may adopt a CD-ROM and include program code, and may be run on a computing device. However, the program product of the present application is not limited thereto, and in this document, a readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with a command execution system, device or device.

[0215] The readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in combination with a command execution system, device, or device. Although the preferred embodiments of the present application have been described, once the basic creative concept is known to those skilled in the art, additional changes and modifications may be made to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0216] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A vertical search method, characterized in that: include: receiving a vertical search request, wherein the vertical search request includes at least one target search keyword; Input the at least one target search keyword into the trained semantic matching model to obtain a search result set of the set search type; wherein the semantic matching model is obtained by training based on an extended data set, and the extended data set is obtained in the following manner: obtain an initial training data set, in which each training data contains at least one historical search keyword and a corresponding historical search result; based on the entity association relationship contained in the preset service knowledge graph, determine from the service knowledge graph the first adjacent entity set associated with each historical search keyword and the second adjacent entity set associated with each historical search result contained in the initial training data set; based on the obtained first adjacent entity sets and second adjacent entity sets, form at least one set of extended data; the at least one set of extended data is different from each training data; based on the at least one set of extended data, obtain the extended data set; Returning a search result page containing the search result set; Among them, the first adjacent entity set associated with each historical search keyword is obtained in the following manner: for each historical search keyword contained in the initial training data set, the following operations are performed respectively: based on the entity association relationship contained in the preset service knowledge graph, a candidate adjacent entity set associated with a historical search keyword is obtained, wherein the candidate adjacent entity set associated with the historical search keyword includes each adjacent entity of the historical search keyword in the service knowledge graph; from each candidate adjacent entity included in the obtained candidate adjacent entity set, the target adjacent entity is deleted, and the target adjacent entity is: based on the service knowledge graph, an entity with a lower entity in the candidate adjacent entity set; the candidate adjacent entity set with the target adjacent entity deleted is used as the first adjacent entity set associated with the historical search keyword.

2. The method according to claim 1, characterized in that The step of composing at least one set of extended data based on each of the first adjacent entity sets and each of the second adjacent entity sets obtained includes at least one of the following operations: For each piece of training data included in the initial training data set, the following operations are performed respectively: Combine at least one historical search keyword contained in a piece of training data with each second adjacent entity in a second adjacent entity set associated with the corresponding historical search result to obtain a first set of extended data; Combine each first adjacent entity in the first adjacent entity set associated with each of the historical search results contained in a piece of training data and the corresponding at least one historical search keyword to obtain a second set of extended data; Combine a historical search result included in a piece of training data with each second adjacent entity in a second adjacent entity set associated with the historical search result to obtain a third set of extended data; Combining at least one historical search keyword contained in a piece of training data with some of the second adjacent entities in the second adjacent entity set associated with the corresponding historical search result to obtain a fourth set of extended data; A historical search result included in a piece of training data is combined with some of the first adjacent entities in the first adjacent entity set associated with the corresponding at least one historical search keyword to obtain a fifth set of extended data.

3. The method according to claim 2, characterized in that The step of combining at least one historical search keyword included in a piece of training data with some of the second adjacent entities in the second adjacent entity set associated with the corresponding historical search result to obtain a fourth set of extended data includes at least one of the following operations: Determine the target domain type to which the historical search results contained in the piece of training data belong, and filter out some second adjacent entities whose domain types are different from the target domain type from the set of second adjacent entities associated with the historical search results, and combine the at least one historical search keyword with the some second adjacent entities, respectively, to obtain the fourth set of extended data; Determine the target upper entity associated with the historical search results contained in the training data, and filter out some second adjacent entities whose associated upper entities are the same as the target upper entity from the set of second adjacent entities associated with the historical search results, and combine the at least one historical search keyword with the some second adjacent entities respectively to obtain the fourth set of extended data.

4. The method according to claim 2, characterized in that After combining the historical search results included in a piece of training data with each second adjacent entity in the second adjacent entity set associated with the historical search results to obtain a third set of extended data, before obtaining the extended data set based on the at least one set of extended data, the method further includes: If the entity association relationship between the historical search result and a second adjacent entity indicates that the historical search result and the second adjacent entity are in a hierarchical relationship, then the label of the extended data obtained by combining the historical search result and the second adjacent entity is set to be first-level related; If the entity association relationship between the historical search result and a second adjacent entity indicates that the historical search result and the second adjacent entity are in a synonym relationship, then the label of the extended data obtained by combining the historical search result and the second adjacent entity is set as a secondary correlation; Each tag is used to represent the real relevance between the corresponding historical search keyword and the corresponding historical search result, and the relevance of the first-level relevance representation is lower than the relevance of the second-level relevance representation.

5. The method according to claim 2, characterized in that The obtaining the extended data set based on the at least one set of extended data comprises: Extracting a fifth set of training data from the first set of extended data and the second set of extended data according to the set extraction ratio; The extended data set is obtained based on the third set of extended data, the fourth set of extended data and the fifth set of extended data.

6. The method according to claim 1, characterized in that The step of inputting the at least one target search keyword into the trained semantic matching model to obtain a search result set of a set search type further includes: Obtaining the relevance of each search result contained in the search result set and the at least one target search keyword, each relevance being used to characterize the relevance between the corresponding search result and the at least one target search keyword; Then the returning of the search result page containing the search result set includes: Based on the obtained relevances, the search results are ranked, and a search result page containing the ranked search results is returned.

7. A vertical search device, characterized in that: include: A receiving unit, configured to receive a vertical search request, wherein the vertical search request includes at least one target search keyword; A search unit, configured to input the at least one target search keyword into a trained semantic matching model to obtain a search result set of a set search type; wherein the semantic matching model is trained based on an extended data set; The extension unit is used to obtain the extended data set in the following manner: obtain an initial training data set, in which each piece of training data contains at least one historical search keyword and a corresponding historical search result; based on the entity association relationship contained in a preset service knowledge graph, determine from the service knowledge graph the first adjacent entity set associated with each historical search keyword and the second adjacent entity set associated with each historical search result contained in the initial training data set; compose at least one set of extended data based on each first adjacent entity set and each second adjacent entity set obtained; wherein the at least one set of extended data is different from each piece of training data; and obtain the extended data set based on the at least one set of extended data; A sending unit, configured to return a search result page including the search result set; The expansion unit is specifically used to obtain the first adjacent entity set associated with each historical search keyword in the following manner: for each historical search keyword included in the initial training data set, the following operations are performed respectively: based on the entity association relationship included in the preset service knowledge graph, a candidate adjacent entity set associated with a historical search keyword is obtained, wherein the candidate adjacent entity set associated with the historical search keyword includes each adjacent entity of the historical search keyword in the service knowledge graph; from each candidate adjacent entity included in the obtained candidate adjacent entity set, a target adjacent entity is deleted, wherein the target adjacent entity is an entity that has a lower entity in the candidate adjacent entity set based on the service knowledge graph; the candidate adjacent entity set with the target adjacent entity deleted is used as the first adjacent entity set associated with the historical search keyword.

8. The device according to claim 7, characterized in that When at least one set of extended data is obtained based on each of the first adjacent entity sets and each of the second adjacent entity sets, the extension unit is used to perform at least one of the following operations: For each piece of training data included in the initial training data set, the following operations are performed respectively: Combine at least one historical search keyword contained in a piece of training data with each second adjacent entity in a second adjacent entity set associated with the corresponding historical search result to obtain a first set of extended data; Combine each first adjacent entity in the first adjacent entity set associated with each of the historical search results contained in a piece of training data and the corresponding at least one historical search keyword to obtain a second set of extended data; Combine a historical search result included in a piece of training data with each second adjacent entity in a second adjacent entity set associated with the historical search result to obtain a third set of extended data; Combining at least one historical search keyword contained in a piece of training data with some of the second adjacent entities in the second adjacent entity set associated with the corresponding historical search result to obtain a fourth set of extended data; A historical search result included in a piece of training data is combined with some of the first adjacent entities in the first adjacent entity set associated with the corresponding at least one historical search keyword to obtain a fifth set of extended data.

9. The device according to claim 8, characterized in that When combining at least one historical search keyword included in a piece of training data with some of the second adjacent entities in the second adjacent entity set associated with the corresponding historical search result to obtain a fourth set of extended data, the extension unit is used to perform at least one of the following operations: Determine the target domain type to which the historical search results contained in the piece of training data belong, and filter out some second adjacent entities whose domain types are different from the target domain type from the set of second adjacent entities associated with the historical search results, and combine the at least one historical search keyword with the some second adjacent entities, respectively, to obtain the fourth set of extended data; Determine the target upper entity associated with the historical search results contained in the training data, and filter out some second adjacent entities whose associated upper entities are the same as the target upper entity from the set of second adjacent entities associated with the historical search results, and combine the at least one historical search keyword with the some second adjacent entities respectively to obtain the fourth set of extended data.

10. An electronic device, characterized in that: It comprises a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of any one of the methods of claims 1 to 6.

11. A computer-readable storage medium, characterized in that: It includes a computer program. When the computer program is run on an electronic device, the computer program is used to enable the electronic device to execute the steps of any method described in claims 1 to 6.

Citation Information

Patent Citations

  • Method and device for determining search intention of user

    CN111680207A

  • Search method and device based on brand protection

    CN111782942A