Training method, electronic device and storage medium for search system
By adopting a training method based on pre-trained language model in the search system, and using a deep neural network model to convert different types of data into unified semantic vectors, the problems of diversity of search results and poor user experience in the existing technology are solved, and efficient multimodal data retrieval and improved user experience are achieved.
Patent Information
- Application Number
- CN202111308794.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-05
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-11-05
AI Technical Summary
Existing search systems are difficult to effectively process different types of data (such as text, pictures, and videos), resulting in poor search results and poor user experience.
The search system training method based on pre-trained language model is adopted, and the end-to-end deep neural network basic model composed of cascaded recall model and sorting model is used to convert different types of data into unified semantic vectors to realize unified retrieval of multimodal data.
It improves search performance, improves the diversity and user experience of search content, and can effectively handle multimodal, multilingual, and multi-resource search needs.
Smart Images

Figure CN114036322B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to the field of intelligent search technology, and specifically to a training method for a search system, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] Artificial intelligence is a discipline that studies how to use computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.). It includes both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include computer vision technology, speech recognition technology, natural language processing technology, as well as machine learning / deep learning, big data processing technology, knowledge graph technology, and other major directions.
[0003] Data search is one of the basic services on the Internet, which can provide search results that meet user needs based on user search requests.
[0004] The methods described in this section are not necessarily methods that have been previously conceived or employed. Unless otherwise indicated, it should not be assumed that any method described in this section is considered to be prior art simply because it is included in this section. Similarly, unless otherwise indicated, the issues mentioned in this section should not be considered to have been recognized in any prior art. Summary of the invention
[0005] The present disclosure provides a training method for a search system based on a pre-trained language model, an electronic device, a computer-readable storage medium, and a computer program product.
[0006] According to one aspect of the present disclosure, a training method for a search system based on a pre-trained language model is provided, wherein the search system includes an end-to-end deep neural network base model composed of a cascade of a recall model and a ranking model, and wherein the recall model is constructed based on a dual encoder, and the ranking model is constructed based on a cross encoder, and the method includes: receiving a sample data set, wherein the sample data in the sample data set includes a sample search request and a first target output data set; initializing multiple parameters in the recall model and the ranking model; for each sample data, performing the following operations: converting the sample search request in the sample data into a first request semantic vector by the first encoder in the recall model; converting multiple candidate data of different types into corresponding multiple first data semantic vectors by the second encoder in the recall model, wherein the different The multiple candidate data of the type include at least text, picture and video; the first similarities between the first request semantic vector and the multiple first data semantic vectors are calculated respectively, so as to select a first number of first data semantic vectors, wherein the first similarities between the first number of first data semantic vectors and the first request semantic vector both meet a preset condition; the sample search request and the candidate data corresponding to each of the first data semantic vectors in the first number of first data semantic vectors are input into the cross encoder of the sorting model as the first joint input value in sequence, so as to sort the candidate data corresponding to the first number of first data semantic vectors respectively; the loss function is calculated based on the sorted candidate data and the first target output data set; and the multiple parameters in the recall model and the sorting model are adjusted based on the loss function.
[0007] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned training method of the search system based on the pre-trained language model.
[0008] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the above-mentioned training method for the search system based on the pre-trained language model.
[0009] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein the computer program implements the above-mentioned training method of the search system based on the pre-trained language model when executed by a processor.
[0010] According to one or more embodiments of the present disclosure, the training of a search system based on a pre-trained language model may be implemented.
[0011] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings exemplarily illustrate the embodiments and constitute a part of the specification, and together with the text description of the specification, are used to explain the exemplary implementation of the embodiments. The embodiments shown are for illustrative purposes only and do not limit the scope of the claims. In all drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0013] Figure 1 A schematic diagram of a model structure based on a cross encoder structure and a schematic diagram of a model structure based on a dual encoder structure according to an exemplary embodiment of the present disclosure are shown;
[0014] Figure 2 A flowchart of a training method for a search system based on a pre-trained language model according to an exemplary embodiment of the present disclosure is shown;
[0015] Figure 3 A flowchart of a training method for a search system based on a pre-trained language model according to an exemplary embodiment of the present disclosure is shown;
[0016] Figure 4 A structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0017] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted in the following description.
[0018] In the present disclosure, unless otherwise specified, the use of the terms "first", "second", etc. to describe various elements is not intended to limit the positional relationship, timing relationship, or importance relationship of these elements, and such terms are only used to distinguish one element from another element. In some examples, the first element and the second element may refer to the same instance of the element, and in some cases, based on the description of the context, they may also refer to different instances.
[0019] The terms used in the description of various examples in this disclosure are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element can be one or more. In addition, the term "and / or" used in this disclosure covers any one of the listed items and all possible combinations.
[0020] Data search is one of the basic services on the Internet, which can provide search results that meet user needs based on user search requests.
[0021] The inventor creatively proposed a training method for a search system, which concentrates network resources in various forms including text, pictures, videos, tables, etc. in a unified resource database in the form of unified vector expressions to convert different types of data into corresponding data semantic vectors with unified specifications. Therefore, by comparing the semantic vector corresponding to the user's search request with the semantic vector corresponding to at least one data, the search results for the user's search request can be recalled from at least one data. The search system trained in this way can directly match the semantic vectors corresponding to different types of data for similarity by mapping different types of data to the same semantic vector space, and obtain multimodal data that matches the user's search request, which is conducive to improving search performance, specifically improving the diversity of retrieved content, and improving user experience.
[0022] The attributes of the data include at least one of the following: modality, language, and data structure. The modality of the data includes text, pictures, and videos, the language includes various types of languages, such as Chinese and English, and the data structure includes structured data (such as tables, graphs) and unstructured data. Therefore, the solution of the embodiment of the present application can realize a unified search method for multi-modality, multi-language, and multi-resources.
[0023] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0024] According to one aspect of the present disclosure, a training method for a search system based on a pre-trained language model is provided, wherein the search system includes an end-to-end deep neural network base model composed of a cascade of a recall model and a ranking model, and wherein the recall model is based on a dual encoder construction and the ranking model is based on a cross encoder construction.
[0025] Figure 1The schematic diagram of the model structure based on the cross encoder construction and the schematic diagram of the model structure based on the dual encoder construction are shown. The question sentence (i.e., the search request input by the user) and the passage sentence (i.e., the data in the search resource database) are input into the recall model based on the dual encoder construction. The two sentences are passed into two independent encoding networks, and the two encoding networks can be the same network. The two encoding networks respectively output the semantic vectors E corresponding to the two sentences q (q) and E p (p), so that similarity can be calculated based on the semantic vector. Based on the similarity, one or more passage sentences matching the question sentence are obtained.
[0026] The question sentence and one or more passage sentences recalled by the recall model are sequentially input into the recommendation model constructed based on the cross encoder, and the similarity score between the question sentence and the one or more passage sentences is output. Based on the similarity score, the one or more passage sentences are sorted. Among them, the input recommendation model is the semantic vector corresponding to the question sentence and the one or more passage sentences recalled.
[0027] Figure 2A training method for a search system based on a pre-trained language model according to an exemplary embodiment of the present disclosure is shown, comprising: step S201, receiving a sample data set, wherein the sample data in the sample data set includes a sample search request and a first target output data set; step S202, initializing multiple parameters in the recall model and the sorting model; for each sample data, performing the following operations: step S203, converting the sample search request in the sample data into a first request semantic vector by a first encoder in the recall model; step S204, converting multiple candidate data of different types into corresponding multiple first data semantic vectors by a second encoder in the recall model, wherein the multiple candidate data of different types include at least text, picture and video, and wherein the multiple first data semantic vectors have a uniform specification; step S205 5. Calculate the first similarity between the first request semantic vector and the plurality of first data semantic vectors respectively to obtain a first number of first data semantic vectors, wherein the first similarity between the first number of first data semantic vectors and the first request semantic vector satisfies a preset condition; step S206, sequentially input the sample search request and the candidate data corresponding to each of the first data semantic vectors in the first number of first data semantic vectors as the first joint input value into the cross encoder of the sorting model to sort the candidate data corresponding to the first number of first data semantic vectors respectively; step S207, calculate the loss function based on the sorted candidate data and the first target output data set; and step S208, adjust multiple parameters in the recall model and the sorting model based on the loss function. The similarity between vectors can be, for example, but not limited to, cosine similarity.
[0028] According to the training method of the embodiment of the present disclosure, the training of the search system based on the pre-trained language model can be implemented.
[0029] According to some embodiments, the plurality of first data semantic vectors have a uniform specification. Thus, a uniform search of different types of data can be achieved. For example, different types of data can be uniformly converted into a 1000-dimensional data semantic vector, and the search request should be converted into a 1000-dimensional vector of the same specification in the same mapping manner.
[0030] According to some embodiments, in addition to text, pictures and videos, the different types of multiple candidate data also include at least tables and knowledge graphs. It is understandable that the different types of data may further include other types of data, such as maps, animations, etc. More types of data can further enrich the search resource database, thereby further improving the diversity of search results, better meeting user needs, and improving user experience.
[0031] According to some embodiments, at least one text or video data among the multiple candidate data of different types is obtained by fine-grained division of the original complete data, thereby enabling a deeper understanding of the data content, thereby achieving fine-grained indexing, and obtaining search results that better meet user needs.
[0032] Exemplarily, the obtaining of at least one text or video data by fine-grained division of the original complete data may be fine-grained division of the original complete data according to semantics. In some embodiments, the fine-grained division of the original complete data may include semantic segmentation of the original complete data to obtain at least one text or video data. Taking web page text data as an example, the original complete web page text data may include multiple paragraphs, each paragraph may have different semantic features, then the data semantic vector corresponding to the corresponding complete web page text data cannot fully express the different semantic features of each paragraph, and the first request semantic vector reflecting the user's needs cannot be matched with the semantics of each paragraph during the search process. By dividing the original complete web page text data, it can be divided into multiple fragments with different semantic features, each fragment corresponding to one of the at least one text. Each fragment is converted into a corresponding data semantic vector, which can be matched with the first request semantic vector reflecting the user's needs during the search process, and a search result that better meets the user's needs is obtained. Similarly, fine-grained division can be performed based on the video text data corresponding to the video, and the specific principle and process are similar to the web page text data.
[0033] According to some embodiments, each of the plurality of first data semantic vectors includes a dimension related to the content quality of the corresponding candidate data. The dimension related to the content quality of the corresponding data may be a content quality score of the corresponding data, but is not limited thereto. Thus, the search system trained using the method can further consider the content quality of the data and improve the quality of the search results.
[0034] According to some embodiments, each of the plurality of first data semantic vectors includes a dimension related to the release time of the corresponding candidate data. The dimension related to the release time of the corresponding data may be the release time of the corresponding data, but is not limited thereto. Thus, the search system trained using the method can further consider the timeliness of the data and improve the quality of the search results.
[0035] According to some embodiments, each of the plurality of first data semantic vectors includes a dimension related to the source credibility of the corresponding candidate data. The dimension related to the source credibility of the corresponding data may be the source website type of the corresponding data and the credibility of the corresponding website type, but is not limited thereto. Thus, the search system trained using the method can further consider the authority of the data and improve the quality of the search results.
[0036] According to some embodiments, each data semantic vector in the semantic space vector includes at least two of the following dimensions of the corresponding data: a dimension related to content quality, a dimension related to release time, and a dimension related to source credibility. In the search process, by adding multiple dimensions of data, the quality of search results can be further improved.
[0037] Based on the same principle, according to some embodiments, the first request semantic vector includes context information related to the user's search, and the context information includes at least one of time, location, and the user's previous search. Thus, the search accuracy of the search system trained using the method can be further improved.
[0038] It is understandable that the user's direct search needs can be described more accurately based on the context information related to the user's search. For example, when the user inputs the search information "how is tomorrow's weather", the first request semantic vector can include the user's location, such as Beijing, so as to provide the user with search results related to "tomorrow's weather in Beijing", more accurately meet the user's needs and improve the user experience.
[0039] According to some embodiments, the first joint input value includes at least one of the content quality, release time, and source credibility of the corresponding candidate data. Thus, the content quality, release time, source credibility, etc. of the candidate data can be fully considered in the sorting process, thereby obtaining a higher quality sorting result, and further enabling the search system trained using the method to generate search results that better meet user needs, thereby improving user experience.
[0040] According to some embodiments, the system further comprises a recommendation model, and wherein the sample data in the sample data set further comprises a second target output data set, such as Figure 3As shown, the training method also includes: step S301, initializing multiple parameters in the recommendation model; for each sample data, performing the following operations: step S302, sequentially inputting the sample search request and the candidate data corresponding to each of the first number of first data semantic vectors as the second joint input value into the cross encoder of the recommendation model to sort the candidate data corresponding to the first number of first data semantic vectors respectively; step S303, calculating the loss function based on the sorted candidate data and the second target output data set; and step S304, adjusting multiple parameters in the target model and the recommendation model based on the loss function.
[0041] According to some embodiments, the second joint input value includes semantic relevance features and perceptual relevance features of the corresponding candidate data. The semantic relevance features are used to describe the direct semantics of the candidate data, and the perceptual relevance features focus on dimensions related to user needs and interests. Thus, the accuracy of the recommendation model trained using the method can be further improved, and the potential needs of users can be better met.
[0042] For example, for a webpage containing content introducing public figure A, the semantic relevance feature dimension in the data semantic vector corresponding to the webpage is used to describe the direct semantics of the webpage content. The perceptual relevance feature dimension in the data semantic vector corresponding to the webpage focuses on describing the user's possible interests around public figure A. For example, the user may be interested in who is the wife of public figure A, what works public figure A has, etc., and the perceptual relevance feature dimension can include the corresponding content. In this way, the accuracy of the search system trained using the method can be further improved, and the potential needs of users can be better met.
[0043] According to some embodiments, the data volume of the second target output data set is smaller than the data volume of the first target output data set. By setting up two target output data sets, the search and recommendation can be more targeted. The first target output data set is used to meet the large recall requirements of correlation search, and the second target output data set is used to meet the precise recall requirements of user potential needs, thereby better improving the quality of query results and further improving user experience.
[0044] Exemplarily, the amount of data in the first target output data set may be in the order of tens of billions or hundreds of billions, so as to cover more content resources and more comprehensively cover the content needs of users. Correspondingly, the amount of data in the second target output data set may be in the order of millions.
[0045] According to some embodiments, the data in the second target output data set used for recommendation is selected according to a predetermined quality standard, thereby being able to provide users with higher quality recommended content, better meet users' extended needs, and enhance user experience.
[0046] The search system trained by the above technical solution can respond to the user's search request and generate search results, and can determine the user's potential search intention and generate recommended results, so as to accurately meet the user's direct needs, while also expanding the field of vision and meeting the user's extended needs.
[0047] According to another aspect of the present disclosure, an electronic device is also provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned training method of the search system based on the pre-trained language model.
[0048] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is also provided, wherein the computer instructions are used to enable the computer to execute the above-mentioned training method for the search system based on the pre-trained language model.
[0049] According to another aspect of the present disclosure, a computer program product is also provided, including a computer program, wherein the computer program implements the above-mentioned training method of the search system based on the pre-trained language model when executed by a processor.
[0050] refer to Figure 4 , a block diagram of an electronic device 400 that can be used as a server or client of the present disclosure will now be described, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0051] like Figure 4As shown, the device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0052] Multiple components in the device 400 are connected to the I / O interface 405, including: an input unit 406, an output unit 407, a storage unit 408, and a communication unit 409. The input unit 406 can be any type of device that can input information to the device 400. The input unit 406 can receive input digital or character information, and generate key signal input related to user settings and / or function control of the electronic device, and can include but is not limited to a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone and / or a remote control. The output unit 407 can be any type of device that can present information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator and / or a printer. The storage unit 408 can include but is not limited to a disk, an optical disk. The communication unit 409 allows the device 400 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver and / or a chipset, such as Bluetooth TM devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0053] The computing unit 401 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 401 performs the various methods and processes described above, such as the above-described method for data search or the training method for a search system based on a pre-trained language model. For example, in some embodiments, the above-described method for data search or the training method for a search system based on a pre-trained language model may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the method for data search or the training method for a search system based on a pre-trained language model described above may be performed. Alternatively, in other embodiments, the computing unit 401 may be configured in any other appropriate manner (for example, by means of firmware) to execute the above-mentioned method for data search or the training method of the search system based on the pre-trained language model.
[0054] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0055] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0056] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0057] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0058] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0059] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0060] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0061] Although the embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above-mentioned methods, systems and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but only by the claims after authorization and their equivalent scope. Various elements in the embodiments or examples can be omitted or replaced by their equivalent elements. In addition, each step can be performed in an order different from that described in the present disclosure. Further, the various elements in the embodiments or examples can be combined in various ways. It is important that with the evolution of technology, many elements described herein can be replaced by equivalent elements that appear after the present disclosure.
Claims
1. A training method for a search system based on a pre-trained language model, wherein: The search system includes an end-to-end deep neural network base model composed of a recall model and a ranking model cascaded, wherein the recall model is constructed based on a dual encoder and the ranking model is constructed based on a cross encoder, and the method includes: receiving a sample data set, wherein the sample data in the sample data set includes a sample search request and a first target output data set; Initializing multiple parameters in the recall model and the ranking model; For each sample data, perform the following operations: The first encoder in the recall model converts the sample search request in the sample data into a first request semantic vector; The second encoder in the recall model converts the multiple candidate data of different types into corresponding multiple first data semantic vectors respectively, wherein the multiple candidate data of different types include at least text, picture and video; Respectively calculating first similarities between the first request semantic vector and the plurality of first data semantic vectors to obtain a first number of first data semantic vectors, wherein the first similarities between the first number of first data semantic vectors and the first request semantic vector both satisfy a preset condition; sequentially inputting the sample search request and the candidate data corresponding to each of the first data semantic vectors of the first number as first joint input values into the cross encoder of the sorting model to sort the candidate data corresponding to the first data semantic vectors respectively; Calculating a loss function based on the sorted candidate data and the first target output data set; and A plurality of parameters in the recall model and the ranking model are adjusted based on the loss function.
2. The method of claim 1, wherein: The first joint input value includes at least one of content quality, release time, and source credibility of corresponding candidate data.
3. The method according to claim 1 or 2, wherein: The system further includes a recommendation model, and wherein the sample data in the sample data set further includes a second target output data set, The method further comprises: Initializing multiple parameters in the recommendation model; For each sample data, perform the following operations: sequentially inputting the sample search request and the candidate data corresponding to each of the first data semantic vectors of the first number as second joint input values into the cross encoder of the recommendation model to sort the candidate data corresponding to the first data semantic vectors respectively; Calculating a loss function based on the sorted candidate data and the second target output data set; and A plurality of parameters in the target model and the recommendation model are adjusted based on the loss function.
4. The method of claim 3, wherein: The second joint input value includes semantic relevance features and perceptual relevance features of the corresponding candidate data.
5. The method according to claim 1 or 2, wherein: Each of the plurality of first data semantic vectors includes a dimension associated with content quality of corresponding candidate data.
6. The method according to claim 1 or 2, wherein: Each of the plurality of first data semantic vectors includes a dimension associated with a release time of the corresponding candidate data.
7. The method according to claim 1 or 2, wherein: Each of the plurality of first data semantic vectors includes a dimension associated with a source credibility of the corresponding candidate data.
8. The method according to claim 1 or 2, wherein: The first request semantic vector includes context information related to the user's search, and the context information includes at least one of time, location, and a previous search of the user.
9. The method according to claim 1 or 2, wherein: The multiple candidate data of different types also include at least a table and a knowledge graph.
10. The method according to claim 1 or 2, wherein: At least one text or video data among the multiple candidate data of different types is obtained by performing fine-grained division on the original complete data.
11. The method according to claim 1 or 2, wherein: The plurality of first data semantic vectors have a uniform specification.
12. The method of claim 3, wherein: The data volume of the second target output data set is smaller than the data volume of the first target output data set.
13. The method of claim 12, wherein: The data in the second target output data set is selected according to a predetermined quality standard.
14. An electronic device comprising: at least one processor; as well as a memory communicatively coupled to the at least one processor; in The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 13.
15. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 13.
16. A computer program product comprising a computer program, wherein: The computer program implements the method according to any one of claims 1 to 13 when executed by a processor.
Citation Information
Patent Citations
Semantic analysis model training method and device, electronic equipment and storage medium
CN112560496A
Search term recommendation method, target model training method, device and equipment
CN112650907A