Data search method and device, equipment and storage medium

By obtaining the feature vectors of user historical search data and the search terms entered by the user, the problem that traditional search engines cannot understand user needs is solved, and higher search results relevance and accuracy are achieved.

CN120336614APending Publication Date: 2025-07-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410078714.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Traditional search engine models based on keyword matching and semantic matching cannot fully understand user needs, resulting in low correlation between search results.

Method used

During data search, the preamble search data within the user's historical time period is obtained, feature extraction is performed, feature vectors are generated, and candidate data is recalled based on the feature vector and the search terms entered by the user, and the most relevant search results are selected.

Benefits of technology

It improves the matching degree between search results and user actual needs and improves the accuracy of data search.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336614A_ABST
    Figure CN120336614A_ABST
Patent Text Reader

Abstract

The invention provides a data search method and device, equipment and a storage medium, can be applied to the technical fields of short videos, artificial intelligence, intelligent vehicle-mounted devices and the like, and comprises the steps that a search word input by an object is acquired, and N pieces of preorder search data consumed by the object in a historical time period are acquired in response to the search word input by the object; performing feature extraction on P pieces of preorder search data in the N pieces of preorder search data to obtain a first feature vector; based on the search terms, M pieces of candidate search data are recalled; and based on the first feature vector, selecting at least one piece of search data from the M pieces of candidate search data as a search result. In the data search process, the actual consumption demand of the object is understood based on the preorder search data of the object, and then the search result is recalled based on the feature information and the search word of the preorder search data, so that the matching degree of the search result and the actual demand of the object can be improved, and the accuracy of data search is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of computer technologies, and in particular, to a data search method, apparatus, device, and storage medium. Background Art

[0002] With the rapid development of the Internet, digital content, such as video content, has become an important way for users to obtain information, entertainment, and learning. To meet the growing personalized needs of users, a search engine needs to be able to accurately understand the needs of users and provide highly relevant search results for users. However, traditional search engine models based on keyword matching often fail to fully understand the needs of users, resulting in low relevance of search results.

[0003] To improve the relevance of search results, those skilled in the art have proposed a search model based on semantic matching, that is, based on semantic analysis and matching of search terms and content to be searched, recall search results. However, the matching degree between the search results of the current search model based on semantic matching and the actual needs of users is still not high. Summary of the Invention

[0004] The present application provides a data search method, apparatus, device, and storage medium. When performing data search, N pieces of previous search data consumed by an object in a historical time period are considered to judge the potential search needs of the object, and the feature information of the N pieces of previous search data is injected into the search process, so that the matching degree between the search results and the actual needs of the object is high, thereby improving the accuracy of data search.

[0005] In a first aspect, the present application provides a data search method, including:

[0006] Obtain a search term input by an object, and in response to the search term input by the object, obtain N pieces of previous search data consumed by the object in a historical time period, where N is a positive integer;

[0007] Extract features from P pieces of the N pieces of previous search data to obtain a first feature vector, where P is a positive integer less than or equal to N;

[0008] Based on the search term, recall M pieces of candidate search data, where M is a positive integer greater than 1;

[0009] Based on the first feature vector and the search term, select at least one piece of candidate search data from the M pieces of candidate search data as a search result.

[0010] In a second aspect, the present application provides a data search apparatus, including:

[0011] An acquisition unit, configured to acquire a search term input by an object, and in response to the search term input by the object, acquire N pieces of previous search data consumed by the object within a historical time period, where N is a positive integer;

[0012] A feature extraction unit, configured to perform feature extraction on P pieces of the N pieces of previous search data to obtain a first feature vector, where P is a positive integer less than or equal to N;

[0013] A recall unit, configured to recall M pieces of candidate search data based on the search term, where M is a positive integer greater than 1;

[0014] A processing unit, configured to select at least one piece of candidate search data from the M pieces of candidate search data as a search result based on the first feature vector and the search term.

[0015] In some embodiments, when P is less than N, the feature extraction unit is specifically configured to, based on the search term, select the P pieces of previous search data from the N pieces of previous search data; perform feature extraction on the P pieces of previous search data to obtain the first feature vector.

[0016] In some embodiments, the feature extraction unit is specifically configured to, for the i-th piece of the N pieces of previous search data, determine the matching degree between the i-th piece of previous search data and the search term, where i is a positive integer less than or equal to N; based on the matching degree between each piece of previous search data and the search term, select the P pieces of previous search data from the N pieces of previous search data.

[0017] In some embodiments, the feature extraction unit is specifically configured to extract the core word of the search term, and extract the text information of the i-th piece of previous search data; based on the core word of the search term and the text information of the i-th piece of previous search data, determine the hit rate of the search term in the i-th piece of previous search data; based on the hit rate of the search term in the i-th piece of previous search data, determine the matching degree between the i-th piece of previous search data and the search term.

[0018] In some embodiments, when the i-th piece of previous search data is video search data, then in some embodiments, the feature extraction unit performs at least one of the following:

[0019] Extract the text information of the video title of the i-th piece of previous search data;

[0020] Perform text recognition on the video cover image of the i-th piece of previous search data to obtain the text information of the video cover image;

[0021] Perform semantic recognition on the speech content of the i-th piece of pre-search data to obtain the text information of the speech content

[0022] Perform text recognition on the video image content of the i-th piece of pre-search data to obtain the text information of the video image content.

[0023] In some embodiments, the feature extraction unit is specifically configured to represent each piece of the P pieces of pre-search data as a feature vector; based on the feature vector representations of each piece of the P pieces of pre-search data, determine the first feature vector.

[0024] In some embodiments, the feature extraction unit is specifically configured to perform feature transformation on the feature vector representations of the P pieces of pre-search data through a fully connected layer to obtain the first feature vector.

[0025] In some embodiments, the first feature vector is a one-dimensional feature vector.

[0026] In some embodiments, the processing unit is specifically configured to select at least one piece of candidate search data from the M pieces of candidate search data as the search result based on the first feature vector and the search term.

[0027] In some embodiments, the processing unit is specifically configured to, for the j-th piece of candidate search data among the M pieces of candidate search data, determine the score of the j-th piece of candidate search data based on the first feature vector, the search term, and the j-th piece of candidate search data, where j is a positive integer less than or equal to M; based on the scores of each piece of candidate search data among the M pieces of candidate search data, select at least one piece of candidate search data from the M pieces of candidate search data as the search result.

[0028] In some embodiments, the processing unit is specifically configured to determine the text information of the j-th piece of candidate search data; through a semantic matching model, process the first feature vector, the search term, and the text information of the j-th piece of candidate search data to obtain the score of the j-th piece of candidate search data.

[0029] In some embodiments, the processing unit is specifically configured to map each character in the search term to a token to obtain the token representation of the search term; map each character in the text information of the j-th piece of candidate search data to a token to obtain the token representation of the j-th piece of candidate search data; through the semantic matching model, process the first feature vector, the token representation of the search term, and the token representation of the j-th piece of candidate search data to obtain the score of the j-th piece of candidate search data.

[0030] In some embodiments, the processing unit is specifically configured to splice the first feature vector, the token representation of the search term, and the token representation of the j-th candidate search data in a preset order to obtain the input data of the semantic matching model; process the input data through the semantic matching model to obtain the score of the j-th candidate search data.

[0031] In some embodiments, the preset order is that the token representation of the search term is spliced after the first feature vector, and the token representation of the j-th candidate search data is spliced after the token representation of the search term.

[0032] In some embodiments, when the j-th candidate search data is video search data, the processing unit 14 is configured to perform at least one of the following;

[0033] Extract the text information of the video title of the j-th candidate search data;

[0034] Perform text recognition on the video cover image of the j-th candidate search data to obtain the text information of the video cover image;

[0035] Perform semantic recognition on the voice content of the j-th candidate search data to obtain the text information of the voice content

[0036] Perform text recognition on the video image content of the j-th candidate search data to obtain the text information of the video image content.

[0037] In some embodiments, the acquisition unit is specifically configured to acquire N pieces of previous search data consumed by the object during the historical time period based on the consumption behavior data of the object, where the consumption behavior data includes at least one of browsing duration, browsing times, liking, commenting, and sharing.

[0038] In a third aspect, an electronic device is provided, including a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the methods in the first aspect and its various implementation manners above.

[0039] In a fourth aspect, a chip is provided for implementing the methods in any one of the first aspect and its various implementation manners above. Specifically, the chip includes: a processor for calling and running a computer program from a memory, so that a device installed with the chip executes the methods in any one of the first aspect and its various implementation manners above.

[0040] In a fifth aspect, a computer-readable storage medium is provided for storing a computer program, and the computer program causes a computer to execute the methods in the first aspect and its various implementation manners above.

[0041] In a sixth aspect, a computer program product is provided, including computer program instructions, which cause a computer to execute the methods in the first aspect and its various implementation manners as described above.

[0042] In a seventh aspect, a computer program is provided, which causes a computer to execute the methods in the first aspect and its various implementation manners as described above when running on the computer.

[0043] In summary, when performing data search in this application, the previous consumption behavior of the object is considered. Specifically, the search term input by the object is obtained, and in response to the search term input by the object, N pieces of previous search data consumed by the object within a historical time period are obtained, where N is a positive integer. Feature extraction is performed on P pieces of the N pieces of previous search data to obtain a first feature vector, where P is a positive integer less than or equal to N. At the same time, based on the search term, M pieces of candidate search data are recalled, where M is a positive integer greater than 1. Then, based on the first feature vector, at least one piece of search data is selected from the M pieces of candidate search data and presented to the object as the search result. It can be seen that in the data search process of the embodiment of this application, in response to the search term input by the object, the previous search data consumed by the object recently is obtained, and based on this previous search data, the actual consumption demand of the object is understood. Then, based on the feature information of the previous search data and the search term, the search result is recalled, which can improve the matching degree between the search result and the actual demand of the object, and further improve the accuracy of data search. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 It is a schematic diagram of an implementation environment related to the embodiment of this application;

[0046] Figure 2 It is a schematic flowchart of a data search method provided by an embodiment of this application;

[0047] Figure 3 It is an example diagram for determining the first feature vector;

[0048] Figure 4 It is an example diagram for predicting the score of the j-th candidate search data through a semantic matching model;

[0049] Figure 5A It is a schematic diagram of a semantic matching model;

[0050] Figure 5B Another schematic diagram of an idol for the semantic matching model;

[0051] Figure 6 A schematic flowchart of the data search method provided by an embodiment of the present application;

[0052] Figure 7 An example diagram for searching videos;

[0053] Figure 8 A schematic diagram of an application example;

[0054] Figure 9 A schematic block diagram of the data search device provided by an embodiment of the present application;

[0055] Figure 10 A schematic block diagram of the electronic device provided by an embodiment of the present application. Specific embodiments

[0056] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0057] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In the embodiments of the present invention, "B corresponding to A" means that B is associated with A. In one implementation, B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices. In the description of the present application, unless otherwise specified, "a plurality" means two or more than two.

[0058] The technical solutions proposed in the present application can be applied to technical fields such as short videos, artificial intelligence, and intelligent vehicles, to improve the matching degree between search results and the actual needs of objects, and thus improve the accuracy of data search.

[0059] The relevant concepts involved in the embodiments of this application will be introduced below.

[0060] Short video: That is, a short film video, which is a way of spreading content on the Internet. Generally, it is a video with a duration within n minutes spread on new Internet media. With the popularization of mobile terminals and the speed increase of the network, short, flat, and fast large-traffic dissemination content has gradually gained the favor of the public. In some embodiments, the short video in the embodiments of this application may refer to a video number product or other short video platform products.

[0061] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, and to perceive the environment, acquire knowledge, and use knowledge to obtain the best results of theory, methods, techniques, and application systems. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machine to have the functions of perception, reasoning, and decision-making.

[0062] Artificial intelligence technology is an interdisciplinary subject involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0063] Machine Learning (ML) is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how a computer simulates or realizes human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve its own performance. Machine learning is the core of artificial intelligence and the fundamental way to make a computer intelligent. Its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0064] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0065] In the embodiments of the present application, artificial intelligence technology is applied to data search to improve the accuracy of data search.

[0066] With the rapid development of the Internet, digital content, such as video content, has become an important way for users to obtain information, entertainment, and learning. To meet the growing personalized needs of users, search engines need to be able to accurately understand the needs of the object (such as the user) and provide highly relevant search results for the object. However, traditional search engine models based on keyword matching often fail to fully understand the needs of the object, resulting in low relevance of search results.

[0067] To improve the relevance of search results, those skilled in the art have proposed a search model based on semantic matching, that is, based on semantic analysis and matching of search terms and content to be searched, recall search results. However, the matching degree of the current search results of the search model based on semantic matching with the actual needs of the object is still not high, resulting in unsatisfactory search accuracy.

[0068] To solve this technical problem, the data search method provided in the embodiments of the present application considers the object's previous consumption behavior during data search. Specifically, obtain the search term input by the object, and in response to the search term input by the object, obtain N pieces of previous search data consumed by the object within the historical time period. Extract features from P pieces of the N pieces of previous search data to obtain a first feature vector. At the same time, recall M pieces of candidate search data based on the search term. Then, based on the first feature vector and the search term, select at least one piece of search data from the M pieces of candidate search data as the search result to be presented to the object. It can be seen that in the data search process of the embodiments of the present application, in response to the search term input by the object, obtain the previous search data consumed by the object recently, and based on this previous search data, understand the object's actual consumption needs, and then recall the search results based on the feature information of the previous search data and the search term, which can improve the matching degree of the search results with the object's actual needs, and thus improve the accuracy of data search.

[0069] The following introduces the implementation environment of the data search method provided in the embodiments of the present application.

[0070] Figure 1A schematic diagram of an implementation environment related to an embodiment of the present application, including a terminal device 101 and a server 102. The terminal device 101 and the server 102 can be communicatively connected by wired or wireless means.

[0071] In an embodiment of the present application, an object can interact with the terminal device 101 to exchange data. For example, the object can input a search term on the terminal device 101. The terminal device 101 can also present the search results corresponding to the search term to the object.

[0072] The embodiment of the present application does not limit the specific execution entity of the data search method.

[0073] In some embodiments, the data search method of the embodiment of the present application is mainly completed by the terminal device. For example, the terminal device 101 obtains the search term input by the object. In the embodiment of the present application, the terminal device locally stores the search data consumed by the object in the recent period of time. In this way, in response to the search term input by the object, the terminal device 101 selects N pieces of previous search data from the search data consumed by the object in the recent period of time stored locally. Then, the terminal device extracts features from P pieces of the N pieces of previous search data to obtain a first feature vector. At the same time, the terminal device 101 of the embodiment of the present application also sends the search term input by the object to the server 102. The server 102 recalls M pieces of candidate search data based on the search term and sends the recalled M pieces of candidate search data to the terminal device 101. In this way, the terminal device 101 can select at least one search data from the M pieces of candidate search data recalled based on the search term as the search result based on the first feature vector corresponding to the N pieces of previous search data and the search term input by the object, and present the search result to the object.

[0074] In some embodiments, the data search method of the embodiments of the present application is mainly completed by the server. For example, the terminal device 101 obtains the search term input by the object and sends a search request to the server 102. The search request includes the search term and the identification information of the object. The server 102 of the embodiments of the present application locally stores the search data consumed by the object in the recent period of time. After receiving the search request sent by the terminal device 101, the server 102 selects N pieces of previous search data from the search data consumed by the object in the recent period of time stored locally based on the identification information of the object included in the search request. Then, the server 102 extracts features from P pieces of the N pieces of previous search data to obtain a first feature vector. At the same time, the server 102 also recalls M pieces of candidate search data based on the search term included in the search request. In this way, the server 102 can select at least one search data from the M pieces of candidate search data recalled based on the search term as the search result based on the first feature vector corresponding to the N pieces of previous search data and the search term input by the object. The server 102 sends the search result to the terminal device 101, and the terminal device 101 presents the search result to the object.

[0075] The embodiments of the present application do not limit the specific type of the terminal device 101. In some embodiments, the terminal device 101 may include but is not limited to: mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, wearable intelligent devices, medical devices, and so on. The device is often configured with a display device, and the display device may also be a monitor, a display screen, a touch screen, etc. The touch screen may also be a touch panel, a touch screen panel, etc.

[0076] In some embodiments, there may be one or more servers. When there are multiple servers, at least two servers are used to provide different services, and / or at least two servers are used to provide the same service, such as providing the same service in a load balancing manner. The embodiments of the present application do not limit this. Among them, the above-mentioned server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The server can also become a node of the blockchain.

[0077] In the embodiments of the present application, the terminal device 101 and the server 102 can be directly or indirectly connected through wired communication or wireless communication methods, and the present application does not limit this.

[0078] It should be noted that the implementation environment of the embodiments of the present application includes, but is not limited to, Figure 1 as shown.

[0079] The technical solutions of the embodiments of the present application will be described in detail below through some embodiments. These embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0080] Figure 2 FIG. is a schematic flowchart of a data search method provided by an embodiment of the present application. The execution subject of the embodiments of the present application is a device with a data search function, such as a data search device. In some embodiments, the execution subject of the embodiments of the present application may be Figure 1 the server in Figure 1 or the terminal device in Figure 1 or a system composed of a server and a terminal device in. For the sake of convenience of description, the embodiments of the present application are described by taking the execution subject as an electronic device as an example.

[0081] As Figure 2 shown, the data search process of the embodiments of the present application includes:

[0082] S101. Obtain the search term input by the object, and in response to the search term input by the object, obtain N pieces of previous search data consumed by the object within the historical time period.

[0083] Wherein, N is a positive integer.

[0084] The data search method of the embodiments of the present application can be applied to any type of data search, such as music search, text search, video search, or mixed data search, etc.

[0085] In the embodiments of the present application, during data search, based on the previous search data of the object, the actual consumption demand of the object is understood, thereby improving the matching degree between the search result and the actual demand of the object and enhancing the accuracy of data search. Specifically, the electronic device obtains the search term input by the object, and in response to the search term input by the object, obtains N pieces of previous search data consumed by the object within the historical time period, and performs subsequent data search based on the search term and the N pieces of previous search data.

[0086] The embodiments of the present application do not limit the specific manner in which the electronic device obtains the search term input by the object.

[0087] In some embodiments, the electronic device may be a terminal device, such as a user terminal device. Exemplary user terminal devices may be any user terminal device with a data search function, such as a smart phone, a notebook, or a vehicle-mounted terminal. At this time, the subject may input a search word on the electronic device, such as the subject inputting the search word on the electronic device by voice, or the subject inputting the search word on the electronic device by an input method, or by gestures. That is, the embodiment of the present application does not limit the input method of the subject inputting the search word.

[0088] In some embodiments, the electronic device may be a server, which is connected to the terminal device. For example, the terminal device may be a user terminal device such as a smart phone, a notebook, or a vehicle-mounted terminal. At this time, the subject may input a search word on the terminal device, for example, the subject may input the search word on the terminal device by voice, or the subject may input the search word on the terminal device by an input method. Then, the terminal device sends the search word input by the subject to the server.

[0089] The embodiment of the present application does not limit the triggering conditions of the subject entering the search term on the terminal device. Exemplarily, the triggering conditions include at least one of the following: the subject initiates the search after browsing the recommended data of the recommendation system, the subject clicks on the comment background word to initiate the search, and the subject clicks on the recommended term to send the search. This can exclude the scenario where the subject directly initiates the search without generating the previous consumption data.

[0090] In an embodiment of the present application, when the electronic device detects a search term input by an object, in response to the input operation of the search term, obtains N preceding search data consumed by the object within a historical time period.

[0091] In the embodiment of the present application, the electronic device obtains N preceding search data consumed by the object in the historical time period in at least the following ways:

[0092] Case 1: If the electronic device is a terminal device, illustratively, a client of a search application is installed on the terminal device, and the object can search for data on the client. For example, a client of a video playback application is installed on the terminal device, and the object can search for and play videos on the client of the video playback application.

[0093] In an example of this Case 1, the terminal device stores the historical consumption records of the object, that is, it stores multiple search data consumed by the object in the search application during a recent historical time period. For example, it stores multiple videos watched by the object during a recent historical time period. In the embodiments of this application, for the convenience of description, the search data consumed by the object previously is denoted as the previous search data. In this example, when the object inputs a search term on the client of the terminal device, the terminal device, in response to the search term input by the object, obtains N pieces of previous search data from the multiple pieces of previous search data consumed by the object in the search application during a recent historical time period stored locally on the terminal device.

[0094] In another example of this Case 1, the server side of the search application stores the historical consumption records of the object, that is, it stores multiple pieces of previous search data consumed by the object in the search application during a recent historical time period. In this example, when the object inputs a search term on the client of the terminal device, the terminal device, in response to the search term input by the object, sends the search term to the server. The server obtains N pieces of previous search data from the multiple pieces of previous search data consumed by the object in the search application during a recent historical time period stored, and sends the N pieces of previous search data to the terminal device.

[0095] Case 2, if the electronic device is a server. Exemplarily, a client of a search application is installed on the terminal device corresponding to the object, and the object can perform data search on the client. For example, a client of a video playback application is installed on the terminal device, and the object can perform video search and video playback on the client of the video playback application. The server can be understood as the background server of the search application, which is used to provide search resources.

[0096] In an example of this Case 2, the terminal device stores the historical consumption records of the object, that is, it stores multiple pieces of previous search data consumed by the object in the search application during a recent historical time period. In this example, when the object inputs a search term on the client of the terminal device, the terminal device, in response to the search term input by the object, obtains N pieces of previous search data from the multiple pieces of previous search data consumed by the object in the search application during a recent historical time period stored locally on the terminal device, and sends the obtained N pieces of previous search data to the server.

[0097] In this case 2, the server side of the search application stores the historical consumption records of the object, that is, it stores multiple pieces of previous search data consumed by the object in the search application during a recent historical time period. In this example, when the object enters a search term on the client side of the terminal device, the terminal device sends the search term to the server in response to the search term input by the object. The server obtains N pieces of previous search data from the multiple pieces of previous search data consumed by the object in the search application during the recent historical time period that are stored.

[0098] In some embodiments, the specific manner of obtaining N pieces of previous search data from the multiple pieces of previous search data consumed by the object in the search application during the recent historical time period can be to randomly select N pieces of previous search data from the multiple pieces of previous search data consumed by the object in the search application during the recent historical time period.

[0099] In some embodiments, the specific manner of obtaining N pieces of previous search data from the multiple pieces of previous search data consumed by the object in the search application during the recent historical time period can be to obtain N pieces of previous search data consumed by the object during the historical time period based on the consumption behavior data of the object. The consumption behavior data includes at least one of browsing duration, number of browsing times, likes, comments, and shares. For example, from the multiple pieces of previous search data consumed by the object in the search application during the recent historical time period, select the previous search data that meets at least one of the conditions of longer browsing duration, more browsing times, more like times, more comment times, and more share times, so as to obtain N pieces of previous search data.

[0100] In some embodiments, in order to further obtain forward search data related to the actual needs of the object, when the embodiment of the present application obtains N pieces of previous search data from the multiple pieces of previous search data consumed by the object in the search application during the recent historical time period, first, from the multiple pieces of previous search data consumed by the object in the search application during the recent historical time period, collect multiple pieces of previous search data consumed by the object in the recommendation application within a preset time window from the current search, for example, collect multiple pieces of previous search data consumed by the object in the recommendation application within a 1-hour time window from the current search. Then, select N pieces of previous search data from the multiple pieces of previous search data consumed by the object within this preset time window. For example, based on the consumption behavior data of the object, select N pieces of previous search data from the multiple pieces of previous search data consumed by the object within this preset time window.

[0101] After the electronic device obtains the search term input by the object and obtains N pieces of previous search data consumed by the object during the historical time period based on the above steps, it executes the following step S102.

[0102] S102 . Perform feature extraction on P pieces of the N pieces of the preceding search data to obtain a first feature vector.

[0103] Wherein, P is a positive integer less than or equal to N.

[0104] In an embodiment of the present application, in order to make the results of data search more in line with the actual needs of the object, N previous search data consumed by the object in a historical time period are obtained during the data search, and then the actual needs of the object are judged based on the N previous search data, so as to assist the subsequent data recall, so that the search results obtained are more in line with the actual needs of the object, so as to improve the accuracy of data search.

[0105] Specifically, after the electronic device obtains N pieces of preceding search data consumed by the object in a historical time period, it performs feature extraction on P pieces of preceding search data among the N pieces of preceding search data to obtain a first feature vector.

[0106] In some embodiments, P is equal to N, that is, the electronic device performs feature extraction on N pre-order search data to obtain a first feature vector. For example, N pre-order search data are input into a feature extraction model for feature extraction to obtain a feature vector corresponding to the N pre-order search data, and the feature vector is recorded as the first feature vector. For another example, the feature vector of each pre-order search data in the N pre-order search data is extracted, and then the feature vectors corresponding to the N pre-order search data are respectively fused to obtain a feature vector recorded as the first feature vector.

[0107] In some embodiments, in order to improve the accuracy of data search, it is necessary to remove the previous search data that is not related to the current search term from the N previous search data. At this time, P is less than N, and the above S102 includes the following steps S102-A and S102-B:

[0108] S102-A, based on the search term, select P pieces of preceding search data from N pieces of preceding search data;

[0109] S102-B, extract features from P pieces of preceding search data to obtain a first feature vector.

[0110] In this implementation, when the electronic device obtains N pieces of pre-order search data consumed by the object in the historical time period, in order to improve the accuracy of data search, the electronic device removes the pre-order search data that is not relevant to the current search term from the N pieces of pre-order search data. Specifically, based on the search term currently input by the object, the electronic device selects P pieces of pre-order search data with a high correlation with the search term from the N pieces of pre-order search data consumed by the object in the historical time period for subsequent data search.

[0111] The following introduces the specific process of the electronic device selecting P pieces of previous search data from N pieces of previous search data based on the search term.

[0112] In a possible implementation, for each piece of previous search data among the N pieces of previous search data, semantic analysis is performed on the search term and this piece of previous search data to determine the semantic relevance between this piece of previous search data and the search term. Exemplarily, the electronic device obtains a semantic matching model, which is used to predict the semantic relevance between two input text messages. In this way, the electronic device can obtain the text information of this forward search data, and then input the search term and the text information of this forward search data into the semantic matching model for semantic matching to determine the relevance between the search term and this forward search data. Based on this method, the electronic device can determine the relevance between each piece of previous search data among the N pieces of previous search data and the search term, and then based on the relevance, select P pieces of previous search data with a relatively high relevance to the search term from the N pieces of previous search data.

[0113] In a possible implementation, the above S102-A includes the following steps S102-A1 and S102-A2:

[0114] S102-A1. For the i-th piece of previous search data among the N pieces of previous search data, determine the matching degree between the i-th piece of previous search data and the search term, where i is a positive integer less than or equal to N;

[0115] S102-A2. Based on the matching degree between each piece of previous search data and the search term, select P pieces of previous search data from the N pieces of previous search data.

[0116] In this implementation, the electronic device determines the matching degree between each piece of previous search data among the N pieces of previous search data and the search term, and then based on the matching degree, selects P pieces of previous search data from the N pieces of previous search data.

[0117] In the embodiments of the present application, the methods for the electronic device to determine the matching degree between each piece of previous search data among the N pieces of previous search data and the search term are basically the same. For the convenience of description, here, taking the determination of the matching degree between the i-th piece of previous search data among the N pieces of previous search data and the search term as an example for illustration.

[0118] The embodiments of the present application do not limit the specific manner in which the electronic device determines the matching degree between the i-th piece of previous search data and the search term.

[0119] In some examples, the electronic device obtains a matching model, which can predict the matching degree between two input data. Based on this, the electronic device can input the i-th previous search data and the search term into the matching model for matching prediction to obtain the matching degree between the i-th previous search data and the search term.

[0120] In some examples, the electronic device determines the matching degree between the i-th previous search data and the search term through the following steps S102-A11 to S102-A13:

[0121] S102-A11: Extract the core word of the search term and extract the text information of the i-th previous search data;

[0122] S102-A12: Based on the core word of the search term and the text information of the i-th previous search data, determine the hit rate of the search term in the i-th previous search data;

[0123] S102-A13: Based on the hit rate of the search term in the i-th previous search data, determine the matching degree between the i-th previous search data and the search term.

[0124] In this implementation manner, when the electronic device calculates the matching degree between the i-th previous search data and the search term, it extracts the core word of the search term. For example, through search term analysis technology, the fourth-level and fifth-level core words in the search term are extracted, and the fourth-level and fifth-level core words represent the most important words for expressing the search term. In this way, when determining the matching degree between the search term and the i-th previous search data based on the core word of the search term, the amount of calculation data can be reduced while ensuring the accuracy of the matching degree calculation, and the speed of calculating the matching degree between the search term and the i-th previous search data can be improved.

[0125] While extracting the core word of the search term, the electronic device also extracts the text information of the i-th previous search data.

[0126] In one example, if the i-th previous search data includes text data, then the text data in the i-th previous search data is used as the text information of the i-th previous search data or as a part of the text information of the i-th previous search data.

[0127] In one example, if the i-th previous search data includes an image and the image includes text, then the text in the image is recognized, and the recognized text is used as the text information of the i-th previous search data or as a part of the text information of the i-th previous search data.

[0128] In one example, if the i-th previous search data is video search data, then the text information of the i-th previous search data is extracted, including at least one of the following:

[0129] Extract the text information of the video title of the i-th piece of pre-search data;

[0130] Perform text recognition on the video cover image of the i-th piece of pre-search data to obtain the text information of the video cover image;

[0131] Perform semantic recognition on the speech content of the i-th piece of pre-search data to obtain the text information of the speech content

[0132] Perform text recognition on the video image content of the i-th piece of pre-search data to obtain the text information of the video image content.

[0133] Among them, the video title of the i-th piece of pre-search data is usually text data, so the text information of the video title of the i-th piece of pre-search data can be directly obtained.

[0134] If the video cover image of the i-th piece of pre-search data includes text, then recognize the text in the video cover image to obtain the text information of the video cover image.

[0135] If the i-th piece of pre-search data also includes speech information, then perform semantic recognition on the speech content of the i-th piece of pre-search data. For example, through Automatic Speech Recognition (ASR) technology, perform semantic recognition on the speech content of the i-th piece of pre-search data to obtain the text information of the speech content of the i-th piece of pre-search data.

[0136] In this example, the i-th piece of pre-search data is video search data, and the video search data includes multiple video images, such as including multiple video image frames. For each video image content, recognize the text in the video image content. For example, use Optical Character Recognition (OCR) technology to recognize the text in the video image content to obtain the text of the video image content. Based on this method, text recognition can be performed on all video image contents included in the video search content to obtain the text information of the video image content of the i-th piece of pre-search data.

[0137] In this example, when the i-th previous search data is video search data, the electronic device extracts the text information of the video title of the i-th previous search data, and / or performs text recognition on the video cover image of the i-th previous search data to obtain the text information of the video cover image, and / or performs semantic recognition on the speech content of the i-th previous search data to obtain the text information of the speech content, and / or performs text recognition on the video image content of the i-th previous search data to obtain the text information of the video image content. In this way, at least one of the text information of the video title of the i-th previous search data, the text information of the video cover image, the text information of the speech content, and the text information of the video image content is used as the text information of the i-th previous search data.

[0138] Based on the above steps, after the electronic device extracts the core word of the search term and extracts the text information of the i-th previous search data, it executes the steps of S102-A12 above, and determines the hit rate of the search term in the i-th previous search data based on the core word of the search term and the text information of the i-th previous search data.

[0139] The embodiments of the present application do not limit the specific manner in which the electronic device determines the hit rate of the search term in the i-th previous search data based on the core word of the search term and the text information of the i-th previous search data.

[0140] In a possible implementation manner, the electronic device determines the number of core words included in the search term that are hit in the text information of the i-th previous search data. For example, the search term includes 3 core words. Assume that the first core word is hit in the text information of the i-th previous search data, that is, the text information of the i-th previous search data includes the first core word. Assume that the second core word is not hit in the text information of the i-th previous search data, that is, the text information of the i-th previous search data does not include the second core word. Assume that the third core word is hit in the text information of the i-th previous search data, that is, the text information of the i-th previous search data includes the third core word. At this time, it can be determined that 2 of the 3 core words in the search term are hit in the text information of the i-th previous search data. In this way, it can be determined that the hit rate of the 3 core words in the search term in the text information of the i-th previous search data is 2 / 3, and then this hit rate is determined as the hit rate of the search term in the i-th previous search data.

[0141] Exemplarily, the electronic device can determine the hit rate of the search term in the i-th previous search data through the following formula (1):

[0142]

[0143] Among them, term_hitrate represents the hit rate of the search term in the i-th previous search data, term_num represents the number of core terms included in the search term, and ∑ term is_hit represents the number of core terms in the core terms that hit the text information of the i-th previous search data.

[0144] After the electronic device determines the naming rate of the search term in the i-th previous search data based on the above steps, based on this hit rate, it determines the matching degree between the i-th previous search data and the search term. For example, the naming rate of the search term in the i-th previous search data is determined as the matching degree between the i-th previous search data and the search term. Of course, it is also possible to perform preset processing on the naming of the search term in the i-th previous search data, such as multiplying by a number greater than 0, to obtain the matching degree between the i-th previous search data and the search term. It should be noted that the naming rate of the search term in the i-th previous search data is positively correlated with the matching degree between the search term and the i-th previous search data, that is, the higher the naming rate of the search term in the i-th previous search data, the higher the matching degree between the i-th previous search data and the search term.

[0145] The above text introduces the specific process of the electronic device determining the matching degree between the i-th previous search data and the search term. The electronic device can use the same method to determine the matching degree between each of the N previous search data and the search term. Then, the electronic device can select P previous search data from the N previous search data based on the matching degree between each of the N previous search data and the search term. For example, select the P previous search data with the largest matching degree with the search term from the N previous search data.

[0146] After the electronic device selects P previous search data from the N previous search data based on the above steps, it executes the steps of S102-B above to perform feature extraction on the P previous search data to obtain the first feature vector.

[0147] The embodiments of the present application do not limit the specific manner in which the electronic device performs feature extraction on the P previous search data to obtain the first feature vector.

[0148] In some embodiments, the electronic device obtains a feature extraction model, inputs the P previous search data into the feature extraction model for feature extraction, and obtains the first feature vector. The embodiments of the present application do not limit the specific network structure of the feature extraction model. Optionally, the feature extraction model includes at least one convolutional layer.

[0149] In some embodiments, the above S102-B includes the following steps of S102-B1 and S102-B2:

[0150] S102 - B1. Perform eigenvector representation on each piece of the P pieces of pre - order search data;

[0151] S102 - B2. Determine the first eigenvector based on the eigenvector representation of each piece of the P pieces of pre - order search data.

[0152] In this implementation manner, when the electronic device determines the first eigenvector corresponding to the P pieces of pre - order search data, it performs vector representation on each piece of the P pieces of pre - order search data to obtain the eigenvector representation of each piece of the P pieces of pre - order search data.

[0153] For example, for each piece of the P pieces of pre - order search data, a low - dimensional, dense, floating - point eigenvector representation (embedding representation) is used. In one example, a single - tower representation model based on transformer can be used to perform vector representation on each piece of the P pieces of pre - order search data. For example, the pre - order search data can be represented as an eigenvector representation of m×n. The specific values of m and n in the embodiments of this application are not limited. For example, m = 1 and n is equal to 768. That is to say, each piece of the P pieces of pre - order search data is represented as an eigenvector representation of 1×768.

[0154] After the electronic device performs eigenvector representation on each piece of the P pieces of pre - order search data, it determines the first eigenvector based on the eigenvector representation of each piece of the P pieces of pre - order search data.

[0155] The embodiments of this application do not limit the specific manner in which the electronic device determines the first eigenvector based on the eigenvector representation of each piece of the P pieces of pre - order search data.

[0156] In one possible implementation manner, the electronic device fuses the eigenvector representations of each piece of the P pieces of pre - order search data to obtain the first eigenvector. Since the dimensions of the eigenvector representations of each piece of the P pieces of pre - order search data are the same, in one example, the electronic device can add the eigenvector representations of each piece of the P pieces of pre - order search data to obtain the first eigenvector. In one example, the electronic device can splice the eigenvector representations of each piece of the P pieces of pre - order search data to obtain the first eigenvector. In one example, the electronic device can multiply the eigenvector representations of each piece of the P pieces of pre - order search data to obtain the first eigenvector.

[0157] In a possible implementation, the electronic device performs feature transformation on the feature vector representations of P pieces of previous search data through a fully connected layer to obtain a first feature vector. Specifically, as Figure 3 shown, the electronic device inputs the feature vector representations of P pieces of previous search data into the fully connected layer for feature transformation to obtain a first feature vector.

[0158] The embodiments of the present application do not limit the specific scale size of the first feature vector.

[0159] In an example, the first feature vector is a one-dimensional feature vector. For example, the first feature vector is a feature vector with a size of 1×768.

[0160] As can be seen from the above, the first feature vector in the embodiments of the present application includes the feature information of the previous search data of the object in hours within the historical time period. In this way, when performing subsequent data searches based on this first feature vector, the accuracy of data searches can be improved.

[0161] S103. Recall M pieces of candidate search data based on the search term.

[0162] Where M is a positive integer greater than 1.

[0163] It should be noted that the embodiments of the present application do not limit the specific execution order of the electronic device to recall M pieces of candidate search data based on the search term, obtain N pieces of previous search data consumed by the object within the historical time period in response to the search term input by the object, and perform feature extraction on P pieces of the N pieces of previous search data to obtain the first feature vector.

[0164] In an example, the electronic device first obtains N pieces of previous search data consumed by the object within the historical time period in response to the search term input by the object, and then recalls M pieces of candidate search data based on the search term.

[0165] In an example, the electronic device obtains N pieces of previous search data consumed by the object within the historical time period in response to the search term input by the object, and at the same time recalls M pieces of candidate search data based on the search term.

[0166] In an example, after the electronic device performs feature extraction on P pieces of the N pieces of previous search data to obtain the first feature vector, it recalls M pieces of candidate search data based on the search term.

[0167] In an example, when the electronic device performs feature extraction on P pieces of the N pieces of previous search data to obtain the first feature vector, it recalls M pieces of candidate search data based on the search term.

[0168] The embodiments of the present application for an electronic device to recall M candidate search data based on a search term include at least the following situations:

[0169] Situation 1: If the electronic device is a terminal device, the electronic device sends the search term input by the user to the server, and the server recalls M candidate search data based on the search term. For example, the server analyzes and understands the search term, and at the same time analyzes and understands each search data (i.e., item) in the search database. Then, the server calculates the correlation between the search term and each search data based on the understanding result of the search term and the understanding result of each search data, and then recalls M candidate search data from the search database based on the correlation. Then, the server sends the recalled M candidate search data to the terminal device.

[0170] Situation 2: If the electronic device is a server, the electronic device sends the search term input by the user to the server, and the server recalls M candidate search data based on the search term. For example, the server analyzes and understands the search term, and at the same time analyzes and understands each search data (i.e., item) in the search database. Then, the server calculates the correlation between the search term and each search data based on the understanding result of the search term and the understanding result of each search data, and then recalls M candidate search data from the search database based on the correlation.

[0171] S104. Select at least one search data from the M candidate search data as the search result based on the first feature vector.

[0172] In the embodiments of the present application, after the electronic device detects that the user inputs a search term, it obtains N previous search data consumed by the user in the historical time period, and extracts features from P of the N previous search data to obtain the first feature vector. At the same time, based on the search term, M candidate search data are recalled. Then, the electronic device selects at least one search data from the M candidate search data as the search result based on the first feature vector. Since the first feature vector includes the feature information of the previous search data consumed by the user in the historical time period, based on this first feature vector, at least one candidate search data that meets the actual needs of the user can be selected from the M candidate search data as the search result, thereby improving the accuracy of data search.

[0173] The embodiments of the present application do not limit the specific manner in which the electronic device selects at least one search data from the M candidate search data as the search result based on the first feature vector.

[0174] In some embodiments, for each piece of candidate search data among M pieces of candidate search data, a score of the candidate search data is determined based on the first feature vector and the candidate search data. For example, the first feature vector and the candidate search data are input into a semantic matching model for score calculation to obtain the score of the candidate search data. Furthermore, at least one piece of candidate search data with the highest score among the M pieces of candidate search data is determined as the search result corresponding to the search term.

[0175] In some embodiments, the above S104 includes the following steps of S104-A:

[0176] S104-A: Select at least one piece of candidate search data from the M pieces of candidate search data as the search result based on the first feature vector and the search term.

[0177] In this implementation manner, when selecting at least one piece of candidate search data from the M pieces of candidate search data, the electronic device not only considers the first feature vector but also the search term, making the search result more accurate.

[0178] In some embodiments, after the electronic device performs feature fusion on the first feature vector and the search term, a second feature vector is obtained. For each piece of candidate search data among the M pieces of candidate search data, a score of the candidate search data is determined based on the second feature vector and the candidate search data. For example, the second feature vector and the candidate search data are input into a semantic matching model for score calculation to obtain the score of the candidate search data. Furthermore, at least one piece of candidate search data with the highest score among the M pieces of candidate search data is determined as the search result corresponding to the search term.

[0179] In some embodiments, the above S104-A includes the following steps of S104-A1 and S104-A2:

[0180] S104-A1: For the j-th piece of candidate search data among the M pieces of candidate search data, determine the score of the j-th piece of candidate search data based on the first feature vector, the search term, and the j-th piece of candidate search data, where j is a positive integer less than or equal to M;

[0181] S104-A2: Select at least one piece of candidate search data from the M pieces of candidate search data as the search result based on the scores of each piece of candidate search data among the M pieces of candidate search data.

[0182] In this implementation manner, for each piece of candidate search data among the M pieces of candidate search data, the electronic device determines the score of the candidate search data based on the first feature vector, the search term, and the j-th piece of candidate search data. Furthermore, based on the scores, at least one piece of candidate search data is selected from the M pieces of candidate search data as the search result, achieving search accuracy.

[0183] In the embodiments of the present application, the specific process of the electronic device determining the scores of each of the M candidate search data is basically the same. For ease of description, the process of determining the score of the j-th candidate search data is taken as an example for illustration herein.

[0184] In the embodiments of the present application, the electronic device determines the score of the j-th candidate search data based on the first feature vector, the search term, and the j-th candidate search data. That is, in the embodiments of the present application, when calculating the score of the j-th candidate search data, not only the search term is considered, but also the first feature vector including the feature information of the previous search data consumed by the object in the historical time period is considered, so that the accurate calculation of the score of the j-th candidate search data can be realized.

[0185] The embodiments of the present application do not limit the specific manner in which the electronic device determines the score of the j-th candidate search data based on the first feature vector, the search term, and the j-th candidate search data.

[0186] In a possible implementation manner, the electronic device calculates the matching degree between the j-th candidate search data and the search term, and at the same time calculates the matching degree between the j-th candidate search data and the first feature vector. Then, based on the matching degree between the j-th candidate search data and the search term, and the matching degree between the j-th candidate search data and the first feature vector, the score of the j-th candidate search data is determined. For example, the sum of the matching degree between the j-th candidate search data and the search term, and the matching degree between the j-th candidate search data and the first feature vector is determined as the score of the j-th candidate search data.

[0187] In a possible implementation manner, the electronic device can determine the score of the j-th candidate search data through the following steps S104-A11 and S104-A12:

[0188] S104-A11: Determine the text information of the j-th candidate search data;

[0189] S104-A12: Through the semantic matching model, process the first feature vector, the search term, and the text information of the j-th candidate search data to obtain the score of the j-th candidate search data.

[0190] In this implementation manner, for candidate search data in different scenarios, the manner of obtaining the text information of the candidate search data is different.

[0191] In an example, if the j-th candidate search data includes text data, the text data in the j-th candidate search data is used as the text information of the j-th candidate search data or as a part of the text information of the j-th candidate search data.

[0192] In one example, if the j-th candidate search data includes an image and the image includes text, then the text in the image is recognized, and the recognized text is used as the text information of the j-th candidate search data or as a part of the text information of the j-th candidate search data.

[0193] In one example, if the j-th candidate search data is video search data, then the text information of the j-th candidate search data is extracted, including at least one of the following:

[0194] Extract the text information of the video title of the j-th candidate search data;

[0195] Perform text recognition on the video cover image of the j-th candidate search data to obtain the text information of the video cover image;

[0196] Perform semantic recognition on the voice content of the j-th candidate search data to obtain the text information of the voice content

[0197] Perform text recognition on the video image content of the j-th candidate search data to obtain the text information of the video image content.

[0198] Among them, the video title of the j-th candidate search data is usually text data, so the text information of the video title of the j-th candidate search data can be directly obtained.

[0199] If the video cover image of the j-th candidate search data includes text, then the text in the video cover image is recognized to obtain the text information of the video cover image.

[0200] If the j-th candidate search data also includes voice information, then perform semantic recognition on the voice content of the j-th candidate search data. For example, through ASR technology, perform semantic recognition on the voice content of the j-th candidate search data to obtain the text information of the voice content of the j-th candidate search data.

[0201] In this example, the j-th candidate search data is video search data, and the video search data includes multiple video images, such as including multiple video image frames. For each video image content, the text in the video image content is recognized. For example, OCR technology is used to recognize the text in the video image content to obtain the text of the video image content. Based on this method, text recognition can be implemented for all video image contents included in the video search content, and the text information of the video image content of the j-th candidate search data can be obtained.

[0202] In this example, when the j-th candidate search data is video search data, the electronic device extracts the text information of the video title of the j-th candidate search data, and / or performs text recognition on the video cover image of the j-th candidate search data to obtain the text information of the video cover image, and / or performs semantic recognition on the speech content of the j-th candidate search data to obtain the text information of the speech content, and / or performs text recognition on the video image content of the j-th candidate search data to obtain the text information of the video image content. In this way, at least one of the text information of the video title of the j-th candidate search data, the text information of the video cover image, the text information of the speech content, and the text information of the video image content is used as the text information of the j-th candidate search data.

[0203] Based on the above steps, the electronic device determines the text information of the j-th candidate search data. As Figure 4 shown, the first feature vector, the search term, and the text information of the j-th candidate search data are used as the input of the semantic matching model, and the score of the j-th candidate search data output by the semantic matching model is obtained.

[0204] The specific network structure of the semantic matching model is not limited in the embodiments of the present application.

[0205] In a possible implementation manner, as Figure 5A shown, the semantic matching model is a transformer encoder(). Assume that the search term in the embodiments of the present application is "recommend good-looking notebooks", and the text information of the j-th candidate search data is "extremely good-looking...". In this way, as Figure 5A shown, the electronic device inputs the first feature vector, the search term, and the text information of the j-th candidate search data as the input information of the semantic matching model into the semantic matching model. The semantic matching model processes the first feature vector, the search term, and the text information of the j-th candidate search data to obtain the score of the j-th candidate search data.

[0206] In a possible implementation manner, as Figure 5B shown, the semantic matching model is a BERT model. In this way, as Figure 5B shown, the electronic device inputs the first feature vector, the search term, and the text information of the j-th candidate search data as the input information of the semantic matching model into the semantic matching model. The semantic matching model processes the first feature vector, the search term, and the text information of the j-th candidate search data to obtain the score of the j-th candidate search data.

[0207] In some embodiments, the above S104-A12 includes the following steps of S104-A121 to S104-A123:

[0208] S104 - A121. Map each character in the search term to a token to obtain the token representation of the search term;

[0209] S104 - A122. Map each character in the text information of the j - th candidate search data to a token to obtain the token representation of the j - th candidate search data;

[0210] S104 - A123. Through a semantic matching model, process the first feature vector, the token representation of the search term, and the token representation of the j - th candidate search data to obtain the score of the j - th candidate search data.

[0211] It should be noted that there is no order distinction between the above S104 - A122 and S104 - A121 in specific implementation, that is, S104 - A121 can be executed after S104 - A122, or before S104 - A122, or executed synchronously with S104 - A122.

[0212] In this implementation manner, before the electronic device inputs the first feature vector, the search term, and the j - th candidate search data into the semantic matching model, it first maps each character in the search term to a token to obtain the token representation of the search term. For example, map each character in the search term to a token representation (token emb) of size m×n.

[0213] Meanwhile, map each character in the text information of the j - th candidate search data to a token to obtain the token representation of the j - th candidate search data. For example, map each character in the text information of the j - th candidate search data to a token representation (token emb) of size m×n.

[0214] Then, through the semantic matching model, process the first feature vector, the token representation of the search term, and the token representation of the j - th candidate search data to obtain the score of the j - th candidate search data.

[0215] In some embodiments, the electronic device concatenates the first feature vector, the token representation of the search term, and the token representation of the j - th candidate search data in a preset order to obtain the input data of the semantic matching model, and then processes the input data through the semantic matching model to obtain the score of the j - th candidate search data.

[0216] The embodiments of this application do not limit the above - mentioned preset order.

[0217] In one example, the preset order is that the token representation of the search term is concatenated after the first feature vector, and the token representation of the j - th candidate search data is concatenated after the token representation of the search term. For example Figure 5A and 5BAs shown, the electronic device places the first feature vector (i.e., the pre-order search data representation) of size m×n (e.g., 1×768) after the CLS vector, concatenates the token representation of the search term after the first feature vector, and concatenates the token representation of the j-th candidate search data after the token representation of the search term.

[0218] In some embodiments, as shown in Figure 5A and 5B the figure, the token representation of the search term and the token representation of the j-th candidate search data are separated by the [SEP] separator.

[0219] In some embodiments, the first feature vector and the token representation of the search term can also be separated by the [SEP] separator.

[0220] Based on the above steps, the electronic device can determine the score of the j-th candidate search data. Similarly, the electronic device can use the same method to determine the score of each candidate search data among the M candidate search data. In this way, the electronic device can select at least one candidate search data from the M candidate search data as the search result based on the scores of each candidate search data among the M candidate search data.

[0221] In one example, assume that the recommendation system sets the search result to include Q search data. In this way, the electronic device can select the Q candidate search data with higher scores from the M candidate search data as the search result. Among them, the Q candidate search data are sorted in descending order according to the scores.

[0222] In one example, the recommendation system does not limit the number of search data included in the search result. At this time, the electronic device can sort the M candidate search data based on the scores and use the sorted candidate search data as the search result.

[0223] The training process of the above semantic matching model in the embodiments of the present application will be introduced below.

[0224] In the embodiments of the present application, the training data of the semantic matching model is the click processing log table, that is, for a single search (query), the search data clicked by the object is determined as the true search result of the search term. Then, based on the search term, multiple candidate search data are recalled, and at the same time, multiple previous search data consumed by the object before triggering the search term are obtained. Using the same method as above, the feature vectors corresponding to the multiple previous search data are determined, denoted as the previous feature vectors. In this way, the search term, the previous feature vectors, and the text of the recalled candidate search data are input into the semantic matching model to obtain the scores of each candidate search data input into the semantic matching model. Then, based on the scores of each candidate search data and the true search result of the search term, the model loss of the semantic matching model is determined. Furthermore, the model parameters in the semantic matching model are adjusted based on the model loss. Repeat the above steps until the model training end condition is reached. The evaluation metrics of the model are accuracy and AUC, where AUC is called the Area Under Curve.

[0225] The embodiments of the present application do not limit the type of loss function used for training the semantic matching model.

[0226] In one example, the loss function of the semantic matching model is a pointwise loss function. For example, the loss function of the semantic matching model is shown in Equation (2):

[0227]

[0228] where y t represents the label of sample t, that is, whether the target result user clicks. Clicking is 1, and not clicking is 0. p t represents the probability value that the predicted label of sample t is 1, T represents the number of samples in each batch. x represents the input sample.

[0229] The data search method provided by the embodiments of the present application takes into account the object's previous consumption behavior during data search. Specifically, the search term input by the object is obtained, and in response to the search term input by the object, N pieces of previous search data consumed by the object within a historical time period are obtained, where N is a positive integer. Feature extraction is performed on P pieces of the N pieces of previous search data to obtain a first feature vector, where P is a positive integer less than or equal to N. At the same time, based on the search term, M candidate search data are recalled, where M is a positive integer greater than 1. Then, based on the first feature vector and the search term, at least one search data is selected from the M candidate search data and presented to the object. It can be seen that during the data search process of the embodiments of the present application, in response to the search term input by the object, the previous search data consumed by the object recently is obtained, and based on this previous search data, the actual consumption demand of the object is understood. Then, based on the feature information of the previous search data and the search term, the search results are recalled, which can improve the matching degree between the search results and the actual needs of the object, and thus improve the accuracy of data search.

[0230] The above provides an overall introduction to the data search process of the embodiments of the present application. The following combines Figure 6 to further introduce the data search method provided by the embodiments of the present application by taking video search as an example.

[0231] Figure 6 It is a schematic flowchart of the data search method provided by an embodiment of the present application. This embodiment takes video search as an example. Figure 6 The shown embodiment can be understood as a specific implementation manner of the above Figure 2 shown embodiment.

[0232] As Figure 6 shown, the data search method of the embodiments of the present application includes the following steps:

[0233] S201. Obtain the search term input by the object, and in response to the search term input by the object, obtain N pieces of previous video data consumed by the object within a historical time period.

[0234] Taking video search as an example in the embodiments of the present application, the N pieces of previous video data consumed by the corresponding object within a historical time period are previous video data, that is, the video data consumed by the object in the early stage in the video search application.

[0235] The specific implementation process of the above S201 can refer to the relevant description of S101 above, and will not be elaborated here.

[0236] S202. Based on the search term, select P pieces of previous video data from the N pieces of previous video data.

[0237] For example, for the i-th piece of pre-order video data among N pieces of pre-order video data, determine the matching degree between the i-th piece of pre-order video data and the search term, where i is a positive integer less than or equal to N. Exemplarily, extract the core word of the search term, and extract the text information of the i-th piece of pre-order video data. For example, extract the text information of the video title of the i-th piece of pre-order video data, perform text recognition on the video cover image of the i-th piece of pre-order video data to obtain the text information of the video cover image, perform semantic recognition on the speech content of the i-th piece of pre-order video data to obtain the text information of the speech content, and perform text recognition on the video image content of the i-th piece of pre-order video data to obtain the text information of the video image content. Then, based on the core word of the search term and the text information of the i-th piece of pre-order video data, determine the hit rate of the search term in the i-th piece of pre-order video data. Furthermore, based on the hit rate of the search term in the i-th piece of pre-order video data, determine the matching degree between the i-th piece of pre-order video data and the search term. Finally, based on the matching degree between each piece of pre-order video data and the search term, select P pieces of pre-order video data from the N pieces of pre-order video data.

[0238] The specific implementation process of the above S202 can refer to the relevant description of the above S102, which will not be elaborated here.

[0239] S203. Extract features from the P pieces of pre-order video data to obtain a first feature vector.

[0240] For example, perform feature vector representation on each piece of pre-order video data among the P pieces of pre-order video data. Then, based on the feature vector representation of each piece of pre-order video data among the P pieces of pre-order video data, determine the first feature vector. Exemplarily, perform feature transformation on the feature vector representation of the P pieces of pre-order video data through a fully connected layer to obtain the first feature vector.

[0241] Optionally, the above first feature vector is a one-dimensional feature vector.

[0242] The specific implementation process of the above S203 can refer to the relevant description of the above S102, which will not be elaborated here.

[0243] S204. Recall M pieces of candidate video data based on the search term.

[0244] Where M is a positive integer greater than 1.

[0245] In the embodiment of the present application, taking video search as an example, the candidate video data recalled based on the search term is candidate video data.

[0246] The specific implementation process of the above S204 can refer to the relevant description of the above S103, which will not be elaborated here.

[0247] S205. For the j-th candidate video data among the M candidate video data, determine the text information of the j-th candidate video data.

[0248] Where j is a positive integer less than or equal to M.

[0249] In some embodiments, determining the text information of the j-th candidate video data includes at least one of the following:

[0250] Extract the text information of the video title of the j-th candidate video data;

[0251] Perform text recognition on the video cover image of the j-th candidate video data to obtain the text information of the video cover image;

[0252] Perform semantic recognition on the speech content of the j-th candidate video data to obtain the text information of the speech content, and perform text recognition on the video image content of the j-th candidate video data to obtain the text information of the video image content.

[0253] S206. Process the first feature vector, the search term, and the text information of the j-th candidate video data through a semantic matching model to obtain the score of the j-th candidate video data.

[0254] In some embodiments, the electronic device maps each character in the search term to a token to obtain the token representation of the search term; maps each character in the text information of the j-th candidate video data to a token to obtain the token representation of the j-th candidate video data; and processes the first feature vector, the token representation of the search term, and the token representation of the j-th candidate video data through a semantic matching model to obtain the score of the j-th candidate video data.

[0255] In some embodiments, the electronic device concatenates the first feature vector, the token representation of the search term, and the token representation of the j-th candidate video data in a preset order to obtain the input data of the semantic matching model; and processes the input data through the semantic matching model to obtain the score of the j-th candidate video data.

[0256] The embodiments of the present application do not limit the preset order.

[0257] In one example, the preset order is that the token representation of the search term is concatenated after the first feature vector, and the token representation of the j-th candidate video data is concatenated after the token representation of the search term.

[0258] Exemplarily, such as Figure 7As shown, the electronic device splices the token representation of the search term after the first feature vector, and splices the token representation of the j-th candidate video data after the token representation of the search term to obtain the input data of the semantic matching model, and then inputs the input data into the semantic matching model for processing to obtain the score of the j-th candidate video data.

[0259] The specific implementation process of the above S206 can refer to the relevant description of the above S104 and will not be elaborated here.

[0260] S207. Based on the scores of each of the M candidate video data, at least one candidate video data is selected from the M candidate video data as the search result.

[0261] For example, according to the scores, the M candidate video data are sorted, and the sorted M candidate video data are presented to the user as the search result.

[0262] In the embodiment of the present application, as Figure 8 shown, take the example of a user searching for videos in a video number recommendation system. As Figure 8 can be seen from the left figure, the user browsed videos about basketball team A previously. After that, when the user searches for the query term "A", the user's demand is clearly for basketball team A at this time. However, if the previous video information consumed by the user in the recommendation system is not analyzed, then directly searching for "A" is an ambiguous term, and its intentions include basketball teams, movie names, game names, the word itself, etc., making the search results not meet the actual needs of the user. In the embodiment of the present application, when searching for videos, the previous consumption behavior of the user is considered, so that the actual needs of the user can be fully understood. Then, when performing subsequent data searches based on the actual needs of the user, Figure 8 the search results shown on the right can be obtained, making the search results more in line with the actual needs of the user, thereby improving the search accuracy of videos.

[0263] The video search method provided by the embodiment of the present application obtains the search term input by the user, and in response to the search term input by the user, obtains N pieces of previous video data consumed by the user within a historical time period. Feature extraction is performed on P pieces of the N pieces of previous video data to obtain a first feature vector. At the same time, based on the search term, M pieces of candidate video data are recalled. Furthermore, based on the first feature vector and the search term, at least one piece of candidate video data is selected from the M pieces of candidate video data as the search result and presented to the user. It can be seen that in the video search process of the embodiment of the present application, in response to the search term input by the user, the previous video data consumed by the user recently is obtained, and based on this previous video data, the actual video consumption demand of the user is understood. Furthermore, based on the feature information of the previous video data and the search term, the video search result is recalled, which can improve the matching degree between the video search result and the actual demand of the user, and further improve the accuracy of video search.

[0264] As described above in conjunction with Figures 2 to 6 , the embodiment of the model training and image processing method of the present application is described in detail. Below in conjunction with Figure 8 , the device embodiment of the present application is described in detail.

[0265] Figure 9 FIG. is a schematic block diagram of a data search device provided by an embodiment of the present application. The device 10 can be applied to an electronic device.

[0266] As Figure 9 shown, the data search device 10 includes:

[0267] An acquisition unit 11, configured to acquire a search term input by a user, and in response to the search term input by the user, acquire N pieces of previous search data consumed by the user within a historical time period, where N is a positive integer;

[0268] A feature extraction unit 12, configured to perform feature extraction on P pieces of the N pieces of previous search data to obtain a first feature vector, where P is a positive integer less than or equal to N;

[0269] A recall unit 13, configured to recall M pieces of candidate search data based on the search term, where M is a positive integer greater than 1;

[0270] A processing unit 14, configured to select at least one piece of candidate search data from the M pieces of candidate search data as a search result based on the first feature vector and the search term.

[0271] In some embodiments, when P is less than N, the feature extraction unit 12 is specifically configured to select the P pieces of previous search data from the N pieces of previous search data based on the search term; perform feature extraction on the P pieces of previous search data to obtain the first feature vector.

[0272] In some embodiments, the feature extraction unit 12 is specifically configured to determine the matching degree between the i-th pre-search data among the N pre-search data and the search term, where i is a positive integer less than or equal to N; and select the P pre-search data from the N pre-search data based on the matching degree between each pre-search data and the search term.

[0273] In some embodiments, the feature extraction unit 12 is specifically configured to extract the core word of the search term and the text information of the i-th pre-search data; determine the hit rate of the search term in the i-th pre-search data based on the core word of the search term and the text information of the i-th pre-search data; and determine the matching degree between the i-th pre-search data and the search term based on the hit rate of the search term in the i-th pre-search data.

[0274] In some embodiments, when the i-th pre-search data is video search data, in some embodiments, the feature extraction unit 12 performs at least one of the following:

[0275] Extract the text information of the video title of the i-th pre-search data;

[0276] Perform text recognition on the video cover image of the i-th pre-search data to obtain the text information of the video cover image;

[0277] Perform semantic recognition on the voice content of the i-th pre-search data to obtain the text information of the voice content

[0278] Perform text recognition on the video image content of the i-th pre-search data to obtain the text information of the video image content.

[0279] In some embodiments, the feature extraction unit 12 is specifically configured to represent each of the P pre-search data as a feature vector; and determine the first feature vector based on the feature vector representation of each of the P pre-search data.

[0280] In some embodiments, the feature extraction unit 12 is specifically configured to perform feature transformation on the feature vector representation of the P pre-search data through a fully connected layer to obtain the first feature vector.

[0281] In some embodiments, the first feature vector is a one-dimensional feature vector.

[0282] In some embodiments, the processing unit 14 is specifically configured to select at least one candidate search data from the M candidate search data as the search result based on the first feature vector and the search term.

[0283] In some embodiments, for the j-th candidate search data among the M candidate search data, the processing unit 14 is specifically configured to determine the score of the j-th candidate search data based on the first feature vector, the search term, and the j-th candidate search data, where j is a positive integer less than or equal to M; and select at least one candidate search data from the M candidate search data as the search result based on the scores of each candidate search data among the M candidate search data.

[0284] In some embodiments, the processing unit 14 is specifically configured to determine the text information of the j-th candidate search data; and process the first feature vector, the search term, and the text information of the j-th candidate search data through a semantic matching model to obtain the score of the j-th candidate search data.

[0285] In some embodiments, the processing unit 14 is specifically configured to map each character in the search term to a token to obtain the token representation of the search term; map each character in the text information of the j-th candidate search data to a token to obtain the token representation of the j-th candidate search data; and process the first feature vector, the token representation of the search term, and the token representation of the j-th candidate search data through the semantic matching model to obtain the score of the j-th candidate search data.

[0286] In some embodiments, the processing unit 14 is specifically configured to splice the first feature vector, the token representation of the search term, and the token representation of the j-th candidate search data in a preset order to obtain the input data of the semantic matching model; and process the input data through the semantic matching model to obtain the score of the j-th candidate search data.

[0287] In some embodiments, the preset order is that the token representation of the search term is spliced after the first feature vector, and the token representation of the j-th candidate search data is spliced after the token representation of the search term.

[0288] In some embodiments, when the j-th candidate search data is video search data, the processing unit 14 is configured to perform at least one of the following:

[0289] Extract the text information of the video title of the j-th candidate search data;

[0290] Perform text recognition on the video cover image of the j-th candidate search data to obtain the text information of the video cover image;

[0291] Perform semantic recognition on the voice content of the j-th candidate search data to obtain the text information of the voice content

[0292] Perform text recognition on the video image content of the j-th candidate search data to obtain the text information of the video image content.

[0293] In some embodiments, the obtaining unit 11 is specifically configured to obtain N pieces of previous search data consumed by the object during the historical time period based on the consumption behavior data of the object, where the consumption behavior data includes at least one of browsing duration, browsing times, liking, commenting, and sharing.

[0294] It should be understood that the device embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, they will not be elaborated here. Specifically, Figure 9 The shown device can execute the embodiments of the above illegal video account recognition method, and the foregoing and other operations and / or functions of each module in the device respectively implement the corresponding method embodiments of the electronic device. For the sake of brevity, they will not be elaborated here.

[0295] The device of the embodiments of the present application has been described above from the perspective of functional modules in combination with the drawings. It should be understood that the functional modules can be implemented in the form of hardware, can also be implemented by instructions in the form of software, and can also be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiments in the present application can be completed by the integrated logic circuit of the hardware in the processor and / or instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.

[0296] Figure 10 is a schematic block diagram of the electronic device provided by the embodiments of the present application, Figure 10 The electronic device can be used to execute the above method embodiments.

[0297] Such as Figure 10 As shown, the electronic device 30 may include:

[0298] A memory 31 and a processor 32. The memory 31 is used to store a computer program 33 and transmit the program code 33 to the processor 32. In other words, the processor 32 can call and run the computer program 33 from the memory 31 to implement the method in the embodiments of the present application.

[0299] For example, the processor 32 can be used to execute the steps in the above method according to the instructions in the computer program 33.

[0300] In some embodiments of the present application, the processor 32 may include, but is not limited to:

[0301] A general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and so on.

[0302] In some embodiments of the present application, the memory 31 includes, but is not limited to:

[0303] A volatile memory and / or a non-volatile memory. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0304] In some embodiments of the present application, the computer program 33 may be divided into one or more modules, which are stored in the memory 31 and executed by the processor 32 to complete the method for recording pages provided in the present application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 33 in the electronic device.

[0305] As Figure 10 shown, the electronic device 30 may further include:

[0306] A transceiver 34, which may be connected to the processor 32 or the memory 31.

[0307] Among them, the processor 32 can control the transceiver 34 to communicate with other devices. Specifically, it can send information or data to other devices, or receive information or data sent by other devices. The transceiver 34 may include a transmitter and a receiver. The transceiver 34 may further include an antenna, and the number of antennas may be one or more.

[0308] It should be understood that the various components in the electronic device 30 are connected through a bus system. Among them, the bus system includes not only a data bus, but also a power bus, a control bus, and a status signal bus.

[0309] According to one aspect of the present application, there is provided a computer storage medium, on which a computer program is stored. When the computer program is executed by the computer, the computer can execute the method in the above method embodiments. Or rather, the embodiments of the present application further provide a computer program product containing instructions. When the instructions are executed by the computer, the computer executes the method in the above method embodiments.

[0310] According to another aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method in the above method embodiments.

[0311] In other words, when implemented using software, it can be implemented in the form of a computer program product in whole or in part. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0312] Those of ordinary skill in the art will realize that the modules and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0313] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or modules can be in electrical, mechanical, or other forms.

[0314] The module described as a separate component may or may not be physically separated. The component shown as a module may or may not be a physical module, that is, it may be located in one place or distributed across multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of this application, each functional module can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.

[0315] The above content is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A data search method, characterized in that, Including: Obtaining the search term input by the object, and in response to the search term input by the object, obtaining N pieces of previous search data consumed by the object within a historical time period, where N is a positive integer; Performing feature extraction on P pieces of the N pieces of previous search data to obtain a first feature vector, where P is a positive integer less than or equal to N; Based on the search term, recalling M pieces of candidate search data, where M is a positive integer greater than 1; Based on the first feature vector and the search term, selecting at least one piece of candidate search data from the M pieces of candidate search data as the search result.

2. The method according to claim 1, wherein When P is less than N, the performing feature extraction on P pieces of the N pieces of previous search data to obtain the first feature vector includes: Based on the search term, selecting the P pieces of previous search data from the N pieces of previous search data; Performing feature extraction on the P pieces of previous search data to obtain the first feature vector.

3. The method according to claim 2, wherein The selecting the P pieces of previous search data from the N pieces of previous search data based on the search term includes: For the i-th piece of previous search data among the N pieces of previous search data, determining the matching degree between the i-th piece of previous search data and the search term, where i is a positive integer less than or equal to N; Based on the matching degree between each piece of previous search data and the search term, selecting the P pieces of previous search data from the N pieces of previous search data.

4. The method according to claim 3, characterized in that, The determining the matching degree between the i-th piece of previous search data and the search term includes: Extracting the core word of the search term, and extracting the text information of the i-th piece of previous search data; Based on the core word of the search term and the text information of the i-th piece of previous search data, determining the hit rate of the search term in the i-th piece of previous search data; Based on the hit rate of the search term in the i-th piece of previous search data, determining the matching degree between the i-th piece of previous search data and the search term.

5. The method according to claim 4, wherein When the i-th piece of previous search data is video search data, the extracting the text information of the i-th piece of previous search data includes at least one of the following: Extracting the text information of the video title of the i-th piece of previous search data; Performing text recognition on the video cover image of the i-th piece of previous search data to obtain the text information of the video cover image; Performing semantic recognition on the voice content of the i-th piece of previous search data to obtain the text information of the voice content Performing text recognition on the video image content of the i-th piece of previous search data to obtain the text information of the video image content.

6. The method according to claim 2, wherein The performing feature extraction on the P pieces of previous search data to obtain the first feature vector includes: Performing feature vector representation on each piece of the P pieces of previous search data; Based on the feature vector representation of each piece of the P pieces of previous search data, determining the first feature vector.

7. The method according to claim 6, wherein The determining the first feature vector based on the feature vector representation of each piece of the P pieces of previous search data includes: Perform feature transformation on the feature vector representations of the P pieces of previous search data through a fully connected layer to obtain the first feature vector.

8. The method according to claim 7, wherein The first feature vector is a one-dimensional feature vector.

9. The method according to any one of claims 1-8, characterized in that Selecting at least one candidate search data from the M candidate search data as the search result based on the first feature vector includes: Selecting at least one candidate search data from the M candidate search data as the search result based on the first feature vector and the search term.

10. The method according to claim 9, characterized in that, Selecting at least one candidate search data from the M candidate search data as the search result based on the first feature vector and the search term includes: For the j-th candidate search data among the M candidate search data, determine the score of the j-th candidate search data based on the first feature vector, the search term, and the j-th candidate search data, where j is a positive integer less than or equal to M; Select at least one candidate search data from the M candidate search data as the search result based on the scores of each candidate search data among the M candidate search data.

11. The method according to claim 10, wherein, Determining the score of the j-th candidate search data based on the first feature vector, the search term, and the j-th candidate search data includes: Determine the text information of the j-th candidate search data; Process the first feature vector, the search term, and the text information of the j-th candidate search data through a semantic matching model to obtain the score of the j-th candidate search data.

12. The method according to claim 11, wherein Processing the first feature vector, the search term, and the text information of the j-th candidate search data through a semantic matching model to obtain the score of the j-th candidate search data includes: Map each word in the search term to a token to obtain the token representation of the search term; Map each word in the text information of the j-th candidate search data to a token to obtain the token representation of the j-th candidate search data; Process the first feature vector, the token representation of the search term, and the token representation of the j-th candidate search data through the semantic matching model to obtain the score of the j-th candidate search data.

13. The method according to claim 12, wherein, Processing the first feature vector, the token representation of the search term, and the token representation of the j-th candidate search data through the semantic matching model to obtain the score of the j-th candidate search data includes: Concatenate the first feature vector, the token representation of the search term, and the token representation of the j-th candidate search data in a preset order to obtain the input data of the semantic matching model; Process the input data through the semantic matching model to obtain the score of the j-th candidate search data.

14. The method according to claim 13, wherein The preset order is that the token representation of the search term is concatenated after the first feature vector, and the token representation of the j-th candidate search data is concatenated after the token representation of the search term.

15. The method according to claim 10, wherein If the j-th candidate search data is video search data, then determining the text information of the j-th candidate search data includes at least one of the following: Extracting the text information of the video title of the j-th candidate search data; Performing text recognition on the video cover image of the j-th candidate search data to obtain the text information of the video cover image; Performing semantic recognition on the voice content of the j-th candidate search data to obtain the text information of the voice content Performing text recognition on the video image content of the j-th candidate search data to obtain the text information of the video image content.

16. The method according to claim 1, characterized in that, The obtaining of the N pieces of previous search data consumed by the object within the historical time period includes: Based on the consumption behavior data of the object, obtaining the N pieces of previous search data consumed by the object within the historical time period, where the consumption behavior data includes at least one of browsing duration, number of browsing times, likes, comments, and shares.

17. A data search device, characterized in that, Including: An obtaining unit, configured to obtain a search term input by an object, and in response to the search term input by the object, obtain N pieces of previous search data consumed by the object within a historical time period, where N is a positive integer; A feature extraction unit, configured to extract features from P pieces of the N pieces of previous search data to obtain a first feature vector, where P is a positive integer less than or equal to N; A recall unit, configured to recall M pieces of candidate search data based on the search term, where M is a positive integer greater than 1; A processing unit, configured to select at least one piece of candidate search data from the M pieces of candidate search data as a search result based on the first feature vector and the search term.

18. A computer device, including a processor and a memory; The memory is configured to store a computer program; The processor is configured to execute the computer program to implement the method according to any one of claims 1 to 16 above.

19. A computer-readable storage medium, characterized in that, For storing a computer program; The computer program causes the computer to execute the method according to any one of claims 1 to 16 above.