A data processing method, apparatus, device, and medium
By constructing a sample sequence based on historical search data and training a word vector model, the problem of insufficient sample number is solved, and the training effect of the network model and the matching degree of search results is improved.
Patent Information
- Application Number
- CN202110004769.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-04
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-01-04
AI Technical Summary
In the prior art, constructing sample data based on session records of a single user results in insufficient sample size, affecting the training process and performance of the network model.
By obtaining historical search data, a sample sequence corresponding to the historical search field is constructed, and it is added to the sample set. The word vector model is trained using the sample set to obtain the trained word vector model.
It enriches the number of samples, improves the training effect and performance of the network model, and ensures the matching and accuracy of the search results.
Smart Images

Figure CN113392310B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, specifically to the field of artificial intelligence, and particularly to a data processing method, a data processing device, a data processing equipment, and a computer-readable storage medium. Background Art
[0002] In search application scenarios, search engines (such as Baidu Search, Sogou Search, Toutiao Search, etc.) play a relatively important role. Search engines mainly rely on network models to implement search functions. In the prior art, the sample data used to train network models is constructed based on the one-time or multiple complete session records of a single user. However, it is found in practice that the method of constructing sample data based on the session records of a single user will have the problem of insufficient sample quantity due to the integrity and quantity of session records, thus affecting the training process of network models and the performance of the trained network models. Summary of the Invention
[0003] Embodiments of this application provide a data processing method, device, equipment, and medium, which can solve the problem of insufficient sample quantity, ensure the training process of network models, and ensure the performance of the trained network models.
[0004] On the one hand, embodiments of this application provide a data processing method, which includes:
[0005] Obtain historical search data, where the historical search data includes the first historical search field generated within a historical time period, the account data of N co-occurrence documents triggered based on the first historical search field, and the trigger time of each co-occurrence document, and N is a positive integer; a co-occurrence document refers to a search result triggered among one or more search results that match the first historical search field;
[0006] Construct a first sample sequence corresponding to the first historical search field according to the historical search data, where the first sample sequence includes the account data of N co-occurrence documents arranged in sequence according to the order of trigger time;
[0007] Add the first sample sequence to a sample set, where the sample set includes M sample sequences, M historical search fields are generated within the historical time period, and the first historical search field is any one of the M historical search fields; one sample in the sample set corresponds to one historical search field generated within the historical time period, and M is a positive integer;
[0008] Train a word vector model using the sample set to obtain a trained word vector model.
[0009] On the other hand, embodiments of this application provide a data processing device, which includes:
[0010] An acquisition unit, configured to acquire historical search data, where the historical search data includes a first historical search field generated within a historical time period, account data of N co-occurring documents triggered based on the first historical search field, and the trigger time of each co-occurring document, and N is a positive integer; a co-occurring document refers to a search result triggered among one or more search results that match the first historical search field;
[0011] A processing unit, configured to construct a first sample sequence corresponding to the first historical search field according to the historical search data, where the first sample sequence includes the account data of N co-occurring documents arranged in sequence according to the order of trigger time; add the first sample sequence to a sample set, where the sample set includes M sample sequences, M historical search fields are generated within the historical time period, and the first historical search field is any one of the M historical search fields; one sample sequence in the sample set corresponds to one historical search field generated within the historical time period, and M is a positive integer; and, train a word vector model using the sample set to obtain a trained word vector model.
[0012] In an implementation manner, the processing unit is further configured to:
[0013] In response to a search request carrying a target search field, call the trained word vector model to process the target search field to obtain a word vector of the target search field;
[0014] Extract the account data of a first document to be matched from a database, and call the trained word vector model to process the account data of the first document to obtain a word vector corresponding to the account data of the first document;
[0015] Calculate a matching degree between the word vector of the target search field and the word vector corresponding to the account data of the first document;
[0016] If the matching degree is higher than a threshold, determine the first document as a target document that matches the target search field.
[0017] In an implementation manner, the processing unit is further configured to:
[0018] In response to a search request carrying a target search field, search for P second documents that match the target search field in a database, and P is a positive integer;
[0019] Call the trained word vector model to process the target search field to obtain a word vector of the target search field; and, call the trained word vector model to process the account data of the P second documents to obtain word vectors of the account data of each second document;
[0020] Calculate the matching degree between the word vectors of the account data of each second document and the word vector of the target search field respectively;
[0021] Sort the P second documents in descending order of the matching degree.
[0022] In one implementation, the processing unit is further configured to:
[0023] Display a document search page of the application, where the document search page includes a search box, search options, and a query result display area;
[0024] When there is an input operation on the search box, display the input target search field in the search box;
[0025] When a search request carrying the target search field is generated by selecting the search options, sequentially display the P second documents in the query result display area.
[0026] In one implementation, if the account data includes the identifier of the application to which the document belongs and the identifier of the social public service account to which the document belongs; then the processing unit is further configured to:
[0027] When the account data of any second document displayed in the query result display area is triggered, jump to the service interface corresponding to the account data of the triggered second document; the service interface includes the service interface of the application to which the triggered second document belongs, or the service interface of the social public service account to which the triggered second document belongs.
[0028] In one implementation, the processing unit is specifically configured to:
[0029] Obtain the first historical search field generated within a historical time period;
[0030] Query multiple matching documents that match the first historical search field from the database;
[0031] Determine the N triggered matching documents among the multiple matching documents as co-occurrence documents; and,
[0032] Obtain the account data of the N co-occurrence documents and the trigger time of each co-occurrence document.
[0033] In one implementation, the word vector model is a word vector model constructed by using the first prediction method; the reference word vectors of the account data of each co-occurrence document in the first sample sequence are further included in the sample set; the processing unit is specifically configured to:
[0034] Input the account data of the \(i\)-th co-occurring document in the first sample sequence into the word vector model for prediction processing to obtain the word vectors of the account data of other co-occurring documents in the first sample sequence except the account data of the \(i\)-th co-occurring document;
[0035] Optimize the word vector model based on the difference between the predicted word vectors of the account data of other co-occurring documents and the reference word vectors of the account data of other co-occurring documents;
[0036] Use each sample sequence in the sample set to perform iterative optimization training on the word vector model to obtain the trained word vector model.
[0037] In one implementation, the word vector model is a word vector model constructed using the second prediction method; the sample set also includes the reference word vectors of the account data of each co-occurring document in the first sample sequence; the processing unit is specifically configured to:
[0038] Input the account data of other co-occurring documents in the first sample sequence except the account data of the \(i\)-th co-occurring document into the word vector model for prediction processing to obtain the word vector of the account data of the \(i\)-th co-occurring document;
[0039] Optimize the word vector model based on the difference between the predicted word vector of the account data of the \(i\)-th co-occurring document and the reference word vector of the account data of the \(i\)-th co-occurring document;
[0040] Use each sample sequence in the sample set to perform iterative optimization training on the word vector model to obtain the trained word vector model.
[0041] In one implementation, the account data includes any one of the following: the identifier of the document, the identifier of the publisher of the document, the identifier of the application to which the document belongs, the identifier of the social public service account to which the document belongs; the co-occurring document includes any one of the following: text, picture, video or audio.
[0042] On the other hand, an embodiment of the present application provides a data processing device, and the device includes:
[0043] A processor, adapted to implement one or more instructions; and,
[0044] A computer-readable storage medium, storing one or more instructions, and the one or more instructions are adapted to be loaded and executed by the processor to perform the data processing method as described above.
[0045] On the other hand, an embodiment of the present application provides a computer-readable storage medium, and the computer-readable storage medium stores one or more instructions, and the one or more instructions are adapted to be loaded and executed by the processor to perform the above data processing method.
[0046] On the other hand, an embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned data processing method.
[0047] In an embodiment of the present application, a sample set can be constructed based on M historical search fields generated within a historical time period and account data of all co-occurring documents triggered by the M historical search fields; since the M historical search fields and the corresponding co-occurring documents are data generated when one or more users perform search operations within the historical time period, the sample set constructed based on these data has a relatively rich number of samples. Using this sample set to train a word vector model can obtain a trained word vector model with better performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0049] Figure 1 FIG. shows a schematic structural diagram of a data processing system provided by an exemplary embodiment of the present application;
[0050] Figure 2 FIG. shows a schematic flowchart of a data processing method proposed by an exemplary embodiment of the present application;
[0051] Figure 3 FIG. shows a schematic diagram of constructing a sample sequence provided by an exemplary embodiment of the present application;
[0052] Figure 4 FIG. shows a schematic diagram of the network structure of a skip-gram model provided by an exemplary embodiment of the present application;
[0053] Figure 5 FIG. shows a schematic diagram of the network structure of a CBOW model provided by an exemplary embodiment of the present application;
[0054] Figure 6 FIG. shows a schematic flowchart of another data processing method provided by an exemplary embodiment of the present application;
[0055] Figure 7 FIG. shows a schematic diagram of a document search page provided by an exemplary embodiment of the present application;
[0056] Figure 8 Shows a schematic flowchart of a process for searching for a target search field provided by an exemplary embodiment of the present application;
[0057] Figure 9 Shows a schematic diagram of a second document order display provided by an exemplary embodiment of the present application;
[0058] Figure 10 Shows a schematic diagram of a service interface corresponding to account data displayed by an exemplary embodiment of the present application;
[0059] Figure 11 Shows a schematic structural diagram of a data processing device provided by an exemplary embodiment of the present application;
[0060] Figure 12 Shows a schematic structural diagram of a data processing device provided by an exemplary embodiment of the present application. Detailed implementation manners
[0061] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0062] An embodiment of the present application proposes a data processing solution. This data processing solution involves technologies such as machine learning in artificial intelligence. Among them:
[0063] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0064] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence, and its applications cover all fields of artificial intelligence. Machine learning and deep learning generally include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. Machine learning can be regarded as a task whose goal is to enable machines (computers in a broad sense) to obtain human-like intelligence through learning. For example, humans can play Go, and computer programs (AlphaGo or AlphaGoZero) are designed to master Go knowledge and be able to play Go. Among them, multiple methods can be used to achieve the task of machine learning, such as neural networks, linear regression, decision trees, support vector machines, Bayesian classifiers, reinforcement learning, probabilistic graphical models, clustering, and so on.
[0065] Among them, the Neural Network is a method to achieve the task of machine learning. When talking about neural networks in the field of machine learning, it generally refers to "neural network learning". It is a network structure composed of many simple elements. This network structure is similar to the biological nervous system and is used to simulate the interaction between organisms and the natural environment. Moreover, the more network structures there are, the richer the functions of the neural network tend to be. The neural network is a relatively broad concept. For different learning tasks such as speech, text, and images, neural network models more suitable for specific learning tasks have emerged, such as the Recurrent Neural Network (RNN), Convolutional Neural Network (CNN), fully convolutional neural network (FCNN), and so on.
[0066] Embodiments of the present application relate to a search system. The essence of the search system refers to the process of obtaining search results that match the given search field (query) by running a search engine for the search field given by the user. The search field given by the user can be composed of one or more characters, and the characters can include at least one of the following: Chinese characters (i.e., Chinese words), English characters (i.e., letters), numbers, etc.; for example, the search field is "Braised Pork", which is composed of three Chinese characters. The search system can be set in various Internet products that support the search function, and the Internet products here can include but are not limited to: websites, applications, plugins in applications, or subroutines in applications, etc. For example, the WeChat application is equipped with a search system that can support the search function, including but not limited to: the search function for mini-programs in the WeChat application, the search function for public service accounts in the WeChat application, the search function for social accounts in the WeChat application, and the search function for pictures, videos, texts, etc. generated in the WeChat application.
[0067] The search function is supported by a network model running in the search system. The network model can be obtained by training a machine learning model (e.g., a neural network model, including: a recurrent neural network model, a convolutional neural network model, etc.) with sample data. The more the number of sample data used to train the machine learning model, the better the performance of the trained network model, and the more matching the search results obtained based on the trained network model are with the search fields. An embodiment of the present application provides a data processing solution, which supports using the data of multiple users collected within a historical time period as sample data to enrich the number of sample data. The data processing solution may include: ① Obtaining historical search data, which may include a first historical search field generated within the historical time period, account data of N co-occurring documents triggered based on the first historical search field, and the trigger time of each co-occurring document, where N is a positive integer; a co-occurring document refers to a document that appears together; a co-occurring document based on the first historical search field refers to a document that appears together with the first historical search field, that is, one or more search results that match the first historical search field obtained by searching based on the first historical search field. It should be particularly noted that the N co-occurring documents in the embodiment of the present application specifically refer to N search results triggered (e.g., clicked, selected, etc.) among the one or more search results that match the first historical search field; ② Constructing a first sample sequence corresponding to the first historical search field according to the historical search data, where the first sample sequence includes the account data of the N co-occurring documents arranged in the order of the trigger time; ③ Adding the first sample sequence to the sample set, where the sample set contains M sample sequences, M historical search fields are generated within the historical time period, the first historical search field is any one of the M historical search fields, and one sample sequence in the sample set corresponds to one historical search field generated within the historical time period, where M is a positive integer; ④ Training a word vector model (i.e., the above-mentioned network model) with the sample set to obtain a trained word vector model. In the above process, the sample set constructed based on the M historical search fields and the corresponding co-occurring documents has a relatively rich number of samples. Using this sample set to train the word vector model can obtain a trained word vector model with better performance.
[0068] The data processing solution provided by the embodiments of the present application can be applied to the cold start scenario of a search system. Among them, the cold start scenario of the search system can include the following scenarios: ① User cold start scenario: It refers to the scenario where when a user newly registers as a member in an application program with a search function, the search system recommends documents (such as videos, pictures, and audios) to the user; in this scenario, the search system lacks the user's historical data (such as browsing records, etc.), and cannot accurately recommend satisfactory search results to the user. ② Content cold start scenario: It refers to the scenario where when a certain document is first published to the application program, the search system distributes the document to users; in this scenario, the search system lacks relevant data of the document (such as the number of shares, the number of favorites, etc.), and it is difficult to accurately recommend the document to interested users. ③ Application program (such as a search application program) cold start scenario: It refers to the scenario where when the application program is first launched, the search system trains a network model; in this scenario, the search system lacks user data and cannot well train a network model with good performance to achieve accurate search. Based on the above description, it can be seen that the search system often faces the problem of missing sample data (such as the behavior data of newly registered users does not exist in the database, and the behavior data of any users does not exist for the newly launched application program, etc.) in the cold start scenario. Therefore, adopting the data processing solution provided by the embodiments of the present application in the cold start scenario of the search system can enrich the quantity of sample data, so that a word vector model with better performance can be trained based on the rich sample data. It should be noted that the embodiments of the present application are applicable not only to the cold start scenario of the search system but also to other scenarios lacking sample data. The above description takes the cold start scenario of the search system as an example and does not limit the embodiments of the present application.
[0069] To better understand the data processing solution proposed by the embodiments of the present application, the data processing solution involved in the embodiments of the present application will be introduced below in combination with an actual search application scenario. Please refer to Figure 1 , Figure 1 shows a schematic architecture diagram of a data processing system provided by an exemplary embodiment of the present application; as Figure 1As shown in the figure, the data processing system may include, but is not limited to: a server 101 and a terminal 102. The embodiments of the present application do not limit the naming and quantity of the server, nor the naming and quantity of the terminal. The terminal 102 may refer to a device for running an application program (i.e., an application program with a search function, such as a WeChat application program). The terminal 102 may include, but is not limited to: a PC (Personal Computer), a PDA (tablet computer), a mobile phone, a wearable intelligent device, etc. The terminal is often configured with a display device, and the display device may also be a monitor, a display screen, a touch screen, etc. The touch screen may also be a touch panel, etc. The display device may be used to display search results. The server 101 may be a background server of the application program in the terminal, used to interact with the terminal to provide computing and application service support for the application program in the terminal. The server 101 may also be a device for constructing a network model. The server 101 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal and the server may be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this here.
[0070] When the terminal 102 detects a search field given by the user, it calls the search system loaded by the search application to perform search analysis on the search field (such as performing word segmentation processing on the search field, etc.); the terminal 102 roughly recalls a large number of search results from the database of the server 101 based on the analysis results, so that the search system can perform fine sorting operations on these search results (such as screening and sorting the search results recalled roughly), etc., to obtain search results that match the search field and are sorted in order.
[0071] It should be noted that a trained word vector model runs in the search system, and some operations (such as rough recall and fine sorting) performed by the search system on the search field can be completed by the trained word vector model. The embodiments of the present application support using a sample set composed of historical search data to train the word vector model to obtain a trained word vector model with better performance; when the trained word vector model is used to execute the above search process, search results with a higher matching degree to the search field can be obtained. The following will be combined with the attached Figure 2 to illustrate the training steps of the word vector model.
[0072] Please refer to Figure 2 , Figure 2The figure shows a schematic flowchart of a data processing method proposed by an exemplary embodiment of the present application; this data processing method can be executed by a data processing device, which can be Figure 1 the terminal or server shown in the figure, and the data processing device can also be any device with model training function other than the terminal and the server. This solution includes but is not limited to steps S201 - S204, where:
[0073] S201. Obtain historical search data.
[0074] The historical search data is the data generated by the user's search operations during a historical time period. Herein, the historical time period can refer to a certain historical time period before the current moment. For example, the historical time period can be a certain day or multiple days in history, etc. The historical search data may include the first historical search field generated during the historical time period (i.e., a certain search term during the historical time period), the account data of N co - occurrence documents triggered based on the first historical search field (i.e., those search results clicked by the user among one or more search results matching the first historical search field), and the trigger time of each co - occurrence document, where N is a positive integer. Among them, the co - occurrence document can include any one of the following: text, picture, video, or audio, etc. The account data of the co - occurrence document can include any one of the following: the identifier of the document (such as the number of the video, a field used to uniquely identify the video (including characters such as numbers or letters)), the identifier of the publisher of the document (such as the account ID of the video publisher), the identifier of the application to which the document belongs (such as the identifier of the application that publishes or reposts the video), the identifier of the social public service account to which the document belongs (such as the identifier of the public account or service account where the video is published). The trigger time of the co - occurrence document refers to the moment when the co - occurrence document is clicked during the historical time period.
[0075] It should be noted that the historical search data can be sourced from all or some of the users who generated search operations during the historical time period. For example, the search application includes users A, B, C, D... and the historical time period is the target time period before the current moment; if during the target time period, user A performs a search operation on search field A, user B performs a search operation on search field B, and user D performs a search operation on search field A, and user C does not perform any search on any field during the target time period; then, the historical search data can come from one or more of users A, B, and D.
[0076] The process of obtaining historical search data may include: (1) Obtaining the first historical search field generated within a historical time period. The so-called first historical search field refers to the search terms or sentences of any registered user within the historical time period; for example, a user inputs the search field "Braised Pork" in the search engine within the historical time period, and at this time, "Braised Pork" is taken as a historical search field. (2) Querying multiple matching documents from the database that match the first historical search field. There is an association relationship between these matching documents and the first historical search field, and this association relationship can be reflected as semantic relevance, etc.; for example, when the matching degree between the semantic information of a certain document and the semantic information of the first historical search field is greater than a certain threshold, it indicates that the semantics of the document may be similar to the semantics of the first historical search field, and it is determined that the document can be used as a matching document. All matching documents can be displayed on the display screen for users to browse, click, etc. (3) Determining the N matching documents triggered (such as clicked) among the multiple matching documents as co-occurrence documents. In other words, when the first historical search field is searched by different users, multiple matching documents can be displayed on the display screens of each user, then the part of the documents triggered by the users among these matching documents is taken as co-occurrence documents. Here, the users can refer to multiple users who perform search operations according to the first historical search field and click on the matching documents. For example: The first historical search field is Braised Pork, and the multiple matching documents displayed on the display screen and matching the first historical search field are the first matching document, the second matching document, the third matching document, and the fourth matching document. When the first matching document and the third matching document are triggered (i.e., clicked) within the historical time period, the first matching document and the third matching document are taken as co-occurrence documents. (4) Obtaining the account data of the N co-occurrence documents and the trigger time of each co-occurrence document.
[0077] In summary, by performing steps (1)-(4) on multiple historical search fields within a historical time period, multiple historical search data can be obtained, which enriches the source of sample data and makes the subsequent generated sample set contain relatively rich sample data.
[0078] S202. Constructing a first sample sequence corresponding to the first historical search field according to the historical search data.
[0079] S203. Adding the first sample sequence to the sample set.
[0080] In steps S202 - S203, if the historical search data contains the trigger times and account data of N co - occurring documents, then arranging the account data of the N co - occurring documents of the first historical search field in the order of the trigger times, the first sample sequence corresponding to the first historical search field can be obtained. In other words, the first sample sequence is obtained by arranging one by one account data in the order of the trigger times. It can be understood that if M historical search fields (the first historical search field is any one of them) are generated within the historical time period, then arranging the account data of the co - occurring documents of each historical search field among the M historical search fields in the order of the trigger times, M sample sequences can be obtained. The M sample sequences form a sample set. Using the account data of the co - occurring documents of M historical search fields within the historical time period to construct the sample set increases the number of sample sequences included in the sample set.
[0081] The following combines the attached Figure 3 , and taking the account data as the identifier of the application to which the document belongs as an example, the construction of the first sample sequence will be introduced in detail. Please refer to Figure 3 , Figure 3 shows a schematic diagram of constructing a sample sequence provided by an exemplary embodiment of the present application; as Figure 3 shown, assume that the historical time period is from 00:00 to 24:00 on December 8th, the first historical search field is braised pork, and the 4 co - occurring documents matching the first historical search field are: the first co - occurring document, the second co - occurring document, the third co - occurring document, and the fourth co - occurring document; among them, the first co - occurring document belongs to application A, the account data of the first co - occurring document is account identifier A, and the trigger time of the first co - occurring document is 2:00 on December 8th, the second co - occurring document belongs to application B, the account data of the second co - occurring document is account identifier B, and the trigger time of the second co - occurring document is 8:00 on December 8th, the third co - occurring document belongs to application C, the account data of the third co - occurring document is account identifier C, and the trigger time of the third co - occurring document is 5:00 on December 8th, the fourth co - occurring document belongs to application D, the account data of the fourth co - occurring document is account identifier D, and the trigger time of the fourth co - occurring document is 9:00 on December 8th; then arranging the account data of the co - occurring documents in the order of the trigger times, the first sample sequence can be obtained: account identifier A -> account identifier C -> account identifier B -> account identifier D.
[0082] It should be noted that if the same co - occurring document is triggered multiple times within the historical time period, then the account data of this co - occurring document is also recorded multiple times. Refer to Figure 3Continue to illustrate: Assume that the first historical search field is XXX, and the triggering situations of the 4 co-occurring documents matching the first historical search field are as follows: The fifth co-occurring document (account identifier E) is triggered 2 times, and the triggering times are 2:00 and 4:00 on December 8th respectively; the second co-occurring document (account identifier F) is triggered 2 times, and the triggering times are 3:00 and 3:50 on December 8th respectively; the third co-occurring document (account identifier F) is triggered 1 time, and the triggering time is 5:00 on December 8th; the fourth co-occurring document (account identifier H) is triggered 1 time, and the triggering time is 9:00 on December 8th. Then, arranging the account data of the co-occurring documents in the order of the triggering time, the sample sequence can be obtained (i.e., Figure 3 the second sample sequence shown): Account identifier E -> Account identifier F -> Account identifier F -> Account identifier E -> Account identifier G -> Account identifier H.
[0083] S204. Use the sample set to train the word vector model to obtain the trained word vector model.
[0084] The trained word vector model can be used to calculate the word vectors (Word Embedding) of fields, and the distances between the word vectors of each field can be used to represent the similarities between each field; for example: if the distance between the word vectors of two fields is less than the distance threshold, then these two fields are relatively similar (such as similar grammatical structures, similar parts of speech, similar semantics, etc.), if the distance between the word vectors of two fields is greater than or equal to the distance threshold, then the similarity degree of these two fields is relatively low. Therefore, it is particularly important to construct a network model for calculating the word vectors of fields. Common word vector models may include, but are not limited to: word2vec model, GloVe model, ELMo model, etc. In the embodiments of this application, the word vector model is taken as the word2vec model as an example to introduce the process of training the word vector model with a sample set, which does not limit the embodiments of this application, and is hereby explained. Among them, the word2vec model is a model that learns semantic knowledge in an unsupervised manner from a large amount of text corpora. The learning methods of the word2vec model include the first prediction method and the second prediction method. The so-called first prediction method refers to the method of predicting the word vectors of the context (that is, all or part of the account data in the sample sequence to which the specific word belongs except the specific word) given a specific word (such as a certain account data in the sample sequence), and the so-called second prediction method refers to the method of predicting the word vector of a specific word given the context. The word2vec model can use the skip-gram method or the CBOW method to complete the modeling; the first prediction method can refer to the skip-gram method. Using the skip-gram method to complete the modeling, the obtained word vector model is called the skip-gram model; the second prediction method can refer to the CBOW method. Using the CBOW method to complete the modeling, the obtained word vector model can be called the CBOW model.
[0085] The following takes the skip-gram method and the CBOW method as examples to introduce the modeling process based on the first prediction method and the modeling process based on the second prediction method respectively, where:
[0086] I. The word vector model constructed by using the first prediction method (that is, the skip-gram model). The network structure diagram of the skip-gram model can be seen in Figure 4 , Figure 4 shows a schematic diagram of the network structure of a skip-gram model provided by an exemplary embodiment of this application; as Figure 4 shown, the process of training the word vector model based on this network structure may include steps s11-s13, where:
[0087] s11: Input the account data of the \(i\)-th co-occurring document in the first sample sequence into the word vector model (i.e., the skip-gram model) for prediction processing to obtain the word vectors of the account data of other co-occurring documents in the first sample sequence except for the account data of the \(i\)-th co-occurring document. Specifically, Figure 4 assuming that the window size is \(m\), and the window size is used to set the quantity and range of the context output by the word vector model. When the input of the word vector model is the account data of the \(i\)-th co-occurring document, the output result of the word vector model is: the word vectors of the account data of the context of the account data of the \(i\)-th co-occurring document; the context here may include: the account data within the window range of the account data belonging to the \(i\)-th co-occurring document. For example, if the window size \(m = 2\), the context includes: the account data of the co-occurring documents at positions 2 before and 2 after the account data of the \(i\)-th co-occurring document in the sorting position in the first sample sequence.
[0088] s12: Optimize the word vector model based on the difference between the predicted word vectors of the account data of other co-occurring documents and the reference word vectors of the account data of other co-occurring documents. It can be understood that in addition to the first sample sequence, the sample set for training the word vector model also includes the reference word vectors of the account data of each co-occurring document in the first sample sequence, and the reference word vector is a pre-given and relatively accurate word vector. By calculating the loss value (or difference value) between the predicted word vector of the account data and the reference word vector, the training quality of the current word vector model can be judged; when the loss value between the predicted word vector of the account data and the reference word vector is less than the loss threshold, it means that the word vector model has met the training requirements, and a relatively accurate word vector of the field can be obtained by using this word vector model; conversely, when the loss value between the predicted word vector of the account data and the reference word vector is greater than or equal to the loss threshold, it means that the word vector model has not met the training requirements, and the word vector of the field obtained by using this word vector model is not accurate enough. Then, the parameters in the word vector model need to be adjusted according to the loss value to optimize the word vector model until the word vector model meets the training requirements. It should be noted that the loss thresholds of different network models may be different, which will not be elaborated here.
[0089] s13: Use each sample sequence in the sample set to perform iterative optimization training on the word vector model, and obtain the trained word vector model. That is to say, all or part of the \(M\) sample sequences included in the sample set perform step s12 to achieve iterative optimization training of the word vector model, and obtain a word vector model with higher accuracy and better prediction quality.
[0090] II. The word vector model constructed by the second prediction method (i.e., the CBOW model). The network structure diagram of the CBOW model can be seen in Figure 5 ,Figure 5 The figure shows a schematic diagram of the network structure of a CBOW model provided by an exemplary embodiment of the present application; as Figure 5 shown, the process of training a word vector model based on this network structure may include steps S21 - S23, where:
[0091] S21: Input the account data of other co-occurring documents in the first sample sequence except for the account data of the i-th co-occurring document into the word vector model for prediction processing to obtain the word vector of the account data of the i-th co-occurring document. In combination Figure 5 with this, when the input of the word vector model is the account data of other co-occurring documents in the first sample sequence except for the account data of the i-th co-occurring document, the output result of the word vector model is: the word vector of the account data of the i-th co-occurring document. Among them, the account data of other co-occurring documents input except for the account data of the i-th co-occurring document is evenly distributed in the front and back positions of the account data of the i-th co-occurring document in the first sample sequence.
[0092] S22: Optimize the word vector model based on the difference between the predicted word vector of the account data of the i-th co-occurring document and the reference word vector of the account data of the i-th co-occurring document. It should be noted that the specific implementation process of this step S22 can refer to the relevant description of the specific implementation process shown in step S12 of the foregoing embodiment describing the skip-gram model, and will not be elaborated here.
[0093] S23: Use each sample sequence in the sample set to perform iterative optimization training on the word vector model to obtain the trained word vector model. It should be noted that the specific implementation process of this step S23 can refer to the relevant description of the specific implementation process shown in step S13 of the foregoing embodiment describing the skip-gram model, and will not be elaborated here.
[0094] In the embodiment of the present application, a sample set can be constructed based on M historical search fields generated within a historical time period and the account data of all co-occurring documents triggered by these M historical search fields; since these M historical search fields and the corresponding co-occurring documents are data generated when one or more users perform search operations within a historical time period, the sample set constructed based on these data has a relatively rich number of samples. Using this sample set to train the word vector model can obtain a trained word vector model with better performance.
[0095] Please refer to Figure 6 , Figure 6 which shows a schematic flowchart of another data processing method provided by an exemplary embodiment of the present application; this data processing solution can Figure 1 be executed through the interaction between the terminal and the server as shown, and this solution includes but is not limited to steps S601 - S603, where:
[0096] S601. Display the document search page of the application. The document search page includes a search box, search options, and a query result display area.
[0097] The document search page is used to implement the search for the target search field and display the search results. When the user opens and uses the application, the application displays the document search page, and the search box, search options, and query result display area are shown on the document search page. Among them, the search box can be used to receive the target search field written by the user; the search options serve as the search entry, and when the search options are triggered, it means that the search is triggered. The query result display area is used to display the search results that match the target search field. Please refer to Figure 7 , Figure 7 shows a schematic diagram of a document search page provided by an exemplary embodiment of the present application; as Figure 7 shown, the document search page 701 is displayed in the application, and the document search page 701 includes a search box 7011, search options 7012, and a query result display area 7013.
[0098] It should be noted that (1) the display forms of the document search pages of different applications may not be the same. The embodiments of the present application take the Figure 7 page shown in the appendix as an example to introduce the document search page of the application, and it does not limit the embodiments of the present application. (2) Since the display area of the query result display area is limited, some of the search results included in the query result display area may be hidden. In this case, the query result display area may include a scroll bar 7014. By operating the scroll bar 7014, the hidden search results can be scrolled and displayed. Of course, in addition to scrolling and displaying the content in the query result display area through the scroll bar 7014, it is also possible to scroll and display the search results in the query result display area by pressing any position of the query result display area. The embodiments of the present application do not limit this.
[0099] S602. When there is an input operation on the search box, display the input target search field in the search box.
[0100] S603. When the search options are selected to generate a search request carrying the target search field, sequentially display P second documents in the query result display area.
[0101] The implementation process described in steps S602 - S603 can be referred to Figure 8 the flow schematic diagram shown, please refer to Figure 8 , Figure 8 shows a flow schematic diagram of searching for the target search field provided by an exemplary embodiment of the present application; as Figure 8As shown, a document search page 701 is displayed in the application; when the user triggers the search box 7011 on the document search page 701, it indicates that the user wants to perform a search operation. At this time, a keyboard area 7015 is displayed on the document search page 701; the user can write the target search field in the search box 7011 by triggering the characters included in the keyboard area 7015; correspondingly, the target search field written by the user is displayed in the search box 7011. When the user selects the search option 7012, a search request is generated based on the target search field, and one or more first documents to be matched are obtained in response to the search request; and it is detected whether the first document to be matched is a document that matches the target search field. If so, the first document is determined as the target document, and the target document is a document that can be displayed on the display screen. At the same time, in response to the search request, P second documents (or target documents) that match the target search field are also sorted, and the P second documents are displayed in the query result display area 7013 in the sorted order.
[0102] The way of sequentially displaying the P second documents here may include: sorting in the order from the highest to the lowest matching degree; the matching degree refers to the matching degree between the word vectors of the account data of the second document and the word vectors of the target search field. Sorting the P second documents in this way enables the second documents with a higher matching degree with the target search field (i.e., search results) to be displayed in the front position, which helps the user quickly obtain the desired search results and meet the user's search needs.
[0103] Please refer to Figure 9 , Figure 9 shows a schematic diagram of the sequential display of a second document provided by an exemplary embodiment of the present application; as Figure 9As shown, the target search field is "XXX", and the second documents that match "XXX" include: Second Document A, Second Document B, Second Document C, and Second Document D. Among them, the account data of Second Document A is the identifier of the application "XXX Paradise", the account data of Second Document B is the identifier of the application "XXX World", the account data of Second Document C is the identifier of the application "Laughing XXX", and the account data of Second Document D is the identifier of the application "Food World". It can be seen that among the above 4 second documents, the matching degree between the account data "XXX Paradise" of Second Document A and the target search field "XXX" is higher than the matching degree between the account data "Anime XXX" of Second Document B and the target search field "XXX"; the matching degree between the account data "Anime XXX" of Second Document B and the target search field "XXX" is higher than the matching degree between the account data "Laughing XXX" of Second Document C and the target search field "XXX"; the matching degree between the account data "Laughing XXX" of Second Document C and the target search field "XXX" is higher than the matching degree between the account data "Food World" of Second Document D and the target search field "XXX". Among them, the matching degree between the account data and the target search field is reflected by the matching degree between the word vector of the account data "XXX Paradise" and the word vector of the target search field "XXX". In summary, the documents displayed from top to bottom in the query result display area 7013 are: Second Document A -> Second Document B -> Second Document C -> Second Document D.
[0104] As described above, the contents of both Second Document A and Second Document B are related to the target search field. However, the matching degree between the account data "XXX Paradise" of Second Document A and the target search field "XXX" is higher than the matching degree between the account data "XXX World" of Second Document B and the target search field "XXX", indicating that the quality of Second Document A published with the account data "XXX Paradise" is better than that of Second Document B published with the account data "XXX World". Introducing the matching degree between the word vector of the account data of the document and the word vector of the target search field to sort the documents increases the matching dimension compared to only considering the relationship between the document itself and the target search field, and can obtain search results that are more matched to the target search field, meeting the search needs of users.
[0105] Next, Figure 8 the two steps of detecting whether the first document to be matched is a document that matches the target search field and sorting the P second documents that match the target search field will be described in detail.
[0106] (1) The process of detecting whether the first document to be matched is a document that matches the target search field includes: calling the trained word vector model to process the target search field to obtain the word vector of the target search field; at the same time, extracting the account data of the first document to be matched from the database (such as the identifier of the video to be matched, etc.), and calling the trained word vector model to process the account data of the first document to obtain the word vector corresponding to the account data of the first document; calculating the matching degree between the word vector of the target search field and the word vector corresponding to the account data of the first document. If the matching degree between the word vector of the target search field and the word vector corresponding to the account data of the first document is greater than the threshold, it means that the distance between the word vector of the target search field and the word vector of the account data of the first document is less than the distance threshold, then the first document is determined as the target document that matches the target search field; otherwise, if the matching degree between the word vector of the target search field and the word vector corresponding to the account data of the first document is less than or equal to the threshold, it means that the distance between the word vector of the target search field and the word vector of the account data of the first document is greater than or equal to the distance threshold, then it is determined that the first document does not match the target search field.
[0107] For example, assume that the threshold (i.e., the matching degree threshold) is 60% (or in the form of 0.6, etc.). The first documents to be matched obtained include: the first document A, the first document B, the first document C, and the first document D. After calculation, it is obtained that: the matching degree between the word vector of the account data of the first document A and the word vector of the target search field is 30%, the matching degree between the word vector of the account data of the second document B and the word vector of the target search field is 70%, the matching degree between the word vector of the account data of the third document C and the word vector of the target search field is 65%, and the matching degree between the word vector of the account data of the fourth document A and the word vector of the target search field is 78%. It can be seen that 78% > 70% > 65% > 30%, then the first document B, the first document C, and the first document D are determined as the target documents that match the target search field.
[0108] (2) The process of sorting P second documents that match the target search field includes: searching in the database for P second documents that match the target search field. These P second documents can be determined in the manner described in step (1), or can be determined by other means. The embodiments of the present application do not limit this; calling the trained word vector model to process the target search field to obtain the word vector of the target search field; and, calling the trained word vector model to process the account data of the P second documents to obtain the word vector of the account data of each second document; respectively calculating the matching degree between the account identifier of each second document and the word vector of the target search field, and sorting the P second documents in descending order of the matching degree.
[0109] For example, assume that the second documents matching the target search field include: Second Document A, Second Document B, Second Document C, and Second Document D. After calculation, the matching degree between the word vector of the account data of Second Document A and the word vector of the target search field is 73%, the matching degree between the word vector of the account data of Second Document B and the word vector of the target search field is 69%, the matching degree between the word vector of the account data of Second Document C and the word vector of the target search field is 86%, and the matching degree between the word vector of the account data of Second Document D and the word vector of the target search field is 79%. It can be seen that 86% > 79% > 73% > 69%. Then, according to the order of the matching degree from high to low, the sequence of Second Document A, Second Document B, Second Document C, and Second Document D is: Second Document C -> Second Document D -> Second Document A -> Second Document B. This sequence can be displayed on the interface as: Second Document C is displayed at the top, and then Second Document D, Second Document A, and Second Document B are displayed in sequence after Second Document C. It can be understood that if the P second documents are obtained in the manner described in step (1), the matching degrees between the word vectors of the account data of each second document and the target search field have been obtained. Then, directly sort the P second documents according to the matching degree.
[0110] It should be noted that steps (1)-(2) can be executed by the terminal or the server or the interaction between the terminal and the server; for example: the terminal sends a search request to the server, and the server executes the above steps; another example: the terminal obtains the first document and P second documents from the server, and the terminal executes the above steps. Steps (1)-(2) can also be executed by the terminal; for example: the memory space of the terminal stores the first document and P second documents, and the terminal executes the above steps. The embodiments of the present application do not limit this.
[0111] In addition, the embodiments of the present application also support jumping from the document search page to the service interface corresponding to the account data of the second document. In this way, other content associated with the account data can be browsed in the service interface corresponding to the account data, which helps users quickly obtain the content they are interested in. For example, when the account data of the second document includes any of the following identifiers, triggering any of the account data of the second documents displayed in the query result display area will jump to the service interface corresponding to the triggered account data of the second document. The content included in the service interface and the display form of each part of the content are displayed according to the type of the account data. For example: when the triggered account data of the second document is the identifier of the application program to which the second document belongs, the service interface includes the service interface of the application program to which the triggered second document belongs; another example: when the triggered account data of the second document is the identifier of the social public service account to which the second document belongs, the service interface includes the service interface of the social public service account to which the triggered second document belongs; still another example: when the triggered account data of the second document is the identifier of the publisher who published the second document, the service interface includes the service interface of the publisher who published the triggered second document.
[0112] The following will introduce the above-described page jump in conjunction with the Figure 10 accompanying drawings. Please refer to Figure 10 , Figure 10 which shows a schematic diagram of a service interface corresponding to account data provided by an exemplary embodiment of the present application; as Figure 10 shown, in the document search page 701 of the application program, multiple second documents and relevant information of each second document (such as account data, cover image, etc.) are displayed; when the account data 1001 of the second document is triggered, it jumps from the document search page 701 to the service interface 1002 corresponding to the account data 1001. In this example, taking the identifier of the publisher who published the second document as the account data as an example, the service interface 1002 corresponding to the account data 1001 is given, which does not limit the embodiments of the present application; in the service interface 1002, the identifier of the publisher and historical publishing information are displayed. Through the above process, it is convenient for users to browse other document content published by the publisher and improves the promotion efficiency of the documents.
[0113] In the embodiments of the present application, a word vector model generated according to the Figure 2 embodiment shown is used to search the target search field, and multiple second documents with a relatively high matching degree with the target search field can be obtained; moreover, each second document is displayed in the order of the matching degree between the word vector of the account data of the second document and the word vector of the target search field, and the documents with a high matching degree among the multiple second documents can be preferentially displayed at the front position in the document search page, which helps users quickly obtain the desired search results and improves the user's search experience.
[0114] Figure 11 The structure diagram of a data processing device provided by an exemplary embodiment of the present application is shown; the data processing device can be a computer program (including program code) running in the terminal 102. For example, the data processing device can be a search application in the terminal 102; the data processing device can be used to execute Figure 2 、 Figure 6 Some or all of the steps in the method embodiments shown. Please refer to Figure 11 The data processing device includes the following units:
[0115] An acquisition unit 1101, configured to acquire historical search data, where the historical search data includes a first historical search field generated within a historical time period, account data of N co-occurring documents triggered based on the first historical search field, and the trigger time of each co-occurring document, where N is a positive integer; a co-occurring document refers to a search result triggered among one or more search results matching the first historical search field;
[0116] A processing unit 1102, configured to construct a first sample sequence corresponding to the first historical search field according to the historical search data, where the first sample sequence includes the account data of N co-occurring documents arranged in chronological order of trigger time; add the first sample sequence to a sample set, where the sample set includes M sample sequences, M historical search fields are generated within the historical time period, and the first historical search field is any one of the M historical search fields; a sample in the sample set corresponds to a historical search field generated within the historical time period, and M is a positive integer; and, train a word vector model using the sample set to obtain a trained word vector model.
[0117] In an implementation manner, the processing unit 1102 is further configured to:
[0118] In response to a search request carrying a target search field, call the trained word vector model to process the target search field to obtain a word vector of the target search field;
[0119] Extract the account data of the first document to be matched from the database, and call the trained word vector model to process the account data of the first document to obtain a word vector corresponding to the account data of the first document;
[0120] Calculate the matching degree between the word vector of the target search field and the word vector corresponding to the account data of the first document;
[0121] If the matching degree is higher than a threshold, determine the first document as a target document matching the target search field.
[0122] In an implementation manner, the processing unit 1102 is further configured to:
[0123] In response to a search request carrying a target search field, P second documents that match the target search field are found from a database, where P is a positive integer;
[0124] The trained word vector model is called to process the target search field to obtain the word vector of the target search field; and, the trained word vector model is called to process the account data of the P second documents to obtain the word vectors of the account data of each second document;
[0125] The matching degrees between the word vectors of the account data of each second document and the word vector of the target search field are calculated respectively;
[0126] The P second documents are sorted in descending order of the matching degree.
[0127] In one implementation, the processing unit 1102 is further configured to:
[0128] Display the document search page of the application, where the document search page includes a search box, search options, and a query result display area;
[0129] When there is an input operation on the search box, the input target search field is displayed in the search box;
[0130] When a search request carrying the target search field is generated by selecting the search options, the P second documents are sequentially displayed in the query result display area.
[0131] In one implementation, if the account data includes the identifier of the application to which the document belongs and the identifier of the social public service account to which the document belongs; then the processing unit 1102 is further configured to:
[0132] When the account data of any second document displayed in the query result display area is triggered, jump to the service interface corresponding to the account data of the triggered second document; the service interface includes the service interface of the application to which the triggered second document belongs, or the service interface of the social public service account to which the triggered second document belongs.
[0133] In one implementation, the processing unit 1102 is specifically configured to:
[0134] Obtain the first historical search field generated within a historical time period;
[0135] Query multiple matching documents that match the first historical search field from the database;
[0136] Determine the N triggered matching documents among the multiple matching documents as co-occurrence documents; and,
[0137] Obtain the account data of N co-occurring documents and the triggering time of each co-occurring document.
[0138] In one embodiment, the word vector model is a word vector model constructed using a first prediction method; the sample set further includes the reference word vectors of the account data of each co-occurring document in the first sample sequence; the processing unit 1102 is specifically configured to:
[0139] Input the account data of the i-th co-occurring document in the first sample sequence into the word vector model for prediction processing to obtain the word vectors of the account data of other co-occurring documents in the first sample sequence except the account data of the i-th co-occurring document;
[0140] Optimize the word vector model based on the difference between the predicted word vectors of the account data of other co-occurring documents and the reference word vectors of the account data of other co-occurring documents;
[0141] Iteratively optimize and train the word vector model using each sample sequence in the sample set to obtain the trained word vector model.
[0142] In one embodiment, the word vector model is a word vector model constructed using a second prediction method; the sample set further includes the reference word vectors of the account data of each co-occurring document in the first sample sequence; the processing unit 1102 is specifically configured to:
[0143] Input the account data of other co-occurring documents in the first sample sequence except the account data of the i-th co-occurring document into the word vector model for prediction processing to obtain the word vector of the account data of the i-th co-occurring document;
[0144] Optimize the word vector model based on the difference between the predicted word vector of the account data of the i-th co-occurring document and the reference word vector of the account data of the i-th co-occurring document;
[0145] Iteratively optimize and train the word vector model using each sample sequence in the sample set to obtain the trained word vector model.
[0146] In one embodiment, the account data includes any one of the following: the identifier of the document, the identifier of the publisher of the document, the identifier of the application to which the document belongs, the identifier of the social public service account to which the document belongs; the co-occurring documents include any one of the following: text, picture, video, or audio.
[0147] According to an embodiment of the present application, Figure 11Each unit in the data processing device shown can be separately or entirely combined into one or several other units to form, or a certain one (or some) of the units can be further split into multiple smaller units with more specific functions to form. This can achieve the same operations without affecting the realization of the technical effects of the embodiments of this application. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of this application, the data processing device can also include other units. In practical applications, these functions can also be assisted by other units and can be realized through the cooperation of multiple units. According to another embodiment of this application, it can be achieved by running a computer program (including program code) that can execute each step involved in the corresponding method shown in Figure 2 、 Figure 6 on a general computing device such as a computer that includes processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), to construct a data processing device as shown in Figure 11 and to implement the data processing method of the embodiments of this application. The computer program can be recorded on, for example, a computer-readable recording medium, loaded into the above computing device through the computer-readable recording medium, and run therein.
[0148] In the embodiments of this application, after the acquisition unit 1101 acquires the historical search data, it can send the historical search data to the processing unit 1102; correspondingly, the processing unit 1102 constructs a first sample sequence corresponding to the first historical search field according to the historical search data, and the processing unit 1102 also adds the first sample sequence to the sample set. The sample set includes M sample sequences, M historical search fields are generated within the historical time period, the first historical search field is any one of the M historical search fields, and one sample sequence in the sample set corresponds to one historical search field generated within the historical time period. In the above process, since the M historical search fields and the corresponding co-occurrence documents are data generated by one or more users performing search operations within the historical time period, the sample set constructed based on these data has a relatively rich number of samples. Using this sample set to train the word vector model can obtain a trained word vector model with better performance.
[0149] Figure 12 shows a schematic structural diagram of a data processing device provided by an exemplary embodiment of this application. Please refer to Figure 12, the data processing device includes a processor 1201, a communication interface 1202, and a computer-readable storage medium 1203. Among them, the processor 1201, the communication interface 1202, and the computer-readable storage medium 1203 can be connected through a bus or other means. Among them, the communication interface 1202 is used to receive and send data. The computer-readable storage medium 1203 can be stored in the memory of the data processing device. The computer-readable storage medium 1203 is used to store a computer program, and the computer program includes program instructions. The processor 1201 is used to execute the program instructions stored in the computer-readable storage medium 1203. The processor 1201 (or CPU (Central Processing Unit, central processor)) is the computing core and control core of the data processing device, and is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function.
[0150] The embodiment of the present application also provides a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in the data processing device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the data processing device and, of course, the extended storage medium supported by the data processing device. The computer-readable storage medium provides a storage space, and this storage space stores the processing system of the data processing device. And, one or more instructions suitable for being loaded and executed by the processor 1201 are also stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located far from the aforementioned processor.
[0151] In one embodiment, the data processing device can be Figure 1 the terminal 102 or the server 101 shown in the figure. The data processing device can also be any other device with a model training function other than the terminal 102 and the server 101; one or more instructions are stored in the computer-readable storage medium; the processor 1201 loads and executes one or more instructions stored in the computer-readable storage medium to implement the corresponding steps in the above data processing method embodiment; in a specific implementation, one or more instructions in the computer-readable storage medium are loaded and executed by the processor 1201 as follows:
[0152] Obtain historical search data, where the historical search data includes the first historical search field generated within a historical time period, the account data of N co-occurring documents triggered based on the first historical search field, and the triggering time of each co-occurring document, where N is a positive integer; a co-occurring document refers to a search result triggered among one or more search results that match the first historical search field;
[0153] Construct a first sample sequence corresponding to the first historical search field based on the historical search data, where the first sample sequence includes the account data of N co-occurring documents arranged in chronological order of the triggering time;
[0154] Add the first sample sequence to a sample set, where the sample set includes M sample sequences, M historical search fields are generated within the historical time period, and the first historical search field is any one of the M historical search fields; one sample sequence in the sample set corresponds to one historical search field generated within the historical time period, and M is a positive integer;
[0155] Train a word vector model using the sample set to obtain a trained word vector model.
[0156] In one implementation, one or more instructions in a computer-readable storage medium are loaded and further executed by a processor 1201 as follows:
[0157] In response to a search request carrying a target search field, call the trained word vector model to process the target search field to obtain a word vector of the target search field;
[0158] Extract the account data of a first document to be matched from a database, and call the trained word vector model to process the account data of the first document to obtain a word vector corresponding to the account data of the first document;
[0159] Calculate the matching degree between the word vector of the target search field and the word vector corresponding to the account data of the first document;
[0160] If the matching degree is higher than a threshold, determine the first document as a target document that matches the target search field.
[0161] In one implementation, one or more instructions in a computer-readable storage medium are loaded and further executed by a processor 1201 as follows:
[0162] In response to a search request carrying a target search field, search for P second documents that match the target search field in a database, where P is a positive integer;
[0163] Call the trained word vector model to process the target search field to obtain the word vector of the target search field; and call the trained word vector model to process the account data of the P second documents to obtain the word vectors of the account data of each second document.
[0164] Calculate the matching degree between the word vector of the account data of each second document and the word vector of the target search field respectively.
[0165] Sort the P second documents in descending order of the matching degree.
[0166] In one implementation, one or more instructions in the computer-readable storage medium are loaded and executed by the processor 1201, and the following steps are further performed:
[0167] Display the document search page of the application. The document search page includes a search box, search options, and a query result display area.
[0168] When there is an input operation on the search box, display the input target search field in the search box.
[0169] When the search option is selected to generate a search request carrying the target search field, sequentially display the P second documents in the query result display area.
[0170] In one implementation, if the account data includes the identifier of the application to which the document belongs and the identifier of the social public service account to which the document belongs; then one or more instructions in the computer-readable storage medium are loaded and executed by the processor 1201, and the following steps are further performed:
[0171] When the account data of any second document displayed in the query result display area is triggered, jump to the service interface corresponding to the account data of the triggered second document; the service interface includes the service interface of the application to which the triggered second document belongs, or the service interface of the social public service account to which the triggered second document belongs.
[0172] In one implementation, one or more instructions in the computer-readable storage medium are loaded and executed by the processor 1201 when obtaining historical search data, and the following steps are specifically performed:
[0173] Obtain the first historical search field generated within the historical time period.
[0174] Query multiple matching documents that match the first historical search field from the database.
[0175] Determine the N triggered matching documents among the multiple matching documents as co-occurrence documents; and
[0176] Obtain the account data of N co-occurring documents and the triggering time of each co-occurring document.
[0177] In one implementation, the word vector model is a word vector model constructed using the first prediction method; the sample set further includes the reference word vectors of the account data of each co-occurring document in the first sample sequence.
[0178] When one or more instructions in the computer-readable storage medium are loaded by the processor 1201 and the word vector model is trained using the sample set to obtain the trained word vector model, the following steps are specifically executed:
[0179] Input the account data of the i-th co-occurring document in the first sample sequence into the word vector model for prediction processing to obtain the word vectors of the account data of other co-occurring documents in the first sample sequence except the account data of the i-th co-occurring document.
[0180] Optimize the word vector model based on the difference between the predicted word vectors of the account data of other co-occurring documents and the reference word vectors of the account data of the other co-occurring documents.
[0181] Iteratively optimize and train the word vector model using each sample sequence in the sample set to obtain the trained word vector model.
[0182] In one implementation, the word vector model is a word vector model constructed using the second prediction method; the sample set further includes the reference word vectors of the account data of each co-occurring document in the first sample sequence.
[0183] When one or more instructions in the computer-readable storage medium are loaded by the processor 1201 and the word vector model is trained using the sample set to obtain the trained word vector model, the following steps are specifically executed:
[0184] Input the account data of other co-occurring documents in the first sample sequence except the account data of the i-th co-occurring document into the word vector model for prediction processing to obtain the word vector of the account data of the i-th co-occurring document.
[0185] Optimize the word vector model based on the difference between the predicted word vector of the account data of the i-th co-occurring document and the reference word vector of the account data of the i-th co-occurring document.
[0186] Iteratively optimize and train the word vector model using each sample sequence in the sample set to obtain the trained word vector model.
[0187] In one embodiment, the account data includes any one of the following: the identifier of the document, the identifier of the publisher of the document, the identifier of the application to which the document belongs, the identifier of the social public service account to which the document belongs; the co-occurring documents include any one of the following: text, picture, video or audio.
[0188] In the embodiment of the present application, when the processor 1201 obtains the historical search data, it can construct a first sample sequence corresponding to the first historical search field according to the historical search data, and add the first sample sequence to the sample set. The sample set includes M sample sequences. M historical search fields are generated within the historical time period. The first historical search field is any one of the M historical search fields. One sample sequence in the sample set corresponds to one historical search field generated within the historical time period. In the above process, since the M historical search fields and the corresponding co-occurring documents are data generated when one or more users perform search operations within the historical time period, the sample set constructed based on these data has a relatively rich number of samples. Using this sample set to train the word vector model can obtain a trained word vector model with better performance.
[0189] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but this implementation should not be considered to exceed the scope of the present application.
[0190] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)), etc.
[0191] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A data processing method, characterized in that, Including: Obtain historical search data, where the historical search data includes a first historical search field generated within a historical time period, account data of N co-occurring documents triggered based on the first historical search field, and the trigger time of each co-occurring document, N being a positive integer; the first historical search field has been searched by multiple users within the historical time period, and the co-occurring documents refer to the searched results triggered among one or more search results matching the first historical search field; the N co-occurring documents include: the searched results triggered by each user among one or more search results matching the first historical search field; Construct a first sample sequence corresponding to the first historical search field according to the historical search data, where the first sample sequence includes the account data of the N co-occurring documents arranged in the order of trigger time; Add the first sample sequence to a sample set, where the sample set includes M sample sequences, M historical search fields are generated within the historical time period, and the first historical search field is any one of the M historical search fields; one sample sequence in the sample set corresponds to one historical search field generated within the historical time period, M being a positive integer; Use the sample set to train a word vector model to obtain a trained word vector model.
2. The method according to claim 1, wherein The method further includes: In response to a search request carrying a target search field, call the trained word vector model to process the target search field to obtain the word vector of the target search field; Extract the account data of a first document to be matched from a database, and call the trained word vector model to process the account data of the first document to obtain the word vector corresponding to the account data of the first document; Calculate the matching degree between the word vector of the target search field and the word vector corresponding to the account data of the first document; If the matching degree is higher than a threshold, determine the first document as a target document matching the target search field.
3. The method according to claim 1, characterized in that The method further includes: In response to a search request carrying a target search field, search for P second documents matching the target search field in a database, P being a positive integer; Call the trained word vector model to process the target search field to obtain the word vector of the target search field; and call the trained word vector model to process the account data of the P second documents to obtain the word vectors of the account data of each second document; Calculate the matching degree between the word vector of the account data of each second document and the word vector of the target search field respectively; Sort the P second documents in descending order of the matching degree.
4. The method according to claim 3, characterized in that, The method further includes: Display a document search page of an application, where the document search page includes a search box, search options, and a query result display area; When there is an input operation on the search box, display the input target search field in the search box; When the search option is selected to generate a search request carrying the target search field, the P second documents are sequentially displayed in the query result display area.
5. The method according to claim 4, wherein If the account data includes the identifier of the application to which the document belongs and the identifier of the social public service account to which the document belongs; then the method further includes: When the account data of any second document displayed in the query result display area is triggered, jump to the service interface corresponding to the account data of the triggered second document; the service interface includes the service interface of the application to which the triggered second document belongs, or the service interface of the social public service account to which the triggered second document belongs.
6. The method according to claim 1, wherein The obtaining of the historical search data includes: Obtaining the first historical search field generated within the historical time period; Querying from the database a plurality of matching documents that match the first historical search field; Determining the N matching documents triggered among the plurality of matching documents as the co-occurrence documents; and, Obtaining the account data of the N co-occurrence documents and the trigger time of each co-occurrence document.
7. The method according to claim 1, wherein The word vector model is a word vector model constructed by the first prediction method; the sample set further includes the reference word vectors of the account data of each co-occurrence document in the first sample sequence; The training of the word vector model using the sample set to obtain the trained word vector model includes: Inputting the account data of the i-th co-occurrence document in the first sample sequence into the word vector model for prediction processing to obtain the word vectors of the account data of the other co-occurrence documents in the first sample sequence except the account data of the i-th co-occurrence document; Optimizing the word vector model based on the difference between the predicted word vectors of the account data of the other co-occurrence documents and the reference word vectors of the account data of the other co-occurrence documents; Performing iterative optimization training on the word vector model using each sample sequence in the sample set to obtain the trained word vector model.
8. The method according to claim 1, wherein The word vector model is a word vector model constructed by the second prediction method; the sample set further includes the reference word vectors of the account data of each co-occurrence document in the first sample sequence; The training of the word vector model using the sample set to obtain the trained word vector model includes: Inputting the account data of the other co-occurrence documents in the first sample sequence except the account data of the i-th co-occurrence document into the word vector model for prediction processing to obtain the word vector of the account data of the i-th co-occurrence document; Optimizing the word vector model based on the difference between the predicted word vector of the account data of the i-th co-occurrence document and the reference word vector of the account data of the i-th co-occurrence document; Performing iterative optimization training on the word vector model using each sample sequence in the sample set to obtain the trained word vector model.
9. The method according to claim 1, characterized in that, The account data includes any one of the following: the identifier of the document, the identifier of the publisher of the document, the identifier of the application to which the document belongs, the identifier of the social public service account to which the document belongs; the co-occurrence document includes any one of the following: text, picture, video or audio.
10. A data processing device, characterized in that, Including: An acquisition unit, configured to acquire historical search data, where the historical search data includes a first historical search field generated within a historical time period, account data of N co-occurring documents triggered based on the first historical search field, and the triggering time of each co-occurring document, and N is a positive integer; the first historical search field has been searched by multiple users within the historical time period, and the co-occurring document refers to a search result triggered in one or more search results matching the first historical search field; the N co-occurring documents include: search results triggered by each user in one or more search results matching the first historical search field; A processing unit, configured to construct a first sample sequence corresponding to the first historical search field according to the historical search data, where the first sample sequence includes the account data of the N co-occurring documents arranged in chronological order of the triggering time; add the first sample sequence to a sample set, where the sample set includes M sample sequences, and M historical search fields are generated within the historical time period, and the first historical search field is any one of the M historical search fields; one sample sequence in the sample set corresponds to one historical search field generated within the historical time period, and M is a positive integer; and, use the sample set to train a word vector model to obtain a trained word vector model.
11. A data processing device, characterized in that, Including: A processor, adapted to implement one or more instructions; and, A computer-readable storage medium storing one or more instructions, where the one or more instructions are adapted to be loaded and executed by the processor to perform the data processing method according to any one of claims 1-9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more instructions, where the one or more instructions are adapted to be loaded and executed by a processor to perform the data processing method according to any one of claims 1-9.
Citation Information
Patent Citations
Similarity mining method and device
CN107193832A
Dynamic predictive similarity grouping based on vectorization of merchant data
US20190295124A1