A data processing method, apparatus, computer device, and storage medium
Through the multi-level feature extraction and splicing processing of the multimodal matching model, the problem of low accuracy of prediction results in the existing text matching methods is solved, and a higher degree of matching prediction accuracy is achieved.
Patent Information
- Application Number
- CN202011261127.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-12
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-11-12
AI Technical Summary
When the existing text matching method performs single text matching, it leads to large errors in the target matching data, reducing the accuracy of the prediction results.
Using a multimodal matching model, through the text feature learner and multimodal feature learner in the feature learner, multi-level feature extraction and splicing processing are performed on the search service data and the service data to be matched, and vector splicing results are generated to indicate the matching degree.
The accuracy of the matching prediction between the search service data and the service data to be matched is improved, errors are reduced, and the accuracy of the prediction results is improved.
Smart Images

Figure CN112231347B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a data processing method, apparatus, computer device, and storage medium. Background Art
[0002] Currently, in the business search scenario, a user can enter business data of interest (e.g., text data a) in an application client. At this time, a computer device often searches for target matching data (e.g., business data b) with a high text matching degree with the text data a for the user through text matching. It can be understood that in the process of matching the text data a with the business data b, the computer device needs to determine the feature vector 1 of the text data a and the feature vector 2 of the text (e.g., title text) in the business data b, and then can determine the similarity between the text data a and the business data b by determining the similarity distance between the feature vector 1 and the feature vector 2.
[0003] It can be seen that the existing text matching method needs to extract the feature vector of the title text from the business data b and directly use the feature vector of the title text as the feature vector for representing the entire business data b. Therefore, in the process of single text matching, there will be a large error in the finally searched target matching data, thus reducing the accuracy of the prediction result. Summary of the Invention
[0004] Embodiments of this application provide a data processing method, apparatus, computer device, and storage medium, which can improve the accuracy of the prediction result.
[0005] On the one hand, an embodiment of this application provides a data processing method, including:
[0006] Obtain a multi-modal matching model for matching search business data and to-be-matched business data; the multi-modal matching model includes a feature learner and a prediction generator; the to-be-matched business data includes first-modal business data and second-modal business data;
[0007] Perform a first learning process on the first feature extraction vector of the search business data and the second feature extraction vector of the first-modal business data through the text feature learner in the feature learner to obtain a first learning result; the learning vector in the first learning result is obtained from a text global information vector and a text local fine-grained vector; the text global information vector is obtained based on the first multi-scale convolution kernel in the first global feature learning layer of the text feature learner; the text local fine-grained vector is obtained based on the first local feature learning layer of the text feature learner;
[0008] Through the multi-modal feature learner in the feature learner, perform a second learning process on the first feature extraction vector and the third feature extraction vector of the second-modal service data to obtain a second learning result; the learning vector in the second learning result is obtained from the multi-modal global information vector and the multi-modal local fine-grained vector; the multi-modal global information vector is obtained based on the second multi-scale convolution kernel in the second global feature learning layer of the multi-modal feature learner; the multi-modal local fine-grained vector is obtained based on the second local feature learning layer of the multi-modal feature learner.
[0009] Through the prediction generator, splice the learning vector in the first learning result and the learning vector in the second learning result to obtain a vector splicing result; the vector splicing result is used to indicate the prediction of the matching degree between the search service data and the data to be matched service data.
[0010] One aspect of the embodiments of the present application provides a data processing method, including:
[0011] Obtain a sample data group for training a multi-modal training model; the sample data group includes a first type of sample data group and a second type of sample data group; the first type of sample data group is a sample data group with sample label information; the second type of sample data group is a sample data group without sample label information; the sample label information is used to indicate the matching degree between the first type of sample data groups.
[0012] Input the sample data group into the multi-modal training model, and the multi-modal training model outputs the prediction result between the sample data groups, and use the prediction result as the prediction label information; the multi-modal training model includes a sample feature extractor, a sample feature learner, and a sample prediction generator.
[0013] Obtain the sample splicing vector corresponding to the sample data group, and determine the optimal perturbation amount of the sample data group based on the sample splicing vector, the model loss function of the multi-modal training model, and the expected conditions associated with the multi-modal training model.
[0014] Generate adversarial sample data corresponding to the sample data group based on the optimal perturbation amount and the sample splicing vector, and perform iterative training on the multi-modal training model based on the adversarial sample data and the model loss function to obtain a model training result.
[0015] When the model training result indicates that the multi-modal training model after iterative training meets the model convergence condition, use the multi-modal training model that meets the model convergence condition as the multi-modal matching model for predicting the matching degree between business data groups.
[0016] One aspect of the embodiments of the present application provides a data processing device, including:
[0017] A model acquisition module for acquiring a multi-modal matching model used to match search service data and service data to be matched; the multi-modal matching model includes a feature learner and a prediction generator; the service data to be matched includes first-modal service data and second-modal service data;
[0018] A first learning and processing module for performing first learning and processing on a first feature extraction vector of search service data and a second feature extraction vector of first-modal service data through a text feature learner in the feature learner to obtain a first learning result; the learning vector in the first learning result is obtained from a text global information vector and a text local fine-grained vector; the text global information vector is obtained based on a first multi-scale convolution kernel in a first global feature learning layer of the text feature learner; the text local fine-grained vector is obtained based on a first local feature learning layer of the text feature learner;
[0019] A second learning and processing module for performing second learning and processing on the first feature extraction vector and a third feature extraction vector of second-modal service data through a multi-modal feature learner in the feature learner to obtain a second learning result; the learning vector in the second learning result is obtained from a multi-modal global information vector and a multi-modal local fine-grained vector; the multi-modal global information vector is obtained based on a second multi-scale convolution kernel in a second global feature learning layer of the multi-modal feature learner; the multi-modal local fine-grained vector is obtained based on a second local feature learning layer of the multi-modal feature learner;
[0020] A splicing processing module for splicing the learning vector in the first learning result and the learning vector in the second learning result through the prediction generator to obtain a vector splicing result; the vector splicing result is used to indicate the prediction of the matching degree between the search service data and the service data to be matched.
[0021] Wherein, the device further includes:
[0022] A request acquisition module for acquiring a service search request including search service data sent by a user terminal; the service search request is generated when the user terminal responds to a trigger operation on a search control in an application client; the search service data is obtained by the user terminal from a search area of a search display interface;
[0023] A data acquisition module for, based on the service search request, acquiring service data of a first service type from a video database, using the service data of the first service type as the first-modal service data, and acquiring service data of a second service type from the video database, using the service data of the second service type as the second-modal service data; the first service type is different from the second service type;
[0024] The business data to be matched determination module is used to use the business data jointly mapped by the first-modal business data and the second-modal business data as the business data to be matched.
[0025] Among them, if the business type of the search business data is the first business type and the first business type belongs to the text type, then the second business type includes at least one of the following business types: video type or picture type; the multimodal matching model includes a feature extractor; the feature extractor includes a word vector extraction network and a residual network;
[0026] The device further includes:
[0027] The text data to be encoded determination module is used to use the search business data and the first-modal business data as the text data to be encoded;
[0028] The vector extraction module is used to extract a feature extraction vector from the text data to be encoded through the word vector extraction network; the feature extraction vector includes a first feature extraction vector extracted from the search business data and a second feature extraction vector extracted from the first-modal business data;
[0029] The frame extraction processing module is used to perform frame extraction processing on the second-modal business data to obtain video frames, input the video frames into the residual network, and the residual network extracts a third feature extraction vector corresponding to the second-modal business data.
[0030] Among them, the vector extraction module includes:
[0031] The preprocessing unit is used to preprocess the text data to be encoded, use the preprocessed text data to be encoded as the text data to be matched, perform character segmentation processing on the text data to be matched according to the text vocabulary, and obtain a word information sequence and a word position sequence corresponding to the text data to be matched; the total number of words in the text data to be matched is H; H is a positive integer;
[0032] The target word acquisition unit is used to traverse and acquire the word information corresponding to the k-th word of the text data to be matched from the word information sequence, use the acquired word information as the target word information, acquire the word position information corresponding to the target word information from the word position sequence, and use the acquired word position information as the target word position information; k is a positive integer less than or equal to H;
[0033] The vector extraction unit is used to input the target word information into the word vector extraction network, and the word vector extraction network extracts a target word information vector corresponding to the k-th word, input the target word position information into the word vector extraction network, and the word vector extraction network extracts a target word position vector corresponding to the k-th word; the word vector extraction network is trained based on the text vocabulary;
[0034] A feature extraction vector determination unit, configured to obtain a feature extraction vector corresponding to the k-th word based on the target word information vector and the target word position vector, and until the value of k is H, obtain the feature extraction vector corresponding to the text data to be matched.
[0035] Wherein, the feature learner includes a first multi-layer perceptron associated with the text feature learner; the text feature learner includes a first bidirectional hidden encoding layer, a first global feature learning layer, and a first local feature learning layer;
[0036] The first learning processing module includes:
[0037] A text initial vector determination unit, configured to input the first feature extraction vector of the search service data and the second feature extraction vector of the first-modal service data into the first bidirectional hidden encoding layer respectively, to obtain a first initial hidden vector corresponding to the first feature extraction vector and a second initial hidden vector corresponding to the second feature extraction vector;
[0038] A text global vector determination unit, configured to obtain a first global information vector corresponding to the first feature extraction vector and a second global information vector corresponding to the second feature extraction vector based on the first initial hidden vector, the second initial hidden vector, and the first global feature learning layer, and use the first global information vector and the second global information vector as the text global information vector;
[0039] A text local vector determination unit, configured to obtain a first local fine-grained vector corresponding to the first feature extraction vector and a second local fine-grained vector corresponding to the second feature extraction vector based on the first initial hidden vector, the second initial hidden vector, and the first local feature learning layer, and use the first local fine-grained vector and the second local fine-grained vector as the text local fine-grained vector;
[0040] A text output vector determination unit, configured to obtain a first output vector corresponding to the first feature extraction vector and a second output vector corresponding to the second feature extraction vector based on the text global information vector and the text local fine-grained vector;
[0041] A first learning result determination unit, configured to input the first output vector into the first multi-layer perceptron to obtain a first learning vector corresponding to the first feature extraction vector, and input the second output vector into the first multi-layer perceptron to obtain a second learning vector corresponding to the second feature extraction vector, and use the first learning vector and the second learning vector as the first learning result.
[0042] Wherein, the text global vector determination unit includes:
[0043] An initial hidden vector determination subunit, configured to use the first initial hidden vector and the second initial hidden vector as the initial hidden vectors corresponding to the text data to be matched respectively; the initial hidden vector is a hidden vector matrix with H rows; H is obtained from the total number of words in the text data to be matched; the hidden vector matrix includes a hidden vector p k ; the hidden vector p k is the hidden vector corresponding to the k-th word obtained by traversing the text data to be matched; k is a positive integer less than or equal to H;
[0044] A convolution kernel acquisition subunit, configured to input the initial hidden vector into a first global feature learning layer to obtain a first multi-scale convolution kernel associated with the first global feature learning layer; the first multi-scale convolution kernel includes N first-type convolution kernels and (N - 1) second-type convolution kernels; N is a positive integer greater than 1;
[0045] A convolution feature determination subunit, configured to input the initial hidden vector into N first-type convolution kernels respectively to obtain N first convolution features, and obtain first-type convolution features and second-type convolution features from the N first convolution features; perform convolution processing on the second-type convolution features through (N - 1) second-type convolution kernels respectively to obtain (N - 1) second convolution features;
[0046] A pooling feature determination subunit, configured to input the first-type convolution features and (N - 1) second convolution features into an average pooling layer to obtain the pooling feature corresponding to the k-th word, until the value of k is H, to obtain the pooling features corresponding to each word in the text data to be matched respectively;
[0047] A global vector determination subunit, configured to input the pooling features corresponding to each word in the text data to be matched into a connection layer to obtain a text global information vector corresponding to the text data to be matched; the text global information vector includes a first global information vector corresponding to a first feature extraction vector and a second global information vector corresponding to a second feature extraction vector.
[0048] Wherein, the first initial hidden vector is a hidden vector matrix with m rows; the second initial hidden vector is a hidden vector matrix with n rows; m is obtained from the total number of words associated with the search service data; n is obtained from the total number of words associated with the first-modal service data;
[0049] The text local vector determination unit includes:
[0050] A hidden vector acquisition subunit, configured to input the first initial hidden vector and the second initial hidden vector into a first local feature learning layer, and traverse and obtain the hidden vector p corresponding to the i-th word from the first initial hidden vector associated with the service search data aiand the hidden vector p corresponding to the u-th character au Traverse and obtain the hidden vector p corresponding to the j-th character from the second initial hidden vector associated with the first-modal service data bj and the hidden vector p corresponding to the v-th character bv ; both i and u are positive integers less than or equal to m; both j and v are positive integers less than or equal to n;
[0051] A local weight determination subunit for determining the hidden vector p ai and the hidden vector p bj The first local weight e between them ij Determine the hidden vector p ai and the hidden vector p bv The second local weight e between them iv Determine the hidden vector p au and the hidden vector p bj The third local weight e between them uj ;
[0052] A first local vector determination subunit for determining the first intermediate hidden vector corresponding to the i-th character based on the first local weight e ij , the second local weight e iv and the hidden vector p bj until the value of i is m, obtaining m first intermediate hidden vectors, and based on the m first intermediate hidden vectors, obtaining the first local fine-grained vector corresponding to the first feature extraction vector; Until the value of i is m, obtain m first intermediate hidden vectors, and based on the m first intermediate hidden vectors, obtain the first local fine-grained vector corresponding to the first feature extraction vector;
[0053] A second local vector determination subunit for determining the second intermediate hidden vector corresponding to the j-th character based on the first local weight e ij , the third local weight e uj and the hidden vector p ai until the value of j is n, obtaining n second intermediate hidden vectors, and based on the n second intermediate hidden vectors, obtaining the second local fine-grained vector corresponding to the second feature extraction vector; Until the value of j is n, obtain n second intermediate hidden vectors, and based on the n second intermediate hidden vectors, obtain the second local fine-grained vector corresponding to the second feature extraction vector;
[0054] A text local vector determination subunit for using the first local fine-grained vector and the second local fine-grained vector as the text local fine-grained vector.
[0055] Among them, the feature learner includes a second multi-layer perceptron associated with the multi-modal feature learner; the multi-modal feature learner includes a second bidirectional hidden encoding layer, a second global feature learning layer, and a second local feature learning layer;
[0056] This second learning and processing module includes:
[0057] A multi-modal initial vector determination unit, configured to respectively input a first feature extraction vector and a third feature extraction vector of second-modal service data into a second bidirectional hidden encoding layer, to obtain a third initial hidden vector corresponding to the third feature extraction vector and a fourth initial hidden vector corresponding to the first feature extraction vector;
[0058] A multi-modal global vector determination unit, configured to obtain a third global information vector corresponding to the third feature extraction vector and a fourth global information vector corresponding to the first feature extraction vector based on the third initial hidden vector, the fourth initial hidden vector, and a second global feature learning layer, and use the third global information vector and the fourth global information vector as multi-modal global information vectors;
[0059] A multi-modal local vector determination unit, configured to obtain a third local fine-grained vector corresponding to the third feature extraction vector and a fourth local fine-grained vector corresponding to the first feature extraction vector based on the third initial hidden vector, the fourth initial hidden vector, and a second local feature learning layer, and use the third local fine-grained vector and the fourth local fine-grained vector as multi-modal local fine-grained vectors;
[0060] A multi-modal output vector determination unit, configured to obtain a third output vector corresponding to the third feature extraction vector and a fourth output vector corresponding to the first feature extraction vector based on the multi-modal global information vectors and the multi-modal local fine-grained vectors;
[0061] A second learning result determination unit, configured to input the third output vector into a second multi-layer perceptron to obtain a third learning vector corresponding to the third feature extraction vector, and input the fourth output vector into the second multi-layer perceptron to obtain a fourth learning vector corresponding to the first feature extraction vector, and use the third learning vector and the fourth learning vector as the second learning result.
[0062] Wherein, the learning vectors in the first learning result include a first learning vector corresponding to the first feature extraction vector and a second learning vector corresponding to the second feature extraction vector; the learning vectors in the second learning result include a third learning vector corresponding to the third feature extraction vector and a fourth learning vector corresponding to the first feature extraction vector;
[0063] The splicing processing module includes:
[0064] A splicing processing unit, configured to splice the first learning vector and the fourth learning vector through a prediction generator to obtain a first spliced vector, and splice the second learning vector and the third learning vector to obtain a second spliced vector;
[0065] A splicing result determination unit, configured to use the first spliced vector and the second spliced vector as the vector splicing result.
[0066] Wherein, the apparatus further comprises:
[0067] A search result determination module, configured to use the to-be-matched service data as the service search result corresponding to the service search request when the matching degree between the search service data and the to-be-matched service data indicates that the search service data matches the to-be-matched service data successfully;
[0068] A search result push module, configured to push the service search result to a user terminal, so that the user terminal switches the display interface from a search display interface to a service data display interface, and outputs the service search result to the service data display interface.
[0069] An embodiment of the present application provides a data processing apparatus, comprising:
[0070] A sample acquisition module, configured to acquire a sample data group for training a multi-modal training model; the sample data group includes a first type of sample data group and a second type of sample data group; the first type of sample data group is a sample data group with sample label information; the second type of sample data group is a sample data group without sample label information; the sample label information is used to indicate the matching degree between the first type of sample data groups;
[0071] A prediction result output module, configured to input the sample data group into the multi-modal training model, output a prediction result between the sample data groups by the multi-modal training model, and use the prediction result as prediction label information; the multi-modal training model includes a sample feature extractor, a sample feature learner, and a sample prediction generator;
[0072] An optimal perturbation amount determination module, configured to acquire a sample splicing vector corresponding to the sample data group, and determine an optimal perturbation amount of the sample data group based on the sample splicing vector, a model loss function of the multi-modal training model, and an expected condition associated with the multi-modal training model;
[0073] An iterative training module, configured to generate an adversarial sample data corresponding to the sample data group based on the optimal perturbation amount and the sample splicing vector, and perform iterative training on the multi-modal training model based on the adversarial sample data and the model loss function to obtain a model training result;
[0074] A model determination module, configured to use the multi-modal training model that satisfies the model convergence condition as a multi-modal matching model for predicting the matching degree between service data groups when the model training result indicates that the multi-modal training model after iterative training satisfies the model convergence condition.
[0075] Wherein, the optimal perturbation amount determination module includes:
[0076] An acquisition unit, configured to acquire a sample splicing vector corresponding to the sample data group, and acquire model parameters of the multi-modal training model;
[0077] An initial perturbation amount determination unit, configured to determine an initial perturbation amount corresponding to a sample data group based on a sample splicing vector, prediction label information, model parameters, and a model loss function of a multi-modal training model;
[0078] An optimal perturbation amount determination unit, configured to obtain an expected condition associated with the multi-modal training model, and when it is detected that there is an initial perturbation amount in the initial perturbation amounts that satisfies the expected condition, use the initial perturbation amount that satisfies the expected condition as the optimal perturbation amount of the sample data group.
[0079] On the one hand, the present application provides a computer device, including: a processor, a memory, and a network interface;
[0080] The processor is connected to the memory and the network interface. Among them, the network interface is used to provide a data communication function, the memory is used to store a computer program, and the processor is used to call the computer program to execute the method in the above-mentioned one aspect in the embodiment of the present application.
[0081] On the one hand, the present application provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions, and when the program instructions are executed by a processor, the method in the above-mentioned one aspect in the embodiment of the present application is executed.
[0082] On the one hand, the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method in the above-mentioned one aspect.
[0083] In an embodiment of the present application, when a computer device obtains service search data and service data to be matched including first-modal service data and second-modal service data, it can obtain a multi-modal matching model. The multi-modal matching model here may include a feature learner and a prediction generator. Further, the computer device can perform a first learning process on a first feature extraction vector of the search service data and a second feature extraction vector of the first-modal service data through a text feature learner in the feature learner to obtain a first learning result, so as to fully learn the text features between the search service data and the first-modal service data. At the same time, the computer device can also perform a second learning process on the first feature extraction vector and a third feature extraction vector of the second-modal service data through a multi-modal feature learner in the feature learner to obtain a second learning result, so as to learn the multi-modal features between the search service data and the second-modal service data. Further, the computer device splices the learning vectors in the first learning result and the learning vectors in the second learning result, so that the feature vectors of the search service data and the service data to be matched (i.e., the vector splicing result) can be quickly and accurately represented. Furthermore, when predicting the matching degree between the search service data and the service data to be matched based on the vector splicing result, the accuracy of the prediction result can be effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0085] Figure 1 is a schematic structural diagram of a network architecture provided by an embodiment of the present application;
[0086] Figure 2 is a schematic diagram of a data interaction scenario provided by an embodiment of the present application;
[0087] Figure 3 is a schematic flowchart of a data processing method provided by an embodiment of the present application;
[0088] Figure 4 is a schematic diagram of a scenario for determining search service data provided by an embodiment of the present application;
[0089] Figure 5a is a schematic structural diagram of a text feature learner provided by an embodiment of the present application;
[0090] Figure 5bIt is a schematic structural diagram of a global feature learning layer provided by an embodiment of the present application;
[0091] Figure 6 It is a schematic diagram of a scenario for displaying business search results provided by an embodiment of the present application;
[0092] Figure 7 It is a schematic flowchart of a method for training a multimodal matching model provided by an embodiment of the present application;
[0093] Figure 8 It is a schematic structural diagram of a data processing device provided by an embodiment of the present application;
[0094] Figure 9 It is a schematic diagram of a computer device provided by an embodiment of the present application;
[0095] Figure 10 It is a schematic structural diagram of a data processing device provided by an embodiment of the present application;
[0096] Figure 11 It is a schematic diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0097] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0098] Please refer to Figure 1 , Figure 1 It is a schematic structural diagram of a network architecture provided by an embodiment of the present application. As Figure 1 shown, the network architecture may include a server 10 and a user terminal cluster. The user terminal cluster may include one or more user terminals, and the number of user terminals will not be limited here. As Figure 1 shown, it may specifically include user terminal 100a, user terminal 100b, user terminal 100c,..., user terminal 100n. As Figure 1 shown, user terminal 100a, user terminal 100b, user terminal 100c,..., user terminal 100n may be respectively network-connected to the above-mentioned server 10, so that each user terminal can perform data interaction with the server 10 through this network connection.
[0099] Among them, each user terminal in the user terminal cluster may include: intelligent terminals with data processing functions such as smart phones, tablet computers, laptop computers, desktop computers, wearable devices, smart homes, and head-mounted devices. It should be understood that, as Figure 1 shown, each user terminal in the user terminal cluster may be installed with a target application (i.e., application client). When the application client runs on each user terminal, it can perform data interaction with the above-mentioned Figure 1 shown server 10 respectively. Among them, the application client may include application clients with service search functions such as social clients, multimedia clients (e.g., video clients), entertainment clients (e.g., game clients), education clients, live broadcast clients, and shopping clients. Among them, the application client may be an independent client or an embedded sub-client integrated in a certain client (e.g., social client, education client, and multimedia client, etc.), which is not limited here.
[0100] As Figure 1 shown, the server 10 in the embodiment of the present application may be the server corresponding to the application client. The server 10 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0101] For ease of understanding, in the embodiment of the present application, one user terminal can be selected from the multiple user terminals shown in Figure 1 as the target user terminal. For example, in the embodiment of the present application, the user terminal 100a shown in Figure 1 can be used as the target user terminal, and the target application (i.e., application client) with the service search function can be integrated in the target user terminal. At this time, the target user terminal can realize data interaction with the server 10 through the service data platform corresponding to the application client.
[0102] It should be understood that the data processing method in the embodiment of the present application may involve the machine learning direction in the field of artificial intelligence. It can be understood that so-called artificial intelligence (AI for short) refers to using digital computers or computer devices controlled by data computers (e.g., Figure 1A new technical science that uses the server 10) shown to simulate, extend, and expand the theory, methods, techniques, and application systems of human intelligence. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0103] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields involved, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0104] Among them, Machine Learning (ML) is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0105] It can be understood that the multi-modal matching model in the embodiments of this application has important application value in search scenarios. It should be understood that this multi-modal matching model can be used to predict the matching degree between business data groups (for example, a data group composed of two types of business data: search business data and business data to be matched). Among them, the business data matched by this multi-modal matching model can include at least one of the following business types: text type, video type, and picture type, etc.
[0106] For example, in a shopping search scenario, a target user terminal (e.g., user terminal 100a) can obtain search service data (e.g., text data or picture data associated with a thermos cup) input by the target user in an application client (e.g., a shopping client), and then can send the search service data to the server corresponding to the shopping client (e.g., server 10). At this time, the server 10 can quickly and accurately search for service data (e.g., a purchase link data of a certain thermos cup) that matches the search service data through the multimodal matching model, and push the searched service data to the target user terminal so that the target user can select a desired product.
[0107] Optionally, in a reading search scenario, a target user terminal (e.g., user terminal 100a) can obtain search service data (e.g., title text data associated with a basketball) input by the target user in an application client (e.g., a social client), and then can send the search service data to the server corresponding to the social client (e.g., server 10). At this time, the server 10 can quickly and accurately search for service data (e.g., video data associated with a basketball) that matches the search service data through the multimodal matching model, and push the searched service data to the target user terminal so that the target user can view video data of interest, thereby improving user stickiness.
[0108] Optionally, in a video search scenario, a target user terminal (e.g., user terminal 100a) can obtain video data that the target user is watching in an application client (e.g., a video client), and use the watched video data as search service data (e.g., video data 1 associated with a cat), and then can send the search service data to the server corresponding to the video client (e.g., server 10). At this time, the server 10 can quickly and accurately search for service data (e.g., video data 2 associated with a cat) that matches the search service data through the multimodal matching model, and push the searched service data to the target user terminal, thereby improving the search experience of the target user.
[0109] Furthermore, please refer to Figure 2 , Figure 2 which is a schematic diagram of a data interaction scenario provided by an embodiment of the present application. As Figure 2 shown, the computer device in the embodiment of the present application can be a server 2B as shown in Figure 2 , and the server 2B can be the server 10 as shown in the above Figure 1 . The user terminal 2A in the embodiment of the present application can be any one of the user terminal clusters as shown in the above Figure 1 , for example, user terminal 100a.
[0110] It should be understood that the user corresponding to the user terminal 2A can search for the content of their interest (e.g., text data 1) in the search display interface of the application client of this user terminal. The user terminal 2A can use the content searched by the user as search service data, and then generate a service search request based on this search service data. Further, the user terminal 2A can send this service search request to Figure 2 the server 2B as shown. At this time, the server 2B can obtain the service data to be matched including the first-modal service data (e.g., text data 2) and the second-modal service data (e.g., video data 3) from the video database based on this service search request. Among them, the text data 2 can be the title text corresponding to the video data 3.
[0111] Further, the server 2B can obtain a multi-modal matching model for matching the search service data and the service data to be matched. Among them, as Figure 2 shown, this multi-modal matching model can include a feature extractor (Featureextraction, abbreviated as FE), a feature learner and perceptron (abbreviated as FLP), and a prediction and generato (abbreviated as PG). It should be understood that the computer device can extract a first feature extraction vector from this search service data through the word vector extraction network in the feature extractor, and extract a second feature extraction vector from this first-modal service data. At the same time, the computer device can extract a third feature extraction vector from this second-modal service data through the residual network in the feature extractor.
[0112] As Figure 2As shown, the feature learner may include a text feature learner (learner and perceptron for text, abbreviated as FLPT) and a multi-layer perceptron (Multilayer Perceptron, abbreviated as MLP) associated with the text feature learner. Among them, in the embodiments of this application, the multi-layer perceptron associated with the text feature learner may be referred to as the first multi-layer perceptron. It should be understood that the computer device may perform a first learning process on the first feature extraction vector and the second feature extraction vector through the text feature learner and the first multi-layer perceptron to obtain a first learning result. Among them, the learning vectors in the first learning result may include a first learning vector corresponding to the first feature extraction vector and a second learning vector corresponding to the second feature extraction vector. The learning vectors in the first learning result are obtained from the text global information vector and the text local fine-grained vector; the text global information vector here (for example, the first global information vector corresponding to the first feature extraction vector and the second global information vector corresponding to the second feature extraction vector) is obtained based on the first multi-scale convolution kernel in the first global feature learning layer of the text feature learner; the text local fine-grained vector here (for example, the first local fine-grained vector corresponding to the first feature extraction vector and the second local fine-grained vector corresponding to the second feature extraction vector) is obtained based on the first local feature learning layer of the text feature learner.
[0113] As Figure 2As shown, the feature learner may further include a multi-modal feature learner (Feature learner and perceptron for multi-type, abbreviated as FLPM) and a multi-layer perceptron associated with the multi-modal feature learner. Among them, in the embodiments of the present application, the multi-layer perceptron associated with the multi-modal feature learner may be referred to as the second multi-layer perceptron. It should be understood that the computer device may perform a second learning process on the first feature extraction vector and the third feature extraction vector through the multi-modal feature learner and the second multi-layer perceptron to obtain a second learning result. Among them, the learning vectors in the second learning result may include a third learning vector corresponding to the third feature extraction vector and a fourth learning vector corresponding to the first feature extraction vector. The learning vectors in the second learning result may be obtained from a multi-modal global information vector and a multi-modal local fine-grained vector; the multi-modal global information vector here (for example, the third global information vector corresponding to the third feature extraction vector and the fourth global information vector corresponding to the first feature extraction vector) is obtained based on the second multi-scale convolution kernel in the second global feature learning layer of the multi-modal feature learner; the multi-modal local fine-grained vector here (for example, the third local fine-grained vector corresponding to the third feature extraction vector and the fourth local fine-grained vector corresponding to the first feature extraction vector) is obtained based on the second local feature learning layer of the multi-modal feature learner.
[0114] Furthermore, the computer device may splice the learning vectors in the first learning result and the learning vectors in the second learning result through the multi-modal vector splicing layer in the prediction generator to obtain a vector splicing result. It can be understood that the computer device may splice the first learning vector and the fourth learning vector to obtain a first splicing vector, and splice the second learning vector and the third learning vector to obtain a second splicing vector. At this time, the computer device may input the first splicing vector and the second splicing vector into the multi-layer perceptron in the prediction generator, so as to predict the matching degree between the search service data and the to-be-matched service data and obtain a prediction result. Among them, in the embodiments of the present application, the multi-layer perceptron in the prediction generator may be referred to as the third multi-layer perceptron.
[0115] It can be understood that when the prediction result indicates that the search service data matches the service data to be matched successfully, the computer device may use the service data to be matched as the service search result corresponding to the service search request, and then may return the service search result to the user terminal 2A, so that the user terminal 2A can output the service search result on the service data display interface of the application client. When the prediction result indicates that the search service data fails to match the service data to be matched, the computer device may continue to obtain a new service data to be matched from the video database to match the search service data with this new service data to be matched.
[0116] It can be seen that the server 2B in the embodiment of the present application can, through the multimodal matching model, deeply learn the semantic information between the search service data and the service data to be matched, effectively map the semantic information of service data such as text type, video type or picture type to the same semantic space, respectively extract the semantic features of the search service data and the service data to be matched, and further improve the accuracy of the match between the search service data and the service data to be matched, thereby improving the accuracy of the prediction result and enabling the user to more accurately search for the service search result that matches the search service data.
[0117] Among them, the specific implementation manner in which the computer device predicts the matching degree between the search service data and the service data to be matched through the multimodal matching model can be referred to in the following Figures 3 - 7 corresponding embodiment.
[0118] Further, please refer to Figure 3 , Figure 3 which is a schematic flowchart of a data processing method provided by an embodiment of the present application. As Figure 3 shown, this method can be executed by a computer device with a matching degree prediction function. The computer device can be a user terminal (for example, the user terminal 100a shown above Figure 1 ), or can be a server (for example, the server 10 shown above Figure 1 ), which is not limited herein. For ease of understanding, this application takes the example that this method is executed by the server. This method may at least include the following steps S101-step S104:
[0119] Step S101, obtain a multimodal matching model for matching the search service data and the service data to be matched.
[0120] Specifically, a computer device with a matching degree prediction function can obtain service data that matches the search service data from a video database when it acquires the search service data. Herein, in the embodiments of the present application, the service data that matches the search service data can be referred to as service data to be matched. Further, the computer device can load a multi-modal matching model to match the search service data and the service data to be matched.
[0121] It should be understood that a user corresponding to a user terminal running an application client (for example, Figure 1 the user terminal 100a shown) can perform a triggering operation on a search page switching control (for example, a "Search" control) in the application display interface of the application client when accessing the application client, so that the display interface of the user terminal switches from the application display interface to a search display interface. The triggering operation can include contact operations such as clicking and long pressing, and can also include non-contact operations such as voice and gestures, which are not limited herein. Further, the user terminal can obtain the search service data that the user is interested in. The search service data can be service data input by the user into the search area by voice or clicking, etc. Optionally, the search service data can also be service data determined after the user performs a triggering operation on a certain interesting hot topic title in the hot list on the search display interface. When the user performs a triggering operation on the search control in the search display interface, the user terminal can respond to the triggering operation, generate a service search request including the search service data, and then send the service search request to the computer device.
[0122] For ease of understanding, further, please refer to Figure 4 , Figure 4 which is a schematic diagram of a scenario for determining search service data provided by the embodiments of the present application. As Figure 4 shown, the user terminal in the embodiments of the present application can be a user terminal running an application client (for example, a social client), and the user terminal can be any one of the user terminal clusters shown above Figure 1 such as the user terminal 100a.
[0123] It should be understood that the application display interface 400a in the embodiments of the present application can include a search page switching control (for example, a "Search" control). When the user corresponding to the user terminal needs to search for service data of interest, the user can perform a triggering operation on the search page switching control. At this time, the user terminal can respond to the triggering operation and switch the display interface from the application display interface 400a to the search display interface 400b. As Figure 4As shown, the user can determine the business data of interest in the search area of the search display interface 400b by means of voice or click input, etc. Furthermore, when the input is completed, the user can perform a trigger operation on the search control, so that the user terminal can respond to this trigger operation, take the business data in the search area as the search business data, and generate a business search request for sending to the computer device corresponding to the application client.
[0124] It should be understood that a popular list composed of currently popular titles can also be displayed in the search display interface 400b. Among them, the popular list can include multiple titles, such as Figure 4 As shown, the popular list can specifically include Title 1, Title 2, Title 3, and Title 4. Optionally, the user can also directly perform a trigger operation on a piece of business data of interest (for example, Title 1) in the popular list. Furthermore, the user terminal can respond to this trigger operation, so that Title 1 associated with this trigger operation can be taken as the search business data, and a business search request for sending to the computer device corresponding to the application client can be generated. For example, Title 1 can be "I love China".
[0125] Furthermore, when receiving the business search request, the computer device can obtain the business data to be matched based on the business search request. Among them, the business data to be matched can be text-type business data, or picture-type business data, or business data including text type and video type. There is no limitation here. In this embodiment of the application, the business data to be matched can take the business data including text type and video type as an example to illustrate the matching of the matching degree between the search business data and the business data to be matched by means of the multimodal matching model.
[0126] It can be understood that the computer device can obtain the business data with the first business type from the video database, take the business data with the first business type (for example, text type) as the first-modal business data, and obtain the business data with the second business type (for example, video type) from the video database, take the business data with the second business type as the second-modal business data. Among them, the first business type is different from the second business type. Furthermore, the computer device can take the business data jointly mapped by the first-modal business data and the second-modal business data as the business data to be matched. At this time, the computer device can obtain a multimodal matching model for matching the search business data and the business data to be matched. Among them, the multimodal matching model can include a feature extractor, a feature learner, and a prediction generator.
[0127] Among them, if the service type of the search service data is the first service type, and the first service type belongs to the text type, then the second service type includes at least one of the following service types: video type or picture type. For example, the search service data in the embodiments of the present application may be text data 1, the first modal service data may be text data 2, and the second modal service data may be video data 3. Among them, the text data 2 may be the title text corresponding to the video data 3. The service data jointly mapped by the text data 2 and the video data 3 may be used as the service data to be matched.
[0128] It can be understood that the computer device may use the search service data and the first modal service data as the text data to be encoded, and then may extract a feature extraction vector from the text data to be encoded through the word vector extraction network in the feature extractor. Among them, the feature extraction vector may include a first feature extraction vector extracted from the search service data and a second feature extraction vector extracted from the first modal service data. At the same time, the computer device may also perform frame extraction processing on the second modal service data to obtain video frames, and then may input the video frames into the residual network in the feature extractor, and the residual network extracts a third feature extraction vector corresponding to the second modal service data.
[0129] It should be understood that the computer device may use the search service data and the first modal service data as the text data to be encoded, and then may preprocess the text data to be encoded to use the preprocessed text data to be encoded as the text data to be matched. Among them, the text data to be matched may include the text data to be matched a obtained by preprocessing the search service data (for example, text data 1) and the text data to be matched b obtained by preprocessing the first modal service data (for example, text data 2).
[0130] Among them, the preprocessing here may include special symbol processing, English case conversion, and simplification and unification of complex and simplified fonts. Further, in terms of the feature representation of the text data to be matched, the computer device may consider the character-level features of the text data to be matched, may load a word segmentation model and a text word list (for example, Word2Vec word list), and perform character segmentation processing on the text data to be matched to obtain the character information sequence and character position sequence corresponding to the text data to be matched. Among them, the word segmentation model here may be a pre-trained word segmentation model or other types of word segmentation models such as the qq word segmentation model, which is not limited here.
[0131] It can be understood that the feature extractor of the multimodal matching model may include a word vector extraction network and a residual network. The word vector extraction network here can be used to extract features from business data of text type. For example, the word vector extraction network can be a network composed of a Word2Vec model, a GloVe model, or a fastText model, etc. The residual network here can be used to extract features from business data of picture type. For example, the residual network can be a Resnet152 neural network. Therefore, the construction of the feature extractor can effectively extract the feature extraction vectors corresponding to business data of multiple business types (such as text type, video type, and picture type), so as to obtain a richer feature representation, and further improve the accuracy of subsequent matching.
[0132] Among them, the total number of words in the text data to be matched is H; H here can be a positive integer. It is worth noting that if the length of the text data to be matched is too long (that is, the total number of words is too large), it may cause the computer device to have too large or too small gradients during feature extraction. To solve this problem, the computer device can set the maximum length of the text data to be matched (for example, 128). When the total number of words in the text data to be matched is greater than the maximum length, the computer device can truncate the text data to be matched to obtain multiple sequences, and then splice the feature extraction vectors of each sequence extracted to obtain the feature extraction vector of the text data to be matched. Optionally, the computer device can also perform summary extraction on the text data to be matched, etc., to compress the text data to be matched within the maximum length, and then perform feature extraction on the compressed text data to be matched.
[0133] It should be understood that the computer device can traverse and obtain the word information corresponding to the k-th word of the text data to be matched from the word information sequence of the text data to be matched, and can use the obtained word information as the target word information. At the same time, the computer device can also obtain the word position information corresponding to the target word information from the word position sequence, and can use the obtained word position information as the target word position information. Among them, k here can be a positive integer less than or equal to H. It should be understood that the computer device can input the target word information into the word vector extraction network, and the word vector extraction network extracts the target word information vector corresponding to the k-th word, and inputs the target word position information into the word vector extraction network, and the word vector extraction network extracts the target word position vector corresponding to the k-th word. Among them, the word vector extraction network can be trained based on the text vocabulary. Further, the computer device can obtain the feature extraction vector corresponding to the k-th word based on the target word information vector and the target word position vector, until the value of k is H, to obtain the feature extraction vector corresponding to the text data to be matched.
[0134] For example, the text data to be matched here can be "I love China", and the character information sequence obtained through character segmentation processing can be "I", "love", "China", "country". The computer device can traverse the character information sequence to obtain the character information corresponding to the first character in the text data to be matched (for example, "I"), and can use the obtained character information as the target character information. At the same time, the computer device can also obtain the character position information corresponding to the target character information (for example, "1") from the character position sequence obtained after character segmentation processing, and can use the obtained character position information as the target character position information. At this time, the computer device can input the target character information "I" into the word vector extraction network, and the word vector extraction network extracts the target character information vector corresponding to "I" (for example, 300-dimensional). Further, the computer device can input the target character position information "1" corresponding to "I" into the word vector extraction network, and the word vector extraction network extracts the target character position vector corresponding to "I". At this time, the computer device can perform superposition summation and averaging processing on the target character information vector corresponding to "I" and the target character position vector corresponding to "I" to obtain the feature extraction vector corresponding to "I".
[0135] It can be seen from this that the feature extractor can extract the character information features (i.e., character information vectors) of the text data to be matched (for example, the text data to be matched a or the text data to be matched b). In other words, the computer device can extract the complete digital and English word features in the text data to be matched, avoiding the loss of semantic information caused by splitting numbers and English words. In addition, the computer device can also extract the character position features (i.e., character position vectors) of the text data to be matched, enabling more accurate matching of search service data and the text data to be matched in the future.
[0136] It should be understood that the computer device can also perform frame extraction on the second-modal service data (for example, video data 3) to obtain video frames, and then can input the video frames into the residual network in the feature extractor, and the residual network extracts the third feature extraction vector corresponding to the second-modal service data.
[0137] For example, the computer device can extract one frame every 1 second to obtain the video frames corresponding to video data 3. It can be understood that when the total number of the video frames is too large, it will also cause the situation that the gradient of the computer device during feature extraction is too large or too small. To solve this problem, the computer device can set the maximum number of frames of the video frames (for example, 128). When the total number of video frames obtained after the computer device performs frame extraction on video data 3 is greater than the maximum number of frames, the computer device can uniformly and equally spaced extract the video frames to be deleted from the video frames and delete the video frames to be deleted.
[0138] Among them, the computer device can determine the ratio of the total number of video frames to the number of video frames to be deleted. If the ratio is an odd number, the computer device can use the middle video frame as the video frame to be deleted. For example, when the total number of video frames obtained by the computer device after the video data 3 is extracted is 200 frames, it can be determined that the number of video frames to be deleted is 72 frames. At this time, the computer device can determine the ratio of the total number to the number of video frames to be deleted (for example, 2.7, which is rounded to 3). In other words, the computer device can determine one video frame to be deleted in every 3 frames of 200 video frames. For example, these 3 video frames can be video frame 1, video frame 2 and video frame 3 respectively. At this time, the computer device can use the middle video frame (i.e., video frame 2) as the video frame to be deleted. Optionally, if the ratio is an even number (for example, 4), the computer device can arbitrarily select the video frame on the left or right side of the middle as the video frame to be deleted. For example, the four video frames may be video frame 1, video frame 2, video frame 3 and video frame 4 respectively. At this time, the computer device may select any one of video frame 2 or video frame 3 as the video frame to be deleted.
[0139] Therefore, after the feature extractor of the multimodal matching model, the computer device can extract the first feature extraction vector from the search service data (for example, text a). For example, the first feature extraction vector can be represented as w1, w2, ..., w m . Wherein, m is obtained by the total number of words associated with text a. The computer device can extract a second feature extraction vector from the first modal service data (e.g., text b). The second feature extraction vector can be represented as q1, q2, ..., q n . Where n is obtained by the total number of words associated with text a. The computer device can extract a third feature extraction vector from the second modality service data (e.g., video c). The third feature extraction vector can be represented as r1, r2, ..., r o . where o is obtained from the total number of frames associated with video c.
[0140] Step S102: Perform a first learning process on the first feature extraction vector of the search business data and the second feature extraction vector of the first modal business data through a text feature learner in the feature learner to obtain a first learning result.
[0141] Among them, the feature learner in the multi-modal matching model may include a text feature learner and a first multi-layer perceptron associated with the text feature learner. The text feature learner may include a first bidirectional hidden encoding layer, a first global feature learning layer, and a first local feature learning layer. Specifically, the computer device may input the first feature extraction vector of the search service data and the second feature extraction vector of the first-modal service data into the first bidirectional hidden encoding layer respectively, to obtain a first initial hidden vector corresponding to the first feature extraction vector and a second initial hidden vector corresponding to the second feature extraction vector. Further, the computer device may, based on the first initial hidden vector, the second initial hidden vector, and the first global feature learning layer, obtain a first global information vector corresponding to the first feature extraction vector and a second global information vector corresponding to the second feature extraction vector, and use the first global information vector and the second global information vector as text global information vectors; at the same time, the computer device may, based on the first initial hidden vector, the second initial hidden vector, and the first local feature learning layer, obtain a first local fine-grained vector corresponding to the first feature extraction vector and a second local fine-grained vector corresponding to the second feature extraction vector, and use the first local fine-grained vector and the second local fine-grained vector as text local fine-grained vectors. Further, the computer device may, based on the text global information vector and the text local fine-grained vector, obtain a first output vector corresponding to the first feature extraction vector and a second output vector corresponding to the second feature extraction vector. At this time, the computer device may input the first output vector into the first multi-layer perceptron to obtain a first learning vector corresponding to the first feature extraction vector, and may input the second output vector into the first multi-layer perceptron to obtain a second learning vector corresponding to the second feature extraction vector, and then may use the first learning vector and the second learning vector as the first learning result.
[0142] Among them, since the text feature learner involves two input sources, namely the first feature extraction vector and the second feature extraction vector, the embodiment of the present application may adopt a two-tower model framework that can distinguish the two input sources structurally, and may make improvements on the basis of the two-tower model framework, so that the text feature learner can obtain better learning effects and learning efficiency.
[0143] Further, please refer to Figure 5a , Figure 5a which is a schematic structural diagram of a text feature learner provided by an embodiment of the present application. As Figure 5a shown, the text feature learner in the embodiment of the present application may include a bidirectional hidden encoding layer 511, a global feature learning layer 512, and a local feature learning layer 513.
[0144] Among them, it can be understood that Figure 5aThe bidirectional hidden encoding layer 511 (i.e., the first bidirectional hidden encoding layer) in the text feature learner shown can be used to encode the feature extraction vector into a hidden state. For example, the computer device can encode the first feature extraction vector associated with the search service data into a first initial hidden vector, and encode the second feature extraction vector associated with the service data to be matched into a second initial hidden vector. This first bidirectional hidden encoding layer can better learn the vector representation (i.e., the initial hidden vector) of the hidden state between the service data (search service data or service data to be matched) input to the multi-modal matching model, thereby making the learned semantic features more abstract and more robust. For example, the bidirectional hidden encoding layer can be a bidirectional long short-term memory network (Bi-directional Long Short-Term Memory, abbreviated as BiLSTM network), a bidirectional gated recurrent unit (Bi-directional Gated Recurrent Unit, abbreviated as BiGRU), or a convolutional neural network (abbreviated as CNN network), etc.
[0145] It should be understood that Figure 5a The global feature learning layer 512 shown can be an enhanced network in network structure (ENIN structure) for learning global features. Among them, the text feature learner in the embodiment of the present application can use the ENIN structure to increase the large-scale convolution kernel, so that the learned features have stronger abstraction and robustness. In other words, the computer device can input the initial hidden vectors (i.e., the first initial hidden vector and the second initial hidden vector) obtained by the bidirectional hidden encoding layer 511 into the global feature learning layer 512 to obtain the text global information vector. The text global information vector here can include the first global information vector corresponding to the first feature extraction vector and the second global information vector corresponding to the second feature extraction vector.
[0146] Figure 5a The local feature learning layer 513 shown can be an attention learning mechanism layer for learning local fine-grained features. For example, the computer device can input the initial hidden vectors (i.e., the first initial hidden vector and the second initial hidden vector) obtained by the bidirectional hidden encoding layer 511 into the local feature learning layer 513 to obtain the text local fine-grained vector. The text local fine-grained vector here can include the first local fine-grained vector corresponding to the first feature extraction vector and the second local fine-grained vector corresponding to the second feature extraction vector.
[0147] Further, the computer device may perform a superposition summation process on the text global information vector obtained through the global information feature learning layer 512 and the text local fine-grained vector of the corresponding feature extraction vector obtained through the local feature learning layer 513 to obtain an output vector. The output vector here may include a first output vector corresponding to the first feature extraction vector and a second output vector corresponding to the second feature extraction vector. Among them, the first output vector may be obtained by the computer device performing a superposition summation process on the first global information vector and the first local fine-grained vector, and the second output vector may be obtained by the computer device performing a superposition summation process on the second global information vector and the second local fine-grained vector.
[0148] Among them, it can be understood that the calculation formula for the computer device to input the feature extraction vector into the first bidirectional hidden encoding layer to obtain the corresponding initial hidden vector can be as shown in the following formulas (1) and (2):
[0149]
[0150] Among them, w i refers to the word vector corresponding to the i-th word obtained from the first feature extraction vector corresponding to the text data a to be matched. m may be the total number of words in the text data a to be matched. p ai refers to the hidden vector corresponding to the i-th word in the text data a to be matched. The first initial hidden vector in the embodiment of the present application may be a hidden vector matrix p a .
[0151]
[0152] Among them, q j refers to the word vector corresponding to the j-th word obtained from the second feature extraction vector corresponding to the text data b to be matched. n may be the total number of words in the text data b to be matched. p bj refers to the hidden vector corresponding to the j-th word in the text data b to be matched. The second initial hidden vector in the embodiment of the present application may be a hidden vector matrix p b .
[0153] Further, the computer device may use the first initial hidden vector and the second initial hidden vector as the initial hidden vectors corresponding to the text data to be matched respectively, and input the initial hidden vector into the first global feature learning layer to obtain the global information vector corresponding to the text data to be matched. Specifically, the calculation formula for the computer device to obtain the global information vector corresponding to the text data to be matched can be as shown in the following formula (3):
[0154] x = ENIN(p), (3)
[0155] Wherein, p may be the initial hidden vector corresponding to the text data to be matched. x may represent the global information vector corresponding to the text data to be matched.
[0156] It can be understood that the initial hidden vector may be a hidden vector matrix with H rows and D columns. Here, H may be obtained from the total number of words in the text data to be matched. Here, D may be the vector dimension obtained by feature extraction of the text data to be matched. The hidden vector matrix may include the hidden vector p k . The hidden vector p k may be the hidden vector corresponding to the k-th word obtained by traversing the text data to be matched; here, k may be a positive integer less than or equal to H.
[0157] It should be understood that the computer device may input the initial hidden vector into the first global feature learning layer to obtain the first multi-scale convolution kernel associated with the first global feature learning layer. Specifically, the formula for the computer device to input the initial hidden vector into the first multi-scale convolution kernel for convolution calculation may be as shown in the following formula (4):
[0158]
[0159] Wherein, w represents the size of the convolution kernel, W f is the weight size specified for the corresponding convolution kernel, b is the specified bias, and ReLU() is the activation function. p k:k+w-1 represents the vector matrix formed by the hidden vector corresponding to the k-th word to the hidden vector corresponding to the k+w-1-th word in the text data to be matched. represents the convolution feature obtained after the k-th word in the text data to be matched is subjected to convolution calculation in the convolution kernel. It should be understood that if the initial hidden vector is a hidden vector matrix with H rows and D columns, the convolution feature obtained after convolution calculation by the convolution kernel shown in the above formula (4) is a matrix with (H-w+1) rows and (D-w+1) columns.
[0160] Wherein, the first multi-scale convolution kernel obtained by the computer device may include N first-type convolution kernels and (N-1) second-type convolution kernels. Here, N may be a positive integer greater than 1. It should be understood that the first-type convolution kernel may be a convolution kernel that does not change the convolution size. For example, a 1*1 convolution kernel. The second-type convolution kernel may be a convolution kernel that changes the convolution size. For example, a 2*2 convolution kernel, a 3*3 convolution kernel, or a 4*4 convolution kernel, etc.
[0161] Further, the computer device may input the initial hidden vector into the N first-type convolutional kernels respectively to obtain N first convolutional features. At this time, the computer device may obtain a first convolutional feature from these N first convolutional features for directly inputting into the average pooling layer, and may refer to this obtained first convolutional feature as the first-type convolutional feature. Meanwhile, the computer device may obtain the (N - 1) first convolutional features other than the first-type convolutional feature, and then may refer to these (N - 1) first convolutional features as the second-type convolutional features.
[0162] It should be understood that the computer device may perform convolutional processing on the second-type convolutional features through the (N - 1) second-type convolutional kernels respectively to obtain (N - 1) second convolutional features. Further, the computer device may input the first-type convolutional feature and the (N - 1) second convolutional features into the average pooling layer, so as to obtain the pooling feature corresponding to the k-th word. Until the value of k is H, the pooling features corresponding to each word in the text data to be matched are obtained. It can be understood that the computer device may input the pooling features corresponding to each word in the text data to be matched into the connection layer to obtain the global information vector corresponding to the text data to be matched.
[0163] For ease of understanding, further, please refer to Figure 5b , Figure 5b which is a schematic structural diagram of a global feature learning layer provided by an embodiment of the present application. As Figure 5b shown, the global feature learning layer in the embodiment of the present application may be the global feature learning layer 512 shown above Figure 5a shown. The first multi-scale convolutional kernels associated with the global feature learning layer in the embodiment of the present application may include N first-type convolutional kernels and (N - 1) second-type convolutional kernels. Taking N = 4 as an example, Figure 5b the first multi-scale convolutional kernel shown may include 4 first-type convolutional kernels and 3 second-type convolutional kernels. These 4 first-type convolutional kernels are all 1*1 convolutional kernels, and may specifically include convolutional kernel 51a, convolutional kernel 52a, convolutional kernel 53a, and convolutional kernel 54a. These 3 second-type convolutional kernels may specifically include convolutional kernel 55b (for example, 2*2 convolutional kernel), convolutional kernel 56b (for example, 3*3 convolutional kernel), and convolutional kernel 57b (for example, 4*4 convolutional kernel).
[0164] It should be understood that the computer device can input the initial hidden vectors corresponding to the text data to be matched (for example, a 10*300 hidden vector matrix) into these 4 first-type convolution kernels respectively. Until the value of k is 10, 4 corresponding first convolution features of 10*300 are obtained. Among them, these 4 first convolution features can include convolution feature 1 obtained through convolution kernel 51a, convolution feature 2 obtained through convolution kernel 52a, convolution feature 3 obtained through convolution kernel 53a, and convolution feature 4 obtained through convolution kernel 54a.
[0165] At this time, the computer device can obtain convolution feature 1 from these 4 first convolution features as the first-type convolution feature, and regard the other 3 first convolution features (for example, convolution feature 2, convolution feature 3, and convolution feature 4) except convolution feature 1 as the second-type convolution feature. Further, the computer device can input convolution feature 2 into convolution kernel 55b. Until the value of k is 10, convolution feature 5 of 9*299 is obtained. Similarly, the computer device can input convolution feature 3 into convolution kernel 56b. Until the value of k is 10, convolution feature 6 of 8*298 is obtained, and convolution feature 4 can be input into convolution kernel 57b. Until the value of k is 10, convolution feature 7 of 7*297 is obtained.
[0166] Further, the computer device can input convolution feature 1, convolution feature 5, convolution feature 6, and convolution feature 7 into Figure 5b the average pooling layer shown, obtain 4 convolution features with the same convolution size by filling with the number 0, and then can perform average processing on the 4 filled convolution features, so as to obtain the pooling features corresponding to each word in the text data to be matched. At this time, the computer device can input the pooling features corresponding to each word in the text data to be matched into the connection layer to obtain the global information vector corresponding to the text data to be matched.
[0167] At the same time, the computer device can obtain a first local fine-grained vector associated with the search service data and a second local fine-grained vector associated with the first-modal service data based on the first initial hidden vector, the second initial hidden vector, and the first local feature learning layer.
[0168] Among them, the first initial hidden vector can be a hidden vector matrix with m rows. Here, m is obtained from the total number of words associated with the search service data. In other words, here m is the total number of words in the text data a to be matched. The second initial hidden vector can be a hidden vector matrix with n rows; here, n can be obtained from the total number of words associated with the first-modal service data. In other words, here n is the total number of words in the text data b to be matched.
[0169] The computer device can input the first initial hidden vector and the second initial hidden vector into the first local feature learning layer, and traverse and obtain the hidden vector p corresponding to the i-th word from the first initial hidden vector associated with the business search data ai and the hidden vector p corresponding to the u-th word au . Wherein, both i and u here can be positive integers less than or equal to m. At the same time, the computer device can traverse and obtain the hidden vector p corresponding to the j-th word from the second initial hidden vector associated with the first-modal business data bj and the hidden vector p corresponding to the v-th word bv . Both j and v here are positive integers less than or equal to n.
[0170] Specifically, the calculation formula for the computer device to determine the local weight can be as shown in the following formula (5):
[0171]
[0172] Where p ai refers to the hidden vector corresponding to the i-th word in the text data a to be matched, and p bj refers to the hidden vector corresponding to the j-th word in the text data b to be matched. e ij refers to the local weight between the hidden vector p ai and the hidden vector p bj .
[0173] It can be understood that the computer device can determine the first local weight e ai between the hidden vector p bj and the hidden vector p ij through the above formula (5), the second local weight e ai between the hidden vector p bv and the hidden vector p iv , and the third local weight e au between the hidden vector p bj and the hidden vector p uj .
[0174] It should be understood that the computer device can determine the first intermediate hidden vector corresponding to the i-th word based on the first local weight e ij , the second local weight e iv , and the hidden vector p bj Until the value of i is m, m first intermediate hidden vectors are obtained. Further, based on the m first intermediate hidden vectors, a first local fine-grained vector associated with the search service data can be obtained. Similarly, the computer device can be based on the first local weight e ij , the third local weight e uj and the hidden vector p ai , determine the second intermediate hidden vector corresponding to the j-th word Until the value of j is n, n second intermediate hidden vectors are obtained. Further, based on the n second intermediate hidden vectors, a second local fine-grained vector associated with the first-modal service data can be obtained.
[0175] Specifically, the calculation formula for the computer device to determine the local fine-grained vector can be as shown in the following formulas (6) to (9):
[0176]
[0177]
[0178]
[0179]
[0180] Among them, represents the intermediate hidden vector corresponding to the i-th word in the text data a to be matched. represents the intermediate hidden vector corresponding to the j-th word in the text data b to be matched. y a represents the local fine-grained vector corresponding to the text data a to be matched, y b represents the local fine-grained vector corresponding to the text data b to be matched.
[0181] It should be understood that the computer device can perform a superposition summation process on the first global information vector and the first local fine-grained vector to obtain the first output vector corresponding to the first feature extraction vector, and can perform a superposition summation process on the second global information vector and the second local fine-grained vector to obtain the second output vector corresponding to the second feature extraction vector.
[0182] Specifically, the calculation formula for the computer device to determine the output vector can be as shown in the following formulas (10) to (11):
[0183] f a =[y a +x a , (10)
[0184] f b =[y b +x b, (11)
[0185] Among them, y a represents the local fine-grained vector corresponding to the text data a to be matched, and x a represents the global information vector corresponding to the text data a to be matched, and f a is the output vector obtained by the text data a to be matched through the text feature learning machine. y b represents the local fine-grained vector corresponding to the text data b to be matched, and x b represents the global information vector matrix corresponding to the text data b to be matched, and f b is the output vector obtained by the text data b to be matched through the text feature learning machine.
[0186] Furthermore, the computer device can input the first output vector into the first multi-layer perceptron to obtain the first learning vector corresponding to the first feature extraction vector, and input the second output vector into the first multi-layer perceptron to obtain the second learning vector corresponding to the second feature extraction vector. Furthermore, the first learning vector and the second learning vector can be used as the first learning result.
[0187] Specifically, the calculation formula for the computer device to determine the learning vector can be as shown in the following formulas (12) and (13):
[0188] z a = MLP(f a ), (12)
[0189] z b = MLP(f b ), (13)
[0190] Among them, f a is the output vector obtained by the text data a to be matched through the text feature learning machine, and z a is the learning vector obtained by the text data a to be matched through the first multi-layer perceptron, and f b is the output vector obtained by the text data b to be matched through the text feature learning machine, and z b is the learning vector obtained by the text data b to be matched through the first multi-layer perceptron.
[0191] Step S103, through the multi-modal feature learning machine in the feature learning machine, perform second learning processing on the first feature extraction vector and the third feature extraction vector of the second modal service data to obtain a second learning result.
[0192] Among them, the feature learner in the multimodal matching model may further include a multimodal feature learner and a second multi-layer perceptron associated with the multimodal feature learner. The multimodal feature learner may include a second bidirectional hidden encoding layer, a second global feature learning layer, and a second local feature learning layer. Specifically, the computer device inputs the first feature extraction vector and the third feature extraction vector into the second bidirectional hidden encoding layer respectively, to obtain a third initial hidden vector corresponding to the third feature extraction vector and a fourth initial hidden vector corresponding to the first feature extraction vector. Further, the computer device may, based on the third initial hidden vector, the fourth initial hidden vector, and the second global feature learning layer, obtain a third global information vector corresponding to the third feature extraction vector and a fourth global information vector corresponding to the first feature extraction vector, and use the third global information vector and the fourth global information vector as multimodal global information vectors. At the same time, the computer device may, based on the third initial hidden vector, the fourth initial hidden vector, and the second local feature learning layer, obtain a third local fine-grained vector corresponding to the third feature extraction vector and a fourth local fine-grained vector corresponding to the first feature extraction vector, and then use the third local fine-grained vector and the fourth local fine-grained vector as multimodal local fine-grained vectors. Further, the computer device may, based on the multimodal global information vectors and the multimodal local fine-grained vectors, obtain a third output vector corresponding to the third feature extraction vector and a fourth output vector corresponding to the first feature extraction vector. At this time, the computer device may input the third output vector into the second multi-layer perceptron to obtain a third learning vector corresponding to the third feature extraction vector, and input the fourth output vector into the second multi-layer perceptron to obtain a fourth learning vector corresponding to the first feature extraction vector, and then use the third learning vector and the fourth learning vector as the second learning result.
[0193] Among them, for the specific implementation manner of the computer device to obtain the third learning vector and the fourth learning vector, reference may be made to the specific implementation manner of obtaining the first learning vector and the second learning vector in step S102 above, and details will not be elaborated here.
[0194] Step S104, the prediction generator splices the learning vectors in the first learning result and the learning vectors in the second learning result to obtain a vector splicing result.
[0195] Among them, the learned vectors in the first learning result may include the first learned vector corresponding to the first feature extraction vector and the second learned vector corresponding to the second feature extraction vector; the learned vectors in the second learning result may include the third learned vector corresponding to the third feature extraction vector and the fourth learned vector corresponding to the first feature extraction vector. It should be understood that the computer device can splice the first learned vector and the fourth learned vector through the multi-modal vector splicing layer in the prediction generator, so as to obtain the first spliced vector corresponding to the search service data. Similarly, the computer device can splice the second learned vector and the third learned vector, so as to obtain the second spliced vector corresponding to the service data to be matched. The vector splicing result here can be used to indicate the prediction of the matching degree between the search service data and the service data to be matched.
[0196] It should be understood that the calculation formulas for the computer device to obtain the spliced vector can be seen in the following formulas (14) and (15):
[0197] l a =[z a ;s a , (14)
[0198] l b =[z b ;s c , (15)
[0199] Among them, z a is the learned vector (i.e., the first learned vector) obtained by the first multi-layer perceptron associated with the text feature learner for the to-be-matched text data a associated with the search service data, s a is the learned vector (i.e., the fourth learned vector) obtained by the second multi-layer perceptron associated with the multi-modal feature learner for the to-be-matched text data a, l a refers to the spliced vector (i.e., the first spliced vector) associated with the search service data. z b is the learned vector (i.e., the second learned vector) obtained by the first multi-layer perceptron associated with the text feature learner for the to-be-matched text data b associated with the first-modal service data, s c is the learned vector (i.e., the third learned vector) obtained by the second multi-layer perceptron associated with the multi-modal feature learner for the video frame associated with the second-modal service data, l b refers to the spliced vector (i.e., the second spliced vector) associated with the service data to be matched.
[0200] Among them, through the text feature learner, the computer device can learn the features between the search service data (e.g., text data 1) and the first-modal service data (e.g., text data 2), and through the multi-modal feature learner, it can learn the features between the search service data and the second-modal service data (e.g., video data 3). Therefore, the computer device can effectively learn multi-modal information, and then can more accurately predict the matching degree between the search service data and the service data to be matched.
[0201] Furthermore, the computer device can input the first splicing vector and the second splicing vector into the third multi-layer perceptron in the prediction generator to predict the matching degree between the search service data and the service data to be matched. Further, the computer device can obtain the prediction result (i.e., the classification result) output by the third multi-layer perceptron. If the matching degree indicated by the classification result is the first matching degree (e.g., 1), the computer device can determine that the search service data matches the service data to be matched successfully. If the matching degree indicated by the classification result is the second matching degree (e.g., 0), the computer device can determine that the search service data fails to match the service data to be matched.
[0202] It should be understood that the classification result g of the computer device for determining the search service data and the service data to be matched can be as shown in the following formula (16):
[0203] g = MLP[l a ; l b , (16)
[0204] where l a refers to the splicing vector associated with the search service data (i.e., the first splicing vector). l b refers to the splicing vector associated with the service data to be matched (i.e., the second splicing vector).
[0205] It can be understood that when the matching degree between the search service data and the service data to be matched indicates that the search service data matches the service data to be matched successfully, the computer device can use the service data to be matched as the service search result corresponding to the service search request, and then can push the service search result to the user terminal. At this time, the user terminal can switch the display interface from the search display interface to the service data display interface and output the service search result to the service data display interface.
[0206] Optionally, the computer device can obtain multiple popular topics in the current popular list offline, and then, through a multimodal matching model, respectively determine popular video data matching these multiple popular topics from the video database. Further, the computer device can generate a popular list based on the popular video data. When the business data predicted by the computer device and matching the search business data (for example, business data X) exists in the popular list, at this time, the computer device can preferentially push the business data X to the user terminal, so that the user terminal can preferentially display the business data X on the business data display interface of the user terminal.
[0207] For ease of understanding, further, please refer to Table 1, which is a popular list provided by an embodiment of the present application. Among them, Table 1 may include a topic identifier, a cluster identifier, a title, a popularity score, a popularity factor, a publication time, an update time, etc., which are not limited herein.
[0208] Table 1
[0209]
[0210] For ease of understanding, further, please refer to Figure 6 , Figure 6 which is a schematic diagram of a scenario for displaying business search results provided by an embodiment of the present application. The computer device in the embodiment of the present application can be, for example, Figure 6 the server 6B shown in the figure, and the server 6B can be the server 10 shown above. Figure 1 The user terminal 6A in the embodiment of the present application can be any one of the user terminal clusters shown above. Figure 1 For example, the user terminal 100a.
[0211] As Figure 6 shown, in the search display interface 600 of the application client running in the user terminal 6A, a popular list composed of currently more popular titles can be displayed. Among them, the popular list may include multiple titles. As Figure 6 shown, the popular list may specifically include Title 1, Title 2, Title 3, and Title 4. It can be understood that the user a corresponding to the user terminal 6A can directly perform a triggering operation on a certain piece of business data of interest (for example, Title 1) in the popular list, and then the user terminal 6A can respond to the triggering operation, so that Title 1 associated with the triggering operation can be used as the search business data to generate a business search request to be sent to the server 6B corresponding to the application client. For example, Title 1 can be "Singers Performing at the New Year's Eve Party".
[0212] When the server 6B receives the service search request, it can search for the service data to be matched that matches the search service data of the title 1. The service data to be matched here can be of text type, video type or picture type. Among them, the service data to be matched in the embodiments of the present application can take the service data of video type as an example. For example, when the server 6B obtains the search service data sent by the user terminal 6A (for example, "Singers performing at the New Year's Eve party"), it can obtain the service data to be matched from the video database, and then can obtain a multi-modal matching model to predict the matching degree between the search service data and the service data to be matched.
[0213] It can be understood that when the matching degree between the search service data and the service data to be matched indicates that the search service data matches the service data to be matched successfully, the server 6B can use the service data to be matched as the service search result corresponding to the service search request. Among them, the service search results determined by the server 6B can include multiple (taking 2 as an example), specifically including service search result 1 (for example, a singing video of a certain singer X at the New Year's Eve party) and service search result 2 (for example, a singing video of a certain singer Y at the New Year's Eve party).
[0214] At this time, the server 6B can push the two service search results, service search result 1 and service search result 2, to the user terminal 6A, so that the user terminal 6A can switch the display interface from the search display interface 600 to the service data display interface (for example, Figure 6 the service data display interface 610 shown), and output these two service search results to the service data display interface 600, so that the user a can conveniently obtain the service data of the video type that he is interested in, thereby improving the user's search experience.
[0215] Further, please refer to Figure 7 , Figure 7 which is a schematic flowchart of a method for training a multi-modal matching model provided by an embodiment of the present application. As Figure 7 shown, this method can be executed by a computer device with a matching degree prediction function. The computer device can be a user terminal (for example, the user terminal 100a shown above Figure 1 ), or it can be a server (for example, the server 10 shown above Figure 1 ), which is not limited here. This method can at least include the following steps S201 - step S205:
[0216] Step S201, obtain a sample data group for training the multi-modal training model.
[0217] Specifically, the computer device can obtain a sample data set for training a multi-modal training model. Among them, the sample data set can include a first type of sample data set and a second type of sample data set. Here, the first type of sample data set can be a sample data set with sample label information; here, the second type of sample data set can be a sample data set without sample label information. The sample label information is used to indicate the matching degree between the first type of sample data sets.
[0218] It should be understood that in order to obtain a better recognition effect, the embodiment of the present application can use the first type of sample data set to initially train the multi-modal training model, so as to obtain the multi-modal training model after the initial training is completed (for example, multi-modal training model 1). Further, the computer device can use the multi-modal training model 1 to predict the second type of sample data set, and use the predicted result as the sample label information of the second type of sample data set. At this time, the computer device can mix the predicted second type of sample data set and the first type of sample data set according to a certain ratio (for example, ratios such as 1:9, 1:4, 3:7, etc.), and use the mixed sample data set as a new training sample data set to retrain the multi-modal training model 1. Repeating this process multiple times can improve the prediction effect of the model.
[0219] Step S202: Input the sample data set into the multi-modal training model, and the multi-modal training model outputs the prediction result between the sample data sets, and use the prediction result as the prediction label information.
[0220] Among them, the multi-modal training model can include a sample feature extractor, a sample feature learner, and a sample prediction generator.
[0221] Step S203: Obtain the sample splicing vector corresponding to the sample data set, and determine the optimal perturbation amount of the sample data set based on the sample splicing vector, the model loss function of the multi-modal training model, and the expected conditions associated with the multi-modal training model.
[0222] Specifically, the computer device can obtain the sample splicing vector corresponding to the sample data set, and obtain the model parameters of the multi-modal training model; furthermore, it can determine the initial perturbation amount corresponding to the sample data set based on the sample splicing vector, the prediction label information, the model parameters, and the model loss function of the multi-modal training model. Further, the computer device can obtain the expected conditions associated with the multi-modal training model. When it detects that there is an initial perturbation amount in the initial perturbation amounts that meets the expected conditions, it uses the initial perturbation amount that meets the expected conditions as the optimal perturbation amount of the sample data set.
[0223] It should be understood that, in order to effectively improve the robustness of the multi-modal training model and the generalization ability of the model, the computer device can improve the model loss function in the multi-modal training model and add an adversarial training process to the adversarial training learning layer. Among them, the key in adversarial training is to find adversarial samples, and the adversarial samples here are usually constructed by adding a certain perturbation to the sample concatenation vector, and then the model is trained, so that the finally obtained multi-modal matching model has the ability to recognize adversarial samples.
[0224] Specifically, the expected conditions involved in the embodiments of the present application can be shown as the following formula (17):
[0225]
[0226] Among them, the formula can be divided into two parts, one is the maximization of the internal model loss function, and the other is the minimization of the external risk. The internal max is to find the most effective perturbation to make the model make mistakes (attack), and the external min is to adapt based on this attack to find the most robust model parameters. Here is the initial perturbation amount obtained by perturbing the multi-modal vector concatenation layer, D represents the sample data group, g is the predicted label information of the sample data group, and L(l o +Δl o , g; θ) refers to the loss function of a single sample data group, Ω is the perturbation space, θ is the model parameter of the multi-modal training model, and l o is the sample concatenation vector of the sample data group. Here E is the mathematical expectation.
[0227] Step S204, based on the optimal perturbation amount and the sample concatenation vector, generate the adversarial sample data corresponding to the sample data group, and based on the adversarial sample data and the model loss function, perform iterative training on the multi-modal training model to obtain the model training result.
[0228] It can be understood that the computer device can adjust the model parameters of the multi-modal training model by using a suitable optimizer. The optimizer can be any one or more optimizers such as the GD optimizer, SGD optimizer, Momentum optimizer, RMSProp optimizer, and Adam optimizer. Among them, the embodiments of the present application can take the Adam optimizer with a relatively fast training speed as an example to update the model parameters of the multi-modal training model, so that the expectation of the entire data distribution is still the smallest. The Adam optimizer mainly acts on the output layer of the model network structure. The Adam optimizer not only has the advantages of simple implementation, high computational efficiency, and low memory requirements, but also is applicable to problems with sparse gradients or large noise in gradients. Among them, the maximum number of steps of the feature learner in the multi-modal training model can be set to 10, and the learning rate can be set to 0.015.
[0229] Step S205, when the model training result indicates that the multi-modal training model after iterative training meets the model convergence condition, the multi-modal training model that meets the model convergence condition is used as the multi-modal matching model for predicting the matching degree between business data groups.
[0230] Specifically, when the model training result indicates that the multi-modal training model after iterative training meets the model convergence condition, the computer device can use the multi-modal training model that meets the model convergence condition as the multi-modal matching model for predicting the matching degree. When the model training result indicates that the model loss function of the multi-modal training model after iterative training does not meet the model convergence condition, the computer device uses the multi-modal training model after iterative training as the multi-modal transition model, and then can adjust the model parameters of the multi-modal transition model based on the model loss function that does not meet the model convergence condition until the adjusted multi-modal transition model meets the model convergence condition. At this time, the computer device can use the multi-modal transition model that meets the model matching condition as the multi-modal matching model for predicting the matching degree.
[0231] Among them, the embodiment of the present application adopts a semi-supervised learning mechanism, mixes the predicted second-type sample data group and the first-type sample data group according to a certain ratio, and uses the mixed sample data group as the new training sample data group to retrain the multi-modal training model. When the training is completed, a multi-modal matching model that meets the model convergence condition can be obtained. The multi-modal matching model obtained in this way can learn more features in a larger sample space, thereby making the model parameters of the multi-modal matching model more robust and effectively improving the model accuracy.
[0232] Further, please refer to Figure 8 , Figure 8 which is a schematic structural diagram of a data processing device provided by an embodiment of the present application. As Figure 8 shown, the data processing device 1 can be a computer program (including program code) running in a computer device. For example, the data processing device 1 is an application software; the data processing device 1 can be used to execute the corresponding steps in the method provided by the embodiment of the present application. As Figure 8 shown, the data processing device 1 can run on a computer device with a matching degree prediction function. The data processing device 1 can include: a model acquisition module 11, a first learning and processing module 12, a second learning and processing module 13, a splicing processing module 14, a request acquisition module 15, a data acquisition module 16, a to-be-matched business data determination module 17, a to-be-encoded text data determination module 18, a vector extraction module 19, a frame extraction processing module 20, a search result determination module 21, and a search result push module 22.
[0233] The model acquisition module 11 is configured to acquire a multimodal matching model for matching search service data and service data to be matched; the multimodal matching model includes a feature learner and a prediction generator; the service data to be matched includes first-modal service data and second-modal service data;
[0234] The first learning and processing module 12 is configured to perform first learning and processing on a first feature extraction vector of the search service data and a second feature extraction vector of the first-modal service data through a text feature learner in the feature learner to obtain a first learning result; the learning vector in the first learning result is obtained from a text global information vector and a text local fine-grained vector; the text global information vector is obtained based on a first multi-scale convolution kernel in a first global feature learning layer of the text feature learner; the text local fine-grained vector is obtained based on a first local feature learning layer of the text feature learner.
[0235] Wherein, the feature learner includes a first multi-layer perceptron associated with the text feature learner; the text feature learner includes a first bidirectional hidden encoding layer, a first global feature learning layer, and a first local feature learning layer;
[0236] The first learning and processing module 12 includes: a text initial vector determination unit 121, a text global vector determination unit 122, a text local vector determination unit 123, a text output vector determination unit 124, and a first learning result determination unit 125.
[0237] The text initial vector determination unit 121 is configured to input the first feature extraction vector of the search service data and the second feature extraction vector of the first-modal service data into the first bidirectional hidden encoding layer respectively to obtain a first initial hidden vector corresponding to the first feature extraction vector and a second initial hidden vector corresponding to the second feature extraction vector;
[0238] The text global vector determination unit 122 is configured to obtain a first global information vector corresponding to the first feature extraction vector and a second global information vector corresponding to the second feature extraction vector based on the first initial hidden vector, the second initial hidden vector, and the first global feature learning layer, and use the first global information vector and the second global information vector as the text global information vector.
[0239] Wherein, the text global vector determination unit 122 includes: an initial hidden vector determination subunit 1221, a convolution kernel acquisition subunit 1222, a convolution feature determination subunit 1223, a pooling feature determination subunit 1224, and a global vector determination subunit 1225.
[0240] The initial hidden vector determination subunit 1221 is configured to use the first initial hidden vector and the second initial hidden vector as the initial hidden vectors corresponding to the text data to be matched respectively; the initial hidden vector is a hidden vector matrix with H rows; H is obtained from the total number of words in the text data to be matched; the hidden vector matrix includes a hidden vector p k ; the hidden vector p k is the hidden vector corresponding to the k-th word obtained by traversing the text data to be matched; k is a positive integer less than or equal to H;
[0241] The convolution kernel acquisition subunit 1222 is configured to input the initial hidden vector into the first global feature learning layer to obtain a first multi-scale convolution kernel associated with the first global feature learning layer; the first multi-scale convolution kernel includes N first-type convolution kernels and (N - 1) second-type convolution kernels; N is a positive integer greater than 1;
[0242] The convolution feature determination subunit 1223 is configured to input the initial hidden vector into N first-type convolution kernels respectively to obtain N first convolution features, and obtain the first-type convolution feature and the second-type convolution feature from the N first convolution features; perform convolution processing on the second-type convolution feature through (N - 1) second-type convolution kernels respectively to obtain (N - 1) second convolution features;
[0243] The pooling feature determination subunit 1224 is configured to input the first-type convolution feature and (N - 1) second convolution features into the average pooling layer to obtain the pooling feature corresponding to the k-th word, until the value of k is H, to obtain the pooling features corresponding to each word in the text data to be matched respectively;
[0244] The global vector determination subunit 1225 is configured to input the pooling features corresponding to each word in the text data to be matched respectively into the connection layer to obtain the text global information vector corresponding to the text data to be matched; the text global information vector includes the first global information vector corresponding to the first feature extraction vector and the second global information vector corresponding to the second feature extraction vector.
[0245] Among them, the specific implementation manners of the initial hidden vector determination subunit 1221, the convolution kernel acquisition subunit 1222, the convolution feature determination subunit 1223, the pooling feature determination subunit 1224, and the global vector determination subunit 1225 can refer to the description of the text global information vector in the corresponding embodiments above Figure 5b and will not be elaborated here.
[0246] The text local vector determination unit 123 is configured to obtain a first local fine-grained vector corresponding to a first feature extraction vector and a second local fine-grained vector corresponding to a second feature extraction vector based on a first initial hidden vector, a second initial hidden vector, and a first local feature learning layer, and use the first local fine-grained vector and the second local fine-grained vector as the text local fine-grained vector.
[0247] Wherein, the first initial hidden vector is a hidden vector matrix with m rows; the second initial hidden vector is a hidden vector matrix with n rows; m is obtained from the total number of words associated with the search service data; n is obtained from the total number of words associated with the first modality service data;
[0248] The text local vector determination unit 123 includes: a hidden vector acquisition subunit 1231, a local weight determination subunit 1232, a first local vector determination subunit 1233, a second local vector determination subunit 1234, and a text local vector determination subunit 1235.
[0249] The hidden vector acquisition subunit 1231 is configured to input the first initial hidden vector and the second initial hidden vector into the first local feature learning layer, and traverse and obtain the hidden vector p corresponding to the i-th word from the first initial hidden vector associated with the service search data ai and the hidden vector p corresponding to the u-th word au , and traverse and obtain the hidden vector p corresponding to the j-th word from the second initial hidden vector associated with the first modality service data bj and the hidden vector p corresponding to the v-th word bv ; both i and u are positive integers less than or equal to m; both j and v are positive integers less than or equal to n;
[0250] The local weight determination subunit 1232 is configured to determine a first local weight e between the hidden vector p ai and the hidden vector p bj , determine a second local weight e between the hidden vector p ij and the hidden vector p ai , determine a third local weight e between the hidden vector p bv and the hidden vector p iv ; au bj uj ;
[0251] The first local vector determination subunit 1233 is configured to determine a first intermediate hidden vector corresponding to the i-th word based on the first local weight e ij , the second local weight e iv , and the hidden vector p bj ; bj Until the value of i is m, m first intermediate hidden vectors are obtained. Based on the m first intermediate hidden vectors, a first local fine-grained vector corresponding to the first feature extraction vector is obtained;
[0252] The second local vector determination subunit 1234 is used to determine, based on the first local weight e ij , the third local weight e uj and the hidden vector p ai , the second intermediate hidden vector corresponding to the j-th word; Until the value of j is n, n second intermediate hidden vectors are obtained. Based on the n second intermediate hidden vectors, a second local fine-grained vector corresponding to the second feature extraction vector is obtained;
[0253] The text local vector determination subunit 1235 is used to use the first local fine-grained vector and the second local fine-grained vector as the text local fine-grained vector.
[0254] Among them, the specific implementation manners of the hidden vector acquisition subunit 1231, the local weight determination subunit 1232, the first local vector determination subunit 1233, the second local vector determination subunit 1234, and the text local vector determination subunit 1235 can refer to the descriptions of the local fine-grained vectors in the corresponding embodiments above Figure 5b and will not be elaborated here.
[0255] The text output vector determination unit 124 is used to obtain a first output vector corresponding to the first feature extraction vector and a second output vector corresponding to the second feature extraction vector based on the text global information vector and the text local fine-grained vector;
[0256] The first learning result determination unit 125 is used to input the first output vector into the first multi-layer perceptron to obtain a first learning vector corresponding to the first feature extraction vector, and input the second output vector into the first multi-layer perceptron to obtain a second learning vector corresponding to the second feature extraction vector, and use the first learning vector and the second learning vector as the first learning result.
[0257] Among them, the specific implementation manners of the text initial vector determination unit 121, the text global vector determination unit 122, the text local vector determination unit 123, the text output vector determination unit 124, and the first learning result determination unit 125 can refer to the descriptions of step S102 in the corresponding embodiments above Figure 3 and will not be elaborated here.
[0258] The second learning and processing module 13 is configured to perform second learning and processing on the first feature extraction vector and the third feature extraction vector of the second-modal service data through the multi-modal feature learner in the feature learner, so as to obtain a second learning result; the learning vector in the second learning result is obtained from the multi-modal global information vector and the multi-modal local fine-grained vector; the multi-modal global information vector is obtained based on the second multi-scale convolution kernel in the second global feature learning layer of the multi-modal feature learner; the multi-modal local fine-grained vector is obtained based on the second local feature learning layer of the multi-modal feature learner.
[0259] Wherein, the feature learner includes a second multi-layer perceptron associated with the multi-modal feature learner; the multi-modal feature learner includes a second bidirectional hidden encoding layer, a second global feature learning layer, and a second local feature learning layer;
[0260] The second learning and processing module 13 includes: a multi-modal initial vector determination unit 131, a multi-modal global vector determination unit 132, a multi-modal local vector determination unit 133, a multi-modal output vector determination unit 134, and a second learning result determination unit 135.
[0261] The multi-modal initial vector determination unit 131 is configured to input the first feature extraction vector and the third feature extraction vector of the second-modal service data into the second bidirectional hidden encoding layer in the multi-modal feature learner respectively, so as to obtain a third initial hidden vector corresponding to the third feature extraction vector and a fourth initial hidden vector corresponding to the first feature extraction vector;
[0262] The multi-modal global vector determination unit 132 is configured to obtain a third global information vector corresponding to the third feature extraction vector and a fourth global information vector corresponding to the first feature extraction vector based on the third initial hidden vector, the fourth initial hidden vector, and the second global feature learning layer, and use the third global information vector and the fourth global information vector as the multi-modal global information vector;
[0263] The multi-modal local vector determination unit 133 is configured to obtain a third local fine-grained vector corresponding to the third feature extraction vector and a fourth local fine-grained vector corresponding to the first feature extraction vector based on the third initial hidden vector, the fourth initial hidden vector, and the second local feature learning layer, and use the third local fine-grained vector and the fourth local fine-grained vector as the multi-modal local fine-grained vector;
[0264] The multi-modal output vector determination unit 134 is configured to obtain a third output vector corresponding to the third feature extraction vector and a fourth output vector corresponding to the first feature extraction vector based on the multi-modal global information vector and the multi-modal local fine-grained vector;
[0265] The second learning result determination unit 135 is configured to input the third output vector into the second multi-layer perceptron to obtain a third learning vector corresponding to the third feature extraction vector, and input the fourth output vector into the second multi-layer perceptron to obtain a fourth learning vector corresponding to the first feature extraction vector, and use the third learning vector and the fourth learning vector as the second learning result.
[0266] Among them, the specific implementation manners of the multi-modal initial vector determination unit 131, the multi-modal global vector determination unit 132, the multi-modal local vector determination unit 133, the multi-modal output vector determination unit 134, and the second learning result determination unit 135 can refer to the description of step S103 in the corresponding embodiment above, and will not be elaborated here. Figure 3 The description of step S103 in the corresponding embodiment above will not be elaborated here.
[0267] The splicing processing module 14 is configured to splice the learning vectors in the first learning result and the learning vectors in the second learning result through a prediction generator to obtain a vector splicing result; the vector splicing result is used to indicate the prediction of the matching degree between the search service data and the service data to be matched.
[0268] Among them, the learning vectors in the first learning result include a first learning vector corresponding to the first feature extraction vector and a second learning vector corresponding to the second feature extraction vector; the learning vectors in the second learning result include a third learning vector corresponding to the third feature extraction vector and a fourth learning vector corresponding to the first feature extraction vector;
[0269] The splicing processing module 14 includes: a splicing processing unit 141 and a splicing result determination unit 142.
[0270] The splicing processing unit 141 is configured to splice the first learning vector and the fourth learning vector through a prediction generator to obtain a first splicing vector, and splice the second learning vector and the third learning vector to obtain a second splicing vector;
[0271] The splicing result determination unit 142 is configured to use the first splicing vector and the second splicing vector as the vector splicing result.
[0272] Among them, the specific implementation manners of the splicing processing unit 141 and the splicing result determination unit 142 can refer to the description of step S104 in the corresponding embodiment above, and will not be elaborated here. Figure 3 The description of step S104 in the corresponding embodiment above will not be elaborated here.
[0273] The request acquisition module 15 is configured to acquire a service search request including search service data sent by a user terminal; the service search request is generated when the user terminal responds to a trigger operation on a search control in an application client; the search service data is acquired by the user terminal from a search area of a search display interface.
[0274] The data acquisition module 16 is configured to, based on the service search request, acquire service data of a first service type from a video database, use the service data of the first service type as first-modal service data, and acquire service data of a second service type from the video database, use the service data of the second service type as second-modal service data; the first service type is different from the second service type.
[0275] The service data to be matched determination module 17 is configured to use the service data commonly mapped by the first-modal service data and the second-modal service data as the service data to be matched.
[0276] Wherein, if the service type of the search service data is the first service type and the first service type belongs to the text type, the second service type includes at least one of the following service types: video type or picture type; the multi-modal matching model includes a feature extractor; the feature extractor includes a word vector extraction network and a residual network.
[0277] The text data to be encoded determination module 18 is configured to use the search service data and the first-modal service data as the text data to be encoded.
[0278] The vector extraction module 19 is configured to extract a feature extraction vector from the text data to be encoded through a word vector extraction network; the feature extraction vector includes a first feature extraction vector extracted from the search service data and a second feature extraction vector extracted from the first-modal service data.
[0279] Wherein, the vector extraction module 19 includes: a preprocessing unit 191, a target word acquisition unit 192, a vector extraction unit 193, and a feature extraction vector determination unit 194.
[0280] The preprocessing unit 191 is configured to preprocess the text data to be encoded, use the preprocessed text data to be encoded as the text data to be matched, perform character segmentation processing on the text data to be matched according to a text vocabulary, and obtain a word information sequence and a word position sequence corresponding to the text data to be matched; the total number of words in the text data to be matched is H; H is a positive integer.
[0281] The target character acquisition unit 192 is configured to traverse and acquire the character information corresponding to the k-th character of the text data to be matched from the character information sequence, use the acquired character information as the target character information, acquire the character position information corresponding to the target character information from the character position sequence, and use the acquired character position information as the target character position information; k is a positive integer less than or equal to H.
[0282] The vector extraction unit 193 is configured to input the target character information into a word vector extraction network, and the word vector extraction network extracts the target character information vector corresponding to the k-th character, input the target character position information into the word vector extraction network, and the word vector extraction network extracts the target character position vector corresponding to the k-th character; the word vector extraction network is trained based on a text vocabulary.
[0283] The feature extraction vector determination unit 194 is configured to obtain the feature extraction vector corresponding to the k-th character based on the target character information vector and the target character position vector, and until the value of k is H, obtain the feature extraction vector corresponding to the text data to be matched.
[0284] Among them, for the specific implementation manners of the preprocessing unit 191, the target character acquisition unit 192, the vector extraction unit 193, and the feature extraction vector determination unit 194, reference can be made to the description of the feature extractor in the corresponding embodiments above. Figure 3 Details will not be elaborated here.
[0285] The frame extraction processing module 20 is configured to perform frame extraction processing on the second-modal service data to obtain video frames, input the video frames into a residual network, and the residual network extracts the third feature extraction vector corresponding to the second-modal service data.
[0286] The search result determination module 21 is configured to use the service data to be matched as the service search result corresponding to the service search request when the matching degree between the search service data and the service data to be matched indicates that the search service data and the service data to be matched are successfully matched.
[0287] The search result pushing module 22 is configured to push the service search result to the user terminal, so that the user terminal switches the display interface from the search display interface to the service data display interface, and outputs the service search result to the service data display interface.
[0288] Among them, for the specific implementation manners of the model acquisition module 11, the first learning processing module 12, the second learning processing module 13, the splicing processing module 14, the request acquisition module 15, the data acquisition module 16, the service data to be matched determination module 17, the text data to be encoded determination module 18, the vector extraction module 19, the frame extraction processing module 20, the search result determination module 21, and the search result pushing module 22, reference can be made to the above Figure 3The descriptions of steps S101 to S104 in the corresponding embodiments will not be elaborated here. In addition, the descriptions of the beneficial effects of using the same method will not be elaborated either.
[0289] Further, please refer to Figure 9 , Figure 9 which is a schematic diagram of a computer device provided by an embodiment of the present application. As Figure 9 shown, the computer device 1000 may be the server 2B in the above-mentioned Figure 2 corresponding embodiment. The computer device 1000 may include: at least one processor 1001, such as a CPU, at least one network interface 1004, a user interface 1003, a memory 1005, and at least one communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. Among them, the user interface 1003 may include a display screen (Display) and a keyboard (Keyboard). The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The memory 1005 may optionally also be at least one storage device located far from the aforementioned processor 1001. As Figure 9 shown, the memory 1005, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.
[0290] In Figure 9 the computer device 1000 shown, the network interface 1004 is mainly used for network communication with a user terminal; the user interface 1003 is mainly used to provide an input interface for a user; and the processor 1001 may be used to call the device control application program stored in the memory 1005 to achieve:
[0291] obtaining a multimodal matching model for matching search service data and service data to be matched; the multimodal matching model includes a feature learner and a prediction generator; the service data to be matched includes first-modal service data and second-modal service data;
[0292] Through the text feature learner in the feature learner, perform a first learning process on the first feature extraction vector of the search service data and the second feature extraction vector of the first-modal service data to obtain a first learning result; the learning vector in the first learning result is obtained from the text global information vector and the text local fine-grained vector; the text global information vector is obtained based on the first multi-scale convolution kernel in the first global feature learning layer of the text feature learner; the text local fine-grained vector is obtained based on the first local feature learning layer of the text feature learner.
[0293] Through the multi-modal feature learner in the feature learner, perform a second learning process on the first feature extraction vector and the third feature extraction vector of the second-modal service data to obtain a second learning result; the learning vector in the second learning result is obtained from the multi-modal global information vector and the multi-modal local fine-grained vector; the multi-modal global information vector is obtained based on the second multi-scale convolution kernel in the second global feature learning layer of the multi-modal feature learner; the multi-modal local fine-grained vector is obtained based on the second local feature learning layer of the multi-modal feature learner.
[0294] The prediction generator splices the learning vector in the first learning result and the learning vector in the second learning result to obtain a vector splicing result; the vector splicing result is used to indicate the prediction of the matching degree between the search service data and the service data to be matched.
[0295] It should be understood that the computer device 1000 described in the embodiments of the present application can execute the description of the data processing method in the corresponding embodiments mentioned above Figure 3 and Figure 7 and can also execute the description of the data processing device 1 in the corresponding embodiments mentioned above Figure 8 which will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either.
[0296] In addition, it should be pointed out here that: the embodiments of the present application also provide a computer-readable storage medium, and the computer-readable storage medium stores the computer program executed by the data processing device 1 mentioned above, and the computer program includes program instructions. When the processor executes the program instructions, it can execute the above Figure 3 or Figure 7The description of the data processing method in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated. For the technical details not disclosed in the embodiment of the computer-readable storage medium involved in this application, please refer to the description of the method embodiment of this application. As an example, the program instructions can be deployed to be executed on one computing device, or on multiple computing devices located at one location, or on multiple computing devices distributed at multiple locations and interconnected by a communication network. The multiple computing devices distributed at multiple locations and interconnected by a communication network can form a blockchain system.
[0297] Further, please refer to Figure 10 , Figure 10 which is a schematic structural diagram of a data processing device provided by an embodiment of the present application. As Figure 10 shown, the data processing device 2 can be a computer program (including program code) running in a computer device. For example, the data processing device 2 is an application software; the data processing device 2 can be used to execute the corresponding steps in the method provided by the embodiment of the present application. As Figure 10 shown, the data processing device 2 can run on a computer device. The data processing device 2 may include: a sample acquisition module 100, a prediction result output module 200, an optimal perturbation amount determination module 300, an iterative training module 400, and a model determination module 500.
[0298] The sample acquisition module 100 is used to acquire a sample data group for training a multi-modal training model; the sample data group includes a first type of sample data group and a second type of sample data group; the first type of sample data group is a sample data group with sample label information; the second type of sample data group is a sample data group without sample label information; the sample label information is used to indicate the matching degree between the first type of sample data groups;
[0299] The prediction result output module 200 is used to input the sample data group into the multi-modal training model, and the multi-modal training model outputs the prediction result between the sample data groups, and uses the prediction result as the prediction label information; the multi-modal training model includes a sample feature extractor, a sample feature learner, and a sample prediction generator;
[0300] The optimal perturbation amount determination module 300 is used to obtain the sample splicing vector corresponding to the sample data group, and determine the optimal perturbation amount of the sample data group based on the sample splicing vector, the model loss function of the multi-modal training model, and the expected conditions associated with the multi-modal training model.
[0301] Among them, the optimal perturbation amount determination module 300 includes: an acquisition unit 3010, an initial perturbation amount determination unit 3020, and an optimal perturbation amount determination unit 3030.
[0302] The obtaining unit 3010 is configured to obtain the sample splicing vector corresponding to the sample data group and obtain the model parameters of the multi-modal training model.
[0303] The initial perturbation amount determining unit 3020 is configured to determine the initial perturbation amount corresponding to the sample data group based on the sample splicing vector, the predicted label information, the model parameters, and the model loss function of the multi-modal training model.
[0304] The optimal perturbation amount determining unit 3030 is configured to obtain the expected conditions associated with the multi-modal training model, and when it detects that there is an initial perturbation amount in the initial perturbation amounts that satisfies the expected conditions, use the initial perturbation amount that satisfies the expected conditions as the optimal perturbation amount of the sample data group.
[0305] Among them, the specific implementation manners of the obtaining unit 3010, the initial perturbation amount determining unit 3020, and the optimal perturbation amount determining unit 3030 can refer to the description of step S203 in the corresponding embodiment above, and will not be elaborated here. Figure 7 The description of the corresponding embodiment of step S203 will not be continued here.
[0306] The iterative training module 400 is configured to generate adversarial sample data corresponding to the sample data group based on the optimal perturbation amount and the sample splicing vector, and perform iterative training on the multi-modal training model based on the adversarial sample data and the model loss function to obtain the model training result.
[0307] The model determining module 500 is configured to, when the model training result indicates that the multi-modal training model after iterative training meets the model convergence condition, use the multi-modal training model that meets the model convergence condition as the multi-modal matching model for predicting the matching degree between business data groups.
[0308] Among them, the specific implementation manners of the sample obtaining module 100, the prediction result output module 200, the optimal perturbation amount determining module 300, the iterative training module 400, and the model determining module 500 can refer to the descriptions of steps S201 - S205 in the corresponding embodiments above, and will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either. Figure 7 The description of the corresponding embodiments of steps S201 - S205 will not be continued here. In addition, the description of the beneficial effects of using the same method will not be elaborated either.
[0309] Further, please refer to Figure 11 , Figure 11 which is a schematic diagram of a computer device provided by an embodiment of the present application. As Figure 11The computer device 3000 shown may include: at least one processor 3001, such as a CPU, at least one network interface 3004, a user interface 3003, a memory 3005, and at least one communication bus 3002. Among them, the communication bus 3002 is used to realize the connection and communication between these components. Among them, the network interface 3004 may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface). The memory 3005 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The memory 3005 may optionally also be at least one storage device located far from the aforementioned processor 3001. As Figure 11 shown, the memory 3005, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.
[0310] In Figure 11 the computer device 3000 shown, the network interface 3004 is mainly used to provide network communication functions; the user interface 3003 is mainly used to provide an input interface for users; and the processor 3001 may be used to call the device control application program stored in the memory 3005 to implement:
[0311] Obtain a sample data group for training a multi-modal training model; the sample data group includes a first type of sample data group and a second type of sample data group; the first type of sample data group is a sample data group with sample label information; the second type of sample data group is a sample data group without sample label information; the sample label information is used to indicate the matching degree between the first type of sample data groups;
[0312] Input the sample data group into the multi-modal training model, and the multi-modal training model outputs a prediction result between the sample data groups, and use the prediction result as prediction label information; the multi-modal training model includes a sample feature extractor, a sample feature learner, and a sample prediction generator;
[0313] Obtain a sample splicing vector corresponding to the sample data group, and determine the optimal perturbation amount of the sample data group based on the sample splicing vector, the model loss function of the multi-modal training model, and the expected conditions associated with the multi-modal training model;
[0314] Generate adversarial sample data corresponding to the sample data group based on the optimal perturbation amount and the sample splicing vector, and perform iterative training on the multi-modal training model based on the adversarial sample data and the model loss function to obtain a model training result;
[0315] When the model training result indicates that the multi-modal training model after iterative training meets the model convergence condition, the multi-modal training model that meets the model convergence condition is used as the multi-modal matching model for predicting the matching degree between business data groups.
[0316] It should be understood that the computer device 3000 described in the embodiments of the present application can execute the description of the data processing method in the corresponding embodiments mentioned above, and can also execute the description of the data processing device 2 in the corresponding embodiments mentioned above, which will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either. Figure 7 The description of the data processing method in the corresponding embodiments mentioned above, and can also execute the description of the data processing device 2 in the corresponding embodiments mentioned above, which will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either. Figure 10 It should be understood that the computer device 3000 described in the embodiments of the present application can execute the description of the data processing method in the corresponding embodiments mentioned above, and can also execute the description of the data processing device 2 in the corresponding embodiments mentioned above, which will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either.
[0317] In addition, it should be pointed out here that: the embodiments of the present application also provide a computer-readable storage medium, and the computer-readable storage medium stores the computer program executed by the computer device 3000 mentioned above, and the computer program includes program instructions. When the above-mentioned processor executes the above-mentioned program instructions, it can execute the description of the above-mentioned data processing method in the corresponding embodiments mentioned above. Therefore, it will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in the present application, please refer to the description of the method embodiments of the present application. Figure 7 The description of the data processing method in the corresponding embodiments mentioned above, and can also execute the description of the data processing device 2 in the corresponding embodiments mentioned above, which will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in the present application, please refer to the description of the method embodiments of the present application.
[0318] On the one hand, the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device can execute the description of the data processing method in the corresponding embodiments mentioned above. Figure 3 Or Figure 7 The description of the data processing method in the corresponding embodiments mentioned above, which will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either.
[0319] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The above program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above methods. Among them, the above storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0320] The above disclosure is only for the preferred embodiments of the present application. Of course, it cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. A data processing method, characterized in that, Including: Obtain a multi-modal matching model for matching search service data and service data to be matched; the multi-modal matching model includes a feature learner and a prediction generator; the service data to be matched includes first-modal service data and second-modal service data; Through the text feature learner in the feature learner, perform a first learning process on the first feature extraction vector of the search service data and the second feature extraction vector of the first-modal service data to obtain a first learning result; the learning vector in the first learning result is obtained from a text global information vector and a text local fine-grained vector; the text global information vector is obtained based on the first multi-scale convolution kernel in the first global feature learning layer of the text feature learner; the text local fine-grained vector is obtained based on the first local feature learning layer of the text feature learner; the learning vector in the first learning result includes a first learning vector corresponding to the first feature extraction vector and a second learning vector corresponding to the second feature extraction vector; Through the multi-modal feature learner in the feature learner, perform a second learning process on the first feature extraction vector and the third feature extraction vector of the second-modal service data to obtain a second learning result; the learning vector in the second learning result is obtained from a multi-modal global information vector and a multi-modal local fine-grained vector; the multi-modal global information vector is obtained based on the second multi-scale convolution kernel in the second global feature learning layer of the multi-modal feature learner; the multi-modal local fine-grained vector is obtained based on the second local feature learning layer of the multi-modal feature learner; the learning vector in the second learning result includes a third learning vector corresponding to the third feature extraction vector and a fourth learning vector corresponding to the first feature extraction vector; Through the prediction generator, splice the first learning vector and the fourth learning vector to obtain a first spliced vector, and splice the second learning vector and the third learning vector to obtain a second spliced vector; Use the first spliced vector and the second spliced vector as the vector splicing result; the vector splicing result is used to indicate the prediction of the matching degree between the search service data and the service data to be matched.
2. The method according to claim 1, characterized in that, The method further includes: Obtain a service search request including search service data sent by a user terminal; the service search request is generated when the user terminal responds to a trigger operation on a search control in an application client; the search service data is obtained by the user terminal from the search area of the search display interface; Based on the business search request, obtain business data with a first business type from the video database, use the business data with the first business type as the first-modal business data, and obtain business data with a second business type from the video database, use the business data with the second business type as the second-modal business data; the first business type is different from the second business type; Use the business data jointly mapped by the first-modal business data and the second-modal business data as the business data to be matched.
3. The method according to claim 2, characterized in that, If the business type of the searched business data is the first business type and the first business type belongs to the text type, then the second business type includes at least one of the following business types: video type or picture type; The multi-modal matching model includes a feature extractor; The feature extractor includes a word vector extraction network and a residual network; The method further includes: Use the searched business data and the first-modal business data as the text data to be encoded; Through the word vector extraction network, extract a feature extraction vector from the text data to be encoded; the feature extraction vector includes a first feature extraction vector extracted from the searched business data and a second feature extraction vector extracted from the first-modal business data; Perform frame extraction on the second-modal business data to obtain video frames, input the video frames into the residual network, and extract a third feature extraction vector corresponding to the second-modal business data by the residual network.
4. The method according to claim 3, wherein The step of extracting a feature extraction vector from the text data to be encoded through the word vector extraction network includes: Preprocess the text data to be encoded, use the preprocessed text data to be encoded as the text data to be matched, perform character segmentation processing on the text data to be matched according to the text vocabulary, and obtain a character information sequence and a character position sequence corresponding to the text data to be matched; the total number of characters in the text data to be matched is H; H is a positive integer; Traverse and obtain the character information corresponding to the k-th character of the text data to be matched from the character information sequence, use the obtained character information as the target character information, obtain the character position information corresponding to the target character information from the character position sequence, and use the obtained character position information as the target character position information; k is a positive integer less than or equal to H; Input the target character information into the word vector extraction network, and extract a target character information vector corresponding to the k-th character by the word vector extraction network. Input the target character position information into the word vector extraction network, and extract a target character position vector corresponding to the k-th character by the word vector extraction network; the word vector extraction network is trained based on the text vocabulary; Based on the target character information vector and the target character position vector, obtain a feature extraction vector corresponding to the k-th character. When the value of k is H, obtain the feature extraction vector corresponding to the text data to be matched.
5. The method according to claim 1, wherein The feature learner includes a first multi-layer perceptron associated with the text feature learner; the text feature learner includes a first bidirectional hidden encoding layer, a first global feature learning layer, and a first local feature learning layer; Performing a first learning process on the first feature extraction vector of the search service data and the second feature extraction vector of the first-modal service data through the text feature learner in the feature learner to obtain a first learning result, including: Respectively inputting the first feature extraction vector of the search service data and the second feature extraction vector of the first-modal service data into the first bidirectional hidden encoding layer to obtain a first initial hidden vector corresponding to the first feature extraction vector and a second initial hidden vector corresponding to the second feature extraction vector; Based on the first initial hidden vector, the second initial hidden vector, and the first global feature learning layer, obtaining a first global information vector corresponding to the first feature extraction vector and a second global information vector corresponding to the second feature extraction vector, and using the first global information vector and the second global information vector as text global information vectors; Based on the first initial hidden vector, the second initial hidden vector, and the first local feature learning layer, obtaining a first local fine-grained vector corresponding to the first feature extraction vector and a second local fine-grained vector corresponding to the second feature extraction vector, and using the first local fine-grained vector and the second local fine-grained vector as text local fine-grained vectors; Based on the text global information vectors and the text local fine-grained vectors, obtaining a first output vector corresponding to the first feature extraction vector and a second output vector corresponding to the second feature extraction vector; Inputting the first output vector into the first multi-layer perceptron to obtain a first learning vector corresponding to the first feature extraction vector, and inputting the second output vector into the first multi-layer perceptron to obtain a second learning vector corresponding to the second feature extraction vector, and using the first learning vector and the second learning vector as the first learning result.
6. The method according to claim 5, wherein The step of obtaining a first global information vector corresponding to the first feature extraction vector and a second global information vector corresponding to the second feature extraction vector based on the first initial hidden vector, the second initial hidden vector, and the first global feature learning layer, and using the first global information vector and the second global information vector as text global information vectors, includes: Use the first initial hidden vector and the second initial hidden vector as the initial hidden vectors corresponding to the text data to be matched respectively; the initial hidden vector is a hidden vector matrix with H rows; H is obtained from the total number of words in the text data to be matched; the hidden vector matrix includes a hidden vector p k ; the hidden vector p k is the hidden vector corresponding to the k-th word obtained by traversing the text data to be matched; k is a positive integer less than or equal to H; Inputting the initial hidden vector into the first global feature learning layer to obtain a first multi-scale convolution kernel associated with the first global feature learning layer; the first multi-scale convolution kernel includes N first-type convolution kernels and (N - 1) second-type convolution kernels; N is a positive integer greater than 1; The initial hidden vector is respectively input into the N first-type convolutional kernels to obtain N first convolutional features, and a first-type convolutional feature and a second-type convolutional feature are obtained from the N first convolutional features; the second-type convolutional features are respectively subjected to convolutional processing by the (N - 1) second-type convolutional kernels to obtain (N - 1) second convolutional features; The first-type convolutional feature and the (N - 1) second convolutional features are input into an average pooling layer to obtain the pooling feature corresponding to the k-th word. Until the value of k is H, the pooling features corresponding to each word in the text data to be matched are obtained; The pooling features corresponding to each word in the text data to be matched are input into a connection layer to obtain a text global information vector corresponding to the text data to be matched; the text global information vector includes a first global information vector corresponding to the first feature extraction vector and a second global information vector corresponding to the second feature extraction vector.
7. The method according to claim 5, wherein The first initial hidden vector is a hidden vector matrix with m rows; the second initial hidden vector is a hidden vector matrix with n rows; m is obtained from the total number of words associated with the search service data; n is obtained from the total number of words associated with the first-modal service data; Based on the first initial hidden vector, the second initial hidden vector, and the first local feature learning layer, obtaining a first local fine-grained vector corresponding to the first feature extraction vector and a second local fine-grained vector corresponding to the second feature extraction vector, and using the first local fine-grained vector and the second local fine-grained vector as text local fine-grained vectors, includes: Input the first initial hidden vector and the second initial hidden vector into the first local feature learning layer, and traverse to obtain the hidden vector p corresponding to the i-th word from the first initial hidden vector associated with the search service data ai and the hidden vector p corresponding to the u-th word au , and traverse to obtain the hidden vector p corresponding to the j-th word from the second initial hidden vector associated with the first-modal service data bj and the hidden vector p corresponding to the v-th word bv ; both the i and the u are positive integers less than or equal to m; both the j and the v are positive integers less than or equal to n; Determine the hidden vector p ai With the hidden vector p bj The first local weight e between them ij To determine the hidden vector p ai With the hidden vector p bv The second local weight e between them iv To determine the hidden vector p au With the hidden vector p bj The third local weight e between them uj ; Based on the first local weight e ij , the second local weight e iv and the hidden vector p bj , determine the first intermediate hidden vector corresponding to the i-th word Until the value of i is m, m first intermediate hidden vectors are obtained. Based on the m first intermediate hidden vectors, the first local fine-grained vector corresponding to the first feature extraction vector is obtained; Based on the first local weight e ij , the third local weight e uj and the hidden vector p ai , determine the second intermediate hidden vector corresponding to the j-th word Until the value of j is n, n second intermediate hidden vectors are obtained. Based on the n second intermediate hidden vectors, the second local fine-grained vector corresponding to the second feature extraction vector is obtained; Using the first local fine-grained vector and the second local fine-grained vector as text local fine-grained vectors.
8. The method according to claim 1, wherein The feature learner includes a second multi-layer perceptron associated with the multi-modal feature learner; the multi-modal feature learner includes a second bidirectional hidden encoding layer, a second global feature learning layer, and a second local feature learning layer; The second learning process is performed on the first feature extraction vector and the third feature extraction vector of the second-modal service data through the multi-modal feature learner in the feature learner to obtain a second learning result, including: The first feature extraction vector and the third feature extraction vector are respectively input into the second bidirectional hidden encoding layer to obtain a third initial hidden vector corresponding to the third feature extraction vector of the second-modal service data and a fourth initial hidden vector corresponding to the first feature extraction vector; Based on the third initial hidden vector, the fourth initial hidden vector, and the second global feature learning layer, obtaining a third global information vector corresponding to the third feature extraction vector and a fourth global information vector corresponding to the first feature extraction vector, and using the third global information vector and the fourth global information vector as multi-modal global information vectors; Based on the third initial hidden vector, the fourth initial hidden vector, and the second local feature learning layer, obtain the third local fine-grained vector corresponding to the third feature extraction vector and the fourth local fine-grained vector corresponding to the first feature extraction vector, and use the third local fine-grained vector and the fourth local fine-grained vector as the multi-modal local fine-grained vectors; Based on the multi-modal global information vector and the multi-modal local fine-grained vectors, obtain the third output vector corresponding to the third feature extraction vector and the fourth output vector corresponding to the first feature extraction vector; Input the third output vector into the second multi-layer perceptron to obtain the third learning vector corresponding to the third feature extraction vector, and input the fourth output vector into the second multi-layer perceptron to obtain the fourth learning vector corresponding to the first feature extraction vector. Use the third learning vector and the fourth learning vector as the second learning result.
9. A data processing method, characterized in that, The method is executed by a computer device with a matching degree prediction function, and the method includes: Obtain a sample data group for training a multi-modal training model; the sample data group includes a first type of sample data group and a second type of sample data group; the first type of sample data group is a sample data group with sample label information; the second type of sample data group is a sample data group without sample label information; the sample label information is used to indicate the matching degree between the first type of sample data groups; Input the sample data group into the multi-modal training model, and the multi-modal training model outputs a prediction result between the sample data groups, and use the prediction result as prediction label information; the multi-modal training model includes a sample feature extractor, a sample feature learner, and a sample prediction generator; the sample feature learner is used to learn the feature extraction vectors extracted by the sample feature extractor to obtain learning vectors associated with the sample data group, and the sample prediction generator is used to perform vector splicing on the learning vectors associated with the sample data group to obtain a sample splicing vector corresponding to the sample data group; Obtain the sample splicing vector corresponding to the sample data group, and determine the optimal perturbation amount of the sample data group based on the sample splicing vector, the model loss function of the multi-modal training model associated with the prediction label information, and the expected conditions associated with the multi-modal training model; Generate adversarial sample data corresponding to the sample data group based on the optimal perturbation amount and the sample splicing vector, and perform iterative training on the multi-modal training model based on the adversarial sample data and the model loss function to obtain a model training result; When the model training result indicates that the multi-modal training model after iterative training meets the model convergence condition, the multi-modal training model that meets the model convergence condition is used as a multi-modal matching model for predicting the matching degree between business data groups. The business data groups include search business data and to-be-matched business data in a search scenario. The computer device loads the multi-modal matching model and predicts the matching degree between the search business data and the to-be-matched business data through the multi-modal matching model.
10. The method according to claim 9, wherein The obtaining of the sample splicing vector corresponding to the sample data group, and determining the optimal perturbation amount of the sample data group based on the sample splicing vector, the model loss function of the multi-modal training model associated with the prediction label information, and the expected conditions associated with the multi-modal training model includes: Obtaining the sample splicing vector corresponding to the sample data group, and obtaining the model parameters of the multi-modal training model; Based on the sample splicing vector, the prediction label information, the model parameters, and the model loss function of the multi-modal training model, determining the initial perturbation amount corresponding to the sample data group; Obtaining the expected conditions associated with the multi-modal training model. When it is detected that there is an initial perturbation amount in the initial perturbation amounts that meets the expected conditions, the initial perturbation amount that meets the expected conditions is used as the optimal perturbation amount of the sample data group.
11. A data processing device, characterized in that, Including: A model acquisition module, configured to acquire a multi-modal matching model for matching search business data and to-be-matched business data; the multi-modal matching model includes a feature learner and a prediction generator; the to-be-matched business data includes first-modal business data and second-modal business data; A first learning processing module, configured to perform first learning processing on the first feature extraction vector of the search business data and the second feature extraction vector of the first-modal business data through the text feature learner in the feature learner, to obtain a first learning result; the learning vector in the first learning result is obtained from a text global information vector and a text local fine-grained vector; the text global information vector is obtained based on the first multi-scale convolution kernel in the first global feature learning layer of the text feature learner; the text local fine-grained vector is obtained based on the first local feature learning layer of the text feature learner; the learning vector in the first learning result includes the first learning vector corresponding to the first feature extraction vector and the second learning vector corresponding to the second feature extraction vector; A second learning processing module, configured to perform second learning processing on the first feature extraction vector and the third feature extraction vector of the second-modal service data through the multi-modal feature learner in the feature learner, to obtain a second learning result; the learning vector in the second learning result is obtained from a multi-modal global information vector and a multi-modal local fine-grained vector; the multi-modal global information vector is obtained based on a second multi-scale convolution kernel in the second global feature learning layer of the multi-modal feature learner; the multi-modal local fine-grained vector is obtained based on the second local feature learning layer of the multi-modal feature learner; the learning vector in the second learning result includes a third learning vector corresponding to the third feature extraction vector and a fourth learning vector corresponding to the first feature extraction vector; A splicing processing module, configured to splice the first learning vector and the fourth learning vector through the prediction generator to obtain a first spliced vector, and splice the second learning vector and the third learning vector to obtain a second spliced vector; The splicing processing module is further configured to use the first spliced vector and the second spliced vector as a vector splicing result; the vector splicing result is used to indicate a prediction of the matching degree between the search service data and the to-be-matched service data.
12. A data processing device, characterized in that, The device runs on a computer device with a matching degree prediction function, and the device includes: A sample acquisition module, configured to acquire a sample data group for training a multi-modal training model; the sample data group includes a first type of sample data group and a second type of sample data group; the first type of sample data group is a sample data group with sample label information; the second type of sample data group is a sample data group without sample label information; the sample label information is used to indicate the matching degree between the first type of sample data groups; A prediction result output module, configured to input the sample data group into the multi-modal training model, and the multi-modal training model outputs a prediction result between the sample data groups, and use the prediction result as prediction label information; the multi-modal training model includes a sample feature extractor, a sample feature learner, and a sample prediction generator; the sample feature learner is configured to learn the feature extraction vector extracted by the sample feature extractor to obtain a learning vector associated with the sample data group, and the sample prediction generator is configured to perform vector splicing on the learning vector associated with the sample data group to obtain a sample spliced vector corresponding to the sample data group; An optimal perturbation amount determination module, configured to acquire the sample spliced vector corresponding to the sample data group, and determine the optimal perturbation amount of the sample data group based on the sample spliced vector, the model loss function of the multi-modal training model associated with the prediction label information, and the expected conditions associated with the multi-modal training model; An iterative training module, configured to generate adversarial sample data corresponding to the sample data group based on the optimal perturbation amount and the sample splicing vector, and iteratively train the multimodal training model based on the adversarial sample data and the model loss function to obtain a model training result; A model determination module, configured to, when the model training result indicates that the multimodal training model after iterative training meets the model convergence condition, use the multimodal training model that meets the model convergence condition as a multimodal matching model for predicting the matching degree between business data groups, where the business data group includes search business data and to-be-matched business data in a search scenario, the computer device loads the multimodal matching model, and predicts the matching degree between the search business data and the to-be-matched business data through the multimodal matching model.
13. A computer device, characterized in that, Comprising: A processor, a memory, and a network interface; The processor is connected to the memory and the network interface. Among them, the network interface is used to provide data communication functions, the memory is used to store computer programs, and the processor is used to call the computer programs to execute the method according to any one of claims 1-10.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the method according to any one of claims 1-10 is executed.
15. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1-10 is implemented.
Citation Information
Patent Citations
Cross-media retrieval method and system
CN111026887A
Text-to-video cross-modal retrieval method based on multistage coding
CN111309971A