Data processing method, device and equipment and readable storage medium

Through the feature routing layer, the cross-connect weight of the feature interaction layer is automatically calculated and the relevant feature vectors are selected, which solves the problem that the feature cross-connection method in the prior art cannot adapt to different data distributions, and achieves higher model accuracy and adaptability.

CN120296395APending Publication Date: 2025-07-11TENCENT TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510458725.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing feature cross-connection method cannot automatically select the most important features based on different input data, resulting in poor performance and lack of flexibility and accuracy when facing different data distribution and task requirements.

Method used

The feature routing layer automatically calculates the features of the cross-connection required by the feature interaction layer, generates the cross-connection weight, selects feature vectors with high correlation, and outputs the service prediction results through the feature interaction layer to achieve automatic feature cross-connection.

Benefits of technology

It improves the accuracy and adaptability of the model, can accurately select features under different data distribution and task requirements, capture user personalized needs, reduce model complexity, and improve training and reasoning efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296395A_ABST
    Figure CN120296395A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method, device and equipment and a readable storage medium. The method comprises the following steps: inputting M object business data into a business cross-connection model; coding the M object business data through a coding layer to obtain M unit domain feature vectors, and inputting the M unit domain feature vectors into a feature interaction layer connected with the coding layer; inputting the M unit domain feature vectors into a feature routing layer, and generating a cross-connection weight between each unit domain feature vector and a feature interaction layer through the feature routing layer; and generating a cross-connection feature vector according to the cross-connection weight, inputting the cross-connection feature vector into a feature interaction layer connected with the feature routing layer, and outputting a feature vector for generating a service prediction result associated with the M pieces of object service data through the feature interaction layer. With the adoption of the method and the device, the cross-linked features of the feature interaction layer can be accurately selected through the feature routing layer, so that the accuracy of a service estimation result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a data processing method, device, equipment and readable storage medium. Background Art

[0002] In the application of the model, in order to enhance the importance of certain features, it is often necessary to directly transfer and connect the features of a certain layer to a subsequent layer. This mechanism is also called feature cross-connection. Feature cross-connection enhances the model's ability to utilize features by retaining and reusing low-level features, thereby improving the model's expressiveness and performance.

[0003] The existing feature cross-linking method is manual feature cross-linking, which requires pre-determining the cross-linking path between network layers and the cross-linking features. When faced with different input data, manual cross-linking cannot determine the most important features in different input data, and can only truncate some features according to the pre-set rules, resulting in feature loss. Moreover, the cross-linking features are pre-selected. When faced with different data distributions and task requirements, it is impossible to accurately select the features that need to be cross-linked, resulting in poor model performance, lack of flexibility, and low accuracy. Summary of the invention

[0004] The embodiments of the present application provide a data processing method, apparatus, device and readable storage medium, which can accurately select the features spanned by the feature interaction layer through the feature routing layer, thereby improving the accuracy of the business prediction results.

[0005] On the one hand, an embodiment of the present application provides a data processing method, including:

[0006] Obtain M object business data, and input the M object business data into the business cross-connection model; the business fields corresponding to the M object business data are different from each other; the business cross-connection model includes a coding layer, a feature interaction layer and a feature routing layer; M is a positive integer;

[0007] The M object business data are respectively encoded through the encoding layer to obtain M unit domain feature vectors, and the M unit domain feature vectors are input into the feature interaction layer connected to the encoding layer;

[0008] The M unit domain feature vectors are input into the feature routing layer, and the cross-connection weights between each unit domain feature vector and the feature interaction layer are generated through the feature routing layer; the cross-connection weights are used to characterize the correlation between the unit domain feature vector and the feature interaction layer;

[0009] Select N unit domain feature vectors from the M unit domain feature vectors according to the cross - connection weights, generate a cross - connection feature vector based on the N unit domain feature vectors, input the cross - connection feature vector into a feature interaction layer connected to the feature routing layer, and output, through the feature interaction layer, a feature vector for generating a service prediction result associated with the M object service data; N is a positive integer less than or equal to M.

[0010] Among them, input the M unit domain feature vectors into the feature routing layer, and generate the cross - connection weights between each unit domain feature vector and the feature interaction layer through the feature routing layer, including:

[0011] Input the M unit domain feature vectors into the feature routing layer, perform dimensionality reduction processing on the M unit domain feature vectors to obtain M reduced - dimensional feature vectors; the vector dimensions of the M reduced - dimensional feature vectors are the same;

[0012] Obtain a routing input matrix and a routing output matrix associated with the feature interaction layer in the feature routing layer, perform dot - product operations on the M reduced - dimensional feature vectors and the routing input matrix respectively to obtain M routing feature vectors; the routing input matrix and the routing output matrix are jointly used to characterize the network layer attributes of the feature interaction layer;

[0013] Generate linear activation vectors corresponding to the M routing feature vectors respectively through a rectified linear activation function, and perform dot - product operations on the M linear activation vectors and the routing output matrix respectively to obtain M linear activation parameters;

[0014] Perform piece - wise parameter mapping on the M linear activation parameters through a piece - wise activation function to obtain the cross - connection weights between each unit domain feature vector and the feature interaction layer.

[0015] Among them, performing dimensionality reduction processing on the M unit domain feature vectors to obtain M reduced - dimensional feature vectors includes:

[0016] Obtain the feature dimension values corresponding to the M unit domain feature vectors respectively, and based on the M feature dimension values, obtain the projection parameter matrices corresponding to the M unit domain feature vectors respectively;

[0017] Perform projection processing on the corresponding unit domain feature vectors through the M projection parameter matrices respectively to obtain M reduced - dimensional feature vectors.

[0018] Among them, the cross - connection weights include a first weight value, a second weight value, and a third weight value; performing piece - wise parameter mapping on the M linear activation parameters through a piece - wise activation function to obtain the cross - connection weights between each unit domain feature vector and the feature interaction layer includes:

[0019] Obtain a first activation threshold and a second activation threshold corresponding to the piece - wise activation function;

[0020] If the linear activation parameter corresponding to the unit domain eigenvector is less than the first activation threshold, set the cross-connection weight corresponding to the unit domain eigenvector to the first weight value;

[0021] If the linear activation parameter corresponding to the unit domain eigenvector is greater than the first activation threshold and less than the second activation threshold, generate a second weight value based on the linear activation parameter, and set the cross-connection weight corresponding to the unit domain eigenvector to the second weight value;

[0022] If the linear activation parameter corresponding to the unit domain eigenvector is greater than the second activation threshold, set the cross-connection weight corresponding to the unit domain eigenvector to the third weight value;

[0023] Among them, the degrees of correlation between the unit domain eigenvectors represented by the first weight value, the second weight value, and the third weight value and the feature interaction layer are different from each other; the degree of correlation corresponding to the first weight value is less than the degree of correlation corresponding to the second weight value, and the degree of correlation corresponding to the second weight value is less than the degree of correlation corresponding to the third weight value.

[0024] Among them, the cross-connection weight includes the weight values corresponding to M unit domain eigenvectors respectively; selecting N unit domain eigenvectors from the M unit domain eigenvectors according to the cross-connection weight includes:

[0025] Determine the N unit domain eigenvectors as the unit domain eigenvectors with cross-connection weights greater than the weight threshold among the M unit domain eigenvectors;

[0026] Generating a cross-connection eigenvector according to the N unit domain eigenvectors includes:

[0027] Generate a mask weight matrix according to the weight values corresponding to the N unit domain eigenvectors respectively, and obtain the reduced-dimensional eigenvectors corresponding to the N unit domain eigenvectors respectively; the reduced-dimensional eigenvector is obtained by performing dimensionality reduction processing on the unit domain eigenvector;

[0028] Perform a dot product operation on the sequence composed of the N reduced-dimensional eigenvectors and the mask weight matrix to obtain a cross-connection eigenvector.

[0029] Among them, the number of feature interaction layers and feature routing layers in the service cross-connection model is both T. Each feature interaction layer is respectively connected to a feature routing layer, and the feature routing layers connected to each feature interaction layer are different from each other. The T feature interaction layers are connected in series; the input of each feature routing layer is M unit domain eigenvectors, and the output of each feature routing layer is used as the input to the connected feature interaction layer; the T feature interaction layers include feature interaction layer H i and feature interaction layer H i-1 , where i is a positive integer less than T; the T feature routing layers include the feature routing layer B connected to the feature interaction layer H i i; Output a feature vector for generating a service prediction result associated with the service data of M objects through the feature interaction layer, including:

[0030] In the feature interaction layer H i Among them, based on the target feature vector corresponding to the feature interaction layer H i And the cross-connection feature vector input by the feature routing layer B i Output a feature vector; if the feature interaction layer H i Is the feature interaction layer connected to the encoding layer, the target feature vector is M unit domain feature vectors; if the feature interaction layer H i Is connected to the feature interaction layer H i-1 Then the target feature vector is the feature vector output by the feature interaction layer H i-1 ;

[0031] The method further includes:

[0032] If the feature interaction layer H i Is the last feature interaction layer among the T feature interaction layers, then generate a service prediction result associated with the service data of M objects according to the feature vector output by the feature interaction layer H i ;

[0033] Among them, the service cross-connection model includes a gating network layer and S expert network layers, S is a positive integer, and each expert network layer includes a candidate feature interaction layer and a candidate feature routing layer, and the service domains corresponding to the S expert network layers are different from each other; before inputting the M unit domain feature vectors into the feature routing layer, the method further includes:

[0034] Input the M unit domain feature vectors into the gating network layer, and generate gating activation values respectively corresponding to each expert network layer through the gating network layer; the gating activation values are used to represent the correlation degree between the expert network layer and the M unit domain feature vectors;

[0035] Based on the S gating activation values, sort the S expert network layers, determine the target expert network layer among the sorted S expert network layers, determine the candidate feature interaction layer in the target expert network layer as the feature interaction layer, and determine the candidate feature routing layer in the target expert network layer as the feature routing layer.

[0036] Among them, outputting a feature vector for generating a service prediction result associated with the service data of M objects through the feature interaction layer includes:

[0037] In the feature interaction layer, obtain the media content feature vectors respectively corresponding to Q media contents in the media content database, and perform vector splicing on the Q media content feature vectors to obtain a media content feature sequence; Q is a positive integer;

[0038] Fuse the cross-connected feature vector and the M unit domain feature vectors to obtain a fused feature vector, and perform attention processing on the media content feature sequence and the fused feature vector to obtain a feature vector containing the attention result vector;

[0039] The method further includes:

[0040] Based on the attention scores corresponding to the Q media contents in the attention result vector, generate the estimated interaction parameters corresponding to the Q media contents respectively, and determine the estimated interaction parameters corresponding to the Q media contents respectively as the service prediction results associated with the M object service data.

[0041] Among them, performing attention processing on the media content feature sequence and the fused feature vector to obtain a feature vector containing the attention result vector includes:

[0042] Obtain the query parameter matrix, key parameter matrix, and value parameter matrix in the feature interaction layer; the query parameter matrix, key parameter matrix, and value parameter matrix are all matrices composed of learnable parameters;

[0043] Perform a dot product operation on the fused feature vector and the query parameter matrix to obtain a query vector, perform a dot product operation on the media content feature sequence and the key parameter matrix to obtain a key vector, and perform a dot product operation on the media content feature sequence and the value parameter matrix to obtain a value vector;

[0044] Generate an attention score vector based on the query vector and the key vector, perform dimensionality reduction processing on the attention score vector based on the vector length of the key vector, perform normalization processing on the dimensionality-reduced attention score vector to obtain an attention weight vector, and perform a dot product operation on the attention weight vector and the value vector to obtain a feature vector containing the attention result vector.

[0045] Among them, fusing the cross-connected feature vector and the M unit domain feature vectors to obtain a fused feature vector includes:

[0046] If the cross-connected feature vector and the M unit domain feature vectors meet the isomorphic feature condition, perform vector addition on the cross-connected feature vector and the M unit domain feature vectors to obtain a fused feature vector; the isomorphic feature condition means that the cross-connected feature vector and the M unit domain feature vectors belong to the same semantic space;

[0047] If the cross-connected feature vector and the M unit domain feature vectors do not meet the isomorphic feature condition, perform vector splicing on the cross-connected feature vector and the M unit domain feature vectors to obtain a fused feature vector.

[0048] On the one hand, an embodiment of the present application provides another data processing method, including:

[0049] Obtain the business data of M sample objects, and input the business data of M sample objects into the initial cross-connection model; the business fields corresponding to the business data of M sample objects are different from each other; the initial cross-connection model includes an initial encoding layer, an initial feature interaction layer, and an initial feature routing layer; M is a positive integer;

[0050] Perform encoding processing on the business data of M sample objects respectively through the initial encoding layer to obtain M sample unit domain feature vectors, and input the M sample unit domain feature vectors into the initial feature interaction layer connected to the initial encoding layer;

[0051] Input the M sample unit domain feature vectors into the initial feature routing layer, and generate the sample cross-connection weights between each sample unit domain feature vector and the initial feature interaction layer through the initial feature routing layer; the sample cross-connection weights are used to characterize the correlation degree between the sample unit domain feature vector and the sample feature interaction layer;

[0052] Select N sample unit domain feature vectors from the M sample unit domain feature vectors according to the sample cross-connection weights, generate a sample cross-connection feature vector according to the N sample unit domain feature vectors, and input the sample cross-connection feature vector into the initial feature interaction layer connected to the initial feature routing layer; N is a positive integer less than or equal to M;

[0053] In the initial feature interaction layer, generate an initial prediction result associated with the business data of M sample objects through the sample cross-connection feature vector and the M sample unit domain feature vectors;

[0054] Based on the annotation results and the initial prediction results corresponding to the business data of M sample objects respectively, adjust the model parameters of the initial encoding layer, the initial feature interaction layer, and the initial feature routing layer to obtain a business cross-connection model composed of an encoding layer, a feature interaction layer, and a feature routing layer; the business cross-connection model is used to generate a business prediction result corresponding to the object business data.

[0055] One aspect of the embodiments of the present application provides a data processing device, including:

[0056] A data acquisition module, configured to acquire the business data of M objects, and input the business data of M objects into the business cross-connection model; the business fields corresponding to the business data of M objects are different from each other; the business cross-connection model includes an encoding layer, a feature interaction layer, and a feature routing layer; M is a positive integer;

[0057] A feature encoding module, configured to perform encoding processing on the business data of M objects respectively through the encoding layer to obtain M unit domain feature vectors, and input the M unit domain feature vectors into the feature interaction layer connected to the encoding layer;

[0058] The weight processing module is used to input M unit domain feature vectors into the feature routing layer, and generate the cross-connection weights between each unit domain feature vector and the feature interaction layer through the feature routing layer; the cross-connection weights are used to represent the correlation degree between the unit domain feature vector and the feature interaction layer.

[0059] The cross-connection processing module is used to select N unit domain feature vectors from the M unit domain feature vectors according to the cross-connection weights, generate cross-connection feature vectors based on the N unit domain feature vectors, input the cross-connection feature vectors into the feature interaction layer connected to the feature routing layer, and output, through the feature interaction layer, the feature vectors used to generate the service prediction results associated with the M object service data; N is a positive integer less than or equal to M.

[0060] In a possible implementation manner, when the weight processing module is used to input M unit domain feature vectors into the feature routing layer and generate the cross-connection weights between each unit domain feature vector and the feature interaction layer through the feature routing layer, it is specifically used to perform the following operations:

[0061] Input the M unit domain feature vectors into the feature routing layer, perform dimensionality reduction processing on the M unit domain feature vectors to obtain M dimensionality-reduced feature vectors; the vector dimensions of the M dimensionality-reduced feature vectors are the same.

[0062] Obtain the routing input matrix and routing output matrix associated with the feature interaction layer in the feature routing layer, perform dot product operations on the M dimensionality-reduced feature vectors and the routing input matrix respectively to obtain M routing feature vectors; the routing input matrix and the routing output matrix are jointly used to represent the network layer attributes of the feature interaction layer.

[0063] Generate the linear activation vectors corresponding to the M routing feature vectors through the rectified linear activation function, and perform dot product operations on the M linear activation vectors and the routing output matrix respectively to obtain M linear activation parameters.

[0064] Perform piecewise parameter mapping on the M linear activation parameters through the piecewise activation function to obtain the cross-connection weights between each unit domain feature vector and the feature interaction layer.

[0065] In a possible implementation manner, when the weight processing module is used to perform dimensionality reduction processing on the M unit domain feature vectors to obtain M dimensionality-reduced feature vectors, it is specifically used to perform the following operations:

[0066] Obtain the feature dimension values corresponding to the M unit domain feature vectors respectively, and based on the M feature dimension values, obtain the projection parameter matrices corresponding to the M unit domain feature vectors respectively.

[0067] Perform projection processing on the corresponding unit domain feature vectors through the M projection parameter matrices respectively to obtain M dimensionality-reduced feature vectors.

[0068] In a possible implementation, the cross-connection weights include a first weight value, a second weight value, and a third weight value; when the weight processing module is used to perform piecewise parameter mapping on M linear activation parameters through a piecewise activation function to obtain the cross-connection weights between each unit domain feature vector and the feature interaction layer, it is specifically used to perform the following operations:

[0069] Obtain the first activation threshold and the second activation threshold corresponding to the piecewise activation function;

[0070] If the linear activation parameter corresponding to the unit domain feature vector is less than the first activation threshold, set the cross-connection weight corresponding to the unit domain feature vector to the first weight value;

[0071] If the linear activation parameter corresponding to the unit domain feature vector is greater than the first activation threshold and less than the second activation threshold, generate a second weight value based on the linear activation parameter, and set the cross-connection weight corresponding to the unit domain feature vector to the second weight value;

[0072] If the linear activation parameter corresponding to the unit domain feature vector is greater than the second activation threshold, set the cross-connection weight corresponding to the unit domain feature vector to the third weight value;

[0073] Among them, the degrees of correlation between the unit domain feature vectors represented by the first weight value, the second weight value, and the third weight value and the feature interaction layer are different; the degree of correlation corresponding to the first weight value is less than the degree of correlation corresponding to the second weight value, and the degree of correlation corresponding to the second weight value is less than the degree of correlation corresponding to the third weight value.

[0074] In a possible implementation, the cross-connection weights include the weight values corresponding to M unit domain feature vectors respectively; when the cross-connection processing module is used to select N unit domain feature vectors from the M unit domain feature vectors according to the cross-connection weights, it is specifically used to perform the following operations:

[0075] Determine the N unit domain feature vectors from the M unit domain feature vectors whose cross-connection weights are greater than the weight threshold;

[0076] Generating a cross-connection feature vector based on the N unit domain feature vectors includes:

[0077] Generate a mask weight matrix according to the weight values corresponding to the N unit domain feature vectors respectively, and obtain the reduced-dimensional feature vectors corresponding to the N unit domain feature vectors respectively; the reduced-dimensional feature vectors are obtained by performing dimensionality reduction processing on the unit domain feature vectors;

[0078] Perform a dot product operation on the sequence composed of the N reduced-dimensional feature vectors and the mask weight matrix to obtain the cross-connection feature vector.

[0079] In a possible implementation, the number of feature interaction layers and feature routing layers in the service cross-connection model is both T. Each feature interaction layer is respectively connected to a feature routing layer, and the feature routing layers connected to each feature interaction layer are different from each other. The T feature interaction layers are connected in series; the input of each feature routing layer is M unit domain feature vectors, and the output of each feature routing layer is used to be input into the connected feature interaction layer; the T feature interaction layers include feature interaction layer H i and feature interaction layer H i-1 , where i is a positive integer less than T; the T feature routing layers include feature routing layer B i connected to feature interaction layer H i ; when the cross-connection processing module is used to output a feature vector for generating a service prediction result associated with M object service data through the feature interaction layer, it is specifically used to perform the following operations:

[0080] In feature interaction layer H i , based on the target feature vector corresponding to feature interaction layer H i and the cross-connection feature vector input by feature routing layer B i , output a feature vector; if feature interaction layer H i is the feature interaction layer connected to the encoding layer, the target feature vector is M unit domain feature vectors; if feature interaction layer H i is connected to feature interaction layer H i-1 , the target feature vector is the feature vector output by feature interaction layer H i-1 ;

[0081] The cross-connection processing module is further used to perform the following operations:

[0082] If feature interaction layer H i is the last feature interaction layer among the T feature interaction layers, generate a service prediction result associated with M object service data according to the feature vector output by feature interaction layer H i .

[0083] In a possible implementation, the service cross-connection model includes a gating network layer and S expert network layers, where S is a positive integer. Each expert network layer includes a candidate feature interaction layer and a candidate feature routing layer, and the service domains corresponding to the S expert network layers are different from each other; the cross-connection processing module is further used to perform the following operations:

[0084] Input the M unit domain feature vectors into the gating network layer, and generate gating activation values corresponding to each expert network layer through the gating network layer; the gating activation values are used to represent the correlation degree between the expert network layer and the M unit domain feature vectors;

[0085] Based on the S gating activation values, sort the S expert network layers, determine the target expert network layer among the sorted S expert network layers, determine the candidate feature interaction layer in the target expert network layer as the feature interaction layer, and determine the candidate feature routing layer in the target expert network layer as the feature routing layer.

[0086] In a possible implementation, when the cross-connection processing module is used to output a feature vector for generating a service prediction result associated with the service data of M objects through the feature interaction layer, it is specifically used to perform the following operations:

[0087] In the feature interaction layer, obtain the media content feature vectors corresponding to Q media contents in the media content database respectively, perform vector concatenation on the Q media content feature vectors to obtain a media content feature sequence; Q is a positive integer;

[0088] Perform fusion processing on the cross-connection feature vector and the M unit domain feature vectors to obtain a fusion feature vector, and perform attention processing on the media content feature sequence and the fusion feature vector to obtain a feature vector containing the attention result vector;

[0089] The cross-connection processing module is further used to perform the following operations:

[0090] Based on the attention scores corresponding to the Q media contents in the attention result vector, generate the estimated interaction parameters corresponding to the Q media contents respectively, and determine the estimated interaction parameters corresponding to the Q media contents respectively as the service prediction results associated with the service data of M objects.

[0091] In a possible implementation, when the cross-connection processing module is used to perform attention processing on the media content feature sequence and the fusion feature vector to obtain a feature vector containing the attention result vector, it is specifically used to perform the following operations:

[0092] Obtain the query parameter matrix, key parameter matrix, and value parameter matrix in the feature interaction layer; the query parameter matrix, key parameter matrix, and value parameter matrix are all matrices composed of learnable parameters;

[0093] Perform a dot product operation on the fusion feature vector and the query parameter matrix to obtain a query vector, perform a dot product operation on the media content feature sequence and the key parameter matrix to obtain a key vector, and perform a dot product operation on the media content feature sequence and the value parameter matrix to obtain a value vector;

[0094] Generate an attention score vector based on the query vector and the key vector, perform dimensionality reduction processing on the attention score vector based on the vector length of the key vector, perform normalization processing on the dimensionality-reduced attention score vector to obtain an attention weight vector, and perform a dot product operation on the attention weight vector and the value vector to obtain a feature vector containing the attention result vector.

[0095] In a possible implementation, when the cross-connection processing module is used to perform fusion processing on the cross-connection feature vector and M unit domain feature vectors to obtain a fusion feature vector, it is specifically used to perform the following operations:

[0096] If the cross-connection feature vector and the M unit domain feature vectors meet the isomorphic feature condition, the cross-connection feature vector and the M unit domain feature vectors are subjected to vector addition to obtain a fusion feature vector; the isomorphic feature condition means that the cross-connection feature vector and the M unit domain feature vectors belong to the same semantic space;

[0097] If the cross-connection feature vector and the M unit domain feature vectors do not meet the isomorphic feature condition, the cross-connection feature vector and the M unit domain feature vectors are subjected to vector concatenation to obtain a fusion feature vector.

[0098] On the one hand, an embodiment of the present application provides another data processing device, including:

[0099] A sample data acquisition module, configured to acquire M sample object service data and input the M sample object service data into an initial cross-connection model; the service fields corresponding to the M sample object service data are different from each other; the initial cross-connection model includes an initial encoding layer, an initial feature interaction layer, and an initial feature routing layer; M is a positive integer;

[0100] A sample feature encoding module, configured to respectively perform encoding processing on the M sample object service data through the initial encoding layer to obtain M sample unit domain feature vectors, and input the M sample unit domain feature vectors into the initial feature interaction layer connected to the initial encoding layer;

[0101] A sample weight processing module, configured to input the M sample unit domain feature vectors into the initial feature routing layer, and generate sample cross-connection weights between each sample unit domain feature vector and the initial feature interaction layer through the initial feature routing layer; the sample cross-connection weights are used to characterize the correlation degree between the sample unit domain feature vectors and the sample feature interaction layer;

[0102] A sample cross-connection processing module, configured to select N sample unit domain feature vectors from the M sample unit domain feature vectors according to the sample cross-connection weights, generate a sample cross-connection feature vector according to the N sample unit domain feature vectors, and input the sample cross-connection feature vector into the initial feature interaction layer connected to the initial feature routing layer; N is a positive integer less than or equal to M;

[0103] A result prediction module, configured to generate an initial prediction result associated with the M sample object service data through the sample cross-connection feature vector and the M sample unit domain feature vectors in the initial feature interaction layer;

[0104] A model training module, which is used to adjust the model parameters of an initial encoding layer, an initial feature interaction layer, and an initial feature routing layer based on the annotation results and initial prediction results respectively corresponding to the business data of M sample objects, so as to obtain a business cross-connection model composed of an encoding layer, a feature interaction layer, and a feature routing layer; the business cross-connection model is used to generate business prediction results corresponding to the object business data.

[0105] On the one hand, an embodiment of the present application provides a computer device, including: a processor, a memory, and a network interface;

[0106] The processor is connected to the memory and the network interface. Among them, the network interface is used to provide a data communication function, and the memory is used to store a computer program. When the computer program is executed by the processor, the computer device executes the method provided by the embodiment of the present application.

[0107] On the one hand, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded and executed by a processor so that a computer device having the processor executes the method provided by the embodiment of the present application.

[0108] On the one hand, an embodiment of the present application provides a computer program product. The computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the method provided by the embodiment of the present application.

[0109] In the embodiments of the present application, the encoding layer in the service cross-connection model encodes the service data of M object services with different service domains to obtain M unit domain feature vectors. The service cross-connection model further includes a feature interaction layer and a feature routing layer. The M unit domain feature vectors are input into the feature interaction layer connected to the encoding layer, and a corresponding feature routing layer is configured for the feature interaction layer. The M unit domain feature vectors are input into the feature routing layer. The cross-connection weights between each unit domain feature vector and the feature interaction layer are generated through the feature routing layer, and the correlation degree between the unit domain feature vector and the feature interaction layer is represented by the cross-connection weights. Furthermore, N unit domain feature vectors are selected from the M unit domain feature vectors through the cross-connection weights to generate cross-connection feature vectors. The cross-connection feature vectors are input into the feature interaction layer connected to the feature routing layer, and the feature vectors for generating service prediction results associated with the service data of the M objects are output through the feature interaction layer. It can be seen that in the embodiments of the present application, an automatic feature cross-connection structure is implemented through the feature routing layer. The feature routing layer can automatically calculate the cross-connection weights between the input features and the connected feature interaction layer, and then can select the features cross-connected by the feature interaction layer through the cross-connection weights, realizing the automatic screening of the cross-connected features, so that the service cross-connection model can better capture the correlation degree of the cross-connected features. Secondly, the automatic feature cross-connection can avoid the fixed mode of manual cross-connection. When facing different data distributions and task requirements, the feature cross-connection scheme can be accurately selected. When different users have different object service data, the feature routing layer can still accurately select the features matching the feature routing layer, so as to better capture the personalized needs of users and improve the accuracy of the service estimation results generated by the service cross-connection model. BRIEF DESCRIPTION OF THE DRAWINGS

[0110] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0111] Figure 1 is a schematic diagram of a network architecture provided by an embodiment of the present application;

[0112] Figure 2 is a schematic diagram of a data processing scenario provided by an embodiment of the present application;

[0113] Figure 3 is a flowchart of a data processing method provided by an embodiment of the present application Figure 1 ;

[0114] Figure 4It is a flowchart of a data processing method provided by an embodiment of the present application Figure 2 ;

[0115] Figure 5 It is a schematic diagram of a model structure provided by an embodiment of the present application Figure 1 ;

[0116] Figure 6 It is a schematic diagram of a model structure provided by an embodiment of the present application Figure 2 ;

[0117] Figure 7 It is a flowchart of a data processing method provided by an embodiment of the present application Figure 3 ;

[0118] Figure 8 It is a schematic diagram of the structure of a data processing device provided by an embodiment of the present application Figure 1 ;

[0119] Figure 9 It is a schematic diagram of the structure of a data processing device provided by an embodiment of the present application Figure 2 ;

[0120] Figure 10 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0121] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without making creative efforts shall fall within the protection scope of the present application.

[0122] It can be understood that in the specific implementation manners of the present application, for the object business data and other data related to the object, when the present application and the following embodiments are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant regions.

[0123] Wherein, if it is necessary to collect data of an object (such as a user, etc.) in this application, a prompt interface or a pop-up window is displayed before and during the collection. The prompt interface or the pop-up window is used to prompt the user that some data is currently being collected. Only after obtaining the user's confirmation operation on the prompt interface or the pop-up window, the relevant steps for data acquisition are started; otherwise, the process ends. Moreover, the obtained user data will be used in reasonable, legal scenarios or for legitimate purposes, etc. Optionally, in some scenarios where user data needs to be used but the user's authorization has not been obtained, authorization can also be requested from the user, and the user data will be used only when the authorization is passed.

[0124] Please refer to Figure 1 , Figure 1 which is a schematic diagram of a network architecture provided by an embodiment of this application. As Figure 1 shown, the network architecture may include a service server 100 and a cluster of terminal devices. The cluster of terminal devices may include terminal devices 10a, 10b,..., 10n. Among them, any terminal device in the cluster of terminal devices may have a communication connection with the service server 100. For example, there is a communication connection between terminal device 10a and service server 100, and there is a communication connection between terminal device 10b and service server 100. Among them, the above communication connection is not limited to the connection method and can be directly or indirectly connected through a wired communication method, or can be directly or indirectly connected through a wireless communication method, or can also be connected through other methods, which are not limited in this application.

[0125] Among them, each terminal device in the cluster of terminal devices may include: smart phones, tablet computers, laptop computers, desktop computers, intelligent voice interaction devices, smart home appliances (such as smart TVs), wearable devices, vehicle-mounted terminals, aircraft, etc., which are intelligent terminals with data processing functions. Among them, the vehicle-mounted terminal may be a terminal device in the intelligent transportation scenario and the assisted driving scenario. It should be understood that Figure 1 each terminal device in the cluster of terminal devices shown may be installed with an application client with data processing functions. When the application client runs on each terminal device, it can perform data interaction with the above-mentioned Figure 1 shown service server 100 respectively.

[0126] Among them, the application client may specifically include: in-vehicle client, smart home client, entertainment client (e.g., game client), multimedia client (e.g., video client), social client, and information client (e.g., news client), etc. Among them, the application client in the embodiments of the present application may be integrated in a certain client (e.g., social client), or may be an independent client (e.g., news client). The embodiments of the present application do not limit the type of the application client. Among them, the service server 100 may be the server corresponding to the application client, and the service server 100 may be an independent physical server.

[0127] For ease of understanding, take the terminal device 10a in the terminal device cluster as an example for illustration. As Figure 1 shown, the terminal device 10a may send M object service data of the object (user) in the application client to the service server 100. Among them, the object service data may be obtained by dividing the data generated by the object in the application client according to the actual use (such as the business field), or may be divided according to the media type of the data. The embodiments of the present application do not limit this here. For example, the M object service data may include data related to articles, data related to tags, and data related to interaction operations (such as likes, collections), etc.

[0128] A service cross-connection model may be deployed in the service server 100. The service server 100 may input the M object service data into the service cross-connection model. The service cross-connection model may include an encoding layer, a feature interaction layer, and a feature routing layer. The service server 100 may perform encoding processing on the M object service data respectively through the encoder in the service cross-connection model to obtain M unit domain feature vectors, and input the M unit domain feature vectors into the feature interaction layer connected to the encoding layer and the feature routing layer connected to the encoding layer.

[0129] The feature routing layer may automatically calculate the features required for cross-connection of the connected feature interaction layer, and then select N unit domain feature vectors from the M unit domain feature vectors, generate a cross-connection feature vector through the N unit domain feature vectors, and input the cross-connection feature vector into the feature interaction layer to complete feature cross-connection. Among them, feature cross-connection refers to the mechanism of directly transmitting and connecting the features of a certain layer to a subsequent certain layer. For the content of the service cross-connection model and the automatic feature cross-connection structure, reference may be made to the content of the corresponding embodiments below Figure 3 corresponding embodiments.

[0130] In the service cross - connection model, the service server 100 can generate service prediction results corresponding to M object service data through cross - connection feature vectors and M unit domain feature vectors, and send the service prediction results to the terminal device 10a. The service prediction results include media data streams recommended to the object, and the terminal device 10a can display the above - mentioned media data streams to the object in the application client.

[0131] In the embodiment of the present application, the feature routing layer is used to automatically screen out N unit domain feature vectors from M unit domain feature vectors, and generate cross - connection feature vectors through the N unit domain feature vectors, so as to accurately select the features cross - connected by the feature interaction layer, realize an automatic feature cross - connection structure, and ensure that the service cross - connection model always maintains a good performance state. It provides a cross - connection method that better meets the actual needs for different service scenarios, improves the accuracy of the service estimation results generated by the service cross - connection model, and greatly improves the pertinence and practicality of the model.

[0132] Please refer to Figure 2 , Figure 2 It is a schematic diagram of a data - processing scenario provided by the embodiment of the present application. As Figure 2 shown, the computer device can input M object service data into the recall module. The recall module can be a service cross - connection model, and the service cross - connection model can include an encoding layer, a feature interaction layer, and a feature routing layer. The computer device can respectively perform encoding processing on the M object service data through the encoder in the service cross - connection model to obtain M unit domain feature vectors, and input the M unit domain feature vectors into the feature interaction layer connected to the encoding layer and the feature routing layer connected to the encoding layer. Among them, the computer device can be the Figure 1 service server 100 in the corresponding embodiment above or any terminal device in the terminal device cluster.

[0133] The feature routing layer can automatically calculate the features that need to be cross-connected for the connected feature interaction layer, that is, it can select N unit domain feature vectors from M unit domain feature vectors, generate cross-connected feature vectors through the N unit domain feature vectors, and input the cross-connected feature vectors into the feature interaction layer to complete feature cross-connection. In the feature interaction layer, the computer device can obtain the media content feature vectors corresponding to Q media contents in the media content database. Among them, the media content can be text data (such as news articles, media tweets, etc.), video data (short videos, long videos, etc.), image data (photographic images, comics, etc.) and their combinations, and the embodiments of the present application do not limit this here. The computer device can perform vector splicing on the global media content feature vectors to obtain a media content feature sequence, perform fusion processing on the cross-connected feature vectors and M unit domain feature vectors to obtain fusion feature vectors, and perform attention processing on the media content feature sequence and the fusion feature vectors to obtain feature vectors containing attention result vectors. The computer device can generate estimated interaction parameters corresponding to Q media contents through the attention scores in the attention result vectors, and determine the media contents in the Q media contents whose attention scores or estimated interaction parameters are greater than or equal to a certain threshold as the contents recalled by the service cross-connection model in the media content database. Among them, the estimated interaction parameters can refer to advertising parameters such as CTR (Click-Through Rate), CVR (Conversion Rate), and Return on Investment (ROI).

[0134] The computer device can input the recalled media content into the selection module. The selection module can also be a prediction model, which is used to further predict the estimated interaction parameters of the media content screened by the recall module, and perform fine sorting through the estimated interaction parameters. The media content sorted by the selection module can be input into the rearrangement module to rearrange the selected media content, and finally recommended to the user. Among them, rearrangement can refer to shuffling and mixing media contents of different theme types. The computer device can display the rearranged media data in the application client. As shown in the scenario page 101, the rearranged media content can include media content 1, media content 2,..., media content n. As shown in the scenario page 102, the user can click on media content 1 to view media content 1.

[0135] In the embodiment of the present application, through the proposed automatic feature cross-connection structure, that is, the feature routing layer can automatically select N unit domain feature vectors from M unit domain feature vectors, and generate cross-connection feature vectors through the N unit domain feature vectors, so as to accurately select the features cross-connected by the feature interaction layer, enabling the service cross-connection model to better adapt to various situations. Secondly, automatic cross-connection can be tailored to individual needs. When facing different data distributions and task requirements, it can accurately select the feature cross-connection scheme. When different users have different object service data, the feature routing layer can still accurately select the features that match the feature routing layer, and customize the feature cross-connection scheme for each user, so as to better capture the personalized needs of users and improve the recommendation accuracy of the model.

[0136] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of a data processing method provided by an embodiment of the present application. Figure 1 This data processing method can be executed by a computer device, and the computer device can be any terminal device in the service server 100 or the terminal device cluster as shown in Figure 1 . For example, it can be the terminal device 10a. Hereinafter, this data processing method being executed by a computer device will be taken as an example for description. Among them, this data processing method can at least include the following steps S101-S104:

[0137] Step S101, obtain M object service data, and input the M object service data into the service cross-connection model; the service fields corresponding to the M object service data are different from each other; the service cross-connection model includes an encoding layer, a feature interaction layer, and a feature routing layer; M is a positive integer.

[0138] Specifically, the computer device can obtain M object service data. The object service data can be obtained by dividing the data generated by the object in the application client (the application installed on the computer device) according to the actual use (such as the service field), or can be divided according to the media type of the data. The embodiment of the present application does not limit this here. The service fields corresponding to the M object service data are different from each other. For example, the M object service data can include data related to articles, data related to tags, and data related to interaction operations (such as likes, collections, etc.).

[0139] The computer device can input M object service data into the service cross-connection model. The service cross-connection model can be obtained by constructing an automatic cross-connection structure for each feature interaction layer in the baseline model on the basis of the baseline model in any service field. The service cross-connection model can include an encoding layer, a feature interaction layer, and a feature routing layer. Among them, the feature interaction layer can be a network layer that requires multi-level features for task processing, and the feature routing layer can include a network layer of a multi-layer perceptron (MLP) and a binary classification activation function (such as a sigmoid activation function, a hard sigmoid activation function).

[0140] Step S102, respectively perform encoding processing on the M object service data through the encoding layer to obtain M unit domain feature vectors, and input the M unit domain feature vectors into the feature interaction layer connected to the encoding layer;

[0141] Specifically, the computer device can respectively perform encoding processing on the M object service data through the encoding layer to obtain M unit domain feature vectors, and input the M unit domain feature vectors into the feature interaction layer connected to the encoding layer. Taking the object service data as text data as an example, the computer device can split and encode the object service data into tokens through a Tokenizer. For example, if the object service data is “Let’s do tokenization!”, the corresponding token sequence of the object service data can be expressed as [Let,’s,do,token,ization]. Among them, [Let,’s], [do], [token], and [ization] are all tokens. The method of text splitting can be word-based, character-based, or subword-based. The embodiments of the present application do not limit this here.

[0142] Step S103, input the M unit domain feature vectors into the feature routing layer, and generate cross-connection weights between each unit domain feature vector and the feature interaction layer through the feature routing layer; the cross-connection weights are used to characterize the correlation degree between the unit domain feature vector and the feature interaction layer;

[0143] Specifically, the computer device can input M unit domain feature vectors into the feature routing layer. Through the multi-layer perceptron and binary classification activation function in the feature routing layer, the cross-connection weights between each unit domain feature vector and the feature interaction layer are generated. The cross-connection weights can be used to represent the correlation degree between the unit domain feature vector and the feature interaction layer, that is, the importance of feature cross-connection. For example, the cross-connection weight can be a value between the interval [0, 1]. When the cross-connection weight is 0, it means that the unit domain feature vector corresponding to this cross-connection weight is not relevant to the feature interaction layer. When the cross-connection weight is greater than 0, it means that the unit domain feature vector corresponding to this cross-connection weight is relevant to the feature interaction layer. The greater the cross-connection weight, the higher the correlation degree between the unit domain feature vector and the feature interaction layer.

[0144] Step S104, select N unit domain feature vectors from the M unit domain feature vectors according to the cross-connection weights, generate a cross-connection feature vector according to the N unit domain feature vectors, input the cross-connection feature vector into the feature interaction layer connected to the feature routing layer, and output, through the feature interaction layer, a feature vector for generating a service prediction result associated with the M object service data; N is a positive integer less than or equal to M.

[0145] Specifically, the computer device can select N unit domain feature vectors from the M unit domain feature vectors through the M cross-connection weights. For example, it can be the unit domain feature vectors with cross-connection weights greater than 0 among the M unit domain feature vectors, which are determined as the N unit domain feature vectors. The computer device can generate a cross-connection feature vector through the N unit domain feature vectors and input the cross-connection feature vector into the feature interaction layer connected to the feature routing layer. Among them, the feature interaction layer is used to generate a feature vector for the service prediction result associated with the M object service data.

[0146] For example, in the feature interaction layer, the computer device can obtain the media content feature vectors corresponding to Q media contents in the media content database. Among them, the media content can be text data (such as news articles, media tweets, etc.), video data (short videos, long videos, etc.), image data (photographic images, comics, etc.) and their combinations. The embodiments of the present application do not make limitations here. The computer device can splice the global media content feature vectors to obtain a media content feature sequence, perform fusion processing on the cross-connection feature vector and the M unit domain feature vectors to obtain a fusion feature vector, and perform attention processing on the media content feature sequence and the fusion feature vector to obtain a feature vector containing an attention result vector. The computer device can generate estimated interaction parameters corresponding to the Q media contents respectively through the attention scores in the attention result vector, and determine the estimated interaction parameters corresponding to the Q media contents respectively as the service prediction result. Among them, the estimated interaction parameters can refer to advertising parameters such as CTR, CVR, and ROI.

[0147] The embodiment of the present application encodes M object business data with different business fields through the encoding layer in the business cross-connection model to obtain M unit domain feature vectors. The business cross-connection model also includes a feature interaction layer and a feature routing layer. The M unit domain feature vectors are input into the feature interaction layer connected to the encoding layer, and a corresponding feature routing layer is configured for the feature interaction layer. The M unit domain feature vectors are input into the feature routing layer, and the cross-connection weights between each unit domain feature vector and the feature interaction layer are generated through the feature routing layer. The correlation between the unit domain feature vector and the feature interaction layer is represented by the cross-connection weights. Then, N unit domain feature vectors are selected from the M unit domain feature vectors through the cross-connection weights to generate a cross-connection feature vector, and the cross-connection feature vector is input into the feature interaction layer connected to the feature routing layer. The feature interaction layer outputs a feature vector for generating a business prediction result associated with the M object business data. It can be seen that in the embodiment of the present application, an automatic feature cross-connection structure is realized through the feature routing layer. The feature routing layer can automatically calculate the cross-connection weights of the input features and the connected feature interaction layer, and then the features cross-connected by the feature interaction layer can be selected through the cross-connection weights, so as to realize automatic screening of the cross-connected features, so that the business cross-connection model can better capture the relevance of the cross-connected features. Secondly, automatic feature cross-connection can avoid the fixed mode of manual cross-connection. When faced with different data distributions and task requirements, the feature cross-connection scheme can be accurately selected. When different users have different object business data, the feature routing layer can still accurately select features that match the feature routing layer, thereby better capturing the personalized needs of users and improving the accuracy of the business prediction results generated by the business cross-connection model.

[0148] See also Figure 4 , Figure 4 This is a flow diagram of a data processing method provided in an embodiment of the present application. Figure 2 The data processing method can be executed by a computer device, which can be Figure 1 The service server 100 shown or any terminal device in the terminal device cluster may be, for example, the terminal device 10a. The following will take the data processing method executed by a computer device as an example for explanation. The data processing method may at least include the following steps S201 to S206:

[0149] Step S201, obtaining M object service data;

[0150] For details, please refer to the above Figure 3 The specific content of step S101 of the corresponding embodiment will not be repeated here in this embodiment of the present application.

[0151] Step S202: Encode each of the M object service data through an encoding layer to obtain M unit domain feature vectors, and input the M unit domain feature vectors into a feature interaction layer connected to the encoding layer;

[0152] Specifically, the computer device can input the M object service data into the service cross-connection model. Figure 5 , Figure 5 It is a schematic diagram of a model structure provided by an embodiment of the present application. Figure 1 , such as Figure 5 shown, the service cross-connection model may include an encoding layer, T feature interaction layers, T feature routing layers, and a decoding layer. Among them, the T feature interaction layers may include Feature Interaction Layer 1, Feature Interaction Layer 2,..., Feature Interaction Layer T. The T feature routing layers may include Feature Routing Layer 1 corresponding to Feature Interaction Layer 1, Feature Routing Layer 2 corresponding to Feature Interaction Layer 2,..., Feature Routing Layer T corresponding to Feature Interaction Layer T. The feature interaction layer may be a network layer that requires multi-level features for task processing. The feature routing layer may include a multi-layer perceptron and a network layer with a binary classification activation function (such as a sigmoid activation function, a hard sigmoid activation function). Each feature interaction layer is respectively connected to a feature routing layer, and the feature routing layers connected to each feature interaction layer are different from each other. The T feature interaction layers are connected in series. The input of each feature routing layer is the same, and the output of each feature routing layer is used to input into the connected feature interaction layer.

[0153] The computer device can encode each of the M object service data through the encoding layer to obtain M unit domain feature vectors. Taking the object service data as image data as an example, the computer device can divide the object service data into M unit image data (patches). Taking the size of the object service data as 48 pixels × 48 pixels as an example, the computer device can divide the object service data into 9 unit image data, and the 9 unit image data may include Unit Image Data 1, Unit Image Data 2, Unit Image Data 3,..., Unit Image Data 9. Each image unit may be in a ratio of 16 pixels × 16 pixels. The computer device can embed (embedding) the 9 unit image data into vectors of the same size through an encoder to obtain unit feature vectors corresponding to each unit image data respectively. The computer device can add a vector containing position information to the unit feature vectors, that is, perform position encoding on the position information of each unit image data in the sample image data to obtain (Position Embedding) corresponding to each unit image data respectively. The computer device can splice the unit feature vectors and the position feature vectors to obtain unit domain feature vectors.

[0154] Step S203: Input the M unit domain feature vectors into the feature routing layer, perform dimensionality reduction processing on the M unit domain feature vectors to obtain M dimensionality-reduced feature vectors; the vector dimensions of the M dimensionality-reduced feature vectors are the same; obtain the routing input matrix and the routing output matrix associated with the feature interaction layer in the feature routing layer, perform dot product operations on the M dimensionality-reduced feature vectors and the routing input matrix respectively to obtain M routing feature vectors; the routing input matrix and the routing output matrix are jointly used to characterize the network layer attributes of the feature interaction layer; through the rectified linear activation function, generate the linear activation vectors corresponding to the M routing feature vectors respectively, perform dot product operations on the M linear activation vectors and the routing output matrix respectively to obtain M linear activation parameters; perform piecewise parameter mapping on the M linear activation parameters through the piecewise activation function to obtain the cross-connection weights between each unit domain feature vector and the feature interaction layer.

[0155] Specifically, please also participate Figure 5 , such as Figure 5 As shown, the computer device can input the M unit domain feature vectors into each feature routing layer, and each feature routing layer can perform dimensionality reduction processing on the M unit domain feature vectors to obtain M dimensionality-reduced feature vectors. The dimensionality reduction processing can refer to converting the M unit domain feature vectors into feature dimensions that match the feature interaction layer connected to the feature routing layer. Among them, the vector dimensions of the M dimensionality-reduced feature vectors are the same.

[0156] For ease of understanding, take the feature routing layer 1 in the T feature routing layers as an example. The feature routing layer 1 can obtain the feature dimension values corresponding to the M unit domain feature vectors respectively, and the feature dimension value is the vector length of the unit domain feature vector. The feature routing layer can obtain the projection parameter matrices corresponding to the M unit domain feature vectors respectively based on the M feature dimension values, and perform projection processing on the corresponding unit domain feature vectors through the M projection parameter matrices to obtain M dimensionality-reduced feature vectors. Taking the conversion of the unit domain feature vector of dimension A to the dimensionality-reduced feature vector of dimension B as an example, the process of dimensionality reduction processing can be as shown in formula (1):

[0157]

[0158] Among them, i = 1, 2,..., B, j = 1, 2,..., A. α i is the i-th vector element of the dimensionality-reduced feature vector, β j is the j-th vector element of the unit domain feature vector, represents the element value of the j-th row and the i-th column of the projection parameter matrix W AB , and the dimension of the projection parameter matrix W AB is A×B. The dimensionality-reduced feature vector can be expressed as {α1, α2,..., α B}. The dimensionality reduction process can be achieved through a multi-layer perceptron. Therefore, during the dimensionality reduction process, the multi-layer perceptron can add a bias term to each vector element obtained after dimensionality reduction. The embodiments of the present application do not limit this here.

[0159] It can be understood that through the dimensionality reduction process, the embodiments of the present application can convert the features processed by each feature routing layer into the same vector dimension. The dimensionality reduction process can map high-dimensional features into a low-dimensional space. High-dimensional data may contain some noise and irrelevant information. By dimensionality reduction, these interference factors can be filtered out, enabling the model to focus more on learning the core patterns and rules in the data and enhancing the generalization ability of the model. By mapping high-dimensional features into a low-dimensional space through dimensionality reduction coding, the dimension of the data can be reduced without losing important information, reducing the time consumption and computational complexity during the calculation process, and improving the training efficiency and generalization ability of the model. At the same time, through the dimensionality reduction process, the total dimension of the features after cross-connection is also limited, avoiding the situation where too many cross-connected features "take over the spotlight" and ensuring the stability and reliability of the model.

[0160] The computer device can obtain the routing input matrix W associated with the feature interaction layer in the feature routing layer 1 and the routing output matrix W 2 . Both the routing input matrix and the routing output matrix are weight matrices composed of learnable parameters of the feature routing layer, and can be jointly used to represent the network layer attributes of the feature interaction layer. The computer device can perform dot product operations on the M dimensionality reduction feature vectors and the routing input matrix respectively to obtain M routing feature vectors, and generate linear activation vectors corresponding to the M routing feature vectors through a rectified linear unit activation function. The process can be as shown in formula (2):

[0161] h = ReLU(W 1 ·x + b1) Formula (2)

[0162] where h is the linear activation vector corresponding to the routing feature vector, x is the dimensionality reduction feature vector, ReLU is the rectified linear unit activation function, and b1 is the bias term.

[0163] The computer device can perform dot product operations on the M linear activation vectors and the routing output matrix respectively to obtain M linear activation parameters. The process can be as shown in formula (3):

[0164] s = W 2 ·h + b2 Formula (3)

[0165] where s is the linear activation parameter and b2 is the bias term.

[0166] It can be understood that the computer device can perform piecewise parameter mapping on the M linear activation parameters through a piecewise activation function to obtain the cross-connection weights between each unit domain feature vector and the feature interaction layer. The process can be shown as in formula (4):

[0167] hard sigmoid(s) = max(0, min(1, γ·s + δ)) Formula (4)

[0168] Among them, hard sigmoid(s) is the cross-connection weight. γ is the slope parameter, δ is the offset parameter, and both γ and δ are pre-set parameters.

[0169] The piecewise activation function can also be expressed as a piecewise function. The computer device can obtain the first activation threshold and the second activation threshold corresponding to the piecewise activation function. If the linear activation parameter corresponding to the unit domain feature vector is less than the first activation threshold, the cross-connection weight corresponding to the unit domain feature vector is set to the first weight value. If the linear activation parameter corresponding to the unit domain feature vector is greater than the first activation threshold and less than the second activation threshold, the second weight value is generated based on the linear activation parameter, and the cross-connection weight corresponding to the unit domain feature vector is set to the second weight value. If the linear activation parameter corresponding to the unit domain feature vector is greater than the second activation threshold, the cross-connection weight corresponding to the unit domain feature vector is set to the third weight value. The process can be shown as in formula (5):

[0170]

[0171] Among them, ε1 is the first activation threshold, and ε2 is the second activation threshold. The first weight value can be 0. The second weight value can be γ·s + δ, and γ·s + δ can be less than 1. For example, it can be 0.2s + 0.3. The third weight value can be 1. The correlation degrees of the unit domain feature vectors characterized by the first weight value, the second weight value, and the third weight value with the feature interaction layer are different from each other. The correlation degree corresponding to the first weight value is less than the correlation degree corresponding to the second weight value, and the correlation degree corresponding to the second weight value is less than the correlation degree corresponding to the third weight value.

[0172] It can be understood that the binary classification activation function in the feature routing layer of the embodiments of the present application can be the hard sigmoid activation function. The hard sigmoid activation function can be more easily quantized in fixed-point numbers, making the cross-connection weights more concentrated on fixed 0 or 1, and making the features screened by feature cross-connection clearer and more stable.

[0173] Step S204, determine the N unit domain feature vectors from the M unit domain feature vectors whose cross-connection weights are greater than the weight threshold;

[0174] Specifically, the cross-connection weights include the weight values corresponding to M unit domain feature vectors respectively. The computer device can determine N unit domain feature vectors from the M unit domain feature vectors whose cross-connection weights are greater than the weight threshold. Among them, the weight threshold can be a preset value, for example, it can be 0.

[0175] Step S205: Generate a mask weight matrix according to the weight values corresponding to the N unit domain feature vectors respectively, and obtain the dimensionality-reduced feature vectors corresponding to the N unit domain feature vectors respectively; the dimensionality-reduced feature vectors are obtained by performing dimensionality reduction processing on the unit domain feature vectors; perform a dot product operation on the sequence composed of the N dimensionality-reduced feature vectors and the mask weight matrix to obtain the cross-connection feature vector.

[0176] Specifically, the computer device can generate a mask weight matrix through the weight values corresponding to the N unit domain feature vectors respectively. The element value of the mask weight matrix corresponding to the unit domain feature vector can be the weight value corresponding to the unit domain feature vector. For example, the mask weight matrix can be a matrix composed of 0, γ·s + δ, and 1. The computer device can obtain the dimensionality-reduced feature vectors corresponding to the N unit domain feature vectors respectively, and perform a dot product operation on the sequence composed of the N dimensionality-reduced feature vectors and the mask weight matrix to obtain the cross-connection feature vector. Among them, by performing a dot product operation with the element value 0 in the mask weight matrix, the computer device can exclude the dimensionality-reduced feature vectors that the feature interaction layer does not need, and then select the dimensionality-reduced feature vectors that need to perform cross-connection. And the dimensionality-reduced feature vectors that need to perform cross-connection are weighted by the weight values, indicating the degree of relevance between the dimensionality-reduced feature vectors and the feature routing layer.

[0177] Step S206: Input the cross-connection feature vector into the feature interaction layer connected to the feature routing layer, and output, through the feature interaction layer, the feature vector used to generate the service prediction result associated with the M object service data;

[0178] Specifically, please also refer to Figure 5 , as Figure 5 shown, the computer device can input the cross-connection feature vector into the feature interaction layer connected to the feature routing layer. The input of each feature routing layer is M unit domain feature vectors, and the output of each feature routing layer is used to input into the connected feature interaction layer. The T feature interaction layers include the feature interaction layer H i and the feature interaction layer H i-1 , where i is a positive integer less than T; the T feature routing layers include the feature routing layer B i connected to the feature interaction layer H i ;

[0179] In the feature interaction layer H i , the computer device can be based on the feature interaction layer H iThe corresponding target feature vector and feature routing layer B i The input cross-connected feature vector, output feature vector, if the feature interaction layer H i is a feature interaction layer connected to the encoding layer, the target feature vector is an M-unit domain feature vector; if the feature interaction layer H i is connected to the feature interaction layer H i-1 then the target feature vector is the feature vector output by the feature interaction layer H i-1 If the feature interaction layer H i is the last feature interaction layer among the T feature interaction layers, that is, the feature interaction layer H i is a feature interaction layer connected to the decoding layer, the computer device can generate a service prediction result associated with M object service data through the feature vector output by the feature interaction layer H i Output feature vector.

[0180] For ease of understanding, taking the T feature interaction layers including feature interaction layer 1, feature interaction layer 2,..., feature interaction layer T as an example, the computer device can input the cross-connected feature vector output by the feature routing layer 1 to the feature interaction layer 1, and the feature interaction layer 1 is the feature interaction layer connected to the encoding layer. The feature interaction layer 1 can generate the feature vector 1 through the cross-connected feature vector output by the feature routing layer 1 and the M unit domain feature vectors, and the feature vector 1 is the vector output by the feature interaction layer 1.

[0181] The feature interaction layer 1 can input the feature vector 1 to the feature interaction layer 2, and the feature interaction layer 2 can generate the feature vector 2 through the cross-connected feature vector output by the feature routing layer 2 and the feature vector 1 output by the feature interaction layer 1. The feature vector 2 can be used to input to the feature interaction layer connected to the feature interaction layer 2 until the feature vector T output by the feature interaction layer T is obtained. The feature vector T can be used to input to the decoding layer, and the feature vector T can be decoded in the decoding layer to obtain the service prediction result. The feature vector T is the feature vector used to generate the service prediction result associated with the M object service data.

[0182] It can be understood that the automatic cross-connection structure avoids the complex and fixed feature connections in the manual feature cross-connection scheme, reduces the complexity of the model to a certain extent, makes the training and inference of the model more efficient, and improves the generalization ability and adaptability of the model.

[0183] Each feature interaction layer can process the input vectors for tasks. For ease of understanding, take the feature interaction layer 1 in the T feature interaction layers as an example. In the feature interaction layer 1, the feature interaction layer 1 can obtain the media content feature vectors corresponding to Q media contents in the media content database, and splice the Q media content feature vectors to obtain a media content feature sequence. Among them, the media content can be text data (such as news articles, media tweets, etc.), video data (short videos, long videos, etc.), image data (photographic images, comics, etc.) and their combinations, which are not limited in the embodiments of the present application.

[0184] The feature interaction layer 1 can perform fusion processing on the cross-connected feature vector and M unit domain feature vectors to obtain a fusion feature vector. The process of fusion processing can be: if the cross-connected feature vector and the M unit domain feature vectors meet the isomorphic feature condition, then perform vector addition on the cross-connected feature vector and the M unit domain feature vectors to obtain a fusion feature vector; if the cross-connected feature vector and the M unit domain feature vectors do not meet the isomorphic feature condition, then splice the cross-connected feature vector and the M unit domain feature vectors to obtain a fusion feature vector. Among them, the isomorphic feature condition means that the cross-connected feature vector and the M unit domain feature vectors belong to the same semantic space, and the specific process of the fusion processing in the embodiments of the present application is not limited here.

[0185] The feature interaction layer 1 can perform attention processing on the media content feature sequence and the fusion feature vector to obtain a feature vector containing the attention result vector. The process of attention processing can be: obtain the query parameter matrix, key parameter matrix, and value parameter matrix in the feature interaction layer; the query parameter matrix, key parameter matrix, and value parameter matrix are all matrices composed of learnable parameters; perform a dot product operation on the fusion feature vector and the query parameter matrix to obtain a query vector, perform a dot product operation on the media content feature sequence and the key parameter matrix to obtain a key vector, and perform a dot product operation on the media content feature sequence and the value parameter matrix to obtain a value vector; generate an attention score vector based on the query vector and the key vector, perform dimensionality reduction processing on the attention score vector based on the vector length of the key vector, perform normalization processing on the dimensionality-reduced attention score vector to obtain an attention weight vector, and perform a dot product operation on the attention weight vector and the value vector to obtain a feature vector containing the attention result vector.

[0186] Specifically, the computer device can obtain the query parameter matrix W Q 、key parameter matrix W K and value parameter matrix W V in the feature interaction layer. Among them, the query parameter matrix W Q 、key parameter matrix W K and value parameter matrix W V are all matrices composed of learnable parameters.

[0187] The computer device combines the feature vector with the query parameter matrix W Q to perform a dot product operation to obtain a query vector Q. It performs a dot product operation on the media content feature sequence and the key parameter matrix W K to obtain a key vector K. It performs a dot product operation on the media content feature sequence and the value parameter matrix W V to obtain a value vector V. The computer device can generate an attention score vector based on the query vector Q and the key vector K, perform dimensionality reduction processing on the attention score vector based on the dimension d of the key vector, and perform normalization processing (Softmax) on the dimensionally reduced attention score vector to obtain a normalized vector. It performs a dot product operation on the normalized vector and the value vector V to obtain the feature vector z output by the feature interaction layer 1. The process can be as shown in formula (6):

[0188]

[0189] where, K T is the transposed key vector, and d k is the dimension of the attention key vector K.

[0190] When obtaining the feature vector T output by the feature interaction layer T, the feature vector T can be an attention result vector. The computer device can generate predicted interaction parameters corresponding to Q media contents based on the attention scores corresponding to Q media contents in the attention result vector, and determine the predicted interaction parameters corresponding to Q media contents as the business prediction results associated with M object service data. Among them, the predicted interaction parameters can refer to advertising parameters such as CTR, CVR, and ROI.

[0191] Optionally, the business cross-connection model can be a mixture of experts model ((Mixture of Experts, MoE), MoE), and please also refer to Figure 6 , Figure 6 which is a schematic diagram of a model structure provided by an embodiment of the present application Figure 2 ,as Figure 6 shown, the business cross-connection model can include a gating network layer and S expert network layers. The S expert network layers can include expert network 1, expert network 2,..., expert network S. Each expert network layer includes a candidate feature interaction layer and a candidate feature routing layer. The business domains corresponding to the S expert network layers are different from each other.

[0192] The computer device can input M unit domain feature vectors into the gating network layer, and generate gating activation values corresponding to each expert network layer through the gating network layer. The gating activation values can be used to represent the correlation degree between the expert network layer and the M unit domain feature vectors. The computer device can sort the S expert network layers based on the S gating activation values, determine the target expert network layer from the sorted S expert network layers, determine the candidate feature interaction layer in the target expert network layer as the feature interaction layer, and determine the candidate feature routing layer in the target expert network layer as the feature routing layer. The processing process of the target expert network layer can refer to the specific content of steps S203 to S206 above and will not be elaborated here. The target expert network layer can output an initial prediction result, and the computer device can determine the initial prediction result output by the target expert network layer as the service prediction result. Optionally, the computer device can also select and activate several expert network layers from the S expert network layers. Each expert network can focus on different task domains, and each expert network can output an initial prediction result. The computer device can aggregate the initial prediction results output by each expert network to obtain the service prediction result, improving the accuracy of the service prediction structure.

[0193] It can be understood that in the actual model construction process, different layers may have different functions and processing capabilities. For example, in a deep neural network, the shallow network is usually responsible for extracting low-level features of the data; while the deep network is more proficient in learning high-level semantic information and abstract features of the data. By performing cross-link operations on eligible domain features and connecting them to the corresponding network layers in the model. The embodiments of the present application can break the fixed mode of feature transmission in traditional models and achieve dynamic cross-linking and personalized adjustment of features. In this way, the features of each layer can complement and cooperate with each other, giving full play to their respective advantages. This cross-link operation can also be adjusted personalized for different data sets and tasks. Different data sets may have different feature distributions and rules. By flexibly selecting the domain features that need to be cross-linked, the embodiments of the present application can make the model better adapt to the specific data set and task requirements.

[0194] After rigorous experimental verification, the automatic cross-connection structure proposed in the embodiments of this application can effectively improve the model performance. This improvement is reflected in multiple aspects. For example, obvious improvements are shown in evaluation metrics such as AUC (Area Under the ROC Curve) and MAE (Mean Absolute Error). Moreover, the improvement of the model performance is directly reflected in the enhancement of business metrics, bringing a positive impact to the relevant business. Experimental verification has been carried out in the news recommendation system. By comparing with traditional methods, the method in the embodiments of this application has achieved significant improvements in various metrics; the online experiment has also obtained good business benefits, and key business metrics such as click-through rate, user stay duration, and user satisfaction have all been effectively improved.

[0195] In the embodiments of this application, the encoding layer in the business cross-connection model encodes M object business data with different business domains to obtain M unit domain feature vectors. The business cross-connection model also includes a feature interaction layer and a feature routing layer. The M unit domain feature vectors are input into the feature interaction layer connected to the encoding layer, and a corresponding feature routing layer is configured for the feature interaction layer. The M unit domain feature vectors are input into the feature routing layer. The feature routing layer generates the cross-connection weights between each unit domain feature vector and the feature interaction layer, and the cross-connection weights represent the correlation degree between the unit domain feature vector and the feature interaction layer. Furthermore, N unit domain feature vectors are selected from the M unit domain feature vectors through the cross-connection weights to generate cross-connection feature vectors. The cross-connection feature vectors are input into the feature interaction layer connected to the feature routing layer, and the feature interaction layer outputs the feature vectors used to generate the business prediction results associated with the M object business data. It can be seen that in the embodiments of this application, an automatic feature cross-connection structure is realized through the feature routing layer. The feature routing layer can automatically calculate the cross-connection weights between the input features and the connected feature interaction layer, and then can select the features cross-connected by the feature interaction layer through the cross-connection weights, realizing the automatic screening of the cross-connected features, so that the business cross-connection model can better capture the correlation degree of the cross-connected features. Secondly, the automatic feature cross-connection can avoid the fixed mode of manual cross-connection. When facing different data distributions and task requirements, the feature cross-connection scheme can be accurately selected. When different users have different object business data, the feature routing layer can still accurately select the features that match the feature routing layer, so as to better capture the personalized needs of users and improve the accuracy of the business prediction results generated by the business cross-connection model.

[0196] On the other hand, the automatic cross-connection structure avoids the complex and fixed feature connections in the manual feature cross-connection solution, reduces the complexity of the model to a certain extent, makes the training and inference of the model more efficient, and improves the generalization ability and adaptability of the model. At the same time, through dimensionality reduction processing, the embodiments of the present application can convert the features processed by each feature routing layer into the same vector dimension, map the high-dimensional features into the low-dimensional space through dimensionality reduction encoding, so as to reduce the dimension of the data without losing important information, reduce the time consumption and computational complexity in the calculation process, and improve the training efficiency and generalization ability of the model. At the same time, through dimensionality reduction processing, the total dimension of the features after cross-connection is also limited, avoiding the situation of "the guest usurping the host" caused by excessive cross-connection features, and ensuring the stability and reliability of the model.

[0197] Please refer to Figure 7 , Figure 7 which is a schematic flowchart of a data processing method provided by an embodiment of the present application Figure 3 , and this data processing method can be executed by a computer device, and the computer device can be any one of the service server 100 or the terminal device cluster as shown in Figure 1 , for example, it can be the terminal device 10a. The following will take this data processing method being executed by a computer device as an example for description. Among them, this data processing method can at least include the following steps S301-step S306:

[0198] Step S301, obtain M sample object service data, and input the M sample object service data into the initial cross-connection model; the service fields corresponding to the M sample object service data are different from each other; the initial cross-connection model includes an initial encoding layer, an initial feature interaction layer, and an initial feature routing layer; M is a positive integer;

[0199] Specifically, reference can be made to the specific content of step S201 in the corresponding embodiment above Figure 4 , and the embodiments of the present application will not elaborate here.

[0200] Step S302, respectively perform encoding processing on the M sample object service data through the initial encoding layer to obtain M sample unit domain feature vectors, and input the M sample unit domain feature vectors into the initial feature interaction layer connected to the initial encoding layer;

[0201] Specifically, reference can be made to the specific content of step S202 in the corresponding embodiment above Figure 4 , and the embodiments of the present application will not elaborate here.

[0202] Step S303: Input the M sample unit domain feature vectors into the initial feature routing layer, and generate the sample cross-connection weights between each sample unit domain feature vector and the initial feature interaction layer through the initial feature routing layer; the sample cross-connection weights are used to characterize the correlation degree between the sample unit domain feature vector and the sample feature interaction layer.

[0203] Specifically, reference can be made to the specific content of step S203 in the corresponding embodiment above. The embodiments of the present application will not elaborate here. Figure 4 Specifically, reference can be made to the specific content of step S203 in the corresponding embodiment above. The embodiments of the present application will not elaborate here.

[0204] Step S304: Select N sample unit domain feature vectors from the M sample unit domain feature vectors according to the sample cross-connection weights, generate a sample cross-connection feature vector based on the N sample unit domain feature vectors, and input the sample cross-connection feature vector into the initial feature interaction layer connected to the initial feature routing layer; N is a positive integer less than or equal to M.

[0205] Specifically, reference can be made to the specific content of steps S204 and S205 in the corresponding embodiment above. The embodiments of the present application will not elaborate here. Figure 4 Specifically, reference can be made to the specific content of steps S204 and S205 in the corresponding embodiment above. The embodiments of the present application will not elaborate here.

[0206] Step S305: In the initial feature interaction layer, generate an initial prediction result associated with the M sample object service data through the sample cross-connection feature vector and the M sample unit domain feature vectors.

[0207] Specifically, reference can be made to the specific content of step S206 in the corresponding embodiment above. The embodiments of the present application will not elaborate here. Figure 4 Specifically, reference can be made to the specific content of step S206 in the corresponding embodiment above. The embodiments of the present application will not elaborate here.

[0208] Step S306: Based on the annotation results and the initial prediction results respectively corresponding to the M sample object service data, adjust the model parameters of the initial coding layer, the initial feature interaction layer, and the initial feature routing layer to obtain a service cross-connection model composed of a coding layer, a feature interaction layer, and a feature routing layer; the service cross-connection model is used to generate a service prediction result corresponding to the object service data.

[0209] Specifically, the computer device can generate a model loss value based on the annotation results and initial prediction results respectively corresponding to the business data of M sample objects. The loss function for calculating the model loss value can be Binary Cross-Entropy Loss (BCE), Focal Loss (FL), etc., and the embodiments of the present application do not limit this here. The computer device can adjust the model parameters of the initial encoding layer, initial feature interaction layer, and initial feature routing layer through the model loss value. When the initial encoding layer, initial feature interaction layer, and initial feature routing layer all meet the training convergence conditions, an encoding layer trained from the initial encoding layer, a feature interaction layer trained from the initial feature interaction layer, and a feature routing layer trained from the initial feature routing layer are obtained, and then a business cross-connection model composed of the encoding layer, feature interaction layer, and feature routing layer is obtained. Among them, the business cross-connection model is used to generate the business prediction results corresponding to the object business data. The training convergence condition can refer to that during the training process, if the loss value of the loss function does not decrease significantly after multiple training batches (training rounds) or the training batch reaches the preset maximum value, etc., or a preset performance metric (such as recall rate, accuracy, etc.) is reached. The embodiments of the present application do not limit the training convergence condition here.

[0210] In the embodiments of the present application, an initial cross-connection model is trained into a service cross-connection model, and the service cross-connection model may include an encoding layer, a feature routing layer, and a feature interaction layer. When the task domain processed by the feature interaction layer in the service cross-connection model changes, the feature routing layer and the feature interaction layer can be retrained to achieve the adaptive change of the feature routing layer and adjust the features cross-connected by the feature interaction layer. The encoding layer in the service cross-connection model encodes M object service data with different service domains to obtain M unit domain feature vectors. The M unit domain feature vectors are input into the feature interaction layer connected to the encoding layer, and a corresponding feature routing layer is configured for the feature interaction layer. The M unit domain feature vectors are input into the feature routing layer, and the cross-connection weights between each unit domain feature vector and the feature interaction layer are generated through the feature routing layer. The cross-connection weights represent the correlation degree between the unit domain feature vector and the feature interaction layer. Furthermore, N unit domain feature vectors are selected from the M unit domain feature vectors through the cross-connection weights to generate cross-connection feature vectors. The cross-connection feature vectors are input into the feature interaction layer connected to the feature routing layer, and the feature vectors used to generate the service prediction results associated with the M object service data are output through the feature interaction layer. It can be seen that in the embodiments of the present application, an automatic feature cross-connection structure is realized through the feature routing layer. The feature routing layer can automatically calculate the cross-connection weights between the input features and the connected feature interaction layer, and then can select the features cross-connected by the feature interaction layer through the cross-connection weights, realizing the automatic screening of the cross-connected features, so that the service cross-connection model can better capture the correlation degree of the cross-connected features. Secondly, the automatic feature cross-connection can avoid the fixed mode of manual cross-connection. When facing different data distributions and task requirements, the feature cross-connection scheme can be accurately selected. When different users have different object service data, the features matching the feature routing layer can still be accurately selected through the feature routing layer, so as to better capture the personalized needs of users and improve the accuracy of the service estimation results generated by the service cross-connection model.

[0211] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of a data processing device provided by the embodiments of the present application Figure 1 . As Figure 8 shown, the data processing device 1 includes a data acquisition module 810, a feature encoding module 820, a weight processing module 830, and a cross-connection processing module 840.

[0212] The data acquisition module 810 is configured to acquire M object service data and input the M object service data into the service cross-connection model; the service domains corresponding to the M object service data are different from each other; the service cross-connection model includes an encoding layer, a feature interaction layer, and a feature routing layer; M is a positive integer;

[0213] A feature encoding module 820, which is used to encode M pieces of object service data respectively through an encoding layer to obtain M unit domain feature vectors, and input the M unit domain feature vectors into a feature interaction layer connected to the encoding layer;

[0214] A weight processing module 830, which is used to input the M unit domain feature vectors into a feature routing layer, and generate cross-connection weights between each unit domain feature vector and the feature interaction layer through the feature routing layer; the cross-connection weights are used to represent the correlation degree between the unit domain feature vector and the feature interaction layer;

[0215] A cross-connection processing module 840, which is used to select N unit domain feature vectors from the M unit domain feature vectors according to the cross-connection weights, generate a cross-connection feature vector according to the N unit domain feature vectors, input the cross-connection feature vector into a feature interaction layer connected to the feature routing layer, and output a feature vector for generating a service prediction result associated with the M pieces of object service data through the feature interaction layer; N is a positive integer less than or equal to M.

[0216] In a possible implementation manner, when the weight processing module 830 is used to input the M unit domain feature vectors into the feature routing layer and generate cross-connection weights between each unit domain feature vector and the feature interaction layer through the feature routing layer, it is specifically used to perform the following operations:

[0217] Input the M unit domain feature vectors into the feature routing layer, perform dimensionality reduction processing on the M unit domain feature vectors to obtain M dimensionality-reduced feature vectors; the vector dimensions of the M dimensionality-reduced feature vectors are the same;

[0218] Obtain a routing input matrix and a routing output matrix associated with the feature interaction layer in the feature routing layer, perform dot product operations on the M dimensionality-reduced feature vectors and the routing input matrix respectively to obtain M routing feature vectors; the routing input matrix and the routing output matrix are jointly used to represent the network layer attributes of the feature interaction layer;

[0219] Generate linear activation vectors corresponding to the M routing feature vectors respectively through a rectified linear activation function, and perform dot product operations on the M linear activation vectors and the routing output matrix respectively to obtain M linear activation parameters;

[0220] Perform piecewise parameter mapping on the M linear activation parameters through a piecewise activation function to obtain cross-connection weights between each unit domain feature vector and the feature interaction layer.

[0221] In a possible implementation manner, when the weight processing module 830 is used to perform dimensionality reduction processing on the M unit domain feature vectors to obtain M dimensionality-reduced feature vectors, it is specifically used to perform the following operations:

[0222] Obtain the characteristic dimension values corresponding to the M unit domain feature vectors respectively, and based on the M characteristic dimension values, obtain the projection parameter matrices corresponding to the M unit domain feature vectors respectively;

[0223] Perform projection processing on the corresponding unit domain feature vectors through the M projection parameter matrices respectively to obtain M dimensionality-reduced feature vectors.

[0224] In a possible implementation, the cross-connection weights include a first weight value, a second weight value, and a third weight value; when the weight processing module 830 is used to perform piecewise parameter mapping on the M linear activation parameters through a piecewise activation function to obtain the cross-connection weights between each unit domain feature vector and the feature interaction layer respectively, it is specifically used to perform the following operations:

[0225] Obtain the first activation threshold and the second activation threshold corresponding to the piecewise activation function;

[0226] If the linear activation parameter corresponding to the unit domain feature vector is less than the first activation threshold, set the cross-connection weight corresponding to the unit domain feature vector to the first weight value;

[0227] If the linear activation parameter corresponding to the unit domain feature vector is greater than the first activation threshold and less than the second activation threshold, generate a second weight value based on the linear activation parameter, and set the cross-connection weight corresponding to the unit domain feature vector to the second weight value;

[0228] If the linear activation parameter corresponding to the unit domain feature vector is greater than the second activation threshold, set the cross-connection weight corresponding to the unit domain feature vector to the third weight value;

[0229] Among them, the degrees of correlation between the unit domain feature vectors represented by the first weight value, the second weight value, and the third weight value and the feature interaction layer are different from each other; the degree of correlation corresponding to the first weight value is less than the degree of correlation corresponding to the second weight value, and the degree of correlation corresponding to the second weight value is less than the degree of correlation corresponding to the third weight value.

[0230] In a possible implementation, the cross-connection weights include the weight values corresponding to the M unit domain feature vectors respectively; when the cross-connection processing module 840 is used to select N unit domain feature vectors from the M unit domain feature vectors according to the cross-connection weights, it is specifically used to perform the following operations:

[0231] Determine the unit domain feature vectors with cross-connection weights greater than the weight threshold among the M unit domain feature vectors as the N unit domain feature vectors;

[0232] Generate cross-connection feature vectors based on the N unit domain feature vectors, including:

[0233] Generate a masked weight matrix according to the weight values corresponding to N unit domain feature vectors, and obtain the reduced-dimensional feature vectors corresponding to the N unit domain feature vectors respectively; the reduced-dimensional feature vectors are obtained by performing dimensionality reduction processing on the unit domain feature vectors.

[0234] Perform a dot product operation on the sequence composed of N reduced-dimensional feature vectors and the masked weight matrix to obtain cross-connected feature vectors.

[0235] In a possible implementation, the number of feature interaction layers and feature routing layers in the service cross-connection model is both T. Each feature interaction layer is respectively connected to a feature routing layer, and the feature routing layers connected to each feature interaction layer are different from each other. The T feature interaction layers are connected in series; the input of each feature routing layer is M unit domain feature vectors, and the output of each feature routing layer is used as the input to the connected feature interaction layer; the T feature interaction layers include feature interaction layer H i and feature interaction layer H i-1 , where i is a positive integer less than T; the T feature routing layers include the feature routing layer B i connected to the feature interaction layer H i ; when the cross-connection processing module 840 is used to output the feature vectors for generating service prediction results associated with M object service data through the feature interaction layer, it is specifically used to perform the following operations:

[0236] In the feature interaction layer H i , based on the target feature vector corresponding to the feature interaction layer H i and the cross-connected feature vectors input by the feature routing layer B i , output feature vectors; if the feature interaction layer H i is the feature interaction layer connected to the encoding layer, the target feature vector is M unit domain feature vectors; if the feature interaction layer H i is connected to the feature interaction layer H i-1 , the target feature vector is the feature vector output by the feature interaction layer H i-1 ;

[0237] The cross-connection processing module 840 is further used to perform the following operations:

[0238] If the feature interaction layer H i is the last feature interaction layer among the T feature interaction layers, then generate service prediction results associated with M object service data according to the feature vectors output by the feature interaction layer H i .

[0239] In a possible implementation, the service cross-connection model includes a gating network layer and S expert network layers, where S is a positive integer. Each expert network layer includes a candidate feature interaction layer and a candidate feature routing layer, and the service domains corresponding to the S expert network layers are different from each other; the cross-connection processing module 840 is further configured to perform the following operations:

[0240] Input M unit domain feature vectors into the gating network layer, and generate gating activation values corresponding to each expert network layer through the gating network layer; the gating activation values are used to characterize the correlation degree between the expert network layer and the M unit domain feature vectors;

[0241] Based on the S gating activation values, sort the S expert network layers, determine the target expert network layer among the sorted S expert network layers, determine the candidate feature interaction layer in the target expert network layer as the feature interaction layer, and determine the candidate feature routing layer in the target expert network layer as the feature routing layer.

[0242] In a possible implementation, when the cross-connection processing module 840 is used to output a feature vector for generating a service prediction result associated with M object service data through the feature interaction layer, it is specifically configured to perform the following operations:

[0243] In the feature interaction layer, obtain the media content feature vectors corresponding to Q media contents in the media content database, and perform vector concatenation on the Q media content feature vectors to obtain a media content feature sequence; Q is a positive integer;

[0244] Perform fusion processing on the cross-connection feature vector and the M unit domain feature vectors to obtain a fusion feature vector, and perform attention processing on the media content feature sequence and the fusion feature vector to obtain a feature vector containing an attention result vector;

[0245] The cross-connection processing module 840 is further configured to perform the following operations:

[0246] Generate estimated interaction parameters corresponding to the Q media contents based on the attention scores corresponding to the Q media contents in the attention result vector, and determine the estimated interaction parameters corresponding to the Q media contents as the service prediction results associated with the M object service data.

[0247] In a possible implementation, when the cross-connection processing module 840 is used to perform attention processing on the media content feature sequence and the fusion feature vector to obtain a feature vector containing an attention result vector, it is specifically configured to perform the following operations:

[0248] Obtain the query parameter matrix, key parameter matrix, and value parameter matrix in the feature interaction layer; the query parameter matrix, key parameter matrix, and value parameter matrix are all matrices composed of learnable parameters;

[0249] Perform a dot product operation on the fused feature vector and the query parameter matrix to obtain a query vector, perform a dot product operation on the media content feature sequence and the key parameter matrix to obtain a key vector, and perform a dot product operation on the media content feature sequence and the value parameter matrix to obtain a value vector;

[0250] Generate an attention score vector based on the query vector and the key vector, perform dimensionality reduction processing on the attention score vector based on the vector length of the key vector, perform normalization processing on the dimensionally reduced attention score vector to obtain an attention weight vector, and perform a dot product operation on the attention weight vector and the value vector to obtain a feature vector containing the attention result vector.

[0251] In a possible implementation manner, when the cross-connection processing module 840 is used to perform fusion processing on the cross-connection feature vector and the M unit domain feature vectors to obtain a fused feature vector, it is specifically used to perform the following operations:

[0252] If the cross-connection feature vector and the M unit domain feature vectors meet the isomorphic feature condition, perform vector addition on the cross-connection feature vector and the M unit domain feature vectors to obtain a fused feature vector; the isomorphic feature condition means that the cross-connection feature vector and the M unit domain feature vectors belong to the same semantic space;

[0253] If the cross-connection feature vector and the M unit domain feature vectors do not meet the isomorphic feature condition, perform vector concatenation on the cross-connection feature vector and the M unit domain feature vectors to obtain a fused feature vector.

[0254] In the embodiments of the present application, the encoding layer in the service cross-connection model encodes the service data of M object services with different service domains to obtain M unit domain feature vectors. The service cross-connection model further includes a feature interaction layer and a feature routing layer. The M unit domain feature vectors are input into the feature interaction layer connected to the encoding layer, and a corresponding feature routing layer is configured for the feature interaction layer. The M unit domain feature vectors are input into the feature routing layer. The cross-connection weights between each unit domain feature vector and the feature interaction layer are generated through the feature routing layer, and the correlation degree between the unit domain feature vector and the feature interaction layer is represented by the cross-connection weights. Furthermore, N unit domain feature vectors are selected from the M unit domain feature vectors through the cross-connection weights to generate cross-connection feature vectors. The cross-connection feature vectors are input into the feature interaction layer connected to the feature routing layer, and the feature vectors for generating service prediction results associated with the service data of M objects are output through the feature interaction layer. It can be seen that in the embodiments of the present application, an automatic feature cross-connection structure is implemented through the feature routing layer. The feature routing layer can automatically calculate the cross-connection weights between the input features and the connected feature interaction layer, and then the features cross-connected by the feature interaction layer can be selected through the cross-connection weights, realizing the automatic screening of the cross-connected features, so that the service cross-connection model can better capture the correlation degree of the cross-connected features. Secondly, the automatic feature cross-connection can avoid the fixed mode of manual cross-connection. When facing different data distributions and task requirements, the feature cross-connection scheme can be accurately selected. When different users have different object service data, the features matching the feature routing layer can still be accurately selected through the feature routing layer, so as to better capture the personalized needs of users and improve the accuracy of the service prediction results generated by the service cross-connection model.

[0255] On the other hand, the automatic cross-connection structure avoids the complicated and fixed feature connections in the manual feature cross-connection scheme, reducing the complexity of the model to a certain extent, making the training and inference of the model more efficient, and improving the generalization ability and adaptability of the model. At the same time, through dimensionality reduction processing, the embodiments of the present application can convert the features processed by each feature routing layer into the same vector dimension, and map the high-dimensional features to the low-dimensional space through dimensionality reduction encoding, so as to reduce the dimension of the data without losing important information, reduce the time consumption and computational complexity in the calculation process, and improve the training efficiency and generalization ability of the model. At the same time, through dimensionality reduction processing, the total dimension of the cross-connected features is also limited, avoiding the situation of "the guest overpowering the host" caused by too many cross-connected features, and ensuring the stability and reliability of the model.

[0256] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the functions of that module or unit.

[0257] Please refer to Figure 9 , Figure 9 which is a structural schematic diagram of a data processing device provided by the embodiments of the present application. Figure 2 . As Figure 9 shown, the data processing device 2 includes a sample data acquisition module 910, a sample feature encoding module 920, a sample weight processing module 930, a sample cross-connection processing module 940, a result prediction module 950, and a model training module 960.

[0258] The sample data acquisition module 910 is configured to acquire service data of M sample objects, and input the service data of the M sample objects into an initial cross-connection model; the service fields corresponding to the service data of the M sample objects are different from each other; the initial cross-connection model includes an initial encoding layer, an initial feature interaction layer, and an initial feature routing layer; M is a positive integer;

[0259] The sample feature encoding module 920 is configured to respectively perform encoding processing on the service data of the M sample objects through the initial encoding layer to obtain M sample unit domain feature vectors, and input the M sample unit domain feature vectors into the initial feature interaction layer connected to the initial encoding layer;

[0260] The sample weight processing module 930 is configured to input the M sample unit domain feature vectors into the initial feature routing layer, and generate sample cross-connection weights between each sample unit domain feature vector and the initial feature interaction layer through the initial feature routing layer; the sample cross-connection weights are used to characterize the correlation degree between the sample unit domain feature vectors and the sample feature interaction layer;

[0261] The sample cross-connection processing module 940 is configured to select N sample unit domain feature vectors from the M sample unit domain feature vectors according to the sample cross-connection weights, generate a sample cross-connection feature vector according to the N sample unit domain feature vectors, and input the sample cross-connection feature vector into the initial feature interaction layer connected to the initial feature routing layer; N is a positive integer less than or equal to M;

[0262] The result prediction module 950 is configured to generate an initial prediction result associated with the service data of the M sample objects through the sample cross-connection feature vector and the M sample unit domain feature vectors in the initial feature interaction layer;

[0263] A model training module 960, configured to adjust model parameters of an initial encoding layer, an initial feature interaction layer, and an initial feature routing layer based on annotation results and initial prediction results respectively corresponding to M sample object service data, so as to obtain a service cross-connection model composed of an encoding layer, a feature interaction layer, and a feature routing layer; the service cross-connection model is used to generate a service prediction result corresponding to the object service data.

[0264] In the embodiment of the present application, the initial cross-connection model is trained into a service cross-connection model, and the service cross-connection model may include an encoding layer, a feature routing layer, and a feature interaction layer. When the task domain processed by the feature interaction layer in the service cross-connection model changes, the feature routing layer and the feature interaction layer can be retrained to achieve adaptive changes in the feature routing layer and adjust the features cross-connected by the feature interaction layer. The encoding layer in the service cross-connection model performs encoding processing on M object service data with different service domains to obtain M unit domain feature vectors. The M unit domain feature vectors are input into the feature interaction layer connected to the encoding layer, and a corresponding feature routing layer is configured for the feature interaction layer. The M unit domain feature vectors are input into the feature routing layer, and the cross-connection weights between each unit domain feature vector and the feature interaction layer are generated through the feature routing layer. The cross-connection weights represent the correlation degree between the unit domain feature vector and the feature interaction layer. Furthermore, N unit domain feature vectors are selected from the M unit domain feature vectors through the cross-connection weights to generate cross-connection feature vectors. The cross-connection feature vectors are input into the feature interaction layer connected to the feature routing layer, and the feature interaction layer outputs a feature vector used to generate a service prediction result associated with the M object service data. It can be seen that in the embodiment of the present application, an automatic feature cross-connection structure is implemented through the feature routing layer. The feature routing layer can automatically calculate the cross-connection weights between the input features and the connected feature interaction layer, and then can select the features cross-connected by the feature interaction layer through the cross-connection weights, realizing automatic screening of the cross-connected features, so that the service cross-connection model can better capture the correlation degree of the cross-connected features. Secondly, the automatic feature cross-connection can avoid the fixed mode of manual cross-connection. When facing different data distributions and task requirements, the feature cross-connection scheme can be accurately selected. When different users have different object service data, the feature routing layer can still accurately select the features matching the feature routing layer, so as to better capture the personalized needs of users and improve the accuracy of the service estimation results generated by the service cross-connection model.

[0265] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0266] Please refer to Figure 10 , Figure 10 which is a schematic structural diagram of a computer device provided by the embodiments of the present application. As Figure 10 shown, the computer device 1000 may include: a processor 1001, a network interface 1004, and a memory 1005. In addition, the above computer device 1000 may further include: a user interface 1003 and at least one communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. Among them, the user interface 1003 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. Optionally, the memory 1005 may also be at least one storage device located far from the aforementioned processor 1001. As Figure 10 shown, the memory 1005, as a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.

[0267] In the computer device 1000 as Figure 10 shown, the network interface 1004 can provide a network communication network element; while the user interface 1003 is mainly used to provide an input interface for the user; and the processor 1001 can be used to call the device control application program stored in the memory 1005.

[0268] When the computer device is a data processing device 1, it is used to implement:

[0269] Obtain M object service data, and input the M object service data into a service cross-connection model; the service fields corresponding to the M object service data are different from each other; the service cross-connection model includes an encoding layer, a feature interaction layer, and a feature routing layer; M is a positive integer;

[0270] The encoding layer performs encoding processing on the M object service data respectively to obtain M unit domain feature vectors, and inputs the M unit domain feature vectors into a feature interaction layer connected to the encoding layer;

[0271] The M unit domain feature vectors are input into the feature routing layer, and the feature routing layer generates cross-connection weights between each unit domain feature vector and the feature interaction layer respectively; the cross-connection weights are used to represent the correlation degree between the unit domain feature vector and the feature interaction layer;

[0272] N unit domain feature vectors are selected from the M unit domain feature vectors according to the cross-connection weights, a cross-connection feature vector is generated according to the N unit domain feature vectors, the cross-connection feature vector is input into a feature interaction layer connected to the feature routing layer, and the feature interaction layer outputs a feature vector for generating a service prediction result associated with the M object service data; N is a positive integer less than or equal to M.

[0273] When the computer device is a data processing device 2, it is used to implement:

[0274] Obtain M sample object service data, and input the M sample object service data into an initial cross-connection model; the service fields corresponding to the M sample object service data are different from each other; the initial cross-connection model includes an initial encoding layer, an initial feature interaction layer and an initial feature routing layer; M is a positive integer;

[0275] The initial encoding layer performs encoding processing on the M sample object service data respectively to obtain M sample unit domain feature vectors;

[0276] The M sample unit domain feature vectors are input into the initial feature routing layer, and the initial feature routing layer generates sample cross-connection weights between each sample unit domain feature vector and the initial feature interaction layer respectively; the sample cross-connection weights are used to represent the correlation degree between the sample unit domain feature vector and the sample feature interaction layer;

[0277] N sample unit domain feature vectors are selected from the M sample unit domain feature vectors according to the sample cross-connection weights, a sample cross-connection feature vector is generated according to the N sample unit domain feature vectors, and the sample cross-connection feature vector is input into an initial feature interaction layer connected to the initial feature routing layer; N is a positive integer less than or equal to M;

[0278] In the initial feature interaction layer, an initial prediction result associated with the M sample object service data is generated through the sample cross-connection feature vector and the M sample unit domain feature vectors;

[0279] Based on the annotation results corresponding to the service data of the M sample objects and the initial prediction results, the model parameters of the initial encoding layer, the initial feature interaction layer, and the initial feature routing layer are adjusted to obtain a service cross-connection model composed of an encoding layer, a feature interaction layer, and a feature routing layer; the service cross-connection model is used to generate a service prediction result corresponding to the object service data.

[0280] It should be understood that the computer device 1000 described in the embodiments of the present application can execute the description of the data processing method in any of the foregoing Figure 3 、 Figure 4 and Figure 7 corresponding embodiments, which will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either.

[0281] In addition, it should be pointed out here that: the embodiments of the present application also provide a computer-readable storage medium, and a computer program is stored in the above computer-readable storage medium. When the above processor executes the above computer program, it can execute the description of the above data processing method in any of the foregoing Figure 3 、 Figure 4 and Figure 7 corresponding embodiments, so it will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in the present application, please refer to the description of the method embodiments of the present application.

[0282] The above computer-readable storage medium may be the data processing device provided in any of the foregoing embodiments or the internal storage unit of the above computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store the data that has been displayed or will be displayed.

[0283] In addition, it should be pointed out here that: the embodiments of the present application also provide a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the foregoingFigure 3 , Figure 4 and Figure 7 any of the methods provided by the corresponding embodiments.

[0284] The terms "first", "second", etc. in the specification, claims and drawings of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product or device that includes a series of steps or units is not limited to the listed steps or modules, but optionally further includes steps or modules not listed, or optionally further includes other step units inherent to these processes, methods, apparatuses, products or devices.

[0285] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of network elements in the above description. Whether these network elements are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described network elements for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0286] The methods and related apparatuses provided by the embodiments of the present application are described with reference to the method flowcharts and / or structural schematic diagrams provided by the embodiments of the present application. Specifically, each process and / or block of the method flowchart and / or structural schematic diagram, and the combination of the processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1The functions specified in one or more boxes. These computer program instructions can also be loaded onto a computer or other programmable device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing in the process Figure 1 steps of one or more processes and / or structural diagrams that illustrate the functions specified in one or more boxes.

[0287] The steps in the method embodiments of this application can be adjusted, combined, and deleted according to actual needs.

[0288] The modules in the device embodiments of this application can be combined, divided, and deleted according to actual needs.

[0289] The foregoing disclosure is only for the better embodiments of this application. Of course, it cannot be used to limit the scope of rights of this application. Therefore, equivalent changes made according to the claims of this application still fall within the scope covered by this application.

Claims

1. A data processing method, characterized in that, Including: Obtain business data of M objects, and input the business data of the M objects into a business cross-connection model; the business fields corresponding to the business data of the M objects are different from each other; the business cross-connection model includes an encoding layer, a feature interaction layer, and a feature routing layer; M is a positive integer; Perform encoding processing on the business data of the M objects respectively through the encoding layer to obtain M unit domain feature vectors, and input the M unit domain feature vectors into a feature interaction layer connected to the encoding layer; Input the M unit domain feature vectors into the feature routing layer, and generate cross-connection weights between each unit domain feature vector and the feature interaction layer through the feature routing layer; the cross-connection weights are used to characterize the correlation degree between the unit domain feature vectors and the feature interaction layer; Select N unit domain feature vectors from the M unit domain feature vectors according to the cross-connection weights, generate a cross-connection feature vector according to the N unit domain feature vectors, input the cross-connection feature vector into a feature interaction layer connected to the feature routing layer, and output a feature vector for generating a business prediction result associated with the business data of the M objects through the feature interaction layer; N is a positive integer less than or equal to M.

2. The method according to claim 1, wherein The step of inputting the M unit domain feature vectors into the feature routing layer and generating cross-connection weights between each unit domain feature vector and the feature interaction layer through the feature routing layer includes: Input the M unit domain feature vectors into the feature routing layer, and perform dimensionality reduction processing on the M unit domain feature vectors to obtain M dimensionality-reduced feature vectors; the vector dimensions of the M dimensionality-reduced feature vectors are the same; Obtain a routing input matrix and a routing output matrix associated with the feature interaction layer in the feature routing layer, and perform dot product operations on the M dimensionality-reduced feature vectors and the routing input matrix respectively to obtain M routing feature vectors; the routing input matrix and the routing output matrix are jointly used to characterize the network layer attributes of the feature interaction layer; Generate linear activation vectors corresponding to the M routing feature vectors respectively through a rectified linear activation function, and perform dot product operations on the M linear activation vectors and the routing output matrix respectively to obtain M linear activation parameters; Perform piecewise parameter mapping on the M linear activation parameters through a piecewise activation function to obtain cross-connection weights between each unit domain feature vector and the feature interaction layer.

3. The method according to claim 2, wherein The step of performing dimensionality reduction processing on the M unit domain feature vectors to obtain M dimensionality-reduced feature vectors includes: Obtain the feature dimension values corresponding to the M unit domain feature vectors respectively, and obtain the projection parameter matrices corresponding to the M unit domain feature vectors respectively based on the M feature dimension values; Perform projection processing on the corresponding unit domain feature vectors respectively through the M projection parameter matrices to obtain M dimensionality-reduced feature vectors.

4. The method according to claim 2, wherein The cross-connection weights include a first weight value, a second weight value, and a third weight value; the step of performing piecewise parameter mapping on the M linear activation parameters through a piecewise activation function to obtain cross-connection weights between each unit domain feature vector and the feature interaction layer includes: Obtain the first activation threshold and the second activation threshold corresponding to the piecewise activation function; If the linear activation parameter corresponding to the unit domain feature vector is less than the first activation threshold, set the cross-connection weight corresponding to the unit domain feature vector to the first weight value; If the linear activation parameter corresponding to the unit domain feature vector is greater than the first activation threshold and less than the second activation threshold, generate the second weight value based on the linear activation parameter, and set the cross-connection weight corresponding to the unit domain feature vector to the second weight value; If the linear activation parameter corresponding to the unit domain feature vector is greater than the second activation threshold, set the cross-connection weight corresponding to the unit domain feature vector to the third weight value; Wherein, the degrees of correlation between the unit domain feature vectors represented by the first weight value, the second weight value, and the third weight value and the feature interaction layer are different from each other; the degree of correlation corresponding to the first weight value is less than the degree of correlation corresponding to the second weight value, and the degree of correlation corresponding to the second weight value is less than the degree of correlation corresponding to the third weight value.

5. The method according to claim 1, wherein The cross-connection weight includes the weight values corresponding to the M unit domain feature vectors respectively; the selecting N unit domain feature vectors from the M unit domain feature vectors according to the cross-connection weight includes: Determine the N unit domain feature vectors as the unit domain feature vectors with cross-connection weights greater than the weight threshold among the M unit domain feature vectors; The generating a cross-connection feature vector according to the N unit domain feature vectors includes: Generate a mask weight matrix according to the weight values corresponding to the N unit domain feature vectors respectively, and obtain the reduced-dimensional feature vectors corresponding to the N unit domain feature vectors respectively; the reduced-dimensional feature vectors are obtained by performing dimensionality reduction processing on the unit domain feature vectors; Perform a dot product operation on the sequence composed of the N reduced-dimensional feature vectors and the mask weight matrix to obtain a cross-connection feature vector.

6. The method according to claim 1, wherein The number of the feature interaction layer and the feature routing layer in the service cross-connection model is both T. Each feature interaction layer is respectively connected to a feature routing layer, and the feature routing layers connected to each feature interaction layer are different from each other. The T feature interaction layers are connected in series; the input of each feature routing layer is the M unit domain feature vectors, and the output of each feature routing layer is used for input to the connected feature interaction layer; The T feature interaction layers include the feature interaction layer H i and the feature interaction layer H i-1 , where i is a positive integer less than T; the T feature routing layers include the feature routing layer B i connected to the feature interaction layer H i ; the output of the feature interaction layer to generate a feature vector associated with the business data of the M objects for the business prediction result includes: In the feature interaction layer H i Based on the target feature vector corresponding to the feature interaction layer H i and the cross-connected feature vector input by the feature routing layer B i an output feature vector is output; if the feature interaction layer H i is a feature interaction layer connected to the encoding layer, the target feature vector is the M unit domain feature vectors; if the feature interaction layer H i is connected to the feature interaction layer H i-1 the target feature vector is the feature vector output by the feature interaction layer H i-1 ; The method further includes: If the feature interaction layer H i is the last one among the T feature interaction layers, then a service prediction result associated with the M object service data is generated according to the feature vector output by the feature interaction layer H i ​ 7. The method according to claim 1, wherein The service cross-connection model includes a gating network layer and S expert network layers, S is a positive integer, and each expert network layer includes a candidate feature interaction layer and a candidate feature routing layer. The service domains corresponding to the S expert network layers are different from each other; Before inputting the M unit domain feature vectors to the feature routing layer, the method further includes: Input the M unit domain feature vectors to the gating network layer, and generate the gating activation values corresponding to each expert network layer respectively through the gating network layer; the gating activation values are used to represent the degree of correlation between the expert network layer and the M unit domain feature vectors; Based on the S gating activation values, sort the S expert network layers, determine the target expert network layer among the sorted S expert network layers, determine the candidate feature interaction layer in the target expert network layer as the feature interaction layer, and determine the candidate feature routing layer in the target expert network layer as the feature routing layer.

8. The method according to claim 1, wherein The output of the feature interaction layer for generating a feature vector associated with the M object service data includes: In the feature interaction layer, obtain the media content feature vectors corresponding to Q media contents in the media content database, perform vector concatenation on the Q media content feature vectors to obtain a media content feature sequence; Q is a positive integer; Perform fusion processing on the cross-connected feature vector and the M unit domain feature vectors to obtain a fusion feature vector, and perform attention processing on the media content feature sequence and the fusion feature vector to obtain a feature vector containing an attention result vector; The method further includes: Based on the attention scores corresponding to the Q media contents in the attention result vector, generate estimated interaction parameters corresponding to the Q media contents, and determine the estimated interaction parameters corresponding to the Q media contents as the service prediction results associated with the M object service data.

9. The method according to claim 8, wherein The performing attention processing on the media content feature sequence and the fusion feature vector to obtain a feature vector containing an attention result vector includes: Obtain the query parameter matrix, key parameter matrix, and value parameter matrix in the feature interaction layer; the query parameter matrix, the key parameter matrix, and the value parameter matrix are all matrices composed of learnable parameters; Perform a dot product operation on the fusion feature vector and the query parameter matrix to obtain a query vector, perform a dot product operation on the media content feature sequence and the key parameter matrix to obtain a key vector, and perform a dot product operation on the media content feature sequence and the value parameter matrix to obtain a value vector; Generate an attention score vector based on the query vector and the key vector, perform dimensionality reduction processing on the attention score vector based on the vector length of the key vector, perform normalization processing on the dimensionally reduced attention score vector to obtain an attention weight vector, and perform a dot product operation on the attention weight vector and the value vector to obtain a feature vector containing an attention result vector.

10. The method according to claim 8, characterized in that, The performing fusion processing on the cross-connected feature vector and the M unit domain feature vectors to obtain a fusion feature vector includes: If the cross-connected feature vector and the M unit domain feature vectors satisfy the isomorphic feature condition, perform vector addition on the cross-connected feature vector and the M unit domain feature vectors to obtain a fusion feature vector; the isomorphic feature condition means that the cross-connected feature vector and the M unit domain feature vectors belong to the same semantic space; If the cross-connected feature vector and the M unit domain feature vectors do not satisfy the isomorphic feature condition, perform vector concatenation on the cross-connected feature vector and the M unit domain feature vectors to obtain a fusion feature vector.

11. A data processing method, characterized in that, Includes: Obtain the business data of M sample objects, and input the business data of the M sample objects into the initial cross-connection model; the business fields corresponding to the business data of the M sample objects are different from each other; the initial cross-connection model includes an initial encoding layer, an initial feature interaction layer, and an initial feature routing layer; M is a positive integer; Through the initial encoding layer, encode the business data of the M sample objects respectively to obtain M sample unit domain feature vectors, and input the M sample unit domain feature vectors into the initial feature interaction layer connected to the initial encoding layer; Input the M sample unit domain feature vectors into the initial feature routing layer, and generate the cross-connection weights between each sample unit domain feature vector and the initial feature interaction layer through the initial feature routing layer; the cross-connection weights are used to characterize the correlation degree between the sample unit domain feature vector and the sample feature interaction layer; Select N sample unit domain feature vectors from the M sample unit domain feature vectors according to the sample cross-connection weights, generate a sample cross-connection feature vector according to the N sample unit domain feature vectors, and input the sample cross-connection feature vector into the initial feature interaction layer connected to the initial feature routing layer; N is a positive integer less than or equal to M; In the initial feature interaction layer, generate an initial prediction result associated with the business data of the M sample objects through the sample cross-connection feature vector and the M sample unit domain feature vectors; Based on the annotation results corresponding to the business data of the M sample objects and the initial prediction results, adjust the model parameters of the initial encoding layer, the initial feature interaction layer, and the initial feature routing layer to obtain a business cross-connection model composed of an encoding layer, a feature interaction layer, and a feature routing layer; the business cross-connection model is used to generate a business prediction result corresponding to the object business data.

12. A data processing device, characterized in that, Include: A data acquisition module, configured to acquire the business data of M objects, and input the business data of the M objects into the business cross-connection model; the business fields corresponding to the business data of the M objects are different from each other; the business cross-connection model includes an encoding layer, a feature interaction layer, and a feature routing layer; M is a positive integer; A feature encoding module, configured to encode the business data of the M objects respectively through the encoding layer to obtain M unit domain feature vectors, and input the M unit domain feature vectors into the feature interaction layer connected to the encoding layer; A weight processing module, configured to input the M unit domain feature vectors into the feature routing layer, and generate cross-connection weights between each unit domain feature vector and the feature interaction layer through the feature routing layer; the cross-connection weights are used to characterize the correlation degree between the unit domain feature vector and the feature interaction layer; A cross - connection processing module, configured to select N unit - domain feature vectors from the M unit - domain feature vectors according to the cross - connection weights, generate a cross - connection feature vector based on the N unit - domain feature vectors, input the cross - connection feature vector into a feature interaction layer connected to the feature routing layer, and output, through the feature interaction layer, a feature vector for generating a service prediction result associated with the M object service data; N is a positive integer less than or equal to M.

13. A data processing device, characterized in that, It includes: A sample data acquisition module, configured to acquire M sample object service data and input the M sample object service data into an initial cross - connection model; the service fields corresponding to the M sample object service data are different from each other; the initial cross - connection model includes an initial encoding layer, an initial feature interaction layer, and an initial feature routing layer; M is a positive integer; A sample feature encoding module, configured to respectively perform encoding processing on the M sample object service data through the initial encoding layer to obtain M sample unit - domain feature vectors, and input the M sample unit - domain feature vectors into an initial feature interaction layer connected to the initial encoding layer; A sample weight processing module, configured to input the M sample unit - domain feature vectors into the initial feature routing layer, and generate sample cross - connection weights between each sample unit - domain feature vector and the initial feature interaction layer through the initial feature routing layer; the sample cross - connection weights are used to characterize the correlation degree between the sample unit - domain feature vectors and the sample feature interaction layer; A sample cross - connection processing module, configured to select N sample unit - domain feature vectors from the M sample unit - domain feature vectors according to the sample cross - connection weights, generate a sample cross - connection feature vector based on the N sample unit - domain feature vectors, and input the sample cross - connection feature vector into an initial feature interaction layer connected to the initial feature routing layer; N is a positive integer less than or equal to M; A result prediction module, configured to generate an initial prediction result associated with the M sample object service data through the sample cross - connection feature vector and the M sample unit - domain feature vectors in the initial feature interaction layer; A model training module, configured to adjust model parameters of the initial encoding layer, the initial feature interaction layer, and the initial feature routing layer based on the annotation results corresponding to the M sample object service data and the initial prediction results, so as to obtain a service cross - connection model composed of an encoding layer, a feature interaction layer, and a feature routing layer; the service cross - connection model is used to generate a service prediction result corresponding to object service data.

14. A computer device, characterized in that, It includes: A processor, a memory, and a network interface; The processor is connected to the memory and the network interface. Among them, the network interface is used to provide data communication functions, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method according to any one of claims 1 - 11.

15. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and the computer program is adapted to be loaded and executed by a processor so that a computer device having the processor executes the method according to any one of claims 1-11.

16. A computer program product, characterized in that, The computer program product includes a computer program, the computer program is stored in a computer-readable storage medium, and is adapted to be read and executed by a processor so that a computer device having the processor executes the method according to any one of claims 1-11.