Method and system based on distributed multivariate data analysis

By performing hidden operations on distributed multivariate data and dynamic update node extraction, combined with marker processing and feature enhancement, the problems of data privacy protection and dynamic change capture are solved, and data analysis efficiency and prediction capabilities of machine learning models are improved.

CN119918075BActive Publication Date: 2025-08-15广东柏瑞数据技术有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411988993.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-08-15
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

The prior art is difficult to effectively protect the privacy in distributed multivariate data, it is difficult to capture dynamic changing characteristics, and the machine learning model training is inefficient, making it difficult to adapt to complex and diverse distributed multivariate data.

Method used

By performing hidden operations on derivative field data and dynamic update node extraction, hidden distributed multivariate data are generated, and training templates for machine learning network models are built through tag processing and feature enhancement.

Benefits of technology

Effectively protect data privacy, accurately capture dynamic changes moments, improve data analysis efficiency and accuracy, and enhance the prediction accuracy and generalization capabilities of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119918075B_ABST
    Figure CN119918075B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a method and system based on distributed multivariate data analysis, which accurately captures the key change moments in the data by hiding the derived field data and further extracting the nodes through dynamic updates, providing more detailed time dimension information for data analysis. The labeling processing of the derived field data simplifies the complexity of the data and improves the efficiency and accuracy of subsequent analysis. The feature enhancement of hidden data based on the target derived template knowledge data, dynamically updated node data and derived field labeling data not only enriches the feature dimension of the data, but also incorporates business logic and knowledge experience, significantly improving the depth and breadth of data analysis. Finally, by constructing a training template for the machine learning network model, the model can learn data features more accurately, improving the prediction accuracy and generalization ability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and system based on distributed multivariate data analysis. Background Art

[0002] With the rapid development of information technology, all types of data are growing at an unprecedented rate, especially multi-dimensional data in distributed environments, such as mobile payment data, personnel database data, scene monitoring data, regulatory platform data, and behavior log platform data. This data comes from a wide range of sources and in various formats, and contains rich business information and value. However, directly analyzing this complex and massive distributed multi-dimensional data faces many challenges, including but not limited to insufficient data privacy protection, difficulty in capturing dynamic changes, difficulty in extracting data features, and inefficient model training.

[0003] First, distributed multivariate data often contains a large amount of sensitive information, such as personal identity information and transaction amounts. If this information is not properly processed and used directly for analysis, it will pose a serious risk of privacy leakage. Therefore, how to effectively utilize this data for in-depth analysis while protecting data privacy has become a pressing issue.

[0004] Secondly, some fields in distributed multivariate data (such as transaction time and monitoring time) are dynamically changing. These dynamic fields are crucial for understanding business trends and predicting future trends. However, traditional data analysis methods often struggle to effectively capture the key nodes in these dynamic changes, resulting in delayed or biased analysis results.

[0005] Finally, when building machine learning models for data analysis, the design of the training template directly impacts model performance. Traditional training templates are often designed based on a single type of data and are difficult to adapt to the complexity and diversity of distributed multivariate data. Therefore, designing appropriate training templates for machine learning network models based on the characteristics of distributed multivariate data to improve the model's predictive accuracy and generalization capabilities is a major challenge facing the current data analysis field. Summary of the Invention

[0006] In order to at least overcome the above-mentioned deficiencies in the prior art, the embodiment of the present application aims to provide a method and system based on distributed multivariate data analysis.

[0007] According to one aspect of the present application, a method based on distributed multivariate data analysis is provided, the method comprising:

[0008] Performing a hiding operation on the to-be-derived field data in the target distributed multivariate data to generate hidden distributed multivariate data, wherein the target distributed multivariate data includes at least two types of data selected from the group consisting of mobile payment data, personnel database data, scene monitoring data, supervision platform data, and behavior log platform data;

[0009] Dynamically updating node extraction is performed on the dynamic field data in the target distributed multivariate data to generate dynamic updating node data; the field data to be derived includes the dynamic field data to be derived in the dynamic field data;

[0010] Marking the to-be-derived field data in the target distributed multivariate data to generate derived field marked data;

[0011] Based on the target derived template knowledge data, the dynamically updated node data and the derived field tag data, feature enhancement is performed on the hidden distributed multivariate data to generate derived distributed multivariate data that combines the derived feature dimensions expressed by the target derived template knowledge data;

[0012] A training template for a corresponding machine learning network model is constructed based on the derived distributed multivariate data and the target distributed multivariate data.

[0013] According to one aspect of an embodiment of the present application, a computer program product or computer program is provided. The computer program product or computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in various optional implementations of the above three aspects.

[0014] In the technical solutions provided in some embodiments of the present application, the embodiments of the present application effectively protect sensitive information in the data, enhance the security of the data, and avoid the risk of privacy leakage by hiding the derived field data. The method further accurately captures the key change moments in the data by dynamically updating node extraction, providing more refined time dimension information for data analysis. The labeling processing of derived field data simplifies the complexity of the data and improves the efficiency and accuracy of subsequent analysis. Feature enhancement of hidden data based on target derived template knowledge data, dynamically updated node data and derived field labeling data not only enriches the feature dimension of the data, but also incorporates business logic and knowledge experience, significantly improving the depth and breadth of data analysis. Finally, by constructing a training template for the machine learning network model, the model can learn data features more accurately, improving the prediction accuracy and generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required to be used in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be extracted in combination with these drawings without creative work.

[0016] Figure 1 A flow chart of a method based on distributed multivariate data analysis provided in an embodiment of the present application;

[0017] Figure 2 A schematic block diagram of the architecture of a system based on distributed multivariate data analysis provided in an embodiment of the present application. DETAILED DESCRIPTION

[0018] The following description is intended to enable one of ordinary skill in the art to practice and incorporate the present application, and is provided in the context of a specific application scenario and its requirements. It will be apparent to one of ordinary skill in the art that various modifications may be made to the disclosed embodiments, and that the general principles defined herein may be applied to other embodiments and application scenarios without departing from the principles and scope of the present application. Therefore, the present application is not limited to the described embodiments, but should be accorded the broadest scope consistent with the claims.

[0019] Figure 1 This is a flow chart of a method based on distributed multivariate data analysis provided by an embodiment of the present application. The method based on distributed multivariate data analysis is introduced in detail below.

[0020] Step S110, performing a hiding operation on the field data to be derived in the target distributed multivariate data to generate hidden distributed multivariate data, wherein the target distributed multivariate data includes at least two types of data among mobile payment data, personnel database data, scene monitoring data, supervision platform data and behavior log platform data.

[0021] Specifically, the target distributed multivariate data refers to a collection of data from various sources and types, originating from different systems or platforms. This data includes, but is not limited to, mobile payment data, personnel database data, scenario monitoring data, regulatory platform data, and behavioral log platform data. These data typically have different structures and formats, but all serve a common business objective or analytical need.

[0022] The field data to be derived are those fields in the target distributed multivariate data that require further processing or transformation to extract more information or features. These fields to be derived may contain sensitive information or require specific processing to realize their potential value. The hiding operation is the process of encrypting, replacing or obfuscating sensitive or protected field data to be derived to ensure the security or privacy of the data while retaining its possibility for analysis. The hidden distributed multivariate data is the target distributed multivariate data after the hiding operation, in which the field data to be derived has been encrypted, replaced or obfuscated to protect sensitive information.

[0023] For example, in this embodiment, the target distributed multivariate data received by the server includes mobile payment data, personnel database data, and scene monitoring data. The mobile payment data includes fields such as transaction amount, transaction time, and payment method; the personnel database data includes fields such as name, age, and ID number; and the scene monitoring data includes fields such as monitoring location, monitoring time, and personnel flow.

[0024] Among these data, the field data to be derived is set to the transaction amount in the mobile payment data (for example, because the amount-related information may need special processing to protect privacy or perform special feature engineering) and the age field in the personnel database data.

[0025] The server performs a hiding operation. For example, for the transaction amount field, an encryption algorithm (such as the AES encryption algorithm) is used to encrypt each transaction amount value. For the age field, a replacement strategy can be used to replace the actual age with a random number generated according to a certain rule (such as a number randomly generated within a range of 5 years above or below the actual age). After this operation, the transaction amount and age fields are hidden in the generated hidden distributed multivariate data, existing in an encrypted or disguised form, while the other fields remain unchanged.

[0026] Step S120: extract dynamic update nodes from the dynamic field data in the target distributed multivariate data to generate dynamic update node data. The field data to be derived includes the dynamic field data to be derived in the dynamic field data.

[0027] Specifically, dynamic field data refers to time-varying data fields within the target distributed multivariate data, such as transaction time and monitoring time. This data reflects the real-time status or activity of the system. Dynamically updated node data is key time points or time periods extracted from the dynamic field data. These time points or time periods represent moments of significant data change or special significance.

[0028] For example, in this embodiment, still taking the above-mentioned target distributed multivariate data as an example, the dynamic field data includes the transaction time in the mobile payment data (because the transaction time will be continuously updated with each transaction) and the personnel flow in the scene monitoring data (the personnel flow at different times is dynamically changing).

[0029] The server dynamically updates the transaction time in the mobile payment data and extracts nodes. For example, a day is divided into 24 nodes by hour, and the number of transactions in each hour is counted. When the number of transactions exceeds a certain threshold (such as 100 times), this hour is marked as a dynamic update node. For the flow of people in the scene monitoring data, the server counts it in time periods of 15 minutes. If the change in the flow of people in two adjacent 15-minute time periods exceeds 50 people, this time point is marked as a dynamic update node. Through such operations, dynamic update node data is generated, which can reflect the key node information of the data in the dynamic change process.

[0030] Step S130 : performing labeling processing on the to-be-derived field data in the target distributed multivariate data to generate derived field label data.

[0031] Continuing to use the target distributed multivariate data mentioned above, the server performs labeling processing on the derived field data (i.e., the transaction amount in the mobile payment data and the age in the personnel database data).

[0032] For example, for the transaction amount field, if the amount is less than 100 yuan, it is marked as a "small transaction" and represented by the number 1; if the amount is between 100 yuan and 1,000 yuan, it is marked as a "medium transaction" and represented by the number 2; if the amount is greater than 1,000 yuan, it is marked as a "large transaction" and represented by the number 3. For the age field, if the age is less than 18, it is marked as a "minor" and represented by the number 1; if the age is between 18 and 60, it is marked as an "adult" and represented by the number 0; if the age is greater than 60, it is marked as a "senior" and represented by the number 1. This generates derived field labeled data, which can play an auxiliary role in subsequent operations such as feature enhancement.

[0033] Step S140, based on the target derived template knowledge data, the dynamically updated node data and the derived field tag data, feature enhancement is performed on the hidden distributed multivariate data to generate derived distributed multivariate data that combines the derived feature dimensions expressed by the target derived template knowledge data.

[0034] Specifically, the target-derived template knowledge data is constructed based on business logic and data analysis experience. It contains knowledge patterns about the relationships between fields in the target distributed multivariate data and is used to guide the feature enhancement process. Feature enhancement refers to the process of processing hidden distributed multivariate data using the target-derived template knowledge data, dynamically updated node data, and derived field tag data to add new features or improve existing features. This helps improve the accuracy and effectiveness of data analysis.

[0035] The derived distributed multivariate data is hidden distributed multivariate data that has been subjected to feature enhancement processing, and is combined with the derived feature dimensions expressed by the target derived template knowledge data, and has richer and more valuable information.

[0036] For example, in this embodiment, it is assumed that the target derived template knowledge data is constructed in advance based on business needs and data analysis experience, and includes, for example, some relationship patterns between transaction amount, age and other data features, such as the knowledge pattern "large transactions and adult groups are more inclined to make mobile payments during specific time periods on weekdays."

[0037] The server first embeds the target-derived template knowledge data, hidden distributed multivariate data, and dynamically updated node data. For example, for the target-derived template knowledge data, the relational patterns therein are converted into a first embedding feature vector using a word vector model (such as a variant of Word2Vec, suitable for processing such non-textual but logically relational data). For the hidden distributed multivariate data, each field (after the hiding operation) is extracted and converted into a second embedding feature vector using a multi-layer neural network (such as a fully connected neural network with three hidden layers). For the dynamically updated node data, it is converted into a third embedding feature vector using a time series-based embedding model (such as a variant of Transformer).

[0038] Then, based on the first embedded feature vector, the second embedded feature vector, the third embedded feature vector and the derived field label data, the randomly generated distributed multivariate data is iteratively optimized. For example, the randomly generated distributed multivariate data are initially some numerical vectors randomly generated according to a certain distribution (such as a normal distribution), representing the hypothetical feature data. The first embedded feature vector is feature transformed using a feature transformation model (assuming that this model contains 5 first feature focusing units) to generate feature transformation data. Each first feature focusing unit may focus on a different relationship part in the target derived template knowledge data. For example, one unit focuses on the relationship between transaction amount and time, and another unit focuses on the relationship between age and payment method, etc.

[0039] Using the second embedded feature vector, the third embedded feature vector, and the derived field tag data as iterative transfer conditions (assuming this also includes the business logic description document of the target distributed multivariate data, which describes the data source, usage, and business logic relationships between the data), an iterative optimization model (including five second feature focusing units) is used for iterative optimization. For example, three rounds (Y = 3) of optimization iterations are performed.

[0040] First round of optimization iteration: Obtain the reference embedding feature vector of the first round of optimization iteration. This vector is generated by fusing the second embedding feature vector, the third embedding feature vector, the derived field tag data, and the random embedding feature vector corresponding to the randomly generated distributed multivariate data.

[0041] In the first second feature focusing unit, interactive calculations are performed based on the first feature focusing result generated by the first first feature focusing unit (for example, the first focusing index points to the relationship between the transaction amount and the payment method, and the first focusing feature is a quantitative representation of this relationship) and the second feature focusing result of the first second feature focusing unit for the first round of optimization iteration (determined based on the reference embedded feature vector, for example, the index points to the part related to age, and the feature is a quantitative representation related to age). For example, an interactive operation (such as addition operation) is performed on the first focusing index and the second focusing index of the first second feature focusing unit for the first round of optimization iteration to generate a target focusing index of the first second feature focusing unit for the first round of optimization iteration; an interactive operation (such as multiplication operation) is performed on the first focusing feature and the second focusing feature of the first second feature focusing unit for the first round of optimization iteration to generate a target focusing feature of the first second feature focusing unit for the first round of optimization iteration; based on the target interaction index (assuming it is a coefficient index pre-set according to business logic) and the target focusing index, an interaction factor is determined (such as by looking up a coefficient table); based on the interaction factor, an interactive calculation (such as multiplication) is performed on the corresponding target focusing feature to generate an interactive embedding feature vector generated by the first second feature focusing unit for the first round of optimization iteration.

[0042] The interactive embedding feature vector is optimized iteratively (eg, by adjusting certain element values in the vector to make it more consistent with the expected optimization target) to generate the first temporary optimized iterative embedding feature vector for the first round of optimization iteration.

[0043] Second round of optimization iteration: Obtain the reference embedding feature vector for the second round of optimization iteration. This vector is the optimized embedding feature vector for the first round of optimization iteration. Similar operations are performed in the second feature focusing unit to perform interactive calculations, generate interactive embedding feature vectors, and perform optimization iterations to generate the second temporary optimized embedding feature vector.

[0044] Optimization iteration 3: Obtain the reference embedding feature vector for the 3rd optimization iteration. This vector is the optimized embedding feature vector for the 2nd optimization iteration. Similar operations are performed in the 5th second feature focusing unit. Finally, the 5th temporary optimized embedding feature vector for the 3rd optimization iteration is used as the derived embedding feature vector.

[0045] Finally, the derived embedded feature vector is restored, for example, through an inverse transformation (the inverse of the previous embedding representation operation), to generate derived distributed multivariate data that incorporates the derived feature dimensions expressed by the target derived template knowledge data. This derived distributed multivariate data not only contains the hidden information of the original data, but also incorporates the relational patterns and feature dimensions of the target derived template knowledge data.

[0046] Step S150: constructing a training template of a corresponding machine learning network model based on the derived distributed multivariate data and the target distributed multivariate data.

[0047] In this embodiment, the server constructs a training template of a corresponding machine learning network model based on the previously generated derived distributed multivariate data and target distributed multivariate data.

[0048] For example, a neural network model can be used as the machine learning network model, with some features from the derived distributed multivariate data as one input layer, and some related features from the target distributed multivariate data as another input layer. Several hidden layers (e.g., three hidden layers, each containing 100 neurons) are then set in the middle. Based on the business objective (e.g., predicting whether a transaction behavior is risky), the output layer is set to a binary classification output (0 for no risk, 1 for risk).

[0049] When building a training template, define a loss function (such as the cross-entropy loss function) to measure the difference between the model's predicted output and the true label (assuming the actual trading risk situation is obtained through manual annotation or historical data verification). At the same time, set an optimization algorithm (such as the Adam optimization algorithm) to adjust the network model parameters (such as the connection weights and bias terms between neurons). This allows the model to be continuously optimized during training to adapt to the feature relationships between the derived distributed multivariate data and the target distributed multivariate data, thereby improving prediction accuracy.

[0050] Based on the above steps, the embodiment of the present application effectively protects sensitive information in the data, enhances the security of the data, and avoids the risk of privacy leakage by hiding the derived field data. The method further accurately captures the key change moments in the data by dynamically updating node extraction, providing more refined time dimension information for data analysis. The labeling processing of the derived field data simplifies the complexity of the data and improves the efficiency and accuracy of subsequent analysis. The feature enhancement of hidden data based on the target derived template knowledge data, dynamically updated node data and derived field labeling data not only enriches the feature dimension of the data, but also incorporates business logic and knowledge experience, significantly improving the depth and breadth of data analysis. Finally, by constructing a training template for the machine learning network model, the model can learn data features more accurately, improving the prediction accuracy and generalization ability of the model.

[0051] In a possible implementation, step S140 includes:

[0052] Step S141, embedding the target derived template knowledge data, the hidden distributed multivariate data and the dynamically updated node data respectively to generate a first embedded feature vector of the target derived template knowledge data, a second embedded feature vector of the hidden distributed multivariate data and a third embedded feature vector of the dynamically updated node data.

[0053] Step S142: performing iterative optimization processing on the randomly generated distributed multivariate data according to the first embedded feature vector, the second embedded feature vector, the third embedded feature vector and the derived field tag data to generate a derived embedded feature vector.

[0054] Step S143 , performing embedding restoration on the derived embedded feature vector to generate derived distributed multivariate data that combines the derived feature dimensions expressed by the target derived template knowledge data.

[0055] In this embodiment, it is assumed that the target derived template knowledge data in the server contains some predefined rules for mobile payment behavior, such as "during peak hours on weekdays (9:10 and 18:19), users aged between 25 and 45 are more likely to use mobile payment for shopping."

[0056] The server quantizes these rules for embedded representation. For example, weekdays are represented by the number 15, peak hours can be represented by encoding the 24 hours of a day and then using corresponding numerical intervals (for example, 9:10 is encoded as 90100, and 18:19 is encoded as 180190), and the age range of 25-45 is also encoded accordingly.

[0057] The server then uses a specially designed embedding model (e.g., a neural network-based embedding layer structure) to convert these quantized rules into a first embedding feature vector. This embedding model may be pre-trained and can convert these business logic-related data into a vector form suitable for subsequent calculations.

[0058] Taking the previously processed hidden distributed multivariate data as an example, the transaction amount of the mobile payment data was encrypted and hidden, and the age in the personnel database data was replaced and hidden.

[0059] The server extracts features from this hidden data. For mobile payment data, in addition to the hidden transaction amount, other fields include payment method (assuming cash payment is coded as 0 and electronic payment is coded as 1) and transaction location (coded according to geographic location). For personnel database data, names can be encoded as text, and ID numbers can be encoded after desensitization.

[0060] The multi-layer perceptron (MLP) model is used as an embedding representation tool to transform these processed hidden distributed multivariate data into a second embedded feature vector. The input layer of the MLP receives the encoded value of each data field, and after nonlinear transformations through multiple hidden layers, the output layer outputs the second embedded feature vector.

[0061] Consider the dynamically updated node data previously extracted from transaction time in mobile payment data and personnel flow in scene monitoring data.

[0062] For dynamically updated nodes related to transaction time, for example, we count transactions in half-hour windows and mark them as dynamically updated nodes when the number of transactions exceeds a certain threshold. The server encodes this node information, such as marking the first half-hour window as 1, the second half-hour window as 2, and so on. For dynamically updated nodes related to personnel flow, encoding is performed based on different flow threshold ranges.

[0063] The dynamic update node data is converted into a third embedding feature vector using time series embedding technology (similar to the positional encoding concept in the Transformer architecture for processing time series data). This embedding technology can capture the characteristics of the dynamic update nodes in terms of time and value changes.

[0064] Next, we assume that the feature transformation model consists of multiple submodules, each of which corresponds to a type of rule or relation in the target derived template knowledge data.

[0065] For example, for the relationship rules between mobile payment time, age and payment probability mentioned earlier, one submodule specifically handles time-related feature conversion, and the other submodule handles age-related feature conversion.

[0066] When performing feature transformation on the first embedded feature vector, each submodule operates on the corresponding portion of the vector. For example, the time-dependent submodule may transform the time-dependent portion of the first embedded feature vector according to a predefined time transformation function (e.g., transforming the time encoding according to a periodic function), thereby obtaining a portion of the feature-transformed data. Similarly, the age-dependent submodule transforms the age-dependent portion, ultimately combining the two to obtain the complete feature-transformed data.

[0067] Here, the randomly generated distributed multivariate data are assumed to be initially random numeric vectors representing hypothetical features associated with the target data. For example, the dimensions of these vectors may be similar to the dimensions of the hidden distributed multivariate data, but the initial values are randomly generated.

[0068] The iterative optimization model contains multiple second feature focusing units, assuming 3 units here.

[0069] During the first round of iterative optimization, a reference embedding feature vector for the first round of optimization iteration is first generated. The second embedding feature vector (containing the features of the hidden distributed multivariate data), the third embedding feature vector (dynamically updated node data features), the derived field tag data (such as the previously labeled age group and transaction amount level), and the random embedding feature vector corresponding to the randomly generated distributed multivariate data are fused. For example, simple vector concatenation or weighted summation can be used for fusion to obtain the reference embedding feature vector.

[0070] In the first second feature focusing unit, interactive calculations are performed based on the first feature focusing result generated by the first first feature focusing unit (assuming that the first feature focusing result includes the focusing index and feature representation of the relationship between age and payment probability in the target derived template knowledge data) and the second feature focusing result of the first second feature focusing unit for the first round of optimization iteration (age-related feature indices and feature values determined based on the reference embedded feature vector).

[0071] For the interactive operation of the focus index, it is assumed that the age-related index in the first feature focus result and the age-related index in the second feature focus result are added to obtain a new target focus index.

[0072] For the interactive operation of the focused features, the age-payment probability relationship feature value in the first feature focused result is multiplied by the age-related feature value in the second feature focused result to obtain the target focused feature.

[0073] Based on a pre-set target interaction index (eg, a weight index of age in different scenarios determined according to business logic) and a target focus index, an interaction factor is determined (eg, by searching a predefined interaction factor table).

[0074] Finally, the target focused features are interactively calculated (such as multiplication) based on the interaction factor to generate the interactive embedding feature vector generated by the first second feature focusing unit for the first round of optimization iteration.

[0075] The interactive embedding feature vector is optimized iteratively. For example, by adjusting the values of certain elements in the vector according to a predefined optimization rule (such as making the elements in the vector closer to the ideal feature values in the target derived template knowledge data), a first temporary optimized iterative embedding feature vector for the first round of optimization iteration is generated.

[0076] During the second round of iterative optimization, the reference embedding feature vector of the second round of optimization iteration is obtained. This vector is the optimized embedding feature vector of the first round of optimization iteration.

[0077] In the second second feature focusing unit, similar operations to the first round are repeated. Based on the feature focusing result generated by the second first feature focusing unit (assuming it is related to the payment method) and the second feature focusing result of the second round of optimization iteration (features related to the payment method determined based on the reference embedded feature vector), interactive calculation is performed to generate an interactive embedded feature vector, and then optimization iteration is performed to obtain the second temporary optimized iterative embedded feature vector.

[0078] During the third round of iterative optimization, the reference embedding feature vector of the third round of optimization iteration is obtained. This vector is the optimized embedding feature vector of the second round of optimization iteration.

[0079] In the third second feature focusing unit, the operation is performed according to the previous process, and finally the third temporary optimized iteration embedding feature vector for the third round of optimization iteration is used as the derived embedding feature vector.

[0080] Finally, the server performs embedding restoration on the generated derivative embedding feature vector.

[0081] Assuming that a multi-layer perceptron (MLP) was used to transform the original data into an embedded feature vector when embedding, the corresponding inverse operation is now used. For example, if the MLP performed a linear transformation (such as matrix multiplication and addition of a bias term) when embedding, then the embedding is restored by inverse matrix multiplication and subtraction of the bias term.

[0082] For quantitative encoding operations performed during the embedding process, such as encoding of transaction time and age ranges, reverse decoding is performed according to the encoding rules during embedding restoration. For example, time encoding is restored to the actual time interval, and age encoding is restored to the age range.

[0083] Through this embedding-reduction operation, the derived embedded feature vector is converted into derived distributed multivariate data that incorporates the derived feature dimensions expressed by the target derived template knowledge data. This derived distributed multivariate data not only contains the hidden information in the original hidden distributed multivariate data but also incorporates the relational patterns and feature dimensions in the target derived template knowledge data, such as the association features between different ages, payment times, and payment methods.

[0084] In a possible implementation, step S142 includes:

[0085] Step S1421: Perform feature conversion on the first embedded feature vector using a feature conversion model to generate feature conversion data.

[0086] Step S1422, using the second embedded feature vector, the third embedded feature vector and the derived field tag data as iterative transfer conditions, using an iterative optimization model, and iteratively optimizing the randomly generated distributed multivariate data according to the feature conversion data to generate the derived embedded feature vector.

[0087] In a possible implementation, the feature conversion model includes X first feature focusing units, the iterative optimization model includes X second feature focusing units, and the feature conversion data includes first feature focusing results loaded by each of the first feature focusing units.

[0088] Step S1422 includes:

[0089] The second embedded feature vector, the third embedded feature vector and the derived field label data are used as iterative transfer conditions, and an iterative optimization model is used to perform Y rounds of optimization iterations on the randomly generated distributed multivariate data according to the feature conversion data, and the optimized embedded feature vector generated by the Y-th round of optimization iteration is used as the derived embedded feature vector.

[0090] The specific steps of the optimization iteration of round a include:

[0091] Obtain a reference embedding feature vector for the a-th optimization iteration, where the reference embedding feature vector for the first optimization iteration is generated by fusing the second embedding feature vector, the third embedding feature vector, the derived field tag data, and the random embedding feature vector corresponding to the randomly generated distributed multivariate data. When 1 < a ≤ Y, the reference embedding feature vector for the a-th optimization iteration is the optimized embedding feature vector for the a-1-th optimization iteration.

[0092] In the bth second feature focusing unit, an interactive calculation is performed based on the first feature focusing result generated by the bth first feature focusing unit and the second feature focusing result of the bth second feature focusing unit for the ath round of optimization iteration to generate an interactive embedded feature vector generated by the bth second feature focusing unit for the ath round of optimization iteration. 1≤b≤X, b is a positive integer. The second feature focusing result of the first second feature focusing unit for the ath round of optimization iteration is determined based on the reference embedded feature vector of the ath round of optimization iteration.

[0093] The interactive embedding feature vector generated by the b-th second feature focusing unit is optimized and iterated to generate a b-th temporary optimized iterative embedding feature vector for the a-th round of optimization iteration.

[0094] If b is less than X, determine the second feature focusing result of the b+1th second feature focusing unit for the ath round of optimization iteration based on the bth temporary optimization iteration embedded feature vector for the ath round of optimization iteration, count and accumulate b, and return to execute the step of interactively calculating in the bth second feature focusing unit based on the first feature focusing result generated by the bth first feature focusing unit and the second feature focusing result of the bth second feature focusing unit for the ath round of optimization iteration, to generate the interactive embedded feature vector generated by the bth second feature focusing unit for the ath round of optimization iteration.

[0095] If b=X, the bth temporary optimized iteration embedding feature vector for the ath round of optimization iteration is used as the optimized embedding feature vector for the ath round of optimization iteration.

[0096] In this embodiment, it is assumed that the feature conversion model in the server includes three first feature focusing units (X=3), and the first embedded feature vector includes the quantitative representation of the target derived template knowledge data mentioned above, such as the quantitative information of the relationship between age and payment probability, payment time and payment amount in the mobile payment scenario.

[0097] The first feature focusing unit 1 focuses on feature transformation related to the relationship between age and payment probability. It contains a predefined age-based payment probability mapping table, which is derived from a large amount of historical mobile payment data and user age data. For example, the probability of a high-consumption payment (amount greater than 1,000 yuan) for the 20-30 age group is 0.2, the probability of a medium-consumption payment (1,000-1,000 yuan) is 0.5, and the probability of a low-consumption payment (less than 100 yuan) is 0.3.

[0098] When processing the first embedded feature vector, it identifies the portion of the vector related to age and payment probability. Assume that the age portion of the first embedded feature vector is coded numerically to represent the age range (e.g., 25 is coded as 25), and the payment probability portion is represented as a decimal. Unit 1 converts this data according to the mapping table, for example, readjusting the high, medium, and low consumption payment probabilities corresponding to age 25 (possibly with some fine-tuning based on business logic, such as increasing the high consumption payment probability by 0.05). This results in a portion of the first feature focus result, including the adjusted age-payment probability relationship index (pointing to the new mapping relationship) and the new feature value.

[0099] The first feature focusing unit 2 focuses on the relationship between payment time and payment amount. It has a payment time-payment amount relationship model derived from time series analysis. For example, during the afternoon hours (1:15 PM) on weekdays, the proportion of small payments (less than 100 RMB) is higher, while in the evening hours (7:21 PM), the proportion of large payments (greater than 1,000 RMB) is relatively higher.

[0100] For the portion of the first embedded feature vector related to payment time and payment amount, this unit transforms it according to the relational model. If the payment time in the vector is coded as 14 (corresponding to 14 points) and the payment amount is coded as a medium amount (assuming 500 yuan), unit 2 may adjust the expected payment amount based on the model (for example, adjusting it to favor small payments based on time trends, adjusting the amount to 80 yuan) and simultaneously update the relational index, forming the first feature focus result for this portion.

[0101] The first feature focusing unit 3 is responsible for processing the relationship between payment methods and age. For example, young people (aged 18-30) are more inclined to use electronic payments, while the elderly (over 60) may prefer cash payments or specific elderly-friendly payment methods.

[0102] The first embedded feature vector is then searched for the portion related to payment method (e.g., electronic payment is coded as 1) and age (e.g., 22 years old is coded as 22) and transformed according to the predefined relationship schema. If the current relationship indicates that the probability of a 22-year-old using electronic payment is 0.8, unit 3 might adjust the probability to 0.9 based on new market trends (e.g., young people are more receptive to new electronic payment methods), update the relationship index, and generate the first feature-focused result for this portion.

[0103] The operation results of the three units are integrated to obtain feature conversion data, which includes the first feature focusing results loaded by each first feature focusing unit, and covers the information after conversion and adjustment of different relationships in the target derived template knowledge data.

[0104] Next, assuming that the iterative optimization model also includes three second feature focusing units (X=3), and is set to perform three rounds of optimization iterations (Y=3), the randomly generated distributed multivariate data are initially some random numerical vectors, for example, the dimension of each vector is 10, and the value is randomly generated between 0 and 1.

[0105] During the first round (a=1) of optimization iteration, the server fuses the second embedded feature vector (containing the features of hidden distributed multivariate data, such as hidden transaction amounts, personnel database data, etc.), the third embedded feature vector (dynamically updated node data features, such as dynamic node information of transaction time and personnel flow), derived field tag data (such as previously marked age groups and transaction amount levels), and the random embedded feature vector corresponding to the randomly generated distributed multivariate data.

[0106] For example, for each element in the second embedded feature vector, a weighted sum is performed with the elements at the corresponding positions in the third embedded feature vector, the derived field marker data, and the random embedded feature vector. Assuming that the second embedded feature vector element is v1, the corresponding element in the third embedded feature vector is v2, the derived field marker data element is v3, and the random embedded feature vector element is v4, the fused element is w = 0.3*v1+0.2*v2+0.1*v3+0.4*v4. This is how the reference embedded feature vector for the first round of optimization iteration is generated.

[0107] During the operation in the first second feature focusing unit (b=1), the second feature focusing result of the first second feature focusing unit for the first optimization iteration is determined based on the reference embedded feature vector of the first optimization iteration. Assume that this unit mainly focuses on the relationship between age and payment method.

[0108] It extracts the parts related to age and payment method from the reference embedded feature vector, for example, age coded as 25 and payment method coded as electronic payment. Based on a predefined age-payment method relationship model (similar to the relationship model in the previous feature conversion unit, but possibly different in details), it obtains the initial second feature focus results, including an index pointing to a specific part of the age-payment method relationship (such as the specific region in the model where a 25-year-old uses electronic payment) and the corresponding feature value (such as a quantitative feature of using electronic payment in this case, assuming it is 0.7).

[0109] At the same time, interactive calculations are performed based on the first feature focusing result generated by the first first feature focusing unit (the adjusted result of the relationship between age and payment method obtained in the previous feature conversion stage, assuming that the adjusted eigenvalue is 0.8) and the second feature focusing result of the first round of optimization iteration by the first second feature focusing unit.

[0110] For the interactive operation of indexes, a certain logical operation is performed on the age-payment method relationship index in the first feature focusing result and the index in the second feature focusing result, such as a bitwise AND operation (if the index is represented in binary), to obtain the target focusing index.

[0111] For the interactive operation of eigenvalues, the eigenvalue 0.8 in the first feature focusing result is multiplied by the eigenvalue 0.7 in the second feature focusing result to obtain the target focused eigenvalue 0.56.

[0112] Based on the pre-set target interaction index (assuming the weight index of the age payment method in different scenarios determined by the business logic, which is 0.6 here) and the target focus index, the interaction factor is determined (such as 0.8 obtained by looking up the predefined interaction factor table).

[0113] Finally, the target focused eigenvalue is interactively calculated (eg, multiplied) based on the interaction factor, 0.56*0.8=0.448, to generate the interactive embedding feature vector generated by the first second feature focusing unit for the first round of optimization iteration.

[0114] This interactive embedding feature vector is optimized iteratively. For example, the server uses a pre-set optimization rule to adjust each element in the vector to better align with business logic. If the element value is less than 0.5, it is increased by 0.1 (within a certain range), resulting in the first interim optimized iterative embedding feature vector for the first round of optimization iteration.

[0115] During the operation of the second second feature focusing unit (b=2), the second feature focusing result of the second second feature focusing unit for the first optimization iteration is determined based on the first temporary optimized iterative embedded feature vector for the first optimization iteration. For example, this unit focuses on the relationship between payment time and transaction amount, extracts information related to payment time and transaction amount from the temporary optimized iterative embedded feature vector, and obtains the initial second feature focusing result based on a predefined payment time and transaction amount relationship model.

[0116] Then, based on the first feature focusing result generated by the second first feature focusing unit (the result after adjustment of the relationship between payment time and transaction amount obtained in the previous feature conversion stage) and the second feature focusing result of the second second feature focusing unit for the first round of optimization iteration, interactive calculations are performed. The steps are similar to the operations in the first second feature focusing unit. An interactive embedded feature vector is generated and optimized iteratively to obtain the second temporary optimized iterative embedded feature vector.

[0117] During the operation in the third second feature focusing unit (b=3), in accordance with the previous method, the second feature focusing result of the third second feature focusing unit for the first round of optimization iteration is determined based on the third temporary optimization iterative embedded feature vector, and interactive calculations are performed based on the first feature focusing result generated by the third first feature focusing unit and the second feature focusing result of the third second feature focusing unit for the first round of optimization iteration to generate an interactive embedded feature vector and perform optimization iteration.

[0118] Since b=X=3, the third temporary optimized iteration embedding feature vector for the first round of optimization iteration is used as the optimized embedding feature vector for the first round of optimization iteration.

[0119] During the second round (a=2) of optimization iteration, the reference embedding feature vector is the optimized embedding feature vector of the first round of optimization iteration.

[0120] Repeat the operation process of the first round in each second feature focusing unit, starting from the first second feature focusing unit (b=1), and perform interactive calculation, generate interactive embedding feature vectors, and optimize iterations in sequence, until the third second feature focusing unit (b=3) uses the third temporary optimized iterative embedding feature vector for the second round of optimization iteration as the optimized embedded feature vector for the second round of optimization iteration.

[0121] During the third round (a=3) of optimization iteration, the reference embedding feature vector is the optimized embedding feature vector of the second round of optimization iteration.

[0122] Repeat the previous process again. After the third second feature focusing unit operation in the third round is completed, the third temporary optimized iterative embedding feature vector of the third optimization iteration is used as the derived embedding feature vector. This derived embedding feature vector has undergone three rounds of iterative optimization and integrates information from the target derived template knowledge data, hidden distributed multivariate data, dynamically updated node data, and derived field tag data. It is continuously adjusted through iterative optimization to meet the requirements of business logic and data characteristics.

[0123] In a possible implementation, the first feature focusing result includes a first focusing index and a first focusing feature, and the second feature focusing result includes a second focusing index, a second focusing feature, and a target interaction index.

[0124] In the b-th second feature focusing unit, interactive calculation is performed based on the first feature focusing result generated by the b-th first feature focusing unit and the second feature focusing result of the b-th second feature focusing unit for the a-th round of optimization iteration, to generate the interactive embedding feature vector generated by the b-th second feature focusing unit for the a-th round of optimization iteration, including:

[0125] In the bth second feature focusing unit, the first focusing index and the second focusing index of the bth second feature focusing unit for the ath round of optimization iteration are interactively operated to generate the target focusing index of the bth second feature focusing unit for the ath round of optimization iteration.

[0126] Perform interactive operations on the first focusing feature and the second focusing feature of the bth second feature focusing unit for the ath round of optimization iteration to generate a target focusing feature of the bth second feature focusing unit for the ath round of optimization iteration.

[0127] Based on the target interaction index of the bth second feature focusing unit for the ath round of optimization iteration and the target focusing index, the interaction factor of the bth second feature focusing unit for the ath round of optimization iteration is determined.

[0128] Based on the interaction factor of the b-th second feature focusing unit for the a-th round of optimization iteration, the corresponding target focusing feature is interactively calculated to generate an interactive embedding feature vector generated by the b-th second feature focusing unit for the a-th round of optimization iteration.

[0129] In this embodiment, it is assumed that the second round (a=2) of optimization iteration is in progress and the first second feature focusing unit (b=1) is currently in progress.

[0130] When the first embedded feature vector was previously transformed using the feature conversion model, the first first feature focusing unit generated a first feature focusing result. Assume that the first focusing index in this first feature focusing result is a code representing the relationship between age and payment method. For example, this code is "0101", which may correspond to the relationship mapping between the age between 20 and 30 years old and a certain electronic payment method (such as Alipay payment) in a specific business scenario. The first focused feature is a quantitative value related to the relationship, such as 0.8, which represents the strength of a certain business feature under this age-payment method relationship (which can be understood as the tendency of this age group to use this payment method).

[0131] The second feature focusing result of the first second feature focusing unit for the second round of optimization iteration includes a second focusing index, a second focusing feature, and a target interaction index. Assuming that the second focusing index is determined based on the reference embedded feature vector of the current round, this index is "0110", which may be related to a specific representation of the relationship between age and payment method obtained after the fusion of hidden distributed multivariate data, dynamically updated node data, and derived field tag data. The second focusing feature is 0.6, which is another quantitative value related to the relationship extracted from the current round of data. The target interaction index is pre-set according to the business logic, for example, 0.5, which represents a weight adjustment factor when different data features interact in this specific business scenario.

[0132] Specifically, in the first second feature focusing unit, the first focusing index "0101" and the second focusing index "0110" of the first second feature focusing unit for the second round of optimization iteration are interactively operated. A bitwise XOR operation is used here (this is a possible interactive operation method, and the specific operation method depends on the design of the business logic and data representation). The result of the bitwise XOR operation is "0011", which generates the target focusing index of the first second feature focusing unit for the second round of optimization iteration. This target focusing index will be used for the subsequent determination of the interaction factor and the interactive calculation of the target focusing feature.

[0133] An interactive operation is performed on the first focused feature 0.8 and the second focused feature 0.6 of the first second feature focusing unit for the second optimization iteration. A multiplication operation is used here (again, the operation method depends on the business logic), resulting in a target focused feature of 0.8 * 0.6 = 0.48. This target focused feature represents a new quantitative relationship after the interaction between the first and second focused features.

[0134] Based on the target interaction index of 0.5 and the target focus index "0011" for the first second feature focusing unit for the second optimization iteration, the interaction factor of the first second feature focusing unit for the second optimization iteration is determined. Assume that the server has a predefined interaction factor lookup table that determines the interaction factor based on different combinations of the target interaction index and the target focus index. For the case where the target interaction index is 0.5 and the target focus index is "0011", the interaction factor is found in the lookup table to be 0.7. This interaction factor will be used for the final interaction calculation of the target focused feature.

[0135] Based on the interaction factor 0.7 of the first second feature focusing unit for the second optimization iteration, the corresponding target focused feature 0.48 is interactively calculated. Here, a multiplication operation is used to obtain 0.48 * 0.7 = 0.336. This result will be used as part of the interactive embedding feature vector generated by the first second feature focusing unit for the second optimization iteration (if the vector is multidimensional, this is only the calculation result of one dimension; other dimensions may be calculated according to similar logic or combined with other related calculation logic).

[0136] Assume that the optimization iteration continues for the second round (a=2) and reaches the second second feature focusing unit (b=2).

[0137] The first focus index in the first feature focus result generated by the second first feature focus unit is assumed to be "1010." This likely corresponds to a relationship encoding between payment time (e.g., 14:16 PM) and transaction amount (e.g., a medium-amount transaction range) in the target derived template knowledge data. The first focus feature is 0.7, indicating the strength of a business characteristic associated with this relationship (e.g., the expected probability of a medium-amount transaction occurring within this time period).

[0138] The second focus index in the second feature focus result of the second round of optimization iteration for the second second feature focusing unit is assumed to be "1001." This is a specific representation related to the payment time and transaction amount relationship, determined based on the reference embedded feature vector of the current round. The second focus feature is 0.5, and the target interaction index is 0.6 (pre-set according to the business logic).

[0139] Specifically, the first focusing index "1010" and the second focusing index "1001" of the second second feature focusing unit for the second round of optimization iteration are interactively operated. Here, a bitwise AND operation is used to obtain the result "1000", which is the target focusing index of the second second feature focusing unit for the second round of optimization iteration.

[0140] The first focusing feature 0.7 is multiplied by the second focusing feature 0.5 of the second second feature focusing unit for the second round of optimization iteration, and the target focusing feature is obtained as 0.7*0.5=0.35.

[0141] According to the target interaction index 0.6 and the target focus index "1000", the interaction factor is found to be 0.8 from the predefined interaction factor lookup table.

[0142] Multiplying the target focused feature 0.35 by the interaction factor 0.8 yields 0.35*0.8=0.28. This result will be used as part of the interaction embedding feature vector generated by the second feature focusing unit for the second round of optimization iterations, and together with the calculation results of other dimensions (if any), it constitutes the complete interaction embedding feature vector.

[0143] Following this process, similar interactive calculations are performed in each second feature focusing unit (b) of each round (a), thereby gradually generating interactive embedding feature vectors for each round and each unit. After subsequent optimization iterative operations, these vectors ultimately generate derived embedding feature vectors, providing the basis for subsequent operations such as embedding restoration and generation of derived distributed multivariate data.

[0144] In a possible implementation, the iterative transfer condition further includes a business logic description document for the target distributed multivariate data.

[0145] In one possible implementation, the method further includes:

[0146] Step A110: Acquire template-derived template knowledge data, template-hidden distributed multivariate data, template-derived field tag data, and template-dynamically updated node data. The template-derived template knowledge data is determined based on the template-to-be-derived field data in the first template-distributed multivariate data. The template-hidden distributed multivariate data is generated by hiding the template-to-be-derived field data in the first template-distributed multivariate data. The template-derived field tag data is the tagging result of the template-to-be-derived field data. The template-dynamically-updated node data is obtained by extracting the dynamic update nodes of the template-dynamic field data in the first template-distributed multivariate data. The template-to-be-derived field data includes the dynamic field data to be derived in the template-dynamic field data.

[0147] Step A120, respectively embed the template-derived template knowledge data, the template-hidden distributed multivariate data, and the template-dynamically-updated node data to generate a first template-embedded feature vector of the template-derived template knowledge data, a second template-embedded feature vector of the template-hidden distributed multivariate data, and a third template-embedded feature vector of the template-dynamically-updated node data.

[0148] Step A130: Perform feature conversion on the first template embedded feature vector using a feature conversion model to generate template feature conversion data.

[0149] Step A140, using the second template embedded feature vector, the third template embedded feature vector and the template derived field mark data as iterative transfer conditions, using the iterative optimization model, based on the template feature conversion data, iteratively optimizes the template randomly generated distributed multivariate data to generate a template derived embedded feature vector.

[0150] Step A150 , performing embedding restoration on the template-derived embedded feature vector to generate template-derived distributed multivariate data.

[0151] Step A160: Training the feature conversion model and the iterative optimization model according to the loss function value between the template-derived distributed multivariate data and the first template distributed multivariate data until the first model convergence requirement is met.

[0152] In this embodiment, for the target distributed multivariate data (including mobile payment data, personnel database data, scene monitoring data, etc.), the server has a corresponding business logic description document. This document describes in detail the logical relationship and business rules between the data.

[0153] For example, in terms of mobile payment data, the document states that there is a correlation between transaction amounts and user age, payment time, and payment location. Specifically, users aged 20-30 may make relatively high amounts of mobile payments at night (18:00-22:00), and payments in commercial centers are generally higher than in other areas. Regarding personnel database data, it is associated with scene monitoring data. For example, when the flow of people monitored in certain specific scenes (such as large event venues) is large, certain attributes of the relevant personnel in the personnel database (such as activity tags) may be updated.

[0154] Throughout the iterative optimization process (such as the previously mentioned iterative optimization of randomly generated distributed multivariate data based on the first and second embedded eigenvectors), the server references this business logic description document. For example, when determining interaction factors or adjusting eigenvalues, according to the rules in the document, if the current iteration involves the relationship between age and payment amount, and the age is 20-30 and the payment time is in the evening (18:22), the relevant eigenvalues will be adjusted or the corresponding weight bias will be given when determining the interaction factor, based on the logic mentioned in the document that the payment amount is relatively high in this case.

[0155] Assume that the first template distributed multivariate data includes simulated mobile payment, personnel database, and scenario monitoring data. The template's to-be-derived field data, such as the transaction amount in the mobile payment data and age in the personnel database data, are identified as template's to-be-derived field data (because they may have potential derivative relationships with other data features).

[0156] The server determines the template-derived template knowledge data based on the template's to-be-derived field data. For example, it analyzes the relationship between transaction amount, payment time, payment location, and the person's age. If it finds that users aged 30-40 in a business district typically pay between 200 and 500 yuan at noon on weekdays (12:00-14:00), this relationship pattern will be constructed as part of the template-derived template knowledge data.

[0157] The template's derived field data (such as transaction amount and age) in the first template's distributed multivariate data is hidden. Taking the transaction amount as an example, the server can use a simple offset encryption method to add a random number (within a certain range, such as 0100) to each transaction amount to hide it. For the age field, the actual age can be replaced with a fixed code, such as replacing all ages with a random number between 0100, to generate template-hidden distributed multivariate data.

[0158] Next, the data for the template's derived fields is labeled. For transaction amounts, if the amount is less than 100 yuan, it is marked as "low" and represented by the number 1; 100-500 yuan is marked as "medium" and represented by the number 2; and over 500 yuan is marked as "high" and represented by the number 3. For age, those under 18 are marked as "minor" and represented by 1; those aged 18-60 are marked as "adult" and represented by 0; and those over 60 are marked as "elderly" and represented by 1. This yields the labeled data for the template's derived fields.

[0159] Next, in the first template's distributed multivariate data, transaction time, for example, is a dynamic field in mobile payment data. The server counts transactions in half-hour windows. When the number of transactions exceeds 100, the half-hour window is marked as a dynamic update node. For the flow of people in the scene monitoring data, when the change in flow exceeds 50 people between two consecutive 10-minute time periods, this time point is marked as a dynamic update node, thereby generating the template's dynamic update node data.

[0160] Next, the server quantifies the relational patterns in the template-derived template knowledge data. For example, for the aforementioned relationship between users aged 30-40 and those who pay in the business district at noon on weekdays (12:14) and the amount of money paid, the server quantifies the relational patterns in the template-derived template knowledge data. The server quantifies the relationship between the age range, payment time interval, payment location category, and amount range, respectively, by assigning numerical codes to the age range, e.g., 30-40 is coded as 3040, 12:14 is coded as 1214, the business district is coded as 1 (assuming different codes exist for different areas), and 200-500 yuan is coded as 200-500.

[0161] A neural network-based embedding model (such as a multi-layer perceptron with specific structure and parameters) is used to convert these quantized relational patterns into the first template embedding feature vector. This embedding model maps the input quantized relational pattern into a low-dimensional vector space by learning the intrinsic relationship in the data, thereby obtaining the first template embedding feature vector.

[0162] For template-hidden distributed multivariate data, let's take mobile payment data after hiding as an example. In addition to the hidden transaction amount (e.g., the value encrypted using offset), other fields include payment method (assuming electronic payment is coded as 1 and cash payment is coded as 0) and transaction location (coded according to geographic location).

[0163] This data is embedded using an autoencoder (a neural network structure that automatically learns to represent data features). The autoencoder's input layer receives the encoding of each field of the template hidden distributed multivariate data. Through the encoding and decoding process of the hidden layer, the output layer outputs a second template embedded feature vector, which captures the characteristic information of the hidden data.

[0164] For template dynamic update node data, such as the dynamic update nodes extracted from transaction time and personnel flow, the dynamic update nodes of transaction time are coded in chronological order, for example, the first dynamic update node is coded as 1, the second is coded as 2, and so on. For the dynamic update nodes of personnel flow, they are coded according to different flow threshold ranges, for example, the first node with a flow change between 50 and 100 people is coded as 11, the second is coded as 12, and so on.

[0165] An embedding technology specifically for time series and discrete event coding (such as an embedding algorithm based on time series analysis) is used to convert these template dynamic update node data into a third template embedding feature vector, which can represent the characteristics of the dynamic update node in time and event logic.

[0166] Assume that the feature conversion model contains three first feature focusing units.

[0167] The first unit focuses on feature transformation related to the relationship between age and payment amount. It contains an age-payment-amount mapping table derived from statistical analysis. For example, it displays the probability distribution of different payment amount ranges for the 30-40 age group. When processing the portion of the first template's embedded feature vector related to age and payment amount, the features are adjusted according to the mapping table, resulting in the first feature focus result for this portion, including the adjusted age-payment-amount relationship index and the new feature value.

[0168] The second unit focuses on the relationship between payment time and payment location, and transforms the relevant part of the first template embedded feature vector according to the pre-built payment time and payment location relationship model to obtain the corresponding first feature focusing result.

[0169] The third unit processes the relationship between payment method and age, and obtains the first feature focus result of this part in a similar way, and finally combines them to obtain the template feature conversion data.

[0170] In addition, the iterative optimization model also includes three second feature focusing units.

[0171] The second template embedded feature vector, the third template embedded feature vector and the template derived field mark data are used as iterative transfer conditions, and the template randomly generates distributed multivariate data (initial random numerical vectors, for example, each vector dimension is 8, and the value is randomly generated between 01) according to the template feature conversion data for iterative optimization processing.

[0172] For example, in the first round of optimization iteration, a reference embedding feature vector is first generated, and the second template embedding feature vector, the third template embedding feature vector, the template-derived field label data, and the random embedding feature vector corresponding to the template randomly generated distributed multivariate data are fused (such as weighted summation).

[0173] In the first second feature focusing unit, interactive calculations are performed based on the first feature focusing result generated by the first first feature focusing unit and the second feature focusing result of the first second feature focusing unit for the first round of optimization iteration (such as the index interaction operation, feature interaction operation, determination of interaction factors and final feature interaction calculation described previously) to generate an interactive embedding feature vector, and then optimization iteration is performed to obtain the first temporary optimized iterative embedding feature vector.

[0174] A similar process is followed in the second and third second feature focusing units. After multiple rounds (assuming 3 rounds) of iterative optimization, the optimized embedding feature vector generated by the last round of optimization iteration is used as the template to derive the embedding feature vector.

[0175] Assume that in the previous embedding representation process, the template-derived template knowledge data was converted using a neural network embedding model, the template hidden distributed multivariate data was embedded using an autoencoder, and the template dynamically updated node data was embedded using a time series analysis embedding algorithm.

[0176] When restoring the embedding of the template-derived embedding feature vector, the portion converted by the neural network embedding model is restored using inverse neural network operations. For example, if the neural network embedding model performs matrix multiplication and nonlinear activation function operations during forward propagation, the restoration is performed using the inverse matrix multiplication and inverse activation function operations (if any).

[0177] For the portion processed by the autoencoder, the decoder of the autoencoder performs a restoration, restoring the encoded features to their original representation in the data space. For the portion processed by the time series analysis embedding algorithm, the inverse process of the algorithm restores the encoding to the original dynamically updated node data representation, thereby generating template-derived distributed multivariate data.

[0178] Thus, a loss function value is calculated between the template-derived distributed multivariate data and the first template distributed multivariate data. For example, the mean square error (MSE) is used as the loss function to calculate the sum of squared errors between the predicted values (such as the predicted transaction amount, age, etc.) in the template-derived distributed multivariate data and the true values in the first template distributed multivariate data.

[0179] Based on this loss function value, the feature conversion model and the iterative optimization model are trained. If the loss function value is large, it means that the prediction results of the model are significantly different from the actual data, and the model parameters need to be adjusted. For example, the internal parameters of each first feature focusing unit in the feature conversion model (such as the probability value in the age payment amount mapping table, the coefficient in the relationship model, etc.) are adjusted, and the parameters of each second feature focusing unit in the iterative optimization model (such as the weight in the interactive calculation, the step size of the optimization iteration, etc.) are adjusted.

[0180] This process is repeated continuously, calculating the loss function value and adjusting the model parameters until the first model convergence requirement is met. For example, when the loss function value is less than a pre-set threshold (such as 0.01), the model is considered to have converged and training is stopped.

[0181] In a possible implementation, step A110 includes:

[0182] Step A111 : extracting derivable fields from the first template distributed multivariate data to determine template field data to be derived from the first template distributed multivariate data.

[0183] Step A112: Marking the template's to-be-derived field data to generate the template's derived field marking data.

[0184] Step A113 : performing a hiding operation on the template-to-be-derived field data in the first template distributed multivariate data according to the template-derived field tag data, to generate template-hidden distributed multivariate data.

[0185] Step A114: Process the first template distributed multivariate data according to the template-derived field tag data to generate the template-derived template knowledge data.

[0186] Step A115 , extracting dynamic update nodes from the template dynamic field data in the first template distributed multivariate data to generate the template dynamic update node data.

[0187] In a possible implementation, step A113 includes:

[0188] Step A1131 : Regularly convert the first template distributed multivariate data into a first eigenvalue interval to generate first temporary template distributed multivariate data.

[0189] Step A1132: Regularly convert the template-derived field tag data into a second eigenvalue interval to generate first template-derived field tag data, where the second eigenvalue interval is within the first eigenvalue interval.

[0190] Step A1133: fuse the first temporary template distributed multivariate data and the first template derived field tag data element by element to generate second temporary distributed multivariate data.

[0191] Step A1134: perform inverse regularization conversion on the second temporary distributed multivariate data to generate the template hidden distributed multivariate data.

[0192] In one possible implementation, step A114 includes:

[0193] Step A1141 : Regularly convert the template-derived field tag data into a second characteristic value interval to generate second template-derived field tag data.

[0194] Step A1142: fuse the second template-derived field tag data and the first template distributed multivariate data element by element to generate the template-derived template knowledge data.

[0195] In a possible implementation, step A111 includes:

[0196] Step A1111: perform structural feature conversion on the first template distributed multivariate data using a first feature mapping structure to generate a template structure embedding feature vector.

[0197] Step A1112: Perform logical feature conversion on the first template distributed multivariate data using a second feature mapping structure to generate a template logical embedding feature vector.

[0198] Step A1113: Perform dimensionality reduction processing on the template structure embedding feature vector to generate a reduced-dimensionality structure embedding feature vector.

[0199] Step A1114: perform dimensionality-increasing processing on the template logic embedding feature vector to generate a dimensionality-increasing logic embedding feature vector.

[0200] Step A1115: Combine the template logic embedding feature vector and the dimension reduction structure embedding feature vector to generate a first template combined embedding feature vector.

[0201] Step A1116: Combine the template structure embedding feature vector with the dimension-raising logic embedding feature vector to generate a second template combination embedding feature vector.

[0202] Step A1117: Merge the first template combination embedding feature vector and the second template combination embedding feature vector to generate a template embedding feature vector.

[0203] Step A1118: Determine the template-to-be-derived field data in the first template distributed multivariate data based on the template-embedded feature vector.

[0204] In this embodiment, it is further assumed that the first template distributed multivariate data includes mobile payment data, personnel database data, and scene monitoring data.

[0205] For mobile payment data, the first feature mapping structure can be a neural network-based feature extractor. For example, given input fields such as transaction amount, payment time, and payment method from mobile payment data, this neural network maps the values of each field into a high-dimensional space. For example, a transaction amount of 100 yuan might be mapped to a 10-dimensional vector [0.1, 0.2, -0.1, 0.3, ...], while payment time might be mapped to another different 10-dimensional vector. Similar mapping operations are performed for fields such as name, age, and gender in the personnel database data. Finally, these mapped vectors are combined to form a template structure embedding feature vector. Assuming the dimension of the mapped vector for mobile payment data is 100, the dimension of the mapped vector for personnel database data is 80, and the dimension of the mapped vector for scene monitoring data is 60, the combined template structure embedding feature vector has a dimension of 240.

[0206] The second feature mapping structure is a logical rule-based mapper. For mobile payment data, it might map based on the logical relationship between payment amount and payment time. For example, if the payment amount is greater than 500 yuan and the payment time is in the evening (18:22), a logical value of 1 is assigned; otherwise, it is 0. For personnel database data, mapping is performed based on the logical relationship between age and gender. For example, if the age is between 18:30 and the person is male, a logical value of 1 is assigned. These logical values are combined to form a template logical embedding feature vector. Assume that the dimension of this vector is 50, which is determined by the number of logical relationships.

[0207] Next, the template structural embedding feature vector undergoes dimensionality reduction. For example, using the principal component analysis (PCA) algorithm, the 240-dimensional template structural embedding feature vector is reduced to 120 dimensions, resulting in a reduced-dimensional structural embedding feature vector. For the template logical embedding feature vector, a neural network-based dimensionality increase technique (such as a specific multilayer perceptron structure) is used to increase the 50-dimensional vector to 100 dimensions, generating a increased-dimensional logical embedding feature vector.

[0208] The template logic embedding feature vector (100 dimensions) after dimensionality increase is combined with the template structure embedding feature vector (120 dimensions) after dimensionality reduction to obtain the first template combination embedding feature vector with a dimension of 220. At the same time, the original template structure embedding feature vector (240 dimensions) is combined with the template logic embedding feature vector (100 dimensions) after dimensionality increase to obtain the second template combination embedding feature vector with a dimension of 340.

[0209] The two combined embedding feature vectors are then merged. For example, the first template combined embedding feature vector and the second template combined embedding feature vector are sequentially concatenated to obtain a template embedding feature vector with a dimension of 560.

[0210] The server determines the template's field data to be derived based on the numerical features in the template's embedded feature vector and predefined rules. For example, if certain dimensions in the template's embedded feature vector have high weights or specific numerical patterns related to payment amount and age, the payment amount in the mobile payment data and age in the person database data are determined as the template's field data to be derived.

[0211] For the previously determined template field data to be derived, such as the payment amount in mobile payment data and the age in personnel database data.

[0212] For the payment amount, the server marks it according to the following rules: if the payment amount is less than 100 yuan, it is marked as "small amount" and represented by the number 1; if the amount is between 100 yuan and 500 yuan, it is marked as "medium amount" and represented by the number 2; if the amount is greater than 500 yuan, it is marked as "large amount" and represented by the number 3.

[0213] Age is marked according to the following rules: if the age is less than 18 years old, it is marked as "minor" and represented by 1; if the age is between 18 and 60 years old, it is marked as "adult" and represented by 0; if the age is greater than 60 years old, it is marked as "elderly" and represented by 1.

[0214] Through such tagging processing, template-derived field tag data is generated.

[0215] For the mobile payment data in the first template distributed multivariate data, it is assumed that the transaction amount ranges from 0 to 1,000 yuan, the payment time is from 0 to 24 hours (in hours), and the payment method is represented by 0 (cash) and 1 (electronic payment). The server converts the transaction amount to the first eigenvalue interval of 0 to 1 according to the linear transformation rule. For example, a transaction amount of 100 yuan is converted to 0.1 (100 / 1000). The payment time is also converted similarly, such as 12 o'clock is converted to 0.5 (12 / 24). For the age in the personnel database data, assuming that the age range is 0 to 100 years old, the age value is converted to the interval of 0 to 1. For example, 30 years old is converted to 0.3 (30 / 100), thus generating the first temporary template distributed multivariate data.

[0216] For the previously generated template-derived field tag data, such as the payment amount tags (1, 2, 3) and age tags (1, 0, 1), these tags are converted to a smaller second feature value range, such as 0 to 0.5. If the payment amount tag is 1 (small amount), it is converted to 0.1; if it is 2 (medium amount), it is converted to 0.3; if it is 3 (large amount), it is converted to 0.5. For the age tag, 1 (minor) is converted to 0.1, 0 (adult) is converted to 0.3, and 1 (elderly) is converted to 0.5, thus obtaining the first template-derived field tag data.

[0217] Each element in the first temporary template distributed multivariate data (e.g., 0.1 after the transaction amount is converted, 0.5 after the payment time is converted, etc.) is fused element by element with the corresponding element in the first template derived field tag data (e.g., 0.1 after the payment amount tag is converted). For example, a simple addition operation can be used to add 0.1 after the transaction amount is converted and 0.1 after the payment amount tag is converted to obtain 0.2. The same operation is performed on the age-related elements in the personnel database data to generate the second temporary distributed multivariate data.

[0218] Perform an inverse regularization transformation on each element in the second temporary distributed multivariate data. For example, the value 0.2, previously obtained by fusing the transaction amount and payment amount tags, is restored using the inverse of the previous regularization transformation. Since the transaction amount was previously converted from 0 to 1000 yuan to 0 to 1, the inverse transformation now reverses the process. The corresponding transaction amount for 0.2 is 200 yuan (0.2 * 1000). Similar inverse transformations are performed on elements such as payment time and age, ultimately resulting in the template-hidden distributed multivariate data.

[0219] Similar to the previous operation, the template-derived field tag data (such as the payment amount tags 1, 2, 3 and the age tags 1, 0, 1) are converted to the second eigenvalue range of 0 to 0.5 to obtain the second template-derived field tag data (such as the payment amount tags are converted to 0.1, 0.3, 0.5, and the age tags are converted to 0.1, 0.3, 0.5).

[0220] For each element in the distributed multivariate data of the first template, such as the transaction amount, payment time, and payment method in the mobile payment data, and the name, age, and gender in the personnel database data, element-by-element fusion is performed with the corresponding element in the derived field tag data of the second template. For example, for a transaction amount of 100 yuan, the corresponding payment amount tag is converted to 0.1, and the result after fusion can be a new value (such as using a simple weighted sum, 100*0.9+0.1*100=91). For the age of 30 in the personnel database data, the corresponding age tag is converted to 0.3, and the result after fusion may be 30*0.7+0.3*30=30 (this is just an example of a fusion method). Through such an element-by-element fusion operation, template-derived template knowledge data is generated.

[0221] It is assumed that the template dynamic field data in the first template distributed multivariate data includes transaction time in mobile payment data and personnel flow in scenario monitoring data.

[0222] For transaction times in mobile payment data, the server counts transactions in 15-minute windows. If the number of transactions in a 15-minute window exceeds 100, the window is marked as a dynamic update node. For example, if there are 120 transactions in the window from 10:00 AM to 10:15 AM, the window is marked as a dynamic update node.

[0223] For the flow of people in scene monitoring data, we count changes in flow every 10 minutes. If the change in flow exceeds 50 people between two consecutive 10-minute time periods, this time point is marked as a dynamic update node. For example, if the flow from 9:00 to 9:10 is 100 people, and from 9:10 to 9:20 is 160 people, the change in flow is 60 people, exceeding 50 people. Therefore, 9:10 is marked as a dynamic update node. This operation generates template dynamic update node data.

[0224] In a possible implementation, the iterative transfer condition further includes a business logic description document generated by a business logic construction model based on the target distributed multivariate data.

[0225] The method further comprises:

[0226] Step B110: Utilize a data structure conversion unit to perform feature conversion on the second template distributed multivariate data to generate a template data structure vector.

[0227] Step B120: Utilize a logic conversion unit to perform feature conversion on the template logic description document corresponding to the second template distributed multivariate data to generate a template logic conversion vector.

[0228] Step B130: Utilize a determination unit to perform logic relevance determination on the template data structure vector and the template logic description document, and generate a template logic relevance determination result.

[0229] Step B140: Using the business logic to build a model, based on the template data structure vector and the template logic description document, generates a prediction logic description document.

[0230] Step B150, determining the business logic construction error based on the loss function value between the prediction logic description document and the template logic description document, the loss function value between the template data structure vector and the template logic conversion vector, and the template logic correlation judgment result.

[0231] Step B160: Use the business logic construction error to train the data structure conversion unit, the logic conversion unit, the discrimination unit, and the business logic construction model until a second model convergence requirement is met.

[0232] In this embodiment, for target distributed multivariate data (including mobile payment data, personnel database data, scene monitoring data, etc.), the business logic construction model will analyze these data to generate a business logic description document.

[0233] For example, the business logic model analyzes the relationships between transaction amounts, payment times, and payment methods in mobile payment data, and age, gender, and other information in the personnel database. It finds that users aged 20-30 are more likely to use electronic payment methods for small purchases (under 100 yuan) at night (6-22 PM), and records this relationship pattern in the business logic description document. Throughout the entire processing process, this document is used as an iterative transfer condition for related operations. For example, when performing feature enhancement or model training on data, if the processing involves features related to age, payment time, and payment amount, the logical relationships in this document are referenced to ensure that the processing process conforms to the actual business logic of the data.

[0234] Assume that the distributed multivariate data of the second template includes simulated mobile payment and personnel database data. The mobile payment data has fields such as transaction amount, payment time, and payment method, and the personnel database data has fields such as age, gender, and occupation.

[0235] The data structure conversion unit is a neural network based structure, such as a multi-layer perceptron (MLP).

[0236] For the transaction amount in mobile payment data, assuming the original value range is 0-1000 yuan, the data structure conversion unit maps it to a new feature space. For example, through the first hidden layer of the MLP, the transaction amount of 100 yuan may be converted into a 3-dimensional vector [0.1, 0.2, 0.3]. This conversion process is calculated based on the weights and activation functions within the MLP. For the payment time, originally in hours (0-24), after conversion, it may become a 4-dimensional vector, representing the feature representation of different time periods. If the payment method is binary (cash or electronic payment), it may be converted into a 2-dimensional vector.

[0237] For age in the personnel database data, assuming the range is 0-100, it is converted into a 5-dimensional vector using the data structure conversion unit. This vector better represents the characteristic relationship of age within the entire data structure. Gender, if binary (male or female), is converted into a different vector representation. Occupation may be converted into a multidimensional weight vector based on predefined occupational category codes.

[0238] Finally, these converted vectors are combined to form a template data structure vector. For example, if the total dimension of the converted vector of mobile payment data is 10 and the total dimension of the converted vector of personnel database data is 15, then the dimension of the template data structure vector is 25.

[0239] The second template, the distributed multivariate data, corresponds to a logical description document that contains a description of the logical relationships between the data. For example, the document describes that, in a specific scenario, men aged 30-40 are more likely to use a certain payment method to make medium-value transactions (100,500 yuan) at noon on weekdays (12:00-14:00).

[0240] The logical transformation unit uses a technique based on word vectors and logical rule encoding. Each logical element in a document is transformed. For example, "age 30-40" might be encoded as a specific vector [0.1, 0.2, -0.1], "male" as [0.3, -0.2, 0.1], "noon on weekdays (12:00-14:00)" as [0.2, 0.3, 0.4], "a certain payment method" as [0.1, 0.2, 0.3], and "a medium amount (100-500 yuan)" as [0.1, -0.1, 0.2].

[0241] These encoded vectors are then combined in a certain logical order, such as through weighted summation or concatenation, to generate a template logical transformation vector. Assume that the dimension of this vector is ultimately 20 based on the encoding and combination method.

[0242] The discrimination unit receives the template data structure vector (25 dimensions) and the template logic description document (in the form of previously generated logic coding).

[0243] The discriminant unit has a predefined logical relevance discrimination model, which may be based on some statistical rules and logical algorithms. For example, it calculates the distance metric (such as Euclidean distance or cosine similarity) between certain dimensions in the template data structure vector and the corresponding logical element encoding vector in the template logical description document.

[0244] For the age-related dimension in the template data structure vector, calculate its distance to the encoding vector for "age between 30 and 40" in the template's logical description document. If the distance is small, it indicates a high correlation between the age-related aspects of the data structure and the logical description. Similar distance calculations are performed for other elements, such as payment time, payment amount, and gender.

[0245] Based on these distance measurement results, the template logical relevance judgment results are generated. For example, if the distance between most elements is less than a pre-set threshold (such as 0.5), the judgment result is "high relevance", represented by the number 1; if the distance between some elements is larger, the judgment result is "medium relevance", represented by the number 0; if the distance between many elements is large, the judgment result is "low relevance", represented by the number 1.

[0246] The business logic construction model takes the template data structure vector (25 dimensions) and the template logic description document (logical coding form) as input.

[0247] Assume that the business logic construction model is a neural network-based generative model, such as a recurrent neural network (RNN) with an attention mechanism.

[0248] The model first analyzes the template data structure vector and identifies the characteristics of each element (such as age, payment amount, and other related vectors). Then, it combines the logical relationships in the template logic description document to generate a prediction logic description document.

[0249] For example, based on the age vector representation in the template data structure vector and the logical relationship between age and payment behavior in the template logic description document, the model predicts that in a new scenario, users aged 25-35 are likely to use a new payment method for small purchases at specific times (such as weekend afternoons). This prediction logic description document contains a predictive description of the relationship between the data and has a similar structure to the original template logic description document, but the content is the prediction result based on the data structure vector and the original logical relationship.

[0250] Next, the loss function value is calculated between the predicted logical description document and the template logical description document. For example, using the cross-entropy loss function, if a predicted relationship in the predicted logical description document (such as the relationship between age and payment behavior) is inconsistent with the actual relationship in the template logical description document, a large loss value will be generated. For example, if the payment behavior of a user predicted to be 25-35 years old on weekend afternoons is significantly different from the actual relationship in the template logical description document, then this part will contribute significantly to the cross-entropy loss value.

[0251] At the same time, a loss function is calculated between the template data structure vector and the template logistic transformation vector. For example, the mean squared error (MSE) loss function is used to calculate the sum of squared errors between the corresponding elements of these two vectors. If there is a large difference between the age-related vector in the template data structure vector and the age-related encoding vector in the template logistic transformation vector, this MSE loss value will increase.

[0252] Finally, the business logic construction error is determined based on the loss function value between the predicted logic description document and the template logic description document, the loss function value between the template data structure vector and the template logic conversion vector, and the template logic relevance judgment result. For example, if the loss value between the predicted logic description document and the template logic description document is large, the loss value between the template data structure vector and the template logic conversion vector is also large, and the template logic relevance judgment result is "low relevance", then the business logic construction error will be high, indicating that the model performs poorly in constructing business logic.

[0253] Therefore, the business logic construction errors are used to train the data structure conversion unit, logic conversion unit, discrimination unit, and business logic construction model. For the data structure conversion unit (such as the multilayer perceptron), its internal weights and biases are adjusted based on the errors. For example, if age-related conversion results are found to result in large errors, the age-related weights from the input to the hidden layer and from the hidden layer to the output layer are adjusted.

[0254] For the logic conversion unit, the encoding method or weight of the logic element is adjusted according to the error. For example, if the correlation between the encoding vector of a certain logic element and the data structure vector is poor, resulting in a large error, the encoding method of this logic element is adjusted.

[0255] For the discrimination unit, the threshold or internal statistical rules for discriminating logical relevance are adjusted based on the error. For the business logic construction model (such as an RNN with an attention mechanism), the weights of its neural network, the parameters of the attention mechanism, etc. are adjusted based on the error. This training process is repeated until the second model convergence requirements are met. For example, when the loss function value between the prediction logic description document and the template logic description document is less than a pre-set threshold (such as 0.01), and the loss function value between the template data structure vector and the template logic conversion vector is also less than another threshold (such as 0.05), and the template logic relevance discrimination result reaches "high correlation" (such as the discrimination result is 1), it is considered that the second model convergence requirements are met and training is stopped.

[0256] Figure 2 The system 100 based on distributed multivariate data analysis shown includes: a processor 1001 and a memory 1003. The processor 1001 and the memory 1003 are connected, such as through a bus 1002. Optionally, the system 100 based on distributed multivariate data analysis may further include a transceiver 1004, which may be used for data interaction between the server and other servers, such as sending data and / or receiving data. It should be noted that in actual scheduling, the number of transceivers 1004 is not limited to one, and the structure of the system 100 based on distributed multivariate data analysis does not constitute a limitation on the embodiments of the present application.

[0257] The processor 1001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor 1001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0258] The bus 1002 may include a path for transmitting information between the above components. The bus 1002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 1002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 2 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0259] The memory 1003 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store program code and can be read by a computer, without limitation here.

[0260] The memory 1003 is used to store program codes for executing the embodiments of the present application, and the execution is controlled by the processor 1001. The processor 1001 is used to execute the program codes stored in the memory 1003 to implement the steps shown in the above method embodiments.

[0261] In addition, an embodiment of the present application further provides a readable storage medium, in which computer-executable instructions are pre-installed. When a processor executes the computer-executable instructions, the above method based on distributed multivariate data analysis is implemented.

[0262] Similarly, it should be noted that, in order to simplify the description of the present disclosure and thus facilitate the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of the present disclosure, multiple features may sometimes be combined into one embodiment, figure, or description thereof. Similarly, it should be noted that, in order to simplify the description of the present disclosure and thus facilitate the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of the present disclosure, multiple features may sometimes be combined into one embodiment, figure, or description thereof.

Claims

1. A method based on distributed multivariate data analysis, characterized in that: The method comprises: Performing a hiding operation on the to-be-derived field data in the target distributed multivariate data to generate hidden distributed multivariate data, wherein the target distributed multivariate data includes at least two types of data selected from the group consisting of mobile payment data, personnel database data, scene monitoring data, supervision platform data, and behavior log platform data; Dynamically updating node extraction is performed on the dynamic field data in the target distributed multivariate data to generate dynamic updating node data; the field data to be derived includes the dynamic field data to be derived in the dynamic field data; Marking the to-be-derived field data in the target distributed multivariate data to generate derived field marked data; Based on the target derived template knowledge data, the dynamically updated node data and the derived field tag data, feature enhancement is performed on the hidden distributed multivariate data to generate derived distributed multivariate data that combines the derived feature dimensions expressed by the target derived template knowledge data; Constructing a training template of a corresponding machine learning network model based on the derived distributed multivariate data and the target distributed multivariate data; The method of performing feature enhancement on the hidden distributed multivariate data based on the target derived template knowledge data, the dynamically updated node data, and the derived field tag data to generate derived distributed multivariate data that combines the derived feature dimensions expressed by the target derived template knowledge data includes: Embedding the target derived template knowledge data, the hidden distributed multivariate data, and the dynamically updated node data respectively to generate a first embedded feature vector for the target derived template knowledge data, a second embedded feature vector for the hidden distributed multivariate data, and a third embedded feature vector for the dynamically updated node data; performing iterative optimization processing on randomly generated distributed multivariate data based on the first embedded feature vector, the second embedded feature vector, the third embedded feature vector, and the derived field tag data to generate a derived embedded feature vector; The derived embedded feature vector is embedded and restored to generate derived distributed multivariate data that combines the derived feature dimensions expressed by the target derived template knowledge data.

2. The method based on distributed multivariate data analysis according to claim 1, characterized in that: The iterative optimization processing of the randomly generated distributed multivariate data based on the first embedded feature vector, the second embedded feature vector, the third embedded feature vector, and the derived field tag data to generate the derived embedded feature vector includes: Performing feature conversion on the first embedded feature vector using a feature conversion model to generate feature conversion data; The second embedded feature vector, the third embedded feature vector and the derived field label data are used as iterative transfer conditions, and an iterative optimization model is used to iteratively optimize the randomly generated distributed multivariate data according to the feature conversion data to generate the derived embedded feature vector.

3. The method based on distributed multivariate data analysis according to claim 2, characterized in that: The feature conversion model includes X first feature focusing units, and the iterative optimization model includes X second feature focusing units; The feature conversion data includes the first feature focusing results loaded by each of the first feature focusing units; The method of using the second embedded feature vector, the third embedded feature vector, and the derived field tag data as iterative transfer conditions, utilizing an iterative optimization model, and performing iterative optimization processing on the randomly generated distributed multivariate data according to the feature conversion data to generate the derived embedded feature vector includes: Using the second embedded feature vector, the third embedded feature vector, and the derived field tag data as iterative transfer conditions, using an iterative optimization model, performing Y rounds of optimization iterations on the randomly generated distributed multivariate data according to the feature conversion data, and using the optimized embedded feature vector generated by the Y-th round of optimization iteration as the derived embedded feature vector; The specific steps of the optimization iteration of round a include: Obtain a reference embedding feature vector for the a-th round of optimization iteration, wherein the reference embedding feature vector for the first round of optimization iteration is generated by fusing the second embedding feature vector, the third embedding feature vector, the derived field tag data, and the random embedding feature vector corresponding to the randomly generated distributed multivariate data; when 1<a≤Y, the reference embedding feature vector for the a-th round of optimization iteration is the optimized embedding feature vector for the a-1-th round of optimization iteration; In the bth second feature focusing unit, interactive calculation is performed based on the first feature focusing result generated by the bth first feature focusing unit and the second feature focusing result of the bth second feature focusing unit for the ath round of optimization iteration, to generate an interactive embedding feature vector generated by the bth second feature focusing unit for the ath round of optimization iteration; 1≤b≤X, b is a positive integer; the second feature focusing result of the first second feature focusing unit for the ath round of optimization iteration is determined based on the reference embedding feature vector of the ath round of optimization iteration; Performing optimization iteration on the interactive embedding feature vector generated by the b-th second feature focusing unit to generate a b-th temporary optimized iterative embedding feature vector for the a-th round of optimization iteration; If b is less than X, determining the second feature focusing result of the b+1th second feature focusing unit for the ath round of optimization iteration based on the bth temporary optimization iteration embedded feature vector for the ath round of optimization iteration, counting and accumulating b, and returning to execute the step of interactively calculating in the bth second feature focusing unit based on the first feature focusing result generated by the bth first feature focusing unit and the second feature focusing result of the bth second feature focusing unit for the ath round of optimization iteration to generate the interactive embedded feature vector generated by the bth second feature focusing unit for the ath round of optimization iteration; If b=X, the bth temporary optimized iteration embedding feature vector for the ath round of optimization iteration is used as the optimized embedding feature vector for the ath round of optimization iteration.

4. The method based on distributed multivariate data analysis according to claim 3, characterized in that: The first feature focusing result includes a first focusing index and a first focusing feature; the second feature focusing result includes a second focusing index, a second focusing feature and a target interaction index; In the b-th second feature focusing unit, interactive calculation is performed based on the first feature focusing result generated by the b-th first feature focusing unit and the second feature focusing result of the b-th second feature focusing unit for the a-th round of optimization iteration, to generate the interactive embedding feature vector generated by the b-th second feature focusing unit for the a-th round of optimization iteration, including: In the bth second feature focusing unit, interactively operate on the first focusing index and the second focusing index of the bth second feature focusing unit for the ath round of optimization iteration to generate a target focusing index of the bth second feature focusing unit for the ath round of optimization iteration; Performing an interactive operation on the first focusing feature and the second focusing feature of the bth second feature focusing unit for the ath round of optimization iteration to generate a target focusing feature of the bth second feature focusing unit for the ath round of optimization iteration; Determining an interaction factor of the bth second feature focusing unit for the ath round of optimization iteration based on the target interaction index of the bth second feature focusing unit for the ath round of optimization iteration and the target focusing index; Based on the interaction factor of the b-th second feature focusing unit for the a-th round of optimization iteration, the corresponding target focusing feature is interactively calculated to generate an interactive embedding feature vector generated by the b-th second feature focusing unit for the a-th round of optimization iteration.

5. The method based on distributed multivariate data analysis according to claim 2, characterized in that: The iterative transfer condition also includes a business logic description document for the target distributed multivariate data.

6. The method based on distributed multivariate data analysis according to any one of claims 2 to 5, characterized in that: The method further comprises: Acquire template-derived template knowledge data, template-hidden distributed multivariate data, template-derived field tag data, and template-dynamically-updated node data; the template-derived template knowledge data is determined based on the template-to-be-derived field data in the first template-distributed multivariate data; the template-hidden distributed multivariate data is generated by performing a hiding operation on the template-to-be-derived field data in the first template-distributed multivariate data; the template-derived field tag data is a tagging result of the template-to-be-derived field data; the template-dynamically-updated node data is obtained by extracting the template-dynamic field data in the first template-distributed multivariate data through a dynamic update node; the template-to-be-derived field data includes the dynamic-to-be-derived field data in the template-dynamic field data; Embedding the template-derived template knowledge data, the template-hidden distributed multivariate data, and the template-dynamically updated node data respectively to generate a first template-embedded feature vector for the template-derived template knowledge data, a second template-embedded feature vector for the template-hidden distributed multivariate data, and a third template-embedded feature vector for the template-dynamically updated node data; Performing feature conversion on the first template embedded feature vector using a feature conversion model to generate template feature conversion data; Using the second template embedded feature vector, the third template embedded feature vector, and the template-derived field tag data as iterative transfer conditions, using an iterative optimization model, and performing iterative optimization processing on the template randomly generated distributed multivariate data according to the template feature conversion data to generate a template-derived embedded feature vector; Embedding and restoring the template-derived embedded feature vector to generate template-derived distributed multivariate data; The feature conversion model and the iterative optimization model are trained according to the loss function value between the template-derived distributed multivariate data and the first template distributed multivariate data until the first model convergence requirement is met.

7. The method based on distributed multivariate data analysis according to claim 6, characterized in that: The acquisition of template-derived template knowledge data, template-hidden distributed multivariate data, template-derived field tag data, and template-dynamically updated node data includes: Extracting derivable fields from the first template distributed multivariate data to determine template field data to be derived from the first template distributed multivariate data; Marking the template's to-be-derived field data to generate template-derived field marking data; Performing a hiding operation on the template-to-be-derived field data in the first template distributed multivariate data according to the template-derived field tag data to generate template-hidden distributed multivariate data; Processing the first template distributed multivariate data according to the template-derived field tag data to generate the template-derived template knowledge data; Extracting dynamic update nodes from the template dynamic field data in the first template distributed multivariate data to generate the template dynamic update node data; The step of performing a hiding operation on the template-to-be-derived field data in the first template distributed multivariate data based on the template-derived field tag data to generate template-hidden distributed multivariate data includes: Regularly converting the first template distributed multivariate data into a first eigenvalue interval to generate first temporary template distributed multivariate data; Regularly converting the template-derived field tag data into a second characteristic value interval to generate first template-derived field tag data, wherein the second characteristic value interval is within the first characteristic value interval; Fusing the first temporary template distributed multivariate data and the first template derived field tag data element by element to generate second temporary distributed multivariate data; Performing an inverse regularization conversion on the second temporary distributed multivariate data to generate the template hidden distributed multivariate data; The step of processing the first template distributed multivariate data based on the template-derived field tag data to generate the template-derived template knowledge data includes: Regularly converting the template-derived field tag data into a second characteristic value interval to generate second template-derived field tag data; Fusing the second template-derived field tag data and the first template distributed multivariate data element by element to generate the template-derived template knowledge data; The extracting derivable fields from the first template distributed multivariate data to determine the template field data to be derived in the first template distributed multivariate data includes: Performing structural feature conversion on the first template distributed multivariate data using a first feature mapping structure to generate a template structure embedding feature vector; Performing logical feature conversion on the first template distributed multivariate data using a second feature mapping structure to generate a template logical embedding feature vector; Performing dimensionality reduction processing on the template structure embedding feature vector to generate a reduced-dimensionality structure embedding feature vector; Performing dimensionality-increasing processing on the template logic embedding feature vector to generate a dimensionality-increasing logic embedding feature vector; Combining the template logic embedding feature vector with the dimension reduction structure embedding feature vector to generate a first template combination embedding feature vector; Combining the template structure embedding feature vector with the dimension-raising logic embedding feature vector to generate a second template combination embedding feature vector; Merging the first template combination embedding feature vector and the second template combination embedding feature vector to generate a template embedding feature vector; Based on the template embedding feature vector, template-to-be-derived field data in the first template distributed multivariate data is determined.

8. The method based on distributed multivariate data analysis according to any one of claims 2 to 5, characterized in that: The iterative transfer condition also includes a business logic description document generated by the business logic construction model based on the target distributed multivariate data; The method further comprises: Performing feature conversion on the second template distributed multivariate data using a data structure conversion unit to generate a template data structure vector; Performing feature conversion on the template logic description document corresponding to the second template distributed multivariate data using a logic conversion unit to generate a template logic conversion vector; Using a discrimination unit to perform logic relevance discrimination on the template data structure vector and the template logic description document to generate a template logic relevance discrimination result; Using the business logic to build a model based on the template data structure vector and the template logic description document, a prediction logic description document is generated; Determining a business logic construction error based on a loss function value between the prediction logic description document and the template logic description document, a loss function value between the template data structure vector and the template logic conversion vector, and a result of determining the template logic relevance; The data structure conversion unit, the logic conversion unit, the discrimination unit, and the business logic construction model are trained using the business logic construction error until a second model convergence requirement is met.

9. A system based on distributed multivariate data analysis, characterized in that: The system based on distributed multivariate data analysis includes a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions, and the machine-executable instructions are loaded and executed by the processor to implement the method based on distributed multivariate data analysis according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Enterprise big data mining method and system based on artificial intelligence

    CN118094639A

  • Battery carbon footprint distributed calculation method and system for protecting carbon data privacy

    CN118228318A