Labeling Method, Device, Equipment and Storage Medium for Batch Data
Through the combination of self-supervised learning model, graph neural network and multi-layer perceptron model, multi-dimensional data features are extracted and integrated, and the problem of insufficient accuracy and diversity of labeling methods in government affairs systems is solved, and efficient enterprise label generation is achieved.
Patent Information
- Application Number
- CN202411441240.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-16
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-10-16
AI Technical Summary
The prior art labeling method of enterprise data in government affairs systems is difficult to process complex and multi-dimensional data, resulting in insufficient accuracy and diversity of label generation, and low efficiency, making it difficult to meet the real-time processing needs of large-scale data.
The self-supervised learning model, graph neural network model and multi-layer perceptron model are used to extract the deep features of unstructured data, the relationship characteristics of relational network data and the label prediction results of structured data, and the feature fusion and correction are carried out through the full connection layer to generate target label results.
It improves the accuracy and diversity of enterprise label generation, is suitable for labeling processing of large-scale enterprise data, and meets the real-time processing needs of government affairs systems.
Smart Images

Figure CN119377819B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method, apparatus, device, and storage medium for tagging batch data. Background Art
[0002] In modern data processing and management, tagging is a key step. Especially when managing enterprises in a government affairs system, enterprise data is huge and complex, including structured, unstructured, and relational network data. Existing tagging methods are mostly based on a single data source and often only use structured data for simple classification and tag generation, making it difficult to process complex and multi-dimensional data, resulting in insufficient accuracy and diversity of tag generation. In addition, traditional data processing methods are less efficient and difficult to meet the real-time processing requirements of large-scale data in the government affairs system.
[0003] In summary, the problems existing in the prior art need to be solved urgently. Summary of the Invention
[0004] The present invention provides a method, apparatus, device, and storage medium for tagging batch data to solve the defects in the prior art and achieve automatic data tagging.
[0005] The present invention provides a method for tagging batch data, including:
[0006] Obtaining data to be processed, where the data to be processed includes unstructured data, relational network data, and structured data;
[0007] Inputting the unstructured data into a feature extraction model to obtain deep features;
[0008] Inputting the relational network data into a relational feature model to obtain relational features;
[0009] Inputting the structured data into a tag prediction model to obtain a tag prediction result;
[0010] Correcting the tag prediction result according to the deep features and the relational features to obtain a target tag result.
[0011] According to the method for tagging batch data provided by the present invention, the step of correcting the tag prediction result according to the deep features and the relational features to obtain a target tag result specifically includes:
[0012] Inputting the deep features, the relational features, and the tag prediction result into a first fully connected layer for non-linear transformation;
[0013] Input the transformed deep features, relationship features, and label prediction results into the second fully connected layer for feature fusion to obtain the target feature vector;
[0014] Input the target feature vector into the output layer to obtain the target label result.
[0015] According to a method for labeling batch data provided by the present invention, the step of inputting the transformed deep features, relationship features, and label prediction results into the second fully connected layer for feature fusion is specifically implemented in the following manner:
[0016] Z fused = FC2([FC1(Z self-supervised ), FC1(Z GNN ), FC1(Z MLP )])
[0017] where FC1 represents the operation of the first fully connected layer, FC2 represents the operation of the second fully connected layer, Z self-supervised is the deep feature, Z GNN is the relationship feature, and Z MLP is the label prediction result.
[0018] According to a method for labeling batch data provided by the present invention, the feature extraction model is a self-supervised learning model, the relationship feature model is a GNN model, and the label prediction model is an MLP model.
[0019] According to a method for labeling batch data provided by the present invention, the activation function of the output layer is Sigmoid.
[0020] According to a method for labeling batch data provided by the present invention, before the step of obtaining the data to be processed, the method further includes:
[0021] Preprocess the data to be processed, and the preprocessing includes missing value filling, outlier handling, and normalization processing.
[0022] According to a method for labeling batch data provided by the present invention, after the step of correcting the label prediction result according to the deep feature and the relationship feature to obtain the target label result, the method further includes:
[0023] Obtain label correction data;
[0024] Correct the target label result according to the label correction data.
[0025] The present invention also provides a device for labeling batch data, including:
[0026] A data acquisition module, configured to acquire data to be processed, where the data to be processed includes unstructured data, relational network data, and structured data;
[0027] A deep feature module, configured to input the unstructured data into a feature extraction model to obtain deep features;
[0028] A relational feature module, configured to input the relational network data into a relational feature model to obtain relational features;
[0029] A label prediction module, configured to input the structured data into a label prediction model to obtain a label prediction result;
[0030] A result correction module, configured to correct the label prediction result according to the deep features and the relational features to obtain a target label result.
[0031] The present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, the method for labeling batch data as described in any one of the above is implemented.
[0032] The present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method for labeling batch data as described in any one of the above is implemented.
[0033] The present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method for labeling batch data as described in any one of the above is implemented.
[0034] The method, device, equipment, and storage medium for labeling batch data provided by the present invention acquire data to be processed, where the data to be processed includes unstructured data, relational network data, and structured data; input the unstructured data into a feature extraction model to obtain deep features; input the relational network data into a relational feature model to obtain relational features; input the structured data into a label prediction model to obtain a label prediction result; correct the label prediction result according to the deep features and the relational features to obtain a target label result. The method of the present invention extracts deep features from unstructured data, extracts relational features from relational network data, and combines the label prediction result of structured data, making full use of the information of multi-dimensional data, improving the accuracy and diversity of enterprise label generation, and avoiding information loss or deviation that may be caused by relying only on a single data source. Description of the Drawings
[0035] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0036] Figure 1 is a schematic flow chart of the method for tagging batch data provided by the present invention;
[0037] Figure 2 is a schematic structural diagram of the device for tagging batch data provided by the present invention;
[0038] Figure 3 is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0039] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0040] To solve the problems in the prior art, the present invention proposes a method for tagging batch data to achieve automatic tagging of data. The following describes the method for tagging batch data. As Figure 1 shown, it includes but is not limited to the following steps:
[0041] Step 110: Obtain the data to be processed, where the data to be processed includes unstructured data, relationship network data, and structured data.
[0042] In this step, first obtain the enterprise data to be processed from multiple data sources. The data includes:
[0043] Unstructured data: such as the company profile, news reports, audit reports, product descriptions, etc. of an enterprise. This type of data cannot be directly used for structured analysis and needs further processing.
[0044] Relationship network data: such as the equity relationship, cooperation relationship, upstream and downstream supply chain information, etc. between enterprises. This data can be represented by a graph structure and is used to reflect the relevance between enterprises.
[0045] Structured data: such as the registration information, industry classification, business status, financial statements, administrative approval records, etc. of an enterprise. This type of data already has clear fields and formats and is suitable for traditional statistical analysis and modeling.
[0046] These data can be obtained from multiple sources, such as databases in government affairs systems, third-party data platforms, web crawlers, or the official websites of enterprises themselves.
[0047] Step 120: Input the unstructured data into a feature extraction model to obtain deep features.
[0048] In this step, the system inputs the obtained unstructured data into the feature extraction model. The feature extraction model can adopt a self-supervised learning model, such as a model based on the Transformer architecture (e.g., BERT, GPT, etc.). These models can extract data related to labels in the data, and then use encoding models such as BERT to convert the extracted content into deep features.
[0049] Specifically, the unstructured data is first preprocessed, including operations such as removing irrelevant characters, stop word filtering, and word segmentation. The preprocessed text is input into the self-supervised learning model, and through the operation of a multi-layer neural network, deep semantic features are extracted. These deep features can reflect the specific behavior patterns or attributes of the enterprise in the text data, and the generated feature vectors can be used for subsequent labeling analysis.
[0050] Step 130: Input the relational network data into a relational feature model to obtain relational features.
[0051] In this step, the system inputs the relational network data into the relational feature model. This relational feature model is usually a graph neural network (GNN) model, which is used to analyze the complex relational network between enterprises and generate the feature representation of each enterprise in the network.
[0052] Specifically, the system takes enterprises as nodes in the graph and the relationships between enterprises (such as equity relationships, cooperation relationships, upstream and downstream of the supply chain, etc.) as edges in the graph. Through the GNN model, the system can effectively capture the interactions between enterprises and extract the features of each enterprise in the relational network (i.e., relational features). These features reflect the position of the enterprise in the entire economic or industrial ecosystem and the strength of its relationship with other enterprises.
[0053] The GNN model propagates and aggregates the information of neighbor nodes layer by layer through multiple graph convolution operations, thereby generating the feature representation of each node (enterprise).
[0054] Step 140: Input the structured data into a label prediction model to obtain label prediction results.
[0055] In this step, the system inputs structured data into the label prediction model. The label prediction model can adopt a multi-layer perceptron (MLP) model, which is good at processing structured data and can generate preliminary label prediction results based on the basic information of the enterprise.
[0056] Specifically, when implemented, structured data (such as registered capital, number of employees, business industry, administrative approval situation, etc.) is input into the MLP model. After calculation by multiple layers of neural networks, one or more label prediction results are output. These prediction results can be the probability distribution of an enterprise belonging to a certain label category, or a multi-class label indicator.
[0057] Step 150: According to the deep features and the relationship features, correct the label prediction results to obtain the target label results.
[0058] In this step, the system fuses the deep features, the relationship features, and the label prediction results to generate the final target label results.
[0059] Specifically, first, the deep features extracted by self-supervised learning, the relationship features extracted by GNN, and the label prediction results of MLP are input into a fusion module. This fusion module is processed through multiple fully connected layers (i.e., MLP). First, non-linear transformation is performed on the three groups of input features in the first fully connected layer. The transformed features are further fused in the second fully connected layer to generate a unified feature representation.
[0060] The target feature vector after feature fusion is input into the output layer, and the final target label results are generated through an appropriate activation function (such as Sigmoid). The target label results not only combine the label prediction of structured data, but also incorporate the deep information in unstructured data and the upstream and downstream associations in the relationship network, thus more accurately reflecting the characteristics and behavior patterns of the enterprise.
[0061] In this embodiment, by obtaining different types of data, and using self-supervised learning models, GNN models, and MLP models to extract deep features, relationship features, and prediction results of structured data respectively, and finally correcting the label prediction through feature fusion to generate accurate enterprise labels. This method effectively improves the accuracy and diversity of label generation, and is applicable to the labeling process of large-scale enterprise data in fields such as government affairs and finance.
[0062] As a further optional embodiment, the step of correcting the label prediction results according to the deep features and the relationship features to obtain the target label results specifically includes:
[0063] Input the deep features, the relationship features, and the label prediction results into the first fully connected layer for non-linear transformation;
[0064] Input the transformed deep features, relational features, and label prediction results into the second fully connected layer for feature fusion to obtain the target feature vector;
[0065] Input the target feature vector into the output layer to obtain the target label result.
[0066] As a further optional embodiment, the step of inputting the transformed deep features, relational features, and label prediction results into the second fully connected layer for feature fusion is specifically implemented in the following manner:
[0067] Z fused = FC2([FC1(Z self-supervised ), FC1(Z GNN ), FC1(Z MLP )])
[0068] where FC1 represents the operation of the first fully connected layer, FC2 represents the operation of the second fully connected layer, Z self-supervised is the deep feature, Z GNN is the relational feature, and Z MLP is the label prediction result.
[0069] In this step, first input the deep features obtained from the self-supervised learning model, the relational features obtained from the GNN model, and the label prediction results obtained from the MLP model into the first fully connected layer (FC1) for processing respectively. The operation of the fully connected layer realizes the non-linear mapping of features through linear transformation plus an activation function. The purpose of this operation is to perform preliminary dimensionality reduction and non-linear transformation on each type of feature to capture more complex relationships between features.
[0070] In this step, perform a concatenation operation on the three groups of features output by the first fully connected layer to form a unified feature vector. Then, input this feature vector into the second fully connected layer (FC2) for further feature fusion. Through the feature fusion of the fully connected layer, the system can capture the mutual relationships between different types of features, thereby generating a more representative target feature vector.
[0071] The specific operation is as follows:
[0072] Z fused = FC2([FC1(Z self-supervised ), FC1(Z GNN ), FC1(Z MLP )])
[0073] where FC1 represents the operation of the first fully connected layer, FC2 represents the operation of the second fully connected layer, Z self-supervised is the deep feature, ZGNN is a relational feature, Z MLP is the label prediction result.
[0074] In the last step, the fused target feature vector Z fused is input into the output layer. According to the specific task requirements, the output layer can use an appropriate activation function for label prediction. Common activation functions include:
[0075] Sigmoid: Used for binary classification label problems, and the output result is the probability value that the label belongs to a certain class.
[0076] Softmax: Used for multi-classification label problems, and the probability distribution of each class label is output.
[0077] According to the application scenario, the output layer will generate the final target label result. This result has been comprehensively corrected by deep features, relational features, and MLP predictions, and more accurately reflects the multi-dimensional features of the enterprise.
[0078] Through the above steps, this embodiment can effectively perform non-linear transformation and feature fusion on the deep features extracted by self-supervised learning, the relational features extracted by GNN, and the prediction results of MLP. Through the processing of two fully connected layers, the system can comprehensively extract meaningful features from multi-dimensional data and generate accurate label results. This method greatly improves the accuracy and diversity of label generation and is applicable to large-scale enterprise data labeling tasks.
[0079] As a further optional embodiment, the feature extraction model is a self-supervised learning model, the relational feature model is a GNN model, and the label prediction model is an MLP model.
[0080] The self-supervised learning model is used to process unstructured data, such as text, images, or other complex enterprise-related information. This model extracts deep features from the data through pre-training without explicit labels. Self-supervised learning can extract representative high-dimensional feature vectors from a large amount of data by designing specific tasks (such as masked word prediction, contrastive learning, etc.). Finally, these high-dimensional feature vectors can capture the complex semantic information of the enterprise, such as the enterprise's business scope, development history, etc. The role of this model is to generate deep features related to the enterprise and provide the input for subsequent feature fusion.
[0081] Graph Neural Network (GNN) models are mainly used to process relational network data between enterprises, such as equity structures, business cooperation relationships, management cross - overs, upstream and downstream supply chains, etc. between enterprises. These relationships form a graph structure, where nodes represent enterprises and edges represent certain associations between enterprises. Through the propagation and aggregation of information from node neighbors, GNN models can effectively capture the relational characteristics between enterprises. Through multi - layer graph convolution operations, GNN can extract the global representation of nodes (enterprises) in the relational network, that is, relational characteristics. These relational characteristics reflect the mutual influence between enterprises and help to consider the business ecological environment in which the enterprises are located when generating labels.
[0082] Multi - Layer Perceptron (MLP) models are used to process structured data, such as enterprise registration information, annual report data, financial data, number of social security participants, number of patents, etc. These data are usually standardized numerical data. Through a multi - layer fully - connected network structure and combined with non - linear activation functions (such as ReLU), MLP models can extract features from these structured data and generate preliminary label prediction results.
[0083] By learning these structured data, the MLP model can predict the label categories to which the enterprise belongs, such as "high - tech enterprise", "key regulatory object", etc. This label prediction result will serve as the basis for subsequent feature fusion and correction.
[0084] As a further optional embodiment, the activation function of the output layer is Sigmoid.
[0085] The Sigmoid function is an activation function commonly used in binary classification tasks, and its output value range is between 0 and 1. When the vector after feature fusion is input into the output layer, the Sigmoid function will convert it into a probability value, representing the probability that the sample belongs to a certain specific category. A threshold (such as 0.5) is usually set for this probability value to judge whether the label is valid. For example, if the output value is greater than 0.5, the label is 1; if it is less than 0.5, the label is 0.
[0086] As a further optional embodiment, before the step of obtaining the data to be processed, the method further includes:
[0087] Pre - process the data to be processed, and the pre - processing includes missing value filling, outlier handling, and normalization processing.
[0088] In practical applications, some fields in the dataset may have missing values (such as certain financial indicators or administrative information of enterprises). The existence of missing values will affect the learning effect of the model, so these missing data need to be processed during pre - processing. Common processing methods include: mean filling, median filling, mode filling, and prediction filling.
[0089] Outliers are extreme values in a dataset that are far from other data points. They may be caused by data input errors, measurement errors, or other abnormal situations. If outliers are not handled, the model may overfit to these extreme values, thus affecting the accuracy of label prediction. Common outlier handling methods include: direct elimination, range limitation, and scaling.
[0090] Normalization is used to scale data to a unified scale to ensure that different features do not affect the performance of the model due to differences in numerical ranges.
[0091] By preprocessing the data to be processed, including missing value imputation, outlier handling, and normalization, the quality and consistency of the data can be greatly improved, ensuring that the model can better learn the features of the data. The preprocessed data will be more suitable for input into subsequent self-supervised learning models, GNN models, and MLP models, thereby improving the accuracy and robustness of label generation.
[0092] As a further optional embodiment, after the step of correcting the label prediction result based on the deep feature and the relationship feature to obtain the target label result, it further includes:
[0093] Obtain label correction data;
[0094] Correct the target label result according to the label correction data.
[0095] Label correction data refers to the reference data for further adjusting the labels automatically generated by the model based on actual usage feedback, expert annotation, or manual intervention information provided by external data sources during the operation of the system. The ways to obtain label correction data include:
[0096] User feedback: After the enterprise labels are generated, the system can provide an interface for supervisors or other users to feedback on the accuracy of the labels, such as deleting inappropriate labels and adding missing labels.
[0097] Expert annotation: Invite domain experts to manually review the labels of certain enterprises and record the corrected labels recommended by the experts.
[0098] External data source update: By integrating external trusted data sources (such as industry reports, updated policies of regulatory departments, etc.), the existing enterprise label data can be further improved.
[0099] By obtaining this label correction data, a more accurate reference can be provided for the label generation result, ensuring that the enterprise labels match the actual business scenario.
[0100] After obtaining the label correction data, the system will adjust the originally generated label results according to this data. The specific operations may include:
[0101] Label addition: If the correction data indicates that a certain enterprise should add a certain label, the system will add this label to the label set of this enterprise.
[0102] Label deletion: If the feedback indicates that a label of a certain enterprise is incorrect or inappropriate, the system will delete this label from the label set of this enterprise.
[0103] Label replacement: When the correction data indicates that the original label is not accurate enough or the description is not comprehensive enough, the system will replace the label according to the correction data.
[0104] By introducing the label correction data and its corresponding correction steps, the flexibility and accuracy of the label generation system can be effectively improved. Especially in the initial stage of the system or during large-scale deployment, combined with means such as user feedback and expert annotation, the quality of enterprise labels can be gradually optimized to ensure that the labels generated by the system are more in line with actual needs.
[0105] This mechanism can also achieve self-improvement. With the accumulation of more correction data, the label generation system can continuously learn, and the correction rules become more intelligent, thereby greatly improving the accuracy of model prediction and the business adaptability of enterprise labels.
[0106] Next, the labelization device for batch data provided by the present invention will be described. As Figure 2 shown, the labelization device for batch data described below can be correspondingly referred to the labelization method for batch data described above.
[0107] A labelization device for batch data includes:
[0108] A data acquisition module 210, configured to acquire data to be processed, where the data to be processed includes unstructured data, relationship network data, and structured data;
[0109] A deep feature module 220, configured to input the unstructured data into a feature extraction model to obtain deep features;
[0110] A relationship feature module 230, configured to input the relationship network data into a relationship feature model to obtain relationship features;
[0111] A label prediction module 240, configured to input the structured data into a label prediction model to obtain a label prediction result;
[0112] A result correction module 250, configured to correct the label prediction result according to the deep features and the relationship features to obtain a target label result.
[0113] Figure 3 Illustrates a schematic diagram of the physical structure of an electronic device, as Figure 3 shown. The electronic device may include: a processor 310, a communications interface 320, a memory 330, and a communication bus 340. Among them, the processor 310, the communication interface 320, and the memory 330 communicate with each other through the communication bus 340. The processor 310 can call the logical instructions in the memory 330 to execute the method for tagging batch data, and the method includes:
[0114] Obtain the data to be processed, where the data to be processed includes unstructured data, relational network data, and structured data;
[0115] Input the unstructured data into a feature extraction model to obtain deep features;
[0116] Input the relational network data into a relational feature model to obtain relational features;
[0117] Input the structured data into a label prediction model to obtain a label prediction result;
[0118] According to the deep features and the relational features, correct the label prediction result to obtain a target label result.
[0119] In addition, when the logical instructions in the above-mentioned memory 330 are implemented in the form of software function units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. And the aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.
[0120] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program, and the computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for tagging batch data provided by the above-mentioned various methods. The method includes:
[0121] Obtain data to be processed, where the data to be processed includes unstructured data, relational network data, and structured data;
[0122] Input the unstructured data into a feature extraction model to obtain deep features;
[0123] Input the relational network data into a relational feature model to obtain relational features;
[0124] Input the structured data into a label prediction model to obtain a label prediction result;
[0125] According to the deep features and the relational features, correct the label prediction result to obtain a target label result.
[0126] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a method for labeling batch data provided by the above-mentioned various methods. The method includes:
[0127] Obtain data to be processed, where the data to be processed includes unstructured data, relational network data, and structured data;
[0128] Input the unstructured data into a feature extraction model to obtain deep features;
[0129] Input the relational network data into a relational feature model to obtain relational features;
[0130] Input the structured data into a label prediction model to obtain a label prediction result;
[0131] According to the deep features and the relational features, correct the label prediction result to obtain a target label result.
[0132] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0133] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for tagging batch data, characterized in that, Including: Obtain the data to be processed, where the data to be processed includes unstructured data, relationship network data, and structured data. The unstructured data includes enterprise company profiles, enterprise news reports, enterprise audit reports, and enterprise product descriptions. The relationship network data includes enterprise equity relationships, enterprise cooperation relationships, and upstream and downstream supply chain information. The structured data includes industry classifications, operating conditions, financial statements, and administrative approval records; Input the unstructured data into the feature extraction model to obtain deep features; Input the relationship network data into the relationship feature model to obtain relationship features; Input the structured data into the label prediction model to obtain label prediction results; According to the deep features and the relationship features, correct the label prediction results to obtain target label results, and the target label results are used to manage enterprises; The step of correcting the label prediction results according to the deep features and the relationship features to obtain target label results specifically includes: Input the deep features, the relationship features, and the label prediction results into the first fully connected layer for non-linear transformation; Input the transformed deep features, relationship features, and label prediction results into the second fully connected layer for feature fusion to obtain a target feature vector; Input the target feature vector into the output layer to obtain target label results; The step of inputting the transformed deep features, relationship features, and label prediction results into the second fully connected layer for feature fusion is specifically implemented in the following way: Z fused = FC2([FC1(Z self-supervised ), FC1(Z GNN ), FC1(Z MLP )]) Among them, FC1 represents the operation of the first fully connected layer, FC2 represents the operation of the second fully connected layer, Z self-supervised is the deep feature, Z GNN is the relationship feature, Z MLP is the label prediction result; Specifically, perform a concatenation operation on the three groups of features output by the first fully connected layer to form a unified feature vector, and then input the feature vector into the second fully connected layer for further feature fusion.
2. The tagging method for batch data according to claim 1, wherein The feature extraction model is a self-supervised learning model, the relationship feature model is a GNN model, and the label prediction model is an MLP model.
3. The tagging method for batch data according to claim 2, wherein The activation function of the output layer is Sigmoid.
4. The method for tagging batch data according to claim 1, characterized in that, Before the step of obtaining the data to be processed, the method further includes: Preprocess the data to be processed, and the preprocessing includes missing value filling, outlier handling, and normalization processing.
5. The tagging method for batch data according to claim 1, wherein After the step of correcting the label prediction results according to the deep features and the relationship features to obtain target label results, it further includes: Obtain label correction data; Correct the target label results according to the label correction data.
6. A tagging device for batch data, characterized in that, Including: A data acquisition module for obtaining the data to be processed, where the data to be processed includes unstructured data, relationship network data, and structured data. The unstructured data includes enterprise company profiles, enterprise news reports, enterprise audit reports, and enterprise product descriptions. The relationship network data includes enterprise equity relationships, enterprise cooperation relationships, and upstream and downstream supply chain information. The structured data includes industry classifications, operating conditions, financial statements, and administrative approval records; A deep feature module for inputting the unstructured data into the feature extraction model to obtain deep features; A relationship feature module for inputting the relationship network data into a relationship feature model to obtain relationship features; A label prediction module for inputting the structured data into a label prediction model to obtain a label prediction result; A result correction module for correcting the label prediction result according to the deep features and the relationship features to obtain a target label result, where the target label result is used for managing an enterprise; The step of correcting the label prediction result according to the deep features and the relationship features to obtain a target label result specifically includes: Inputting the deep features, the relationship features, and the label prediction result into a first fully connected layer for non-linear transformation; Inputting the transformed deep features, relationship features, and label prediction result into a second fully connected layer for feature fusion to obtain a target feature vector; Inputting the target feature vector into an output layer to obtain a target label result; The step of inputting the transformed deep features, relationship features, and label prediction result into a second fully connected layer for feature fusion is specifically implemented by the following method: Z fused = FC2([FC1(Z self-supervised ), FC1(Z GNN ), FC1(Z MLP )]) Among them, FC1 represents the operation of the first fully connected layer, FC2 represents the operation of the second fully connected layer, Z self-supervised is the deep feature, Z GNN is the relational feature, Z MLP is the label prediction result; Specifically, perform a concatenation operation on the three groups of features output by the first fully connected layer to form a unified feature vector, and then input the feature vector into the second fully connected layer for further feature fusion.
7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method for labeling batch data according to any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method for labeling batch data according to any one of claims 1 to 5.
Citation Information
Patent Citations
Industry classification method and device, terminal equipment and storage medium
CN112487794A
Video tag classification method and system and computer readable storage medium
CN114758283A