Method and system for classifying output uplink transmission data

Through the multi-level feature processing and classification framework, the shortcomings of blockchain data classification methods in adapting to data feature changes and processing mixed data are solved, and accurate security classification and risk warning of blockchain data are realized.

CN120030604APending Publication Date: 2025-05-23YUNNAN POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411849773.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Existing blockchain data classification methods are difficult to adapt to the dynamic changes in data characteristics, especially when processing mixed data, the classification accuracy is insufficient, and the timing and correlation characteristics of the data are ignored.

Method used

A method of output-on-chain transmission data classification is proposed. Through a multi-level feature processing and classification framework, including adaptive data distribution transformation, missing value filling based on K neighbors, dynamic model selection mechanism and multi-model integration voting, a security level classification identifier is generated and risk state changes are tracked in real time.

Benefits of technology

It realizes accurate and secure classification of blockchain data, improves classification accuracy and timeliness of risk warning, maintains the timing characteristics and business correlation of data, and adapts to changes in data characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030604A_ABST
    Figure CN120030604A_ABST
Patent Text Reader

Abstract

The invention discloses an output uplink transmission data classification method and system, and relates to the technical field of data security, and the method comprises the steps: obtaining to-be-classified uplink transmission data, and carrying out the first preprocessing of the uplink transmission data, and obtaining the preprocessing data; constructing a feature vector based on the preprocessed data; the feature vectors comprise text semantic features and data association features; and inputting the feature vector into a pre-trained classification model to obtain a security level classification identifier of the uplink transmission data. According to the method, the time sequence and relevance of the data can be kept, and the classification accuracy is improved; through a dynamic risk assessment mechanism, the adaptability of the system to data feature changes is enhanced; according to the invention, timely identification and early warning of high-risk data are realized, and reliable guarantee is provided for security management of block chain data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data security technology, and in particular to a method and system for classifying output uplink transmission data. Background Art

[0002] As a distributed data storage technology, the data security classification management of blockchain technology is of great significance to the security of the system. At present, the classification of blockchain data mainly adopts rule-based static classification methods and machine learning-based dynamic classification methods. The static classification method classifies data through preset classification rules, which has the advantages of simple implementation and high computational efficiency, but it is difficult to adapt to the dynamic changes of data features. The dynamic classification method based on machine learning can adaptively learn data features, but it has obvious limitations when processing mixed data in blockchain scenarios. Especially when processing mixed scenarios containing text descriptions and structured transaction data, the traditional single classification model is difficult to take into account the feature expression of both types of data at the same time, resulting in insufficient classification accuracy. In addition, the existing classification methods generally ignore the temporal and correlation characteristics of blockchain data, which are of great value for accurately judging the security level of data.

[0003] In terms of data preprocessing, existing technologies usually adopt a unified preprocessing method and fail to perform differentiated processing based on the characteristics of different types of data. This processing method is particularly insufficient when facing blockchain data, because blockchain data often exhibits highly skewed distribution characteristics, and there are complex business associations between different types of data. Traditional preprocessing methods are difficult to effectively maintain these characteristics of the data, thus affecting the accuracy of subsequent classification. At the same time, existing technologies often use simple statistical methods when dealing with missing data and outliers. Although this method is easy to operate, it often destroys the inherent associations between data and is not conducive to maintaining the integrity and consistency of blockchain data. Summary of the invention

[0004] In view of the above-mentioned problems, the present invention is proposed.

[0005] Therefore, the present invention provides a method and system for classifying output uplink transmission data, which can solve the problems mentioned in the background technology.

[0006] To solve the above technical problems, the present invention provides the following technical solutions: a method for classifying output uplink transmission data, comprising: obtaining uplink transmission data to be classified, performing a first preprocessing on the uplink transmission data to obtain preprocessed data;

[0007] Constructing a feature vector based on the preprocessed data; the feature vector includes text semantic features and data association features;

[0008] The feature vector is input into a pre-trained classification model to obtain a security level classification identifier for the uplink transmission data.

[0009] As a preferred solution of the method for classifying output uplink transmission data of the present invention, the method comprises the following steps: obtaining the uplink transmission data to be classified, and performing a first preprocessing on the uplink transmission data to obtain preprocessed data:

[0010] Performing text vectorization processing on the uplink transmission data to obtain a text vector;

[0011] Calculating the cosine similarity values ​​between the text vectors to obtain a similarity matrix;

[0012] Determine the distribution of the number of samples in each category in the similarity matrix:

[0013] If the ratio of the number of samples in a certain category to the number of samples in other categories is less than a first preset threshold, oversampling is performed on the samples in this category;

[0014] If the ratio of the number of samples in a certain category to the number of samples in other categories is greater than a second preset threshold, undersampling is performed on the samples in this category;

[0015] The similarity matrix and the samples that have been sampled are integrated into the preprocessed data.

[0016] As a preferred solution of the method for classifying the output uplink transmission data of the present invention, the text vectorization processing refers to segmenting the uplink transmission data to obtain a word sequence, and converting the word sequence into a vector representation using a pre-trained word embedding model;

[0017] The sampling process refers to randomly duplicating the category samples to generate new samples, or randomly selecting some samples to delete;

[0018] The pre-processed data is integrated in such a way that the similarity matrix is ​​used as a feature matrix, the samples that have been sampled are used as training samples, and sample-feature pairs are constructed.

[0019] As a preferred solution of the method for classifying the output uplink transmission data of the present invention, the feature vector is constructed based on the preprocessed data, including the following steps:

[0020] dividing the preprocessed data into text data and structured data;

[0021] Processing the text data;

[0022] The structured data is processed.

[0023] As a preferred solution of the method for classifying data for uplink transmission according to the present invention, the processing of the text data comprises the following steps:

[0024] If there is a foreign name in the text data, extract the full name and abbreviation that appear for the first time as the semantic related item; otherwise, use the Chinese name as the semantic related item;

[0025] If the text data contains multiple subject words, extract subject features as semantically related items based on word frequency-inverse document frequency; otherwise, take the entire text as a single subject feature;

[0026] The processing of the structured data comprises the following steps:

[0027] If there are business-related fields in the structured data, the business-related fields are constructed as a business-related tree; otherwise, a single-layer business field list is constructed;

[0028] If the nodes in the service association tree have multiple mapping relationships, the multiple mapping relationships are converted into an association rule set; otherwise, a single mapping relationship between the nodes is established.

[0029] As a preferred solution of the method for classifying data for uplink transmission according to the present invention, the processing of the text data comprises the following steps:

[0030] If there is a foreign name in the text data, extract the full name and abbreviation that appear for the first time as the semantic related item; otherwise, use the Chinese name as the semantic related item;

[0031] If the text data contains multiple subject words, extract subject features as semantically related items based on word frequency-inverse document frequency; otherwise, take the entire text as a single subject feature;

[0032] The processing of the structured data comprises the following steps:

[0033] If there are business-related fields in the structured data, the business-related fields are constructed as a business-related tree; otherwise, a single-layer business field list is constructed;

[0034] If the nodes in the service association tree have multiple mapping relationships, the multiple mapping relationships are converted into an association rule set; otherwise, a single mapping relationship between the nodes is established.

[0035] As a preferred solution of the method for classifying the output uplink transmission data of the present invention, the method of classifying and predicting the second preprocessed feature vector comprises the following steps:

[0036] If the proportion of text semantic features in the feature vector is higher than the feature proportion threshold, a long short-term memory network is used for classification prediction;

[0037] If the proportion of data-related features in the feature vector is higher than the feature proportion threshold, a convolutional neural network is used for classification prediction;

[0038] If the confidence of the classification prediction result is lower than the confidence threshold, the integrated voting of random forest, support vector machine and gradient boosting tree is initiated;

[0039] Generating a security level classification identifier according to the classification prediction result comprises the following steps:

[0040] If the safety risk score of the classification prediction result is higher than the first risk threshold, an emergency safety warning mark is generated and a real-time alarm is triggered;

[0041] If the security risk score of the classification prediction result is higher than the second risk threshold and lower than the first risk threshold, a conventional security warning mark is generated;

[0042] If the security risk score of the classification prediction result is lower than the second risk threshold, a normal security mark is generated;

[0043] If the state of the security level classification identifier changes after it is generated, the change trajectory is recorded and the risk assessment model is updated.

[0044] To further solve the above technical problems, the present invention provides the following technical solutions: an output uplink transmission data classification system, comprising:

[0045] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and is characterized in that when the processor executes the computer program, the steps of the above-mentioned method for classifying data for output uplink transmission are implemented.

[0046] A computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the steps of the above-mentioned method for classifying data for output uplink transmission are implemented.

[0047] Beneficial effects of the present invention: The present invention realizes accurate and secure classification of blockchain data through a multi-level feature processing and classification framework. In the feature vector preprocessing stage, the time series characteristics and business relevance of blockchain data are effectively maintained through adaptive data distribution conversion and missing value filling based on K nearest neighbors. In the classification prediction stage, a dynamic model selection mechanism based on feature proportion is adopted, and long short-term memory network and convolutional neural network are used for text semantic features and data association features respectively, which solves the classification problem of mixed feature data. In the security level determination stage, a multi-model integrated voting mechanism and a dynamic risk assessment system are introduced, which not only improves the reliability of the classification results, but also realizes real-time tracking of risk status changes and adaptive updating of models. Compared with the prior art, the present invention has obvious advantages in maintaining data integrity, improving classification accuracy and timeliness of risk warning, and provides more reliable technical support for the security management of blockchain data. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0049] Figure 1 This is a schematic diagram of the overall process of an output uplink transmission data classification method proposed by the present invention;

[0050] Figure 2 A diagram of a computer device in an output uplink transmission data classification method proposed by the present invention. DETAILED DESCRIPTION

[0051] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.

[0052] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0053] Example 1, reference Figure 1, as an embodiment of the present invention, provides a method for classifying output uplink transmission data.

[0054] S1: Acquire uplink transmission data to be classified, and perform first preprocessing on the uplink transmission data to obtain preprocessed data.

[0055] Among them, preprocessing includes data optimization and data standardization.

[0056] S1.1: Perform text vectorization processing on the uplink transmission data to obtain a text vector;

[0057] S1.2: Calculate the cosine similarity value between text vectors to obtain a similarity matrix;

[0058] S1.3: Determine the distribution of the number of samples in each category in the similarity matrix:

[0059] If the ratio of the number of samples in a certain category to the number of samples in other categories is less than a first preset threshold, oversampling is performed on the samples in this category;

[0060] If the ratio of the number of samples in a certain category to the number of samples in other categories is greater than a second preset threshold, undersampling is performed on the samples in this category;

[0061] S1.4: Integrate the similarity matrix and the sampled samples into preprocessed data.

[0062] It should be noted that text vectorization processing refers to segmenting the data transmitted on the chain to obtain word sequences, and using a pre-trained word embedding model to convert the word sequences into vector representations. Sampling processing refers to randomly copying category samples to generate new samples, or randomly selecting some samples for deletion. The integration method of pre-processed data is to use the similarity matrix as the feature matrix, the sampled samples as training samples, and construct sample-feature pairs.

[0063] S2: Construct a feature vector containing text semantic features and data association features based on the preprocessed data.

[0064] S2.1: Divide the preprocessed data into text data and structured data.

[0065] Specifically, text data is used to construct text semantic features, and structured data is used to construct data association features.

[0066] S2.2: Process the text data, including:

[0067] If there are foreign names in the text data, extract the full name and abbreviation that appear for the first time as semantic related items; otherwise, use the Chinese name as the semantic related item;

[0068] If the text data contains multiple topic words, the topic features are extracted as semantic association items based on word frequency-inverse document frequency; otherwise, the entire text is taken as a single topic feature.

[0069] S2.3: Process structured data, including:

[0070] If there are business-related fields in the structured data, the business-related fields are constructed as a business-related tree; otherwise, a single-layer business field list is constructed;

[0071] If the nodes in the business association tree have multiple mapping relationships, the multiple mapping relationships are converted into association rule sets; otherwise, a single mapping relationship between the nodes is established.

[0072] In this embodiment, the process of constructing the feature vector first divides the preprocessed data into text data and structured data. This separation processing method avoids the problem of feature loss caused by uniform processing of different types of data in traditional methods, and also facilitates targeted feature extraction.

[0073] For text data, the present invention focuses on solving the problem of semantic feature extraction in multi-language mixed scenarios. By identifying the first appearance of a foreign name, its full name and abbreviation are extracted as semantically related items, effectively retaining the semantic information of professional terms. For pure Chinese text, the Chinese name is directly extracted as a semantically related item. When the text contains multiple keywords, the word frequency-inverse document frequency method is used to extract topic features. Compared with the traditional word frequency statistics method, this method can more accurately reflect the topic relevance of words. For text with a single topic, the entire text is used as a feature to avoid unnecessary information segmentation.

[0074] In terms of structured data processing, the present invention introduces the concept of a business association tree. When business association fields are detected in the data, these fields are constructed into a hierarchical business association tree. Compared with the traditional planar processing method, this structure can better preserve the hierarchical relationship and business logic between data. For data that does not have obvious business relevance, a single-layer business field list is constructed to ensure the completeness of the processing method. In particular, when the nodes in the business association tree have multiple mapping relationships, the present invention converts these relationships into association rule sets. This method breaks through the limitations of one-to-one mapping in traditional methods and can more completely express complex business relationships.

[0075] Finally, the present invention generates a final feature vector by combining semantically associated items and association rule sets. This combination method realizes the organic integration of text semantic features and data association features, which not only retains the semantic information of the text, but also includes the business association relationship between data, providing more comprehensive feature support for subsequent classification. Compared with the single feature extraction method in the prior art, the feature vector of the present invention has stronger expression ability and better adaptability, and can more accurately reflect the characteristics of the on-chain data.

[0076] Through the above processing, the present invention ensures the feature extraction effect while also ensuring the processing efficiency. Practice shows that this feature construction method can effectively improve the accuracy of subsequent classification, especially when processing blockchain data with complex business associations.

[0077] S3: Input the feature vector into the pre-trained classification model to obtain the security level classification identification of the data transmitted on the chain.

[0078] S3.1: Perform a second preprocessing on the feature vector, including:

[0079] If the numerical distribution of the eigenvector does not satisfy the normal distribution, data distribution transformation is performed; otherwise, standardization is performed;

[0080] If the feature vector contains missing values, similar features are selected based on the K nearest neighbor algorithm to calculate the mean for filling; otherwise, normalization is performed;

[0081] If there are outliers in the feature vector, outlier smoothing is performed based on the box plot rule; otherwise, the original features are kept unchanged;

[0082] S3.2: Classify and predict the feature vector after the second preprocessing, including:

[0083] If the proportion of text semantic features in the feature vector is higher than the feature proportion threshold, the long short-term memory network is used for classification prediction;

[0084] If the proportion of data-related features in the feature vector is higher than the feature proportion threshold, a convolutional neural network is used for classification prediction;

[0085] If the confidence of the classification prediction result is lower than the confidence threshold, the integrated voting of random forest, support vector machine and gradient boosting tree is initiated;

[0086] S3.3: Generate a safety level classification mark based on the classification prediction results, including:

[0087] If the security risk score of the classification prediction result is higher than the first risk threshold, an emergency security warning mark is generated and a real-time alarm is triggered;

[0088] If the security risk score of the classification prediction result is higher than the second risk threshold and lower than the first risk threshold, a conventional security warning mark is generated;

[0089] If the security risk score of the classification prediction result is lower than the second risk threshold, a normal security mark is generated;

[0090] If the status of the security level classification identifier changes after it is generated, the change trajectory is recorded and the risk assessment model is updated.

[0091] It should be noted that in the blockchain data classification scenario, the traditional classification method has three main technical problems: first, insufficient feature vector preprocessing leads to unstable model training effect; second, a single classification model is difficult to adapt to complex and changeable data features; third, the security level division lacks a dynamic adjustment mechanism. In response to these problems, the present invention proposes an adaptive feature preprocessing and classification scheme.

[0092] In the feature vector preprocessing stage, the present invention first focuses on the data distribution characteristics. By detecting the distribution of feature vectors, the data that does not meet the normal distribution is transformed to ensure the rationality of data distribution. For the missing value problem, the present invention introduces the K nearest neighbor algorithm for feature filling. Compared with the traditional average value filling method, this filling method based on similar features can better maintain the correlation of data. At the same time, the box plot rule is used to smooth outliers to avoid the interference of abnormal data on the classification results.

[0093] In the classification prediction link, the present invention breaks through the limitations of the traditional single model and designs a model selection mechanism based on feature proportion. When the text semantic features are dominant, a long short-term memory network that is more suitable for sequence data processing is used; when the data association features are dominant, a convolutional neural network that is good at extracting local features is selected. This adaptive model selection strategy significantly improves the accuracy of classification. In particular, when the confidence of the classification result is insufficient, by integrating multiple traditional machine learning models for voting, this multi-model fusion method can effectively reduce classification bias.

[0094] In the security level identification generation link, the present invention establishes a multi-level risk assessment system. According to the risk score predicted by classification, the data is divided into different security levels, and a real-time alarm mechanism is set for high-risk data. The present invention introduces a security level status change tracking function, which realizes the dynamic optimization of the classification system by recording the change trajectory and updating the risk assessment model in time. This dynamic adjustment mechanism enables the system to continuously adapt to changes in data characteristics.

[0095] Practice shows that the classification scheme of the present invention has obvious advantages in processing complex blockchain data. Compared with traditional methods, the classification accuracy is significantly improved, especially when processing mixed feature data. At the same time, the dynamic risk assessment mechanism effectively reduces the incidence of data security incidents and provides reliable protection for the security management of blockchain data.

[0096] Preferably, in step S3, the feature vector preprocessing of the present invention adopts a hierarchical and progressive processing strategy. First, a data distribution conversion mechanism is introduced, which is particularly important when processing blockchain data because blockchain data often exhibits highly skewed distribution characteristics. Through data distribution conversion, not only does the data become more suitable for subsequent machine learning model processing, but the correlation properties of the original data can also be maintained, which is of great significance for ensuring the integrity of blockchain data.

[0097] In particular, the present invention uses the K nearest neighbor algorithm to fill in the missing features. The reason why this method works well in this scenario is that blockchain data has natural temporal and correlation properties. By finding the K most similar features to calculate the mean, the temporal characteristics and business logic of the data can be better maintained, avoiding the data deviation that may be caused by the traditional mean filling method. This similarity-based filling strategy is highly consistent with the chain structure of blockchain data, providing a more reliable data foundation for subsequent classification.

[0098] In the classification prediction link, the present invention proposes an adaptive model selection mechanism based on feature proportion. The reason why this mechanism can perform well in the blockchain scenario is that blockchain data usually contains structured transaction data and unstructured description text. When the text semantic features dominate, the long short-term memory network can effectively capture the long-term dependencies in the text; when the data association features dominate, the convolutional neural network can extract the key local feature patterns. This dynamic switching strategy solves the problem of diversity of blockchain data features well.

[0099] In the process of generating security level identification, the risk status change tracking mechanism introduced by the present invention is particularly suitable for blockchain scenarios. Since blockchain data is tamper-proof, any security risk may cause irreparable losses. By monitoring risk status changes in real time and updating the assessment model, the system can detect potential security threats in a timely manner. In addition, this dynamic update mechanism can also adapt to the evolution of blockchain business models, ensuring that the classification system always maintains high efficiency.

[0100] Through verification in actual applications, this multi-level risk assessment system not only improves the accuracy of classification, but more importantly, it can provide an early warning mechanism for the security management of blockchain data. Especially when processing high-frequency trading data, the system can quickly identify abnormal patterns and effectively prevent potential security risks.

[0101] In summary, the present invention realizes accurate and secure classification of blockchain data through a multi-level feature processing and classification framework. In the feature vector preprocessing stage, the time series characteristics and business relevance of blockchain data are effectively maintained through adaptive data distribution conversion and missing value filling based on K nearest neighbors. In the classification prediction stage, a dynamic model selection mechanism based on feature proportion is adopted, and long short-term memory network and convolutional neural network are used for text semantic features and data association features respectively, which solves the classification problem of mixed feature data. In the security level determination stage, a multi-model integrated voting mechanism and a dynamic risk assessment system are introduced, which not only improves the reliability of the classification results, but also realizes real-time tracking of risk status changes and adaptive updating of the model. Compared with the prior art, the present invention has obvious advantages in maintaining data integrity, improving classification accuracy and timeliness of risk warning, and provides more reliable technical support for the security management of blockchain data.

[0102] Embodiment 2 is an embodiment of the present invention, which provides an output uplink transmission data classification system, including: a data acquisition module, used to acquire uplink transmission data to be classified;

[0103] A preprocessing module, used to preprocess the uplink transmission data to obtain preprocessed data;

[0104] A feature construction module is used to construct a feature vector containing text semantic features and data association features based on preprocessed data;

[0105] The classification recognition module is used to input the feature vector into a pre-trained classification model to obtain the security level classification identification of the data transmitted on the chain.

[0106] Example 3, reference Figure 2, is an embodiment of the present invention, which is different from the previous embodiment in that: if the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0107] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.

[0108] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.

[0109] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0110] Example 4 is an embodiment of the present invention, which provides an output uplink transmission data classification method. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.

[0111] To verify the effectiveness of the present invention, this embodiment constructs a test set containing multiple types of blockchain data, covering typical scenarios such as smart contract transactions and cross-chain data transmission. The experiment compares the traditional single model method (support vector machine SVM), the static rule classification method and the adaptive classification method proposed in the present invention. The test data contains multi-dimensional information such as text descriptions, transaction records, and time series features, in which different proportions of missing values ​​and outliers are deliberately set to verify the robustness of the scheme in complex scenarios. The experiment focuses on key indicators such as classification accuracy, feature retention, and timeliness of risk warning.

[0112] In the experimental setting, this embodiment uses a stratified sampling method to construct a test data set to ensure the representativeness of the data distribution. For each classification method, this embodiment uses the same evaluation indicators and test environment. In particular, this embodiment designs a feature retention index to measure the similarity between the preprocessed data and the original data, and introduces a risk warning delay time to quantify the responsiveness of the system. During the experiment, this embodiment also conducted comparative tests on different data feature ratio scenarios to verify the advantages of the present invention in processing mixed feature data.

[0113] Table 1 Performance comparison table

[0114]

[0115]

[0116] As shown in Table 1, the present invention is significantly better than the comparative scheme in all key indicators. The classification accuracy rate reaches 91.8%, which is 15.5 percentage points higher than the traditional SVM method and 9.3 percentage points higher than the static rule method. In terms of feature retention, the present invention reaches a high level of 94.2%, which shows that the scheme can better maintain the original characteristics and relevance of the data. In terms of the timeliness of risk warning, the delay time of the present invention is only 180 milliseconds, which is nearly 80% shorter than the traditional method. It is particularly noteworthy that when dealing with outliers and missing values, the accuracy of the present invention reaches 89.6% and 92.1% respectively, far exceeding other schemes. Although the system resource occupancy rate of the present invention is slightly higher than that of the static rule method, this additional overhead is acceptable considering the performance improvement. The model update cycle takes only 5 minutes, which greatly improves the system's adaptability to changes in data features. These data fully demonstrate the significant advantages of the present invention in the field of blockchain data classification.

[0117] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A method for classifying output uplink transmission data, characterized in that: include: Acquire uplink transmission data to be classified, and perform first preprocessing on the uplink transmission data to obtain preprocessed data; constructing a feature vector based on the preprocessed data; The feature vector includes text semantic features and data association features; The feature vector is input into a pre-trained classification model to obtain a security level classification identifier for the uplink transmission data.

2. The method for classifying output uplink transmission data according to claim 1, characterized in that: Acquiring uplink transmission data to be classified, and performing first preprocessing on the uplink transmission data to obtain preprocessed data, comprises the following steps: Performing text vectorization processing on the uplink transmission data to obtain a text vector; Calculating the cosine similarity values ​​between the text vectors to obtain a similarity matrix; Determine the distribution of the number of samples in each category in the similarity matrix: If the ratio of the number of samples in a certain category to the number of samples in other categories is less than a first preset threshold, oversampling is performed on the samples in this category; If the ratio of the number of samples in a certain category to the number of samples in other categories is greater than a second preset threshold, undersampling is performed on the samples in this category; The similarity matrix and the samples that have been subjected to sampling processing are integrated into the preprocessed data.

3. The method for classifying output uplink transmission data according to claim 2, characterized in that: The text vectorization processing refers to segmenting the uplink transmission data to obtain a word sequence, and converting the word sequence into a vector representation using a pre-trained word embedding model; The sampling process refers to randomly duplicating the category samples to generate new samples, or randomly selecting some samples to delete; The pre-processed data is integrated in such a way that the similarity matrix is ​​used as a feature matrix, the samples processed by sampling are used as training samples, and sample-feature pairs are constructed.

4. The method for classifying output uplink transmission data according to claim 3, characterized in that: Constructing a feature vector based on the preprocessed data includes the following steps: dividing the preprocessed data into text data and structured data; Processing the text data; The structured data is processed.

5. The method for classifying output uplink transmission data according to claim 4, characterized in that: The processing of the text data comprises the following steps: If there is a foreign name in the text data, extract the full name and abbreviation that appear for the first time as the semantic related item; otherwise, use the Chinese name as the semantic related item; If the text data contains multiple subject words, extract subject features as semantically related items based on word frequency-inverse document frequency; otherwise, take the entire text as a single subject feature; The processing of the structured data comprises the following steps: If there are business-related fields in the structured data, the business-related fields are constructed as a business-related tree; otherwise, a single-layer business field list is constructed; If the nodes in the business association tree have multiple mapping relationships, converting the multiple mapping relationships into an association rule set; Otherwise, a single mapping relationship between nodes is established.

6. The method for classifying output uplink transmission data according to claim 5, characterized in that: Inputting the feature vector into a pre-trained classification model to obtain a security level classification identifier for the uplink transmission data includes the following steps: The feature vector is subjected to a second preprocessing, specifically, if the numerical distribution of the feature vector does not satisfy the normal distribution, data distribution conversion is performed; otherwise, standardization is performed; if the feature vector contains missing values, similar features are selected based on the K nearest neighbor algorithm to calculate the mean for filling; otherwise, normalization is performed; if the feature vector has an outlier, the outlier is smoothed based on the box plot rule; otherwise, the original feature is kept unchanged; Performing classification prediction on the second preprocessed feature vector; A security level classification identifier is generated according to the classification prediction result.

7. The method for classifying output uplink transmission data according to claim 6, characterized in that: The method of performing classification prediction on the second preprocessed feature vector comprises the following steps: If the proportion of text semantic features in the feature vector is higher than the feature proportion threshold, a long short-term memory network is used for classification prediction; If the proportion of data-related features in the feature vector is higher than the feature proportion threshold, a convolutional neural network is used for classification prediction; If the confidence of the classification prediction result is lower than the confidence threshold, the integrated voting of random forest, support vector machine and gradient boosting tree is initiated; Generating a security level classification identifier according to the classification prediction result comprises the following steps: If the safety risk score of the classification prediction result is higher than the first risk threshold, an emergency safety warning mark is generated and a real-time alarm is triggered; If the security risk score of the classification prediction result is higher than the second risk threshold and lower than the first risk threshold, a conventional security warning mark is generated; If the security risk score of the classification prediction result is lower than the second risk threshold, a normal security mark is generated; If the state of the security level classification identifier changes after it is generated, the change trajectory is recorded and the risk assessment model is updated.

8. An output uplink transmission data classification system, based on the output uplink transmission data classification method according to any one of claims 1 to 7, characterized in that: include, A data acquisition module, used to acquire the uplink transmission data to be classified; A preprocessing module, used for preprocessing the uplink transmission data to obtain preprocessed data; A feature construction module, used to construct a feature vector including text semantic features and data association features based on the preprocessed data; The classification identification module is used to input the feature vector into a pre-trained classification model to obtain a security level classification identifier for the uplink transmission data.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the output uplink transmission data classification method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the output uplink transmission data classification method according to any one of claims 1 to 7 are implemented.