Data real-time processing and trusted distribution system based on AIoT
By using the BERT model to generate semantic vectors in the AIoT system, combining it with the XGBoost algorithm to build a trust assessment model, and dynamically selecting the transmission path, the problem of integrating structured and unstructured data in AIoT data processing is solved, and data analysis capabilities and security are improved.
Patent Information
- Application Number
- CN202510782252.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-12
AI Technical Summary
Existing AIoT data processing methods are unable to effectively integrate structured and unstructured data, resulting in low information utilization, and the data transmission mechanism lacks flexibility and cannot transmit key data in a timely and effective manner.
The data acquisition module is used to collect and preprocess data, the BERT model is used to generate semantic vectors, and the semantic structure hybrid fusion algorithm is combined to identify threat behaviors through the edge reasoning module. The trust assessment model is built using the XGBoost algorithm, and the transmission path is dynamically selected for data distribution based on the trust level.
It achieves a deep understanding and effective use of structured and unstructured data, improves the depth and breadth of data analysis, and enhances the instant response capability and security of edge computing nodes.
Smart Images

Figure CN120639798A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a real-time data processing and trusted distribution system based on AIoT. Background Art
[0002] With the rapid development of the Internet of Things (IoT), various types of smart devices such as temperature sensors, humidity sensors, and anemometers have been widely deployed in various industries to collect and transmit large amounts of data. In recent years, the integration of artificial intelligence and the Internet of Things (AIoT) has become increasingly deeper, driving the transformation of smart devices from simple data collection to intelligent data analysis and processing. By applying advanced machine learning algorithms to the massive data generated by the IoT, researchers can gain a deeper understanding of the changing laws of the physical world and make more accurate predictions. Therefore, real-time data processing methods based on AIoT have shown great application potential and are expected to play an important role in the field of environmental monitoring.
[0003] Despite this, there is still room for improvement in the existing methods of processing AIoT data. Due to the huge differences in the characteristics of structured data and unstructured data, it is usually difficult to effectively integrate the two, resulting in low information utilization and limited processing capabilities, which limits the possibility of extracting valuable information from unstructured data. At the same time, current data transmission mechanisms generally lack flexibility and cannot dynamically adjust transmission strategies according to the security and importance of the data, which may result in key data not being delivered in a timely and effective manner due to network congestion and other reasons. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides an AIoT-based real-time data processing and trusted distribution system to solve the problem of large differences in characteristics between structured data and unstructured data.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: The present invention provides a real-time data processing and trusted distribution system based on AIoT, which includes: Data acquisition module, which collects structured and unstructured data from IoT devices in real time and performs pre-processing; The semantic vector module uses the BERT model to understand natural language based on unstructured data and generate semantic vectors; The fusion module uses a semantic-structure hybrid fusion algorithm based on semantic vectors and structured data to obtain a fused data stream; The edge inference module deploys the TinyLSA model on the edge computing node to receive the fused data stream, perform real-time threat behavior identification, and obtain security data packets; The trust level module uses the XGBoost algorithm to build a trust assessment model and trains it using historical IoT device behavior logs. It then inputs the trained trust assessment model into the security data packet and analyzes its trust level. The data transmission module dynamically selects the transmission path for data distribution based on the trust level of the security data packet.
[0007] As a preferred solution of the AIoT-based data real-time processing and trusted distribution system of the present invention, wherein: the IoT device includes a temperature sensor, a humidity sensor, an anemometer and a text report generator; The structured data includes temperature value, humidity value and wind speed; The unstructured data includes IoT device log text.
[0008] As a preferred solution of the AIoT-based real-time data processing and trusted distribution system described in the present invention, the preprocessing includes using an adaptive filter to denoise structured data and removing punctuation and HTML tags from unstructured data.
[0009] As a preferred solution of the AIoT-based data real-time processing and trusted distribution system of the present invention, wherein: based on unstructured data, the BERT model is used for natural language understanding to generate semantic vectors, specifically, Use a word segmentation tool to segment the IoT device log text in unstructured data into independent vocabulary units. Then use a tokenizer to map each vocabulary unit to an ID to form an input sequence. Add the [CLS] tag and the [SEP] tag to the beginning and end of the input sequence, respectively. Collect historical IoT device log text and add a fully connected classification layer and a Softmax function to the output layer of the BERT model; Input historical IoT device log text into the BERT model to predict the probability distribution value; Calculate the cross entropy loss between the true probability distribution value of the historical IoT device log text and the probability distribution value predicted by the BERT model to obtain the cross entropy loss value; Based on the cross entropy loss value, the chain rule is used to derive the gradient value of the BERT model parameters layer by layer; Based on the gradient values of the BERT model parameters, the optimizer Adam is used to adjust the BERT model parameters to generate a trained BERT model; Import the input sequence into the trained BERT model to obtain the semantic vector.
[0010] As a preferred solution of the AIoT-based data real-time processing and trusted distribution system described in the present invention, wherein: based on semantic vectors and structured data, a semantic structure hybrid fusion algorithm is adopted to obtain a fused data stream, specifically, Use feature concatenation and matrix organization methods to construct a two-dimensional matrix of size d×n, and input the semantic vector and structured data into the two-dimensional matrix to generate a fused tensor matrix; According to the adaptive weighting mechanism, a weight matrix with the same size as the fusion tensor matrix is constructed, and a weight mapping table is designed for the weight matrix; Perform absolute value addition calculation on the semantic vector and structured data of each matrix position in the fused tensor matrix to obtain the sum of the absolute values of each matrix position in the fused tensor matrix; Map the sum of the absolute values of each matrix position in the fused tensor matrix to the weight mapping table, obtain the weight of each matrix position, and fill the weight of each matrix position in the fused tensor matrix into the weight matrix; Perform element-by-element multiplication on the fused tensor matrix and the weight matrix to obtain the fused data stream; As a preferred solution of the AIoT-based data real-time processing and trusted distribution system of the present invention, the TinyLSA model is deployed on the edge computing node to receive the fused data stream and perform real-time threat behavior identification, specifically, Obtain the fused data stream of historical IoT devices as the training set and set the learning rate, input batch size, and number of training rounds in the TinyLSA model; Input the training set into the set TinyLSA model for training; Use ONNX to import the trained TinyLSA model into the edge computing node, and then import the fused data stream into the TinyLSA model for threat identification to generate a threat score A for the fused data stream. As a preferred solution of the AIoT-based data real-time processing and trusted distribution system of the present invention, wherein: and obtaining a security data packet, specifically, Based on the threat score A of the fused data stream, define the threat behavior threshold A1; When A≥A1, the threat level of the output fused data stream is high; When A<A1, the threat level of the output fused data stream is low; The threat level and threat score A of the fused data stream are combined into a safe data packet.
[0011] As a preferred solution of the AIoT-based data real-time processing and trusted distribution system described in the present invention, the trust evaluation model is constructed using the XGBoost algorithm and trained through the behavior logs of historical IoT devices, specifically, Combine the log text semantic vectors, structured data, and trust labels generated during the historical operation of temperature sensors, humidity sensors, and anemometers obtained from IoT devices into a training set. Then, set the learning rate, number of trees, and maximum tree depth in the XGBoost algorithm. The training set is input into the XGBoost algorithm for training to generate a trust assessment model.
[0012] As a preferred solution of the AIoT-based data real-time processing and trusted distribution system of the present invention, wherein: the security data packet is input into the trained trust evaluation model to analyze the trust level of the security data packet, specifically, Deploy the trust assessment model in the local inference engine of the edge computing device to receive the security data packet; Input the threat level and threat score A in the security data package into the trust assessment model to obtain a trust score value; Build trust level classification standards based on business needs, map trust score values to the trust level classification standards, and obtain the trust level of the security data package.
[0013] As a preferred solution of the AIoT-based data real-time processing and trusted distribution system of the present invention, wherein: the transmission path is dynamically selected for data distribution according to the trust level of the security data packet, specifically, Use the QUIC protocol to set QoS priority and establish a high-speed path between the edge computing node and the central server, and use the TLS1.3 encryption protocol to enable forward secrecy and establish an encrypted path; The trust level of the security data packet is matched with the high-speed path and encryption path, and the edge computing device distributes the data of the security data packet.
[0014] The beneficial effects of the present invention are: by utilizing advanced natural language processing methods to convert unstructured data into semantic vectors, a deep understanding and effective utilization of complex information is achieved, and structured and unstructured data are innovatively combined to generate a comprehensive data representation, which greatly enhances the depth and breadth of data analysis. The real-time threat behavior identification mechanism on the edge computing node further enhances the immediate response capability and security. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0016] Figure 1Flowchart for real-time data processing and trusted distribution based on AIoT.
[0017] Figure 2 Schematic diagram of the AIoT-based real-time data processing and trusted distribution system.
[0018] Figure 3 Flowchart for generating semantic vectors.
[0019] Figure 4 Flowchart for generating fused data stream. DETAILED DESCRIPTION
[0020] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0021] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0022] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0023] Reference Figures 1 to 4 , is an embodiment of the present invention, which provides a real-time data processing and trusted distribution system based on AIoT, including the following steps: The data acquisition module collects structured and unstructured data from IoT devices in real time and performs preprocessing.
[0024] In the target monitoring area, such as a weather station, four types of IoT devices are prepared: a temperature sensor for collecting real-time air temperature data; a humidity sensor for measuring water vapor content in the air; an anemometer for detecting air velocity; and a text report generator for recording device operating status, abnormal events, and operation logs. The temperature sensor, humidity sensor, anemometer, and text report generator all have unique identifiers (such as device numbers) and support wireless communication protocols.
[0025] Install the temperature sensor in a well-ventilated location away from direct sunlight. Install the humidity sensor in an area with good air circulation, away from water sources. Install the anemometer in an open area, ensuring the anemometer blades can rotate freely. The text report generator should be integrated into the edge computing node. Connect the temperature sensor, humidity sensor, anemometer, and text report generator to a power source, and request the current time from a standard time server over the network. This automatically calibrates the local clock and ensures that the timestamps of the temperature sensor, humidity sensor, anemometer, and text report generator are consistent.
[0026] The temperature sensor, humidity sensor, and anemometer are set to automatically report data once every minute. The reported content includes the device ID, timestamp, and value. The value units are degrees Celsius (°C), percentage relative humidity (%RH), and meters per second (m / s), respectively.
[0027] The text report generator is configured for event-triggered reporting. This means that a log is automatically generated and sent when the temperature sensor, humidity sensor, or anemometer status changes, exhibits abnormal behavior, or receives external commands. The temperature sensor, humidity sensor, anemometer, and text report generator are wirelessly connected to the edge computing node, which serves as the central receiving end and is responsible for subsequent data processing.
[0028] Edge nodes continuously monitor structured data from temperature sensors, humidity sensors, and anemometers. They check timestamps for validity, such as whether they are older than the device deployment date. They also check whether values exceed physical boundaries. For example, temperatures below -273.15°C (absolute zero) are considered invalid, and humidity above 100% is considered abnormal. If clearly unreasonable structured data is found, it is marked as suspicious. Suspicious samples are not immediately discarded but instead enter a dedicated queue for manual review. All other structured data, except for suspicious samples, is considered qualified.
[0029] Create a fixed-size sliding window buffer with a capacity of 10 data samples. Update the sliding window using the first-in, first-out principle. Add qualified structured data to the sliding window sequentially. If the sliding window is full, remove the oldest qualified structured data and add the latest qualified structured data to the end. Record the timestamps and values of all qualified structured data in the sliding window, and regularly calculate the mean and standard deviation within the sliding window for subsequent filtering and judgment.
[0030] Obtain the latest piece of qualified structured data. Calculate the current mean and standard deviation based on the existing qualified structured data in the sliding window, and determine whether the mean of the latest piece of qualified structured data exceeds the standard deviation (e.g., ±1.5 times the standard deviation). If the latest piece of qualified structured data exceeds ±1.5 times the standard deviation, replace the latest piece of qualified structured data with the mean of the existing qualified structured data in the sliding window. Delete the earliest piece of qualified structured data in the sliding window and add the corrected latest piece of qualified structured data to the sliding window to obtain the denoised structured data.
[0031] The edge computing node listens to the unstructured data from the text report generator, namely the temperature sensor, humidity sensor and anemometer log text. The pattern matching method is used to identify and remove all the content in angle brackets in the log text, such as <log> 、< / log> 、 , and delete all Chinese and English punctuation marks, retain numbers, letters and spaces, merge multiple consecutive spaces, delete line breaks and tabs, and uniformly use single spaces to separate words. Finally, the log text maintains clear semantics and concise structure, thereby generating log text in a standardized format.
[0032] The semantic vector module uses a model to understand natural language based on unstructured data and generate semantic vectors.
[0033] We obtained historical IoT device log text and the preprocessed IoT device log text. We used Google's open-source tokenizer to split the text into multiple independent words based on Chinese semantics. The tokenizer then converted each word into a corresponding numeric ID based on its internal vocabulary, forming an input sequence. The [CLS] and [SEP] tags were added to the beginning and end of the input sequence, respectively. The input sequence of historical IoT device log text was used to train the BERT model, and the preprocessed IoT device log text was used to extract semantic vectors.
[0034] A fully connected layer (DenseLayer) is added to the output of the last layer of the BERT model. This fully connected layer maps the hidden state vector corresponding to the [CLS] tag to the task-related feature space. A Softmax function is added after the fully connected layer to normalize the output value into a probability distribution to facilitate the calculation of the loss function. The output is in the form of two numerical values, representing the predicted probabilities of "normal" and "abnormal," respectively.
[0035] The input sequence of historical IoT device log text is fed into the currently trained BERT model. The BERT model predicts the probability of the input sequence, for example, a normal probability of 0.92 and an abnormal probability of 0.08. The cross-entropy loss function is used to measure the difference between the BERT model's predicted probability and the actual probability of the historical IoT device log text, resulting in a cross-entropy loss value. This cross-entropy loss value reflects the BERT model's current performance; a smaller cross-entropy loss value indicates a more accurate BERT model prediction. Based on this cross-entropy loss value, the error is propagated back through each layer, starting from the BERT model's output layer, to calculate the gradient of the neural network parameters in each layer of the BERT model. The Adam optimizer is used to update the BERT model parameters based on the gradient values. The learning rate is set to 2e-5 to ensure stable training convergence. After multiple rounds of iterations, the BERT model is trained. The fully connected layer and softmax function are removed from the trained BERT model.
[0036] The preprocessed IoT device log text input sequence is input into the trained BERT model. The BERT model outputs the hidden state vector at the [CLS] tag position. This hidden state vector is a real number vector of fixed dimension (usually 768) that contains the overall semantic information of the log text and is therefore called a "semantic vector."
[0037] The fusion module, based on semantic vectors and structured data, adopts a semantic-structure hybrid fusion algorithm to obtain a fused data stream.
[0038] Create an empty two-dimensional matrix with rows equal to the feature dimension d (for example, 768) and columns equal to the number of samples n (for example, 10 data points from the past 10 minutes). For each time point, extract the corresponding semantic vector (768 dimensions) and the corresponding structured data (assuming it is also 768 dimensions). Stack the semantic vector and structured data by column to form a 768×10 column vector. Concatenate the column vectors for all time points into the two-dimensional matrix to form the final fused tensor matrix. The fused tensor matrix size is d×n.
[0039] Set the data intensity range, for example, [0~100): low intensity; [100~500): medium intensity; [500~1000): high intensity; [1000~∞): very high intensity.
[0040] A corresponding weight value is assigned to each data intensity interval, and a weight mapping table is established to indicate the importance of semantic vectors and structured data in the fusion under this intensity, for example, [0~100): weight 0.1; [100~500): weight 0.3; [500~1000): weight 0.6; [1000~∞): weight 0.9.
[0041] Traverse all semantic vectors and structured data in the fused tensor matrix and take the absolute values of all semantic vectors and structured data. Add the absolute values of the semantic vectors and structured data at the same position in the fused tensor matrix to generate a new matrix, called the "intensity matrix." For the semantic vectors and structured data in the intensity matrix, find the intensity interval to which they belong, and find the corresponding weight value according to the weight mapping table. Create a blank two-dimensional matrix of the same size as the fused tensor matrix, fill the found weight values into the corresponding blank two-dimensional matrix, and generate a weight matrix of the same size as the fused tensor matrix. The weight value of each position reflects the relative importance of the data at that position in the fusion process.
[0042] The fused tensor matrix and the weight matrix are element-wise multiplied. Specifically, because the fused tensor matrix and the weight matrix have the same size and dimensions, the semantic vectors and structured vectors at the same position in the fused tensor matrix and the corresponding weights are multiplied together to generate the fused data flow matrix. The fused data flow matrix still has a size of d × n. The fused data flow matrix reflects the combined influence of semantic information and structured data. The fused data flow matrix is concatenated column by column into a long one-dimensional vector. For example, a 768 × 10 fused data flow matrix is flattened into a vector of length 7680, which is the fused data flow.
[0043] The edge inference module deploys the TinyLSA model on the edge computing node to receive the fused data stream, perform real-time threat behavior identification, and obtain security data packets.
[0044] Select a sufficient number of fused data stream samples from the historical database (recommended to be no less than 5,000). Each fused data stream corresponds to a behavior record at a point in time, that is, whether an attack behavior has occurred (binary classification), and finally form a training set.
[0045] Load the pretrained TinyLSA model and set the following hyperparameters (all optimal): learning rate (learning_rate): 0.001, batch size (batch_size): 16, and number of training epochs (epochs): 20. Select binary cross entropy as the loss function for the TinyLSA model, as the task is binary classification (threat or not). Select Adam as the optimizer.
[0046] Start the training loop for each epoch, iterate through each batch in the train_loader, and each batch contains multiple fused data stream samples (e.g., 16). Feed the current batch of training set into the TinyLSA model. The TinyLSA model outputs a predicted floating-point probability value (indicating whether it is a threat behavior). At the same time, the meta-cross-entropy loss function of the TinyLSA model calculates the loss value based on the predicted floating-point probability value. When the loss value does not decrease for 3 consecutive rounds after multiple rounds of training, terminate the training and save the current best TinyLSA model.
[0047] Use a tool to convert the trained TinyLSA model into the ONNX format, that is, the ONNX model. Ensure that the input and output interfaces of the ONNX model are consistent with the edge-side inference program, and test whether the ONNX model runs normally locally. After the ONNX model is correct, upload it to the edge computing node, install the ONNXRuntime environment (such as onnxruntime-linux-x64), write an inference service script, load the ONNX model, configure the input and output interfaces, ensure that it can receive the fused data stream, and start the inference service after configuration to wait for new data.
[0048] The ONNX model receives the fused data stream and performs forward inference to obtain the output result: a value between 0 and 1. Apply the Sigmoid function to this value between 0 and 1 to obtain the threat score A. A close to 1 indicates a high threat, and A close to 0 indicates a low risk. Express the threat score A with a formula. Specifically, ; Among them, represents the threat score, represents the base of the natural logarithm, represents the probability value between 0 and 1 output by the ONNX model; Based on the threat score A, set the threat behavior threshold A1 (e.g., 0.6). The threat behavior threshold is used to determine whether the current IoT device behavior constitutes a threat, and the threat behavior threshold should be dynamically adjusted according to actual business needs. The threat behavior threshold A1 can be updated through an artificial review feedback mechanism.
[0049] Compare the threat score A with the threshold A1. Specifically, when A ≥ A1, the threat level = "high"; when A < A1, the threat level = "low". Finally, obtain two fields, namely the threat score A and the threat level. Package the source of the fused data stream (device ID, timestamp), the threat score A, and the threat level into a structured "security data packet".
[0050] The trust level module uses the XGBoost algorithm to build a trust assessment model and trains it through historical IoT device behavior logs. It inputs the trained trust assessment model into the security data packet to analyze the trust level of the security data packet.
[0051] Collect the semantic vectors, structured data, and corresponding ground truth labels of the log text generated by the temperature, humidity, and anemometer sensors in IoT devices during their historical operation. Combine the semantic vectors and corresponding structured data from the log text of the temperature, humidity, and anemometer sensors into an input feature vector, and add the corresponding trust labels (trustworthy / suspicious) as the target output of the XGBoost algorithm to form a complete training set.
[0052] The XGBoost algorithm was selected as the core algorithm for the trust assessment model, and a set of parameter combinations was set. For example, the learning rate was set to 0.1, the number of decision trees was set to 100, and the maximum depth of a single tree was set to 5. An early stopping mechanism was set during the XGBoost algorithm training process to prevent overfitting, and the output was in the form of probability values to facilitate the subsequent generation of trust scores.
[0053] The training set is fed into the XGBoost algorithm, which outputs a probability value between 0 and 1, indicating the probability that the training set belongs to a certain category (for example, "trusted"). The true label is a known, labeled data label that tells the XGBoost algorithm what the correct answer is. The true label is usually a binary value (0 and 1), where 1 indicates that the sample belongs to the "trusted" category and 0 indicates that the sample belongs to the "untrusted" category. For the known true labels (0 and 1) of the training set and the probability values output by the XGBoost algorithm, a loss function is used to calculate the loss value, specifically, ; in, represents the loss value, represents the true label of the training set, Represents the probability value predicted by the XGBoost algorithm, represents the logarithmic function; Based on the loss value, the XGBoost algorithm will automatically use the gradient boosting method to optimize the structure of the next tree according to the loss value. With each additional tree, the prediction ability of the XGBoost algorithm is enhanced, and parameter adjustment is automatically completed without manual intervention. After training, the validation set is input into the XGBoost algorithm to generate validation set probability values. The loss value between the validation set probability values and the true labels is calculated using the above loss function, and the accuracy and precision of the XGBoost algorithm are calculated. Observe the trend of loss changes on the validation set to determine whether the XGBoost algorithm has converged. When it is found that the loss value on the validation set has not improved significantly for three consecutive rounds, it indicates that the XGBoost algorithm has converged, thus obtaining a trust evaluation model.
[0054] Input the security data packet into the deployed trust evaluation model to obtain a trust score. The trust score is a numerical value in the range of [0, 1] generated after non-linear transformation and Sigmoid normalization of the weights and biases learned by the security data packet within the trust evaluation model, and is used to quantitatively evaluate the behavioral trustworthiness of IoT devices. It is expressed by the formula ; where, represents the trust score, represents the Sigmoid function, [[ID=Install and start the QUIC client unit on the edge computing node, configure the QUIC connection parameters, specify the IP address and listening port of the central server, set the DSCP mark value to EF (Expedited Forwarding) in the QUIC protocol stack, indicating that the traffic has the highest priority, configure the network scheduler on the central server side, identify the EF-marked QUIC data stream, and allocate a dedicated processing thread.
[0057] Install and start the TLS 1.3 client unit on the edge computing nodes, configure the TLS handshake process, use the ECDHE key exchange algorithm (such as x25519) to enable forward secrecy, set up the TLS client to connect to the dedicated port of the central server, configure firewall rules and load balancer on the dedicated port of the central server, and direct TLS traffic to the dedicated decryption and verification unit.
[0058] A detailed path mapping policy document is developed, specifying that high-trust security data packets follow the high-speed path, while medium / low-trust security data packets follow the encrypted path. This path mapping policy is written to the local configuration file of the edge computing node and loaded when the edge computing node service starts. Before each security data packet is sent, the edge computing node searches for the corresponding path exit based on the current trust level of the security packet and sends the data. By building a trust-level-driven dual-path communication mechanism, intelligent distribution of security data packets is achieved between edge computing nodes and the central server. This not only improves the transmission efficiency of trusted data but also enhances the security of suspicious data, forming a new closed-loop, controllable, and dynamically adaptable edge communication architecture.
[0059] In summary, the present invention achieves a deep understanding and effective use of complex information by: utilizing advanced natural language processing methods to convert unstructured data into semantic vectors, innovatively combining structured and unstructured data to generate a comprehensive data representation, greatly improving the depth and breadth of data analysis. The real-time threat behavior identification mechanism on edge computing nodes further enhances immediate response capabilities and security.
[0060] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A real-time data processing and trusted distribution system based on AIoT, characterized by: include, Data acquisition module, which collects structured and unstructured data from IoT devices in real time and performs pre-processing; The semantic vector module uses the BERT model to understand natural language based on unstructured data and generate semantic vectors; The fusion module uses a semantic-structure hybrid fusion algorithm based on semantic vectors and structured data to obtain a fused data stream; The edge inference module deploys the TinyLSA model on the edge computing node to receive the fused data stream, perform real-time threat behavior identification, and obtain security data packets; The trust level module uses the XGBoost algorithm to build a trust assessment model and trains it using historical IoT device behavior logs. It then inputs the trained trust assessment model into the security data packet and analyzes its trust level. The data transmission module dynamically selects the transmission path for data distribution based on the trust level of the security data packet.
2. The AIoT-based real-time data processing and trusted distribution system according to claim 1, characterized in that: The IoT devices include a temperature sensor, a humidity sensor, an anemometer, and a text report generator; The structured data includes temperature value, humidity value and wind speed; The unstructured data includes IoT device log text.
3. The AIoT-based real-time data processing and trusted distribution system according to claim 2, characterized in that: The preprocessing includes using an adaptive filter to remove noise from structured data and removing punctuation marks and HTML tags from unstructured data.
4. The AIoT-based real-time data processing and trusted distribution system according to claim 3, characterized in that: Based on unstructured data, the BERT model is used to understand natural language and generate semantic vectors, specifically, Use a word segmentation tool to segment the IoT device log text in unstructured data into independent vocabulary units. Use a tokenizer to map each vocabulary unit to an ID to form an input sequence. Add the [CLS] tag and the [SEP] tag to the beginning and end of the input sequence, respectively. Collect historical IoT device log text and add a fully connected classification layer and a Softmax function to the output layer of the BERT model; Input historical IoT device log text into the BERT model to predict the probability distribution value; Calculate the cross entropy loss between the true probability distribution value of the historical IoT device log text and the probability distribution value predicted by the BERT model to obtain the cross entropy loss value; Based on the cross entropy loss value, the chain rule is used to derive the gradient value of the BERT model parameters layer by layer; Based on the gradient values of the BERT model parameters, the optimizer Adam is used to adjust the BERT model parameters to generate a trained BERT model; Import the input sequence into the trained BERT model to obtain the semantic vector.
5. The AIoT-based real-time data processing and trusted distribution system according to claim 4, characterized in that: The method adopts a semantic structure hybrid fusion algorithm based on semantic vectors and structured data to obtain a fused data stream, specifically, Use feature concatenation and matrix organization methods to construct a two-dimensional matrix of size d×n, and input the semantic vector and structured data into the two-dimensional matrix to generate a fused tensor matrix; According to the adaptive weighting mechanism, a weight matrix with the same size as the fusion tensor matrix is constructed, and a weight mapping table is designed for the weight matrix; Perform absolute value addition calculation on the semantic vector and structured data of each matrix position in the fused tensor matrix to obtain the sum of the absolute values of each matrix position in the fused tensor matrix; Map the sum of the absolute values of each matrix position in the fused tensor matrix to the weight mapping table, obtain the weight of each matrix position, and fill the weight of each matrix position in the fused tensor matrix into the weight matrix; Perform element-by-element multiplication on the fused tensor matrix and the weight matrix to obtain the fused data stream.
6. The AIoT-based real-time data processing and trusted distribution system according to claim 5, characterized in that: The TinyLSA model is deployed on the edge computing node to receive the fused data stream and perform real-time threat behavior identification, specifically, Obtain the fused data stream of historical IoT devices as the training set and set the learning rate, input batch size, and number of training rounds in the TinyLSA model; Input the training set into the set TinyLSA model for training; Use ONNX to import the trained TinyLSA model into the edge computing node, and import the fused data stream into the TinyLSA model for threat identification to generate a threat score A for the fused data stream.
7. The AIoT-based real-time data processing and trusted distribution system according to claim 6, characterized in that: And obtain the security data package, specifically, Based on the threat score A of the fused data stream, define the threat behavior threshold A1; When A≥A1, the threat level of the output fused data stream is high; When A<A1, the threat level of the output fused data stream is low; The threat level and threat score A of the fused data stream are combined into a safe data packet.
8. The AIoT-based real-time data processing and trusted distribution system according to claim 7, characterized in that: The trust evaluation model is constructed using the XGBoost algorithm and trained through the behavior logs of historical IoT devices. Specifically, Combine the log text semantic vectors, structured data, and trust labels generated during the historical operation of temperature sensors, humidity sensors, and anemometers obtained from IoT devices into a training set. Then, set the learning rate, number of trees, and maximum tree depth in the XGBoost algorithm. The training set is input into the XGBoost algorithm for training to generate a trust assessment model.
9. The AIoT-based real-time data processing and trusted distribution system according to claim 8, characterized in that: The security data package is input into the trained trust assessment model to analyze the trust level of the security data package, specifically, Deploy the trust assessment model in the local inference engine of the edge computing device to receive the security data packet; Input the threat level and threat score A in the security data package into the trust assessment model to obtain a trust score value; Build trust level classification standards based on business needs, map trust score values to the trust level classification standards, and obtain the trust level of the security data package.
10. The AIoT-based real-time data processing and trusted distribution system according to claim 9, characterized in that: The method of dynamically selecting a transmission path for data distribution based on the trust level of the security data packet is as follows: Use the QUIC protocol to set QoS priority and establish a high-speed path between the edge computing node and the central server, and use the TLS1.3 encryption protocol to enable forward secrecy and establish an encrypted path; The trust level of the security data packet is matched with the high-speed path and encryption path, and the edge computing device distributes the data of the security data packet.