Internet of Things data mapping method and device, electronic equipment and storage medium
Through the methods of feature extraction and weight adjustment, the automatic mapping of multimodal IoT data and object model attributes is achieved, solving the problems of data heterogeneity and semantic differences in IoT scenarios, and improving the efficiency and accuracy of data processing.
Patent Information
- Application Number
- CN202510350556.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-08
AI Technical Summary
The heterogeneity and semantic differences in multimodal data in IoT scenarios lead to inefficiency of traditional data processing methods, making it difficult to achieve automated processing and accurate mapping, and increases development costs and access difficulties.
By obtaining multimodal IoT data, performing feature extraction and feature confidence calculation, adjusting weights for fusion, calculating the similarity between the fusion features and the preset object model template, and achieving automatic mapping of multimodal data and object model attributes.
Improve the accuracy and efficiency of multimodal IoT data mapping, and reduce processing and access costs.
Smart Images

Figure CN120277427A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular to an Internet of Things (IoT) data mapping method, device, electronic device, and storage medium. Background Art
[0002] With the rapid development of IoT technology, a large amount of multi-modal data has been generated. These data cover various forms such as text, images, audio, and video, and have significant heterogeneity, complexity, and diversity. The data generated by IoT devices usually comes from different sensors and terminal devices, and there are significant differences in their data formats, structures, and semantics, making it difficult for traditional data processing methods to effectively handle. Especially in large-scale IoT scenarios, the heterogeneity of data makes unified processing and analysis particularly difficult. Traditional manual feature engineering methods are not only inefficient but also difficult to meet the requirements of real-time and large-scale processing.
[0003] In addition, due to the large differences in data formats and semantic expressions between different modalities, the semantic alignment and fusion of multi-modal data are difficult, which easily affects the classification accuracy and information integrity. Related methods often rely on manual intervention or complex rule design, making it difficult to achieve automated processing and prone to introducing errors when processing large-scale data.
[0004] More importantly, there is a lack of an automated object model generation and mapping mechanism in IoT scenarios. The object model is a bridge between IoT devices and upper-layer applications, and its generation and mapping process usually requires a large amount of manual participation, increasing the development cost and docking difficulty. Obviously, there is an urgent need for a new IoT data mapping method to solve at least one of the above problems.
[0005] It should be noted that the above content only provides background technical information related to this application and does not necessarily constitute prior art. Summary of the Invention
[0006] In view of the above-mentioned disadvantages of the prior art, this application provides an IoT data mapping method, device, electronic device, and storage medium to automatically map multi-modal IoT data to corresponding object model attributes, improving the accuracy and efficiency of multi-modal IoT data mapping.
[0007] Other features and advantages of this application will become apparent through the following detailed description, or be learned in part through the practice of this application.
[0008] According to one aspect of the embodiments of the present application, an Internet of Things data mapping method is provided, including: obtaining multimodal Internet of Things data, where the multimodal Internet of Things data includes one or more of text data, image data, and sensor data; performing feature extraction on the multimodal Internet of Things data to obtain initial features corresponding to each modality of Internet of Things data; calculating the feature confidence of each of the initial features respectively, adjusting the weights of each of the initial features based on the feature confidence, and fusing each of the initial features according to the adjusted weights to obtain a fused feature; calculating the target similarity between the fused feature and the object model attributes in a preset object model template, and determining the object model attributes corresponding to the fused feature based on the target similarity, so as to complete the mapping of the multimodal Internet of Things data and the object model attributes.
[0009] In an embodiment of the present application, based on the foregoing solution, after obtaining the multimodal Internet of Things data, the method further includes: if the multimodal Internet of Things data includes text data, performing entity extraction on the text data and performing semantic enhancement on the extracted entities; if the multimodal Internet of Things data includes image data, adjusting the resolution of the image data based on the device computing power; if the multimodal Internet of Things data includes sensor data, calibrating the sensor data and using a sliding window filter to eliminate the instantaneous noise of the sensor data, so as to complete the abnormal smoothing process of the sensor data.
[0010] In an embodiment of the present application, based on the foregoing solution, performing feature extraction on the multimodal Internet of Things data to obtain initial features corresponding to each modality of Internet of Things data includes: if the multimodal Internet of Things data includes text data, converting the text data into a text feature vector through a preset language representation model, and fusing a preset domain entity embedding with the text feature vector to obtain a text initial feature corresponding to the text data, where the text initial feature includes text spatio-temporal features; if the multimodal Internet of Things data includes image data, performing feature extraction on the image data through a preset image recognition model to obtain an image initial feature corresponding to the image data, where the image initial feature includes image spatio-temporal features; if the multimodal Internet of Things data includes sensor data, performing feature extraction on the sensor data through a preset data processing model to obtain a sensor initial feature corresponding to the sensor data, where the sensor initial feature includes sensor spatio-temporal features.
[0011] In one embodiment of the present application, based on the foregoing solution, after extracting features from the multi-modal Internet of Things data to obtain the initial features corresponding to each modality of Internet of Things data, the method further includes: spatially and temporally aligning the text initial features, the image initial features, and the sensor initial features based on the text spatio-temporal features, the image spatio-temporal features, and the sensor spatio-temporal features, so as to process the initial features corresponding to each modality of Internet of Things data in the same spatio-temporal dimension.
[0012] In one embodiment of the present application, based on the foregoing solution, the feature confidence of each of the initial features is calculated respectively, including: if the multi-modal Internet of Things data includes text data, calculating the semantic similarity of multiple text segments in the text data, and using the semantic similarity as the text confidence of the text initial features; if the multi-modal Internet of Things data includes image data, calculating the image similarity between multiple image frames in the image data, and using the image similarity as the image confidence of the image initial features; if the multi-modal Internet of Things data includes sensor data, calculating the frequency-domain signal-to-noise ratio of the sensor data, and using the frequency-domain signal-to-noise ratio as the sensor confidence of the sensor initial features.
[0013] In one embodiment of the present application, based on the foregoing solution, the weights of each of the initial features are adjusted based on the feature confidence, and the initial features are fused according to the adjusted weights to obtain a fused feature, including: adjusting the weights of each of the initial features based on the feature confidence, where the feature confidence is proportional to the weights of each of the initial features; calculating the feature similarity of each of the initial features through a cross-attention mechanism, and adjusting the weights of each of the initial features based on the feature similarity; fusing each of the initial features according to the adjusted weights to obtain the fused feature.
[0014] In one embodiment of the present application, based on the foregoing solution, the preset physical model template is obtained in the following manner: obtaining a physical model instance, and determining a guiding text based on the target task and data format; inputting the physical model instance and the guiding text into a preset large model to generate a physical model template.
[0015] According to one aspect of the embodiments of the present application, there is provided an Internet of Things (IoT) data mapping device, including: a data acquisition module configured to acquire multi-modal IoT data, where the multi-modal IoT data includes one or more of text data, image data, and sensor data; a feature extraction module configured to perform feature extraction on the multi-modal IoT data to obtain initial features corresponding to each modal IoT data; a feature fusion module configured to calculate the feature confidence levels of the respective initial features, adjust the weights of the respective initial features based on the feature confidence levels, and fuse the respective initial features according to the adjusted weights to obtain a fused feature; and a data mapping module configured to calculate the target similarity between the fused feature and the object model attributes in a preset object model template, and determine the object model attributes corresponding to the fused feature based on the target similarity, so as to complete the mapping of the multi-modal IoT data and the object model attributes.
[0016] According to one aspect of the embodiments of the present application, there is provided an electronic device, where the electronic device includes: one or more processors; and a storage device configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device is caused to implement the IoT data mapping method as described in any one of the above embodiments.
[0017] The present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor of a computer, the computer is caused to execute the IoT data mapping method as described in any one of the above embodiments.
[0018] Advantages of the present application: By acquiring multi-modal IoT data, which includes one or more of text data, image data, and sensor data, performing feature extraction on the multi-modal IoT data to obtain initial features corresponding to each modal IoT data, calculating the feature confidence levels of the respective initial features, adjusting the weights of the respective initial features based on the feature confidence levels, fusing the respective initial features according to the adjusted weights to obtain a fused feature, calculating the target similarity between the fused feature and the object model attributes in a preset object model template, and determining the object model attributes corresponding to the fused feature based on the target similarity, the present application completes the mapping of the multi-modal IoT data and the object model attributes, realizes the automatic mapping of the multi-modal IoT data to the corresponding object model attributes, improves the accuracy and efficiency of the multi-modal IoT data mapping, and further reduces the processing and access costs of IoT data.
[0019] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:
[0021] Figure 1 is a schematic diagram of an exemplary system architecture shown in an exemplary embodiment of the present application;
[0022] Figure 2 is a schematic flowchart of an Internet of Things data mapping method shown in an exemplary embodiment of the present application;
[0023] Figure 3 is a schematic flowchart of an Internet of Things data mapping method shown in another exemplary embodiment of the present application;
[0024] Figure 4 is a block diagram of an Internet of Things data mapping system shown in an exemplary embodiment of the present application;
[0025] Figure 5 is a block diagram of an Internet of Things data mapping device shown in an exemplary embodiment of the present application;
[0026] Figure 6 shows a schematic diagram of the structure of a computer system of an electronic device suitable for implementing the embodiments of the present application. Detailed Embodiments
[0027] The following will describe the embodiments of the present application with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be understood that the preferred embodiments are only for explaining the present application, rather than for limiting the protection scope of the present application.
[0028] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present application schematically. Therefore, only the components related to the present application are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0029] In the following description, numerous details are explored to provide a more thorough explanation of the embodiments of the present application. However, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present application difficult to understand.
[0030] First of all, it should be noted that the object model is an abstract representation of objects in the Internet of Things, including object attributes, behaviors, events, etc. Specifically, the object model is the digital representation of entities in the physical space (such as sensors, gateways, buildings, factories, etc.) in the cloud. It describes what the entity is, what it can do, and what information it can provide externally from three dimensions: attributes, services, and events. The definitions of these three dimensions together complete the definition of the product function.
[0031] BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model based on the Transformer architecture. BERT uses a bidirectional Transformer encoder, which can consider the left and right context information of each word in a sentence simultaneously, breaking through the limitations of traditional unidirectional language models.
[0032] ViT (Vision Transformer) is a neural network model based on the Transformer architecture, mainly used to process computer vision tasks. ViT addresses the limitations of traditional convolutional neural networks (CNNs) in dealing with long-range dependencies by introducing the attention mechanism of Transformer. ViT divides an image into a series of image patches, each of which is flattened into a vector and combined with a position encoding vector to form an input sequence. These input sequences are processed by the encoder of Transformer, and the attention mechanism is used to model the relationships between the image patches. Through multiple layers of self-attention mechanisms, ViT can model and understand the global context of the image.
[0033] LoRA (Low-Rank Adaptation) is an efficient model fine-tuning technique that optimizes a small number of parameters of large pre-trained models (such as LLaMA-2) by introducing low-rank matrices to adapt to specific tasks or datasets.
[0034] Figure 1 It is a schematic diagram of an exemplary system architecture shown in an exemplary embodiment of the present application.
[0035] Refer to Figure 1As shown, the system architecture may include a data acquisition device 101 and a computer device 102. Among them, the computer device 102 may be at least one of a desktop graphics processing unit (GPU) computer, a GPU computing cluster, a neural network computer, etc. The data acquisition device 101 is used to acquire multi-modal Internet of Things data, and the multi-modal Internet of Things data includes one or more of multi-modal data such as text data, image data, and sensor data. In this embodiment, after the data acquisition device 101 obtains the above data, it provides them to the computer device 102 for processing. Relevant technicians can use the computer device 102 to extract features from the multi-modal Internet of Things data, obtain the initial features corresponding to each modal Internet of Things data, calculate the feature confidence of each initial feature respectively, adjust the weights of each initial feature based on the feature confidence, fuse each initial feature according to the adjusted weights to obtain a fused feature, calculate the target similarity between the fused feature and the object model attributes in the preset object model template, and determine the object model attributes corresponding to the fused feature based on the target similarity to complete the mapping of the multi-modal Internet of Things data and the object model attributes. It should be noted that the data acquisition device 101 and the computer device 102 provided in this embodiment are only examples and should not impose any limitations on the functions and usage scopes of the embodiments of the present application.
[0036] It should be noted that the Internet of Things data mapping method provided in the embodiments of the present application is generally executed by the computer device 102. Correspondingly, the Internet of Things data mapping device is generally set in the computer device 102.
[0037] Figure 2 It is a schematic flowchart of the Internet of Things data mapping method shown in an exemplary embodiment of the present application. The Internet of Things data mapping method can be executed by a computing processing device, and the computing processing device can be Figure 1 the computer device 102 shown in Figure 2 As shown, the Internet of Things data mapping method at least includes steps S210 to S240, which are introduced in detail as follows:
[0038] In step S210, multi-modal Internet of Things data is acquired.
[0039] In one embodiment of the present application, the multi-modal Internet of Things data includes, but is not limited to, one or more of text data, image data, and sensor data. The Internet of Things devices are accessed in the form of a network interface to obtain the text data in the Internet of Things devices; the image video stream data acquired by the image acquisition device is received in real time through a streaming media server, and key frames are extracted and compression-encoded from the video stream to reduce redundant data for easy storage and transmission; the sensor data of the Internet of Things devices is dynamically collected through an edge gateway, and the multi-modal data streams are aligned based on timestamps. Among them, the sensor data can be time-series data collected by multiple different types of sensors in the Internet of Things. It can be understood that the video stream can also be regarded as image data after key frame extraction.
[0040] In one embodiment of the present application, the process after obtaining the multi-modal Internet of Things data further includes: if the multi-modal Internet of Things data includes text data, entity extraction is performed on the text data, and semantic enhancement is performed on the extracted entities; if the multi-modal Internet of Things data includes image data, the resolution of the image data is adjusted based on the obtained device computing power; if the multi-modal Internet of Things data includes sensor data, the sensor data is calibrated, and a sliding window filter is used to eliminate the instantaneous noise of the sensor data to complete the abnormal smoothing process of the sensor data.
[0041] In this embodiment, scenario-based preprocessing is performed on the collected text data, image data, and time-series sensor data through an industrial Internet of Things platform. Specifically, it includes: performing entity extraction on the text data, and combining the extracted entities with the knowledge graph in the Internet of Things field for semantic enhancement to improve the understandability and analysis value of the data; performing adaptive downsampling on the image data, obtaining the device computing power of the Internet of Things device, and dynamically adjusting the resolution of the image data based on the device computing power to balance data quality and processing efficiency; calibrating and performing abnormal smoothing processing on the sensor data, specifically using a sliding window filter to eliminate instantaneous noise. By calibrating the sensor data, noise and outliers are eliminated to ensure the accuracy and stability of the data.
[0042] In step S220, feature extraction is performed on the multi-modal Internet of Things data to obtain the initial features corresponding to each modal Internet of Things data.
[0043] In an embodiment of the present application, if the multimodal Internet of Things data includes text data, the text data is converted into a text feature vector through a preset language representation model, and the preset domain entity embedding is fused with the text feature vector to obtain the initial text feature corresponding to the text data. The initial text feature includes text spatio-temporal features; if the multimodal Internet of Things data includes image data, the image data is subjected to feature extraction through a preset image recognition model to obtain the initial image feature corresponding to the image data. The initial image feature includes image spatio-temporal features; if the multimodal Internet of Things data includes sensor data, the sensor data is subjected to feature extraction through a preset data processing model to obtain the initial sensor feature corresponding to the sensor data. The initial sensor feature includes sensor spatio-temporal features.
[0044] In this embodiment, a preset language representation model (such as the BERT pre-training model) is used for context encoding to generate high-dimensional text feature vectors, and a domain entity embedding layer is introduced to strengthen the representation of Internet of Things terms, so as to extract the initial text features of the text data. Recurrent neural networks and long short-term memory networks can also be used to extract time series and initial text features by capturing the context information and long-term dependencies of the input sequence. A convolutional neural network is used to extract the initial image features through convolutional operations. The initial image feature information includes size, grayscale, pixels, etc. Through a preset data processing model, methods such as time series analysis and spatial relationship analysis are used to extract the initial sensor features including spatio-temporal features, where the preset data processing model is trained through sample sensor data and sample sensor features.
[0045] In this embodiment, the spatio-temporal feature extraction includes the following methods: extracting target features from the initial text features to obtain text spatio-temporal features, where the target features are features related to time and space; extracting spatio-motion features through a lightweight Vision Transformer and optimizing the calculation efficiency using 3D sparse convolution to extract features from image and video data to obtain image spatio-temporal features; using a temporal convolutional network to extract multi-scale periodic features and fusing self-attention mechanisms to capture long-range dependencies to extract features from sensor data to obtain sensor spatio-temporal features.
[0046] In an embodiment of the present application, after the process of extracting features from the multimodal Internet of Things data to obtain the initial features corresponding to each modality of Internet of Things data, the process further includes: spatially aligning and temporally aligning the initial text features, initial image features, and initial sensor features based on the text spatio-temporal features, image spatio-temporal features, and sensor spatio-temporal features, so as to process the initial features corresponding to each modality of Internet of Things data in the same spatio-temporal dimension.
[0047] In this embodiment, a learnable affine transformation matrix is used to project heterogeneous features, mapping the initial features from different modalities to a common embedding space to complete the spatial alignment of the initial features. The initial features include, but are not limited to, text spatio-temporal features, image spatio-temporal features, and sensor spatio-temporal features. A timestamp mapping mechanism is constructed for the sampling differences of text, images, videos, and sensor data. By using synchronized marker points, the timestamps of data in different modalities are aligned, and the initial features of each modality are aligned in the time dimension to complete the time alignment of the initial features.
[0048] In step S230, the feature confidence of each initial feature is calculated respectively, the weight of each initial feature is adjusted based on the feature confidence, and each initial feature is fused according to the adjusted weight to obtain a fused feature.
[0049] In an embodiment of the present application, if the multi-modal IoT data includes text data, the semantic similarity of multiple text segments in the text data is calculated, and the semantic similarity is used as the text confidence of the text initial feature. If the multi-modal IoT data includes image data, the image similarity between multiple image frames in the image data is calculated, and the image similarity is used as the image confidence of the image initial feature. If the multi-modal IoT data includes sensor data, the frequency-domain signal-to-noise ratio of the sensor data is calculated, and the frequency-domain signal-to-noise ratio is used as the sensor confidence of the sensor initial feature.
[0050] In this embodiment, if it is the text modality, that is, the multi-modal IoT data includes text data, a semantic consistency score is constructed based on the [CLS] token of BERT, and the calculation result is used as the text confidence of the text initial feature. The calculation method can refer to Equation (1):
[0051]
[0052] First, the pre-trained BERT model is used to process the input text to obtain the hidden state vector corresponding to the [CLS] token. This vector captures the context semantic information of the entire text. Then, it is input into a multi-layer perceptron MLP to enhance the discrimination ability of IoT semantics. Finally, the output of the MLP is processed by the sigmoid activation function σ to obtain a consistency score between 0 and 1.
[0053] If it is the visual modality, that is, the multi-modal IoT data includes image data, the SSIM structural similarity index is calculated, and the calculation result is used as the image confidence of the image initial feature. The calculation method can refer to Equation (2):
[0054]
[0055] For each pair of consecutive frames in the video, i.e., image data, the similarity score between them is calculated using the Structural Similarity Index (SSIM) algorithm. The SSIM scores corresponding to all adjacent frames are summed to obtain the cumulative similarity score of the entire video sequence. Finally, the cumulative score is divided by the total number of frames to obtain an average SSIM score, i.e., C. vision and used as the image confidence of the initial image features.
[0056] If it is a sensing modality, i.e., the multi-modal Internet of Things data includes sensor data, the sensor confidence of the initial sensor features is evaluated using the signal-to-noise ratio in the frequency domain. The calculation method can refer to Equation (3):
[0057]
[0058] Separate the signal and noise energy in the frequency domain, where represents the signal energy, represents the noise energy, and the ratio is converted to decibel units through logarithmic operations to evaluate the quality of sensor data.
[0059] In an embodiment of the present application, the weights of each initial feature are adjusted based on the feature confidence, where the feature confidence is proportional to the weights of each initial feature; the feature similarity of each initial feature is calculated through a cross-attention mechanism, and the weights of each initial feature are adjusted based on the feature similarity; the initial features are fused according to the adjusted weights to obtain the fused features.
[0060] In this embodiment, an Adaptive Gated Recurrent Unit is used to optimize the weight allocation, and different weights are assigned according to the high and low feature confidences, referring to Equation (4).
[0061]
[0062] The Gated Recurrent Unit (GRU) is used to process the confidence score C at the current time step t t and the historical hidden state of a specific modality as input to generate a comprehensive state representation g t , providing a basis for subsequent weight calculation.
[0063]
[0064] represents the dot product of the learnable parameter vector W m related to the m-th modality and the comprehensive state representation g t as the score of this modality. By performing a Softmax transformation on the scores of all modalities, i.e., calculating according to the normalized exponential function, the importance weight α of each modality is obtained. (m), which helps to dynamically adjust the influence degree of different modalities during the fusion process.
[0065] The cross-attention mechanism is used to capture the correlation between different modalities, and the feature correlation is evaluated through Equation (6).
[0066] where Q m = h m W Q , K n = h n W k Equation (6)
[0067] In Equation (6), is used to calculate the similarity score matrix between modality m and modality n, where is the dot product of the similarity between modality m and modality n, and d is a scaling factor to avoid excessive numerical values; then a Softmax transformation is performed, which reflects the attention degree of modality m to the features of modality n, and the value vector V n is weighted and summed, and finally the new feature representation h ′ m .
[0068] According to the dynamically adjusted weights, the features of each modality are weighted and summed to form a unified representation, referring to Equation (7):
[0069]
[0070] where α (m) is the importance weight of each modality, and h ′ m is the new feature representation generated by the cross-attention mechanism. According to the weight α (m) the new feature representations h ′ m of all modalities are weighted and summed after applying layer normalization to obtain the final aggregated feature representation z, that is, the fusion feature.
[0071] In step S240, the target similarity between the fusion feature and the physical model attributes in the preset physical model template is calculated, and the physical model attributes corresponding to the fusion feature are determined based on the target similarity to complete the mapping of the multi-modal Internet of Things data and the physical model attributes.
[0072] In an embodiment of the present application, the physical model attribute with the highest target similarity is determined as the physical model attribute corresponding to the fusion feature, and the fusion feature is mapped to the physical model attribute. Since the fusion feature is obtained from the multi-modal Internet of Things data, a mapping relationship between the multi-modal Internet of Things data and the physical model attributes can be established to complete the mapping of the multi-modal Internet of Things data and the physical model attributes.
[0073] In this embodiment, the cosine similarity is used to calculate the semantic similarity between the fusion feature and the physical model template attribute, i.e., the physical model attribute. The calculation method refers to Equation (8):
[0074] S ij =cos(Z i ,ATTR j )·TF-IDF(t k ) Equation (8)
[0075] where Z i is the fusion feature, ATTR j is the physical model template attribute, and cos(Z i ,ATTR j ) is used to measure the similarity between the fusion feature Z i and the physical model template attribute ATTR j . The similarity is weighted and adjusted according to the weight adjustment factor of the term frequency-inverse document frequency TF-IDF(t k ).
[0076] Furthermore, the Hungarian algorithm is used to optimize the optimal matching to identify the most matching attribute mapping relationship. Refer to Equation (9):
[0077] Argmin∑ i,j (1-S ij )·X ij , X ij ∈{0,1}
[0078] Equation (9)
[0079] where S ij represents the similarity. It can also be understood that 1-S ij represents the dissimilarity. That is, when S ij is larger, 1-S ij is smaller, which means a better match. The ultimate goal is to find an allocation scheme to minimize the overall dissimilarity.
[0080] In an embodiment of the present application, the accuracy of the physical model mapping is monitored, and the mapping rule is adjusted based on anomaly detection. An online learning mechanism is implemented to allow the system to continuously collect feedback information during the actual operation process and continuously optimize the performance of the large model in the physical model generation and mapping tasks. The key metrics include mapping accuracy, F1 value, etc.
[0081] In an embodiment of the present application, the preset physical model template is obtained in the following manner: obtaining a physical model instance, and determining the guiding text based on the target task and data format; inputting the physical model instance and the guiding text into a preset large model to generate a physical model template.
[0082] In this embodiment, the generation of the physical model template includes two steps: Step 1, construction of the domain knowledge base: Collect the Internet of Things industry standards, transmission protocols, and public physical model libraries, and combine the fusion characteristics of multimodal data to construct a domain knowledge graph. Then, map the knowledge graph to a high-dimensional semantic space through a knowledge embedding method to provide prior knowledge for the large model. Step 2, generation of the physical model template: Use a preset large model (such as the LLaMA-2 model fine-tuned by LoRA) to parse the domain knowledge base and extract key attributes, methods, and events. Further, guide the large model to generate a physical model template that conforms to industry standards through prompt engineering, including information such as device identification, physical attributes, and data formats.
[0083] In one embodiment of the present application, the model for processing text data can be at least one of models such as BERT and DeepSeek-R1; the model for processing image data can be at least one of models such as ResNet and Vision Transformer; the model for processing multimodal data can be at least one of models such as CLIP and BLIP. BERT is a bidirectional encoder representation model based on the Transformer architecture. It is pre-trained through masked language model (MLM) and next sentence prediction (NSP) tasks and can capture context information in the text. DeepSeek-R1 is an efficient text pre-training model in the DeepSeek series. It may adopt advanced Transformer architectures and optimization techniques to improve the training efficiency and performance of the model. ResNet is a deep convolutional neural network that solves the problem of gradient disappearance in deep neural networks by introducing residual blocks, enabling the network to be constructed deeper. Vision Transformer (ViT) is an image pre-training model based on the Transformer architecture. It divides the image into multiple small blocks and processes these small blocks as the input sequence of the Transformer. CLIP (Contrastive Language-Image Pretraining) is a multimodal pre-training model that maps images and text to the same feature space through contrastive learning, enabling the model to understand the semantic relationship between images and text. BLIP (Bootstrapped Language-Image Pre-training) is another multimodal pre-training model that may adopt more advanced pre-training strategies and optimization techniques to improve the performance of the model in multimodal tasks. It can be understood that the above models are only for illustration, and the present application does not limit this, nor should it bring any limitations to the functions and usage scopes of the embodiments of the present application.
[0084] Figure 3 It is a schematic flow chart of the Internet of Things data mapping method shown in another exemplary embodiment of the present application. Referring to Figure 3 as shown, this Internet of Things data mapping method at least includes steps S1 to S4, which are introduced in detail as follows:
[0085] In step S1, in the multi-modal data acquisition and scene-based preprocessing stage, multi-modal data from different Internet of Things devices is collected, and scene-based data preprocessing operations such as time series alignment, anomaly smoothing, and semantic enhancement are performed;
[0086] In step S2, in the spatio-temporal feature joint extraction and alignment stage, the spatio-temporal dependence of multi-modal Internet of Things data is captured and aligned in the spatio-temporal dimension;
[0087] In step S3, in the dynamic weight cross-modal fusion stage, a differentiable feature weight allocation module is constructed, the fusion weight is dynamically adjusted based on feature confidence, and cross-modal feature integration is performed to form a unified representation;
[0088] In step S4, in the physical model generation and attribute mapping stage, a large model is trained to output a physical model template, and the similarity is calculated to realize the automatic mapping of Internet of Things data and physical model attributes.
[0089] It can be understood that the specific ways of performing the operations in steps S1 to S4 in this embodiment have been described in detail in the foregoing embodiments, and will not be elaborated here.
[0090] The present application first collects one or more modalities of Internet of Things data including text, images, and time series sensor data as raw data through an interface, and preprocesses the raw data in combination with the Internet of Things scenario; initial feature extraction is performed on the preprocessed multi-modal Internet of Things data to obtain initial features of each modality, and the initial features of each modality are data-aligned in the spatio-temporal dimension; through a confidence evaluation and dynamic weight allocation mechanism, the weights of the initial features of each modality are adjusted, and a non-linear and linear collaborative fusion strategy is used to generate fusion features; a multi-dimensional similarity metric model between the fusion features and the predefined physical model attributes is established, and an optimal matching algorithm is used to achieve the precise mapping of multi-modal Internet of Things data and physical model attributes. Through multi-modal data acquisition, cross-modal feature fusion, automatic physical model generation, and automatic mapping of Internet of Things data and physical model attributes, efficient and accurate automatic mapping from the underlying data to the upper-layer physical model is achieved, significantly reducing the Internet of Things data processing and access costs.
[0091] Figure 4 It is a block diagram of the Internet of Things data mapping system shown in an exemplary embodiment of the present application. Referring to Figure 4As shown in the figure, the Internet of Things data mapping system includes an edge collaborative acquisition module, a spatio-temporal feature extraction and alignment module, a differentiable feature weight allocation module, a cross-modal feature fusion module, an object model generation module, an object model attribute mapping gateway, and a result verification module. Specifically, the edge collaborative acquisition module integrates lightweight containers for collecting multi-modal Internet of Things data and supports data alignment and compression on edge devices; the spatio-temporal feature extraction and alignment module is used to extract spatio-temporal data features of different modalities and align each modality of data in the spatio-temporal dimension; the differentiable feature weight allocation module includes a confidence evaluation module and a dynamic weight adjustment module to dynamically adjust the fusion weights based on feature confidence; the cross-modal feature fusion module realizes feature correlation evaluation and weighted fusion of each modality of features to form a unified representation of features; the object model generation module outputs an object model template through a large model and supports object model version management and differential update; the object model attribute mapping gateway is built with a rule engine, a protocol converter, and a semantic verification unit to realize automatic mapping of data and attributes; the result verification module is used to evaluate the classification performance of the model.
[0092] It should be noted that the Internet of Things data mapping system provided in the above embodiment and the Internet of Things data mapping method provided in the above embodiment belong to the same concept. The specific ways in which each module and unit perform operations have been described in detail in the method embodiment and will not be repeated here. In actual application, the Internet of Things data mapping system provided in the above embodiment can, according to needs, allocate the above functions to different functional modules, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above. This is not limited here either.
[0093] Figure 5 is a block diagram of an Internet of Things data mapping device shown in an exemplary embodiment of the present application. This device can be applied to Figure 1 the implementation environment shown in the figure and is specifically configured in the computer device 102. This device can also be applicable to other exemplary implementation environments and is specifically configured in other devices. This embodiment does not limit the implementation environment applicable to this device.
[0094] As Figure 5 shown, this exemplary Internet of Things data mapping device includes: a data acquisition module 510, a feature extraction module 520, a feature fusion module 530, and a data mapping module 540.
[0095] Among them, the data acquisition module 510 is used to obtain multi-modal Internet of Things data, and the multi-modal Internet of Things data includes one or more of text data, image data, and sensor data; the feature extraction module 520 is used to extract features from the multi-modal Internet of Things data to obtain initial features corresponding to each modality of Internet of Things data; the feature fusion module 530 is used to calculate the feature confidence of each initial feature respectively, adjust the weights of each initial feature based on the feature confidence, and fuse each initial feature according to the adjusted weights to obtain a fused feature; the data mapping module 540 is used to calculate the target similarity between the fused feature and the object model attributes in the preset object model template, and determine the object model attributes corresponding to the fused feature based on the target similarity, so as to complete the mapping of the multi-modal Internet of Things data and the object model attributes.
[0096] It should be noted that the Internet of Things data mapping device provided in the above embodiment and the Internet of Things data mapping method provided in the above embodiment belong to the same concept. The specific ways in which each module and unit perform operations have been described in detail in the method embodiment, and will not be repeated here. In practical applications, the Internet of Things data mapping device provided in the above embodiment can, according to needs, allocate the above functions to different functional modules to complete, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above. This is not limited here either.
[0097] An embodiment of the present application further provides an electronic device, including: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device realizes the Internet of Things data mapping method provided in each of the above embodiments.
[0098] Figure 6 The structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown. It should be noted that Figure 6 The computer system 600 of the electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0099] As Figure 6As shown, computer system 600 includes a Central Processing Unit (CPU) 601, which can perform various appropriate actions and processes according to programs stored in Read-Only Memory (ROM) 602 or programs loaded from storage section 608 into Random Access Memory (RAM) 603, such as executing the methods provided in the above various embodiments. In RAM 603, various programs and data required for system operation are also stored. CPU 601, ROM 602, and RAM 603 are connected to each other via bus 604. Input / Output (I / O) interface 605 is also connected to bus 604.
[0100] The following components are connected to I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including, for example, a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), etc. and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to I / O interface 605 as required. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on drive 610 as required so that a computer program read from it can be installed into storage section 608 as required.
[0101] Specifically, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via communication section 609, and / or installed from removable medium 611. When the computer program is executed by the Central Processing Unit (CPU) 601, various functions defined in the system of the present application are executed.
[0102] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0103] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0104] The units involved in the embodiments described in this application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not constitute a limitation on the units themselves in some cases.
[0105] On the other hand, this application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor of a computer, the computer is caused to execute the Internet of Things data mapping method provided in each of the above embodiments. The computer-readable storage medium may be included in the electronic device described in the above embodiments, or may exist separately without being assembled into the electronic device.
[0106] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0107] On the other hand, this application also provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the Internet of Things data mapping method provided in each of the above embodiments.
[0108] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described here can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of this application.
[0109] After considering the specification and practicing the disclosed embodiments here, those skilled in the art will easily think of other implementation manners of this application. This application is intended to cover any variations, uses, or adaptive changes of this application, which follow the general principles of this application and include common general knowledge or conventional technical means in the technical field not disclosed in this application.
[0110] The above embodiments are only illustrative of the principles and effects of the present application, and are not intended to limit the present application. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed in the present application should still be covered by the claims of the present application.
Claims
1. An Internet of Things data mapping method, characterized in that Including: Obtain multi-modal Internet of Things data, where the multi-modal Internet of Things data includes one or more of text data, image data, and sensor data; Extract features from the multi-modal Internet of Things data to obtain initial features corresponding to each modality of Internet of Things data; Calculate the feature confidence degrees of the respective initial features, adjust the weights of the respective initial features based on the feature confidence degrees, and fuse the respective initial features according to the adjusted weights to obtain a fused feature; Calculate the target similarity between the fused feature and the physical model attributes in a preset physical model template, and determine the physical model attributes corresponding to the fused feature based on the target similarity, so as to complete the mapping of the multi-modal Internet of Things data and the physical model attributes.
2. The Internet of Things data mapping method according to claim 1, wherein After obtaining the multi-modal Internet of Things data, the method further includes: If the multi-modal Internet of Things data includes text data, perform entity extraction on the text data and perform semantic enhancement on the extracted entities; If the multi-modal Internet of Things data includes image data, adjust the resolution of the image data based on the device computing power; If the multi-modal Internet of Things data includes sensor data, calibrate the sensor data and use a sliding window filter to eliminate the instantaneous noise of the sensor data, so as to complete the anomaly smoothing process of the sensor data.
3. The Internet of Things data mapping method according to claim 1, characterized in that Extracting features from the multi-modal Internet of Things data to obtain initial features corresponding to each modality of Internet of Things data includes: If the multi-modal Internet of Things data includes text data, convert the text data into a text feature vector through a preset language representation model, and fuse a preset domain entity embedding with the text feature vector to obtain a text initial feature corresponding to the text data, where the text initial feature includes text spatio-temporal features; If the multi-modal Internet of Things data includes image data, extract features from the image data through a preset image recognition model to obtain an image initial feature corresponding to the image data, where the image initial feature includes image spatio-temporal features; If the multi-modal Internet of Things data includes sensor data, extract features from the sensor data through a preset data processing model to obtain a sensor initial feature corresponding to the sensor data, where the sensor initial feature includes sensor spatio-temporal features.
4. The Internet of Things data mapping method according to claim 3, wherein After extracting features from the multi-modal Internet of Things data to obtain initial features corresponding to each modality of Internet of Things data, the method further includes: Perform spatial alignment and temporal alignment on the text initial feature, the image initial feature, and the sensor initial feature based on the text spatio-temporal features, the image spatio-temporal features, and the sensor spatio-temporal features, so as to process the initial features corresponding to each modality of Internet of Things data in the same spatio-temporal dimension.
5. The Internet of Things data mapping method according to any one of claims 1 to 4, characterized in that Calculating the feature confidence degrees of the respective initial features includes: If the multi-modal Internet of Things data includes text data, calculate the semantic similarity of multiple text segments in the text data, and use the semantic similarity as the text confidence degree of the text initial feature; If the multimodal Internet of Things data includes image data, calculate the image similarity between multiple image frames in the image data, and use the image similarity as the image confidence of the initial image features; If the multimodal Internet of Things data includes sensor data, calculate the frequency-domain signal-to-noise ratio of the sensor data, and use the frequency-domain signal-to-noise ratio as the sensor confidence of the initial sensor features.
6. The Internet of Things data mapping method according to any one of claims 1 to 4, characterized in that Based on the feature confidence, adjust the weights of the initial features, and fuse the initial features according to the adjusted weights to obtain fused features, including: Based on the feature confidence, adjust the weights of the initial features, where the feature confidence is proportional to the weights of the initial features; Calculate the feature similarity of the initial features through a cross-attention mechanism, and adjust the weights of the initial features based on the feature similarity; Fuse the initial features according to the adjusted weights to obtain the fused features.
7. The Internet of Things data mapping method according to any one of claims 1 to 4, characterized in that, Obtain the preset physical model template through the following method: Obtain a physical model instance, and determine the guiding text based on the target task and data format; Input the physical model instance and the guiding text into a preset large model to generate a physical model template.
8. An Internet of Things data mapping device, characterized in that, Including: A data acquisition module for obtaining multimodal Internet of Things data, where the multimodal Internet of Things data includes one or more of text data, image data, and sensor data; A feature extraction module for extracting features from the multimodal Internet of Things data to obtain initial features corresponding to each modality of Internet of Things data; A feature fusion module for calculating the feature confidence of each initial feature respectively, adjusting the weights of the initial features based on the feature confidence, and fusing the initial features according to the adjusted weights to obtain fused features; A data mapping module for calculating the target similarity between the fused features and the physical model attributes in the preset physical model template, and determining the physical model attributes corresponding to the fused features based on the target similarity to complete the mapping of the multimodal Internet of Things data and the physical model attributes.
9. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device for storing one or more programs, which when executed by the one or more processors, cause the electronic device to implement the Internet of Things data mapping method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, which when executed by a processor of a computer, causes the computer to execute the Internet of Things data mapping method according to any one of claims 1 to 7.