Cross-domain Internet of Things equipment intelligent collaboration method and system based on semantic knowledge graph
By converting mechanical vibration energy into electrical energy signals, adjusting the sampling frequency and performing cross-modal feature alignment, and constructing a semantic knowledge graph, the problems of asynchronous data fragmentation and multi-modal feature semantic mismatch in cross-domain IoT devices in high-speed vibration scenarios are solved, and efficient monitoring and collaborative control of device status are achieved.
Patent Information
- Application Number
- CN202511240106.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-09-02
AI Technical Summary
In the existing technology, cross-domain IoT devices have problems with motion artifacts and spectral phase shifts caused by time differences in the multimodal collaborative perception of high-speed rotating and vibrating devices, as well as weak semantic correlation of multimodal features, which affects the interpretability of device status and diagnostic accuracy.
By converting mechanical vibration energy into electrical energy signals, synchronously generating level values, adjusting the sampling frequency to obtain image texture and sound spectrum data, using feature-level fusion classifiers to perform cross-modal feature space alignment, constructing a semantic knowledge graph, and dynamically optimizing feature weights to generate collaborative control instructions.
It achieves data synchronization and feature alignment of cross-domain IoT devices in high-speed vibration scenarios, improves the interpretability of device status and diagnostic accuracy, and ensures closed-loop response of device status monitoring and real-time adaptation of collaborative control.
Smart Images

Figure CN120750992A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of electronic technology, and in particular to a method and system for intelligent collaboration of cross-domain Internet of Things devices based on semantic knowledge graphs. Background Art
[0002] In the field of industrial equipment status monitoring, multimodal collaborative perception is needed for cross-domain IoT devices with high-speed rotation and vibration to overcome the time-varying characteristics of mechanical motion, avoid oversampling or feature omission caused by fixed sampling frequency, and support explainable diagnosis of equipment degradation status.
[0003] The existing solution uses adaptive sampling and multimodal decision fusion based on vibration sensors. The specific process is: the device vibration signal is monitored in real time through a piezoelectric sensor. When the vibration intensity exceeds the preset vibration intensity threshold, the camera and microphone are dynamically triggered for sampling; then, the image texture features and sound spectrum data features are extracted respectively, and finally the device status classification results are generated through decision-level fusion.
[0004] However, the existing solutions have two key flaws. First, the visual and acoustic sensors are independently triggered by vibration signals, and the hardware response delay leads to millisecond-level time differences in cross-domain data. In high-speed vibration scenarios, this time difference will cause motion artifacts and spectral phase shifts. Second, when post-fusing image and sound features, there is a lack of a cross-modal feature space alignment mechanism, which results in the forced aggregation of features with weak semantic correlation, weakening the physical interpretability of state representation and hindering accurate diagnosis. Summary of the Invention
[0005] This application provides a cross-domain IoT device intelligent collaboration method and system based on semantic knowledge graph to solve the problems of asynchronous fragmentation of cross-domain data and semantic mismatch of multimodal features in the existing technology.
[0006] In the first aspect, this application provides a cross-domain IoT device intelligent collaboration method based on semantic knowledge graph, including: Convert the mechanical vibration energy of the monitored equipment into electrical energy signals and generate corresponding level values synchronously; Adjusting the sampling frequency of the cross-domain IoT device according to the level value, and obtaining the image texture data and sound spectrum data of the monitored device collected by the cross-domain IoT device after the sampling frequency adjustment is completed; Based on the image texture data and the sound spectrum data, a feature-level fusion classifier is used to perform cross-modal feature space alignment to generate a joint semantic feature vector representing the state of the monitored device; Constructing a semantic knowledge graph based on the joint semantic feature vector; According to the level value, the feature weights of the cross-domain Internet of Things devices in the semantic knowledge graph are dynamically optimized to generate cross-domain Internet of Things device collaborative control instructions based on the dynamically optimized feature weights.
[0007] Optionally, the step of performing cross-modal feature space alignment based on the image texture data and the sound spectrum data using a feature-level fusion classifier to generate a joint semantic feature vector representing the state of the monitored device includes: Performing a mapping operation on the image texture data input into an image projection function to generate a projected image feature, and performing a mapping operation on the sound spectrum data input into a sound projection function to generate a projected sound feature; Based on the projected image features and the projected sound features, performing an inter-modality similarity calculation operation in a shared latent space to generate an inter-modality similarity measurement value set; Based on the inter-modal similarity measurement value set, a feature-level fusion classifier is used to perform an association mapping construction operation to generate a cross-modal feature association mapping table; Reconstructing features according to the cross-modal feature association mapping table to generate feature slices of unified dimension; The feature slices are input into a bidirectional feature interaction module for multiple rounds of aggregation to generate a joint semantic feature vector representing the state of the monitored device.
[0008] Optionally, the performing feature reconstruction according to the cross-modal feature association mapping table to generate feature slices of uniform dimension includes: Performing a block operation on the projected image features to generate an image feature subsegment set, and performing a segmentation operation on the projected sound features to generate a sound feature subsegment set; Performing a key-value matching operation based on the image feature subsegment set, the sound feature subsegment set, and the cross-modal feature association mapping table to generate a matching feature pair set; Performing a cross-aggregation operation on the set of matching feature pairs to generate a fused feature unit sequence; Based on the image feature subsegment set and the sound feature subsegment set, performing an unmatched identifier extraction operation to generate an unmatched image identifier set and an unmatched sound identifier set; performing an interpolation compensation operation based on the unmatched image identifier set and the unmatched sound identifier set to generate a compensated image feature unit and a compensated sound feature unit; A multi-dimensional structure reorganization operation is performed on the fused feature unit sequence, the compensated image feature unit, and the compensated sound feature unit to generate feature slices of uniform dimension.
[0009] Optionally, the performing an association mapping construction operation based on the inter-modality similarity measurement value set using a feature-level fusion classifier to generate a cross-modality feature association mapping table includes: In the feature-level fusion classifier, a threshold screening operation is performed on the set of inter-modality similarity measures to generate a set of candidate mapping pairs; Performing a feature index grouping operation on the projected image features and the projected sound features to generate an image feature index group and a sound feature index group; Performing bidirectional association constraints on the image feature index group and the sound feature index group to generate a valid mapping rule set; Perform rule verification on the candidate mapping pair set and a preset valid mapping rule set to generate a verified mapping pair; A key-value table construction operation is performed on the verified mapping pairs to generate a cross-modal feature association mapping table.
[0010] Optionally, constructing a semantic knowledge graph based on the joint semantic feature vector includes: Extracting state nodes from the joint semantic feature vector to generate a device state node set; Performing a physical relationship modeling operation on the device state node set to generate a physical topology relationship table; Perform edge connection operations based on the physical topology relationship table to generate an initial knowledge graph skeleton; performing eigenvalue analysis on the joint semantic feature vector to generate a vibration eigenvalue set; Injecting the vibration feature value set into the initial knowledge graph skeleton to assign edge weights and generate a weighted knowledge graph; A spatiotemporal constraint binding operation is performed on the weighted knowledge graph to obtain a semantic knowledge graph.
[0011] Optionally, injecting the vibration feature value set into the initial knowledge graph skeleton to assign edge weights to generate a weighted knowledge graph includes: Performing an edge type recognition operation on the initial knowledge graph skeleton to generate a spatial edge set and a temporal edge set; performing a feature dimension separation operation on the vibration feature value set to generate an amplitude feature set and an attenuation feature set; Performing feature assignment operations on the spatial edge set and the amplitude feature set, and on the temporal edge set and the attenuation feature set, respectively, to generate a spatial edge amplitude mapping table and a temporal edge attenuation mapping table; Based on the spatial edge amplitude mapping table and the temporal edge attenuation mapping table, performing spatial weight and temporal weight calculation operations to generate spatial edge weight values and edge temporal weight values; The spatial edge weight value and the temporal edge weight value are injected into the initial knowledge graph skeleton to update the weights and generate a weighted knowledge graph.
[0012] Optionally, dynamically optimizing the feature weights of the cross-domain IoT devices in the semantic knowledge graph according to the level values, and generating cross-domain IoT device collaborative control instructions according to the dynamically optimized feature weights, includes: performing vibration mode analysis on the level value to generate a dominant vibration mode identifier; Performing an edge weight extraction operation on the semantic knowledge graph to generate an original weight distribution table; Inputting the dominant vibration pattern identifier into the original weight distribution table for pattern matching to generate a matching weight subset, and inputting the matching weight subset into the semantic knowledge graph for weight update operation to generate an optimized knowledge graph; Traversing the optimized knowledge graph to retrieve control rules and generate an initial control instruction set; Perform device collaboration constraint operations on the initial control instruction set to generate cross-domain IoT device collaboration control instructions.
[0013] Secondly, this application provides a cross-domain IoT device intelligent collaboration system based on semantic knowledge graph, including: The conversion module is used to convert the mechanical vibration energy of the monitored equipment into an electrical energy signal and synchronously generate a corresponding level value; an acquisition module, configured to adjust a sampling frequency of the cross-domain IoT device according to the level value, and after the sampling frequency adjustment is completed, acquire image texture data and sound spectrum data of the monitored device collected by the cross-domain IoT device; A generation module, configured to perform cross-modal feature space alignment based on the image texture data and the sound spectrum data using a feature-level fusion classifier to generate a joint semantic feature vector representing the state of the monitored device; A construction module, configured to construct a semantic knowledge graph based on the joint semantic feature vector; A dynamic optimization module is used to dynamically optimize the feature weights of cross-domain IoT devices in the semantic knowledge graph according to the level value, so as to generate cross-domain IoT device collaborative control instructions based on the dynamically optimized feature weights.
[0014] In a third aspect, the present application provides a computing device comprising a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement the cross-domain IoT device intelligent collaboration method based on semantic knowledge graph as described in any one of the first aspects.
[0015] In a fourth aspect, the present application provides a computer storage medium storing a computer program. When the computer program is executed by a computer, it implements the cross-domain IoT device intelligent collaboration method based on semantic knowledge graph as described in any one of the first aspects.
[0016] In the present application, a method for intelligent collaboration of cross-domain Internet of Things devices based on a semantic knowledge graph is provided, the method comprising: converting the mechanical vibration energy of a monitored device into an electrical energy signal and synchronously generating a corresponding level value; adjusting the sampling frequency of the cross-domain Internet of Things device according to the level value, and after the sampling frequency adjustment is completed, obtaining image texture data and sound spectrum data of the monitored device collected by the cross-domain Internet of Things device; based on the image texture data and the sound spectrum data, using a feature-level fusion classifier to perform cross-modal feature space alignment to generate a joint semantic feature vector representing the state of the monitored device; constructing a semantic knowledge graph based on the joint semantic feature vector; and dynamically optimizing the feature weights of the cross-domain Internet of Things devices in the semantic knowledge graph according to the level value, so as to generate cross-domain Internet of Things device collaborative control instructions based on the dynamically optimized feature weights.
[0017] This application realizes self-powered sensing and quantitative characterization of vibration intensity by converting mechanical vibration energy into electrical energy signals and synchronously generating level values; dynamically adjusts the sampling frequency based on the level value to ensure that cross-domain equipment automatically optimizes data acquisition efficiency when the vibration working conditions change; after synchronously acquiring image texture and sound spectrum data, uses feature-level fusion classifiers to perform cross-modal spatial alignment to eliminate the semantic gap between heterogeneous data and generate a highly discriminative joint semantic feature vector; constructs a semantic knowledge graph based on this vector to form an interpretable device state topological relationship network; finally, dynamically optimizes the device feature weights in the graph according to the level value, so that the collaborative control instructions adapt to the current vibration intensity in real time, achieving a closed-loop response of state perception and device regulation.
[0018] Furthermore, texture data and spectral data are first mapped into projection features through image and sound projection functions respectively; a set of inter-modal similarity measurement values is calculated in a shared latent space, and a cross-modal feature association mapping table is constructed based on this; the projected image feature blocks are divided into image sub-segment sets, and the sound feature segments are simultaneously divided into sound sub-segment sets; based on the association mapping table, key-value matching is performed on the two types of sub-segment sets to obtain a set of matching feature pairs, and a fusion feature unit sequence is generated through cross-aggregation; at the same time, unmatched image and sound identifier sets are extracted, and compensation feature units are generated through interpolation compensation; finally, the fusion sequence and the compensation unit are multi-dimensionally reorganized to generate feature slices of unified dimension, which are input into the bidirectional feature interaction module for multi-round aggregation to output a joint semantic feature vector. This application can achieve fine-grained cross-modal feature alignment and information preservation, and the projection function mapping ensures that heterogeneous data are represented in a comparable latent space; the key-value matching mechanism realizes the local semantic association between image texture blocks and sound spectrum segments, solving the problem of feature misalignment under non-steady-state signals; unmatched identifier extraction and interpolation compensation retain the unique features of a single modality, avoiding information loss caused by cross-modal matching; multi-dimensional reorganization generates unified dimensional feature slices, providing structurally aligned input for subsequent two-way interaction, and improving the physical interpretability and state discrimination robustness of the joint feature vector.
[0019] These and other aspects of the present application will become more readily apparent from the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0021] Figure 1 A flowchart of a cross-domain IoT device intelligent collaboration method based on a semantic knowledge graph provided in an embodiment of the present application; Figure 2 A schematic diagram of the structure of a cross-domain IoT device intelligent collaboration system based on a semantic knowledge graph provided in an embodiment of the present application; Figure 3 A schematic diagram of the structure of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0022] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0023] In some of the processes described in the specification and claims of this application and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this document or may be executed in parallel. The serial numbers of the operations, such as 11, 12, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to being different types.
[0024] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0025] In order to solve the problems of asynchronous fragmentation of cross-domain data and semantic mismatch of multimodal features in the prior art, an embodiment of the present application provides a cross-domain IoT device intelligent collaboration method based on a semantic knowledge graph. The method adopts the following concept: through the conversion between vibration energy and electrical energy, a level signal is generated in real time. With this signal as the control benchmark, the sampling frequency of the cross-domain IoT device is dynamically adjusted to ensure that data acquisition matches the vibration state of the device; the image texture and sound spectrum data after optimized sampling are synchronously obtained, and the input is a feature-level fusion classifier for cross-modal spatial alignment to generate a joint semantic feature vector representing the health status of the device; an interpretable semantic knowledge graph is constructed based on the vector, and the level signal is used to dynamically optimize the feature weights of different devices in the graph, and finally accurate cross-domain device collaborative control instructions are generated to achieve energy consumption adaptation and status monitoring closed loop.
[0026] Figure 1 The flowchart of the cross-domain IoT device intelligent collaboration method based on semantic knowledge graph provided in the embodiment of this application is as follows: Figure 1 As shown, the method includes: S11. Convert the mechanical vibration energy of the monitored equipment into an electrical energy signal and synchronously generate a corresponding level value.
[0027] The monitored equipment can be any type of industrial equipment. The mechanical vibration energy of the monitored equipment can refer to the kinetic energy generated by bearing wear or rotor imbalance in rotating equipment, which is converted into an electrical signal through piezoelectric materials. The electrical energy signal is a continuous analog current / voltage waveform generated by piezoelectric conversion of mechanical vibration energy, reflecting the instantaneous vibration state of the equipment. The level value is the discrete digital value generated by analog-to-digital conversion of the electrical energy signal, and the value range is positively correlated with the vibration intensity.
[0028] In an embodiment of the present application, the mechanical vibration energy of the monitored equipment is first captured by a piezoelectric conversion device and converted into a continuous analog electrical energy signal; secondly, the electrical energy signal is quantized using an analog-to-digital converter to synchronously generate a digital level value representing the vibration intensity.
[0029] S12. Adjust the sampling frequency of the cross-domain IoT device according to the level value. After the sampling frequency adjustment is completed, obtain the image texture data and sound spectrum data of the monitored device collected by the cross-domain IoT device.
[0030] Among them, cross-domain IoT devices refer to terminals that integrate multimodal sensors such as vision and acoustics, including industrial cameras and directional microphones. They can also include other types of monitoring equipment, which are not specifically limited in the embodiments of this application. The sampling frequency refers to the rate at which IoT devices collect data, and is dynamically adjusted based on the level value to avoid aliasing distortion under high-frequency vibration. Image texture data can refer to the grayscale distribution matrix of the device surface captured by the camera, including degradation features such as cracks and rust.
[0031] In an embodiment of the present application, the level value is first analyzed to determine the current vibration intensity level; secondly, according to the preset intensity-frequency mapping relationship, a sampling frequency adjustment instruction is sent to the cross-domain IoT device; then, after the device completes the sampling frequency reset, the camera is synchronously triggered to collect image texture data on the surface of the device, and the microphone is started to collect sound spectrum data of the device operation.
[0032] S13. Based on the image texture data and the sound spectrum data, a feature-level fusion classifier is used to perform cross-modal feature space alignment to generate a joint semantic feature vector representing the state of the monitored device.
[0033] Sound spectrum data refers to the frequency domain energy distribution generated by Fourier transforming the sound waves collected by the microphone, which characterizes abnormal sound characteristics. A feature-level fusion classifier uses a neural network model with a shared latent space mapping to achieve cross-modal feature alignment and fusion. Cross-modal feature space alignment maps image textures and sound spectra into a unified vector space to eliminate semantic ambiguity. A joint semantic feature vector is a dense vector that fuses bimodal features, with dimensions containing joint semantic information about the device's degradation state.
[0034] In an embodiment of the present application, the image texture data is first input into a pre-trained convolutional neural network to extract spatial features and generate projected image features; at the same time, the sound spectrum data is input into a time series encoder to extract frequency domain features and generate projected sound features; secondly, the cosine similarity matrix of the two types of features is calculated in a shared latent space to generate a set of inter-modal similarity measurement values; then, a cross-modal feature association mapping table is constructed through a feature-level fusion classifier; finally, the features are reconstructed and aligned based on the mapping table, and a joint semantic feature vector is generated by aggregation through a bidirectional long short-term memory network.
[0035] S14. Construct a semantic knowledge graph based on the joint semantic feature vector.
[0036] Among them, the semantic knowledge graph refers to a topological network with device status nodes as entities and physical relationships as edges, and the edge weight represents the strength of fault correlation.
[0037] In an embodiment of the present application, the joint semantic feature vector is first modeled as a graph node to extract the device state node set; secondly, the node connection edges are constructed based on the physical topological relationship of the devices to form an initial knowledge graph skeleton; then, the vibration eigenvalues in the vector are parsed and injected into the edge weight attributes of the skeleton; finally, the spatiotemporal constraint rules are superimposed to generate a weighted semantic knowledge graph.
[0038] S15. Dynamically optimize the feature weights of cross-domain IoT devices in the semantic knowledge graph according to the level values, and generate collaborative control instructions for cross-domain IoT devices according to the dynamically optimized feature weights.
[0039] Feature weights refer to the contribution parameters of nodes and edges in the knowledge graph, which are dynamically adjusted based on vibration intensity to adapt to changing operating conditions. Cross-domain IoT device collaborative control instructions are device linkage instructions generated based on the optimized graph, such as adjusting fan speed or triggering an alarm.
[0040] In an embodiment of the present application, the current dominant vibration mode is first analyzed according to the level value to generate a mode identifier; secondly, the feature weight distribution table is extracted from the semantic knowledge graph; then the weight subset corresponding to the identifier is matched and dynamically optimized and updated; finally, the optimized graph is traversed to generate device control rules, and cross-domain IoT device collaborative control instructions are output after collaborative constraint verification.
[0041] Here's a specific example: In an industrial fan bearing monitoring scenario, the fan's mechanical vibration, such as that caused by bearing wear, is first captured by a vibration sensor and converted into a continuous voltage or current signal. Assume the original signal voltage is measured at 1.5 volts. Next, a digitizer quantizes this analog signal into a discrete digital value, known as a level. The level is calculated as follows: the input voltage divided by the maximum voltage range, multiplied by the maximum digital value. Here, the maximum voltage range is set to 3 volts, and the digitizer has 8 bits, so the maximum digital value is 255. According to the formula: the input voltage of 1.5 volts divided by the maximum voltage range of 3 volts equals 0.5, and 0.5 multiplied by the maximum digital value of 255 equals 127.5, which, after rounding, is approximately 128. This level is proportional to the vibration intensity; larger values indicate stronger vibrations.
[0042] Next, based on the calculated level value of 128, a preset intensity-frequency relationship graph is queried. This graph defines the sensor sampling rates corresponding to different level ranges: for example, a 50 Hz sampling rate is used for levels between 0 and 100, a 100 Hz sampling rate for levels between 100 and 200, and a 200 Hz sampling rate for levels above 200. Therefore, since the current level value of 128 falls within the 100-200 range, the sampling rates of the visual and acoustic sensors are adjusted to 100 Hz. After this adjustment, the intelligent monitoring device, which integrates a camera and microphone, is activated. The camera begins capturing image data from the wind turbine bearing surface, a grayscale matrix containing detailed information such as cracks. The microphone captures the sound waveform of the bearing's operation and, through Fourier transform, converts it into a frequency domain spectrum showing abnormal vibration frequencies.
[0043] The acquired image data is then processed by an image conversion model. The model's computational process is defined as follows: image features are obtained by applying a rectified linear unit function to the maximum pooling result of a two-dimensional convolution operation, ultimately generating a feature vector containing spatial texture information. Simultaneously, the collected sound spectrum data Y is input into a sound coding model. This model first converts the original spectrum into a mel-spectrogram that better reflects human hearing using a mel filter bank. This is then processed using a gated recurrent unit and finally layer normalized to generate a vector reflecting time-frequency domain features. Next, cross-modal feature fusion is performed: the generated image and sound feature vectors are mapped into a shared feature space, and their cosine similarity is calculated within this space. Based on the calculated similarity, highly matching image and sound segment features are recombined into standardized feature segments. For local features that do not match, interpolation and padding are performed using feature values from neighboring spatial or temporal regions. Ultimately, all recombined and padded feature segments are integrated into a unified joint semantic feature vector. This vector incorporates key information from both image and sound modalities, jointly characterizing the current wear state of the bearing.
[0044] Next, a semantic knowledge graph is constructed based on this joint semantic feature vector. First, key equipment status nodes are extracted, such as the node representing "inner race crack" and the node representing "normal wear." Connection rules are then established between the nodes based on the actual physical structure diagram of the wind turbine. For example, an edge is added between the node representing "inner race" and the node representing "ball bearing" to reflect their mechanical connection. These connecting edges are then injected with vibration eigenvalues extracted from the joint feature vector, such as an amplitude value of 0.8. This amplitude value is used to calculate the edge weight, using the formula: the weight is equal to the amplitude divided by the distance between the physical components represented by the nodes. Assuming the calculated physical distance between the node representing the inner race crack and the associated ball bearing node is 10 units, and the amplitude is 0.8, the weight of this edge is 0.8 divided by 10, which equals 0.08.
[0045] Finally, based on the current vibration signal level value of 128, in-depth spectral analysis is performed to identify the dominant vibration mode. For example, frequency feature matching identifies it as a "shaft imbalance mode." Based on this pattern identifier, all edges strongly associated with this pattern and their current weights are retrieved from the constructed semantic knowledge graph to form a subset. The weights of these edges are then updated based on the energy intensity coefficient of the dominant vibration mode (e.g., a coefficient of 0.6). The update formula is: the new weight equals the original weight multiplied by 1 plus the intensity coefficient. For example, the original edge weight is 0.08 × (1 + 0.6) = 0.08 × 1.6 = 0.128. After the weight update, an optimized knowledge graph is obtained. A depth-first traversal of this optimized graph triggers a matching operation based on the pre-set control rule library, generating a preliminary set of equipment control instructions, such as "reduce fan speed by 10%" and "start the lubrication system." To ensure that operating instructions for different devices do not conflict or exceed system limits, the system verifies and optimizes these preliminary instructions based on pre-set coordination constraints (such as the maximum allowable speed difference limit for the fan and the system's total power limit). Ultimately, it generates and outputs coordinated, cross-domain IoT device collaborative control instructions, such as "reduce fan speed by 8% while simultaneously activating the fixed-point lubrication system." This entire process completes a closed-loop process, from adaptive vibration data collection, multimodal information fusion, dynamic state knowledge graph construction, to intelligent collaborative control decision-making, effectively improving industrial equipment monitoring accuracy and maintenance efficiency.
[0046] By executing S11 to S15, the embodiment of the present application adaptively adjusts the data collection efficiency through vibration intensity to eliminate feature omissions under high-speed working conditions; uses cross-modal feature alignment to construct a highly discriminative equipment state representation to support the construction of an interpretable knowledge graph; and generates precise collaborative instructions based on dynamic weight optimization to improve equipment fault response speed and maintenance efficiency.
[0047] In a possible embodiment, S13, based on the image texture data and the sound spectrum data, a feature-level fusion classifier is used to perform cross-modal feature space alignment to generate a joint semantic feature vector representing the state of the monitored device, including: Step 131: Perform a mapping operation on the image texture data input into the image projection function to generate a projected image feature; perform a mapping operation on the sound spectrum data input into the sound projection function to generate a projected sound feature.
[0048] Among them, the image projection function refers to a mapping model implemented by a convolutional neural network, which is used to convert a two-dimensional image texture into a high-dimensional feature vector, including edge detection and texture encoding operations. The projected image feature refers to the tensor generated by the image data after processing by the projection function, and the dimension reflects the abstract semantic level of the spatial texture. The sound projection function refers to a mapping model based on a temporal encoder, which is used to convert a spectrum sequence into a joint time-frequency representation, including a Mel filter bank and a time convolution operation. The projected sound feature refers to the vector sequence generated by the spectrum data through the sound projection function, and the length corresponds to the number of time windows. Exemplarily, the formula of the image projection function can be: ,in, is the image feature after projection, is the rectified linear unit, is the maximum pooling operation, is a two-dimensional convolution operation, is the image texture data, is the weight parameter of the convolutional layer, is the bias parameter of the convolutional layer.
[0049] The formula of the sound projection function can be: ,in, is the sound feature after projection, is layer normalization, is a gated recurrent unit, To convert the original spectrum into Mel spectrum, is the sound spectrum, is the projection weight.
[0050] In an embodiment of the present application, the image texture data is first input into an image projection function based on a convolutional neural network to perform a spatial feature extraction operation, thereby generating a projected image feature containing local texture semantics; at the same time, the sound spectrum data is input into a sound projection function based on a gated recurrent unit to perform a time-frequency feature encoding operation, thereby generating a projected sound feature that characterizes the spectrum evolution law.
[0051] Step 132: Based on the projected image features and the projected sound features, perform an inter-modality similarity calculation operation in the shared latent space to generate an inter-modality similarity measurement value set.
[0052] The shared latent space refers to the mathematical space constructed by the fully connected layers of a neural network, which is used to unify the distance metrics used to represent heterogeneous modal features. The inter-modal similarity calculation operation is the process of performing vector dot products and norm calculations within the latent space, with output values ranging from -1 to 1. The inter-modal similarity metric set is a matrix that records the similarities of all cross-modal feature pairs, with rows and columns corresponding to image feature indices and sound feature indices, respectively.
[0053] In an embodiment of the present application, in an embodiment of the present application, the image texture data is first input into an image projection function based on a convolutional neural network to perform a spatial feature extraction operation, thereby generating a projected image feature containing local texture semantics; at the same time, the sound spectrum data is input into a sound projection function based on a gated recurrent unit to perform a time-frequency feature encoding operation, thereby generating a projected sound feature that characterizes the spectrum evolution law.
[0054] Step 132: Based on the projected image features and the projected sound features, perform an inter-modality similarity calculation operation in the shared latent space to generate an inter-modality similarity measurement value set.
[0055] The shared latent space refers to the mathematical space constructed by the fully connected layers of a neural network, which is used to unify the distance metrics used to represent heterogeneous modal features. The inter-modal similarity calculation operation is the process of performing vector dot products and norm calculations within the latent space, with output values ranging from -1 to 1. The inter-modal similarity metric set is a matrix that records the similarities of all cross-modal feature pairs, with rows and columns corresponding to image feature indices and sound feature indices, respectively.
[0056] In an embodiment of the present application, in the actual application scenario of industrial fan bearing monitoring, the system first collects the mechanical vibration signal generated by bearing wear during the operation of the fan and converts it into a voltage signal. Assume that the voltage value detected by the sensor at this time is 1.5 volts. The analog signal is then input into the analog-to-digital converter for quantization processing, and the digital level value 128 is calculated according to the preset conversion rules. This quantization process follows the standard conversion formula, where 3 volts is the maximum range of the device, 255 corresponds to the maximum digital value of the 8-bit converter, and the final value is calculated by the ratio of the actual voltage to the range.
[0057] When the system begins processing sensor data, it performs key feature extraction operations. Image data collected from the wind turbine bearing surface is fed into a dedicated image processing model, which uses a multi-layer neural network to extract spatial features, including convolution and feature compression. For example, after processing a bearing crack image with a resolution of 1920 by 1080 pixels, the system generates a vector containing 256 feature dimensions. Simultaneously, the collected sound spectrum data is analyzed by a dedicated sound processing model. This model converts the raw audio data into a time-frequency feature sequence through operations such as spectral conversion and time series encoding. For example, 10 seconds of audio data undergoes a 128-order spectral conversion to produce serialized data containing 80 features.
[0058] The system then performs cross-modal correlation analysis within a shared feature space: features extracted from both sensor data types are mapped into a unified dimensional space. Within this space, the system automatically calculates the similarity between each image feature block and the acoustic feature segment. For example, an image feature block representing an inner ring crack has a 92% similarity to the acoustic feature segment representing a 12 kHz vibration noise. The system records this similarity in a 32-row, 40-column similarity matrix.
[0059] Next, a feature association table is constructed based on the similarity results. Specifically, the system sets an 80% similarity threshold as the association threshold to screen for reliable feature correspondences. A neural network is also used to analyze the physical correlations between features, automatically identifying, for example, the correspondence between surface crack textures and high-frequency noise. Ultimately, a feature mapping table containing 256 sets of correspondences is constructed, detailing the feature block pairing information.
[0060] Then, a multimodal feature reconstruction operation is performed. Specifically, the image features are divided into 32 blocks according to spatial relationships, and the sound features are divided into 40 segments according to temporal relationships. The matching feature pairs are then fused, and the five unmatched image blocks are padded with the feature values of the surrounding blocks, and the eight unmatched sound segments are padded with the mean values of adjacent time periods. Finally, a standardized 32 by 40 feature matrix is generated, and the features of each dimension maintain spatiotemporal alignment.
[0061] Finally, the device state vector is synthesized through a deep network. Specifically, the feature input is reconstructed into a neural network module with a bidirectional analysis function. This module captures the crack development trend through forward time series analysis and predicts the wear development path through reverse analysis. After three iterative optimizations, a 32-dimensional joint feature vector is finally generated. Each dimension of this vector reflects a specific state feature. For example, the value of 0.83 in the seventh dimension represents the comprehensive probability of inner ring wear matching high-frequency abnormal noise.
[0062] By executing steps 131 to 135, the embodiment of the present application eliminates the semantic gap between image texture and sound spectrum by sharing latent space; retains the correspondence between local features by using an association mapping table; and generates a highly discriminative joint representation of device status through bidirectional interactive aggregation, providing high-precision feature support for fault diagnosis.
[0063] In a possible embodiment, step 134, performing feature reconstruction according to the cross-modal feature association mapping table to generate feature slices of a unified dimension, includes: Step a1: performing a block operation on the projected image features to generate an image feature sub-segment set, and performing a segmentation operation on the projected sound features to generate a sound feature sub-segment set.
[0064] Among them, the blocking operation may refer to a technique of dividing an image feature matrix into local blocks according to a preset grid size, and the grid parameters are set according to the resolution of the feature map. An image feature sub-segment set refers to a container for storing local features after blocking, and its elements include key-value pairs of block coordinates and feature vectors. A segmentation operation refers to a process of dividing the time series features of a sound according to a fixed time window length, and the window length is dynamically adjusted by the period of spectrum change. A sound feature sub-segment set refers to a container for storing time series features after segmentation, and its elements include a time window starting point and a feature vector. It should be noted that the embodiments of the present application do not specifically limit the size of the preset grid size, fixed time window length, etc.
[0065] In an embodiment of the present application, a grid-based blocking operation is first performed on the projected image features, and the feature tensor is divided into several local blocks along the spatial dimension to generate a set of image feature sub-segments containing position indexes; at the same time, a time window sliding-based segmentation operation is performed on the projected sound features to divide the time series features into segments of equal length to generate a set of sound feature sub-segments with timestamps.
[0066] Step a2: performing a key-value matching operation based on the image feature sub-segment set, the sound feature sub-segment set, and the cross-modal feature association mapping table to generate a matching feature pair set.
[0067] The key-value matching operation refers to the query process of retrieving the associated sound segment identifiers through the image block identifier based on the mapping table. The matching feature pair set is a list of tuples that records the successfully matched image block and sound segment identifier combinations.
[0068] In an embodiment of the present application, first, the cross-modal feature association mapping table is read to establish a key-value correspondence rule; secondly, each block identifier of the image feature sub-segment set is traversed as a key, and the associated sound feature segment identifier is retrieved in the mapping table as a value; then, a matching record is generated for the successfully retrieved key-value pairs; finally, a matching feature pair set containing all matching relationships is output.
[0069] Step a3: perform a cross-aggregation operation on the set of matching feature pairs to generate a fused feature unit sequence.
[0070] The cross-aggregation operation refers to the weighted fusion calculation performed on matching feature pairs, with the weights determined by cross-modal similarity. The fused feature unit sequence refers to the vector sequence generated by aggregation, and the order maintains the original spatial or temporal relationship.
[0071] In an embodiment of the present application, each image sub-segment and sound sub-segment in the set of matching feature pairs is first extracted; secondly, the feature interaction weight is calculated through a cross-modal attention mechanism; then, a weighted sum operation is performed on the image feature vector and the sound feature vector according to the weight; finally, the aggregation results are aggregated in sequential order to generate a fusion feature unit sequence.
[0072] Step a4: Based on the image feature sub-segment set and the sound feature sub-segment set, perform an unmatched identifier extraction operation to generate an unmatched image identifier set and an unmatched sound identifier set.
[0073] The unmatched identifier extraction operation involves identifying feature subsegment identifiers not covered by the mapping table through set difference calculation. The unmatched image identifier set refers to the set of unmatched block coordinate indexes within the image feature subsegment set. The unmatched sound identifier set refers to the set of unmatched time window start point indexes within the sound feature subsegment set.
[0074] In an embodiment of the present application, the unmatched block identifiers in the image feature sub-segment set are first scanned to generate an unmatched image identifier set; secondly, the unmatched segment identifiers in the sound feature sub-segment set are scanned to generate an unmatched sound identifier set; finally, the two sets are deduplicated for redundant identifiers.
[0075] Step a5: performing an interpolation compensation operation based on the unmatched image identifier set and the unmatched sound identifier set to generate a compensated image feature unit and a compensated sound feature unit.
[0076] Interpolation compensation refers to the generation of replacement vectors for unmatched features based on spatial or temporal proximity. Compensated image feature units are virtual features generated by interpolating feature vectors from adjacent blocks, preserving local texture characteristics. Compensated sound feature units are compensation vectors generated by averaging the features of preceding and subsequent time windows, maintaining frequency domain energy continuity.
[0077] In an embodiment of the present application, first, the corresponding feature sub-segment is located according to the unmatched image identifier set, and the compensated image feature unit is generated by linear interpolation of adjacent block features; secondly, the corresponding feature sub-segment is located according to the unmatched sound identifier set, and the compensated sound feature unit is generated by filling the mean of the temporally adjacent segments; finally, the feature dimension normalization processing is performed on the compensation unit.
[0078] Step a6: performing a multi-dimensional structural reorganization operation on the fused feature unit sequence, the compensated image feature unit, and the compensated sound feature unit to generate feature slices of uniform dimension.
[0079] Among them, the multidimensional structure reorganization operation refers to the process of reorganizing heterogeneous features into a tensor of uniform dimension according to the original structure, including channel alignment and position calibration.
[0080] In an embodiment of the present application, the fused feature unit sequence is first sorted according to the original position; secondly, the compensating image feature unit is inserted into the vacant position of the image feature; then, the compensating sound feature unit is inserted into the vacant position of the sound feature; finally, a tensor splicing operation is performed to generate a feature slice with unified channel dimension.
[0081] The following is a specific example: In the feature reconstruction phase of the industrial fan bearing monitoring scenario, the system performs structured processing on the extracted image and sound features. First, the feature block segmentation operation is performed: the 256-dimensional feature vector generated by the bearing surface image is divided into 32 blocks according to the spatial position. Each block corresponds to a 60×60 pixel area and is assigned a unique identifier such as A 32 Represents the block at coordinate (120,90); Synchronously, the 80-dimensional feature sequence generated by the sound spectrum is divided into 40 segments along the time axis, each segment covers 0.25 seconds of audio data, and the identifier is S 28 Mark the segment starting at 7.0 seconds.
[0082] Then the key-value matching association is performed: the system calls the pre-stored 256 sets of feature association mapping tables and traverses all 32 image block identifiers for association retrieval. For example, block A is identified. 15 With sound clip S 22 There is a corresponding relationship, and finally a set of 28 valid matches is formed, such as (A 15 ,S 22 )、(A 21 ,S 30 ) and other paired combinations. In the stage of integrating key features, cross-aggregation is implemented for 28 sets of matching features: weight coefficients are calculated based on cross-modal similarity, such as A 15 With S 22 The similarity is 0.92. The weighted fusion formula is used to synthesize the image features and sound features into a new feature unit at a ratio of 0.6:0.4, generating a 28-dimensional fusion sequence that maintains the original spatiotemporal order. Compensation processing is performed for unmatched features. Specifically: the system identifies 4 unrelated image blocks, such as A3, A7, and A 11 、A 28 , the compensation value is generated by linear interpolation of adjacent block features, for example, the compensation value of block A7 is the average value of the adjacent A6 and A8 features; at the same time, 12 unmatched sound segments are found, such as S1, S8, S 16 、S 32 , the mean of the temporally adjacent segments is used for filling, such as the S8 compensation value is the arithmetic mean of the S7 and S9 features.
[0083] Finally, multi-dimensional reconstruction is performed. Specifically, the 28 groups of fusion units are reordered according to spatial coordinates, compensation units are inserted at the 3rd, 7th, 11th, and 28th positions of the image block sequence, and sound compensation units are inserted at the 1st, 8th, 16th, and 32nd positions of the sound segment sequence, forming a standardized three-dimensional feature slice of 40 time segments and 256 feature channels. Taking the bearing inner ring crack monitoring as an example, when the crack area block A9 successfully matches the high-frequency abnormal sound segment S 14When the crack block A5 is not cracked, the compensation is completed through the feature interpolation of the adjacent blocks A4 and A6, and finally a feature matrix that completely characterizes the spatiotemporal state of the equipment is constructed.
[0084] By executing steps a1 to a6, the embodiment of the present application establishes local feature semantic associations through key-value matching; utilizes an interpolation compensation mechanism to avoid information loss; and generates structured fusion input through multi-dimensional reorganization to improve the discriminability and robustness of state representation.
[0085] In a possible embodiment, step 133, based on the inter-modality similarity metric value set, performing an association mapping construction operation using a feature-level fusion classifier to generate a cross-modal feature association mapping table, includes: Step b1: In the feature-level fusion classifier, a threshold screening operation is performed on the set of inter-modality similarity measure values to generate a set of candidate mapping pairs.
[0086] Among them, the candidate mapping pair set refers to the set of potential matching pairs obtained from the inter-modal similarity measurement value set through threshold screening, including pairing combinations when the similarity between image features and sound features is higher than the threshold, which is used to reflect the preliminary screening results and serve as input for subsequent verification.
[0087] In an embodiment of the present application, first, in a feature-level fusion classifier, a threshold screening operation is performed on a set of inter-modal similarity measure values, where the set of inter-modal similarity measure values includes similarity calculation results between image features and sound features; secondly, each similarity value is compared with a preset threshold, and feature pairs corresponding to values higher than the threshold are screened out; then, these feature pairs are integrated to generate a set of candidate mapping pairs, which includes potential matching image feature and sound feature combinations.
[0088] Step b2: performing a feature index grouping operation on the projected image features and the projected sound features to generate an image feature index group and a sound feature index group.
[0089] The image feature index group refers to the index set formed by grouping the projected image features, including image feature subsets based on feature similarity or spatial indexing. This group structure is used to represent the image features in a bidirectional association constraint. The sound feature index group refers to the index set formed by grouping the projected sound features, including sound feature subsets based on feature similarity or temporal indexing. This group structure is used to represent the sound features in a bidirectional association constraint.
[0090] In an embodiment of the present application, first, a feature index grouping operation is performed on the post-projection image features and the post-projection sound features; secondly, a clustering algorithm and an index-based grouping technology are used to group the post-projection image features according to the index values to form image feature index groups; at the same time, the post-projection sound features are grouped according to the index values to form sound feature index groups; subsequently, image feature index groups and sound feature index groups are generated, and these groups are used for subsequent association analysis.
[0091] Step b3: Perform bidirectional association constraints on the image feature index group and the sound feature index group to generate a valid mapping rule set.
[0092] Bidirectional association constraints are bidirectional consistency requirements imposed on image feature index groups and sound feature index groups, including ensuring that each image group uniquely maps to a sound group and vice versa. These constraints are used to generate valid rules based on inter-group associations. A valid mapping rule set is a set of rules generated through bidirectional association constraints, including logical rules defining valid mappings between image feature index groups and sound feature index groups. These rules are used to evaluate candidate pairs in rule verification.
[0093] In an embodiment of the present application, first, bidirectional association constraints are performed on the image feature index group and the sound feature index group; second, an association rule mining or optimization algorithm is applied to ensure that each group in the image feature index group and the corresponding group in the sound feature index group satisfy bidirectional mapping consistency; then, a valid mapping rule set is generated by analyzing the relationship between the groups, which defines the valid mapping rules between the image feature index group and the sound feature index group.
[0094] Step b4: Verify the candidate mapping pair set with the preset valid mapping rule set to generate verified mapping pairs.
[0095] The "preset valid mapping rule set" refers to a predefined set of valid mapping rules, including standard rules based on domain knowledge or historical data, used as a reference during the rule validation step. The "post-validation mapping pair" refers to a set of reliable mapping pairs after rule validation, including candidate mapping pairs that meet the predefined valid mapping rule set and are used as input data for key-value table construction.
[0096] In an embodiment of the present application, first, the candidate mapping pair set is subjected to rule verification against the preset valid mapping rule set; second, a rule matching algorithm is used to check whether each candidate mapping pair complies with the rules in the preset valid mapping rule set, and invalid pairs that violate the rules are filtered out; then, a verified mapping pair is generated, which includes the verified image feature and sound feature mapping pairs.
[0097] Step b5: perform a key-value table construction operation on the verified mapping pairs to generate a cross-modal feature association mapping table.
[0098] The key-value table construction operation refers to the operation of converting the verified mapping pair into a key-value structure, including constructing a mapping with features as keys and associated features as values, which is used to generate a cross-modal association table for structured storage.
[0099] In an embodiment of the present application, first, a key-value table construction operation is performed on the verified mapping pairs; secondly, the image features in each verified mapping pair are used as keys and the sound features as values, or vice versa, to construct key-value pairs; then, a cross-modal feature association mapping table is generated through data structure operations such as a hash table, which stores the association relationship between features.
[0100] The following is a specific example: First, in the feature-level fusion classifier, a threshold screening operation is performed on the set of inter-modal similarity measurement values of image features and sound features to generate a set of candidate mapping pairs. Secondly, a feature index grouping operation is performed on the projected image features and the projected sound features to form an image feature index group and a sound feature index group. Subsequently, a bidirectional association constraint is applied to the image feature index group and the sound feature index group to generate a valid mapping rule set. Next, the candidate mapping pair set is rule-verified against the preset valid mapping rule set to generate a verified mapping pair. Finally, a key-value table construction operation is performed on the verified mapping pair to generate a cross-modal feature association mapping table.
[0101] By executing steps b1 to b5, the embodiment of the present application implements inter-modal similarity measurement and threshold screening through a feature-level fusion classifier, generates candidate mapping pairs, and then combines feature index grouping and bidirectional association constraints to form an effective rule set. The mapping accuracy is improved through rule verification, and finally a cross-modal association mapping in the form of a key-value table is constructed, thereby enhancing the robustness and reliability of cross-modal feature matching and optimizing the efficiency of multimedia data processing.
[0102] In a possible embodiment, S14, constructing a semantic knowledge graph based on the joint semantic feature vector, includes: Step 141: extract state nodes from the joint semantic feature vector to generate a device state node set.
[0103] State node extraction involves using a clustering algorithm to identify dense regions within the joint feature vector that represent key device states. Each node corresponds to a specific failure mode or operating state. The device state node set is a container for storing core entities in the knowledge graph. Node attributes include state type and confidence parameters.
[0104] In an embodiment of the present application, the high-dimensional feature distribution in the joint semantic feature vector is first analyzed by a clustering algorithm; secondly, the density-concentrated areas in the feature space are identified as key state representation points; then a unique node identifier is assigned to each representation point; finally, a device state node set containing all key state nodes is generated.
[0105] Step 142: Perform a physical relationship modeling operation on the device status node set to generate a physical topology relationship table.
[0106] Among them, the physical topology relationship table refers to a two-dimensional table structure that records node connection rules. The fields include the starting node, the ending node, the relationship type, and the maximum action distance.
[0107] In an embodiment of the present application, the physical structure topology of the monitored device is first loaded; secondly, the spatial position relationship of the device status node set in the physical topology is analyzed; then, the connection rules between the nodes are established, including mechanical transmission relationships and electrical coupling relationships; finally, a physical topology relationship table is generated to record the physical dependency relationships between the nodes.
[0108] Step 143: Perform edge connection operations based on the physical topology relationship table to generate an initial knowledge graph skeleton.
[0109] The edge connection operation refers to the process of adding directed edges between graph nodes based on the physical topology relationship table. Edge types include spatial adjacency and functional coupling. The initial knowledge graph skeleton refers to the topological framework containing only nodes and basic connecting edges, without the addition of dynamic feature weights.
[0110] In an embodiment of the present application, the node connection rules in the physical topology relationship table are first read; secondly, the device status nodes are used as graph nodes; then directed connection edges are added according to the physical dependency relationship; finally, an initial knowledge graph skeleton containing nodes and edges is constructed.
[0111] Step 144 : Perform eigenvalue analysis on the joint semantic feature vector to generate a vibration eigenvalue set.
[0112] Eigenvalue parsing is a technique for separating vibration-related dimensions from a joint semantic feature vector. The parsing targets include time-domain peaks and frequency-domain resonant components. A vibration eigenvalue set is a structured container for storing multidimensional vibration parameters. Its elements include amplitude mean, attenuation coefficient, and frequency band energy fraction.
[0113] In an embodiment of the present application, the vibration-related feature dimensions in the joint semantic feature vector are first located; secondly, the time domain amplitude features and frequency domain energy features are extracted; then the feature values are standardized; finally, a vibration feature value set containing multi-dimensional vibration characteristics is generated.
[0114] Step 145: Inject the vibration feature value set into the initial knowledge graph skeleton to assign edge weights and generate a weighted knowledge graph.
[0115] Among them, the weighted knowledge graph refers to the dynamic topological network formed after the initial skeleton is assigned edge weights, and the edge weights reflect the strength of state transmission between nodes.
[0116] In an embodiment of the present application, the connection edge types in the initial knowledge graph skeleton are first parsed; secondly, the amplitude features in the vibration feature value set are mapped to the spatial edge; at the same time, the attenuation features are mapped to the temporal edge; then the spatial edge weight value and the temporal edge weight value are calculated; finally, the skeleton edge weight is updated to generate a weighted knowledge graph.
[0117] Step 146: Perform spatiotemporal constraint binding operations on the weighted knowledge graph to obtain a semantic knowledge graph.
[0118] Among them, the spatiotemporal constraint binding operation refers to the process of binding geographic location coordinates to nodes and timestamps to edge weights to achieve spatiotemporal alignment between the graph and the physical world.
[0119] In an embodiment of the present application, the geographic location coordinates and timestamp information of the device operation are first obtained; secondly, the spatial coordinates are bound to the graph node position attributes; then the timestamp is bound to the edge weight update time attribute; finally, a semantic knowledge graph that integrates spatiotemporal attributes is generated.
[0120] The following is a specific example: First, the inner ring crack and ball peeling state nodes are extracted from the joint feature vector of wind turbine bearing wear to form a device state node set. Secondly, a physical relationship table between nodes is established based on the wind turbine drive chain topology. The initial knowledge graph skeleton is then constructed using the state nodes as graph nodes and the physical relationships as connecting edges. The vibration features in the joint vector are then parsed to generate a set of vibration feature values, including impact amplitude and high-frequency attenuation. The amplitude features are then injected into the spatial edges to calculate weights, and the attenuation features are injected into the temporal edges to calculate weights, generating a weighted knowledge graph. Finally, the wind turbine's geographic coordinates are bound to the nodes, and the sampling timestamps are bound to the edge weights, forming a semantic knowledge graph that can be traced back in time and space.
[0121] By executing steps 141 to 146, the embodiment of the present application ensures that the graph structure conforms to the actual structure of the equipment through physical topology constraints; vibration feature injection makes the edge weights have physical meaning; and spatiotemporal binding enhances the decision-making support capability of the graph in predictive maintenance.
[0122] In one possible embodiment, step 145 of injecting the vibration feature value set into the initial knowledge graph skeleton to assign edge weights to generate a weighted knowledge graph includes: Step c1: Perform edge type recognition on the initial knowledge graph skeleton to generate a spatial edge set and a temporal edge set.
[0123] Edge type identification is a technique for classifying edges based on their relationship type attributes. Spatial edges represent physical location relationships, while temporal edges represent the order of state evolution. A spatial edge collection is a container for storing edge connections that describe the spatial location relationships of device components in a knowledge graph. Edge attributes include distance and direction parameters. A temporal edge collection is a container for storing edge connections that describe the evolution of device states over time. Edge attributes include time intervals and change trends.
[0124] In an embodiment of the present application, first, all connecting edges of the initial knowledge graph skeleton are traversed; secondly, spatial edges representing spatial position relationships and temporal edges representing temporal evolution relationships are identified based on the relationship type field in the edge attributes; then, the spatial edges are classified into the spatial edge set, and the temporal edges are classified into the temporal edge set; finally, two independent edge sets are output.
[0125] Step c2: perform a feature dimension separation operation on the vibration feature value set to generate an amplitude feature set and an attenuation feature set.
[0126] Feature dimension separation is the process of physically splitting vibration features into independent subsets. The amplitude feature set includes the peak value and RMS value, while the attenuation feature set includes the damping coefficient and energy decay rate. The amplitude feature set is a container for storing vibration intensity parameters, and its elements are strongly correlated with the spatial position of the device. The attenuation feature set is a container for storing vibration energy decay parameters, and its elements are strongly correlated with the time evolution process.
[0127] In an embodiment of the present application, the dimensional labels of the vibration feature value set are first parsed; secondly, the amplitude-related features describing the vibration intensity are separated to form an amplitude feature set; at the same time, the time domain features describing the energy attenuation are separated to form an attenuation feature set; finally, two feature subsets are output.
[0128] Step c3: perform feature assignment operations on the spatial edge set and the amplitude feature set, and on the temporal edge set and the attenuation feature set, respectively, to generate a spatial edge amplitude mapping table and a temporal edge attenuation mapping table.
[0129] The feature assignment operation is the process of establishing a mapping relationship between spectral edges and vibration features. Spatial edges are assigned amplitude features, while temporal edges are assigned attenuation features. The spatial edge amplitude mapping table is a two-dimensional mapping table that records spatial edge identifiers and associated amplitude feature values, supporting weighted retrieval. The temporal edge attenuation mapping table is a two-dimensional mapping table that records temporal edge identifiers and associated attenuation feature values, supporting weighted retrieval.
[0130] In an embodiment of the present application, first, a correspondence rule between spatial edges and amplitude features is established; secondly, each edge of the spatial edge set is traversed, and associated feature values are extracted from the amplitude feature set to generate a spatial edge amplitude mapping table; at the same time, a correspondence rule between temporal edges and attenuation features is established; then, each edge of the temporal edge set is traversed, and associated feature values are extracted from the attenuation feature set to generate a temporal edge attenuation mapping table.
[0131] Step c4: Based on the spatial edge amplitude mapping table and the temporal edge attenuation mapping table, perform spatial weight and temporal weight calculation operations to generate spatial edge weight values and edge temporal weight values.
[0132] The spatial weight calculation operation is the process of calculating edge weights by dividing the amplitude eigenvalue by the spatial distance. The weight value is positively correlated with the vibration intensity and negatively correlated with the distance. The temporal weight calculation operation is the process of calculating edge weights by multiplying the attenuation coefficient by the time interval. The weight value reflects the rate of state change. The spatial edge weight value is the dynamic weight parameter of the spatial edge. A larger value indicates a stronger impact of the spatial relationship on the current state. The temporal edge weight value is the dynamic weight parameter of the temporal edge. A larger value indicates a faster state evolution.
[0133] In an embodiment of the present application, the spatial edge amplitude mapping table is first read; secondly, the spatial edge weight value is calculated by the ratio of the amplitude eigenvalue to the spatial distance; at the same time, the temporal edge attenuation mapping table is read; then the temporal edge weight value is calculated by multiplying the attenuation coefficient and the time interval; finally, two types of weight values are output.
[0134] Step c5: Inject the spatial edge weight values and temporal edge weight values into the initial knowledge graph skeleton to update the weights and generate a weighted knowledge graph.
[0135] Among them, weight update is the operation of injecting the calculated weight value into the edge attribute of the knowledge graph to realize the dynamic construction of the graph.
[0136] In an embodiment of the present application, the spatial edges of the initial knowledge graph skeleton are first located; secondly, the calculated spatial edge weight values are injected to update the edge attributes; at the same time, the temporal edges are located; then the calculated temporal edge weight values are injected to update the edge attributes; and finally, a weighted knowledge graph is generated.
[0137] Here's a specific example: First, identify the spatial edges describing component locations and the temporal edges describing the wear process in the fan bearing knowledge graph, forming independent sets. Next, separate the amplitude and attenuation parameters from the bearing vibration signature to generate corresponding feature sets. Next, establish a mapping table between spatial edges and bearing seat amplitude, and a mapping table between temporal edges and wear attenuation. The spatial edge weight is then calculated by dividing the amplitude by the component distance, and the temporal edge weight is calculated by multiplying the attenuation coefficient by the time interval. Finally, the weight values are injected into the corresponding edges of the graph to generate a dynamically weighted knowledge graph.
[0138] By executing steps c1 to c5, the embodiment of the present application accurately quantifies the impact intensity of spatial relationships through amplitude features, captures the state evolution rate through attenuation features, and enables the graph edge weights to truly reflect the operating conditions of the equipment.
[0139] In a possible embodiment, S15, dynamically optimizing the feature weights of cross-domain IoT devices in the semantic knowledge graph according to the level values, and generating cross-domain IoT device collaborative control instructions according to the dynamically optimized feature weights, includes: Step 151: Perform vibration mode analysis on the level value to generate a dominant vibration mode identifier.
[0140] Vibration pattern analysis involves identifying dominant vibration characteristics within a device's level using spectrum analysis and pattern matching techniques. This includes extracting fundamental frequency resonance and harmonic components. The dominant vibration pattern identifier (DVI) represents a classification label for the device's core vibration type and is derived from a combination of a frequency range code and an energy contribution level.
[0141] In an embodiment of the present application, first, a fast Fourier transform is performed on the level value to extract spectral features; second, the dominant frequency component is identified through a pre-trained vibration pattern classifier; then, the feature template in the device vibration pattern library is matched; finally, a dominant vibration pattern identifier representing the current core vibration type is generated.
[0142] Step 152: Perform edge weight extraction on the semantic knowledge graph to generate an original weight distribution table.
[0143] The edge weight extraction operation traverses the knowledge graph's edge attributes and stores the weights in a structured manner, outputting a set of triples consisting of edge identifier, weight value, and update time. The original weight distribution table is a two-dimensional relational table that records all edge weights in the knowledge graph and supports fast retrieval based on vibration pattern features.
[0144] In an embodiment of the present application, all connected edges of the semantic knowledge graph are first traversed; secondly, the dynamic weight attribute value of each edge is extracted; then the correspondence between the edge identifier and the weight value is recorded; finally, the original weight distribution table containing the weight distribution of the entire graph is generated.
[0145] Step 153: Input the dominant vibration pattern identifier into the original weight distribution table for pattern matching to generate a matching weight subset, and input the matching weight subset into the semantic knowledge graph for weight update operation to generate an optimized knowledge graph.
[0146] Among them, pattern matching refers to the technology of comparing vibration pattern identifiers with characteristic fields in the weight distribution table. The matching criteria include frequency coverage and energy correlation. The matching weight subset refers to the subset of weight records in the original weight distribution table that are strongly correlated with the dominant vibration pattern, which is used for targeted optimization of the graph. The weight update operation refers to the process of dynamically adjusting the graph edge weights according to the vibration pattern intensity coefficient. The adjustment amount is determined by multiplying the pattern energy ratio by the weight sensitivity parameter. The optimized knowledge graph refers to the topological network after the edge weights are adaptively updated by the vibration pattern. Its edge weight distribution reflects the state association strength under the current working conditions.
[0147] In an embodiment of the present application, the characteristic frequency range corresponding to the dominant vibration mode identifier is first parsed; secondly, the edge weight subset associated with the frequency range is retrieved from the original weight distribution table; then the matching weight subset is input into the graph update module; finally, the corresponding edge weight value is adjusted according to the vibration mode intensity coefficient to generate an optimized knowledge graph.
[0148] Step 154: traverse the optimized knowledge graph to retrieve control rules and generate an initial control instruction set.
[0149] Control rule retrieval refers to the query process of traversing a pre-set rule base based on the state of graph nodes. Rule conditions include node threshold states and edge weight thresholds. The initial control instruction set refers to the set of basic device operation instructions generated by rule matching. Elements include device identification, operation type, and strength parameters.
[0150] In an embodiment of the present application, a preset device control rule library is first loaded; secondly, the node path of the optimized knowledge graph is depth-first traversed; then the node status and rule conditions are matched; finally, an initial control instruction set containing device operation instructions is generated.
[0151] Step 155: Perform device collaboration constraint operations on the initial control instruction set to generate cross-domain IoT device collaboration control instructions.
[0152] Among them, device collaborative constraint operation refers to the optimization process of resolving conflicts between multi-device instructions. The constraints include physical space mutual exclusivity, energy supply upper limit and execution timing dependency.
[0153] In an embodiment of the present application, the device operation conflicts of the initial control instruction set are first parsed; secondly, the safety distance constraints and energy allocation constraints between devices are applied; then, the conflicting instructions are resolved through the constraint satisfaction algorithm; finally, cross-domain IoT device collaborative control instructions that can be executed in parallel are generated.
[0154] The following is a specific example: First, the fan bearing level values are spectrally analyzed to generate an identifier for the unbalanced vibration mode of the shaft system. Next, all edge weights are extracted from the semantic knowledge graph to generate an original weight distribution table. Next, a subset of weights related to unbalanced vibration is matched, and the graph is updated based on the vibration intensity to generate an optimized knowledge graph. The graph is then traversed to retrieve nodes with excessive vibration, generating speed reduction instructions and lubrication start instructions. Finally, the equipment's safe speed difference constraint and total power limit are applied to generate a coordinated control instruction to reduce the speed by 30% and simultaneously initiate fixed-point lubrication.
[0155] By executing steps 151 to 155, the embodiment of the present application accurately captures the vibration mode influencing path through dynamic optimization of the knowledge graph; generates safe and efficient cross-device linkage instructions based on constraint resolution, and improves the system response reliability under complex working conditions.
[0156] Figure 2 A schematic diagram of the structure of a cross-domain IoT device intelligent collaboration system based on a semantic knowledge graph provided in an embodiment of the present application is shown in FIG. Figure 2 As shown, the system includes: The conversion module 21 is used to convert the mechanical vibration energy of the monitored device into an electrical energy signal and synchronously generate a corresponding level value.
[0157] The acquisition module 22 is used to adjust the sampling frequency of the cross-domain Internet of Things device according to the level value, and after the sampling frequency adjustment is completed, obtain the image texture data and sound spectrum data of the monitored device collected by the cross-domain Internet of Things device.
[0158] The generation module 23 is used to perform cross-modal feature space alignment based on the image texture data and the sound spectrum data using a feature-level fusion classifier to generate a joint semantic feature vector representing the state of the monitored device.
[0159] The construction module 24 is used to construct a semantic knowledge graph based on the joint semantic feature vector.
[0160] The dynamic optimization module 25 is used to dynamically optimize the feature weights of cross-domain IoT devices in the semantic knowledge graph according to the level values, so as to generate cross-domain IoT device collaborative control instructions based on the dynamically optimized feature weights.
[0161] Figure 2 The cross-domain IoT device intelligent collaboration system based on semantic knowledge graph can be executed Figure 1 The implementation principle and technical effects of the cross-domain IoT device intelligent collaboration method based on semantic knowledge graph described in the illustrated embodiment will not be repeated here. The specific manner in which each module and unit performs operations in the cross-domain IoT device intelligent collaboration system based on semantic knowledge graph in the above embodiment has been described in detail in the embodiment of the method and will not be elaborated here.
[0162] In one possible design, Figure 2 The cross-domain IoT device intelligent collaboration system based on semantic knowledge graph of the embodiment shown can be implemented as a computing device, such as Figure 3 As shown, the computing device may include a storage component 31 and a processing component 32 .
[0163] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 32 .
[0164] The processing component 32 is used to: convert the mechanical vibration energy of the monitored device into an electrical energy signal and synchronously generate a corresponding level value; adjust the sampling frequency of the cross-domain Internet of Things device according to the level value, and after the sampling frequency adjustment is completed, obtain the image texture data and sound spectrum data of the monitored device collected by the cross-domain Internet of Things device; based on the image texture data and the sound spectrum data, use a feature-level fusion classifier to perform cross-modal feature space alignment to generate a joint semantic feature vector representing the state of the monitored device; construct a semantic knowledge graph based on the joint semantic feature vector; dynamically optimize the feature weights of the cross-domain Internet of Things devices in the semantic knowledge graph according to the level value, so as to generate cross-domain Internet of Things device collaborative control instructions based on the dynamically optimized feature weights.
[0165] The processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above method.
[0166] The storage component 31 is configured to store various types of data to support operations on the terminal. The storage component can be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as random access memory (RAM), static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0167] Of course, a computing device may also include other components, such as input / output interfaces, display components, communication components, etc.
[0168] The input / output interface provides an interface between the processing component and the peripheral interface module, which can be an output device, an input device, etc.
[0169] The communication component is configured to facilitate, among other things, wired or wireless communications between the computing device and other devices.
[0170] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. In this case, the computing device can refer to a cloud server, and the above-mentioned processing components, storage components, etc. can be basic server resources rented or purchased from the cloud computing platform.
[0171] The present application also provides a computer storage medium storing a computer program, wherein the computer program can achieve the above-mentioned Figure 1 The illustrated embodiment is a cross-domain IoT device intelligent collaboration method based on a semantic knowledge graph.
[0172] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0173] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0174] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0175] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A cross-domain IoT device intelligent collaboration method based on semantic knowledge graph, characterized by: include: Convert the mechanical vibration energy of the monitored equipment into electrical energy signals and generate corresponding level values synchronously; Adjusting the sampling frequency of the cross-domain IoT device according to the level value, and obtaining the image texture data and sound spectrum data of the monitored device collected by the cross-domain IoT device after the sampling frequency adjustment is completed; Based on the image texture data and the sound spectrum data, a feature-level fusion classifier is used to perform cross-modal feature space alignment to generate a joint semantic feature vector representing the state of the monitored device; Constructing a semantic knowledge graph based on the joint semantic feature vector; According to the level value, the feature weights of the cross-domain Internet of Things devices in the semantic knowledge graph are dynamically optimized to generate cross-domain Internet of Things device collaborative control instructions based on the dynamically optimized feature weights.
2. The method according to claim 1, characterized in that The method of performing cross-modal feature space alignment based on the image texture data and the sound spectrum data using a feature-level fusion classifier to generate a joint semantic feature vector representing the state of the monitored device includes: Performing a mapping operation on the image texture data input into an image projection function to generate a projected image feature, and performing a mapping operation on the sound spectrum data input into a sound projection function to generate a projected sound feature; Based on the projected image features and the projected sound features, performing an inter-modality similarity calculation operation in a shared latent space to generate an inter-modality similarity measurement value set; Based on the inter-modal similarity measurement value set, a feature-level fusion classifier is used to perform an association mapping construction operation to generate a cross-modal feature association mapping table; Reconstructing features according to the cross-modal feature association mapping table to generate feature slices of unified dimension; The feature slices are input into a bidirectional feature interaction module for multiple rounds of aggregation to generate a joint semantic feature vector representing the state of the monitored device.
3. The method according to claim 2, characterized in that The step of reconstructing features according to the cross-modal feature association mapping table to generate feature slices of uniform dimensions includes: Performing a block operation on the projected image features to generate an image feature subsegment set, and performing a segmentation operation on the projected sound features to generate a sound feature subsegment set; Performing a key-value matching operation based on the image feature subsegment set, the sound feature subsegment set, and the cross-modal feature association mapping table to generate a matching feature pair set; Performing a cross-aggregation operation on the set of matching feature pairs to generate a fused feature unit sequence; Based on the image feature subsegment set and the sound feature subsegment set, performing an unmatched identifier extraction operation to generate an unmatched image identifier set and an unmatched sound identifier set; performing an interpolation compensation operation based on the unmatched image identifier set and the unmatched sound identifier set to generate a compensated image feature unit and a compensated sound feature unit; A multi-dimensional structure reorganization operation is performed on the fused feature unit sequence, the compensated image feature unit, and the compensated sound feature unit to generate feature slices of uniform dimension.
4. The method according to claim 2, characterized in that The method of performing an association mapping construction operation based on the inter-modal similarity measurement value set and utilizing a feature-level fusion classifier to generate a cross-modal feature association mapping table includes: In the feature-level fusion classifier, a threshold screening operation is performed on the set of inter-modality similarity measures to generate a set of candidate mapping pairs; Performing a feature index grouping operation on the projected image features and the projected sound features to generate an image feature index group and a sound feature index group; Performing bidirectional association constraints on the image feature index group and the sound feature index group to generate a valid mapping rule set; Perform rule verification on the candidate mapping pair set and a preset valid mapping rule set to generate a verified mapping pair; A key-value table construction operation is performed on the verified mapping pairs to generate a cross-modal feature association mapping table.
5. The method according to claim 1, wherein The step of constructing a semantic knowledge graph based on the joint semantic feature vector includes: Extracting state nodes from the joint semantic feature vector to generate a device state node set; Performing a physical relationship modeling operation on the device state node set to generate a physical topology relationship table; Perform edge connection operations based on the physical topology relationship table to generate an initial knowledge graph skeleton; performing eigenvalue analysis on the joint semantic feature vector to generate a vibration eigenvalue set; Injecting the vibration feature value set into the initial knowledge graph skeleton to assign edge weights and generate a weighted knowledge graph; A spatiotemporal constraint binding operation is performed on the weighted knowledge graph to obtain a semantic knowledge graph.
6. The method according to claim 5, characterized in that Injecting the vibration feature value set into the initial knowledge graph skeleton to assign edge weights to generate a weighted knowledge graph includes: Performing an edge type recognition operation on the initial knowledge graph skeleton to generate a spatial edge set and a temporal edge set; performing a feature dimension separation operation on the vibration feature value set to generate an amplitude feature set and an attenuation feature set; Performing feature assignment operations on the spatial edge set and the amplitude feature set, and on the temporal edge set and the attenuation feature set, respectively, to generate a spatial edge amplitude mapping table and a temporal edge attenuation mapping table; Based on the spatial edge amplitude mapping table and the temporal edge attenuation mapping table, performing spatial weight and temporal weight calculation operations to generate spatial edge weight values and edge temporal weight values; The spatial edge weight value and the temporal edge weight value are injected into the initial knowledge graph skeleton to update the weights and generate a weighted knowledge graph.
7. The method according to claim 1, characterized in that The dynamically optimizing the feature weights of the cross-domain IoT devices in the semantic knowledge graph according to the level values, and generating cross-domain IoT device collaborative control instructions according to the dynamically optimized feature weights, includes: performing vibration mode analysis on the level value to generate a dominant vibration mode identifier; Performing an edge weight extraction operation on the semantic knowledge graph to generate an original weight distribution table; Inputting the dominant vibration pattern identifier into the original weight distribution table for pattern matching to generate a matching weight subset, and inputting the matching weight subset into the semantic knowledge graph for weight update operation to generate an optimized knowledge graph; Traversing the optimized knowledge graph to retrieve control rules and generate an initial control instruction set; Perform device collaboration constraint operations on the initial control instruction set to generate cross-domain IoT device collaboration control instructions.
8. A cross-domain IoT device intelligent collaboration system based on semantic knowledge graph, characterized by: include: The conversion module is used to convert the mechanical vibration energy of the monitored equipment into an electrical energy signal and synchronously generate a corresponding level value; an acquisition module, configured to adjust a sampling frequency of the cross-domain IoT device according to the level value, and after the sampling frequency adjustment is completed, acquire image texture data and sound spectrum data of the monitored device collected by the cross-domain IoT device; A generation module, configured to perform cross-modal feature space alignment based on the image texture data and the sound spectrum data using a feature-level fusion classifier to generate a joint semantic feature vector representing the state of the monitored device; A construction module, configured to construct a semantic knowledge graph based on the joint semantic feature vector; A dynamic optimization module is used to dynamically optimize the feature weights of cross-domain IoT devices in the semantic knowledge graph according to the level value, so as to generate cross-domain IoT device collaborative control instructions based on the dynamically optimized feature weights.
9. A computing device, characterized in that It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement the cross-domain Internet of Things device intelligent collaboration method based on semantic knowledge graph as described in any one of claims 1 to 7.
10. A computer storage medium, characterized in that A computer program is stored, and when the computer program is executed by a computer, the cross-domain Internet of Things device intelligent collaboration method based on the semantic knowledge graph as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Multi-modal knowledge graph method based on power grid dispatching
CN117171358A
Practical training method and system based on multi-mode Internet of Things perception and virtual-real symbiosis
CN118862648A
Electric power cross-modal knowledge fusion multi-agent cooperative processing method and system
CN119477235A
Complex equipment fault diagnosis method and system based on multi-modal knowledge graph
CN120217264A
Knowledge graph construction method for multi-modal data
CN120296652A
Cited By
Rate matching puncturing method based on semantic importance
CN121441450A
Industrial operation data visual display method and system based on artificial intelligence platform
CN121478877A
Process collaborative optimization method and system for textile heterogeneous equipment group
CN121809923A