Artificial Intelligence-Based Computer Room Status Data Processing Method, Device, and Storage Medium
By preprocessing, feature extraction, alignment and fusion of multimodal data and timing data in the computer room, the judgment errors caused by data fragmentation in the computer room are solved, and the accuracy and reliability of safety incident judgments are improved.
Patent Information
- Application Number
- CN202510248130.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-03-04
AI Technical Summary
In the prior art, multimodal data and timing data analysis of the computer room are separated, resulting in errors in determining the equipment status and inability to accurately judge the computer room safety incident in the case of sensor failure or abnormality.
Through an artificial intelligence-based method, the multimodal data and timing data in the computer room are preprocessed, feature extraction, feature alignment and fusion features are generated to determine the current status of the computer room.
It improves the accuracy of machine room safety incident judgment, reduces false alarms and missed reports, and provides operation and maintenance personnel with more reliable decision-making basis.
Smart Images

Figure CN119720111B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method, device and storage medium for processing computer room status data based on artificial intelligence. Background Art
[0002] In the context of the rapid development of big data and artificial intelligence, the processing and analysis of multi-modal data and time-series data have become increasingly important. Multi-modal data usually includes various types of data such as images, audio, and text, while time-series data emphasizes the change of data over time.
[0003] In the traditional computer room operation and maintenance scenario, multi-modal data mainly refers to video data and audio data in the computer room, etc., while time-series data mainly refers to status monitoring data obtained by various sensors in the computer room. By analyzing the multi-modal data and time-series data of the computer room, the change trend of the computer room environment and equipment status can be intuitively understood. Existing data analysis methods often separate these two types of data, and the two are not linked. That is, multi-modal data such as video and audio is usually only used for post-event viewing, while time-series data is only used for analysis when the sensors are working properly, and often due to sensor failures or abnormalities, wrong judgments are made on the equipment status. Summary of the Invention
[0004] The present application provides a method, device and storage medium for processing computer room status data based on artificial intelligence, which can improve the accuracy of judging computer room security events.
[0005] On the one hand, the present application provides a method for processing computer room status data based on artificial intelligence, the method comprising:
[0006] Preprocessing multi-modal data and time-series data related to the computer room respectively to obtain preprocessed multi-modal data and preprocessed time-series data;
[0007] Performing feature extraction on the preprocessed multi-modal data and preprocessed time-series data respectively based on an artificial intelligence algorithm to obtain multi-modal data features and time-series data features;
[0008] Aligning the multi-modal data features and time-series data features;
[0009] Fusing the preprocessed multi-modal data and preprocessed time-series data based on the multi-modal data features and time-series data features after feature alignment to obtain fusion features;
[0010] Performing data analysis according to the fusion features to determine the current status of the computer room.
[0011] On the other hand, the present application provides a device for processing computer room status data based on artificial intelligence, the device comprising:
[0012] A preprocessing module is used to preprocess the multimodal data and time series data of the relevant computer room to obtain preprocessed multimodal data and preprocessed time series data respectively;
[0013] A feature extraction module is used to extract features from the preprocessed multimodal data and the preprocessed time series data based on an artificial intelligence algorithm to obtain multimodal data features and time series data features respectively;
[0014] An alignment module, configured to perform feature alignment on the multimodal data features and the time series data features;
[0015] A fusion module, configured to fuse the pre-processed multimodal data and the pre-processed time series data based on the feature-aligned multimodal data features and the time series data features to obtain fused features;
[0016] The determination module is used to perform data analysis based on the fusion features to determine the current status of the computer room.
[0017] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the technical solution of the above-mentioned method for processing computer room status data based on artificial intelligence are implemented.
[0018] In a fourth aspect, the present application provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of the technical solution of the above-mentioned method for processing computer room status data based on artificial intelligence.
[0019] From the technical solution provided by the present application, it can be seen that after extracting features from the pre-processed multimodal data and the pre-processed time series data based on the artificial intelligence algorithm, and obtaining the multimodal data features and time series data features respectively, the multimodal data features and time series data features are aligned, and based on the multimodal data features and time series data features after feature alignment, the pre-processed multimodal data and the pre-processed time series data are fused to obtain fused features. Compared with the existing data analysis methods that often separate these two types of data and do not link the two, the technical solution of the present application improves the accuracy of the judgment of computer room security events through the fusion of data features, reduces false alarms and omissions caused by a single data source, and provides a more reliable decision-making basis for operation and maintenance personnel. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0021] Figure 1 is a flowchart of a method for processing computer room status data based on artificial intelligence provided by an embodiment of the present application;
[0022] Figure 2 is a schematic structural diagram of a device for processing computer room status data based on artificial intelligence provided by an embodiment of the present application;
[0023] Figure 3 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Specific embodiments
[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0025] In this specification, adjectives such as first and second can only be used to distinguish one element or action from another element or action, and do not necessarily require or imply any actual such relationship or order. Where circumstances permit, reference to an element or component or step (etc.) should not be construed as being limited to only one of the elements, components, or steps, but can be one or more of the elements, components, or steps, etc.
[0026] In this specification, for the convenience of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual proportional relationships.
[0027] In traditional computer room operation and maintenance scenarios, multi-modal data mainly refers to video data, audio data, etc. in the computer room, while time-series data mainly refers to status monitoring data obtained by various sensors in the computer room. By analyzing the multi-modal data and time-series data of the computer room, the changing trends of the computer room environment and equipment status can be intuitively understood. Existing data analysis methods often separate these two types of data, and the two are not linked. That is, multi-modal data such as video and audio is usually only used for post-event viewing, while time-series data is only used for analysis when the sensors are working properly, and wrong judgments on the equipment status are often caused by sensor failures or abnormalities.
[0028] In view of the above problems in the prior art, the present application proposes an artificial intelligence-based method for processing computer room status data, and its flowchart is as shown in the appendix Figure 1 as shown, which mainly includes steps S101 to S105, and is described in detail as follows:
[0029] Step S101: Preprocess the multimodal data and time series data related to the computer room respectively to obtain the preprocessed multimodal data and the preprocessed time series data.
[0030] In an embodiment of the present application, the multimodal data related to the computer room includes data such as videos, images, or audios of the computer room workplace, and the time series data related to the computer room includes data generated by some sensors in the computer room workplace (for example, temperature sensors, smoke sensors, monitoring devices, etc.). Preprocessing the multimodal data and time series data related to the computer room mainly includes denoising and unifying the scale of these data. For example, by means of standardization, normalization, etc., the multimodal data and time series data are adjusted to conform to the standard normal distribution and compressed to a specific range, and so on.
[0031] Step S102: Respectively extract features from the preprocessed multimodal data and the preprocessed time series data based on an artificial intelligence algorithm to obtain multimodal data features and time series data features.
[0032] If the preprocessed multimodal data is the preprocessed image, then as an embodiment of the present application, respectively extracting features from the preprocessed multimodal data and the preprocessed time series data based on an artificial intelligence algorithm to obtain multimodal data features and time series data features can be implemented through the following steps S1021 to step S1023, and the detailed description is as follows:
[0033] Step S1021: Extract the features of the preprocessed image based on a deep convolutional neural network to obtain an image data feature map.
[0034] The deep convolutional neural network can be models such as VGG16, ResNet, or EfficientNet. After these models are trained, their multiple convolutional layers and pooling layers are used to perform operations such as convolution and pooling to extract the high-dimensional features of the preprocessed image and obtain an image data feature map.
[0035] Step S1022: Slide a sliding window on the image data feature map to obtain multiple targets on the preprocessed image.
[0036] In a complex environment such as a computer room, the object detection model can effectively identify and locate key objects in an image (such as devices, indicator lights, alarms, etc.). Traditional convolutional neural networks are no longer sufficient for single classification tasks because multiple objects in the image need to be detected and their positions accurately marked. Therefore, in the embodiments of the present application, after obtaining the image data feature map, multiple objects on the preprocessed image can be obtained by sliding a window over the image data feature map. As an embodiment of the present application, obtaining multiple objects on the preprocessed image by sliding a window over the image data feature map can be achieved through steps S1 to S3, which are described in detail as follows:
[0037] Step S1: Slide a window over the image data feature map to generate candidate boxes for multiple objects.
[0038] The sliding window slides along each position of the feature map (i.e., each pixel position), and multiple candidate boxes (anchor boxes) are generated at each position. These candidate boxes have different scales and aspect ratios, aiming to adapt to objects of different sizes and shapes. Assume the size of the sliding window is (usually ), and a window is generated on the feature map each time it slides. The window size usually corresponds to the receptive field of the convolutional layer. Multiple candidate boxes with different scales and aspect ratios are generated at each window position, and these candidate boxes are called anchor boxes. Specifically, assume that at a certain position in the feature map, a window is slid, and K anchor boxes are generated at this position. The size and ratio of each anchor box can be determined by predefined scales and aspect ratios. For example, 3 scales and 3 aspect ratios can be defined, then candidate boxes for objects are generated at each position.
[0039] Step S2: Perform classification and bounding box regression operations on the candidate boxes of multiple objects to obtain the classification score and regression value of each object.
[0040] At each sliding window position, the classification score and bounding box regression value of each anchor box need to be calculated, which can be achieved by performing classification and bounding box regression operations on the candidate boxes of multiple objects. Among them, the classification operation is to judge whether each candidate box contains an object, usually represented by a binary classification (background / foreground) model. For each candidate box of an object, a classification score is output, indicating the probability that this anchor box contains an object. For the k-th anchor box, the classification score represents the probability that this box is the foreground:
[0041]
[0042] Among them, are the features extracted from the convolutional layer, and are the parameters of the classification layer, is the sigmoid function, which is used to map the output to a probability value.
[0043] The bounding box regression operation is used to adjust the position of each anchor box to better fit the true bounding box of the target. For each anchor box , a regression value is output, indicating the adjustment amount required for the anchor box, including: the offset of the center coordinates and the adjustment of the size . For the k-th anchor box, the bounding box regression amount is calculated through the following regression function:
[0044]
[0045] Among them, and are the parameters of the regression layer, is the convolutional feature, represents the adjustment amount required to generate the target box from the anchor box.
[0046] Step S3: Use the non-maximum suppression algorithm to remove the duplicate candidate boxes in the candidate boxes, and obtain the category and position of the target in the preprocessed image.
[0047] Specifically, the boxes with high overlap are removed by calculating the intersection over union (IoU) between each pair of candidate boxes, and the box with the highest score is retained. The calculation formula for the intersection over union is:
[0048]
[0049] Among them, and are two candidate boxes, represents the area of the box, represents the ratio of the overlapping area of the two boxes to the union area.
[0050] Step S1023: Use the LSTM model to extract features from the preprocessed time series data to obtain the time series data features.
[0051] Specifically, an LSTM model is used to extract features from the preprocessed time-series data. The time-series data features obtained can be as follows: Input the preprocessed time-series data into the LSTM (Long Short-Term Memory) model and initialize the state of the LSTM model; at each time step, calculate the forget gate, input gate, candidate memory cell, updated memory cell of the LSTM model, and calculate the output gate, and calculate the hidden state at the current moment; according to the task requirements, use the hidden state at the final moment or the entire hidden state sequence as the time-series data features.
[0052] If the preprocessed multi-modal data is the preprocessed video stream, then as another embodiment of the present application, based on the artificial intelligence algorithm, feature extraction is respectively performed on the preprocessed multi-modal data and the preprocessed time-series data, and the multi-modal data features and the time-series data features can be obtained through the following steps S’1021 to step S’1025, and the detailed description is as follows:
[0053] Step S’1021: Perform a three-dimensional convolution operation on the preprocessed video stream. Among them, the convolution kernel of the three-dimensional convolution is set in the spatial dimension as and in the time dimension with frames as the sliding step.
[0054] It should be noted that the video stream can be regarded as a four-dimensional tensor, and the convolution kernel size of the three-dimensional convolution is where the spatial dimension is used to extract the spatial features of a single video frame of the video stream (such as the device contour, indicator light color, etc.), and the time dimension is used to capture dynamic changes across multiple consecutive frames (such as fan speed changes, temperature diffusion trends, etc.). The sliding step (Stride) in the three-dimensional convolution operation is equal to the time dimension of the convolution kernel size indicating that after processing every frames, the convolution kernel slides to the next position on the time axis.
[0055] In the convolution operation (two-dimensional or three-dimensional convolution operation), the selection of the convolution kernel size is closely related to factors such as computational efficiency, receptive field, and task adaptability. Specifically, a larger-size convolution kernel (such as 5×5) will increase the number of parameters and computational cost, and is prone to overfitting. On the contrary, a smaller-size convolution kernel (such as 1×1) is mainly used for channel dimension adjustment and cannot effectively extract spatial features. In the feature extraction task, edge continuity and local texture details are crucial. A convolution kernel of a suitable size can effectively capture the pixel correlation within the local neighborhood (such as edge direction, color gradient), while avoiding the blurring problem introduced by a large convolution kernel. Based on the consideration of balancing computational efficiency and receptive field and adapting to the feature extraction of the present application, the above convolution kernel size in the spatial dimension Specifically, it can be 3×3, and for the time dimension it can take 5. Performing a three-dimensional convolution operation on the preprocessed video stream can be expressed as:
[0056]
[0057] where W is the convolution kernel weight, is the video frame of the preprocessed video stream, b is the bias term, c is the output channel index, is the output of the three-dimensional convolution operation.
[0058] Step S’1022: Generate an activated feature map by passing the output of the three-dimensional convolution operation through the LeakyReLU activation function.
[0059] Since the traditional activation function ReLU outputs 0 when the input is <0, resulting in a gradient of 0 in the negative region and the inability to update parameters during backpropagation, i.e., gradient disappearance. Therefore, to solve the above gradient disappearance problem, in the embodiments of this application, the output of the three-dimensional convolution operation is passed through the LeakyReLU activation function to generate an activated feature map, where the leakage factor of the LeakyReLU activation function is set to , and the expression of the LeakyReLU activation function is:
[0060]
[0061] The above leakage factor can be determined according to the relationship between the negative activation ratio of the current batch and the preset threshold. Specifically:
[0062]
[0063] where, is the base leakage factor, which can take a value of 0.2, is the Sigmoid function, N is the batch size, is the adjustment amplitude. The above expression indicates that if the negative activation ratio of the current batch is greater than the preset threshold, then (where, is the negative activation ratio), otherwise, is equal to the base leakage factor, i.e., .
[0064] Step S’1023: Perform global average pooling on the activated feature map to calculate the mean statistic of each channel.
[0065] Specifically, perform global average pooling on the activated feature map to calculate the mean statistic of each channel as:
[0066]
[0067] Among them, is the size of the feature map after activation, is the eigenvalue of the c-th channel at position .
[0068] Step S’1024: Input the mean statistic into the Softmax function to generate the normalized attention weight .
[0069] Specifically, , and
[0070] Among them, is the total number of channels of the feature map after activation, .
[0071] Step S’1025: Weight the channels of the feature map after activation according to the formula to obtain the enhanced feature map.
[0072] Step S103: Align the multi-modal data features and the time-series data features.
[0073] The features of different modalities not only have differences in "data form" (for example, images are two-dimensional data, while time-series data has time dependence), but also their representation methods and abstraction levels are different. For example, images are represented by pixel values and color spaces, mainly capturing spatial structure information, text is represented by words in the vocabulary and syntactic structures, mainly focusing on the grammar and semantic information of the language, while time-series data has continuity and time dependence in the time dimension, mainly capturing the changes and trends of the time series. Feature alignment helps the features of different modalities to match each other in the same space, enabling the information of each modality to complement and cooperate with each other. Therefore, after steps S101 and S102, it is necessary to align the multi-modal data features and the time-series data features. As an embodiment of the present application, the alignment of the multi-modal data features and the time-series data features can be achieved through steps S1031 to S1033, which are described in detail as follows:
[0074] Step S1031: Map the multi-modal data features and the time-series data features to a semantic space of the same dimension.
[0075] Suppose represents the multi-modal data features (for example, data features such as images, video streams, etc.), represents the time-series data features, then the multi-modal data features and the time-series data features can be mapped to a semantic space of the same dimension according to the following formula:
[0076]
[0077] Among them, and are weight matrices, and are bias terms, and are the mapped multi-modal data features and temporal data features respectively.
[0078] Step S1032: Calculate the similarity between the multi-modal data features and the temporal data features in the semantic space of the same dimension.
[0079] In the embodiments of the present application, the similarity between the multi-modal data features and the temporal data features in the semantic space of the same dimension can be the cosine similarity.
[0080] Step S1033: Adjust the weights and biases through an optimization algorithm to minimize the distance between the multi-modal data features and the temporal data features.
[0081] Specifically, adjusting the weights and biases through an optimization algorithm to minimize the distance between the multi-modal data features and the temporal data features can be achieved through steps S10331 to S10333, which are described in detail as follows:
[0082] Step S10331: Construct a loss function for measuring the alignment degree between different modal features.
[0083] To maximize the similarity between different modal features, the goal is to make similar sample pairs (such as matching images and temporal data) have a higher similarity, while different sample pairs (such as non-matching images and temporal data) have a lower similarity. Therefore, the loss function for measuring the alignment degree between different modal features can be defined by the contrastive loss, specifically as follows:
[0084]
[0085] Among them, is a label indicating whether the multi-modal data features and the temporal data features match, indicates a match, indicates a non-match, is the cosine similarity between the multi-modal data features and the temporal data features, is a threshold. When the sample pair is a non-match, the similarity cannot exceed this threshold. Usually takes a constant value, indicating that when the similarity of non-matching samples is lower than a certain value, they are considered "non-matching".
[0086] From From the expression: when , it is desired to minimize the difference in similarity, that is, to make their cosine similarity as close to 1 as possible; when , it is desired that the cosine similarity is less than , that is, to maximize the difference in similarity and ensure that they are correctly distinguished.
[0087] Step S10332: Optimize the weights and biases using the gradient descent algorithm.
[0088] Optimize the weight matrix and and the bias and The goal is to minimize the loss function , specifically as follows:
[0089] S1: Calculate the gradients of the weight matrix and and the bias and .
[0090] The core step of gradient descent is to calculate the gradients of the loss function with respect to the model parameters (weights and biases). If the selected optimization algorithm is Stochastic Gradient Descent (SGD), the update rules for the parameters are:
[0091]
[0092] where is the learning rate, which controls the step size of parameter update, and are the gradients of the loss function
[0093] function with respect to the weights, and are the gradients of the loss function with respect to the bias.
[0094] S2: Calculate the specific forms of the gradients.
[0095] To perform gradient calculation, partial derivatives of the loss function with respect to the parameters , , and are required, specifically:
[0096]
[0097] S3: Update the parameters , , and .
[0098] Once the gradient is calculated, the gradient descent algorithm can be used to update , , and parameters such as. Through the backpropagation algorithm, the weights and biases can be adjusted in each iteration to minimize the loss function and thus optimize semantic alignment.
[0099] Step S10333: Iterative training.
[0100] Repeatedly execute the above steps S10331 to S10333, calculate the gradient and update the parameters in each training iteration until the loss function converges to an optimal value.
[0101] As another embodiment of the present application, the feature alignment of multimodal data features and temporal data features can be achieved through steps S'1031 to S'1033, and the detailed description is as follows:
[0102] Step S'1031: Perform dimensionality compression on multimodal data features and temporal data features.
[0103] Since there are significant differences between multimodal data (such as video streams, images, etc.) and temporal data (such as sensor data), the dimensions of multimodal data features and temporal data features may be different, or even vary greatly. For example, multimodal data features may be multimodal features such as a shape, color, or texture feature with a dimension of 256 for video stream data, while temporal data features may be sensor temporal features such as temperature, humidity, or voltage with a dimension of 64. Considering that the prerequisite for calculating the dynamic time warping bending path of the multimodal feature sequence and the temporal feature sequence is that the dimensions of the multimodal data features and the temporal data features need to be the same, dimensionality compression can be performed on the multimodal data features and the temporal data features. Specifically, PCA dimensionality reduction can be performed on the multimodal features, retaining the principal components with a variance contribution rate exceeding a preset threshold, such as 85%, and autoencoder compression can be performed on the temporal data features so that the compressed dimension is equal to the principal component dimension of the multimodal data features.
[0104] Step S'1032: Calculate the dynamic time warping bending path of the multimodal feature sequence and the temporal feature sequence to obtain the optimal bending path.
[0105] Assume that the multimodal feature sequence is , and the temporal feature sequence is . Calculate the dynamic time warping bending path of the multimodal feature sequence and the temporal feature sequence to obtain the optimal bending path, that is, according to the formula:
[0106]
[0107] Find the optimal bending path , such that the cumulative distance of the path is minimized, where is a feature distance metric (e.g., Euclidean distance).
[0108] Step S’1033: Perform cubic spline interpolation on the shorter sequence among the multimodal data features and the time series data features with the same principal component dimension along the optimal bending path, where the interpolation point density of the cubic spline interpolation is dynamically adjusted according to the sampling frequency difference between the two sequences, and the adjustment ratio is the reciprocal of the frequency difference.
[0109] Assume that in the multimodal feature sequence and the time series feature sequence , , then interpolate the multimodal feature sequence, otherwise, interpolate the time series feature sequence. If the sampling frequency of the multimodal data is , and the time series data is , then the interpolation point density of the cubic spline interpolation can be calculated according to the following formula:
[0110]
[0111] where is a smoothing coefficient to prevent division by zero, and the actual number of interpolation points , and L is the original length of the sequence to be interpolated.
[0112] As another embodiment of the present application, the feature alignment of the multimodal data features and the time series data features can be achieved through steps S’’1031 to S’’1036, and the detailed description is as follows:
[0113] Step S’’1031: Map the multimodal data feature X to a query matrix through a linear transformation, and map the time series data feature Y to a key matrix K and a value matrix V.
[0114] Assume that the multimodal data feature (m time steps, with a dimension of ), and the time series data feature (n time steps, with a dimension of ). Map the time series data feature Y to the key matrix K and the value matrix V according to the following linear transformation formula:
[0115] , where , , .
[0116] where is the dimension of the attention head.
[0117] Step S''1032: Calculate the attention score matrix , where is a relative position encoding matrix dynamically adjusted based on the timestamp interval.
[0118] In the embodiment of the present application, the relative position encoding matrix , and its element represents the relative position weight between time step and . A scheme for dynamically adjusting based on the timestamp interval is:
[0119] where is a decay function, .
[0120] Step S''1033: Perform Softmax normalization on the attention score matrix S and weight it with V to obtain the preliminary alignment features .
[0121] Step S''1034: Split the preliminary alignment features Z into multiple attention heads.
[0122] Taking the example of splitting into 8 attention heads with a dimension of 32 for each attention head, independently calculate the output of each attention head as:
[0123] ,
[0124] where
[0125] Step S''1035: Perform LayerNorm normalization on the output of each attention head to obtain .
[0126] Specifically, , where ,
[0127] , and can be loaded from a predefined parameter table according to the type of computer room equipment.
[0128] Step S''1036: Concatenate the outputs of all attention heads and generate the final alignment features through residual connection.
[0129] Specifically, concatenate the outputs of all attention heads according to and perform residual connection Generate the final aligned feature, where is the multimodal data feature, is the output projection matrix, .
[0130] Step S104: Based on the multimodal data feature and the time series data feature after feature alignment, fuse the preprocessed multimodal data and the preprocessed time series data to obtain a fused feature.
[0131] As an embodiment of the present application, fusing the preprocessed multimodal data and the preprocessed time series data based on the multimodal data feature and the time series data feature after feature alignment to obtain a fused feature can be implemented through steps S1041 to S1043, and the detailed description is as follows:
[0132] Step S1041: Calculate the similarity between the multimodal data feature and the time series data feature.
[0133] Specifically, the dot product can be used to calculate the similarity between the multimodal data feature and the time series data feature :
[0134] Where is the feature vector at the position in the multimodal data feature map, C is the channel of the multimodal data feature map (assuming the multimodal data is an image), is the feature representation of the k-th word with the feature dimension D in the time series data.
[0135] Step S1042: Based on the similarity between the multimodal data feature and the time series data feature, use the softmax function to calculate the attention weights between the multimodal data and the time series data.
[0136] Specifically, the attention weights between the multimodal data and the time series data can be calculated by the following formula:
[0137]
[0138] Where L is the length of the time series data corresponding to the time series data feature.
[0139] Step S1043: Calculate the weighted fusion representation between the multimodal data feature and the time series data feature through the attention weights between the multimodal data and the time series data.
[0140] Specifically, calculating the weighted fusion representation between the multimodal data feature and the time series data feature through the attention weights between the multimodal data and the time series data can be implemented by the following formula:
[0141]
[0142] From the weighted fusion representation of the expression, in fact, the multi-modal data features are associated with each word in the weighted sum and the time-series data, generating a spatially aligned representation that combines the common information of the multi-modal data and the time-series data.
[0143] As another embodiment of the present application, based on the multi-modal data features and time-series data features after feature alignment, fusing the preprocessed multi-modal data and preprocessed time-series data to obtain the fusion features can be implemented through steps S'1041 to S'1043, and the detailed description is as follows:
[0144] Step S'1041: Map the multi-modal data features and time-series data features after feature alignment to the first low-dimensional embedding representation and the second low-dimensional embedding representation respectively.
[0145] Step S'1042: Through a shared mapping function, perform cross mapping on the first low-dimensional embedding representation and the second low-dimensional embedding representation to obtain the cross-mapped embedded features.
[0146] Step S'1043: Reconstruct the cross-mapped embedded features back to the feature space of the original modality.
[0147] Considering that the dimension of the fusion features may be too large, even causing the "curse of dimensionality", therefore, after obtaining the fusion features through step S104, the fusion features can also be dimensionally reduced.
[0148] Step S105: Perform data analysis based on the fusion features to determine the current state of the computer room.
[0149] Specifically, as an embodiment of the present application, performing data analysis based on the fusion features to determine the current state of the computer room can be to use classification algorithms, such as support vector machine (SVM), random forest (Random Forest), etc., to analyze the fusion features, judge the security state of the computer room as different levels such as normal, warning, or danger, and learn the feature patterns of different security states according to the training data, and accurately classify the new data. As an embodiment of the present application, performing data analysis based on the fusion features to determine the current state of the computer room can also be to use anomaly detection algorithms, for example, isolation forest (Isolation Forest), local outlier factor (Local Outlier Factor, LOF), etc., to identify the abnormal patterns in the data by analyzing the fusion features, judge whether there is a security event, and send a warning signal when detecting abnormal data points that are significantly different from the normal data pattern.
[0150] From the above appendix Figure 1 As can be seen from the above-described example of the artificial intelligence-based computer room status data processing method, after respectively performing feature extraction on the preprocessed multi-modal data and preprocessed time-series data based on an artificial intelligence algorithm to obtain multi-modal data features and time-series data features, by aligning the multi-modal data features and time-series data features, based on the aligned multi-modal data features and time-series data features, the preprocessed multi-modal data and preprocessed time-series data are fused to obtain fused features. Compared with existing data analysis methods that often separate these two types of data and do not link them together, the technical solution of this application improves the accuracy of judging computer room security events through the fusion of data features, reduces false alarms and missed alarms caused by a single data source, and provides a more reliable decision-making basis for operation and maintenance personnel.
[0151] Please refer to the appendix Figure 2 which is an artificial intelligence-based computer room status data processing device provided by an embodiment of this application. The device may include a preprocessing module 201, a feature extraction module 202, an alignment module 203, a fusion module 204, and a determination module 205, which are described in detail as follows:
[0152] The preprocessing module 201 is configured to respectively preprocess multi-modal data and time-series data related to the computer room to obtain preprocessed multi-modal data and preprocessed time-series data;
[0153] The feature extraction module 202 is configured to respectively perform feature extraction on the preprocessed multi-modal data and preprocessed time-series data based on an artificial intelligence algorithm to obtain multi-modal data features and time-series data features;
[0154] The alignment module 203 is configured to align the multi-modal data features and time-series data features;
[0155] The fusion module 204 is configured to fuse the preprocessed multi-modal data and preprocessed time-series data based on the aligned multi-modal data features and time-series data features to obtain fused features;
[0156] The determination module 205 is configured to perform data analysis based on the fused features to determine the current status of the computer room.
[0157] From the above appendix Figure 2It can be seen from the exemplary AI-based computer room status data processing device that after feature extraction is respectively performed on the preprocessed multi-modal data and preprocessed time-series data based on an AI algorithm, and multi-modal data features and time-series data features are respectively obtained, by aligning the multi-modal data features and time-series data features, the preprocessed multi-modal data and preprocessed time-series data are fused based on the aligned multi-modal data features and time-series data features to obtain fused features. Compared with the existing data analysis methods that often separate these two types of data and do not link them together, the technical solution of this application improves the accuracy of judging computer room security events through the fusion of data features, reduces false alarms and missed alarms caused by a single data source, and provides a more reliable decision-making basis for operation and maintenance personnel.
[0158] Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 3 shown, the electronic device 3 of this embodiment mainly includes: a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30, such as a program for the AI-based computer room status data processing method. When the processor 30 executes the computer program 32, it implements the steps in the above-mentioned embodiment of the AI-based computer room status data processing method, such as Figure 1 the steps S101 to S105 shown. Alternatively, when the processor 30 executes the computer program 32, it implements the functions of each module / unit in the above-mentioned device embodiments, such as Figure 2 the functions of the preprocessing module 201, the feature extraction module 202, the alignment module 203, the fusion module 204, and the determination module 205 shown.
[0159] Exemplarily, the computer program 32 for the method of processing computer room status data based on artificial intelligence mainly includes: respectively preprocessing multimodal data and time-series data related to the computer room to obtain preprocessed multimodal data and preprocessed time-series data; respectively extracting features from the preprocessed multimodal data and preprocessed time-series data based on artificial intelligence algorithms to obtain multimodal data features and time-series data features; performing feature alignment on the multimodal data features and time-series data features; based on the feature-aligned multimodal data features and time-series data features, fusing the preprocessed multimodal data and preprocessed time-series data to obtain fused features; performing data analysis based on the fused features to determine the current state of the computer room. The computer program 32 can be divided into one or more modules / units, and one or more modules / units are stored in the memory 31 and executed by the processor 30 to complete this application. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 32 in the electronic device 3. For example, the computer program 32 can be divided into the functions of a preprocessing module 201, a feature extraction module 202, an alignment module 203, a fusion module 204, and a determination module 205 (modules in the virtual device). The specific functions of each module are as follows: The preprocessing module 201 is used to respectively preprocess multimodal data and time-series data related to the computer room to obtain preprocessed multimodal data and preprocessed time-series data; the feature extraction module 202 is used to respectively extract features from the preprocessed multimodal data and preprocessed time-series data based on artificial intelligence algorithms to obtain multimodal data features and time-series data features; the alignment module 203 is used to perform feature alignment on the multimodal data features and time-series data features; the fusion module 204 is used to, based on the feature-aligned multimodal data features and time-series data features, fuse the preprocessed multimodal data and preprocessed time-series data to obtain fused features; the determination module 205 is used to perform data analysis based on the fused features to determine the current state of the computer room.
[0160] The electronic device 3 may include but is not limited to the processor 30 and the memory 31. Those skilled in the art can understand that Figure 3 merely examples of the electronic device 3, which do not constitute a limitation on the electronic device 3, may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, etc.
[0161] The so-called processor 30 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0162] The memory 31 may be an internal storage unit of the electronic device 3, such as the hard disk or memory of the electronic device 3. The memory 31 may also be an external storage device of the electronic device 3, such as a plug-in hard disk equipped on the electronic device 3, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 31 may also include both the internal storage unit of the electronic device 3 and the external storage device. The memory 31 is used to store computer programs and other programs and data required by the electronic device. The memory 31 may also be used to temporarily store data that has been output or is to be output.
[0163] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above device can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0164] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0165] Those of ordinary skill in the art will recognize that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician may use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0166] In the embodiments provided in this application, it should be understood that the disclosed apparatus / devices and methods can be implemented in other ways. For example, the apparatus / device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another apparatus, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the apparatus or unit can be in electrical, mechanical or other forms.
[0167] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0168] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0169] When the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present application, it can also be completed by instructing relevant hardware through a computer program. The computer program for the method of processing computer room status data based on artificial intelligence can be stored in a storage medium. When the computer program is executed by a processor, it can implement the steps of the above-described various method embodiments, that is, preprocess the multimodal data and time-series data related to the computer room respectively to obtain the preprocessed multimodal data and preprocessed time-series data; perform feature extraction on the preprocessed multimodal data and preprocessed time-series data respectively based on artificial intelligence algorithms to obtain multimodal data features and time-series data features respectively; perform feature alignment on the multimodal data features and time-series data features; based on the feature-aligned multimodal data features and time-series data features, fuse the preprocessed multimodal data and preprocessed time-series data to obtain fused features; perform data analysis based on the fused features to determine the current status of the computer room. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The storage medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, RandomAccess Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the storage medium does not include electrical carrier signals and telecommunication signals.
[0170] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application. The above-described specific implementation manners have further elaborated the purpose, technical solutions, and beneficial effects of the present application. It should be understood that the above is only the specific implementation manners of the present application and is not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application should all be included in the protection scope of the present invention.
Claims
1. An artificial intelligence-based method for processing computer room status data, characterized in that, The method includes: Preprocessing the multimodal data and time-series data related to the computer room respectively to obtain preprocessed multimodal data and preprocessed time-series data; Performing feature extraction on the preprocessed multimodal data and preprocessed time-series data respectively based on an artificial intelligence algorithm to obtain multimodal data features and time-series data features; Aligning the multimodal data features and time-series data features; Fusing the preprocessed multimodal data and preprocessed time-series data based on the multimodal data features and time-series data features after feature alignment to obtain fusion features; Performing data analysis according to the fusion features to determine the current state of the computer room; The aligning the multimodal data features and time-series data features includes: mapping the multimodal data features and time-series data features to a semantic space of the same dimension; calculating the similarity between the multimodal data features and time-series data features in the semantic space of the same dimension, where the similarity includes weights and biases; adjusting the weights and biases through an optimization algorithm to minimize the distance between the multimodal data features and time-series data features; The adjusting the weights and biases through an optimization algorithm to minimize the distance between the multimodal data features and time-series data features includes: constructing a loss function for measuring the alignment degree between different modal features; using the gradient descent algorithm to optimize the weights and biases; performing iterative training, calculating the gradient and updating the parameters in each training iteration until the loss function converges to an optimal value; Performing feature extraction on the preprocessed multi-modal data and preprocessed time-series data respectively based on the artificial intelligence algorithm, and obtaining multi-modal data features and time-series data features respectively, including: performing three-dimensional convolution operation on the preprocessed video stream, wherein the convolution kernel of the three-dimensional convolution is set to , and in the time dimension with frames as the sliding step; generating an activated feature map by passing the output of the three-dimensional convolution operation through the LeakyReLU activation function; performing global average pooling on the activated feature map to calculate the mean statistic of each channel; inputting the mean statistic into the Softmax function to generate normalized attention weights ; weighting the channels of the activated feature map according to the formula to obtain an enhanced feature map, where the is the eigenvalue of the c -th channel at the position .
2. The method for processing computer room status data based on artificial intelligence according to claim 1, wherein, The fusing the preprocessed multimodal data and preprocessed time-series data based on the multimodal data features and time-series data features after feature alignment to obtain fusion features includes: Calculating the similarity between the multimodal data features and time-series data features; Calculating the attention weights between the multimodal data and time-series data using the softmax function based on the similarity between the multimodal data features and time-series data features; Calculating the weighted fusion representation between the multimodal data features and time-series data features through the attention weights between the multimodal data and time-series data.
3. The method for processing computer room status data based on artificial intelligence according to claim 1, wherein The fusing the preprocessed multimodal data and preprocessed time-series data based on the multimodal data features and time-series data features after feature alignment to obtain fusion features includes: Mapping the multimodal data features and time-series data features after feature alignment to a first low-dimensional embedding representation and a second low-dimensional embedding representation respectively; Performing cross-mapping on the first low-dimensional embedding representation and the second low-dimensional embedding representation through a shared mapping function to obtain cross-mapped embedding features; Reconstructing the cross-mapped embedding features back to the feature space of the original modality.
4. The method for processing computer room status data based on artificial intelligence according to any one of claims 1 to 3, characterized in that The method further includes: Reducing the dimension of the fusion features.
5. An artificial intelligence-based computer room status data processing device, characterized in that, The device includes: A preprocessing module for preprocessing the multimodal data and time-series data related to the computer room respectively to obtain preprocessed multimodal data and preprocessed time-series data; A feature extraction module, which is used to perform feature extraction on the preprocessed multi-modal data and preprocessed time-series data respectively based on an artificial intelligence algorithm, and obtain multi-modal data features and time-series data features respectively. The performing feature extraction on the preprocessed multi-modal data and preprocessed time-series data respectively based on an artificial intelligence algorithm and obtaining multi-modal data features and time-series data features respectively includes: performing a three-dimensional convolution operation on the preprocessed video stream, wherein the convolution kernel of the three-dimensional convolution is set to in the spatial dimension and frames as the sliding step in the time dimension; generating an activated feature map by passing the output of the three-dimensional convolution operation through a LeakyReLU activation function; performing global average pooling on the activated feature map to calculate the mean statistic of each channel; inputting the mean statistic into a Softmax function to generate a normalized attention weight ; weighting the channels of the activated feature map according to the formula to obtain an enhanced feature map, where the is the eigenvalue of the c -th channel at the position ; An alignment module for feature alignment of the multi-modal data features and the temporal data features. The feature alignment of the multi-modal data features and the temporal data features includes: mapping the multi-modal data features and the temporal data features to a semantic space of the same dimension; calculating the similarity between the multi-modal data features and the temporal data features in the semantic space of the same dimension, where the similarity includes weights and biases; adjusting the weights and biases through an optimization algorithm to minimize the distance between the multi-modal data features and the temporal data features. The adjusting the weights and biases through an optimization algorithm to minimize the distance between the multi-modal data features and the temporal data features includes: constructing a loss function for measuring the alignment degree between different modal features; using the gradient descent algorithm to optimize the weights and biases; performing iterative training, calculating gradients and updating parameters in each training iteration until the loss function converges to an optimal value; A fusion module for fusing the pre-processed multi-modal data and the pre-processed temporal data based on the multi-modal data features and the temporal data features after feature alignment to obtain fusion features; A determination module for performing data analysis based on the fusion features to determine the current state of the computer room.
6. An electronic device, the device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 4.
7. A storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Method for determining fault of component in industrial equipment and related equipment
CN116933145A
Multi-dimensional machine room management method and system
CN118917605A