Image sample labeling method and system for time sequence fusion feature engineering in electric power transmission scene

The prototype prediction network of the Mamba architecture is used to extract and classify features of image samples in the electric transportation scenario, construct positive and negative sample sets, and update the prototype image feature vector, which solves the problem of high manual labeling costs and achieves efficient and accurate image sample labeling.

CN120689697APending Publication Date: 2025-09-23STATE GRID DIGITAL TECHNOLOGY HOLDING CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510781858.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

In the field of power demand and industrial equipment health monitoring, manual labeling of image samples is costly and inefficient, and existing technologies find it difficult to effectively improve labeling accuracy.

Method used

The prototype prediction network based on Mamba architecture is used to extract and classify the features of labeled image samples, construct positive and negative sample sets, update the prototype image feature vector, and use the updated feature vector to label the unlabeled image samples.

Benefits of technology

It improves the accuracy and efficiency of image sample annotation, provides a solid data foundation, and lays the foundation for subsequent in-depth analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689697A_ABST
    Figure CN120689697A_ABST
Patent Text Reader

Abstract

The invention provides an image sample labeling method and system for time sequence fusion feature engineering in an electric power transmission scene. According to the implementation scheme, on the basis of a prototype prediction network of a Mama framework, feature extraction is carried out on labeled image samples, time sequence category prototype features reflecting all label categories are obtained, prototype image feature vectors of all the label categories are determined, it can be ensured that the prototype vectors can capture the trend of data changing along with time, and the time sequence category prototype features of all the label categories are obtained. The application effect of feature engineering is enhanced; positive and negative sample pairs are constructed by using a plurality of labeled image samples and label categories thereof to update each prototype image feature vector, so that the understanding ability of a prediction network on time series data can be enhanced, and the application depth of feature engineering is improved; and on the basis of each updated prototype image feature vector and the confidence coefficient of each unlabeled image sample, determining and labeling the target unlabeled image sample to obtain a new labeled image sample, so that the image sample labeling efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and data processing technology, as well as the field of power technology, and in particular to an image sample annotation method and system for time series fusion feature engineering. Background Art

[0002] In areas such as power demand and industrial equipment health monitoring, neural network models are commonly used for prediction and processing. These models are pre-trained with a large number of training samples to improve their prediction accuracy. However, since these samples require labeling, manually labeling large numbers of samples results in high labeling costs.

[0003] Therefore, how to use machines to annotate samples with effective annotation resources and improve the annotation accuracy is a technical problem to be solved in this field. Summary of the Invention

[0004] The present invention provides an image sample annotation method and system for time series fusion feature engineering in electric transportation scenarios, which can solve at least one of the above technical problems.

[0005] According to one aspect of the present invention, a method for labeling image samples for time series fusion feature engineering in electric transportation scenarios is provided, comprising: Perform feature engineering processing on the image information of each labeled image sample in the time sequence of the labeled image sample set in the electric transportation scenario to obtain a labeled feature map sample set; Based on the prototype prediction network of the Mamba architecture, feature extraction is performed on the feature map of each labeled feature map sample in the labeled feature map sample set to obtain the prototype image features of the feature map in each labeled feature map sample; Based on the power label category information in each of the labeled feature map samples, each of the prototype image features is classified to obtain a prototype image feature set corresponding to each power label category; Determining, based on the prototype image feature sets corresponding to the respective power label categories, first prototype image feature vectors corresponding to the respective power label categories; Based on some samples in the labeled feature map sample set, constructing a positive sample set and a negative sample set of each power label category; Based on the positive sample sets and negative sample sets of each of the power label categories, the first prototype image feature vectors corresponding to each of the power label categories are updated to obtain the second prototype image feature vectors corresponding to each of the power label categories; based on the second prototype image feature vectors corresponding to each of the power label categories, sample annotation is performed on the unlabeled image samples in the power transportation scenario.

[0006] According to one aspect of the present invention, an image sample annotation device for time series fusion feature engineering in electric transportation scenarios is provided, comprising: A feature engineering processing module is used to perform feature engineering processing on image information in each labeled image sample arranged in time sequence in the labeled image sample set of the electric power transportation scene to obtain a labeled feature map sample set; A feature extraction module is used for extracting features from the feature maps in each labeled feature map sample in the labeled feature map sample set based on the prototype prediction network of the Mamba architecture, so as to obtain the prototype image features of the feature maps in each labeled feature map sample; A feature classification module, configured to classify each of the prototype image features based on the power label category information in each of the labeled feature map samples to obtain a prototype image feature set corresponding to each power label category; a prototype vector determining module, configured to determine a first prototype image feature vector corresponding to each of the power label categories based on a prototype image feature set corresponding to each of the power label categories; A sample set construction module, configured to construct a positive sample set and a negative sample set of each power label category based on some samples in the labeled feature map sample set; a prototype vector updating module, configured to update the first prototype image feature vector corresponding to each power label category based on the positive sample set and the negative sample set of each power label category, to obtain a second prototype image feature vector corresponding to each power label category; The sample labeling module is used to label the unlabeled image samples in the electric transportation scenario based on the second prototype image feature vector corresponding to each of the electric power label categories.

[0007] According to another aspect of the present invention, there is provided an image sample labeling system for time series fusion feature engineering in electric transportation scenarios, comprising: at least one processor, and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the processor is configured to obtain the instructions from the memory and execute the instructions, so that the processor can execute the image sample labeling method for time series fusion feature engineering described in any one of the embodiments of the present invention.

[0008] According to another aspect of the present invention, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to provide to a computer and instruct the computer to execute the image sample labeling method for time series fusion feature engineering described in any one of the embodiments of the present invention.

[0009] By adopting the technical solution of the present invention, feature engineering is performed on the image information in each labeled image sample in the labeled image sample set of the electric power transportation scene to obtain a labeled feature map sample set, wherein the feature engineering includes time series analysis and feature map construction; based on the prototype prediction network of the Mamba architecture, feature extraction is performed on the feature map in each labeled feature map sample in the labeled feature map sample set, and the prototype image features of the feature map in each labeled feature map sample are obtained. In this way, the prototype prediction network is used to extract and process the feature maps with time series after feature engineering. The obtained prototype features can accurately capture the trends and laws of feature changes over time and enhance the application effect of feature engineering. Based on the power label category information in each labeled feature map sample, each prototype image feature is classified to obtain the prototype image feature set corresponding to each power label category; based on the prototype image feature set corresponding to each power label category, the first prototype image feature vector corresponding to each power label category is determined. In this way, the prototype image feature vector corresponding to each power label category can be preliminarily determined. Next, based on a portion of the samples in the labeled feature map sample set, a positive sample set and a negative sample set are constructed for each power label category. Based on the positive sample set and the negative sample set for each power label category, the first prototype image feature vector corresponding to each power label category is updated to obtain the second prototype image feature vector corresponding to each power label category. In this way, the positive and negative sample sets for each power label category can be used to update the preliminarily determined prototype image feature vector. Based on the second prototype image feature vector corresponding to each power label category, the unlabeled image samples are labeled. In this way, using the prototype image feature vectors updated by the positive and negative sample sets for sample labeling can improve the accuracy of sample labeling.

[0010] Moreover, using the new labeled samples in subsequent sample labeling can deepen the application value of feature engineering in time series data analysis and provide a solid data foundation and technical support for subsequent in-depth analysis.

[0011] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings are provided for a better understanding of the present invention and do not constitute a limitation of the present invention. Figure 1 This is a flowchart of an image sample labeling method for time series fusion feature engineering in an electric transportation scenario according to an embodiment of the present invention; Figure 2 is a schematic diagram of an electric transportation scenario according to an embodiment of the present invention; Figure 3 is a structural diagram of a prototype prediction network according to an embodiment of the present invention; Figure 4 This is a structural block diagram of an image sample annotation device for time series fusion feature engineering in electric transportation scenarios according to an embodiment of the present invention; Figure 5 is a block diagram of an electronic device for implementing the method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0013] The following description of exemplary embodiments of the present invention is made in conjunction with the accompanying drawings, and various details of the embodiments of the present invention are included to facilitate understanding. These details should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present invention. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0014] Figure 1 This is a flowchart of an image sample labeling method for time series fusion feature engineering in an electric transportation scenario according to an embodiment of the present invention.

[0015] like Figure 1 As shown, the image sample annotation method for time series fusion feature engineering in electric transportation scenarios may include: S110, performing feature engineering processing on image information in each labeled image sample arranged in time sequence in the labeled image sample set of the electric power transportation scene to obtain a labeled feature map sample set; S120, a prototype prediction network based on the Mamba architecture, performs feature extraction on the feature map in each labeled feature map sample in the labeled feature map sample set to obtain prototype image features of the feature map in each labeled feature map sample; S130, classifying each prototype image feature based on the power label category information in each labeled feature map sample to obtain a prototype image feature set corresponding to each power label category; S140, determining a first prototype image feature vector corresponding to each power label category based on a prototype image feature set corresponding to each power label category; S150, constructing a positive sample set and a negative sample set for each power label category based on some samples in the labeled feature map sample set; S160, based on the positive sample set and the negative sample set of each power label category, updating the first prototype image feature vector corresponding to each power label category to obtain a second prototype image feature vector corresponding to each power label category; S170 , performing sample labeling on unlabeled image samples in the electric power transportation scenario based on the second prototype image feature vector corresponding to each electric power label category.

[0016] For example, the power transportation scenario may be a scenario of the interconnected transmission lines and substations under a designated regional power grid, as well as the location relationship between the transmission lines, substations and surrounding vegetation and buildings. Figure 2 Given a power transportation scenario involving transmission lines and towers, we can capture images of the area or surroundings within the scenario to detect potential transportation risks. These images can include transmission lines, substations, vegetation, or buildings.

[0017] For example, the power tag category information may include power equipment category, obstacle category, etc. The power equipment category may include transmission lines, substations, transmission towers, etc.

[0018] In practical applications, image segmentation models can be used to identify power-related objects and their locations in the surrounding environment image. Based on the identification results, the presence of transportation risks in the power transmission scenario can be determined. For example, is there a risk that vegetation could fall onto the transmission line, impacting power transmission? Or are there risks of external damage to the transmission line or substation?

[0019] For the aforementioned electric power transportation scenario, a small number of manually labeled ambient image samples and a large number of unlabeled ambient image samples are obtained. The unlabeled ambient image samples are then labeled using the method of an embodiment of the present invention in combination with the labeled ambient image samples, resulting in a large number of labeled ambient image samples. These labeled ambient image samples are then used to train the aforementioned image segmentation model.

[0020] It can be understood that step S110 to step S170 can be executed in a loop, and each time step S170 is executed, the process returns to step S110 and is executed from step S110 to step S170.

[0021] Exemplarily, the image information in the labeled image samples is subjected to feature engineering processing operations such as preprocessing, time series decomposition, and feature map construction to obtain labeled feature map samples.

[0022] In practical application scenarios, there is typically a large pool of unlabeled time series data, U, which contains historical data or multi-dimensional time series samples collected in real time. To ensure that the model has a certain degree of discriminative power in the initial stage, the present invention first randomly extracts a small number of samples from this data pool and manually annotates them to construct an initial annotated dataset. The number of initial annotated samples can be determined based on the available manpower or time costs in the actual application scenario, or it can be selected based on the diversity of the data distribution to ensure the representativeness of the initial annotated data.

[0023] Therefore, before inputting the labeled image samples into the prototype prediction network, it is necessary to perform feature engineering operations such as preprocessing, time series decomposition, and feature construction on the original time series image samples to obtain labeled feature map samples, and then input the labeled feature map samples into the prototype prediction network.

[0024] Exemplarily, feature engineering may include time series analysis and feature graph construction.

[0025] Exemplarily, the noise and abnormal values ​​in the original time series image data are preprocessed by filtering, smoothing and other methods, while the missing values ​​are interpolated and eliminated.

[0026] For example, in order to eliminate the influence between different dimensions, the time series image data is normalized or standardized to ensure that each time series feature is compared at the same scale.

[0027] For example, a sliding window is used to divide the normalized or standardized time series data into subsequences of fixed length to capture local time series features, thereby achieving time series decomposition.

[0028] Exemplarily, statistical features such as mean, variance, maximum value, minimum value, skewness and kurtosis are extracted from each subsequence to intuitively describe the characteristic distribution properties of the subsequence.

[0029] Exemplarily, frequency domain features are extracted from each subsequence by using Fourier transform or wavelet transform, which can reflect the periodicity and seasonality information between the subsequences.

[0030] Exemplarily, the feature usage requirements required by feature engineering are adopted to construct corresponding time series features, such as time lag features and periodic features, based on the statistical features and frequency domain features of each subsequence.

[0031] For example, the prototype prediction network can extract features from the feature maps in the labeled feature map samples and classify the extracted features into the category prototype image feature set corresponding to the corresponding power label category based on the annotation information in the labeled feature map samples, namely the power label category. The prototype prediction network input includes the sample body information and sample annotation information of the labeled feature map samples. The samples can be images or text information in the power field.

[0032] It can be understood that the prototype prediction network of the Mamba architecture can extract features that conform to the category specified by the annotation information from the sample ontology information in combination with the annotation information in the labeled feature map sample.

[0033] Exemplarily, the first prototype image feature vector can be calculated using the following formula: Among them, L k represents the class prototype image feature set corresponding to the k-th power label class, (x i ,y i ) represents the labeled feature map sample x i With the labeled feature map sample x i The probability y of belonging to the kth electricity label category i The combination of φ (x i ) represents the prototype prediction network for the labeled feature map sample x i The output features, namely the prototype image features, c k Represents the first prototype image feature vector corresponding to the k-th power label category.

[0034] When initializing the prototype image feature vector, the prototype prediction network of the Mamba architecture can be used to extract features from each labeled feature map sample in the labeled feature map sample set to obtain the category prototype image feature set corresponding to each power label category. Then, the mean features of the category prototype image feature set corresponding to each power label category are respectively calculated to obtain the initial first prototype image feature vector.

[0035] Exemplarily, the features in the category prototype image feature set corresponding to the specified power label category are clustered to obtain the cluster center, and the distance between each feature in the category prototype image feature set and the cluster center is used to obtain multiple features. The mean of the multiple features is calculated to obtain the initial first prototype image feature vector corresponding to the power label category.

[0036] Exemplarily, a small batch of multiple first labeled feature map samples is extracted from the labeled feature map sample set, and thereby a positive sample set and a negative sample set of each power label category are constructed. The positive sample set and the negative sample set of each power label category are used to update the first prototype image feature vector corresponding to each power label category, respectively, to obtain the second prototype image feature vector corresponding to each power label category.

[0037] Exemplarily, a plurality of first labeled feature map samples are determined from the labeled feature map sample set; based on the prototype image features of the feature maps in each of the first labeled feature map samples and the power label category information in each of the first labeled feature map samples, a positive sample set and a negative sample set of each power label category are constructed; It can be understood that the positive sample set of the power label category includes the labeled feature map samples of the same category as the power label category, and the negative sample set of the power label category includes the labeled feature map samples of a different category than the power label category. Each power label category has a positive sample set and a negative sample set.

[0038] Exemplarily, one or more target unlabeled image samples are determined from the unlabeled image sample set, and based on the confidence between each target unlabeled image sample and the second prototype image feature vector corresponding to each power label category, each target unlabeled image sample is labeled separately to obtain each target labeled image sample.

[0039] Exemplarily, a prototype prediction network can be used to extract features from unlabeled image samples and calculate the distance between the feature and each second prototype image feature vector, so as to use the distance to determine the probability distribution, with the probability distribution being the confidence between the unlabeled image sample and each prototype image feature vector.

[0040] For example, for any unlabeled image sample x∈U in the unlabeled data pool U, first use the currently trained prototype prediction network f φ Extract prototype image features f of unlabeled image samples x φ (x), and then calculate the second prototype image feature vector c corresponding to the prototype image feature and each power label category k The Euclidean distance between: d(f φ (x),c k )=||f φ (x)-c k ||2; Using the Softmax function, the calculated distance is calculated to obtain the probability distribution p of the unlabeled image sample x belonging to the kth power label category φ (y=k|x).

[0041] The above probability distribution reflects the classification confidence of the current model for the unlabeled image sample x.

[0042] For example, after obtaining the above confidence, a minimum confidence sampling method can be used to determine the unlabeled image sample with the lowest confidence from the unlabeled image sample set as the target unlabeled image sample. In this way, the sample with the least uncertainty in the prototype prediction network can be selected to improve the decision boundary of the prototype prediction network.

[0043] Alternatively, margin sampling can be used, for example, selecting the sample with the smallest difference between the highest probability and the second highest probability as the target unlabeled image sample. In this way, margin sampling can select the sample that is most difficult for the model to distinguish, thereby improving the model's ability to distinguish similar categories.

[0044] Alternatively, entropy sampling can be used to select the sample with the highest entropy in the classification probability distribution as the target unlabeled image sample for labeling. The greater the entropy of the probability distribution, the less certain the model is about the sample's classification. Entropy sampling aims to select the sample with the highest information content to improve the overall performance of the model.

[0045] According to the above embodiment, feature engineering is performed on the image information in each labeled image sample in the labeled image sample set of the electric power transportation scene to obtain a labeled feature map sample set, wherein the feature engineering includes time series analysis and feature map construction; based on the prototype prediction network of the Mamba architecture, feature extraction is performed on the feature map in each labeled feature map sample in the labeled feature map sample set, and the prototype image features of the feature map in each labeled feature map sample are obtained. In this way, the prototype prediction network is used to extract and process the feature maps with time series characteristics after feature engineering. The obtained prototype features can accurately capture the trends and laws of feature changes over time, thereby enhancing the application effect of feature engineering. Based on the power label category information in each labeled feature map sample, each prototype image feature is classified to obtain the prototype image feature set corresponding to each power label category; based on the prototype image feature set corresponding to each power label category, the first prototype image feature vector corresponding to each power label category is determined. In this way, the prototype image feature vector corresponding to each power label category can be preliminarily determined. Next, based on a portion of the samples in the labeled feature map sample set, a positive sample set and a negative sample set are constructed for each power label category. Based on the positive sample set and the negative sample set for each power label category, the first prototype image feature vector corresponding to each power label category is updated to obtain the second prototype image feature vector corresponding to each power label category. In this way, the positive and negative sample sets for each power label category can be used to update the preliminarily determined prototype image feature vector. Based on the second prototype image feature vector corresponding to each power label category, the unlabeled image samples are labeled. In this way, using the prototype image feature vectors updated by the positive and negative sample sets for sample labeling can improve the accuracy of sample labeling.

[0046] In one embodiment, based on the positive sample set and negative sample set of each power label category, the first prototype image feature vector corresponding to each power label category is updated to obtain the second prototype image feature vector corresponding to each power label category, including: for the positive sample set and negative sample set of the power label category, determining the distance between each positive sample in the positive sample set and each other positive sample and the distance between each positive sample and each negative sample in the negative sample set, and determining the contrast loss function; based on the gradient information in the contrast loss function, updating the first prototype image feature vector corresponding to the power label category to obtain the second prototype image feature vector corresponding to the power label category.

[0047] Exemplarily, the contrast loss function is as follows: Among them, f i =f φ (x i ) is a positive sample xi The prototype image feature, f j =f φ (x j ) is the sample x j Feature representation; f k =f φ (x k ) is the sample x k Prototype image features; P i and N i and x respectively i The sample set belonging to the same category and x i A collection of samples belonging to different categories; d is a distance function, such as the Euclidean distance function, and τ is a temperature hyperparameter.

[0048] Exemplarily, the gradient information in the contrast loss function can be combined with the learning rate of the first prototype image feature vector to update the first prototype image feature vector corresponding to the power label category to obtain the second prototype image feature vector corresponding to the power label category.

[0049] According to the above embodiment, the contrast loss function is used to update the first prototype image feature vector, which can shorten the distance between similar features and increase the distance between different features.

[0050] In one embodiment, the above-mentioned sample labeling of unlabeled image samples in the power transportation scenario based on the second prototype image feature vector corresponding to each of the power label categories includes: determining the target unlabeled image sample from the unlabeled image sample set based on the confidence between the second prototype image feature vector corresponding to each power label category and each unlabeled image sample in the unlabeled image sample set; labeling the target unlabeled image sample based on the confidence between the target unlabeled image sample and the second prototype image feature vector corresponding to each power label category to obtain the target labeled image sample.

[0051] The calculation method of the confidence level and the selection method of the target unlabeled image samples in this example can be referred to the previous example and will not be described in detail here.

[0052] For example, based on the confidence between the target unlabeled image sample and the second prototype image feature vector corresponding to each power label category, the power label category with the highest confidence is determined. The target unlabeled image sample is labeled with this highest confidence power label category to obtain the target labeled image sample. As a result, the target labeled image sample is labeled with the power label category with the highest confidence.

[0053] According to the above embodiment, target unlabeled image samples can be selected according to the confidence level, and the target unlabeled image samples can be labeled using the confidence level between the target unlabeled image samples and the second prototype image feature vectors corresponding to each power label category to obtain target labeled image samples. In this way, only the target unlabeled image samples that most need to be labeled can be labeled, which can improve the labeling efficiency of image samples.

[0054] In one embodiment, the above method also includes: deleting the target unlabeled image sample from the unlabeled image sample set; updating the labeled feature map sample set based on the target labeled sample; and returning to continue executing the image sample labeling method based on the updated labeled feature map sample set to obtain the next target labeled image sample.

[0055] For example, one annotation can be performed on one or several or a preset number of target unlabeled image samples.

[0056] In this example, each time labeling is performed, the process returns to step S110 to step S180 to re-determine the prototype image feature vector for the next labeling. In this way, the prototype image feature vector is continuously updated during the labeling process to avoid overfitting of the labeling effect, thereby improving the labeling accuracy.

[0057] In one embodiment, Figure 3 As shown, the prototype prediction network includes a sequence encoder and a multiple memory bank of a Mamba architecture. The above method further includes: updating the memory block of each power label category in the multiple memory bank based on the second prototype image feature vector corresponding to each power label category.

[0058] In which, the sequence encoder includes multiple network layers, and a temporal attention module is integrated between the first network layer and the second network layer in the multiple network layers; the temporal attention module is used to perform attention weighting on the output image features of the second network layer at the current time step, and fuse the weighted attention results into the input image features of the first network layer at the current time step, so that the first network layer outputs the output image features at the next time step; wherein, the memory block of the power label category corresponding to the input image features of the first network layer at the current time step is used to fuse into the input image features of the first network layer at the current time step, so that the first network layer outputs the output image features at the next time step.

[0059] Exemplarily, the first network layer and the second network layer may be adjacent network layers, or may be two network layers separated by several network layers. From the input to the output direction of the sequence encoder, the first network layer is located before the second network layer.

[0060] Exemplarily, the memory blocks for each power label category in the multiple memory banks are updated with the second prototype image feature vector corresponding to the power label category. Alternatively, the memory block for a given power label category is fused with the second prototype image feature vector corresponding to the power label category to obtain a new memory block for the power label category.

[0061] Exemplarily, an embedding and position encoding layer, i.e., an input layer, may be provided before the sequence encoder to convert the input samples into a feature vector representation that can be processed by the model and to encode time information and position information.

[0062] For example, the sequence encoder may be a Transformer encoder, and a feature aggregation layer may be provided at the output of the Transformer encoder. The feature aggregation layer is configured to perform a global averaging operation on the output of the Transformer encoder to aggregate sequence features into a feature vector having the same format as the sample vector.

[0063] Exemplarily, the feature aggregation layer may be provided with a time series prediction self-supervision head for performing time series self-supervision learning on the results of the global averaging operation.

[0064] Exemplarily, the sequence encoder also integrates the above-mentioned multiple memory banks and temporal attention modules.

[0065] According to the above implementation, the Mamba architecture is combined with a Transformer encoder, and key components such as a dynamic memory gated update mechanism and a temporal attention module are incorporated to achieve efficient feature extraction, dynamic memory updates, and robust classification of time series image sample data. Furthermore, by using the second prototype image feature vector corresponding to each power label category to update the memory blocks for each power label category in the multi-memory bank, the memory blocks can dynamically record the latest image features of each category. This improves the feature extraction accuracy of the prototype prediction network.

[0066] In one embodiment, the above method also includes: inputting the target labeled image sample into the sequence encoder to obtain the output image features of each network layer in the sequence encoder; determining the target memory block of the corresponding category in the multiple memory library based on the power label category of the target labeled image sample; and updating the target memory block based on the output image features of each network layer in the sequence encoder to update the multiple memory library.

[0067] It can be understood that the power label category of the target labeled image sample is the same as the power label category of the target memory block.

[0068] Exemplarily, the output image features of each network layer in the sequence encoder are fused, and the fused features are used to update the target memory block. For example, the output image features of each network layer can be fused using a dot product, and then the target memory block is updated with the fused image features. In another example, the target memory block before the update is further fused with the fused image features to obtain a new target memory block.

[0069] In this example, each time a target labeled image sample is obtained through labeling, the sample is input into the original sequence encoder to obtain the output features of each network layer. The output features of the sample are used to update the memory block of the corresponding category of the sample. In this way, during the labeling process of a large number of samples, the labeling results are used to continuously update multiple memory banks, maintain the memory effect of the memory banks, and further improve the prediction accuracy of the prototype prediction network.

[0070] In one embodiment, based on the output image features of each network layer in the sequence encoder, the target memory block is updated to update the multiple memory banks, including: splicing the output image features of the last network layer in the sequence encoder with the target memory block, and learning the splicing result based on the model parameters of the multiple memory banks to obtain a gating coefficient; performing a dot product on the gating coefficient and the target memory block to obtain a first dot product result; calculating the difference between 1 and the gating coefficient, and performing a dot product on the difference and the layer normalization operation result of the output image feature of the last network layer to obtain a second dot product result; and updating the target memory block based on the first dot product result and the second dot product result.

[0071] For example, the gating coefficients may be as follows: Among them, the gating coefficient g is generated by the Sigmoid function, and its value range is (0,1), which is used to control the feature information f of the new sample φ (x i ) for the original target memory block The update level, W g Represents the model parameters of the multiple memory bank.

[0072] For example, the updated target memory block It can be: Among them, LayerNorm(·) represents the layer normalization operation, W m Represents the model parameters of the layer normalization.

[0073] According to the above embodiment, the retention and update ratio of the memory block, i.e., the gating coefficient, is dynamically adjusted according to the characteristics of the newly annotated image sample. Then, the memory block is updated using the gating coefficient, thereby realizing the adaptive update of the memory library, thereby improving the prototype prediction network's ability to quickly adapt to new image samples.

[0074] In one embodiment, it also includes: determining the cross entropy loss based on the category prototype image feature set corresponding to each power label category output by the prototype prediction network, and each labeled feature map sample in the labeled feature map sample set; determining the memory bank loss based on the updated target memory block in the multiple memory banks and the output features of the last network layer; determining the prototype learning loss based on the category prototype image feature set corresponding to each power label category, and the first prototype image feature vector corresponding to each power label category; determining the prototype contrast regularization loss based on the second prototype image feature vector corresponding to each power label category, and the positive sample pairs and the negative sample pairs; determining the time series prediction self-supervision loss based on the output result of the time series prediction self-supervision head and the corresponding annotation result; determining the target loss based on the cross entropy loss, the memory bank loss, the prototype learning loss, the prototype contrast regularization loss and the time series prediction self-supervision loss; and updating the model parameters of the prototype prediction network based on the target loss.

[0075] For example, the cross entropy loss It is used to measure the difference between the category probability distribution predicted by the model and the true label. Its calculation formula is: Among them, p φ (y=k|x i ) represents the model prediction sample x with parameter φ i The probability of belonging to category k. 1(y i =k) ​​is the indicator function, which means that when y i =k, the value is 1. N is the total number of samples used in this batch training, and K is the total number of categories.

[0076] For example, the prototype learning loss It is used to make the sample features of the same category close to the feature vector of the prototype image of that category. The calculation formula is as follows: Among them, f φ (x i ) represents the sample x extracted by the Mamba encoder i The eigenvector of For category y i The prototype image feature vector of , L is the labeled feature map sample set.

[0077] For example, memory loss It is used to constrain the update of the memory bank so that the memory vector obtained by dynamic memory gating update is close to the intermediate result calculated based on the current sample features. Its calculation formula is as follows: in, For category y i After the memory vector or memory block is updated by dynamic memory gating, W m is the linear transformation matrix used in memory update, which can also be considered as the model parameter of the dynamic memory gating update mechanism.

[0078] For example, the prototype contrast regularization loss It is used to encourage similar samples in the feature space to move closer and dissimilar samples in the feature space to move apart through contrast regularization. The formula for calculating the prototype contrast regularization loss is as follows: Among them, f i =f φ (x i ) is the sample x i The feature representation, f j =f φ (x j ) is the sample x j Feature representation; f k =f φ (x k ) is the sample x k The feature representation of P i and N i and x respectively i The sample set belonging to the same category and x i A collection of samples belonging to different categories; d is a distance function such as the Euclidean distance function, and τ is a temperature hyperparameter.

[0079] For example, the time series prediction self-supervised loss It is used to improve the model's ability to model dynamic characteristics of time series by predicting the signal at the next moment. The calculation formula for time series prediction self-supervision loss is as follows: in, Indicates using the current hidden state h t and the corresponding category prototype c k After splicing, it passes through the linear prediction layer W p The obtained prediction signal, x t+1 Represents the real signal corresponding to the predicted signal, and T is the length of the time series.

[0080] For example, during the training process of the parameter adjustment algorithm, the model parameters φ, the prototype image feature vector c k , memory update parameter W g 、W m And the prediction layer parameters W p Both can be jointly optimized through the back-propagation algorithm.

[0081] For example, the following is a process of adjusting the model parameters based on a gradient descent method (such as the Adam optimizer): For example, at each training iteration, a small batch of samples is sampled from the labeled feature map sample set L To train the model parameters and forward propagate. And calculate the losses Then, the weighted sum of each loss is performed to obtain the total loss. Specifically, the total loss is as follows: Among them, the total loss Back propagation calculation to get the gradient Among them, θ can contain φ, {C k}, W g , W m , W p λ1, λ2, λ3, λ4 and λ5 represent the weights of the corresponding losses respectively.

[0082] Then, the θ parameters are updated according to the optimization algorithm as follows: Here, η represents the learning rate.

[0083] Therefore, through the joint optimization of multi-task losses in this example, the model can not only achieve high accuracy in classification tasks, but also achieve continuous improvement in prototype representation, memory update, and time series prediction, thereby better adapting to the dynamically changing distribution of time series data.

[0084] Figure 4 This is a structural block diagram of an image sample labeling device for time series fusion feature engineering in an electric transportation scenario according to an embodiment of the present invention.

[0085] like Figure 4 As shown, the image sample annotation device for time series fusion feature engineering in electric transportation scenarios includes: A feature engineering processing module 410 is configured to perform feature engineering processing on image information in each of the labeled image samples arranged in time sequence in the labeled image sample set of the electric power transportation scene to obtain a labeled feature map sample set; A feature extraction module 420 is configured to extract features from the feature graphs in each labeled feature graph sample in the labeled feature graph sample set based on the prototype prediction network of the Mamba architecture, and obtain prototype image features of the feature graphs in each labeled feature graph sample; a feature classification module 430 is configured to classify each prototype image feature based on the power label category information in each labeled feature graph sample, and obtain a prototype image feature set corresponding to each power label category; A prototype vector determination module 440 is configured to determine a first prototype image feature vector corresponding to each power label category based on a prototype image feature set corresponding to each power label category; A sample set construction module 450 is used to construct a positive sample set and a negative sample set of each power label category based on some samples in the labeled feature map sample set; A prototype vector updating module 460 is configured to update the first prototype image feature vector corresponding to each power label category based on the positive sample set and the negative sample set of each power label category to obtain a second prototype image feature vector corresponding to each power label category; The sample labeling module 470 is configured to label the unlabeled image samples in the electric transportation scenario based on the second prototype image feature vector corresponding to each electric power label category.

[0086] In one embodiment, the prototype vector updating module 460 includes: a contrast loss function determining unit, configured to determine, for a positive sample set and a negative sample set of the power label category, a distance between each positive sample in the positive sample set and each other positive sample, and a distance between each positive sample and each negative sample in the negative sample set, to determine a contrast loss function; The original vector updating unit is used to update the first prototype image feature vector corresponding to the power label category based on the gradient information in the contrast loss function to obtain the second prototype image feature vector corresponding to the power label category.

[0087] In one embodiment, the sample annotation module 470 includes: a target sample determining unit, configured to determine a target unlabeled image sample from the unlabeled image sample set based on a confidence between a second prototype image feature vector corresponding to each of the power label categories and each unlabeled image sample in the unlabeled image sample set; The sample labeling unit is configured to label the target unlabeled image sample based on the confidence between the target unlabeled image sample and the second prototype image feature vector corresponding to each of the power label categories to obtain a target labeled image sample.

[0088] In one embodiment, the apparatus further comprises: A sample deletion module, configured to delete the target unlabeled image sample from the unlabeled image sample set; A sample set updating module, configured to update the labeled image sample set based on the target labeled image sample; The cyclic labeling module is used to return to continue executing the image sample labeling method based on the updated labeled image sample set to obtain the next target labeled image sample.

[0089] In one embodiment, the prototype prediction network includes a sequence encoder and multiple memory banks of a Mamba architecture, and the apparatus further includes: A first memory block updating module, configured to update the memory blocks of each power label category in the multiple memory banks based on the second prototype image feature vector corresponding to each power label category; The sequence encoder includes multiple network layers, and a temporal attention module is integrated between a first network layer and a second network layer in the multiple network layers. The temporal attention module is used to perform attention weighting on the output image features of the second network layer at the current time step, and fuse the weighted attention result with the input image features of the first network layer at the current time step, so that the first network layer outputs the output image features at the next time step; Among them, the memory block of the power label category corresponding to the input image feature of the first network layer at the current time step is used to be integrated into the input image feature of the first network layer at the current time step, so that the first network layer outputs the output image feature at the next time step.

[0090] In one embodiment, the apparatus further comprises: An image feature acquisition module, configured to input the target labeled image sample into the sequence encoder to obtain output image features of each network layer in the sequence encoder; a target memory block determination module, configured to determine a target memory block of a corresponding category in the multiple memory banks based on the power label category of the target labeled image sample; The second memory block updating module is used to update the target memory block based on the output image features of each network layer in the sequence encoder to update the multiple memory banks.

[0091] In one embodiment, the second memory block update module includes: a gating system determination unit, configured to concatenate the output image features of the last network layer in the sequence encoder with the target memory block, and learn the concatenation result based on the model parameters of the multiple memory banks to obtain a gating coefficient; a first dot product unit, configured to perform a dot product on the gating coefficient and the target memory block to obtain a first dot product result; A second dot product unit is used to calculate the difference between 1 and the gating coefficient, and perform a dot product on the difference with the layer normalization operation result of the output image feature of the last network layer to obtain a second dot product result; A memory block updating unit is configured to update the target memory block based on the first dot product result and the second dot product result.

[0092] For the description of specific functions and examples of each module and submodule of the system in the embodiment of the present invention, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0093] In the technical solution of the present invention, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0094] According to an embodiment of the present invention, the present invention further provides a system and a readable storage medium.

[0095] Figure 5 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0096] like Figure 5 As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0097] Multiple components in the electronic device 800 are connected to the I / O interface 805, including an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0098] The computing unit 801 can be various general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the image sample labeling method for time series fusion feature engineering in electric transportation scenarios. For example, in some embodiments, the image sample labeling method for time series fusion feature engineering in electric transportation scenarios can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the image sample labeling method for time series fusion feature engineering in electric transportation scenarios described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured in any other appropriate manner (e.g., by means of firmware) to perform an image sample labeling method for time series fusion feature engineering in electric transportation scenarios.

[0099] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0100] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0101] In the context of the present invention, machine-readable medium can be a tangible medium that can contain or store a program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0102] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0103] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0104] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0105] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present invention can be achieved. This is not limited herein.

[0106] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. An image sample annotation method for time series fusion feature engineering in electric transportation scenarios, characterized by: include: Perform feature engineering processing on the image information of each labeled image sample in the time sequence of the labeled image sample set in the electric transportation scenario to obtain a labeled feature map sample set; Based on the prototype prediction network of the Mamba architecture, feature extraction is performed on the feature map of each labeled feature map sample in the labeled feature map sample set to obtain the prototype image features of the feature map in each labeled feature map sample; Based on the power label category information in each of the labeled feature map samples, each of the prototype image features is classified to obtain a prototype image feature set corresponding to each power label category; Determining, based on the prototype image feature sets corresponding to the respective power label categories, first prototype image feature vectors corresponding to the respective power label categories; Based on some samples in the labeled feature map sample set, constructing a positive sample set and a negative sample set of each power label category; Based on the positive sample sets and negative sample sets of each of the power label categories, the first prototype image feature vectors corresponding to each of the power label categories are updated to obtain the second prototype image feature vectors corresponding to each of the power label categories; based on the second prototype image feature vectors corresponding to each of the power label categories, sample annotation is performed on the unlabeled image samples in the power transportation scenario.

2. The method according to claim 1, characterized in that The updating of the first prototype image feature vector corresponding to each power label category based on the positive sample set and the negative sample set of each power label category to obtain the second prototype image feature vector corresponding to each power label category includes: For the positive sample set and the negative sample set of the power label category, determining the distance between each positive sample in the positive sample set and each other positive sample, and the distance between each positive sample and each negative sample in the negative sample set, and determining a contrast loss function; Based on the gradient information in the contrast loss function, the first prototype image feature vector corresponding to the power label category is updated to obtain a second prototype image feature vector corresponding to the power label category.

3. The method according to claim 1, characterized in that The step of labeling the unlabeled image samples in the electric transportation scenario based on the second prototype image feature vectors corresponding to the respective electric power label categories includes: Determining a target unlabeled image sample from the unlabeled image sample set based on a confidence between a second prototype image feature vector corresponding to each of the power label categories and each unlabeled image sample in the unlabeled image sample set; Based on the confidence between the target unlabeled image sample and the second prototype image feature vector corresponding to each of the power label categories, the target unlabeled image sample is labeled to obtain a target labeled image sample.

4. The method according to claim 3, characterized in that The method further comprises: Deleting the target unlabeled image sample from the unlabeled image sample set; Based on the target labeled image sample, updating the labeled image sample set; Based on the updated labeled image sample set, the image sample labeling method is returned to be executed again to obtain the next target labeled image sample.

5. The method according to claim 3, characterized in that The prototype prediction network includes a sequence encoder of a Mamba architecture and a multiple memory bank, and the method further includes: updating the memory blocks of each power label category in the multiple memory banks based on the second prototype image feature vector corresponding to each power label category; The sequence encoder includes multiple network layers, and a temporal attention module is integrated between a first network layer and a second network layer in the multiple network layers. The temporal attention module is used to perform attention weighting on the output image features of the second network layer at the current time step, and fuse the weighted attention result with the input image features of the first network layer at the current time step, so that the first network layer outputs the output image features at the next time step; Among them, the memory block of the power label category corresponding to the input image feature of the first network layer at the current time step is used to be integrated into the input image feature of the first network layer at the current time step, so that the first network layer outputs the output image feature at the next time step.

6. The method according to claim 5, characterized in that The method further comprises: Inputting the target labeled image sample into the sequence encoder to obtain output image features of each network layer in the sequence encoder; Determining a target memory block of a corresponding category in the multiple memory banks based on the power label category of the target labeled image sample; Based on the output image features of each network layer in the sequence encoder, the target memory block is updated to update the multiple memory banks.

7. The method according to claim 6, characterized in that The updating of the target memory block based on the output image features of each network layer in the sequence encoder to update the multiple memory banks includes: splicing the output image features of the last network layer in the sequence encoder with the target memory block, and learning the splicing result based on the model parameters of the multiple memory banks to obtain a gating coefficient; Performing a dot product on the gating coefficient and the target memory block to obtain a first dot product result; Calculating a difference between 1 and the gating coefficient, and performing a dot product between the difference and a layer normalization operation result of the output image feature of the last network layer to obtain a second dot product result; The target memory block is updated based on the first dot product result and the second dot product result.

8. An image sample annotation device for time series fusion feature engineering in electric transportation scenarios, characterized by: include: A feature engineering processing module is used to perform feature engineering processing on image information in each labeled image sample arranged in time sequence in the labeled image sample set of the electric power transportation scene to obtain a labeled feature map sample set; A feature extraction module is used for extracting features from the feature maps in each labeled feature map sample in the labeled feature map sample set based on the prototype prediction network of the Mamba architecture, so as to obtain the prototype image features of the feature maps in each labeled feature map sample; A feature classification module, configured to classify each of the prototype image features based on the power label category information in each of the labeled feature map samples to obtain a prototype image feature set corresponding to each power label category; a prototype vector determining module, configured to determine a first prototype image feature vector corresponding to each of the power label categories based on a prototype image feature set corresponding to each of the power label categories; A sample set construction module, configured to construct a positive sample set and a negative sample set of each power label category based on some samples in the labeled feature map sample set; a prototype vector updating module, configured to update the first prototype image feature vector corresponding to each power label category based on the positive sample set and the negative sample set of each power label category, to obtain a second prototype image feature vector corresponding to each power label category; The sample labeling module is used to label the unlabeled image samples in the electric transportation scenario based on the second prototype image feature vector corresponding to each of the electric power label categories.

9. An image sample annotation system for time series fusion feature engineering in electric transportation scenarios, characterized by: include: at least one processor, and a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the processor is used to obtain the instructions from the memory and execute the instructions, so that the processor can execute the method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to be provided to a computer to instruct the computer to execute the method according to any one of claims 1 to 7.