An electromagnetic data annotation method and system based on human-machine hybrid intelligence

Through the electromagnetic data labeling method based on human-computer hybrid intelligence, deep clustering and manual auditing technology are used to solve the problem of high time and low accuracy in the existing technology, and efficient and accurate electromagnetic data labeling is achieved.

CN115659193BActive Publication Date: 2025-06-10SOUTHWEST CHINA RES INST OF ELECTRONICS EQUIP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211361694.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-02
Publication Date
2025-06-10
Estimated Expiration
2042-11-02

AI Technical Summary

Technical Problem

The prior art consumes a lot of manpower, slow labeling speed and low labeling accuracy in electromagnetic data labeling.

Method used

The electromagnetic data annotation method based on human-computer hybrid intelligence is adopted, and the electromagnetic data is automatically separated and annotated by deep clustering and deinterleaving separation method, and the similarity matrix between different types of electromagnetic data after the separation is calculated. Finally, the field experts conduct manual review based on the similarity matrix and label confidence.

Benefits of technology

It realizes the rapid data labeling capability with machine labeling as the main and expert review as the auxiliary, improves the quality and efficiency of electromagnetic data labeling, and solves the problem of adaptive fine labeling of electromagnetic data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115659193B_ABST
    Figure CN115659193B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of electromagnetic data technology, and discloses an electromagnetic data annotation method and system based on human-machine hybrid intelligence. The annotation method automatically separates and annotates electromagnetic data by using a deep clustering deinterleaving method, calculates a similarity matrix between different types of electromagnetic data after separation and annotation, and finally, domain experts manually review the labels of the electromagnetic data after separation and annotation according to the similarity matrix and label confidence, so as to realize electromagnetic data annotation mainly based on machine annotation and supplemented by expert review. The present invention solves the problems existing in the prior art, such as large labor consumption, slow annotation speed, and low annotation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electromagnetic data, and specifically to an electromagnetic data annotation method and system based on human-machine hybrid intelligence. Background Art

[0002] Electromagnetic data with good annotation information can be used to train algorithm models such as modulation recognition, model recognition, state recognition, and user intention recognition of radar communication signals. This has led to a multiple expansion of the demand for electromagnetic data containing annotations. At the same time, the speed of electromagnetic data annotation determines the iteration speed of electromagnetic model development. Improving the efficiency and accuracy of electromagnetic data annotation is a key issue in electromagnetic data processing.

[0003] In recent years, due to its excellent performance, deep learning technology has begun to be widely used in fields such as vision, natural language processing, and unmanned driving. The demand for annotation of massive amounts of data has promoted the development of annotation technology. Large domestic and foreign artificial intelligence companies even have data annotation teams consisting of hundreds or thousands of people. However, these annotation technologies are all for image, video, voice, and text data, and there are significant differences in data format, data volume, and annotation logic compared to electromagnetic data annotation. Traditional electromagnetic data annotation methods rely on the experience and knowledge of domain experts to perform fine data annotation on electromagnetic data pulse by pulse. All these high-quality electromagnetic annotation data are obtained through manual annotation methods, requiring a huge amount of man-hours. Domain experts need to continuously perform repetitive labor on massive amounts of electromagnetic data, which is prone to fatigue and makes it difficult to ensure data consistency. The accuracy of electromagnetic annotation also largely depends on the professional level of the annotator. Therefore, developing an intelligent, efficient, and user-friendly electromagnetic data annotation method and its corresponding device according to the characteristics of electromagnetic data has good application prospects. Summary of the Invention

[0004] To overcome the deficiencies of the prior art, the present invention provides an electromagnetic data annotation method and system based on human-machine hybrid intelligence, which solves the problems in the prior art such as high labor consumption, slow annotation speed, and low annotation accuracy.

[0005] The technical solution adopted by the present invention to solve the above problems is:

[0006] An electromagnetic data annotation method based on human-machine hybrid intelligence automatically separates and annotates electromagnetic data using a deep clustering deinterleaving separation method, calculates the similarity matrix between different types of electromagnetic data after separation and annotation, and finally the domain expert manually reviews the labels of the electromagnetic data after separation and annotation according to the similarity matrix and label confidence, so as to achieve electromagnetic data annotation mainly by machine annotation and supplemented by expert review.

[0007] As a preferred technical solution, it includes the following steps:

[0008] S1, Electromagnetic data preprocessing: First, slice the electromagnetic data to form electromagnetic data frames. Then, remove outliers from each frame of the electromagnetic data frame one by one. Next, perform normalization processing on the electromagnetic data frame and output the preprocessed electromagnetic data frame.

[0009] S2, Deep clustering for deinterleaving and separation annotation: Input the electromagnetic data frames preprocessed in step S1 into a deep long short-term memory network to learn metric feature representations, and perform unsupervised clustering on the learned metric feature representations. Then, use the clustering results as labels and feed them back into the deep long short-term memory network for supervised learning. Repeat this process iteratively until the deep long short-term memory network converges, and output the separation annotation labels and label confidence levels.

[0010] S3, Similarity matrix calculation: Calculate the similarity between different types of electromagnetic data in the current electromagnetic data frame and historical electromagnetic data frames based on the annotation labels generated in step S2, and generate a similarity matrix. Here, the historical electromagnetic data frames refer to all electromagnetic data frames before the current processing time.

[0011] S4, Domain expert review: Manually review and correct the labels of the electromagnetic data after separation annotation based on the label confidence levels generated in step S2 and the similarity matrix generated in step S3.

[0012] As a preferred technical solution, step S1 includes the following steps:

[0013] S11, Electromagnetic data frame generation: Slice the original electromagnetic data sequence using a fixed window to form electromagnetic data frames. Here, the size of the fixed window is w, where w ≥ 1000.

[0014] S12, Outlier filtering: Use an outlier filtering method to remove outliers from the electromagnetic data in the electromagnetic data frame. Here, the outlier filtering method can be the interquartile range outlier filtering method, the z-score method, the density clustering method, or the isolation forest outlier detection method, etc.

[0015] S13, Normalization processing: Calculate the mean μ and standard deviation σ of the electromagnetic data processed in step S12, and perform normalization processing according to I * =(I - μ) / σ. Here, I represents the electromagnetic data processed in step S12, and I * represents the normalized electromagnetic data;

[0016] Then, for each electromagnetic data p i in the electromagnetic data frame after normalization processing, use a one-sided historical time window to slice it with a step of 1, and output the historical window representation x i of the electromagnetic data p i =[p i-m+1 ,...,p i-1 ,pi , among the first m electromagnetic data, if the number of electromagnetic data within the historical time window on one side is less than m, it is filled with 0; where m represents the window size, m≥8, i represents the number of electromagnetic data within the electromagnetic data frame, and 0≤i≤w - 1.

[0017] As a preferred technical solution, in step S12, if the outlier filtering method is used for outlier filtering, the following method is implemented: Arrange the electromagnetic data frames output in step S11 in ascending order of values, and record the value at the 25% position of the electromagnetic data in the arranged electromagnetic data frames as Q 1 , and record the value at the 75% position of the electromagnetic data in the arranged electromagnetic data frames as Q 3 , and the difference between the two is recorded as the interquartile range IQR = Q 3 - Q 1 , and the eigenvalues of the electromagnetic data less than Q 1 - 3IQR or greater than Q 3 + 3IQR are determined as outliers and removed.

[0018] As a preferred technical solution, step S2 includes the following steps:

[0019] S21, construct a deep long short - term memory network. The structure of the deep long short - term memory network is: The input layer is determined by the window size m in step S13 and the input feature number k of the pre - processed electromagnetic data frame, with a shape of m×k; the number of layers l of the long short - term memory layer, l≥2, the number of units in each layer is determined by the window size m, and the output space dimension of each unit takes d, d≥32; the output layer is a fully - connected layer, and the number of neurons in the fully - connected layer is determined by the dimension of the metric feature to be learned;

[0020] S22, forward - propagation unsupervised clustering: Represent the historical window x i (0≤i≤w - 1) of the electromagnetic data in the electromagnetic data frame is input into the deep long short - term memory network structure f θ designed in step S21, and the learned metric feature representation f θ (x i ) is output. The metric feature representation f θ (x i ) is input into the classifier g C , and by iteratively optimizing the objective function:

[0021]

[0022] Update the classifier weight C t , and predict the pseudo - label corresponding to the signal where t represents the current iteration round, L f represents the classifier loss function, θ tThe weight matrix of the deep long short-term memory network at the current iteration round;

[0023] S23, Backpropagation supervised metric feature representation learning: The pseudo-labels output in step S22 along with the electromagnetic data frame x i , are input into the deep long short-term memory network f θ , and the objective function is iteratively optimized:

[0024]

[0025] Update the weight matrix θ of the deep long short-term memory network structure; t where, y t-1 is the clustering result label of the previous forward propagation, and L b represents the backpropagation loss function; the clustering objective function is set as:

[0026]

[0027] where, x i , x j belong to the same class after the previous round of clustering, x k belongs to the nearest neighbor class of x i , α is the learning hyperparameter, α represents the margin, d θ is the learned metric distance, neighbor(i) is the other electromagnetic data in the class to which x i belongs;

[0028] S24, Iterative optimization: Repeat step S22 and step S23 until the clustering objective function converges or reaches the set number of iteration rounds, to obtain the final weight matrix θ of the deep long short-term memory network structure * and the classifier weight C * , and finally output the corresponding label of the electromagnetic data and the label confidence g C* (f θ* (x i )); Denote p b = g C* (f θ* (x i )); where, the value range of p b is 0 ≤ p b ≤ 1.

[0029] As a preferred technical solution, step S3 includes the following steps:

[0030] S31, Similarity distance calculation: Calculate the distribution distance between different electromagnetic data types as the similarity distance between types; Denote the electromagnetic data distribution corresponding to the electromagnetic label r as P r, The electromagnetic data distribution corresponding to the electromagnetic tags is P s , Calculate the similarity distance d between the two types using information cross-entropy:

[0031]

[0032] where z is the sampling value of the distribution P r and N is the number of sampling values;

[0033] S32, Similarity calculation: Convert the similarity distance between the two types according to the similarity conversion formula:

[0034]

[0035] Calculate the similarity between the two types;

[0036] where d is the similarity distance calculated in step S31;

[0037] S33, Similarity matrix calculation: Repeat steps S31 and S32 to calculate the similarity between every two types, and output the similarity matrix.

[0038] As a preferred technical solution, step S4 includes the following steps:

[0039] S41, Expert correction: Visualize the electromagnetic data frames output in step S1 to domain experts with different markings according to the tags output in step S2, and display the machine-labeled tags given in step S2 within the current electromagnetic data frame in descending order of tag confidence in the information column; Domain experts review each type of signal pattern given by the machine labeling according to professional knowledge and experience, and use manual labeling tools to select and correct the labeled tags for samples with labeling errors;

[0040] S42, Tag assignment: Domain experts further modify the digital tags to corresponding character tags according to the training purpose;

[0041] S43, Similar category merging: Visualize and display the electromagnetic similarity map according to the similarity matrix output in step S3, and domain experts merge the electromagnetic data with high similarity belonging to the same type.

[0042] As a preferred technical solution, in step S42, domain experts modify the corresponding character tags according to the training purpose of the data set, modify them to corresponding device type character tags according to the training purpose of model recognition, modify them to corresponding status character tags according to the training purpose of status recognition, and modify them to corresponding modulation character tags according to the training purpose of modulation type recognition.

[0043] As a preferred technical solution, it further includes the following steps:

[0044] S5. Repeat steps S2 to S4 until all the electromagnetic data frames output in step S1 are processed.

[0045] S6. Dataset scoring: Domain experts score the quality of the electromagnetic data and the annotation quality completed in step S5, and output a training dataset in a standard format.

[0046] An electromagnetic data annotation system based on human-machine hybrid intelligence for implementing the electromagnetic data annotation method based on human-machine hybrid intelligence, including the following modules connected in sequence:

[0047] Electromagnetic data preprocessing module: Firstly, slice the electromagnetic data to form electromagnetic data frames, then remove the outliers in each frame of the electromagnetic data frames one by one, and then perform standardization processing on the electromagnetic data frames, and output the preprocessed electromagnetic data frames.

[0048] Deinterleaving and separating annotation module: Input the electromagnetic data frames preprocessed by the electromagnetic data preprocessing module into a deep long short-term memory network to learn the metric feature representation, and perform unsupervised clustering on the learned metric feature representation; then use the clustering result as a label and send it back to the deep long short-term memory network for supervised learning; iterate in this way until the deep long short-term memory network converges, and output the separated annotation labels and label confidence levels.

[0049] Similarity matrix calculation module: Calculate the similarity between different types of electromagnetic data in the current electromagnetic data frame and the historical electromagnetic data frames according to the annotation labels generated by the deinterleaving and separating annotation module, and generate a similarity matrix; where the historical electromagnetic data frames refer to all the electromagnetic data frames before the current processing time.

[0050] Domain expert review module: Manually review and correct the labels of the electromagnetic data after separation annotation according to the label confidence levels generated by the deinterleaving and separating annotation module and the similarity matrix generated by the similarity matrix calculation module.

[0051] Compared with the prior art, the present invention has the following beneficial effects:

[0052] The present invention uses the electromagnetic data machine annotation technology of deep clustering to complete autonomous clustering of electromagnetic targets that cannot be distinguished from traditional feature dimensions or cannot be distinguished by unified manual rules in the context high-dimensional feature representation space, provides a fast data annotation ability mainly based on machine annotation and supplemented by expert review, solves the problem of adaptive fine annotation of electromagnetic data, and improves the quality and efficiency of electromagnetic data annotation. Description of the Drawings

[0053] Figure 1 It is a schematic flow chart of an electromagnetic data annotation method based on human-machine hybrid intelligence according to the present invention;

[0054] Figure 2 One of the partial enlarged views of Figure 1 ;

[0055] Figure 3 One of the partial enlarged views of Figure 1 ;

[0056] Figure 4 The structural schematic diagram of an electromagnetic data annotation system based on human - machine hybrid intelligence according to the present invention;

[0057] Figure 5 One of the partial enlarged views of Figure 4 ;

[0058] Figure 6 One of the partial enlarged views of Figure 4 ; Specific embodiments

[0059] The following combines the embodiments and the attached drawings to make a further detailed description of the present invention, but the implementation manners of the present invention are not limited thereto.

[0060] Embodiment 1

[0061] As Figures 1 to 6 shown, the purpose of the present invention is to save labor costs, speed up the annotation speed, improve the annotation quality, disclose an electromagnetic data annotation method and system based on human - machine hybrid intelligence, simplify, standardize and intelligentize the electromagnetic annotation work, minimize unnecessary simple repetitive labor for electromagnetic field experts, improve the annotation efficiency and quality, provide training support for modulation recognition, model recognition, state recognition, behavior intention recognition, etc. of intelligent electromagnetic data processing, and improve the development efficiency of electromagnetic field algorithm models.

[0062] The method of the present invention first pre - processes the original electromagnetic data to be annotated, de - interleaves and separates the annotated electromagnetic data frames, calculates the similarity matrix between different signals after separation, and finally pushes the separated electromagnetic data, corresponding labels, annotation confidence degrees and similarity matrix to the human - machine interaction interface for review, correction and scoring by domain experts. The method and system of the present invention utilize the electromagnetic data machine annotation technology of deep clustering to autonomously cluster electromagnetic targets that cannot be distinguished from traditional feature dimensions or cannot be distinguished by unified manual rules in the context high - dimensional feature representation space, provide a fast data annotation ability with machine annotation as the main and expert review as the auxiliary, solve the problem of adaptive fine annotation of electromagnetic data, and improve the quality and efficiency of electromagnetic data annotation.

[0063] To achieve the above object, the present invention discloses an electromagnetic data annotation method and system based on human-machine hybrid intelligence. An electromagnetic data annotation method based on human-machine hybrid intelligence first preprocesses the original electromagnetic data to be annotated, deinterleaves and separates the annotated electromagnetic data frames, calculates the similarity matrix between different signals after separation, and finally pushes the separated electromagnetic data, corresponding labels, annotation confidence, and similarity matrix to the human-computer interaction interface for review, correction, and scoring by domain experts.

[0064] Further, the specific implementation method is as follows:

[0065] Step S1, electromagnetic data preprocessing: First, perform equal-length slicing on the electromagnetic data to form electromagnetic data frames, and then remove the outliers in each electromagnetic data frame frame by frame; then perform standardization processing on the electromagnetic data frames, and output the preprocessed electromagnetic data frames.

[0066] Step S2, deep clustering deinterleaving separation annotation: Input the electromagnetic data frames preprocessed in Step S1 into a deep long short-term memory network to learn the metric feature representation, and perform unsupervised clustering on the learned metric feature representation; use the clustering result as a label, and then send it back to the deep long short-term memory network for supervised learning; iterate in this way until the model converges, and output the separated annotation labels and label confidence.

[0067] The deinterleaving separation annotation method referred to in this method is characterized in that it can cluster and separate electromagnetic feature patterns in a high-dimensional metric feature representation space, which is beneficial for domain experts to analyze the working characteristics and context rules of electromagnetic signals.

[0068] The deinterleaving separation annotation method referred to in this method can be the deep clustering method mentioned in Step S2, or other signal separation annotation methods in the field of electromagnetic signal processing, including but not limited to dynamic association method, histogram statistics method, pulse repetition interval transformation method, k-means clustering sorting method, and density clustering sorting method, etc.

[0069] Step S3, similarity matrix calculation: According to the annotation labels generated in Step S2, calculate the similarity between different types of electromagnetic data in the current electromagnetic data frame and the historical electromagnetic data frames pairwise, and generate a similarity matrix; the historical electromagnetic data frames refer to all electromagnetic data frames before the current processing time.

[0070] The similarity calculation referred to in this method can be the distribution distance calculated using information cross-entropy mentioned in Step S3, or the distribution distance calculated using Mahalanobis distance and chi-square distance.

[0071] Step S4, Domain Expert Review: Push the label confidence output from Step S2 and the similarity matrix output from Step S3 to the human-computer interaction interface, and let the domain expert review, confirm, and manually correct the labels of the electromagnetic data after separation and annotation.

[0072] Step S5, Repeat Steps S2 - S4 until all electromagnetic data frames output from Step S1 are processed.

[0073] Step S6, Dataset Scoring: The domain expert scores the quality of the electromagnetic data and the annotation quality completed in Step S5, and outputs a training dataset in a standard format.

[0074] Among them, de-interleaving separation and annotation refers to the process of separating the electromagnetic data sequences corresponding to signal sources from randomly interleaved electromagnetic data; deep clustering refers to jointly using a deep neural network and an unsupervised clustering method to perform de-interleaving separation and annotation on electromagnetic data; high-dimensional metric features refer to the metric features learned by the neural network in deep clustering that are higher than the original feature dimension of the electromagnetic data; the number of neural network layers of the deep long short-term memory network structure is not less than 2 layers.

[0075] Further, the specific method of Step S1 is as follows:

[0076] Step S11, Electromagnetic Data Frame Generation: Slice the original electromagnetic data sequence using a fixed window to form electromagnetic data frames, where the fixed window size is w, and w ≥ 1000. Output the generated electromagnetic data frames frame by frame to the following steps for machine annotation and expert review.

[0077] The length of the electromagnetic data frames generated in Step S11 can be a fixed number of samples or a fixed time length.

[0078] Step S12, Outlier Filtering: Arrange the electromagnetic data in the electromagnetic data frames output from Step S11 in ascending order of numerical value, and take the value corresponding to the 25% (lower quartile) as Q 1 , and take the value corresponding to the 75% (upper quartile) as Q 3 , and the difference between the two is denoted as the interquartile range IQR = Q 3 - Q 1 , and determine the eigenvalue that satisfies less than Q 1 - 3IQR or greater than Q 3 + 3IQR as an outlier and remove it.

[0079] The outlier filtering method in Step S12 includes but is not limited to the quartile outlier filtering method described above, and can also be outlier detection methods such as z-score, density clustering, or isolation forest.

[0080] Step S13. Standardization processing: Calculate the mean μ and standard deviation σ of the electromagnetic data processed in step S12, and perform standardization processing according to I * =(I - μ) / σ, and output the standardized electromagnetic data frame. For each electromagnetic data p i in the standardized electromagnetic data frame, use a one-sided historical time window to slice with a step of 1, and output the historical window of the electromagnetic data p i which is represented as x i =[p i-m ,...,p i-1 ,p i . Among the first m electromagnetic data, if the number of electromagnetic data in the one-sided historical time window is less than m, fill it with 0. Among them, the window slicing size is m, m≥8, i represents the number of the electromagnetic data in the electromagnetic data frame, and 0≤i≤w - 1.

[0081] Furthermore, the specific method of step S2 is as follows:

[0082] Step S21. Construct a deep long short-term memory network: The input layer is determined by the window slicing size m and the number of input features k of the feature vector, and its shape is m×k. The number of layers l of the long short-term memory layer, l≥2, the number of units in each layer is determined by the window slicing size m, and the output space dimension of each unit takes d (d≥32) dimensions. The output layer is a fully connected layer, and the number of neurons in the fully connected layer is determined by the dimension of the metric feature representation to be learned. The activation function of each neuron can be a sigmoid function, or other activation functions such as hyperbolic (tanh), rectified linear unit (ReLu), etc.

[0083] The input features referred to in step S21 can be the carrier frequency, pulse repetition frequency, pulse width, pulse amplitude, azimuth angle, etc. of the full pulse description word, can be the phase component, in-phase component, quadrature component of the original waveform signal, or other expert features extracted from the original waveform, such as the time-frequency image after short-time Fourier transform, the wavelet image after wavelet transform.

[0084] The deep long short-term memory network constructed in step S21 can also be a one-dimensional convolutional neural network, an autoencoder (AutoEncoder), etc. with time series learning ability.

[0085] Step S22. Forward propagation unsupervised clustering: Input the electromagnetic data frame x i output in step S13 into the deep long short-term memory network structure f θ designed in step S21, and output the learned metric feature representation f θ (x i ). Input the metric feature representation f θ (x i ) into the classifier g C, optimize the objective function through iteration:

[0086]

[0087] Update the classifier weight C t , and predict the pseudo-label corresponding to the signal The selection of the classifier includes but is not limited to k-means spatial clustering, hierarchical clustering, density clustering, etc.

[0088] Step S23, Backpropagation Supervised Metric Feature Representation Learning: Input the pseudo-label output by step S22 i , together with the electromagnetic data frame x θ , into the deep long short-term memory network f

[0089]

[0090] Update the weight matrix θ of the deep long short-term memory network structure t . Among them, y t-1 is the clustering result label of the previous forward propagation. The clustering objective function is set as:

[0091]

[0092] Among them, x i , x j belong to the same class after the previous round of clustering, and x k belongs to the nearest neighbor class of x i . α is the learning hyperparameter, representing the heap interval. d θ is the learned metric distance.

[0093] Step S24, Iterative Optimization: Repeat step S22 and step S23 until the clustering objective function converges or reaches the set number of iteration rounds, and obtain the final weight matrix θ of the deep long short-term memory network structure * and the classifier weight C * . Finally, output the corresponding digital label of the metric feature representation and the confidence g of the metric feature representation C* (f θ* (x i )), denoted as p b = g C* (f θ* (x i )); among them, the value range of p b is 0 ≤ p b ≤ 1.

[0094] Furthermore, the specific method of step S3 is:

[0095] Step S31, Similarity Distance Calculation: Calculate the electromagnetic data distribution P corresponding to the electromagnetic tag r r and the electromagnetic data distribution P corresponding to the electromagnetic tag s s according to the information cross-entropy formula:

[0096]

[0097] to obtain the similarity distance between the two types, where z is the sampling value of the distribution P r and N is the number of sampling values.

[0098] Step S32, Similarity Calculation: Calculate the similarity between the two types according to the similarity conversion formula for the similarity distance between the two types:

[0099]

[0100] where d is the similarity distance calculated in Step S31.

[0101] Step S33, Similarity Matrix Calculation: Repeat Steps S31 and S32 to calculate the similarity between all pairs of types, and output the similarity matrix.

[0102] Furthermore, the specific method of Step S4 is as follows:

[0103] Step S41, Expert Correction: Visualize the electromagnetic data frames output in Step S1 with different markings according to the tags output in Step S2 to domain experts, and display the machine-annotated tags given in Step S2 within the current electromagnetic data frame in the information column in descending order of tag confidence. Domain experts review each type of signal pattern given by the machine annotation according to their professional knowledge and experience, and use manual annotation tools to select and correct the marked tags for samples with annotation errors.

[0104] Step S42, Tag Assignment: Domain experts further modify the digital tags to character tags corresponding to the types according to the training purpose of the data set and professional knowledge.

[0105] The character tags corresponding to the types modified according to the training purpose of the data set mentioned in Step S42 can be modified to character tags of the corresponding device types according to the training purpose of model recognition; or can be modified to character tags such as search, tracking, tracking while searching, simultaneous searching and tracking, etc. according to the training purpose of state recognition; or can be modified to character tags such as BPSK, QPSK, 64QAM, etc. according to the training purpose of modulation type recognition.

[0106] Step S43, Merging of the Same Type: Visualize the electromagnetic similarity map according to the similarity matrix output in Step S3, and domain experts merge the electromagnetic data with high similarity belonging to the same type.

[0107] An electromagnetic data annotation device based on human - machine hybrid intelligent annotation includes an electromagnetic data pre - processing module, a de - interleaving annotation separation module, a similarity calculation module, and a domain expert review module.

[0108] Electromagnetic data pre - processing module: Receives the electromagnetic data to be annotated, generates electromagnetic data frames, filters outliers, performs normalization processing, and outputs electromagnetic data frames with a fixed window size. The output end is connected to the de - interleaving annotation separation module.

[0109] De - interleaving annotation separation module: Receives the electromagnetic data frames output by the electromagnetic data pre - processing module, uses the de - interleaving separation algorithm to merge and classify the electromagnetic data according to similarity metrics and evaluation criteria, discovers the internal pattern structure of the electromagnetic data, forms digital tags, and outputs digital tags and tag confidence levels. The output end is connected to the similarity calculation module and the domain expert review module.

[0110] Similarity calculation module: Receives the current - frame machine - annotated data output by the de - interleaving annotation separation module and the historical annotated data output by the domain expert review module, calculates the similarity between different types pairwise, and outputs a similarity matrix. The input end is connected to the de - interleaving annotation separation module and the domain expert review module, and the output end is connected to the domain expert review module. The similarity calculation module can also integrate a type recognition algorithm to sequentially perform type recognition on the received separated annotation data, receive the reviewed historical annotated data output by the domain expert review module for online learning and updating, and output the recognized character tags.

[0111] Domain expert review module: Receives the current - frame machine - annotated data, corresponding digital tags, annotation confidence levels output by the de - interleaving annotation separation module, and the similarity matrix output by the similarity calculation module, presents them to the domain expert in a visual human - machine interaction form, is used to complete operations such as expert correction of machine - annotated data, digital tag assignment, manual merging of highly similar types, dataset scoring, etc., and outputs the standard dataset completed by expert review to the electromagnetic time - series database. The input end is connected to the de - interleaving annotation separation module and the similarity calculation module, and the output end is connected to the electromagnetic time - series database.

[0112] The present invention provides a method and device for electromagnetic data annotation based on human-machine hybrid intelligence. This method utilizes the machine annotation technology of electromagnetic data based on deep clustering to autonomously cluster electromagnetic targets that cannot be distinguished from traditional feature dimensions or by unified manual rules in the context high-dimensional feature representation space, providing fast data annotation capabilities with machine annotation as the main and expert review as the auxiliary, and solving the problem of adaptive fine annotation of electromagnetic data. In the same electromagnetic data annotation task, compared with the traditional manual data annotation method, this method saves hundreds or thousands of times the annotation time while ensuring the quality of data annotation, significantly improving the annotation efficiency for massive electromagnetic data, and solving the problem of high-quality data sources required for rapid iterative learning of algorithms in the electromagnetic field.

[0113] Embodiment 2

[0114] As Figures 1 to 6 shown, as a further optimization of Embodiment 1, on the basis of Embodiment 1, this embodiment further includes the following technical features:

[0115] The implementation process of a method for electromagnetic data annotation based on human-machine hybrid intelligence in this embodiment is as Figure 1 shown. The specific implementation steps are as follows:

[0116] S11 Generation of electromagnetic data frames: The original electromagnetic data sequence is sliced using a fixed window to form electromagnetic data frames, where the size of the fixed window is w, and w≥1000. The generated electromagnetic data frames are output frame by frame to the following steps for machine annotation and expert review.

[0117] S12 Outlier filtering: Arrange the electromagnetic data frames output by S11 in ascending order of values, and take the value corresponding to the 25% (lower quartile) as Q 1 , and take the value corresponding to the 75% (upper quartile) as Q 3 , and the difference between the two is denoted as the interquartile range IQR = Q 3 -Q 1 . Feature values that satisfy being less than Q 1 -3IQR or greater than Q 3 +3IQR are determined as outliers and removed.

[0118] S13 Standardization processing: Calculate the mean μ and standard deviation σ of the electromagnetic data processed by S12, and perform standardization processing according to I * =(I - μ) / σ. For each electromagnetic data p i within the standardized electromagnetic data frame, use a one-sided historical time window to slice with a step of 1, and output the historical window of the electromagnetic data p i as x i =[p i-m ,...,p i-1 ,pi Among the first m electromagnetic data, if the number of electromagnetic data within the historical time window on one side is less than m, it is filled with 0. Wherein, the window size is m, m≥8, i represents the number of electromagnetic data within the electromagnetic data frame, and 0≤i≤w - 1.

[0119] S21 Construct a deep long short - term memory network: In this embodiment, the signal full - pulse description word features, namely carrier frequency, pulse repetition frequency, pulse width, pulse amplitude, and arrival angle, are used as input features. The input layer is determined by the window size m = 64 and the number of input feature k = 5 of the feature vector, with a shape of 64×5. In this embodiment, a 2 - layer long short - term memory network is selected, and the unit output space dimension of both 2 - layer long short - term memory layers is set to 32 dimensions. A dropout layer is added before the input of each long short - term memory layer, and the dropout probability is set to 0.5. The state outputs at all times of the first layer are used as inputs and connected to the second - layer long short - term memory layer. Only the state output at the last time of the second - layer long short - term memory layer is used as the input of the next fully - connected layer. The number of neurons in the fully - connected output layer is set to 64 according to the dimension of the metric feature representation to be learned. The activation function is set to the sigmoid function.

[0120] S22 Forward - propagation unsupervised clustering: The historical window representation x of the electromagnetic data in the electromagnetic data frame output by S13 i (0≤i≤w - 1) is input into the deep long short - term memory network structure f designed in S21 θ , and the learned metric feature representation f θ (x i ) is output. The clustering classifier selects k - means spatial clustering. The metric feature representation f θ (x i ) is input into the classifier g C , and the objective function is iteratively optimized:

[0121]

[0122] The weights C of the classifier are updated t , and finally the pseudo - label corresponding to the signal x i is given where C t is the cluster - center matrix, k is the number of clustering targets, and d = 64 is the dimension of the metric feature representation. is the one - hot encoding representation, satisfying

[0123] S23 Back - propagation supervised metric feature representation learning: The pseudo - label output by S22 along with the pre - processed signal x i is input into the deep long short - term memory network f θ for supervised learning, and the clustering objective function is iteratively optimized:

[0124]

[0125] Update the weight matrix θ of the deep long short - term memory network structure t . Among them, x i , x j belong to the same category after the previous round of clustering, and x k belongs to the nearest neighbor class of x i . α is a hyperparameter representing the heap interval, taking 0.2.

[0126] Repeat S22 and S23 until the clustering objective function converges or reaches the set number of iteration rounds to obtain the final weight matrix θ of the deep long short - term memory network structure * and the classifier weight C * . The final output and g C* (f θ* (x i )) are the digital label given by machine annotation and the corresponding annotation confidence.

[0127] S30 Similarity matrix calculation: Calculate the similarity matrix M according to the similarity formula:

[0128]

[0129] Among them, M rs is the element of the similarity matrix M, corresponding to the similarity between electromagnetic data types r and s, r, s ∈ L, where L is the set of all current annotation labels, and d rs is the similarity distance between electromagnetic data types r and s, which is calculated from the corresponding electromagnetic data distributions P r and P s according to the information cross - entropy formula:

[0130]

[0131] The calculation is as follows. Among them, z is the sampling value of the distribution P r , and N is the number of sampling values.

[0132] Domain expert review in S40: Visualize the machine annotation results with different shapes representing different electromagnetic type data to the domain expert, and display the digital labels given by the machine annotation within the current frame in the information column in descending order of annotation confidence. The domain expert reviews the machine annotation results based on experience and knowledge, and uses a manual annotation tool to select and correct the annotation information for samples with annotation errors; according to the training purpose of the dataset in this embodiment for model identification, the domain expert modifies the digital labels to corresponding character labels, including type A, type B, type C, etc.; display the similarity matrix in the form of an electromagnetic similarity map, and the domain expert merges electromagnetic signals with high similarity belonging to the same type by double-clicking the edge between the two type nodes.

[0133] S50 Traverse and process electromagnetic data frames: Repeat the above steps until all electromagnetic data frames output by S11 are processed.

[0134] S60 Dataset scoring: Mark and score the output dataset according to 1-5 stars, and output the standard dataset.

[0135] Figure 4 The figure shows a system framework diagram of an electromagnetic data annotation device based on human-machine hybrid intelligence according to this embodiment, including an M101 electromagnetic data preprocessing module, an M102 deinterleaving annotation separation module, an M103 similarity calculation module, and an M104 domain expert review module.

[0136] M101 electromagnetic data preprocessing module: Load the original electromagnetic data to be annotated from the local electromagnetic time series database according to this embodiment, generate electromagnetic data frames with a fixed window size of 1000, complete outlier filtering and normalization processing frame by frame, and output electromagnetic data frames with a shape of 1000×64×5. The output end is connected to the M102 deinterleaving annotation separation module.

[0137] M102 deinterleaving annotation separation module: Receive the electromagnetic data frame with a shape of 1000×64×5 output by the M101 electromagnetic data preprocessing module, use deep clustering to deinterleave and separate the electromagnetic data frame, and annotate digital labels, and output the separated electromagnetic data and the corresponding digital labels and label confidences. The output end is connected to the M103 similarity calculation module and the M104 domain expert review module.

[0138] M103 similarity calculation module: Receive the current frame machine annotation data output by the M102 deinterleaving annotation separation module and the historical annotation data output by the M104 domain expert review module, calculate the similarity between different types, and output a similarity matrix. The input end is connected to the M102 deinterleaving annotation separation module and the M104 domain expert review module, and the output end is connected to the M104 domain expert review module.

[0139] M104 Domain Expert Review Module: Receives the current frame electromagnetic data, corresponding digital tags, annotation confidence levels output by the M102 Deinterleaving Annotation Separation Module, and the similarity matrix output by the M103 Similarity Calculation Module, and presents them to domain experts in the form of time series graphs, information lists, and electromagnetic similarity maps for operations such as expert correction of machine-annotated electromagnetic data, digital tag assignment, merging of the same type, and dataset scoring. The standard dataset after expert review is output to the local electromagnetic time series database for storage. The input end is connected to the M102 Deinterleaving Annotation Separation Module and the M103 Similarity Calculation Module, and the output end is connected to the M103 Similarity Calculation Module and the local electromagnetic time series database.

[0140] Figures 1 to 6 Shows the possible architectures, functions, and operations of the systems, methods, and computer program products of various embodiments of the present invention. Each block in the flowchart or system framework diagram may represent a module, program segment, or part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical functions. It should be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, the M102 module and the M103 module may execute in parallel or in the reverse order, depending on the functions involved. It should be noted that each block and combination of blocks in the flowchart or system framework diagram can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0141] As described above, the present invention can be preferably implemented.

[0142] All the features disclosed in all the embodiments in this specification, or all the steps in the methods or processes implicitly disclosed, except for mutually exclusive features and / or steps, can be combined and / or extended, replaced in any way.

[0143] As mentioned above, it is only a preferred embodiment of the present invention, and there is no any form of limitation to the present invention. According to the technical essence of the present invention, any simple modification, equivalent replacement, and improvement made to the above embodiments within the spirit and principle of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. An electromagnetic data annotation method based on human-machine hybrid intelligence, characterized in that, using the deep clustering deinterleaving separation method to automatically separate and annotate electromagnetic data, calculating the similarity matrix between different types of electromagnetic data after separation and annotation, and finally having domain experts manually review the labels of the electromagnetic data after separation and annotation according to the similarity matrix and label confidence, so as to realize the electromagnetic data annotation mainly based on machine annotation and supplemented by expert review; including the following steps: S1, Electromagnetic data preprocessing: First, slice the electromagnetic data to form electromagnetic data frames, then remove the outliers in each electromagnetic data frame one by one, and then perform standardization processing on the electromagnetic data frames, and output the preprocessed electromagnetic data frames; S2, Deep clustering deinterleaving separation annotation: Input the electromagnetic data frames preprocessed in step S1 into a deep long short-term memory network to learn the metric feature representation, and perform unsupervised clustering on the learned metric feature representation; then use the clustering result as a label and send it back to the deep long short-term memory network for supervised learning; iterate in this way until the deep long short-term memory network converges, and output the separation annotation label and label confidence; S3, Similarity matrix calculation: According to the annotation labels generated in step S2, calculate the similarity between different types of electromagnetic data in the current electromagnetic data frame and the historical electromagnetic data frames, and generate a similarity matrix; among them, the historical electromagnetic data frames refer to all electromagnetic data frames before the current processing time; S4, Domain expert review: According to the label confidence generated in step S2 and the similarity matrix generated in step S3, manually review and correct the labels of the electromagnetic data after separation and annotation.

2. The electromagnetic data annotation method based on human-machine hybrid intelligence according to claim 1, characterized in that, step S1 includes the following steps: S11, Electromagnetic data frame generation: Slice the original electromagnetic data sequence using a fixed window to form electromagnetic data frames; where the size of the fixed window is , ; S12, Outlier filtering: Use the outlier filtering method to remove the outliers in the electromagnetic data in the electromagnetic data frame; among them, the outlier filtering method is the quartile outlier filtering method, the z-score method, the density clustering method or the outlier detection method such as the isolation forest; S13, Standardization: Calculate the mean of the electromagnetic data processed in step S12 and the standard deviation , and perform standardization according to ; where represents the electromagnetic data processed in step S12, represents the standardized electromagnetic data; Then, for each piece of electromagnetic data in the standardized electromagnetic data frame , perform window slicing with a single-sided historical time window in steps of 1, and output the historical window representation of the electromagnetic data . For the first pieces of electromagnetic data, if the number of electromagnetic data within the single-sided historical time window is insufficient , pad with 0s; where represents the window size, , , represents the number of the electromagnetic data in the electromagnetic data frame, .

3. The electromagnetic data annotation method based on human-machine hybrid intelligence according to claim 2, characterized in that, In step S12, if the outlier filtering method is adopted for outlier filtering, the following method is implemented: Arrange the electromagnetic data frames output in step S11 in ascending order of values, and record the value at the 25% position of the electromagnetic data in the arranged electromagnetic data frames as , record the value at the 75% position of the electromagnetic data in the arranged electromagnetic data frames as , and the difference between the two is recorded as the interquartile range . The eigenvalue of the electromagnetic data that is less than or greater than is determined as an outlier and removed.

4. The electromagnetic data annotation method based on human-machine hybrid intelligence according to claim 3, characterized in that, step S2 includes the following steps: S21. Construct a deep long short - term memory network. The structure of the deep long short - term memory network is as follows: The input layer is determined by the windowing size in step S13 and the number of input features k of the pre - processed electromagnetic data frame, with a shape of ; The number of long short - term memory layers , . The number of units in each layer is determined by the windowing size . The output space dimension of each unit is taken as , ; The output layer is a fully - connected layer, and the number of neurons in the fully - connected layer is determined by the dimension of the metric feature to be learned; S22, Forward Propagation Unsupervised Clustering: Represent the historical window of electromagnetic data in the electromagnetic data frame ( ) Input it into the designed deep long short-term memory network structure in step S21 , and output the learned metric feature representation , input the metric feature representation into the classifier , and through iterative optimization of the objective function: ; Update the classifier weights and predict the pseudo-labels corresponding to the signals ; where represents the current iteration round represents the classifier loss function represents the weight matrix of the deep long short-term memory network at the current iteration round; S23, Backpropagation Supervised Metric Feature Representation Learning: The pseudo-labels output in step S22 along with the electromagnetic data frame , are input into the deep long short-term memory network , and the objective function is iteratively optimized: ; Update the weight matrix of the deep long short-term memory network structure ; among them, is the clustering result label of the previous forward propagation, represents the backpropagation loss function; the clustering objective function is set as: ; Among them, belong to the same class after the previous round of clustering, belong to the nearest neighbor class of is a learning hyperparameter, represents the heap interval, is the learned metric distance, neighbor(i) is other electromagnetic data in the class to which it belongs; S24, Iterative optimization: Repeat step S22 and step S23 until the clustering objective function converges or the set number of iteration rounds is reached, and obtain the weight matrix of the final deep long short-term memory network structure and the classifier weights , and finally output the corresponding labels of the electromagnetic data and the label confidence , denoted as ; where the value range of .

5. The electromagnetic data annotation method based on human-machine hybrid intelligence according to claim 4, characterized in that, step S3 includes the following steps: S31, Similarity distance calculation: Calculate the distribution distance between different electromagnetic data types as the similarity distance between types; Denote the electromagnetic data distribution corresponding to the electromagnetic tag as , and the electromagnetic data distribution corresponding to the electromagnetic tag s as . Calculate the similarity distance between the two types using information cross-entropy : ; Among them, is the distribution of the sampled values, and is the number of sampled values; S32, Similarity calculation: Calculate the similarity between two types according to the similarity conversion formula for the similarity distance between two types: ; Calculate the similarity between two types; Among them, is the similarity distance calculated in step S31; S33, Similarity matrix calculation: Repeat steps S31 and S32 to calculate the similarity between all pairs of types, and output the similarity matrix.

6. The electromagnetic data annotation method based on human-machine hybrid intelligence according to claim 5, characterized in that, step S4 includes the following steps: S41, Expert Calibration: Visualize the electromagnetic data frames output in step S1 with different marks according to the labels output in step S2 to domain experts, and display the machine-annotated labels given in step S2 within the current electromagnetic data frame in the information bar in descending order of label confidence; Domain experts review each type of signal pattern given by the machine annotation based on professional knowledge and experience, and use manual annotation tools to select and correct the labels for the samples with annotation errors. S42, Label Assignment: Domain experts further modify the digital labels into corresponding character labels according to the training purpose. S43, Similar Type Merging: Visualize the similarity matrix output in step S3 with an electromagnetic similarity map, and domain experts merge the electromagnetic data with high similarity and belonging to the same type.

7. A method for electromagnetic data annotation based on human-machine hybrid intelligence according to claim 6, characterized in that, In step S42, domain experts modify the corresponding type of character labels according to the training purpose of the data set, modify them into corresponding device type character labels according to the training purpose of model identification, modify them into corresponding status character labels according to the training purpose of status identification, and modify them into corresponding modulation character labels according to the training purpose of modulation type identification.

8. A method for electromagnetic data annotation based on human-machine hybrid intelligence according to any one of claims 1 to 7, characterized in that, It further includes the following steps: S5, Repeat steps S2 to S4 until all electromagnetic data frames output in step S1 are processed; S6, Data Set Scoring: Domain experts score the quality of the electromagnetic data and the annotation quality completed in step S5, and output a training data set in a standard format.

9. An electromagnetic data annotation system based on human-machine hybrid intelligence, characterized in that, It is used to implement a method for electromagnetic data annotation based on human-machine hybrid intelligence according to any one of claims 1 to 8, and includes the following modules connected in sequence: Electromagnetic Data Preprocessing Module: First, slice the electromagnetic data to form electromagnetic data frames, then remove the outliers in each electromagnetic data frame one by one, and then perform standardization processing on the electromagnetic data frames, and output the preprocessed electromagnetic data frames; Deinterleaving and Separation Annotation Module: Input the electromagnetic data frames preprocessed by the electromagnetic data preprocessing module into a deep long short-term memory network to learn metric feature representations, and perform unsupervised clustering on the learned metric feature representations; then use the clustering results as labels and send them back to the deep long short-term memory network for supervised learning; Iterate in this way until the deep long short-term memory network converges, and output the separated annotation labels and label confidence; Similarity Matrix Calculation Module: Calculate the similarity between different types of electromagnetic data in the current electromagnetic data frame and historical electromagnetic data frames according to the annotation labels generated by the deinterleaving and separation annotation module, and generate a similarity matrix; where the historical electromagnetic data frames refer to all electromagnetic data frames before the current processing moment; Domain expert review module: Based on the label confidence generated by the deinterleaving and separation annotation module and the similarity matrix generated by the similarity matrix calculation module, manually review and correct the labels of the electromagnetic data after separation and annotation.

Citation Information

Patent Citations

  • Data labeling method based on human-machine cooperative learning

    CN108898225A

  • Electromagnetic signal identification method based on depth long-short-term memory network

    CN111553186A