Small target semantic recognition method for complex industrial scene

By designing semantic encoder and decoder, combined with source channel joint encoding, extracting the context semantic features of small objects in industrial product images, the shortcomings of traditional methods in small object recognition and complex channel processing are solved, and more efficient and reliable small object recognition is achieved.

CN120107968APending Publication Date: 2025-06-06BINZHOU WEIQIAO NATIONAL SCIENCE & TECHNOLOGY ADVANCED TECHNOLOGY RESEARCH INSTITUTE +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510036500.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Traditional industrial anomaly detection methods have shortcomings in small target recognition and complex channel characteristic processing, resulting in limited efficiency and accuracy of identifying small targets in complex industrial scenarios.

Method used

Design a semantic encoder and decoder for small-object recognition tasks, combine the joint encoding of the source and channel, extract the context semantic features of the tiny targets in the industrial product images, and optimize the encoding according to the channel characteristics, thereby improving the transmission reliability of image semantic features and the recognition accuracy of small targets.

Benefits of technology

By extracting context semantic features and source channel joint encoding, the recognition ability of small targets and the reliability of image semantic features transmission in complex industrial scenarios are significantly improved, the misjudgment rate is reduced and the recognition efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107968A_ABST
    Figure CN120107968A_ABST
Patent Text Reader

Abstract

The invention relates to a small target semantic recognition method for a complex industrial scene, and belongs to the field of industrial vision. According to the method, the identification precision for the small target of the industrial product in a complex industrial scene is optimized, and the problem that the detection precision is reduced due to semantic loss caused by the fact that the industrial product image is influenced by a complex channel in the transmission process is solved. According to the method, a small target recognition architecture with end and side cooperation is adopted, a semantic encoder and a source channel joint encoder are deployed at a factory image acquisition end and are used for acquiring and processing industrial product images, and after small target features are extracted, the small target features are encoded into bit streams to be transmitted through a wireless channel; and the source-channel joint decoder, the semantic decoder and the classifier are deployed on the edge server and are used for decoding the bit stream and performing small target identification so as to judge whether the industrial product is abnormal or not. According to the method, the reliability and the real-time performance of a small target identification task are ensured while the resource requirements on terminal computing power, communication and the like are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of industrial vision and relates to a small target semantic recognition method for complex industrial scenes. Background Art

[0002] Industrial anomaly detection plays a key role in promoting industrial intelligence. Its main goal is to identify and locate various defects in the product production process to ensure product quality, improve production efficiency, and maintain production stability. The identification of small targets is particularly important because the defects of many industrial products often appear as small abnormal points or areas, which are easily ignored by traditional methods. Therefore, the ability to accurately identify small targets is the key to improving anomaly detection capabilities.

[0003] In actual industrial production environments, since normal samples account for the vast majority and abnormal samples usually have diverse and unpredictable manifestations, the application of traditional supervised learning methods in anomaly detection is greatly limited. Because abnormal samples are extremely scarce during training, current research mostly uses methods based on self-supervised learning, which generates supervisory signals by utilizing the structure and correlation of the data itself, rather than relying on manual labeling.

[0004] As the requirements for information transmission and processing in industrial scenarios continue to increase, semantic communication has been gradually introduced as an emerging communication paradigm. The core idea of ​​semantic communication is to significantly reduce the amount of data in the transmission process of industrial product images by extracting and transmitting key information related to specific tasks. This method allows effective data compression and bandwidth consumption without sacrificing key information, which is particularly critical in complex industrial scenarios because these scenarios usually require fast and efficient information processing. In complex industrial environments, the use of semantic communication can greatly improve the recognition ability and speed for small targets. After streamlining and semantically optimizing the data, the subsequent anomaly detection model can focus on small targets faster during processing, improving recognition efficiency. This mechanism ensures that even in high-interference and high-complexity environments, the information of industrial product images can still be effectively transmitted, providing reliable data support for subsequent anomaly detection.

[0005] Industrial anomaly detection methods can be divided into two main trends: feature embedding-based and reconstruction-based. The reconstruction-based method assumes that a deep learning network trained only on normal data will fail when trying to reconstruct anomaly areas, and locates anomalies, especially small objects, by comparing the difference between the input image and the reconstructed image. However, this method is prone to overfitting, resulting in both normal and abnormal areas being well reconstructed, resulting in false detection. On the other hand, the feature embedding method extracts the features of the input image with the help of a model pre-trained on the ImageNet dataset, and calculates the feature distance between normal samples and input samples based on the representation of normal features to represent outliers. However, due to the distribution differences between industrial images and the ImageNet dataset, directly applying these networks may cause domain transfer problems and insufficient recognition of small objects.

[0006] In current research, most reconstruction methods use an encoder-decoder architecture to reconstruct the input image, intending to identify abnormal images and small targets through reconstruction errors. Although the powerful learning ability of neural networks can reconstruct abnormal areas including small targets in images, challenges still exist. DRAEM [Zavrtanik V, Kristan M, D.Draem-adiscriminatively trained reconstruction embedding for surface anomaly detection[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision.2021:8330-8339.】A method for generating abnormal samples is designed to obtain samples outside the normal distribution, and the reconstruction and discriminative subnetworks are combined for end-to-end training. OCR-GAN

Liang Y, Zhang J, Zhao S, et al. Omni-frequency channel-selection representations for unsupervised anomaly detection[J]. IEEE Transactions on Image Processing, 2023.

[0007] In the current field of industrial anomaly detection, although existing methods have made some progress in small target recognition, there are still significant shortcomings. Small targets usually occupy a small pixel area in the image, and their contrast with the background is low, which makes detection more difficult. Traditional methods often rely on static feature extraction when processing images, and fail to effectively mine contextual background information. Contextual information can provide additional clues to help the system confirm the existence of small targets. In addition, in complex industrial environments, channel conditions change dynamically, and existing technologies usually do not consider the time-varying characteristics of the channel. Channel interference may cause image information loss and affect the integrity of semantic features, especially when the data volume and transmission speed are high, affecting the detection rate and recognition speed of small targets. Due to the influence of multipath interference, fading and noise, the image transmission quality is unstable and prone to misjudgment. Therefore, the automatic detection system must have stronger robustness so that it can still correctly identify and locate small targets in the face of channel changes. This requires the development of new algorithms to enhance the ability to process context and maintain information integrity under complex propagation conditions.

[0008] In summary, the shortcomings of traditional methods in small target recognition and complex channel characteristic processing limit the effectiveness of their application. More intelligent and adaptive detection systems are urgently needed to meet the challenges of modern complex industrial scenarios. Summary of the invention

[0009] In view of this, the purpose of the present invention is to provide a small target semantic recognition method for complex industrial scenes. By designing a semantic encoder and decoder for small target recognition tasks, the contextual semantic features of small targets in industrial product images are extracted, and the importance of semantic features in small target images is perceived by combining source and channel joint coding. The coding is optimized according to the channel characteristics, thereby improving the reliability of image semantic feature transmission in complex industrial scenes and improving the recognition accuracy of small targets, thereby solving the shortcomings of traditional methods in small target recognition and channel adaptability.

[0010] In order to achieve the above object, the present invention provides the following technical solutions:

[0011] A small object semantic recognition method for complex industrial scenes, comprising:

[0012] S1. At the factory image acquisition end, the collected industrial product images are input into the semantic encoder for context-aware fusion to calculate the distribution area of ​​small targets, and the pixel points in the distribution area of ​​small targets are semantically encoded to form semantic feature codewords;

[0013] S2, inputting the semantic feature codeword into the source-channel joint encoder for source-channel joint encoding, removing the redundant part of the small target semantic feature, and adding redundancy to the semantic features related to the small target recognition task;

[0014] S3, the image acquisition end transmits the small target semantic feature bit stream after joint encoding through a wireless channel, wherein the small target semantic feature bit stream is transmitted through a noisy Gaussian white noise channel;

[0015] S4, at the edge server, inputting the received small target semantic feature bit stream into the source-channel joint decoder to restore the small target semantic feature bit stream jointly encoded by the sender;

[0016] S5, inputting the restored small target semantic feature bit stream into the semantic decoder to restore the small target semantic feature codeword at the image acquisition end;

[0017] S6. Input the small target semantic feature codeword into the classifier to perform the small target recognition task. The edge server constructs a cross entropy loss function to optimize the training of the semantic encoder and semantic decoder again based on the comparison between the recognition result of the classifier and the original input image of the image acquisition end;

[0018] S7. Input the small target semantic feature codeword output by the retrained semantic decoder into the classifier to perform the small target recognition task and determine whether the industrial product has any abnormality.

[0019] Furthermore, in step S1, a positioning method based on optical flow motion features is used to calculate the distribution area of ​​the small target. First, the region of interest is selected through context information, and the optical flow intensity of the small target motion features is calculated. Then, the optical flow curve of the small target is processed using smoothing filtering and monotonic interval control to locate the vertex frame.

[0020] Specifically include:

[0021] 1) Detect the initial target key points of the input image through the regional convolutional neural network, select several key points and connect the selected key points to construct multiple regions of interest;

[0022] 2) Use a filter to smooth each region of interest, extract the coefficients of the polynomial, and minimize the mean square error; where the sliding window size of the filter is 2n+1, and 2n+1 is greater than the order k of the polynomial, then the optical flow intensity m at time t t The fitting equation is: In the formula, a i Indicates the i-th coefficient to be solved; Similarly, equations are established for key points in other regions, and a total of 2n+1 equations can be obtained:

[0023] M=Xα

[0024] M=(m t-n ,...,m t ,...,m t+n )T

[0025]

[0026] α=(a 0 ,a 1 ,a 2 ,…,a k ) T

[0027] In the formula, M represents the optical flow intensity curve of the small target, X represents the coefficient matrix, and α represents the coefficient vector. The optimal fitting polynomial is solved using the least squares method, that is, the optimal polynomial coefficient vector α is found.

[0028] 3) Identify all local maxima of the optical flow intensity curve and establish a candidate set of vertex frames represents the amplitude of the jth candidate peak frame after the curve is smoothed; the monotonic interval of the jth candidate peak frame is extended forward and backward to obtain a symmetric interval centered on j, that is, the maximum monotonic interval of candidate point j; the maximum monotonic interval w of the candidate point must satisfy: w ≥ max[5, λL], L represents the length of each small target sequence, and λ represents a hyperparameter; the width of the monotonic interval, that is, the dimension of the small target context information, is controlled by adjusting the size of λ;

[0029] 4) The candidate point with the largest amplitude is selected as the vertex, and then the distribution area of ​​the small target is determined according to the width of the monotonic interval.

[0030] Furthermore, in step S5, the generator and discriminator of the semantic decoder are first trained with adversarial loss minimization as the optimization goal, and then used to restore the small target semantic feature codewords at the image acquisition end. The training process is:

[0031] 1) The semantic decoder randomly selects two small target images from the training set as input image pairs, which are recorded as the first image and the second image respectively, and extracts the target vector and the irrelevant vector from both the first image and the second image;

[0032] 2) Keep the irrelevant vector unchanged, exchange the target vectors of the two images, concatenate the target vector of the second image with the irrelevant vector of the first image, concatenate the target vector of the first image with the irrelevant vector of the second image, and provide them as input to the generator to generate two new images;

[0033] 3) The discriminator determines the authenticity of the original input image and the generated image and classifies them.

[0034] Among them, the adversarial loss is expressed as:

[0035]

[0036]

[0037] In the formula, represents the loss function of the discriminator, represents the loss function of the generator; G(·) represents the generator, D(·) represents the discriminator; E[·] represents the expected value; l represents the lth feature map of the input sample; Represented by the input sample and The data obtained by linear sampling, is a random number between (0, 1); gp represents the gradient penalty weight.

[0038] The beneficial effects of the present invention are:

[0039] (1) The present invention can extract the contextual semantic features of tiny targets in industrial product images by designing semantic encoders and decoders for small target recognition tasks. The integration of such contextual information significantly improves the recognition capability of small targets, enabling the detection system to accurately identify small targets even in complex backgrounds, effectively reducing the misjudgment rate.

[0040] (2) Aiming at the time-varying characteristics of channel conditions in complex industrial environments, the present invention proposes a joint coding strategy for the source channel, making the image semantic features more reliable during transmission. When channel interference occurs, by sensing the importance of the semantic features of the small target image, the system can adaptively adjust the coding method according to the channel state, effectively reducing information loss and improving the transmission reliability of the image semantic features.

[0041] (3) Aiming at the high real-time and high reliability requirements in complex industrial production scenarios, the present invention proposes an edge-to-edge collaborative small target recognition architecture based on semantic communication and edge computing architecture. The semantic encoder and source-channel joint encoder are deployed at the factory image acquisition end, and the source-channel joint decoder, semantic decoder, and classifier are deployed at the edge server. Through the design of the semantic encoder and decoder, the amount of data of industrial product images to be transmitted is reduced, the transmission delay of industrial product images is reduced, and the real-time performance of small target recognition is improved; through the design of the source-channel joint encoder and decoder, the anti-noise ability of semantic features in channel transmission in complex industrial scenarios is improved.

[0042] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:

[0044] Figure 1 is a structural block diagram of the method of the present invention;

[0045] Figure 2 Schematic diagram of the joint source channel encoder and decoder structure. DETAILED DESCRIPTION

[0046] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0047] An embodiment of the present invention provides a small target semantic recognition method for complex industrial scenarios. The method adopts an end-edge collaborative small target recognition architecture. The semantic encoder and the source-channel joint encoder are deployed at the factory image acquisition end, and the source-channel joint decoder, the semantic decoder, and the classifier are deployed at the edge server. While reducing the requirements for terminal computing power, communication and other resources, the efficiency of the small target recognition task is guaranteed.

[0048] like Figure 1 As shown, the method mainly includes the following steps:

[0049] 1. At the image acquisition end, the industrial product image is input into the semantic encoder for semantic feature extraction, and the important semantic feature codewords related to the small target are extracted in combination with the context information.

[0050] Specifically:

[0051] 1. The image acquisition end inputs the image into the semantic encoder for context-aware fusion to calculate the distribution area of ​​the small target. This embodiment adopts a positioning method based on optical flow motion features, which mainly includes the description of the optical flow motion process and the positioning of the vertex frame. First, the region of interest is selected through context information, and the optical flow intensity of the small target motion feature is calculated; then, the optical flow curve of the small target is processed using smoothing filtering and monotonic interval control to locate the vertex frame. As described below:

[0052] 1) At the image acquisition end, the regional convolutional neural network (R-FCN) is first used to detect the preliminary target key points and select several key points. Then, these key points are connected to construct four regions of interest (ROIs).

[0053] 2) The image acquisition end uses the Savitzky-Golay filter to perform smoothing filtering on the four ROIs and extract the coefficients P of the polynomial K (t), minimize the mean square error, as shown below:

[0054]

[0055] Among them, t is the current time index; a i is the i-th coefficient to be solved; k is the order of the polynomial.

[0056] Assuming that the sliding window size of the filter is 2n+1, and 2n+1 should be greater than the order k of the polynomial, then the optical flow intensity m at time t is t The fitting equation can be expressed as:

[0057]

[0058] Similarly, equations are established for each other point in the region. For another 2n points, 2n+1 equations can be constructed as follows:

[0059] M=Xα(3)

[0060] M=(m t-n ,...,m t ,...,m t+n ) T (4)

[0061]

[0062] α=(a 0 ,a 1 ,a 2 ,…,a k ) T (6)

[0063] Since the number of equations above exceeds the degree of the polynomial, the least squares method is used to solve the optimal fitting polynomial, that is, to find the optimal polynomial coefficient vector α. According to the least squares fitting theory, the optimal α is expressed as follows:

[0064] α=(X T X) -1 X T M (7)

[0065] Through filtering, the optical flow intensity change curve M is smoothed. Although the intuitive method is to directly select the highest point of the curve as the vertex frame, this method is easily affected by local accidental features and random background noise. To this end, first identify all local maxima to establish a candidate set A of vertex frames, as shown in the following formula:

[0066]

[0067] in, is the amplitude of the jth candidate peak frame after the curve is smoothed. The monotonic interval of j is extended forward and backward to obtain a symmetrical interval centered on j, which is the maximum monotonic interval of candidate point j. Then, in order to eliminate local accidental features with large transient values ​​and global background noise with random patterns, the range of the maximum monotonic interval w of the candidate point is limited, and w should satisfy:

[0068] w≥max[5,λL] (9)

[0069] Where L is the length of each small target sequence; λ is a trainable hyperparameter. By adjusting the size of λ, the width of the monotonic interval, that is, the dimension of the small target context information, can be controlled. Therefore, the monotonic interval constraint can adapt to the different sizes of small targets in the image at the image acquisition end. Finally, the candidate point with the largest amplitude is selected as the vertex, and then the distribution area of ​​the small target is determined according to the width of the monotonic interval.

[0070] 2. After the image acquisition end extracts the small target distribution area, it semantically encodes the pixels in the target area to form semantic feature codewords.

[0071] Second, the image acquisition end inputs the extracted small target semantic feature codewords into the source-channel joint encoder for joint source-channel encoding, removes the redundant parts of the small target semantic features, and adds redundancy to the semantic features that are highly relevant to the small target recognition task, forming a joint coded bit stream suitable for channel transmission, and improving the anti-noise ability of the small target semantic feature codewords in the harsh channel environment of complex industrial scenes. The structure of the source-channel joint encoder is as follows: Figure 2 shown.

[0072] The encoder at the image acquisition end is a joint design of the source and channel encoders, with CNN as the core, including a series of convolutional layers, ReLU functions and normalization layers. The convolutional layer extracts the core features of the small target semantic features of the input joint encoder, and the ReLU function learns the nonlinear mapping from the source signal space to the encoding signal space.

[0073] Specifically, the image acquisition end inputs the extracted semantic features into the source channel joint encoder and inputs the small target semantic features Y T The codeword e after being mapped to the joint encoder E can be expressed as:

[0074]

[0075] Where s is the dimension of e, is a set of semantic features, θ 0 are the parameters of the joint encoder E.

[0076] 3. The image acquisition end transmits the jointly encoded small target semantic feature bit stream through the wireless channel. The encoded codeword e is transmitted through a noisy Gaussian white noise channel and can be expressed as:

[0077]

[0078] Where N is the channel noise, the noise is given by Sampling obtained; σ 2 is the noise power, is a complex Gaussian distribution.

[0079] 4. The edge server receives the encoded small target semantic feature bit stream transmitted through the wireless channel, inputs it into the source-channel joint decoder, and restores the small target joint encoded bit stream at the sender.

[0080] The source-channel joint decoder performs the opposite operation to the encoder. Specifically, the edge server maps e′ to the core semantic feature Y through the joint decoder D T ', can be expressed as:

[0081] Y T '=D(e′,τ)=D(E(Y T ,θ 0 )+N,τ) (12)

[0082] Among them, the core semantic feature Y T ' is the semantic feature Y of the small target at the image acquisition end T is the estimated value of , and τ is the parameter of the joint decoder D.

[0083] 5. The edge server inputs the joint encoded bit stream into the semantic decoder to restore the semantic feature codewords of the small target at the image acquisition end. The generator and discriminator of the semantic decoder are trained adversarially, and the adversarial loss function is designed to optimize the semantic decoder.

[0084] Among them, the semantic decoder randomly selects two small target images from the training set as input image pairs (respectively denoted as image I and image II) in each training iteration. First, the edge server extracts the features of the two input images and extracts the target vector and irrelevant vector from each image. Then, the target vectors of the two images are swapped while keeping the irrelevant vector unchanged. The swapped target vectors are concatenated with the corresponding irrelevant vectors (i.e., the target vector of image II is concatenated with the irrelevant vector of image I, and the target vector of image I is concatenated with the irrelevant vector of image II), and are provided as input to the generator to generate two new images. The main task of the discriminator is to judge the authenticity of the input image and the generated image, and to classify them at the same time.

[0085] The edge server trains the semantic decoder through the game between the generator and the discriminator, with the adversarial loss minimization as the optimization goal. After convergence, the generated samples will be very similar to the real samples. The adversarial loss is as follows:

[0086]

[0087]

[0088] Where G(·) is the generator, D(·) is the discriminator; l represents the lth feature map of the input sample; The input sample and The data obtained by linear sampling, is a random number between (0, 1); gp is the gradient penalty weight. After training, the edge server reconstructs the received small target semantic feature codewords.

[0089] 6. The edge server inputs the semantic feature codewords of the small target into the classifier to perform the recognition task of the small target. At the same time, the edge server constructs a cross entropy loss function based on the comparison between the recognition result of the classifier and the original image to optimize the training of the semantic encoder and decoder again. After the training is completed, the semantic feature codewords of the small target restored by the semantic decoder are input into the classifier to perform the recognition task of the small target and determine whether the industrial product has any abnormality.

[0090] Among them, the classifier uses a support vector machine to recognize and classify the input small target semantic feature codewords.

[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.

Claims

1. A small object semantic recognition method for complex industrial scenes, characterized by: The method includes: At the factory image acquisition end, the collected industrial product images are input into the semantic encoder for context-aware fusion to calculate the distribution area of ​​small targets, and the pixel points in the small target distribution area are semantically encoded to form semantic feature codewords; The semantic feature codeword is input into the source-channel joint encoder for source-channel joint encoding, the redundant part of the small target semantic feature is removed, and the redundancy of the semantic features related to the small target recognition task is increased; The image acquisition end transmits the jointly encoded small target semantic feature bit stream through a wireless channel; At the edge server, the received small object semantic feature bit stream is input into the source-channel joint decoder to restore the small object semantic feature bit stream jointly encoded by the sender; The restored small object semantic feature bit stream is input into the semantic decoder to restore the small object semantic feature codeword at the image acquisition end; The semantic feature codewords of small objects are input into the classifier to perform the small object recognition task. The edge server constructs a cross entropy loss function based on the comparison between the recognition result of the classifier and the original input image of the image acquisition end to optimize the training of the semantic encoder and semantic decoder again. The small target semantic feature codewords output by the retrained semantic decoder are input into the classifier to perform the small target recognition task and determine whether the industrial product has any abnormality.

2. The method according to claim 1, characterized in that The distribution area of ​​the small target is calculated by using a positioning method based on optical flow motion features, including: The input image is preliminarily detected for target key points through a regional convolutional neural network, several key points are selected and connected to construct multiple regions of interest. Use a filter to smooth each region of interest, extract the coefficients of the polynomial, and minimize the mean square error; where the sliding window size of the filter is 2n+1, and 2n+1 is greater than the order k of the polynomial, then the optical flow intensity m at time t t The fitting equation is: In the formula, a i Indicates the i-th coefficient to be solved; Similarly, equations are established for key points in other regions, and a total of 2n+1 equations are obtained: M=Xα M=(m t-n ,...,m t ,...,m t+n ) T α=(a0,a1,a2,…,a k ) T In the formula, M represents the optical flow intensity curve of the small target, X represents the coefficient matrix, and α represents the coefficient vector. The optimal fitting polynomial is solved using the least squares method, that is, the optimal polynomial coefficient vector α is found. Identify all local maxima of the optical flow intensity curve and establish a vertex frame candidate set A: represents the amplitude of the jth candidate peak frame after the curve is smoothed; the monotonic interval of the jth candidate peak frame is extended forward and backward to obtain a symmetric interval centered on j, that is, the maximum monotonic interval of candidate point j; the maximum monotonic interval w of the candidate point must satisfy: w ≥ max [5, λL], L represents the length of each small target sequence, and λ represents a hyperparameter; the width of the monotonic interval, that is, the dimension of the small target context information, is controlled by adjusting the size of λ; The candidate point with the largest amplitude is selected as the vertex, and then the distribution area of ​​the small target is determined according to the width of the monotonic interval.

3. The method according to claim 1, characterized in that The generator and discriminator of the semantic decoder are firstly subjected to adversarial training and then used to restore the semantic feature codewords of the small target at the image acquisition end. The training method includes: The semantic decoder randomly selects two small target images from the training set as input image pairs, which are denoted as the first image and the second image respectively, and extracts the target vector and the irrelevant vector from both the first image and the second image; Keep the irrelevant vector unchanged, exchange the target vectors of the two images, concatenate the target vector of the second image with the irrelevant vector of the first image, concatenate the target vector of the first image with the irrelevant vector of the second image, and provide them as input to the generator to generate two new images; The discriminator determines the authenticity of the original input image and the generated image and classifies them.

4. The method according to claim 3, characterized in that The semantic decoder is trained with the minimization of adversarial loss as the optimization objective, and the adversarial loss is expressed as: In the formula, represents the loss function of the discriminator, represents the loss function of the generator; G(·) represents the generator, D(·) represents the discriminator; E[·] represents the expected value; l represents the lth feature map of the input sample; Represented by the input sample and The data obtained by linear sampling, is a random number between (0, 1); gp represents the gradient penalty weight.

5. The method according to claim 1, characterized in that The small target semantic feature bit stream after joint encoding is transmitted through a noisy Gaussian white noise channel.