A signal modulation recognition and positioning method based on time-frequency graph and visual language model

By converting IQ signals into time-frequency graphs and combining them with fine-tuning of YOLO and visual language models, high-precision identification and localization of signal targets are achieved. This solves the problems of insufficient signal recognition accuracy and reliance on manual annotation in existing technologies, and improves the model's generalization ability in complex electromagnetic environments.

CN120763877BActive Publication Date: 2025-11-07ARTIFICIAL INTELLIGENCE INNOVATION RES INST OF ZHEJIANG UNIV OF TECH BINJIANG DISTRICT HANGZHOU
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511277744.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2025-11-07
Estimated Expiration
2045-09-09

AI Technical Summary

Technical Problem

Existing technologies lack sufficient accuracy in signal recognition in complex electromagnetic environments, rely on manual annotation, and have poor generalization ability in the signal domain, lacking a unified spatial localization and semantic understanding framework.

Method used

By converting the original IQ signal into a time-frequency graph, a fine-tuned YOLO model is used for target detection, generating bounding boxes and category labels. Descriptive text is generated by combining a language model, multimodal training pairs are constructed, and the visual language model is fine-tuned to optimize the image-text matching loss and coordinate regression loss, thereby achieving automatic recognition and localization of signal targets.

Benefits of technology

It achieves high-precision identification and spatial positioning of signal targets, improves the semantic understanding of signal images, solves the problems of limited accuracy and reliance on manual annotation in traditional methods, and enhances the generalization ability of the model in complex electromagnetic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120763877B_ABST
    Figure CN120763877B_ABST
Patent Text Reader

Abstract

The application discloses a signal modulation recognition and positioning method based on a time-frequency graph and a visual language model, relates to the technical field of signal processing and artificial intelligence, and aims at the problems of low recognition accuracy, dependence on manual annotation and poor generalization of a general model in a signal scene in the prior art, and generates a time-frequency graph through a short-time Fourier transform; a fine-tuned YOLO model is used to automatically generate the boundary box coordinates and the category label of a signal target; a description text is generated in combination with a language model to construct a multi-modal training pair; a visual language model is fine-tuned by adopting a strategy of combining a graph-text matching loss and a coordinate regression loss, and finally, intelligent recognition and accurate positioning of a signal target in a time-frequency graph are realized. The application is suitable for signal detection and spectrum situation awareness in a complex electromagnetic environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of signal processing and artificial intelligence, and particularly relates to a signal modulation recognition and positioning method based on a time-frequency graph and a visual language model. BACKGROUND

[0002] With the rapid development of wireless communication, radar detection and electronic countermeasures, signals in complex electromagnetic environments are increasingly characterized by high density, wideband and diversification. Rapid and accurate identification of target signal regions in a large amount of signals has become a key issue in the field of signal intelligent processing. As an intersection carrier of signal processing and vision, the time-frequency graph provides the possibility for introducing computer vision and multi-modal semantic modeling methods.

[0003] Currently, target region identification in signal images mainly relies on traditional image processing algorithms or convolutional neural networks for classification and positioning. However, these methods generally have limited recognition accuracy, weak generalization ability, and are not sensitive to weak signals in complex backgrounds. In particular, in the absence of large-scale annotated data, training models with high precision and strong migration ability faces serious challenges.

[0004] In recent years, target detection models such as YOLO have been introduced into signal image analysis due to their end-to-end fast detection capabilities, achieving joint inference of signal existence, time-frequency positioning and modulation category, and significantly improving the efficiency of structured labeling of signals in time-frequency graphs. However, such single-modal visual models still lack the ability to understand signal semantics in complex electromagnetic environments, making it difficult to assist in positioning or explain the recognition results through natural language and other means. On the other hand, multi-modal visual language models have shown strong representation and reasoning capabilities in joint modeling of images and text, providing a new approach to image region reasoning in complex backgrounds.

[0005] However, such models are originally designed for natural images, and when directly migrated to the field of signal images, they often suffer from performance degradation due to lack of structural prior and weak coordinate expression. Existing signal image datasets lack sufficient scene complexity and real channel disturbance, resulting in insufficient generalization ability of multi-modal models, making it difficult to support joint modeling of fine-grained target semantics and location. Therefore, the current technology still faces a series of core challenges such as limited target recognition accuracy, reliance on large amounts of manual annotation, difficulty in migrating general models to the signal field, and lack of a unified spatial positioning and semantic understanding framework. SUMMARY

[0006] To solve the above technical problems, the present application proposes a signal modulation recognition and positioning method based on a time-frequency graph and a visual language model to solve the problems existing in the prior art.

[0007] The first aspect, to achieve the above object, the present application provides a signal modulation recognition and positioning method based on time-frequency graph and visual language model, comprising the following steps:

[0008] S1, the original IQ signal is transformed into a time-frequency graph by time-frequency transformation;

[0009] S2, the time-frequency graph is detected by using a target detection model fine-tuned by a signal data set, and the boundary box coordinates and class labels of the signal target in the time-frequency graph are automatically generated;

[0010] S3, based on the boundary box coordinates and class labels, a corresponding descriptive text is generated by using a language model to construct a multi-modal training pair containing the time-frequency graph, the descriptive text and the boundary box coordinates;

[0011] S4, the multi-modal training pair is used to fine-tune a visual language model, and the fine-tuning process jointly optimizes the graph-text matching loss and the coordinate regression loss;

[0012] S5, the class and coordinate information of the target signal are output by using the fine-tuned visual language model to recognize a new time-frequency graph.

[0013] Optionally, in S1, the process of generating a time-frequency graph by time-frequency transformation of the original IQ signal comprises: performing short-time Fourier transform on the original IQ signal, using a Hanning window function, and setting window length, window overlap rate and Fourier transform point number to generate a single-channel time-frequency graph image.

[0014] Optionally, in S2, the process of detecting by using a target detection model comprises: using a pre-trained YOLO model, fine-tuning using a Wideband-Sig53 data set; scaling the time-frequency graph and inputting the fine-tuned YOLO model to output normalized boundary box coordinates and preset class labels.

[0015] Optionally, in S3, the process of generating a descriptive text comprises: combining the class labels and the time and frequency range information mapped from the boundary box coordinates to generate a natural language description containing signal type and spatial position semantics.

[0016] Optionally, in S4, the process of jointly optimizing the graph-text matching loss and the coordinate regression loss comprises: calculating the information noise contrast estimation loss between image features and text features, and calculating the L1 loss between predicted coordinates and real coordinates, and taking the weighted sum of the two as a joint loss function.

[0017] Optionally, in S5, the process of identifying the new time-frequency map includes: inputting the pre-processed time-frequency map into the fine-tuned visual language model, and receiving a natural language query; through the collaborative work of the visual encoder, the text encoder, the cross-modal attention module and the coordinate regression branch in the model, the structured information containing the signal category, the time range, the frequency range and the bounding box coordinates is output.

[0018] In a second aspect, the present application further provides a signal modulation identification and positioning system based on time-frequency map and visual language model, for implementing a signal modulation identification and positioning method based on time-frequency map and visual language model, the system comprising:

[0019] A signal preprocessing and time-frequency map generation module is configured to generate a time-frequency map from an original IQ signal through time-frequency transformation;

[0020] An automatic annotation module is configured to automatically generate the bounding box coordinates and category labels of the signal targets in the time-frequency map by using a target detection model fine-tuned by a signal dataset to detect the time-frequency map;

[0021] A multi-modal training data construction module is configured to generate corresponding descriptive text by using a language model based on the bounding box coordinates and category labels, so as to construct a multi-modal training pair containing the time-frequency map, the descriptive text and the bounding box coordinates;

[0022] A visual language model fine-tuning module is configured to fine-tune a visual language model using the multi-modal training pair, and the fine-tuning process jointly optimizes the text-image matching loss and the coordinate regression loss;

[0023] A signal identification and output module is configured to identify a new time-frequency map by using the fine-tuned visual language model, and output the category and coordinate information of the target signal.

[0024] In a third aspect, the present application further provides a computer terminal device, comprising:

[0025] One or more processors;

[0026] A memory coupled to the processor, configured to store one or more programs;

[0027] When the one or more programs are executed by the one or more processors, the one or more processors implement the steps of the signal modulation identification and positioning method based on time-frequency map and visual language model in the first aspect.

[0028] In a fourth aspect, the present application further provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps of the signal modulation identification and positioning method based on time-frequency map and visual language model in the first aspect.

[0029] In a fifth aspect, the present application also provides a computer program product comprising a computer program which, when executed by a processor, implements the steps of the signal modulation recognition and positioning method based on the time-frequency graph and the visual language model in the first aspect.

[0030] Compared with the prior art, the present application has the following advantages and technical effects:

[0031] The signal modulation recognition and positioning method based on the time-frequency graph and the visual language model provided by the present application overcomes the defects of limited precision and severe dependence on manual annotation of traditional signal image recognition methods, and solves the problems of poor generalization and insufficient spatial positioning capability of general visual language models in the signal field. By fusing target detection and multi-modal model fine-tuning, automatic recognition and high-precision coordinate positioning of signal target modulation types in the time-frequency graph are realized, which has excellent semantic understanding capability and spatial regression capability, and provides an effective solution for intelligent signal processing in a complex electromagnetic environment. BRIEF DESCRIPTION OF DRAWINGS

[0032] The accompanying drawings, which form a part of the present application, are intended to provide further understanding of the present application, and the illustrative embodiments of the present application and their description serve the purpose of explaining the present application. The accompanying drawings should not be construed as an inappropriate limitation on the present application. In the drawings:

[0033] Figure 1 The method flowchart of the embodiments of the present application;

[0034] Figure 2 The model training framework diagram of the embodiments of the present application. DETAILED DESCRIPTION

[0035] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0036] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a group of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.

[0037] Embodiment one

[0038] In this embodiment, a signal modulation recognition and positioning method based on a time-frequency graph and a visual language model is provided, which comprises:

[0039] S1, generating a time-frequency graph by time-frequency transformation of an original IQ signal;

[0040] S2, detecting the time-frequency graph using the target detection model fine-tuned by the signal data set to automatically generate the bounding box coordinates and class labels of the signal target in the time-frequency graph;

[0041] S3, generating corresponding descriptive text based on the bounding box coordinates and class labels using a language model to construct a multi-modal training pair containing the time-frequency graph, the descriptive text, and the bounding box coordinates;

[0042] S4, fine-tuning the visual language model using the multi-modal training pair, and the fine-tuning process jointly optimizes the image-text matching loss and the coordinate regression loss;

[0043] S5, using the fine-tuned visual language model to recognize new time-frequency graphs and output the class and coordinate information of the target signal.

[0044] Specifically, S1: input the original IQ signal data and generate a time-frequency graph as an image input through short-time Fourier transform (STFT);

[0045] S2: fine-tune the YOLO model using the Wideband-Sig53 data set, and use the fine-tuned YOLO model to detect target signals in the time-frequency graph to automatically generate target bounding boxes and class labels, realizing the structured annotation of image signals;

[0046] S3: use a large language model to construct image and descriptive text pairing samples, and combine YOLO annotation output with signal semantic description to generate a multi-modal training pair; the large language model includes ChatGPT, etc.

[0047] S4: based on existing visual language models (such as Qwen-VL, MiniGPT, etc.), fine-tune them using the above data, introduce image-text matching loss and coordinate regression loss, and realize joint modeling of signal target semantics and spatial position by the model;

[0048] S5: in the inference stage, input a new signal time-frequency graph, and complete the recognition and coordinate output of the target signal region through the fine-tuned visual language model.

[0049] As an implementation in this embodiment, in S1, the process of generating a time-frequency graph from the original IQ signal through time-frequency transformation includes: performing short-time Fourier transform on the original IQ signal, using a Hanning window function, and setting window length, window overlap rate, and Fourier transform point number to generate a single-channel time-frequency graph image.

[0050] Further, the step S1 specifically includes the following contents:

[0051] S1.1: input the original IQ signal (set the length as N sampling points), convert it into a two-dimensional time-frequency graph representation by a time-frequency analysis method (such as short-time Fourier transform STFT), and obtain a time-frequency graph image H and W are the time axis and frequency axis resolution of the time-frequency graph respectively:

[0052] (1)

[0053] More specifically, in S1, the specific operation process is as follows: refer to the attached Figure 1 、 Figure 2 , input the original IQ signal in complex number form (containing in-phase component I and quadrature component Q, the number of sampling points is 32896), and generate a time-frequency graph by short-time Fourier transform (STFT). The specific parameter setting is as follows: a Hanning window function is used, the window length is set to 256 sampling points, the window overlap rate is 50%, and the number of Fourier transform points is 512. After STFT processing, the original 1D IQ signal is converted into a 2D time-frequency graph, the time axis dimension H is calculated from the window length and the overlap rate, the frequency axis dimension W is half of the number of Fourier transform points, and finally a single-channel time-frequency graph image with a size of 256x256x1 is output as the input of the subsequent model.

[0054] As an implementation manner in this embodiment, in S2, the detection process using the target detection model includes: using a pre-trained YOLO model, fine-tuning using a Wideband-Sig53 dataset; scaling the time-frequency graph and inputting it into the fine-tuned YOLO model to output normalized bounding box coordinates and preset class labels.

[0055] Further, the step S2 specifically includes the following contents:

[0056] S2.1: use the Wideband-Sig53 dataset to construct paired samples of the original time-frequency graph, corresponding bounding box coordinates and class labels, input them into the YOLO detection network, and supervise the training of the model to adapt to the detection task of the signal target in the time-frequency graph, with the goal of minimizing the bounding box regression loss and the classification loss.

[0057] S2.2: use the fine-tuned YOLO model to infer the time-frequency graph image output target bounding box coordinates and corresponding class labels , and construct paired samples of the image, the bounding box and the label:

[0058] (2)

[0059] wherein, represents a set of bounding boxes, is the corresponding label, The number of samples is n.

[0060] More specifically, in step 2, the specific operation process is as follows: refer to the attached Figure 1 、 Figure 2 , the pre-trained YOLOv8 model is used as the target detection basic framework, which is pre-trained on the COCO dataset and has strong general target positioning ability. Fine-tune the YOLO model using the Wideband-Sig53 dataset to adapt it to the signal recognition and positioning task. Resize the time-frequency graph image generated in step 1 to 640x640 pixels and input it into the YOLOv8 model for inference. During the inference process, set the confidence threshold to 0.5 (filter low-confidence detection results). The model outputs the bounding box coordinates of the target signal (format [x, y, w, h], where (x, y) is the center point coordinate, (w, h) is the width and height of the bounding box, both are normalized values relative to the size of the time-frequency graph) and the corresponding class label (such as "BPSK" "QPSK" "Radar Pulse" and other preset signal categories). Organize these outputs into structured annotation data and store them in JSON format, including image path, bounding box list, and class label list.

[0061] As an embodiment in this embodiment, in S3, the process of generating descriptive text includes: combining the class label and the time and frequency range information mapped from the bounding box coordinates to generate natural language descriptions containing signal type and spatial location semantics.

[0062] Further, the step S3 includes the following content:

[0063] S3.1: Pair the image obtained in step 2 with the bounding box and label sample Input into a large language model LLM to generate text pairs that describe semantic information such as signal type, spectral features, or possible application scenarios .

[0064] (3)

[0065] S3.2: Pair the image With the natural language description generated based on the label To form multi-modal training data:

[0066] (4)

[0067] ​As an embodiment in this embodiment, in S4, the process of jointly optimizing the image-text matching loss and the coordinate regression loss includes: calculating the information noise contrast estimation loss between the image features and the text features, and calculating the L1 loss between the predicted coordinates and the real coordinates, and taking the weighted sum of the two as the joint loss function.

[0068] More specifically, in step 3, the specific operation process is as follows: refer to the attached Figure 1 、 Figure 2 The structured annotation generated in step 2 is input into a large language model to build image-text paired samples. The text description generation rule is: combine the category label and the bounding box coordinates to generate natural language containing semantic information and spatial information, for example: "This time-frequency diagram contains 1 [category label] signal, its position in the time-frequency diagram is the upper left corner ([x1, y1]) and the lower right corner ([x2, y2]), corresponding to the time range [t_start, t_end] and the frequency range [f_start, f_end] "(where t_start, t_end is mapped to the actual sampling time by the time axis coordinates, and f_start, f_end is mapped to the actual frequency value by the frequency axis coordinates). Each time-frequency diagram image and its corresponding text description form a multi-modal training pair, stored as a triple of "image path-text string-bounding box coordinates", and a dataset for fine-tuning the visual language model is constructed.

[0069] Further, the step S4 includes the following contents:

[0070] S4.1: input image and the corresponding description text , respectively through the visual encoder and the text encoder to extract the features:

[0071] (5)

[0072] S4.2: Calculate the image-text matching loss:

[0073] (6)

[0074] where, is a similarity function, is a temperature coefficient.

[0075] S4.3: input image features into the coordinate regression branch to output the predicted bounding box parameters , and calculate the regression loss with the real box labeled by YOLO:

[0076] (7)

[0077] S4.4: Combining the image-text matching loss with the coordinate prediction loss into a joint loss function:

[0078] (8)

[0079] wherein, is the loss weight, used to balance the optimization objectives of image-text matching and spatial localization.

[0080] More specifically, in step 4, the specific operation process is as follows: refer to the attached Figure 1 、 Figure 2 Qwen-VL is selected as the basic visual language model, which includes a visual encoder (ViT-L / 14 structure), a text encoder (Transformer structure), and a cross-modal attention module. The fine-tuning process adopts the following strategies: input the time-frequency graph image to the visual encoder, and output the image feature vector; input the text description to the text encoder, and output the text feature vector. Calculate the image-text matching loss: use the infoNCE loss function to measure the matching degree of image features and text features through cosine similarity, set the temperature coefficient τ = 0.07, and make the similarity of matched image-text pairs higher than that of non-matched pairs. Introduce the coordinate regression branch: add two fully connected layers (hidden dimension 512, activation function ReLU) after the image features output by the visual encoder, output the predicted bounding box coordinates (normalized to the range [0, 1]), and calculate the L1 loss of the predicted coordinates and the YOLO labeled real coordinates. Training parameter settings: the initial learning rate is 5e-5, the batch size is 16, the training round is 100, the AdamW optimizer is used, and the learning rate is scheduled through the cosine annealing strategy to avoid overfitting and accelerate convergence.

[0081] As an embodiment in this embodiment, in S5, the process of identifying the new time-frequency graph includes: inputting the preprocessed time-frequency graph into the fine-tuned visual language model, and receiving a natural language query; through the cooperation of the visual encoder, text encoder, cross-modal attention module and coordinate regression branch in the model, outputting structured information including signal category, time range, frequency range and bounding box coordinates.

[0082] Further, the step S5 includes the following contents:

[0083] S5.1: Fine-tuned for inference.

[0084] The fine-tuned visual language model can directly accept signal image input and combine natural language queries (such as "what signal does the time-frequency graph contain, and when does it appear") to perform target area recognition and coordinate output, achieving automatic semantic localization of signal images.

[0085] More specifically, in step 5, the specific operation process is as follows: after the model training is completed, the visual encoder, the text encoder, the cross-modal attention module and the coordinate regression branch are retained in the deployment stage, and the auxiliary loss calculation module in the training process is discarded. When reasoning, a new time-frequency graph image (preprocessed: resized to 640x640, normalized to the range of [0, 1]) can be accompanied by a natural language query (such as "what type of signal does the time-frequency graph contain" or "find the coordinates of the radar pulse signal"). The model process is as follows: the visual encoder extracts image features, the text encoder processes query text features, the cross-modal attention module fuses the features of the two, and the coordinate regression branch outputs the predicted bounding box coordinates (normalized values), which are then mapped back to the actual time-frequency coordinates of the time-frequency graph. The output result is structured information containing signal category, time range, frequency range and bounding box coordinates, which can be directly used for downstream signal classification, target tracking and other tasks.

[0086] Based on this, the embodiment of the present application provides a signal modulation identification and positioning method based on a time-frequency graph and a visual language model. The traditional method relies on manual signal image annotation, which is low in efficiency and unstable in accuracy, and cannot fully utilize the semantic understanding ability of a multi-modal large model. Direct application of a general visual language model to a signal scene faces the problems of lack of structural priori and poor coordinate expression ability. The present application innovatively introduces a YOLO target detection model to automatically annotate the signal time-frequency graph, and builds an annotated data set based on this. Further, the visual language model (VLM) is fine-tuned to enable it to accurately identify the modulation type on the time-frequency graph and output the signal target coordinates.

[0087] The method fuses image perception, text encoding and spatial regression ability through the construction of a "target detection + visual language model fine-tuning" joint training framework, realizes semantic recognition and positioning reasoning of the target signal region in the time-frequency graph, introduces image-text matching supervision and coordinate regression loss in the system training process, and significantly improves the identification ability of the multi-modal model for fine-grained regions in the signal image. The method is widely applicable to various complex electromagnetic scenes such as signal detection and spectrum situation recognition, and can provide key support for high-precision signal intelligent processing.

[0088] Embodiment two

[0089] In this embodiment, a computer terminal device is provided, comprising:

[0090] one or more processors;

[0091] a memory coupled to the processor, for storing one or more programs;

[0092] When the one or more programs are executed by the one or more processors, the one or more processors implement the steps of the above-mentioned signal modulation recognition and positioning method based on time-frequency graph and visual language model.

[0093] In the embodiment, a computer readable storage medium is also provided, and the computer readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned signal modulation recognition and positioning method based on time-frequency graph and visual language model are implemented.

[0094] In the embodiment, an electronic device is also provided, and the electronic device includes a memory and a processor. The memory stores a computer program, and the processor is configured to execute the computer program to implement the steps of the above-mentioned signal modulation recognition and positioning method based on time-frequency graph and visual language model.

[0095] In the embodiment, a computer program product is also provided, and the computer program product includes a computer program. When the computer program is executed by a processor, the steps of the above-mentioned signal modulation recognition and positioning method based on time-frequency graph and visual language model are implemented.

[0096] The above-mentioned program can be executed in a processor, or can also be stored in a memory (or called a computer readable medium). The computer readable medium includes permanent and non-permanent, removable and non-removable media, and can be implemented by any method or technology to store information. The information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0097] These computer programs can also be loaded into a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer implemented process, and the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows Figure 1 The steps of the functions specified in one or more flows or one or more blocks can be implemented by different modules. Figure 1 The steps of the functions specified in one or more flows or one or more blocks can be implemented by different modules.

[0098] The embodiment provides such a device or system. The system is called a signal modulation recognition and positioning system based on a time-frequency map and a visual language model, comprising:

[0099] A signal preprocessing and time-frequency map generation module is configured to generate a time-frequency map by performing time-frequency transformation on an original IQ signal;

[0100] An automatic annotation module is configured to detect the time-frequency map by using a target detection model fine-tuned based on a signal dataset, and automatically generate a bounding box coordinate and a category label of a signal target in the time-frequency map;

[0101] A multi-modal training data construction module is configured to generate corresponding descriptive text by using a language model based on the bounding box coordinate and the category label, so as to construct a multi-modal training pair comprising the time-frequency map, the descriptive text and the bounding box coordinate;

[0102] A visual language model fine-tuning module is configured to fine-tune a visual language model by using the multi-modal training pair, and the fine-tuning process jointly optimizes a picture-text matching loss and a coordinate regression loss;

[0103] A signal recognition and output module is configured to recognize a new time-frequency map by using the fine-tuned visual language model, and output a category and a coordinate information of a target signal.

[0104] As an implementation manner of the embodiment, the signal preprocessing and time-frequency map generation module comprises:

[0105] A transformation unit is configured to perform a short-time Fourier transformation on the original IQ signal;

[0106] A parameter configuration unit is configured to configure a Hanning window function, a window length, a window overlap rate and a Fourier transformation point number in the transformation process;

[0107] An image generation unit is configured to generate a single-channel time-frequency map image.

[0108] As an implementation manner of the embodiment, the automatic annotation module comprises:

[0109] A model fine-tuning unit is configured to adopt a pre-trained YOLO model and fine-tune the YOLO model by using a Wideband-Sig53 dataset;

[0110] An inference unit is configured to input the time-frequency map into the fine-tuned YOLO model after scaling;

[0111] An output analysis unit is configured to analyze a normalized bounding box coordinate and a preset category label output by the model.

[0112] As an implementation manner of the embodiment, the multi-modal training data construction module comprises:

[0113] a coordinate mapping unit configured to map the bounding box coordinates into actual time and frequency range information;

[0114] a text generation unit configured to combine the category label and the time and frequency range information to generate a natural language description containing signal type and spatial location semantics;

[0115] a triple construction unit configured to construct a pairing relationship of the time-frequency map, the natural language description and the bounding box coordinates.

[0116] As an implementation in the embodiment, the visual language model fine-tuning module comprises:

[0117] a loss calculation unit configured to calculate an information noise contrast estimation loss between image features and text features, and an L1 loss between predicted coordinates and real coordinates;

[0118] a joint optimization unit configured to train the model with the weighted sum of the two losses as a joint loss function.

[0119] As an implementation in the embodiment, the signal recognition and output module comprises:

[0120] a preprocessing unit configured to perform scaling and normalization preprocessing on a new time-frequency map;

[0121] a query interface unit configured to receive a natural language query;

[0122] a multi-modal inference unit integrated with a visual encoder, a text encoder, a cross-modal attention module and a coordinate regression branch, configured to cooperatively process input information;

[0123] a structured output unit configured to output structured information containing signal category, time range, frequency range and bounding box coordinates.

[0124] The system or device is used to realize the functions of the methods in the above embodiments, each module in the system or device corresponds to each step in the method, and has been described in the method and will not be repeated here.

[0125] Through the above embodiments, the problems of signal modulation recognition and positioning based on time-frequency maps and visual language models in the related art are solved, thereby being able to guarantee the problems in the prior art.

[0126] The above merely describes the preferred embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A signal modulation recognition and localization method based on time-frequency map and visual language model, characterized in that, The method comprises the following steps: S1, generating a time-frequency graph by time-frequency transformation of an original IQ signal; S2, detecting the time-frequency graph using a target detection model fine-tuned by a signal dataset to automatically generate the bounding box coordinates and class labels of signal targets in the time-frequency graph; S3, generating corresponding descriptive text using a language model based on the bounding box coordinates and class labels to construct a multi-modal training pair containing the time-frequency graph, the descriptive text, and the bounding box coordinates; S4, fine-tuning a visual language model using the multi-modal training pair, and the fine-tuning process jointly optimizes the image-text matching loss and the coordinate regression loss; S5, identifying a new time-frequency graph using the fine-tuned visual language model to output the class and coordinate information of the target signal.

2. The method of claim 1, wherein, In S1, the process of generating a time-frequency graph by time-frequency transformation of an original IQ signal includes: performing short-time Fourier transform on the original IQ signal, using a Hanning window function, and setting the window length, window overlap rate, and Fourier transform point number to generate a single-channel time-frequency graph image.

3. The method of claim 1, wherein, In S2, the process of detecting using a target detection model includes: using a pre-trained YOLO model, fine-tuning using a Wideband-Sig53 dataset; scaling the time-frequency graph and inputting it into the fine-tuned YOLO model to output normalized bounding box coordinates and preset class labels.

4. The method of claim 1, wherein, In S3, the process of generating descriptive text includes: combining the class labels and the time and frequency range information mapped from the bounding box coordinates to generate natural language descriptions containing signal type and spatial location semantics.

5. The method of claim 1, wherein, In S4, the process of jointly optimizing the image-text matching loss and the coordinate regression loss includes: calculating the information noise contrast estimation loss between image features and text features, and calculating the L1 loss between predicted coordinates and real coordinates, and taking the weighted sum of the two as the joint loss function.

6. The method of claim 1, wherein, In S5, the process of identifying a new time-frequency graph includes: inputting the preprocessed time-frequency graph into the fine-tuned visual language model, and receiving natural language queries; through the cooperative work of the visual encoder, text encoder, cross-modal attention module, and coordinate regression branch in the model, outputting structured information containing signal class, time range, frequency range, and bounding box coordinates.

7. A signal modulation identification and localization system based on time-frequency map and visual language model, characterized in that, The system comprises: a signal preprocessing and time-frequency graph generation module for generating a time-frequency graph by time-frequency transformation of an original IQ signal; an automatic labeling module for detecting the time-frequency graph using a target detection model fine-tuned by a signal dataset to automatically generate the bounding box coordinates and class labels of signal targets in the time-frequency graph; a multi-modal training data construction module for generating corresponding descriptive text using a language model based on the bounding box coordinates and class labels to construct a multi-modal training pair containing the time-frequency graph, the descriptive text, and the bounding box coordinates; a visual language model fine-tuning module for fine-tuning a visual language model using the multi-modal training pair, and the fine-tuning process jointly optimizes the image-text matching loss and the coordinate regression loss; A signal recognition and output module is configured to recognize a new time-frequency graph by using the fine-tuned visual language model, and output a category and coordinate information of a target signal.

8. A computer terminal device, characterized by The method comprises the steps of: one or more processors; a memory coupled to the processors, the memory storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the steps of the method according to any one of claims 1-6.

9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method according to any one of claims 1-6.

10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Smoking identification method based on multi-modal large model

    CN119027855A

  • Scene graph generation enhancement method based on fine-tuning large language model

    CN120450979A