System and method for automated synthetic radiofrequency scene dataset generation

The automated system generates realistic RF datasets by defining waveform characteristics and scenario rules, addressing the limitations of manual production and existing synthetic datasets, enabling efficient training of machine learning models for complex RF environments.

WO2025171465A9PCT designated stage Publication Date: 2025-10-23QOHERENT INC +5
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CA2024/051552
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-16
Filing Date
2024-11-22
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Manually producing RF datasets for training machine learning models is error-prone, labor-intensive, and costly, and existing synthetic datasets are often unrealistic and lack scalability, failing to model complex RF signal environments with varied impairments.

Method used

A system and method for automated synthetic radiofrequency scene dataset generation that determines waveform characteristics, scenario rules, and environment variables to generate realistic RF recordings, incorporating impairments and metadata, allowing for large-scale, scenario-based training data creation.

Benefits of technology

Enables the production of high-quality, scalable, and realistic synthetic RF datasets that reduce manual labor and hardware requirements, facilitating efficient training of machine learning models for complex RF environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CA2024051552_23102025_PF_FP_ABST
    Figure CA2024051552_23102025_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods for producing synthetic radiofrequency recordings are provided. The method includes: determining a first set of waveform signal characteristics based on permutations of generated transmitters of a signal; determining a first set of rules based on permutations of possible configurations of a set of scenarios for the signal; determining a second set of rules based on environment variables for each scenario; generating individual signals based on the first set of waveform signal characteristics, the first set of rules, and the second set of rules; receiving a set of symbols that is a waveform representation of a bit string corresponding to the individual signals; creating a configuration for the individual signals; and generating and labeling based on at least one of the configurations for the individual signals or constraints to the configurations for use in training or testing machine learning models.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEM AND METHOD FOR AUTOMATED SYNTHETIC RADIOFREQUENCY SCENE DATASET GENERATIONREFERENCE TO PRIORITY APPLICATION

[0001] The present application claims priority to U.S. Provisional Patent Application no. 63 / 554,714, which was filed February 16, 2024, the content of which is incorporated herein by reference in its entirety.FIELD

[0002] Various embodiments are described herein that generally relate to a system for automated synthetic radiofrequency scene dataset generation, as well as the methods therefore.BACKGROUND

[0003] The following paragraphs are provided by way of background to the present disclosure. They are not, however, an admission that anything discussed therein is prior art or part of the knowledge of persons skilled in the art.

[0004] Manually producing datasets for training machine learning (ML) models is an error-prone, time consuming, and labor-intensive process that is generally completed in an ad hoc manner. The dataset resulting from this manual process often contains errors, and these errors cause ML models trained on the dataset to learn incorrect behaviors that adversely affect their performance in the deployment environment (a problem generally known as distributional shift).

[0005] Manually producing datasets of RF recordings for the purpose of training ML models is an especially challenging task, because recording RF requires specialized hardware, and recognizing the contents of an RF recording requires specialist skills. These requirements increase the cost of manual RF dataset production substantially, relative to a simpler task such as recognizingimages of common household items (which requires neither specialized equipment nor specialist skills).

[0006] RF datasets can be synthesized using software based on mathematical models of radio channels, which greatly reduces the labor and hardware requirements for producing such a dataset. However, synthetic RF datasets are often highly simplified and unrealistic compared to actual recordings of RF signals. Recordings of RF spectrum taken in real environments typically show a complex mix of signals that are distorted by a wide variety of impairments, which is presently difficult to model realistically in a synthetic dataset. Currently available tools for RF dataset synthesis are limited in their functionality, being only capable of modeling single signals acted on by a fixed set of impairments. Extending the functionality of these extant synthesis tools is challenging for an end user as the existing tools are not easily scalable

[0007] In some synthetic work, examples produced may not be scenariobased or fit for a specific task (modulation recognition). In some cases, they do not attempt to generalize and are fixed solutions.

[0008] There is a need for a system and method that addresses the challenges and / or shortcomings described above.SUMMARY OF VARIOUS EMBODIMENTS

[0009] Various embodiments of a system and method for automated synthetic radiofrequency scene dataset generation, and computer products for use therewith, are provided according to the teachings herein.

[0010] According to one aspect of the invention, there is disclosed a method for producing synthetic radiofrequency recordings, the method comprising: determining a first set of waveform signal characteristics based at least in part on permutations of generated transmitters of a signal; determining a first set of rules based on permutations of possible configurations of a set of scenarios for the signal; determining a second set of rules based on environment variables for each scenario in the set of scenarios; generating individual signals basedon one or more of the first set of waveform signal characteristics, the first set of rules, and the second set of rules; receiving a set of symbols that is a waveform representation of a bit string corresponding to the individual signals; creating a configuration for the individual signals; and generating and labeling a plurality of recordings based on at least one of the configurations for the individual signals or constraints to the configurations for use in training or testing machine learning models.

[0011] The rules can include for example waveform design rules (including modulation, protocol design stack, error correction, and pulse shaping); impairment rules; channel model rules; message contents and timing rules (or timing structure rules).

[0012] Impairments can include hardware impairments, for example those impairments related to the physical components of a radio emitter, from message receipt to emission of the encoded waveform.

[0013] Other features and advantages of the present application will become apparent from the following detailed description taken together with the accompanying drawings. It should be understood, however, that the detailed description and the specific examples, while indicating preferred embodiments of the application, are given by way of illustration only, since various changes and modifications within the spirit and scope of the application will become apparent to those skilled in the art from this detailed description.BRIEF DESCRIPTION OF THE DRAWINGS

[0014] For a better understanding of the various embodiments described herein, and to show more clearly how these various embodiments may be carried into effect, reference will be made, by way of example, to the accompanying drawings which show at least one example embodiment, and which are now described. The drawings are not intended to limit the scope of the teachings described herein.

[0015] FIG. 1 shows a schematic diagram of an example embodiment of a system for automated synthetic radiofrequency scene dataset generation.

[0016] FIG. 2 is a schematic representation of an example signal synthesizer according.

[0017] FIG. 3 is a schematic representation of an example system for the generation of a signal.

[0018] FIG. 4 is a schematic representation of an example system for generating a signal subjected to rules and impairments.

[0019] FIG. 5 is a schematic representation of an example interaction between a procedure generation executive and a signal synthesizer.

[0020] FIG. 6 is a schematic representation of an example interaction between a procedure generation executive and a signal synthesizer, where the dataset includes metadata.

[0021] FIGs. 7(a)-7(j) are schematic representations of channel impairments where the bottom portion of each representation is a constellation of symbols of individual hardware impairments to a signal, and the top portion of each representation is a spectrogram thereof.

[0022] FIG. 8 is a schematic representation of an example generation of a scenario using multiple signal synthesizers.

[0023] FIG. 9 is a schematic representation of an example collision impairment scenario.

[0024] FIGs. 10(a)-10(h) are schematic time-series spectrograms of example channel impairment scenarios.

[0025] FIG. 11 is a schematic representation of an example configuration for generation of waveforms subjected to channel impairments, where the signal is transmitted over the air.

[0026] FIG. 12 is a schematic representation of example metadata stored with individual synthetic waveforms

[0027] FIG. 13 is a schematic representation of an example time series and spectrogram of a generated signal with associated metadata.

[0028] Further aspects and features of the example embodiments described herein will appear from the following description taken together with the accompanying drawings.DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] Various embodiments in accordance with the teachings herein will be described below to provide an example of at least one embodiment of the claimed subject matter. No embodiment described herein limits any claimed subject matter. The claimed subject matter is not limited to devices, systems, or methods having all of the features of any one of the devices, systems, or methods described below or to features common to multiple or all of the devices, systems, or methods described herein. It is possible that there may be a device, system, or method described herein that is not an embodiment of any claimed subject matter. Any subject matter that is described herein that is not claimed in this document may be the subject matter of another protective instrument, for example, a continuing patent application, and the applicants, inventors, or owners do not intend to abandon, disclaim, or dedicate to the public any such subject matter by its disclosure in this document.

[0030] It will be appreciated that for simplicity and clarity of illustration, where considered appropriate, reference numerals may be repeated among the figures to indicate corresponding or analogous elements. In addition, numerous specific details are set forth in order to provide a thorough understanding of the embodiments described herein. However, it will be understood by those of ordinary skill in the art that the embodiments described herein may be practiced without these specific details. In other instances, well- known methods, procedures, and components have not been described in detail so as not to obscure the embodiments described herein. Also, the description is not to be considered as limiting the scope of the embodiments described herein.

[0031] It should also be noted that the terms “coupled” or “coupling” as used herein can have several different meanings depending in the context in whichthese terms are used. For example, the terms coupled or coupling can have a mechanical or electrical connotation. For example, as used herein, the terms coupled or coupling can indicate that two elements or devices can be directly connected to one another or connected to one another through one or more intermediate elements or devices via an electrical signal, electrical connection, or a mechanical element depending on the particular context.

[0032] It should also be noted that, as used herein, the wording “and / or” is intended to represent an inclusive-or. That is, “X and / or Y” is intended to mean X or Y or both, for example. As a further example, “X, Y, and / or Z” is intended to mean X or Y or Z or any combination thereof.

[0033] It should be noted that terms of degree such as “substantially”, “about” and “approximately” as used herein mean a reasonable amount of deviation of the modified term such that the end result is not significantly changed. These terms of degree may also be construed as including a deviation of the modified term, such as by 1 %, 2%, 5%, or 10%, for example, if this deviation does not negate the meaning of the term it modifies.

[0034] Furthermore, the recitation of numerical ranges by endpoints herein includes all numbers and fractions subsumed within that range (e.g., 1 to 5 includes 1 , 1.5, 2, 2.75, 3, 3.90, 4, and 5). It is also to be understood that all numbers and fractions thereof are presumed to be modified by the term “about” which means a variation of up to a certain amount of the number to which reference is being made if the end result is not significantly changed, such as 1 %, 2%, 5%, or 10%, for example.

[0035] It should also be noted that the use of the term “window” in conjunction with describing the operation of any system or method described herein is meant to be understood as describing a user interface for performing initialization, configuration, or other user operations.

[0036] The example embodiments of the devices, systems, or methods described in accordance with the teachings herein may be implemented as a combination of hardware and software. For example, the embodimentsdescribed herein may be implemented, at least in part, by using one or more computer programs, executing on one or more programmable devices comprising at least one processing element and at least one storage element (i.e., at least one volatile memory element and at least one non-volatile memory element). The hardware may comprise input devices including at least one of a touch screen, a keyboard, a mouse, buttons, keys, sliders, and the like, as well as one or more of a display, a printer, and the like depending on the implementation of the hardware.

[0037] It should also be noted that there may be some elements that are used to implement at least part of the embodiments described herein that may be implemented via software that is written in a high-level procedural language such as object-oriented programming. The program code may be written in C++, C#, JavaScript, Python, or any other suitable programming language and may comprise modules or classes, as is known to those skilled in object-oriented programming. Alternatively, or in addition thereto, some of these elements implemented via software may be written in assembly language, machine language, or firmware as needed. In either case, the language may be a compiled or interpreted language.

[0038] At least some of these software programs may be stored on a computer readable medium such as, but not limited to, a ROM, a magnetic disk, an optical disc, a USB key, and the like that is readable by a device having a processor, an operating system, and the associated hardware and software that is necessary to implement the functionality of at least one of the embodiments described herein. The software program code, when read by the device, configures the device to operate in a new, specific, and predefined manner (e.g., as a specific-purpose computer) in order to perform at least one of the methods described herein.

[0039] At least some of the programs associated with the devices, systems, and methods of the embodiments described herein may be capable of being distributed in a computer program product comprising a computer readable medium that bears computer usable instructions, such as program code, forone or more processing units. The medium may be provided in various forms, including non-transitory forms such as, but not limited to, one or more diskettes, compact disks, tapes, chips, and magnetic and electronic storage. In alternative embodiments, the medium may be transitory in nature such as, but not limited to, wire-line transmissions, satellite transmissions, internet transmissions (e.g., downloads), media, digital and analog signals, and the like. The computer useable instructions may also be in various formats, including compiled and non-compiled code.

[0040] At least some of the embodiments described herein provide a radiofrequency (RF) dataset synthesis tool that incorporates a wide range of signal sources and impairment models within a modular and extensible framework that allow for the creation of more realistic synthetic RF datasets.

[0041] In accordance with the teachings herein, there are provided various embodiments for automated synthetic RF scene dataset generation, and computer products for use therewith. At least some of these embodiments may also be considered to be methods for procedural generation of RF scene recordings according to a user-defined scenario, as well as the generation of annotations associated with a given scene recording.

[0042] Thus, in some embodiments, there is described a system and method for dataset generation via synthetic means, generating high volumes of “scenes” that contain multiple transmitters, based on rules that define what the scene consists of.

[0043] In some embodiments presented the disclosure outlines methods and systems for “collision recognition” and “jamming recognition” training scripts or combined together as a “impairment scenario recognition” problem for a given communications channel.

[0044] In some examples, the procedural dataset generation can be completed purely synthetically or obtained through data transmitted over-the- air on a controlled software-defined radio testbed.

[0045] The instant disclosure is directed thus to procedural dataset generation of RF signals, which can then be used as a training set for machine learning tools. In the context of the embodiments described herein, procedural refers to imposing predetermined constraints or conditions on the RF signals that are to be generated. In some examples, a scene is predetermined. A scene is a series of RF signals generated according to instructions provided to a signal synthesizer, according to rules or constraints imposed on the signal synthesizer and a message to be encoded and transmitted. Each scene comprises a plurality of single synthetic signals, each of which follows a predetermined procedure and ruleset, and each single synthetic signal differs from other synthetic signals through the variation of one or more of the parameters imposed on the waveform, as will be further detailed below.

[0046] A scenario is the generation of a plurality of scenes generated for an overall scenario condition. In other words, in some examples, a scenario could be the generation of a dataset of signals under a theme, and scenes are the generation of individual RF waveforms that cover a wide range of permutations for the RF waveforms.

[0047] As a non-limiting example, a scenario could be for the generation of a dataset where the theme is collision of signals. Each scene would then be the generation of a plurality of RF waveforms subject to permutations, such as different modulations, doppler shift, frequency offset, various degrees of Raleigh fading, etc., as will be further detailed below.

[0048] Another scenario could be the generation of a dataset where the theme is satellite communications between a base station and a satellite. Another scenario could be the generation of a dataset for communication through an urban environment. There are as many scenarios as there are use cases for transmission of RF signals.

[0049] As the scenes and ultimately the scenario is being constructed, relevant metadata is captured. Given that everything captured is inherently synthesized, significant amounts of metadata can be captured.

[0050] When the process is operated iteratively, following the rules of the scene definition procedure, and at scale, and can produce large scale datasets for use in deep learning training for many tasks.

[0051] It is important to note that single signal synthesis may be performed procedurally as well, with its own ruleset. An example embodiment of that is to artificially add impairments to a signal (impairment measurement recognition) or to perform specific emitter identification (RF fingerprinting).

[0052] For example, an embodiment can include the generation of a single signal through multiple individual generation of signals where modifications are made to the constraints.

[0053] However, for most datasets, the process begins by the creation of a scene, which is the encoding of a signal to be transmitted over a channel, where either of the encoding of the signal or the channel are subjected to rules or impairments. The signal is then “transmitted” over the channel, in the sense that channel impairments are applied to the signal. The resulting impaired signal is stored in memory. A scene is then the initial set of impairments that characterize the channel. Examples of such impairments or rules include (but are not limited to): channel bandwidth, number of signal sources, occupied bandwidths of each signal source, transmission duty cycle of each signal source, number of noise sources, etc. A scenario includes a multiplicity of scenes that are each constrained by a theme.

[0054] Unless otherwise specified, the terms rules, scenario, scene, scene recording, and annotation are defined as follows.

[0055] Rules are mathematically derived constraints from events in an RF signal. For example, a “collision” occurs when i.e. the nominal bandwidth of two signals intersect. A bleedover occurs when the occupied bandwidth of one signal intersects with the nominal bandwidth of another signal. A person skilled in the art will readily recognize that many other events can be defined and expressed as mathematical relationships that are then defined as individual rules.

[0056] A scenario is a set of parameters describing a simulated radiofrequency channel, according to a theme. A theme is defined as a common constraint under which all scenes are subjected to. For example, a scenario can be based on the rule of “collision”. A scene is a plurality of RF signals generated for the theme, where the RF signal is subjected to predetermined constraints on the signal itself. These constraints can be for example emitter-related or channel related. In other words, an emitter such as an antenna may inject into a transmitted signal some variations in the output that are inherent to the emitter itself. Examples of such variations include phase noise, IQ imbalance and gain fluctuation. In addition, a transmitted RF signal will be subjected to the conditions of the channel through which the signal propagates. In other words, a scenario is a selected set of rules for a given procedure. A scenario’s parameters may be defined by the user or generated via an external procedure. A scenario may be electronically saved by the user to be loaded again at a different time. The values of a scenario’s parameters may be fixed, or they may be specified to refer to a range of permissible values. In this case, it may be required to select a fixed value from the permissible range of parameter values before a scene can be generated from the scenario.

[0057] A scene is understood as one instance of a simulation of an RF channel over time, defined according to a scenario.

[0058] A scene recording is a synthetic RF recording consisting of a time series of numerical samples that describe a scene over a specified interval of time, as well as accompanying annotations, in the case where annotations are generated along with the RF waveforms. A scene recording may be plotted for display (with or without visible Annotations) via a spectrogram, with orthogonal axes for time, frequency, and amplitude. While the expression “recording” is used throughout the instant description, it will be appreciated that a recording is an electronic representation of a waveform that can be reproduced, and includes all forms of digitally storing a waveform.

[0059] Annotations are a series of descriptive notes that are generated alongside and incorporated into the scene recording, which accurately describethe events taking place within the scene during the duration of the generation of the waveforms. Since annotations are generated contemporaneously with their associated scene recording based on the internal, explicitly defined properties of the scene, annotations can be very accurate in their description of the events in a scene recording. In fact, annotations generated alongside the scene recording are generally more accurate than would otherwise be possible analyzing the recordings through known means without such annotations.

[0060] Annotations may include information to localize an event in time and frequency. An example of the information contained in a typical Annotation would be: “Between times T1 and T2, and within the frequency range bound by F1 and F2, Signal Source A’s transmission collided with Signal Source B’s transmission, causing mutual interference”. Annotations are intended both to enhance human understanding of a Scene Recording, as well as to provide a source of data for training and testing of Machine Learning (ML) models. To facilitate training and testing of ML models, a Scene Recording’s Annotations may be partially or completely hidden from the ML model’s input, forcing the model to rely on the Scene Recording’s samples as its main or sole input from which its output may be calculated. For example, a ML model may be trained to recognize signal collision events within a Scene Recording based off of the recording samples alone, and its accuracy at recognizing such events may be tested by comparing the model’s output to the subset of Annotations that describe all signal collisions contained in the Scene Recording. Annotations can then be used to increase the ML model’s accuracy by retraining the model and augmenting the recording samples with the Annotations.

[0061] Reference is first made to FIG. 1 , showing block diagram of an example embodiment of a system 100 for automated synthetic radiofrequency scene dataset generation. The system 100 includes at least one user device 110 and at least one server 120. The user device 110 and the server 120 may communicate, for example, wirelessly or over the Internet.

[0062] The user device 110 may be a computing device that is operated by a user. The user device 110 may be, for example, a smartphone, a smartwatch,a tablet computer, a laptop, a virtual reality (VR) device, or an augmented reality (AR) device. The user device 110 may also be, for example, a combination of computing devices that operate together, such as a smartphone and a sensor. The user device 110 may also be, for example, a device that is otherwise operated by a user, such as a drone, a robot, or remote-controlled device; in such a case, the user device 110 may be operated, for example, by a user through a personal computing device (such as a smartphone). The user device 110 may be configured to run an application (e.g., a mobile app) that communicates with other parts of the system 100, such as the server 120.

[0063] The server 120 may run on a single computer, including a processor unit 124, a display 126, a user interface 128, an interface unit 130, input / output (I / O) hardware 132, a network unit 134, a power unit 136, and a memory unit (also referred to as “data store”) 138. In other embodiments, the server 120 may have more or less components but generally function in a similar manner. For example, the server 120 may be implemented using more than one computing device, which may or may not be physically proximate, for instance in a cloud computing environment.

[0064] The processor unit 124 may include a standard processor, such as the Intel Xeon processor, for example. Alternatively, there may be a plurality of processors that are used by the processor unit 124, and these processors may function in parallel and perform certain functions. The display 126 may be, but not limited to, a computer monitor or an LCD display such as that for a tablet device. The user interface 128 may be an Application Programming Interface (API) or a web-based application that is accessible via the network unit 134. The network unit 134 may be a standard network adapter such as an Ethernet or 802.11x adapter.

[0065] The processor unit 124 may execute a predictive engine 152 that functions to provide predictions by using machine learning models 146 stored in the memory unit 138. The predictive engine 152 may build a predictive algorithm through machine learning. The training data may include, for example, image data, video data, audio data, and text.

[0066] The processor unit 124 can also execute a graphical user interface (GUI) engine 154 that is used to generate various GUIs. The GUI engine 154 provides data according to a certain layout for each user interface and also receives data input or control inputs from a user. The GUI then uses the inputs from the user to change the data that is shown on the current user interface, or changes the operation of the server 120 which may include showing a different user interface.

[0067] The memory unit 138 may store the program instructions for an operating system 140, program code 142 for other applications, an input module 144, a plurality of machine learning models 146, an output module 148, and a database 150. The machine learning models 146 may include, but are not limited to, image recognition and categorization algorithms based on deep learning models and other approaches. The database 150 may be, for example, a local database, an external database, a database on the cloud, multiple databases, or a combination thereof.

[0068] The programs 142 comprise program code that, when executed, configures the processor unit 124 to operate in a particular manner to implement various functions and tools for the system 100.

[0069] At least some of the embodiments described herein are useful for the creation of many recordings representative of a range of desirable scenarios. Recordings themselves as a collection can be used directly as a dataset, but are not considered a “machine learning ready” dataset.

[0070] At least some of the embodiments described herein provide a methodology and workflow for generating large collections of synthetic and labelled radio frequency (RF) recordings, representative of desirable scenarios applicable to training machine learning models, for inclusion into an RF dataset.

[0071] Users set several permutations of generated transmitters, set several permutations of possible configurations of a set of scenarios, set several environment variables for each scenario, and / or create a configuration for what is generated. The system can be defined with synthesizers directly or can callon third-party software to synthesize RF signals. Examples of such third-party software include but is not limited to GNU Radio or MATLAB. In usage, the system may call on synthesizers defined in the system, then the system automatically generates and labels a large number of recordings based on the configurations or constraints to the configurations that can be used as part of a dataset, or be used directly for training or testing machine learning models.

[0072] The system may use an architecture that attempts to generalize the difficulty around generating scenario-based synthetic RF machine learning data. Some or all elements (transmitters, permutations, configurations, scenarios, impairments, synthesizers) may be modular, interchangeable, and / or implemented as plug-ins.

[0073] The system may be configured so that any parameter within the system can be included as a label (e.g., the nature of transmitters, collisions, interference, and other impairments).

[0074] In at least some of the embodiments described herein, the system may carry out a method for producing large amounts of qualified, labelled / annotated, high-quality synthetic RF recordings in an automated manner, in a short amount of time.

[0075] In the typical implementation, data consists of “scenes” which are a collection of individual signals each subjected to different constraints. Scenes are based on scenarios. It will be apparent to a person skilled in the art that the system 100 provides an adequate user interface to permit a user to set or select the different rules for the generation of RF waveforms. Any component of a scene can be labelled since it is synthetic and can be controlled.

[0076] At least one of the advantages of some of the embodiments described herein is that the embodiments obviate the need for individual RF signals to be programmed one at a time, and then generated manually.

[0077] One of the challenges addressed by some of the embodiments described herein is the requirement for the user to understand the methodology of producing the scenarios, waveforms, etc. and requirement to envisionscenarios they need to produce In some examples, the system and method provide a predetermined set of rules and constraints that can be selected by a user to define a scenario and individual scenes in a user-friendly graphical user interface.

[0078] In addition, while there are limited features and ways to “constrain” the generation of RF waveforms (e.g., constraints inherent to modulation and transmission parameters, all of the other parameters (rules and constraints) can be selected at random.

[0079] Thus, in some examples, the system is configured to generate a library of waveforms . A library of signals is, in some examples, a collection of different signal types categorized by i.e. protocols (5H, LTE, etc.), modulations (APSK, QPSK, etc.) and modes (wideband FM, narrowband FM, AM). Individual waveforms from the library can then be selected by a user or by the procedural engine to add as needed for a scene.

[0080] It will be apparent to a person skilled in the art that while many of the examples described herein refer to the synthetic generation of waveforms using i.e. synthesizers, other examples include using software-defined rations (SDRs), for emitting a signal that is subsequently captured. Examples of media through which the signal can travel include over-the-air or through a cable

[0081] In some embodiments, a system and method for automatic generation of synthetic waveforms includes one or more of the following components:1) Scene Synthesizer: a) A signal configuration list is defined (consisting of modulations, transmission parameters, bandwidths, structure, etc.) per type of signal for a given scenario. b) A signal impairment list is defined, outlining impairments that are added to individual signals. c) A scenario configuration list is defined, outlining a list of scenarios, and any one or more of:i) the expected environment in each scenario (types of present signals, impairments on them, how and where they could be transmitting), ii) whether the scenario is subjected to random or fixed parameters iii) rules dictating what qualifies as the given scenario. d) A configuration for qualified transmitters is defined (e.g., incumbent transmitter and their possible signal configurations, interference transmitter or adjacent transmitter and their possible signal configurations). In some examples, a qualified transmitter is a transmitter that has had rules as set out by the synthesizer or the procedural engine, or both, imposed theron. e) Individual signal synthesizers (signal generators) are defined, taking in inputs based on signal configuration list (a) and are adapted to take an input message, modulate the input message, then produce a time series waveform. This may be manually coded (e.g. coded in python or C++) or it may leverage a synthesizer library produced in a third-party tool (GNU Radio or MATLAB). In all cases, it may be author defined. f) A procedure generation executive contains the instructions provided by elements (a) to (e) above, and: i) produces a number of signals which can be defined by a user, either randomly from a constrained list of configurations, or a fixed configuration (for each signal). ii) each signal is impaired, either randomly from a constrained list, or from a fixed configuration. iii) For each signal, a bounding box label is created, representative of e.g. start time, end time, start frequency, end frequency, indicating a bounding box for all relevant signal dimensions (e.g., nominal bandwidth, occupied bandwidth, peak, center frequency). iv) Generating each signal to create a modified signal;v) Optionally annotating each signal concurrently with its generation; vi) adding all signals for each scene into a dataset vii) Optionally adding Impairments to the resulting dataset;. viii) Additional labels are created and associated with the bounding box based on the rules and constraints that are relevant to the scenario (e.g., in a collision scenario, finding intersections between occupied bandwidths, or one occupied bandwidth and a nominal bandwidth). ix) The synthetic recording and its label metadata is saved to a file.2) Procedure Generation Executive - is a system and method for automated production of a plurality of recordings, executing a plurality of instances of the scene generator generating a plurality of individual waveforms according to the scene rules and constraints.

[0082] The following are example usages, transmitter configurations, scenarios, and modulators.Example usage 1 : python3 synth_rec_ex . py -e 2 -1 5000000 10000000 -r 60000000100000000 -x npyThis particular code will call a module 220 named “synth_rec_ex.py”, which is a computer-implemented method of according to some embodiments. In an example, the module 220 will read the list of possible commands 210 selected by a user as well as the configuration file 230 for some of the constraints to be imposed on signal generation. Module 220 will instruct synthesizer 250 to generate a recording 270 taking into consideration rules 240 such as waveform definitions, impairment definitions, scenario definitions and transmitter definitions. Each synthetic recording produced is saved to a file 280. InExample 1, around 17gb of Numpy files across all defined scenarios (in the Example 1, the scenarios included "collisions", "bleedover", "wbjamming", "ctnbjamming", "Ifmjamming", "birdies", "nojmpairment", and "low_snr"). In this example, 2 random examples per configuration (i.e., in this case, the configuration is the characterization of the permutations of a scene synthesizer - here, the configuration would be “for every permutation of the scene synthesizer, generate two example) are created, of length 5M (mega samples) and length 10M and sample rates 60MS / s (mega samples per second) and 100 MS / s.Example usage 2: python3 synth_rec_ex . py -e 6 -1 500000 1000000 5000000 10000000 -r 40000000 61440000 100000000 -x npyThis example produces around 83gb of Numpy files across all scenarios (In Example 2, the same scenarios were used:"collisions", "bleedover", "wbjamming", "ctnb amming", "Ifmjamming", "birdies", "nojmpairment", and "low_snr"). Example 2 creates 6 random examples per configuration, of length 5M, 10M 50M, 100M and sample rates 40MS / s 61.44MS / S and 100 MS / s.Example transmitter configurations for a DVB-S2 signal synthesizer could be:' frame_length ' : [64800,16200] ,' symbol_rates [7000000,7000000,27500000,30000000,34000000] ,' tx_rate_pct [10, 40, 70, 100] ,'modulations ' : [ "bpsk", "qpsk", "8psk", "16apsk", "32apsk", "64apsk", "128apsk", "256apsk"] ,"amplitudes" : [5, 15,25,40,60, 80, 100] ,"mode": ["constant"] ,"bb_centre_f req" : [ round (x, 2 ) for x in np . arange (-1 , 1.2 , 0.2 ) ] , "roll_off" : [0.20, 0.25, 0.35] #[0.05, 0.10, 0.15, 0.20, 0.25, 0.35] } ## This may also be referred to as "type of interferer" interferer = {' frame_length ' : [ 64800, 16200, 8100] , #whatever' symbol_rates ' : [ 1400000, 3000000, 5000000,7000000, 10000000,20000000,27500000,30000000,34000000] , 'modulations ' : [ "bpsk", "qpsk", "8psk", "16apsk", "32apsk", "64apsk", "1 28apsk", "256apsk", "pam4", "qaml6", "qam32", "qam64", "qaml28", "qam256 "] ,' tx_rate_pct ' : [10, 40, 70, 100] ,"amplitudes" : [5, 15,25,40,60, 80, 100] ,"mode" : [ "constant", "hops", "Ifm"] ,"bb_centre_f req" : [ round (x, 2 ) for x in np . arange (-1 , 1.2 , 0.2 ) ] , "roll_off" : [0.05, 0.10, 0.15, 0.20, 0.25, 0.35] #higher = less interference}Example scenarios: if scenario == "collisions": interf erer_ns = Namespace ( f rame_length n . random. choice ( inter ferer [ "f rame_length" ] ) ,symbol_rate = np . random. choice ( inter ferer [ "symbol_rates " ] ) , modulation = n . random. choice ( inter ferer [ "modulations " ] ) , tx_rate_pct = np . random. choice ( inter ferer [ "tx_rate_pct " ] [ : -1] ) , amplitude = np . random. choice ( inter ferer [ "amplitudes" ] ) , mode = "hops", centre_freq =(inc_ns . centre_f req / inc_ns . sdr_srate) +np . random. choice (np . arange ( -2, 2.5, 0.5 roll_off = np . random. choice (interferer [ "roll_off"] [1:- 1] ) elif scenario == "bleedover": #High-gain adjacent channel user over incumbent transmitter . interf erer_ns = Namespace ( frame_length = np . random. choice ( interferer [ "f rame_length" ] ) , symbol_rate = np . random. choice ( interferer [ "symbol_rates " ] [ 3 : -1 ] ) , modulation = np . random. choice ( interferer [ "modulations " ] ) , tx_rate_pct = np . random. choice ( interferer [ "tx_rate_pct " ] ) , amplitude = np. random. choice ( [75, 80, 85, 90, 95, 100] ) , mode = "constant", centre_freq = inc_ns . centre_f req+0.5*inc_ns . bw*np . random. choice ( [ -1.55, -1.3 , - 1.05, -0.8 , -0.6, 0.7 , 0.95, 1.2 , 1.45 roll_off = np . random. choice ( interferer [ "roll_of f " ] [ : 3 ] ) )elif scenario =="wb_j amming" : interf erer_ns = Namespace ( frame_length n . random. choice ( inter ferer [ "f rame_length" ] ) , symbol_rate np . random. choice ( inter ferer [ "symbol_rates " ] ) , modulation = "wb_jammer", tx_rate_pct = 100, amplitude = np. random. choice ( [70, 80, 85, 90] ) , mode = "constant", centre_freq inc_ns . centre_f req+np . random. choice (np . arange (-2 , 2.5, 0.5) ) , roll_off = np . random. choice ( [ 0.1 , 0.15, 0.2 ] ))Example modulators: def generate_pam4_symbols (num_symbols , variance=0.1 ) : symbols = np . random. choice (np . linspace (-3, 3, 4) , num_symbols) return add_noise ( symbols , variance) def generate_qaml6_symbols (num_symbols, variance=0.1 ) : """Generate QAM16 symbols.""" re_vals = np . array ( [ -1.5 , -0.5, 0.5, 1.5] ) / np.sqrt(lO) im_vals = np . array ( [ -1.5 , -0.5, 0.5, 1.5] ) / np.sqrt(lO) symbols = np . random. choice ( re_vals , num_symbols) + Ij * np . random. choice ( im_vals , num_symbols )

[0083] Figure 3 is a schematic representation of the main sub-components of module 200 and illustrates the data flow. In some embodiments, user device 110 is adapted to present to a user through a graphical user interface 128 the possible commands 210 available for synthesis, along with a set of rules 230. The process starts with the selection of a message 301 (which can begenerated on user device 110 or loaded into memory). The message, i.e. the string of bits that represent the message, is mapped to symbols 303 available in a constellation of symbols. Pulse shaping or generating waveforms 305 from the symbols is then carried out. The waveform is then multiplied 307 to offset the center frequency. This is well known in the art. The waveform is then passed through a channel model 309, where channel constraints are applied. The output of the channel model 309 is then saved, and a plurality of of individual files are saved in a dataset 280. Referring now for Fig. 4, there is shown for some embodiments a representation of how individual signal generation can be modified according to the scene and scenario requirements set out by a user. For example, instructions 401 contain rules as to what signal is to be generated while constraints 403 are the limits within which the generation of the waveform is to remain. More specifically, constraints 403 can include different modulations 411 to be applied to the symbols. Impairments specific to the symbols can be introduced at 413. During waveform generation, different filters or sample windows 415 can be applied. Additional filters or sample windows 417 (the same or different ones from filters or sample windows 415) can be applied during waveform multiplication. In this example, channel models 419 will constrain the behavior of the communication channel.

[0084] While Fig. 4 illustrates one iteration of a signal generation, Fig. 5 schematically represents how a procedure generation executive 501 is configured for a given scenario, and will load the different rules and constraints, specific generation instructions and message, and sequentially call the signal synthesizer 200 to generate a plurality of waveforms according to the various rules and constraints defined for a given scenario. Each individual generated waveform is stored in a dataset 280.

[0085] Fig. 6 adds to the detail of Fig. 5 the generation and storing of metadata related to the generation of each individual waveform generated by the signal synthesizer 200.

[0086] A non-limiting example of the rules and constraints for a given scenario follows. The scenario might be selected to be the generation of adataset of signals where collision of two (or more) different signals occurs. A collision can be defined as the superposition (in time) of an incumbent signal and an interferer signal. The incumbent signal is characterized by its e.g. center frequency, bandwidth, modulations, symbol rates, roll-off, filtering. The incumbent can be similarly characterized. A collision would occur then when the nominal bandwidth bounding boxes of each of the incumbent and interferer signals intersect. Referring now to Fig. 8, leverages multiple signal synthesizers to generate the waveforms of interest.

[0087] Referring now to Fig. 9, there is shown how synthesizer 901 would be configured to generate the incumbent signal 951 (shown in the time domain on the left and in the time and frequency domain on the right) and another synthesizer generates the interferer signal 953. The resulting spectrogram of the output signal of Fig. 9 is shown in Fig. 10 (a).

[0088] For clarity, Fig. 10 illustrates different channel conditions that can be for example individual scenarios for the purposes of synthetic waveform generation. Fig. 10(a) show a spectrogram of the interferer signal 953 colliding with the incumbent signal 951 partially or completely, while Fig. 10(b) shows adjacent channel bleedover from the interferer signal 953. Figure 10(c) shows wideband jamming, while Figure 10(d) shows narrowband or partial jamming. Figure 10(e) is a representation of low frequency modulated jamming, while Figure 10(f) shows what is known as “birdies”, or continuous narrowband spurs. While not strictly speaking impairments to a channel, Figure 10(g) shows a channel unencumbered from impairments, and Fig. 10(h) shows an incumbent transmitter at a low signal-to-noise ratio.

[0089] It will be appreciated that using signal synthesizers, i.e. computer- implemented signal generation and conditioning tools, to synthetically generate the plurality of waveforms within a computing device results in significant time savings. The actual waveforms do not need to be emitted over a physical channel (including air) and received by a receiving for eventual recording, and since the procedure generation executive module contains the rules and constraints, meta data accompanying each generated waveform is more easilygenerated concurrently with the generation of the waveform, and can easily be stored in dataset 280.

[0090] A person skilled in the art will appreciate that the synthesizer can be implemented using any one of a combination of existing methods. For example, there is a signal synthesizer in Python which is typically used for any arbitrary symbol map; random symbol bursts, constant tones, frames, messages and custom coding. Another implementation may use Gnu-Radio 620, which is often used for generic modulations; complete frames containing messages text files, audio, or video; and simple channel models. Another implementation can leverage MATLAB 630, which is typically adapted for rich list modulations; complete frames containing messages; and a rich library of channel model

[0091] It may be useful however to generate RF signals to be transmitted over channel. For example, this may provide a dataset that may include more random impairments. In such a case, and referring now to Fig. 11 , there is shown a transmitter 1110, a channel 1120, an interferer 1130, a receiver 1140, and user device 100. Transmitter 1110 is preferably a software defined radio (SDR), which is responsive to the procedure generation executive to impose non-channel related rules and constraints on the waveform. The waveform is then emitted on the channel 1120 and captured by receiver 1140. The captured waveforms (as there will be many iterations according to the rules and constraints, and specific scenario) are then stored to create dataset 280. To condition the channel 1120, an interferer 1130, controlled by procedure generation executive running on user device 100 to impose channel impairments.

[0092] FIG. 12 shows an example of assigning metadata 1000 to each subset or slice of a signal. Advantageously, metadata 1000 will be captured during generation of the signal according to desired information. In the example of Fig. 12, each subset includes an example waveform identifier 1020 and sublabels characterizing the waveform, the RMS amplitude of the signal and DC offset. Of course, more or fewer data points can be captured according to the requirements of the user. More specifically, in Fig. 10, there are m “examples”1010. The sublabels 1060 captured are based on measurements of signal, observations, or calculations, etc., which can be calculated or inherent to the given subset.

[0093] FIG. 13 shows an example of a recording 2100 produced by the systems and methods described herein, for a single waveform. A time series 2110 of the generated signal is shown, along with a spectrogram 2120 thereof. The spectrogram includes annotation-based metadata (in this case, occupied bandwidth, nominal bandwidth of two colliding signals and where they collide)..

[0094] While the applicant’s teachings described herein are in conjunction with various embodiments for illustrative purposes, it is not intended that the applicant’s teachings be limited to such embodiments as the embodiments described herein are intended to be examples. On the contrary, the applicant’s teachings described and illustrated herein encompass various alternatives, modifications, and equivalents, without departing from the embodiments described herein, the general scope of which is defined in the appended claims.

Claims

CLAIMS:1 . A method for producing synthetic radiofrequency recordings, the method comprising: determining a first set of waveform signal characteristics based at least in part on permutations of generated transmitters of a signal; determining a first set of rules based on permutations of possible configurations of a set of scenarios for the signal; determining a second set of rules based on environment variables for each scenario in the set of scenarios; generating individual signals based on one or more of the first set of waveform signal characteristics, the first set of rules, and the second set of rules; receiving a set of symbols that is a waveform representation of a bit string corresponding to the individual signals; creating a configuration for the individual signals; and generating and labeling a plurality of recordings based on at least one of the configurations for the individual signals or constraints to the configurations for use in training or testing machine learning models.

2. A method for producing a dataset of a plurality of individual synthetic radiofrequency recordings, comprising:(a) defining a scenario for said dataset;(b) defining a plurality of scenes for said scenario, where each scene is a simulation of a radiofrequency channel; where, for each scene,- a message is encoded into a waveform for transmission over a channel according to a set of predetermined rules;- at least one of a non-channel impairment and a channel impairment is imposed on said waveform to create a synthesized waveform; and- each synthesized waveform is recorded;wherein variations of said rules and said at least one of a non-channel impairment and a channel impairment are applied to individual waveforms to form a dataset of a plurality of synthesized waveforms.

3. A method for producing a dataset of a plurality of individual synthetic radiofrequency recordings, comprising:(a) defining a scenario for said dataset;(b) defining a plurality of scenes for said scenario, where each scene is a simulation of a radiofrequency channel;(c) defining a plurality of a first set of parameters associated with nonchannel components;(d) defining a plurality of a second set of parameters associated with a communication channel; where, for each scene,- a message is encoded into a waveform for transmission over a communication channel according to a set of predetermined rules;- at least one of a non-channel impairment and a channel impairment is imposed on said waveform to create a synthesized waveform; and- each synthesized waveform is recorded; wherein at least one parameter of the first set of parameters or at least one parameter or the second set of parameters is varied between each synthesized waveform.

4. The method of claim 3, wherein the non-channel components include transmitters, modulators, and signal sources.

5. The method of claim 3, wherein the communication channel parameters include channel bandwidth, noise levels, and interference patterns.

6. The method of claim 3, further comprising defining a set of rules for encoding the message into the waveform.

7. The method of claim 3, wherein the non-channel impairments include signal distortions and frequency offsets.

8. The method of claim 3, wherein the channel impairments include multipath fading and Doppler shifts.

9. The method of claim 3, further comprising generating metadata for each synthesized waveform, the metadata including information about the applied impairments and parameters.

10. The method of claim 3, wherein the synthesized waveforms are stored in a database for subsequent retrieval and analysis.

11. The method of claim 4, wherein the transmitters include software-defined radios (SDRs) and traditional RF transmitters.

12. A machine-readable medium storing instructions for automated synthetic radiofrequency scene dataset generation, the instructions, which when executed by a processor, cause the processor to perform operations comprising:- defining a scenario for a dataset, wherein the scenario includes a set of parameters describing a simulated radiofrequency channel under a theme;- defining a plurality of scenes for the scenario, where each scene is a simulation of the radiofrequency channel overtime, defined according to the scenario;- determining a first set of waveform signal characteristics based on permutations of generated transmitters of a signal;- determining a first set of rules based on permutations of possible configurations of the set of scenarios for the signal;- determining a second set of rules based on environment variables for each scenario in the set of scenarios;- generating individual signals based on one or more of the first set of waveform signal characteristics, the first set of rules, and the second set of rules;- applying at least one of a non-channel impairment and a channel impairment to the individual signals to create synthesized waveforms;- receiving a set of symbols that is a waveform representation of a bit string corresponding to the individual signals;- creating a configuration for the individual signals;- generating and labeling a plurality of recordings based on at least one of the configurations for the individual signals or constraints to the configurations; and- storing the synthesized waveforms in a database for subsequent retrieval and analysis.

13. A machine-readable medium according to claim 12, where said medium further includes generating metadata for each synthesized waveform, the metadata including information about the applied impairments and parameters.

14. An apparatus for automated synthetic radiofrequency scene dataset generation, comprising:- a user device configured to define scenarios and scenes for the dataset, wherein each scenario includes a set of parameters describing a simulated radiofrequency channel under a theme, and each scene is a simulation of the radiofrequency channel over time defined according to the scenario;- a server in communication with the user device, the server comprising:- a processor unit configured to execute instructions for generating synthetic radiofrequency recordings;- a memory unit storing program instructions, machine learning models, and a database for storing synthesized waveforms and associated metadata;- a graphical user interface (GUI) engine configured to generate user interfaces for defining scenarios and scenes, and receiving user inputs;- a signal synthesizer configured to generate individual signals based on permutations of generated transmitters of a signal, a first set of rules based on permutations of possible configurations of the set of scenarios, and a second set of rules based on environment variables for each scenario;- a procedure generation executive configured to:- apply at least one of a non-channel impairment and a channel impairment to the individual signals to create synthesized waveforms;- generate metadata for each synthesized waveform, the metadata including information about the applied impairments and parameters;- store the synthesized waveforms and their associated metadata in the database for subsequent retrieval and analysis.

15. The apparatus of claim 14, wherein the user device is a smartphone, a tablet computer, or a laptop.

16. The apparatus of claim 14, wherein the server is implemented using a cloud computing environment.

17. The apparatus of claim 14, wherein the server is implemented using a distributed computing environment.

18. The apparatus of claim 14, wherein the graphical user interface (GUI) engine is configured to present a set of predetermined rules and constraints for defining scenarios and scenes.

19. The apparatus of claim 18, wherein the graphical user interface (GUI) engine is configured to allow users to define custom rules and constraints for scenarios and scenes.