Methods and systems for generating radio frequency datasets

The method and system generate RF datasets by capturing, indexing, and labeling RF signal data, addressing the challenge of complex data preparation for machine learning, facilitating efficient training in telecommunications and defense applications.

WO2025171461A1PCT designated stage Publication Date: 2025-08-21QOHERENT INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/CA2024/050912
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-16
Filing Date
2024-07-05
Publication Date
2025-08-21

AI Technical Summary

Technical Problem

Existing methods struggle to gather and process RF signal data effectively for machine learning, requiring significant skill and knowledge, and the complexity of insights has increased, making it difficult to prepare data for model training.

Method used

A method and system for generating RF datasets by capturing recordings, assigning metadata, indexing, locating events, slicing examples, performing quality testing, and labeling them to create subsets for machine learning, using software-defined radio systems.

Benefits of technology

Facilitates the creation of curated RF datasets suitable for machine learning, enabling efficient training of models in applications like telecommunications and defense, reducing the complexity of data preparation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CA2024050912_21082025_PF_FP_ABST
    Figure CA2024050912_21082025_PF_FP_ABST
Patent Text Reader

Abstract

A method and system of creating a dataset for machine learning purposes is provided herein. The method comprises operating one or more processors to: capture at least one recording of an RF signal; capture global metadata associated with the RF signal; assign the global metadata to the corresponding least one recording of the RF signal; index the at least one recording; locate at least one event within the at least one recording; for each of the located events, slice each of the at least one recordings into at least one example; perform quality testing of each of the at least one examples by applying an acceptance criteria / threshold for each example; assign at least one label for each example; and group each of the at least one labeled examples into a subset of the dataset that shares at least one metadata set category.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND SYSTEMS FOR GENERATING RADIO FREQUENCY DATASETSCROSS-REFERENCE TO RELATED APPLICATION

[0001] The present application claims priority to U.S. Provisional Patent Application No. 63 / 554,505, which was filed February 16, 2024, the content of which is incorporated herein by reference in its entirety.FIELD

[0002] The described embodiments relate to systems and methods for generating training radio frequency (RF) signal datasets, for software-defined radio (SDR) applications.BACKGROUND

[0003] Radio frequency (RF) signals are used in a wide range of radio-enabled fields, for transmitting and receiving information and sensing and detecting objects. In many cases, analyzing RF signals can reveal valuable information. For example, in sensing and detecting applications, the analysis of RF signals can help detect objects and in communications applications, the analysis of RF signals can help gain insights on network utilization. Traditionally, these analyses were performed manually, by observing patterns and changes in the RF signals.

[0004] Machine learning algorithms and models have emerged to at least partially automate the analysis process. These models can allow insights to be obtained about collected RF signals. However, as the requirements of these models and the complexity of insights have increased, the skill and knowledge required to develop these models has also increased. In many cases, it can be difficult and / or impractical to gather the appropriate types of data required for model training. Furthermore, once the data is gathered, it can be difficult and / or impractical to prepare or process the data into a format that is suitable for model training or other machine learning tasks.

[0005] There is a need for improved methods and systems for generating and processing the generated data into training datasets for machine learning purposes.SUMMARY

[0006] The various embodiments described herein relate to a method of creating one or more datasets for machine learning purposes. In at least one embodiment, the method comprises operating one or more processors to: Capture at least one recording of an RF signal; Capture global metadata associated with the RF signal; Assign the global metadata to the corresponding least one recording of the RF signal; Index the at least one recording; Locate at least one event within the at least one recording; For each of the located events, slice each of the at least one recordings into at least one example; Perform quality testing of each of the at least one examples by applying an acceptance criteria / threshold for each example; Assign at least one label for each example; and Group each of the at least one labeled examples into a subset of the dataset that shares at least one metadata set category.

[0007] In at least one embodiment, capturing at least one recording of an RF signal comprises synthesizing at least one recording of the RF signal. In at least one embodiment, the least one examples can be of user-defined size, and / or smaller than the at least one recording.

[0008] In at least one embodiment, the method further comprises operating the one or more processors to: annotate the at least one recording, by exposing the metadata for all time and all frequency bands of the at least one recording. In at least one embodiment, annotating the at least one recording by exposing the metadata comprises: adding inferred metadata to each recording, wherein the inferred metadata is inferred after the recording has taken place. In at least one embodiment, capturing the global metadata associated with the RF signal the global metadata comprises capturing information related to all of the data that can be known about an RF signal. For instance, the metadata can include, but is not limited to the modulation, protocol, type of Radio, project name, scenario, description of scenario, and user intentions of the recording. In at least one embodiment, indexing the recordings comprises numbering the recordings.

[0009] In at least one embodiment, locating at least one event within the signals of the recordings comprises operating the one or more processor to determine, for each recording and identifying the events, by at least one of: image recognition, energy detection, RMS and threshold, RMS squared, absolute value, bounding box methods,measurement, and calculation of an attribute. In at least one embodiment, applying an acceptance criteria / threshold for each example comprises at least one of: SNR measurement, energy detection, and machine learning inference having positive result.

[0010] In at least one embodiment, the method further comprises augmenting at least one example from the examples to create new examples from the examples that were previously quality tested. In at least one embodiment, the labels include: user- defined label set, Wi-Fi transmission environment, measurement output, usage info, hi- throughput, lo-throughput; recording-level metadata; global metadata, inferred metadata, output by another model, and classification by another model. In at least one embodiment, grouping each of the at least one labeled examples into a subset of the dataset comprises sharing the at least one metadata category across each subset; and wherein every dataset has separate metadata.

[0011] The various embodiments described relate to a system of creating a dataset for machine learning purposes. In at least one embodiment, the system comprises one or more processors operable to: capture at least one recording of an RF signal; capture global metadata associated with the RF signal; assign the global metadata to the corresponding least one recording of the RF signal; index the at least one recording; locate at least one event within the at least one recording; for each of the located events, slice each of the at least one recordings into at least one example; perform quality testing of each of the at least one examples by applying an acceptance criteria / threshold for each example; assign at least one label for each example; and group each of the at least one labeled examples into a subset of the dataset that shares at least one metadata set category.

[0012] In at least one embodiment, the system further comprises a software defined radio to capture the at least one recording of the RF signal. In at least one embodiment, capturing at least one recording of an RF signal comprises synthesizing at least one recording of the RF signal. In at least one embodiment, the least one examples can be of user-defined size, and / or smaller than the at least one recording.

[0013] In at least one embodiment, the system further comprises the one or more processors operable to: annotate the at least one recording, by exposing the metadatafor all time and all frequency bands of the at least one recording. In at least one embodiment, annotating the at least one recording by exposing the metadata comprises: adding inferred metadata to each recording, wherein the inferred metadata is inferred after the recording has taken place. In at least one embodiment, capturing the global metadata associated with the RF signal the global metadata comprises capturing information related to at least one of: modulation, protocol, type of Radio, project name, scenario, description of scenario, and user intentions. In at least one embodiment, indexing the recordings comprises numbering the recordings.

[0014] In at least one embodiment, locating at least one event within the signals of the recordings comprises operating the one or more processor to determine, for each recording and identifying the events, by at least one of: image recognition, energy detection, RMS and threshold, RMS squared, absolute value, and bounding box methods. In at least one embodiment, applying an acceptance criteria / threshold for each example comprises at least one of: SNR measurement, energy detection, and machine learning inference having positive result. In at least one embodiment, the labels include: user-defined label set, Wi-Fi transmission environment, measurement output, usage info, hi-throughput, lo-throughput; recording-level metadata; global metadata, inferred metadata, output by another model, and classification by another model. In at least one embodiment, grouping each of the at least one labeled examples into a subset of the dataset comprises sharing the at least one metadata category across each subset; and wherein every dataset has separate metadata.BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Several embodiments will be described in detail with reference to the drawings, in which:FIG. 1A is a block diagram of an example RF dataset curation system in communication with external components, in accordance with an embodiment.FIG. 1 B is a block diagram showing example modules of the RF dataset curation system, in accordance with an embodiment.FIG. 1 C is a flowchart of an example method for generating an RF dataset, in accordance with an example embodiment.FIG. 2 is a block diagram of an example RF dataset curation system, in accordance with an embodiment.FIG. 3 is a block diagram showing examples of the dataset generation modules, in accordance with an embodiment.FIG. 4 is a block diagram of an example RF dataset curation system, in accordance with an embodiment.FIG. 5 is a block diagram of an example RF dataset curation system, in accordance with an embodiment.FIG. 6 is a screen capture showing an example step in the method of generating an RF dataset, in accordance with an embodiment.FIG. 7 is an example step in the method of generating an RF dataset, in accordance with an example embodiment.FIG. 8 is a block diagram showing an example step in the method of generating an RF dataset, in accordance with an embodiment.FIG. 9 is a block diagram showing an example step in the method of generating an RF dataset, in accordance with an embodiment.FIG. 10A is a block diagram showing an example step in the method of generating an RF dataset, in accordance with an embodiment.FIG. 10B is a block diagram showing an example step in the method of generating an RF dataset, in accordance with an embodiment.FIG. 11 is a block diagram showing an example step in the method of generating an RF dataset, in accordance with an embodiment.FIG. 12 is a block diagram of an example RF dataset curation system in communication with external components, in accordance with an embodiment.FIG. 13 is a block diagram of an example RF dataset curation system in communication with external components, in accordance with an embodiment.FIG. 1 is a screen capture showing an example step in the method of generating an RF dataset, in accordance with an embodiment.FIG. 15 is a screen capture showing an example step in the method of generating an RF dataset, in accordance with an embodiment.FIG. 16 is a screen capture showing an example step in the method of generating an RF dataset, in accordance with an embodiment.FIG. 17 is a screen capture showing an example step in the method of generating an RF dataset, in accordance with an embodiment.FIG. 18 is a screen capture showing an example step in the method of generating an RF dataset, in accordance with an embodiment.FIG. 19 is a screen capture showing an example step in the method of generating an RF dataset, in accordance with an embodiment.FIG. 20 is a screen capture showing an example step in the method of generating an RF dataset, in accordance with an embodiment.FIG. 21 is a screen capture showing an example step in the method of generating an RF dataset, in accordance with an embodiment.FIG. 22 is a screen capture showing an example step in the method of generating an RF dataset, in accordance with an embodiment.FIG. 23 is a screen capture showing an example step in the method of generating an RF dataset, in accordance with an embodiment.FIG. 24 is a screen capture showing an example step in the method of generating an RF dataset, in accordance with an embodiment.FIG. 25 is a screen capture showing an example step in the method of generating an RF dataset, in accordance with an embodiment.FIG. 26 is a screen capture showing an example step in the method of generating an RF dataset, in accordance with an embodiment.

[0016] The drawings, described below, are provided for purposes of illustration, and not of limitation, of the aspects and features of various examples of embodiments described herein. For simplicity and clarity of illustration, elements shown in the drawings have not necessarily been drawn to scale. The dimensions of some of the elements may be exaggerated relative to other elements for clarity. It will be appreciated that forsimplicity and clarity of illustration, where considered appropriate, reference numerals may be repeated among the drawings to indicate corresponding or analogous elements or steps.DESCRIPTION OF EXAMPLE EMBODIMENTS

[0017] In many applications, analyzing RF signals can provide valuable insights. For example, in sensing and detecting applications, the analysis of RF signals can help detect objects. The detection of objects through analysis of RF signals can be particularly useful in defense applications, or in automotive applications. As another example, in telecommunication applications, the analysis of RF signals can reveal that signals transmitted or received by a satellite are being impaired by weather conditions, or that there is interference due to the presence of another satellite.

[0018] These models can be highly complex and require significant amounts of data to train. However, it can be difficult and / or impractical to gather the appropriate types of data required for model training. Furthermore, once the data is gathered, it can be difficult and / or impractical to prepare or process the data into a format that is suitable for model training or other machine learning tasks.

[0019] The various embodiments described herein allow RF signal datasets to be generated. The datasets generated by the various embodiments described herein can be used to produce sample datasets that mimic the RF environment. The various embodiments of generated datasets described herein can be used by users to train machine learning models. The machine learning models can first complete a signal processing step as a part of a radio system. For example, the various radio system applications can include, but are not limited to conventional RF systems, including 5G, 4G LTE systems, systems using radars, AM / FM systems, consumer electronics, wireless networking, radio imaging, etc.

[0020] The described embodiments can be used in a wide variety of fields in which it is advantageous to analyze radio signals to obtain insights, including but not limited to, telecommunications, including satellite communications, automotive applications, anddefense applications. At least some of the embodiments can be implemented using cloud-based technologies.

[0021] Reference is first made to FIG.1 , which illustrates an example block diagram 100 of an RF dataset generation and curation system 108 (herein referred to as “dataset generation system 108”) in communication with an external data storage 102 and a computing device 106 via a network 104. Although only one computing device 106 is shown in FIG.1 , the dataset generation system 108 may be in communication with a greater number of computing devices 106. The dataset generation system 108 can communicate with the computing device(s) 106 over a wide geographic area via the network 104. While the dataset generation system 108 and the computing device 106 are shown as separate components, in some cases, the dataset generation system 108 or one or more components of the dataset generation system 108 may be implemented within the computing device 106. Alternatively, the dataset generation system 108 can be cloud-based.

[0022] In at least one embodiment, the dataset generation system 108 can be installed to a local computing device. The local computing device can optionally comprise a local Software-Defined-Radio (SDR). An SDR can include an antenna, a general RF tunable front-end component, an Analog to Digital converter, a processor operable to run a software, a communication link (such as USB, ethernet, Bluetooth, Wi-Fi, etc.), and an FPGA to control the firmware. Any SDR or radio computing device that has software to control the radio can be used. The SDR or radio computing device can be located on the host device. In one embodiment, the SDR or radio computing device can be operable to transmit / receive radio signals and should be tunable to operate at different frequencies within a supported range. The tunable front-end component allows the device to be dynamically configured to operate at different frequencies within its supported range. This allows SDRs to adapt to different communication standards and frequency bands without the need for hardware changes.

[0023] In at least one embodiment, the dataset generation system 108 can be installed to a cloud computing device. A local Software-Defined-Radio (SDR) can be coupled to the cloud computing device.

[0024] In at least one embodiment, the dataset generation system 108 can be installed to a local computing device. In this embodiment, the dataset generation system 108 can be connected to an SDR radio that can be controlled remotely, such as via a testbed.

[0025] In at least one embodiment, the dataset generation system 108 can be installed to a cloud computing device. The cloud computing device can be connected to a software application is running on the cloud connected to an SDR radio that can be controlled remotely, such as via a testbed.

[0026] In at least one embodiment, the dataset generation system 108 can comprise a synthesizing application operable to synthesize artificial RF data. The synthesizing application can be cloud-based or be located locally on the computing device.

[0027] In at least one embodiment, the dataset generation system 108 can be installed locally on a computing device and connected to an SDR that is managed by a cloud resource. In this embodiment, the SDR is not controlled by the user, but rather is controlled by the cloud resource.

[0028] The dataset generation system 108 includes a storage component 110, a processor 112, and a communication component 114. The dataset generation system 108 can be implemented with more than one computer server distributed over a wide geographic area and connected via the network 104. The storage component 110, the processor 112 and the communication component 114 may be combined into a fewer number of components or may be separated into further components.

[0029] The processor 112 can be implemented with any suitable processor, controller, digital signal processor, graphics processing unit, application specific integrated circuits (ASICs), and / or field programmable gate arrays (FPGAs) that can provide sufficient processing power for the configuration, purposes, and requirements of the dataset generation system 108. The processor 112 can include more than one processor with each processor being configured to perform different dedicated tasks.

[0030] The communication component 114 can include any interface that enables the dataset generation system 108 to communicate with various devices and othersystems. For example, the communication component 114 can receive inputs (e.g., data recordings, model targets) from the computing device 106 and store the inputs in the storage component 110 or external data storage 102. The processor 112 can then process the inputs according to the methods described herein.

[0031] The communication component 114 can include at least one of a serial port, a parallel port, or a USB port, in some embodiments. The communication component 114 may also include an interface to component via one or more of an Internet, Local Area Network (LAN), Ethernet, Firewire, modem, fiber, or digital subscriber line connection. Various combinations of these elements may be incorporated within the communication component 114. For example, the communication component 114 may receive input from various input devices, such as a mouse, a keyboard, a touch screen, a thumbwheel, a trackpad, a track-ball, a card-reader, voice recognition software and the like depending on the requirements and implementation of the system for generating inference models 108.

[0032] The storage component 110 can include RAM, ROM, one or more hard drives, one or more flash drives, or some other suitable data storage elements such as disk drives. The storage component 110 can include one or more databases for storing data such as, but not limited to related to the inference models including untrained models, the inference models generated by the dataset generation system 108, model tuning parameters and data for generating the inference models, data related to the validation of a trained inference model, including reports generated by the dataset generation system 108 and test results associated with the trained inference model, input data, including input training data, synthetic data generated and stored processed data and model targets, and packaged inference models.

[0033] The external data storage 102 can store data similar to that of the storage component 110. The external data storage 102 can, in some embodiments, be used to store data that is less frequently used and / or older data. In some embodiments, the external data storage 102 can be a third-party data storage stored with input data for analysis by the dataset generation system 108. The data stored in the external datastorage 102 can be retrieved by the computing device 106 and / or the dataset generation system 108 via the network 102.

[0034] The computing device 106 can include any device capable of communicating with other devices through a network such as the network 102. A network device can couple to the network 102 through a wired or wireless connection. The computing device 106 can include a processor and memory, and may be an electronic tablet device, a personal computer, workstation, server, portable computer, mobile device, personal digital assistant, laptop, smart phone, WAP phone, an interactive television, video display terminals, gaming consoles, and portable electronic devices or any combination of these. For example, the computing device 106 can be a user device used for obtaining a trained inference model generated by the dataset generation system 108.

[0035] The network 104 can include any network capable of carrying data, including the Internet, Ethernet, plain old telephone service (POTS) line, public switch telephone network (PSTN), integrated services digital network (ISDN), digital subscriber line (DSL), coaxial cable, fiber optics, satellite, mobile, wireless (e.g. Wi-Fi, WiMAX), SS7 signaling network, fixed line, local area network, wide area network, and others, including any combination of these, capable of interfacing with, and enabling communication between, the dataset generation system 108, the external data storage 102, and the computing device 106.

[0036] Reference is next made to FIG.1 B which show modules of the dataset generation system 108 in accordance with an embodiment. In at least one embodiment, the dataset generation system 108 includes a dataset generation module 120, a dataset curator module 122, a dataset testing module 124, and a model building module 126. Though FIG.1 B show four modules, the modules may be subdivided into sub-modules. Alternatively, some of the modules can be combined and the modules shown can be submodules.

[0037] FIG.1 C provides a flowchart of a method 130 of generating a dataset using one or more processors of the dataset generation system 108, in accordance with an embodiment. At step 132, the dataset generation system 108 captures at least onerecording of an RF signal. At step 134, the dataset generation system 108 captures global metadata associated with the RF signal. At step 136, the dataset generation system 108 assigns the global metadata to the corresponding least one recording of the RF signal. At step 138, the dataset generation system 108 indexes the at least one recording. At step 140, the dataset generation system 108 locates at least one event within the at least one recording. At step 142, the dataset generation system 108 slices, for each of the located events, each of the at least one recordings into at least one example. At step 144, the dataset generation system 108 performs quality testing of each of the at least one examples by applying an acceptance criteria / threshold for each example. At step 146, the dataset generation system 108 assigns at least one label for each example. At step 148, the dataset generation system 108 groups each of the at least one labeled examples into a subset of the dataset that shares at least one metadata set category.

[0038] The method described in FIG.1 C will be herein described in detail. Turning now to FIG.2, which provides a block diagram of an example RF dataset generation system 108, in accordance with an embodiment. The dataset generation system 108 can use synthetic or recorded RF signals 202 as inputs, and output large number of curated datasets 208. The datasets 208 can be used as training data for machine learning applications. The RF dataset generation system 108 can provide a means for taking many RF recordings 202 and producing a dataset, or a plurality of datasets 208. A dataset can refer to a set of sliced signal tapes, organized into labelled classes, and optionally packaged into a format that is compatible with machine learning or inference models.

[0039] FIG. 3 provides a block diagram showing examples of the dataset generation modules, in accordance with an embodiment. In an example embodiment, the RF signals 202 can be synthesized artificially 302, acquired by a testbed 306, and / or observed in the surroundings 304. The signals being captured may be any type of RF (radiofrequency) signals including but not limited to conventional RF systems, including 5G, 4G LTE systems, systems using radars, AM / FM systems, Bluetooth, Wi-Fi, etc.

[0040] In at least one embodiment, the synthesis system 302 can comprise a computing device and a software application for generating synthetic recordings can be used.

[0041] In at least one embodiment, the testbed system 306 can comprise a computing device or multiple computing devices, a software application running on the computing device, and at least one transmitting and / or receiving radio can be used. The testbed system can optionally comprise a complete communication system; and procedures of operation. The procedures of operation can provide the transmitting radio to send a signal and receiving radio to receive a signal. The testbed provides an emulation of a real-world RF environment.

[0042] In at least one embodiment, the observation system 304 comprises at least one computing device and at least one software defined radio to capture signals being transmitted in the real environment. The observation system 304 can optionally comprise a recording device 308.

[0043] FIG. 4 is a block diagram of an example RF dataset curation system, in accordance with an embodiment. At 402, the recordings 202 can be captured. In at least one embodiment, metadata for each of the recordings can be captured and / or saved simultaneously at the time of capturing the RF recordings 202. The metadata can also be assigned to each of the corresponding RF recordings 202. At 402, the recordings 202 can further be indexed. In at least one embodiment, indexing the recordings comprises numbering the recordings. At 404, the indexed recording can be masked, labeled and / or annotated. The masking, labeling and / or annotating steps can be completed for all time and all frequency bands, for all of the recordings. In at least one embodiment, the annotating step can comprise exposing the metadata.

[0044] The dataset generation system locates at least one event within the recordings. At 406, of FIG.4, the annotated RF recordings can be quantified and sliced. Slicing refers to separating the RF recording into smaller recordings around each of the located events. Each of the recordings having events can be sliced into an example. At 408, the dataset generation system 108 groups and packages the examples 410 and performs quality testing on each of the examples 410 by applying an acceptancecriteria / threshold for each example. At 408, the dataset generation system 108 also assigns a label for each example. The system can further comprise an optional step at 412 and 414 to augment, diversify or edit the examples 410 to produce an augmented or diversified dataset. At 416, the dataset generation system 108 groups each of the labeled examples into a subset of the dataset that shares at least one metadata set category to produce a dataset 416. FIG. 5 provides an a block diagram 500 of an example RF dataset curation system, in accordance with an embodiment.

[0045] Turning now to FIGs. 6 to 13, which provide example workflows for the method of generating an RF dataset. FIG. 6 is a screen capture showing the steps of indexing, selecting, and reviewing the RF recordings generated as shown in FIG. 3. In at least one embodiment, capturing at least one recording of an RF signal can comprise capturing at least one spectrogram plot 602. Capturing at least one recording of an RF signal can comprise synthesizing the at least one recording of the RF signal. The spectrogram plots can be captured and subsequently transferred to a NumPy library.

[0046] FIGs. 7 and 8 provides example embodiments showing the steps of masking, labeling, and annotating the RF recordings with metadata. In at least one embodiment, two types of metadata can be defined: global metadata and inferred metadata. Global metadata can be used to refer to metadata that is captured at the time of capturing the RF recording. Examples of global metadata can include modulation, protocol, type of Radio, project name, scenario, description of scenario, user intentions (i.e., 5G). Global metadata can correspond to the RF recording and can be assigned to each of the recorded RF signals as defined in step 136 of FIG.1C. Inferred metadata refers to the metadata which can be inferred after the recording has taken place, such as measurements or information that can be calculated after the recording has taken place. Each recording can optionally be annotated with the inferred metadata manually or automatically. The inferred metadata can be user defined, and optionally automated with machine learning techniques, or analytically-derived signal processing techniques.

[0047] In FIG. 7, a method 700 of masking the data using bounding boxes is shown. In this method, energy peaks can be detected in the RF recordings, and bounding boxes can be drawn around each of these energy peaks. The energy peaks may bepoints of interest, or “events” in the RF recordings, and therefore be helpful to curate for the training dataset. Therefore, the system can locate at least one event within the at least one recording. In at least one embodiment, locating at least one event within the signals of the recordings comprises operating the one or more processor to determine, for each recording and identifying the events, by at least one of: image recognition methods, energy detection, RMS and threshold, RMS squared, absolute value, and bounding box methods.

[0048] In FIG. 8, the steps 800 of labeling and annotating the RF recording is shown. In at least one embodiment, a user can manually apply labels and annotations to bounding boxes or automatically to config-specified time / frequency zones. In another embodiment, a specialized deep learning model or set of models can be used to recognize and label each bounding box. In yet another embodiment, if the recording device is a comms receiver or demodulating, the frame parameters can be used to apply labels. In at least one embodiment, if the software of the radio is controlled, the operating parameters can be captured from the radio and then be applied as labels.

[0049] FIG. 9 is a block diagram 900 showing the optional step of channel model application in the method of generating an RF dataset, in accordance with an embodiment. In at least one embodiment, the method of generating an RF dataset comprises applying a channel model. Channel models can include, but are not limited to time shift, frequency shift, Rayleigh fading channel, IQ imbalance, random resample, AGWN, pulse shaping, multipath fading, impulsive noise, RX filter, CCIR fading, ITU fading, extended fading etc.

[0050] FIG. 10A provides a block diagram 1000A showing the step of quantification and slicing in the method of generating an RF dataset, in accordance with an embodiment. The method includes the step of perform quality testing of each of the at least one examples by applying an acceptance criteria / threshold for each example. In at least one embodiment, applying an acceptance criteria / threshold for each example comprises at least one of: SNR measurement, energy detection, and machine learning inference having positive result. Strategies for removing examples that are only, or are mostly noise, or not qualified for the class, then slicing the examples into appropriatelength pieces for model training. Methods of qualification include: Simple RMS & Threshold based techniques, (e.g. determining if the RMS of the slice over a global threshold for the recording; quantization then threshold; rounding / reducing significant figures, and including slices above / below the threshold); Energy detection-based (e.g. mask based techniques) by generating a PSD for a grouping of slices and use a threshold for the group to decide if to include all or none; Deep learning qualification or redundancy checking techniques by using a deep learning model to determine if each example is relevant, (e.g., cleanlabAI); Metadata / Annotation based methods (e.g., determining if there a bounding box for this slice); Measurement based (e.g., SNR or EVM of a slice); and Intersection based (e.g., determining if bounding boxes intersect).

[0051] The method further comprises the step of assigning at least one label for each example; and grouping each of the labeled examples into a subset of the dataset that shares at least one metadata set category. In at least one embodiment, the labels include: a user-defined label set, Wi-Fi transmission environment, measurement output, usage info, hi-throughput, lo-throughput; recording-level metadata; global metadata, inferred metadata, output by another model, and classification by another model. In at least one embodiment, grouping each of the at least one labeled examples into a subset of the dataset comprises sharing the at least one metadata category across each subset; and wherein every dataset has separate metadata. The dataset generation system 108 can generates a dataset file based on the groupings of unique metadata, for example, as specified by the user.

[0052] The method may optionally comprise the step of augmentation and editing, shown by the augmentation module 1000B. FIG. 10B provides a block diagram showing the augmentation on an image according to an example. Augmentation can be used to create new examples from previously created examples. The image of the apple as shown by 1001 , can be augmented into different images 1002 to 1006 by de-texturizing, de-colorizing, edge enhancing, edge mapping, flipping, rotating, and any other method of augmentation. The new images 1002 to 1006 can then also be used as training data as part of the dataset, giving the model further data to train from.

[0053] The new examples may be qualified meaning that the already pass the quality testing step described in FIG.10A. In at least one embodiment, augmentation can be done by l / Q switching, switching samples in the recording, IQ imbalance, dropping samples, and any other methods of creating new datasets from existing datasets.

[0054] In at least one embodiment, it may be beneficial to diversify the examples to give the model new data to train from, i.e. , providing the model with a different view of existing data. Not diversifying can be disadvantageous as there may be a higher risk of over-fitting the data, and the convergence of the model may be slower. In at least one embodiment, the examples can be modified to add noise. In at least one embodiment, the examples can be diversified by decimation, interpolation, FIR filters, changing center frequency, rotating in-phase, or any other suitable technique. For example, if a model is trained solely on one class of data, and the model has not been provided with enough examples of this class of data, it may not generalize to other similar examples in same class. All of the new examples created should have the same metadata (same class) as the examples that the model used for training.

[0055] In at least one embodiment, augmentation and editing can include at least one of the following: Augment and Fill (i.e., creates randomly augmented examples up to a specified number of examples); Subsampling (i.e., reduces class sizes by a % by randomly dropping examples); homogenization (i.e., Different classes can have a different numbers of examples, and this can vary significantly. Homogenization subsamples excessive class sizes to a user specified length); Drop classes; Join two datasets; Append classes; Append to dataset; Trim examples (i.e., Reduce m of all examples in a given dataset); and Normalization (i.e., normalize all examples in a dataset). FIG. 11 provides a block diagram 1100 showing the homogenization method of homogenizing excess samples to a specified length.Augmentation Methods

[0056] Other methods of augmentation or diversification include: time reversal, Spectral inversion; Channel swap; Amplitude reversal; CutOut; Drop samples; Quantize; Magnitude rescale; PatchShuffle; Identity; Signal Roll-Off; Local Oscillator Drift; Time-Varying Noise; Clip; Add Slope; Random Convolve; Gain Drift; Automatic Gain Control; Mixllp; CutMix; and Expert Feature Transforms. These methods are explained by Boegner et. al. in (Boegner et. al. (2022). Large Scale Radio Frequency Signal Classification. 10.48550 / arXiv.2207.09918.).

[0057] The time reversal augmentation reverses the order of the IQ samples in the input. Since time reversal in the signal domain also results in a spectral inversion, the time reversal augmentation additionally has the option to undo the spectral inversion effect if desired.

[0058] The spectral inversion augmentation inverts the frequency components of the input data by negating the imaginary components of the input.

[0059] The channel swap augmentation switches the real and imaginary components of the input complex data. In the signal domain, this has the same effect as a spectral inversion followed by a static TT / 2 phase shift.

[0060] Amplitude reversal augments the input data by simply multiplying by -1 .

[0061] The drop samples augmentation randomly drops IQ samples from the input data using randomized values for the drop rate, the size of each dropped region, and the fill methods for how to replace the regions with dropped samples. The fill methods can be statically or randomly set to choose the following methods: front fill, back fill, mean, or zero; where front fill replaces each drop regions’ samples with the last previous valid value, back fill replaces each drop regions’ samples with the next valid value, mean replaces each drop regions’ samples with the mean value of the full data example, and zero replaces each drop regions’ samples with zeros.

[0062] The quantize augmentation allows data to be quantized to randomly selected numbers of levels with the quantize data transform, loosely emulating the bitdepth in an analog-to-digital converter (ADC) seen in digital RF systems. The quantization transform also allows for various rounding types between: flooring the observed values to the next-lowest valid quantized value, setting every value in a region to the middle value of the region, or rounding each value to the next largest valid quantized value.

[0063] The magnitude rescaling transform randomly selects a starting point of the input example to rescale the magnitude of the data by multiplying by a random constant. This behavior emulates an RF front end gain adjustment.

[0064] The CutOut transform inputs randomized cut durations and cut types to select how large a region in time should be cut out. The cut out region is then filled with either zeros, ones, low-SNR noise, average-SNR noise, or high-SNR noise.

[0065] The PatchShuffle transform operates solely on in the time domain, randomly shuffling multiple local regions of IQ samples, using a randomized patch size input distribution and a randomized shuffle ratio to discern how many of the patches should undergo random local shuffling.

[0066] The signal roll-off transform applies a lower, upper, or both-sided band-edge RF roll-off effect, simulating front end filtering.

[0067] The local oscillator (LO) drift transform emulates the imperfections of a receiver’s LO. This transform models the LO drift by implementing a random walk in frequency with a drift rate and a max drift set as inputs, where when the max drift is reached, the frequency offset is reset to 0.

[0068] The time-varying transform adds AWGN within a specified low to high SNR range with a specified number of inflection points at which point(s) the slope of the timevarying noise reverses direction.

[0069] The clip transform inputs a percentage that it uses to calculate the max and min values allowable through the clipping transform, setting all values above and below these values to the max and min, respectively.

[0070] The add slope transform computes the slope between every IQ sample in the input with its preceding sample and adds the slope to its current IQ sample. This transform has the effect of amplifying higher frequency components more than the lower frequency components.

[0071] The random convolve transform inputs a random distribution of the number of taps in a filter, where each tap is assigned random values in range 0 to 1 . The randomly generated filter is then convolved with the input IQ data. An alpha value is also input to the transform and it is used to dampen the effect of the randomly filtered data byweighting the newly filtered data and inversely weighting the original data and then summing the results. This random convolution is a relatively cheap form of applying a frequency-selective fading model.

[0072] A gain drift transform is also implemented to complement the LO drift transform’s frequency effects with magnitude effects. The gain drift transform inputs similar max / min drift values and a drift rate, which are used in a random walk of adjusting the magnitudes of the input data IQ samples over time.

[0073] An automatic gain control (AGC) implementation transforms data with an input scaling of default values for randomization across augmentation calls. The default values under random scaling can also be updated with arguments including: an initial gain value, an alpha for averaging the measure signal level, an alpha amount by which to adjust gain when in the tracking state, an alpha value by which to adjust gain when in the overflow state, an alpha value by which to adjust gain when in the acquire state, a reference level specifying the level of intended gain adjustment, a tracking range of allowable deviation before going into the acquire state, a low level which specifies when the AGC is disabled, and a high level which specifies when the AGC enters the overflow state.

[0074] The Mixllp transform inputs a dataset from which to randomly sample the secondary signal to be mixed with the original signal. The transform also inputs an alpha value specifying the logarithmic difference in SNR levels between the two signals. Additionally, since the secondary signal may not be of the same class as the original signal, the class label information is updated to include the added signal’s metadata.

[0075] The CutMix transform implements a modified version of the computer vision domain’s CutMix. Like the MixUp transform, our CutMix transform inputs a dataset to randomly sample the secondary signal from and insert into the original signal, replacing a random region of the signal with the new signal. An additional alpha value is input specifying the relative durations in time to occupy between the original and the newly inserted signal.

[0076] The Expert Feature Transforms can be used for representing the complex IQ signal data in multiple different representations, such as through interleaving the realand imaginary floating point values, converting the complex IQ samples to two channels of real and imaginary parts, computing the complex magnitude, computing the wrapped phase, transforming to the discrete Fourier representation, transforming to variations of a spectrogram, and wavelet transforms.***

[0077] FIGS. 12A and 12B provide a block diagram of an example RF dataset curation system in communication with external components, in accordance with an embodiment. In at least one embodiment 1200A and as shown in FIG. 12A, the system 108 can be paired with a software-defined-radio (SDR) to complete capture, and curation in a complete workflow.

[0078] In at least one embodiment 1200B and as shown in FIG. 12B, the system 108 can be connected with a software-defined-radio (SDR) via the internet or connected via token to complete capture, and curation in a complete workflow.

[0079] In at least one embodiment, a user can use a local software that connects to the system 108 in a cloud-based network for a seamless experience. FIG. 13 is a block diagram 1300 of an example RF dataset curation system in communication with external components, in accordance with an embodiment. In at least one embodiment, an automation application which manages execution of many recording captures, known as a recording executive, can be used to capture a set of frequencies, sample rates, antennas, configs - producing a large number of recordings to curate a dataset that can be used as training data for machine learning purposes.

[0080] FIGS. 14 to 26 provide screen captures showing a step-by-step example method of generating an RF dataset, in accordance with an embodiment. FIGS. 1 A and 14B provide examples of the capture settings 1400A, 1400B used from a software defined radio. In this example, the capture setting includes the IP address, the number of samples, the center frequency, the sampling rate, the gain, and the channel. The metadata being captured simultaneously when the RF recording is being done includes the protocol (i.e., Wi-Fi, AM, FM, Bluetooth), the use case (in this example, the use case is an ambient environment, broadcast), the testbed (in this example, the testbed is a portable blade), the transmitter type (in this example, ambient transmitters), the name ofthe project, and which software defined radio is being used. Once a user fills out this information, the SDR can begin recording the RF signals, as well as the global metadata.

[0081] FIG. 15 provides an example 1500 of assigning the metadata to each of the recordings of the RF signal. In at least one embodiment, two types of metadata can be defined: global metadata and inferred metadata. Global metadata can be used to refer to metadata that is captured at the time of capturing the RF recording. Examples of global metadata can include modulation, protocol, type of Radio, project name, scenario, description of scenario, user intentions (i.e. , 5G). Metadata can be assigned based on measurements of signal, observations, or calculations. Metadata can be calculated or inherent to the RF recording. Global metadata can correspond to the RF recording and can be assigned to each of the recorded RF signals. Metadata can be assigned by labeling each of the plurality of recordings with labels comprising information that may be important for a user for classification. The metadata can be assigned as a plurality of sub-labels.

[0082] Once the recordings are complete and the metadata is assigned, the recordings can be inspected, as shown in FIG. 16. The recording data includes information about the SDR capture settings 1604, the metadata 1606, an image of the RF signal spectrogram 1610, and a summary of the recording 1602. The RF signal spectrogram data 1610 can be a plot showing the recorded frequency (in Hertz) versus the time that the signal was recorded (in milliseconds). The RF signal spectrogram data 1610 can further include a plot showing the IQ versus the time that the signal was recorded (in milliseconds), and overlayed or aligned with the Frequency / Time plot. The spectrogram data may be provided the frequency or time domain. In at least one embodiment, the recorded RF signals can be provided as CSV files, MATLAB files, image files, Python files, SIGMF (signal metadata format) etc. Inspecting the recordings can comprise, in at least one embodiment, splicing or trimming the spectrogram into areas of potential interest, defined by the cut points 1608. For example, in the image of the spectrogram 1610, can be cut into the regions where the frequency shows events of interest, at 1608. This can be done manually by a user selecting the regions of interest or using image processing techniques to detect peaks, or by automating software toidentify and determine the peaks. In one embodiment, the user can draw or select the portions of interest with sliders showing on the user interface. In at least one embodiment, locating at least one event within the signals of the recordings comprises operating the one or more processor to determine, for each recording and identifying the events, by at least one of: image recognition, energy detection, RMS and threshold, RMS squared, absolute value, and bounding box methods. In at least one embodiment, the dataset generation system 108 can locate at least one event within the at least one recording; and for each of the located events, slice each of the at least one recordings into at least one example. FIG. 17 show an example 1700 of an RF recording in which no events of interest were detected.

[0083] Once the events of interest have been identified and the recordings have been sliced into a plurality of examples, FIGS. 18 to 22 provide examples 1800, 1900, 2000, 2100 2200 of the RF recordings which were cut into a plurality of examples. Each of the example recordings provided in FIGs. 18 to 22 can have at least one event therein. Quality testing can then be performed on each of the at least one examples by applying an acceptance criteria / threshold for each example. In at least one embodiment, applying an acceptance criteria / threshold for each example comprises at least one of: SNR measurement, energy detection, and machine learning inference having positive result.

[0084] FIG. 23 provides an embodiment 2300 showing the step of assigning at least one label for each of the accepted examples. In at least one embodiment, the labels include: a user-defined label set, Wi-Fi transmission environment, measurement output, usage info, hi-throughput, lo-throughput; recording-level metadata; global metadata, inferred metadata, output by another model, and classification by another model.

[0085] FIG. 24 provides the resulting dataset statistics 2400 and configuration 2402, according to an embodiment. FIGs. 25 and 26 provide an example of the results 2500, 2600 of the dataset curation method. The plurality of accepted examples obtained can be grouped by each of the at least one labeled examples into a subset of the dataset that shares at least one metadata set category. Some shared metadata categories can include: “FM” or “broadcast” which will prompt the processor to provide a curated dataset comprising a plurality of examples of RF signal recordings having FM or Broadcastingevents. As such, the examples can be grouped into any number of classes sharing a common metadata label or sublabel.

[0086] The curated dataset can then be sent as input training data for a machine learning module. The input training data can include data collected from radio devices. For example, the input training data can include signal data (e.g., IQ samples). As another example, the input training data can include spectrogram data. In some embodiments, the input training data includes labeled training data.

[0087] The input training data can be received for example, from the computing device 106 via the network 104. Alternatively, the dataset generation system 108 can retrieve at least a portion of the input training data from the external data storage 102 or the storage component 110. For example, the input training data can include data previously used for generating an inference model.

[0088] In at least one embodiment, the dataset generation system 108 can also process the input training data received at 310 via the dataset managing module 210 to obtain a processed training dataset. The processed training dataset can be a dataset that can be used as training data for a machine-learning model or as an input to a machine-learning model. In some embodiments, processing the input training data involves generating synthetic data to augment the input training data. In some embodiments, the dataset generation system 108 can receive synthetic data parameters, for example, from the computing device 106 defining a range of parameters for the synthetic data. The synthetic data parameters can for example, be specified by a user. In such embodiments, the dataset generation system 108 can generate synthetic data according to the received synthetic data parameters. In some embodiments, the synthetic data includes modulated versions of the input data and / or modified versions of the input training data modified according to models (e.g., impairment, time shifting, frequency shifting, random resampling, pulse shaping, multipath fading, impulsive noise and other noise models, other fading models, IQ imbalance).

[0089] In some embodiments, processing the input training data involves segmenting the input training data and / or the augmented training dataset and labelling the segmented input training data and / or the augmented training dataset to constructclasses. In some embodiments, at least some of the classes can be combined, based on predetermined variations. For example, the predetermined variations can be received from the computing device 106.

[0090] In some embodiments, processing the input training data involves combining the input training data with one or more existing training datasets. The existing training datasets can be retrieved from memory, for example, from the storage component 110 or the external data storage 102. The existing training datasets can include processed training datasets previously generated. In some embodiments, processing the input training data involves selecting a subset of the input training data to be used for generating an inference model. The subset can be selected based on a result of one or more qualification tests. Selecting a subset of the input training data can allow the dataset generation system 108 to select data meeting quality requirements. For example, the qualification test(s) can include root mean squared (RMS) based qualification test(s), quantization test(s), energy detection test(s), model-based test(s) (including machine learning model-based test(s)) and any combination thereof. In some embodiments, only data meeting certain qualification test(s) may be retained for training the inference model.

[0091] In some embodiments, the dataset generation system 108 receives a selection of the subset. For example, the dataset generation system 108 can receive user selections from the computing device 106.

[0092] The subset of the input training data can be further processed. For example, the subset of the input training data can be further augmented, as explained above. As another example, the subset can be further segmented according to classification labels. Other processing tasks can be performed on the subset, including but not limited to, reducing the size of the subset, removing classes from the subset, homogenizing classes, modifying the data in the subset according to various models.

[0093] The processed training dataset can be stored, for example, in the external data storage 102 or the storage component 110 at one or more stages of the processing. For example, the augmented training data can be stored, and the subset of the inputtraining data can be stored separately. In some embodiments, the processed training dataset is saved for future retrieval.

[0094] It will be appreciated that numerous specific details are set forth in order to provide a thorough understanding of the example embodiments described herein. However, it will be understood by those of ordinary skill in the art that the embodiments described herein may be practiced without these specific details. In other instances, well- known methods, procedures, and components have not been described in detail so as not to obscure the embodiments described herein. Furthermore, this description and the drawings are not to be considered as limiting the scope of the embodiments described herein in any way, but rather as merely describing the implementation of the various embodiments described herein.

[0095] The embodiments of the systems and methods described herein may be implemented in hardware or software, ora combination of both. These embodiments may be implemented in computer programs executing on programmable computers, each computer including at least one processor, a data storage system (including volatile memory or non-volatile memory or other data storage elements or a combination thereof), and at least one communication interface. For example, and without limitation, the programmable computers (referred to below as computing devices) may be a server, network appliance, embedded device, computer expansion module, a personal computer, laptop, personal data assistant, cellular telephone, smart-phone device, tablet computer, a wireless device, or any other computing device capable of being configured to carry out the methods described herein.

[0096] In some embodiments, the communication interface may be a network communication interface. In embodiments in which elements are combined, the communication interface may be a software communication interface, such as those for inter-process communication (IPC). In still other embodiments, there may be a combination of communication interfaces implemented as hardware, software, and combination thereof.

[0097] Program code may be applied to input data to perform the functions described herein and to generate output information. The output information is applied to one or more output devices, in known fashion.

[0098] Each program may be implemented in a high-level procedural or object- oriented programming and / or scripting language, or both, to communicate with a computer system. However, the programs may be implemented in assembly or machine language, if desired. In any case, the language may be a compiled or interpreted language. Each such computer program may be stored on a storage media or a device (e.g., ROM, magnetic disk, optical disc) readable by a general or special purpose programmable computer, for configuring and operating the computer when the storage media or device is read by the computer to perform the procedures described herein. Embodiments of the system may also be considered to be implemented as a non- transitory computer-readable storage medium, configured with a computer program, where the storage medium so configured causes a computer to operate in a specific and predefined manner to perform the functions described herein.

[0099] Furthermore, the system, processes and methods of the described embodiments are capable of being distributed in a computer program product comprising a computer readable medium that bears computer usable instructions for one or more processors. The medium may be provided in various forms, including one or more diskettes, compact disks, tapes, chips, wireline transmissions, satellite transmissions, internet transmission or downloading, magnetic and electronic storage media, digital and analog signals, and the like. The computer useable instructions may also be in various forms, including compiled and non-compiled code.

[0100] Various embodiments have been described herein by way of example only. Various modification and variations may be made to these example embodiments without departing from the spirit and scope of the invention, which is limited only by the appended claims.

Claims

WE CLAIM:1 . A method of creating a dataset for machine learning purposes, the method comprising operating one or more processors to:Capture at least one recording of an RF signal;Capture global metadata associated with the RF signal;Assign the global metadata to the corresponding least one recording of the RF signal;Index the at least one recording;Locate at least one event within the at least one recording;For each of the located events, slice each of the at least one recordings into at least one example;Perform quality testing of each of the at least one examples by applying an acceptance criteria / threshold for each example;Assign at least one label for each example; andGroup each of the at least one labeled examples into a subset of the dataset that shares at least one metadata set category.

2. The method of claim 1 , wherein capturing at least one recording of an RF signal comprises synthesizing at least one recording of the RF signal.

3. The method of claim 1 , wherein the least one examples can be of user-defined size, and / or smaller than the at least one recording.

4. The method of claim 1 , further comprising annotating the at least one recording, by exposing the metadata for all time and all frequency bands of the at least one recording.

5. The method of claim 2, wherein annotating the at least one recording by exposing the metadata comprises: adding inferred metadata to each recording, wherein the inferred metadata is inferred after the recording has taken place.

6. The method of claim 1 , wherein capturing the global metadata associated with the RF signal the global metadata comprises capturing information related to at least one of: modulation, protocol, type of Radio, project name, scenario, description of scenario, and user intentions.

7. The method of claim 2, further comprising augmenting at least one example from the examples to create new examples from the examples that were quality tested.

8. The method of claim 2 or 3, wherein locating at least one event within the signals of the recordings comprises operating the one or more processor to determine, for each recording and identifying the events, by at least one of: image recognition, energy detection, RMS and threshold, RMS squared, absolute value, and bounding box methods.

9. The method of any one of claims 1 to 4, wherein applying an acceptance criteria / threshold for each example comprises at least one of: SNR measurement, energy detection, and machine learning inference having positive result.

10. The method of any one of claims 1 to 6, wherein the labels include: user-defined label set, Wi-Fi transmission environment, measurement output, usage info, hi- throughput, lo-throughput; recording-level metadata; global metadata, inferred metadata, output by another model, and classification by another model.11 . The method of claim 7, wherein grouping each of the at least one labeled examples into a subset of the dataset comprises sharing the at least one metadata category across each subset; and wherein every dataset has separate metadata.

12. A system of creating a dataset for machine learning purposes, the system comprising one or more processors operable to:Capture at least one recording of an RF signal;Capture global metadata associated with the RF signal;Assign the global metadata to the corresponding least one recording of the RF signal;Index the at least one recording;Locate at least one event within the at least one recording;For each of the located events, slice each of the at least one recordings into at least one example;Perform quality testing of each of the at least one examples by applying an acceptance criteria / threshold for each example;Assign at least one label for each example; andGroup each of the at least one labeled examples into a subset of the dataset that shares at least one metadata set category.

13. The system of claim 12, further comprising a software defined radio to capture the at least one recording of the RF signal.

14. The system of claim 12, wherein the least one examples can be of user-defined size, and / or smaller than the at least one recording.

15. The system of claim 12, further comprising annotating the at least one recording, by exposing the metadata for all time and all frequency bands of the at least one recording.

16. The system of claim 12, wherein annotating the at least one recording by exposing the metadata comprises: adding inferred metadata to each recording, wherein the inferred metadata is inferred after the recording has taken place.

17. The system of claim 12, wherein capturing the global metadata associated with the RF signal the global metadata comprises capturing information related to at least one of: modulation, protocol, type of Radio, project name, scenario, description of scenario, and user intentions.

18. The system of claim 12 or 13, wherein locating at least one event within the signals of the recordings comprises operating the one or more processor to determine, for each recording and identifying the events, by at least one of: image recognition, energy detection, RMS and threshold, RMS squared, absolute value, and bounding box methods.

19. The system of any one of claims 12 to 18, wherein applying an acceptance criteria / threshold for each example comprises at least one of: SNR measurement, energy detection, and machine learning inference having positive result.

20. The system of any one of claims 12 to 19, wherein the labels include: user- defined label set, Wi-Fi transmission environment, measurement output, usage info, hi- throughput, lo-throughput; recording-level metadata; global metadata, inferred metadata, output by another model, and classification by another model.21 . The system of claim 12, wherein grouping each of the at least one labeled examples into a subset of the dataset comprises sharing the at least one metadata category across each subset; and wherein every dataset has separate metadata.

Citation Information

Patent Citations

  • Radio frequency band segmentation, signal detection and labelling using machine learning

    US20200143279A1

  • Machine learning based tuning of radio frequency apparatuses

    WO2024030236A1