Audio signal generation method and system, non-transitory computer-readable medium

By combining bone conduction and air conduction sensors and processing audio data using frequency thresholds and weights, the problem of insufficient signal fidelity in public microphones has been solved, generating higher fidelity audio data.

CN114822565BActive Publication Date: 2026-04-17SHENZHEN SHOKZ CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN SHOKZ CO LTD
Filing Date
2019-09-12
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

When using electronic devices for communication in public places, the voice signal captured by the microphone is easily affected by background noise, resulting in insufficient signal fidelity.

Method used

By combining bone conduction and air conduction sensors to collect audio data, the data is divided into segments using frequency thresholds, and then spliced, fused, and combined based on weights to generate third audio data, thereby improving signal fidelity.

Benefits of technology

The generated audio data has higher fidelity, reduced noise, and improved the intelligibility of the speech signal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114822565B_ABST
    Figure CN114822565B_ABST
Patent Text Reader

Abstract

This application relates to an audio signal generation method and system, and a non-transitory computer-readable medium. The audio signal generation method includes: acquiring first audio data collected by a bone conduction sensor; acquiring second audio data collected by an air conduction sensor, wherein the first audio data and the second audio data represent a user's speech, and the first audio data and the second audio data are composed of different frequency components; dividing the first audio data and the second audio data into multiple segments according to one or more frequency thresholds, wherein each segment of the first audio data corresponds to a segment of the second audio data; and splicing, fusing, and / or combining each segment of the multiple segments of the first audio data and the second audio data based on weights to generate third audio data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese patent application filed on September 12, 2019, with application number CN201910864002.8 and the invention title "System and Method for Generating Audio Signals". Technical Field

[0002] This application generally relates to the field of signal processing, and more specifically, to methods and systems for generating audio signals and non-transitory computer-readable media. Background Technology

[0003] With the widespread use of electronic devices, communication between people has become increasingly convenient. When communicating using electronic devices, users can rely on microphones to capture speech signals as they speak. The speech signal captured by the microphone can represent the user's voice. However, due to factors such as the microphone's own performance and noise, it is sometimes difficult to ensure that the speech signal captured by the microphone is fully intelligible (i.e., signal fidelity). Especially in public places such as factories, cars, airplanes, ships, and shopping malls, varying background noise significantly affects communication quality. Therefore, it is desirable to provide systems and methods for generating audio signals with less noise and / or improved fidelity. Summary of the Invention

[0004] This application provides an audio signal generation method, comprising: acquiring first audio data collected by a bone conduction sensor; acquiring second audio data collected by an air conduction sensor, wherein the first audio data and the second audio data represent a user's speech, and the first audio data and the second audio data are respectively composed of different frequency components; dividing the first audio data and the second audio data into multiple segments according to one or more frequency thresholds, wherein each segment of the first audio data corresponds to a segment of the second audio data; and splicing, fusing and / or combining each segment of the multiple segments of the first audio data and the second audio data based on weights to generate third audio data.

[0005] This application also provides an audio signal generation system, including: at least one processor; executable instructions, which can be executed by the at least one processor to cause the system to perform the audio signal generation method as described in the above embodiments.

[0006] This application embodiment also provides an audio signal generation system, including: an acquisition module, configured to acquire first audio data collected by a bone conduction sensor and second audio data collected by an air conduction sensor, wherein the first audio data and the second audio data represent a user's speech, and the first audio data and the second audio data are respectively composed of different frequency components; a weight determination unit, configured to divide the first audio data and the second audio data into multiple segments according to one or more frequency thresholds, wherein each segment of the first audio data corresponds to a segment of the second audio data; and a combination unit, configured to splice, fuse, and / or combine each segment of the multiple segments of the first audio data and the second audio data based on weights to generate third audio data.

[0007] This application also provides a non-transitory computer-readable medium that stores computer instructions, which, when executed, perform the audio signal generation method as described in the above embodiments.

[0008] Some of the additional features of this application will be described in the following description. Some of these additional features will be apparent to those skilled in the art from the study of the following description and the accompanying drawings, or from an understanding of the production or operation of the embodiments. The features of this application can be implemented and achieved through the practice or use of various methods, means, and combinations of the specific embodiments described below. Attached Figure Description

[0009] This application will be further described through exemplary embodiments. These exemplary embodiments will be described in detail with reference to the accompanying drawings. These embodiments are non-limiting exemplary embodiments, in which the same numbers in the figures denote similar structures, wherein:

[0010] Figure 1 This is a schematic diagram of an exemplary audio signal generation system according to some embodiments of this application.

[0011] Figure 2 This is a block diagram of an exemplary processing device according to some embodiments of this application.

[0012] Figure 3 This is a block diagram of an exemplary audio data generation module according to some embodiments of this application.

[0013] Figure 4 This is a flowchart illustrating an exemplary process for generating an audio signal according to some embodiments of this application.

[0014] Figure 5 This is a flowchart illustrating an exemplary process of reconstructing bone conduction audio data using a trained machine learning model, according to some embodiments of this application.

[0015] Figure 6 This is a flowchart illustrating an exemplary process of reconstructing bone conduction audio data using a harmonic correction model, according to some embodiments of this application.

[0016] Figure 7 This is a flowchart illustrating an exemplary process of reconstructing bone conduction audio data using sparse matrix techniques, according to some embodiments of this application.

[0017] Figure 8 This is a flowchart illustrating an exemplary process for generating audio data according to some embodiments of this application.

[0018] Figure 9 This is a flowchart illustrating an exemplary process for generating audio data according to some embodiments of this application.

[0019] Figure 10 The frequency response curves of bone conduction audio data, corresponding reconstructed bone conduction audio data, and corresponding air conduction audio data shown in some embodiments of this application are provided.

[0020] Figure 11 This is a frequency response curve of bone conduction audio data collected by bone conduction sensors located at different parts of the user's body, according to some embodiments of this application.

[0021] Figure 12 This is a frequency response curve of bone conduction audio data collected by bone conduction sensors located at different parts of the user's body, according to some embodiments of this application.

[0022] Figure 13 This is a time-frequency diagram of spliced ​​audio data generated by splicing bone conduction audio data and air conduction audio data according to some embodiments of this application.

[0023] Figure 14 This is a time-frequency diagram of spliced ​​audio data generated from bone conduction audio data spliced ​​at a frequency splicing point of 2000Hz and air conduction audio data after noise reduction using a Wiener filter, according to some embodiments of this application.

[0024] Figure 15 This is a time-frequency diagram of spliced ​​audio data generated from bone conduction audio data spliced ​​at a frequency splicing point of 2000Hz and air conduction audio data after noise reduction using spectral subtraction, according to some embodiments of this application.

[0025] Figure 16 This is a time-frequency diagram of bone conduction audio data according to some embodiments of this application.

[0026] Figure 17This is a time-frequency diagram of air conduction audio data according to some embodiments of this application.

[0027] Figure 18 This is a time-frequency diagram of spliced ​​audio data generated by splicing bone conduction audio data and air conduction audio data according to some embodiments of this application.

[0028] Figure 19 This is a time-frequency diagram of spliced ​​audio data generated by splicing bone conduction audio data and air conduction audio data according to some embodiments of this application.

[0029] Figure 20 This is a time-frequency diagram of spliced ​​audio data generated by splicing bone conduction audio data and air conduction audio data according to some embodiments of this application. Detailed Implementation

[0030] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. However, those skilled in the art should understand that this application can be implemented without these details. In other cases, to avoid unnecessarily obscuring some aspects of this application, well-known methods, procedures, systems, components, and / or circuits have been described at a higher level (without detail). It will be apparent to those skilled in the art that various changes can be made to the disclosed embodiments, and the general principles defined in this application can be applied to other embodiments and application scenarios without departing from the principles and scope of this application. Therefore, this application is not limited to the embodiments shown, but conforms to the broadest scope consistent with the scope of the claims.

[0031] These and other features, characteristics, functions and operating methods of related structural elements, as well as component assembly and manufacturing economics, will become more apparent from the following description of the accompanying drawings, which form part of this application specification. However, it should be understood that the drawings are for illustrative purposes only and are not intended to limit the scope of this application. It should also be understood that the drawings are not drawn to scale.

[0032] Flowcharts are used in this application to illustrate the operations performed by the system according to embodiments of this application. It should be understood that the operations in the flowcharts may not be performed sequentially. Instead, the steps may be processed in reverse order or simultaneously. Furthermore, one or more other operations may be added to these flowcharts. One or more operations may also be deleted from the flowcharts.

[0033] This application provides a system and method for generating audio signals. The system and method can acquire first audio data (also referred to as bone conduction audio data) collected by a bone conduction sensor. The system and method can acquire second audio data (also referred to as air conduction audio data) collected by an air conduction sensor. The bone conduction audio data and air conduction audio data can represent a user's speech, each composed of different frequency components. The system and method can generate audio data based on the bone conduction audio data and air conduction audio data. The frequency components above a certain frequency point in the generated audio data are increased compared to the frequency components above that frequency point in the bone conduction audio data. The system and method can determine target audio data representing the user's speech based on the generated audio data, the target audio data having higher fidelity than the bone conduction audio data and air conduction audio data. According to this application, the audio data generated based on the bone conduction audio data and air conduction audio data has more high-frequency components compared to the bone conduction audio data and less noise compared to the air conduction audio data, which can improve the fidelity and intelligibility of the generated audio data compared to the bone conduction audio data and air conduction audio data. In some embodiments, reconstructed bone conduction audio data can be obtained by adding high-frequency components to the bone conduction audio data. The reconstructed bone conduction audio data is closer to air conduction audio data, and its quality is higher than that of bone conduction audio data, further improving the quality of the generated audio data. In some embodiments, audio data can be generated by splicing bone conduction audio data and air conduction audio data using different frequency splicing points based on factors such as environmental noise. This can reduce noise in the audio data while maintaining its fidelity.

[0034] Figure 1 This is a schematic diagram of an exemplary audio signal generation system 100 according to some embodiments of this application. The audio signal generation system 100 may include an audio acquisition device 110, a server 120, a terminal 130, a storage device 140, and a network 150.

[0035] Audio acquisition device 110 can acquire audio data (e.g., audio signals) by capturing the user's voice or speech while the user is speaking. For example, when a user speaks, the sound emitted by the user causes vibrations in the air around the user's mouth and / or in the user's body tissues (e.g., skull). Audio acquisition device 110 can receive the vibrations and convert them into electrical signals (e.g., analog or digital signals), also known as audio data. The audio data can be transmitted in the form of electrical signals via network 150 to server 120 / terminal 130 and / or storage device 140. In some embodiments, audio acquisition device 110 may include a recorder, headphones (e.g., Bluetooth headphones, wired headphones), hearing aid devices, etc.

[0036] In some embodiments, the audio acquisition device 110 can be connected to the speaker wirelessly (e.g., via network 150) and / or via a wired connection. The audio acquisition device 110 can send the acquired audio data to the speaker to play and / or reproduce the user's voice. In some embodiments, the speaker and the audio acquisition device 110 can be integrated into a single device, such as headphones. In some embodiments, the audio acquisition device 110 and the speaker can be separate from each other. For example, the audio acquisition device 110 can be installed in a first terminal (e.g., headphones), and the speaker can be installed in another terminal (e.g., terminal 130).

[0037] In some embodiments, the audio acquisition device 110 may include a bone conduction microphone 112 and an air conduction microphone 114. The bone conduction microphone 112 may include a bone conduction sensor for acquiring bone conduction audio data. The bone conduction sensor can generate bone conduction audio data by acquiring vibration signals conducted through the user's skeletal (e.g., skull) tissue when the user speaks. In some embodiments, the bone conduction sensor may form a bone conduction sensor array. In some embodiments, the bone conduction microphone 112 may be placed at and / or in contact with a part of the user's body to acquire bone conduction audio data. Parts of the user's body may include the forehead, neck (e.g., throat), face (e.g., the area around the mouth, chin), top of the head, mastoid process, area around the ear, area inside the ear, temple, etc., or any combination thereof. For example, the bone conduction microphone 112 may be placed at and / or in contact with the tragus, auricle, internal auditory canal, external auditory canal, etc. In some embodiments, the acoustic characteristics of the bone conduction audio data may vary depending on the part of the user's body at which the bone conduction microphone 112 is located and / or in contact. For example, bone conduction microphone 112 located in the area around the ear acquires bone conduction audio data with higher energy than bone conduction microphone 112 located on the forehead. Air conduction microphone 114 may include one or more air conduction sensors for acquiring air conduction audio data conducted through air while the user speaks. In some embodiments, the air conduction sensors may form an array. In some embodiments, air conduction microphone 114 may be placed within a certain range (e.g., 0 cm, 1 cm, 2 cm, 5 cm, 10 cm, 20 cm, etc.) from the user's mouth. The acoustic characteristics of the air conduction audio data (e.g., the average amplitude of the air conduction audio data) may vary depending on the distance between the air conduction microphone 114 and the user's mouth. For example, the greater the distance between the air conduction microphone 114 and the user's mouth, the smaller the average amplitude of the air conduction audio data may be.

[0038] In some embodiments, server 120 may be a single server or a group of servers. The server group may be centralized (e.g., a data center) or distributed (e.g., server 120 may be a distributed system). In some embodiments, server 120 may be local or remote. For example, server 120 may access information and / or data stored in terminal 130 and / or storage device 140 via network 150. As another example, server 120 may directly connect to terminal 130 and / or storage device 140 to access stored information and / or data. In some embodiments, server 120 may be implemented on a cloud platform. By way of example only, the cloud platform may include private cloud, public cloud, hybrid cloud, community cloud, distributed cloud, internal cloud, multi-tiered cloud, etc., or any combination thereof.

[0039] In some embodiments, server 120 may include processing device 122. Processing device 122 may process information and / or data related to audio signal generation to perform one or more functions described in this application. For example, processing device 122 may acquire bone conduction audio data collected by bone conduction microphone 112 and air conduction audio data collected by air conduction microphone 114, wherein the bone conduction audio data and air conduction audio data represent the speech of the same user (or user). Processing device 122 may generate target audio data based on the bone conduction audio data and air conduction audio data. As another example, processing device 122 may obtain a trained machine learning model and / or constructed filters from storage device 140 or any other storage device. Processing device 122 may use the trained machine learning model and / or constructed filters to reconstruct the bone conduction audio data. As yet another example, processing device 122 may determine a trained machine learning model by training a primary machine learning model using multiple sets of speech samples (i.e., training data). Each set of speech samples may include bone conduction audio data and air conduction audio data representing the same user's speech. As another example, processing device 122 can perform noise reduction on air conduction audio data to obtain noise-reduced air conduction audio data. Processing device 122 can generate target audio data based on reconstructed bone conduction audio data and noise-reduced air conduction audio data. In some embodiments, processing device 122 may include one or more processing engines (e.g., a single-chip processing engine or a multi-chip processing engine). By way of example only, processing device 122 may include a central processing unit (CPU), application-specific integrated circuit (ASIC), application-specific instruction set processor (ASIP), graphics processing unit (GPU), physical processing unit (PPU), digital signal processor (DSP), field-programmable gate array (FPGA), programmable logic device (PLD), controller, microcontroller unit, reduced instruction set computer (RISC), microprocessor, etc., or any combination thereof.

[0040] In some embodiments, terminal 130 may include mobile device 130-1, tablet computer 130-2, laptop computer 130-3, built-in device in vehicle 130-4, wearable device 130-5, etc., or any combination thereof. In some embodiments, mobile device 130-1 may include smart home device, smart mobile device, virtual reality device, augmented reality device, etc., or any combination thereof. In some embodiments, smart home device may include smart lighting device, smart appliance control device, smart monitoring device, smart TV, smart camera, walkie-talkie, etc., or any combination thereof. In some embodiments, smart mobile device may include smartphone, personal digital assistant (PDA), gaming device, navigation device, point of sale (POS), etc., or any combination thereof. In some embodiments, virtual reality device and / or augmented reality device includes virtual reality helmet, virtual reality glasses, virtual reality headset, augmented reality helmet, augmented reality glasses, augmented reality headset, etc., or any combination thereof. For example, virtual reality device and / or augmented reality device may include Google... TM Glasses, Oculus Rift, HoloLens, GearVR, etc. In some embodiments, in-vehicle device 130-4 may include an in-vehicle computer, in-vehicle TV, etc. In some embodiments, terminal 130 may be a device with positioning technology for locating the position of passengers and / or terminal 130. In some embodiments, wearable device 130-5 may include smart bracelets, smart shoes and socks, smart glasses, smart helmets, smartwatches, smart clothing, smart backpacks, smart accessories, etc., or any combination thereof. In some embodiments, audio acquisition device 110 may be integrated into terminal 130.

[0041] Storage device 140 may store data and / or instructions. For example, storage device 140 may store data of multiple sets of speech samples, one or more machine learning models, trained machine learning models and / or constructed filters, audio data acquired by bone conduction microphone 112 and air conduction microphone 114, etc. In some embodiments, storage device 140 may store data acquired from terminal 130 and / or audio acquisition device 110. In some embodiments, storage device 140 may store data and / or instructions that server 120 may execute for performing the exemplary methods described in this invention. In some embodiments, storage device 140 may include mass storage, removable storage, volatile read-write memory, read-only memory (ROM), etc., or any combination thereof. Exemplary mass storage devices may include disks, optical disks, solid-state drives, etc. Exemplary removable storage may include flash drives, floppy disks, optical disks, memory cards, compact disks, magnetic tapes, etc. Exemplary volatile read-write memory may include random access memory (RAM). Exemplary RAMs may include Dynamic Random Access Memory (DRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM), Static Random Access Memory (SRAM), Thyristor Random Access Memory (T-RAM), and Zero Capacitor Random Access Memory (Z-RAM), etc. Exemplary ROMs may include Mask Read-Only Memory (MROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Optical Disc Read-Only Memory (CD-ROM), and Digital Multifunction Disk Read-Only Memory, etc. In some embodiments, the storage device 140 may execute on a cloud platform. By way of example only, the cloud platform may include private cloud, public cloud, hybrid cloud, community cloud, distributed cloud, internal cloud, multi-tiered cloud, etc., or any combination thereof.

[0042] In some embodiments, storage device 140 may be connected to network 150 to communicate with one or more components of audio signal generation system 100 (e.g., audio acquisition device 110, server 120, and terminal 130). One or more components of audio signal generation system 100 may access data or instructions stored in storage device 140 via network 150. In some embodiments, storage device 140 may be directly connected to or communicate with one or more components of audio signal generation system 100 (e.g., audio acquisition device 110, server 120, and terminal 130). In some embodiments, storage device 140 may be part of server 120.

[0043] Network 150 can facilitate the exchange of information and / or data. In some embodiments, one or more components of the audio signal generation system 100 (e.g., audio acquisition device 110, server 120, terminal 130, and storage device 140) can transmit information and / or data to other components of the audio signal generation system 100 via network 150. For example, server 120 can acquire bone conduction audio data and air conduction audio data from terminal 130 via network 150. In some embodiments, network 150 can be any form of wired or wireless network, or any combination thereof. By way of example only, network 150 may include cable networks, wired networks, fiber optic networks, telecommunications networks, intranets, the Internet, local area networks (LANs), wide area networks (WANs), wireless local area networks (WLANs), metropolitan area networks (MANs), public switched telephone networks (PSTNs), Bluetooth networks, cellular networks, near field communication (NFC) networks, and any combination thereof. In some embodiments, network 150 may include one or more network access points. For example, network 150 may include wired or wireless network access points, such as base stations and / or internet exchange points, through which one or more components of audio signal generation system 100 may connect to network 150 to exchange data and / or information.

[0044] Those skilled in the art will understand that when an element (or component) of the audio signal generation system 100 is executed, it can be performed via electrical and / or electromagnetic signals. For example, when the bone conduction microphone 112 sends bone conduction audio data to the server 120, the processor of the bone conduction microphone 112 can generate an electrical signal encoding the bone conduction audio data. The processor of the bone conduction microphone 112 can then transmit the electrical signal to an output port. If the bone conduction microphone 112 communicates with the server 120 via a wired network, the output port can be physically connected to a cable that can also transmit the electrical signal to an input port of the server 120. If the bone conduction microphone 112 communicates with the server 120 via a wireless network, the output port of the bone conduction microphone 112 can be one or more antennas that convert electrical signals into electromagnetic signals. Similarly, the gas conduction microphone 114 can transmit gas conduction audio data to the server 120 via electrical or electromagnetic signals. In electronic devices such as terminal 130 and / or server 120, instructions and / or actions are performed via electrical signals when their processor processes instructions, issues instructions, and / or performs actions. For example, when a processor retrieves or acquires data from a storage medium, it can send electrical signals to a read / write device on the storage medium, which can read or write structured data into the storage medium. This structured data can be transmitted to the processor in the form of electrical signals via the electronic device's bus. Here, an electrical signal can refer to a single electrical signal, a series of electrical signals, and / or at least two discontinuous electrical signals.

[0045] Figure 2 This is a block diagram illustrating an exemplary processing device according to some embodiments of this application. For example... Figure 2 As shown, the processing device 122 may include an acquisition module 210, a preprocessing module 220, an audio data generation module 230, and a storage module 240. Each of these modules may be hardware circuitry designed to perform certain actions, for example, according to instructions stored in one or more storage media, and / or any combination of hardware circuitry and one or more storage media.

[0046] The acquisition module 210 can be configured to acquire data used to generate an audio signal. For example, the acquisition module 210 can acquire raw audio data, one or more models, training data for training a machine learning model, etc. In some embodiments, the acquisition module 210 can acquire first audio data acquired by a bone conduction sensor. As used herein, a bone conduction sensor can refer to any sensor (e.g., bone conduction microphone 112) capable of acquiring vibrational signals conducted by the user's bone tissue (e.g., skull) while the user is speaking, as described elsewhere in this application (e.g., ...). Figure 1 (and its description). In some embodiments, the first audio data may include an audio signal in the time domain, an audio signal in the frequency domain, etc. The first audio data may include an analog signal or a digital signal. The acquisition module 210 may also acquire second audio data acquired by the gas conduction sensor. The gas conduction sensor may refer to any sensor capable of acquiring vibration signals conducted by air when a user speaks (e.g., gas conduction microphone 114), as described elsewhere in this application (e.g., Figure 1 (and its description). In some embodiments, the second audio data may include audio signals in the time domain and audio signals in the frequency domain, etc. The second audio data may include analog signals or digital signals. In some embodiments, the acquisition module 210 may obtain a trained machine learning model, a constructed filter, a harmonic correction model, etc., for reconstructing the first audio data. In some embodiments, the processing device 122 may acquire one or more models, first audio data, and / or second audio data in real time or periodically from an air conduction sensor (e.g., an air conduction microphone 114), a terminal 130, a storage device 140, or any other storage device via a network 150.

[0047] The preprocessing module 220 can be configured to preprocess first audio data and / or second audio data. The first and second audio data, after preprocessing, can also be referred to as preprocessed first audio data and preprocessed second audio data, respectively. Exemplary preprocessing operations may include domain transformation operations, signal calibration operations, audio reconstruction operations, speech enhancement operations, etc. In some embodiments, the preprocessing module 220 can perform domain transformation operations by performing Fourier transform or inverse Fourier transform. In some embodiments, the preprocessing module 220 can perform a normalization operation on the first and / or second audio data to obtain normalized first audio data and / or normalized second audio data for calibration. In some embodiments, the preprocessing module 220 can perform a speech enhancement operation on the second audio data (or normalized second audio data). In some embodiments, the preprocessing module 220 can perform a noise reduction operation on the second audio data (or normalized second audio data) to obtain noise-reduced second audio data. In some embodiments, the preprocessing module 220 may use a trained machine learning model, a constructed filter, a harmonic correction model, sparse matrix techniques, or any combination thereof to perform an audio reconstruction operation on the first audio data (or normalized first audio data) to generate reconstructed first audio data.

[0048] The audio data generation module 230 can be configured to generate third audio data based on first audio data (or preprocessed first audio data) and second audio data (or preprocessed second audio data). In some embodiments, the noise level associated with the third audio data may be lower than the noise level associated with the second audio data (or preprocessed second audio data). In some embodiments, the audio data generation module 230 can generate third audio data based on the first audio data (or preprocessed first audio data) and second audio data (or preprocessed second audio data) according to one or more frequency thresholds. In some embodiments, the audio data generation module 230 can determine a single frequency threshold. The audio data generation module 230 can concatenate the first audio data (or preprocessed first audio data) and the second audio data (or preprocessed second audio data) in the frequency domain according to the single frequency threshold to generate third audio data.

[0049] In some embodiments, the audio data generation module 230 may determine, at least partially based on a frequency threshold, a first weight and a second weight for the low-frequency portion and the high-frequency portion of the first audio data (or preprocessed first audio data), respectively. The low-frequency portion of the first audio data (or preprocessed first audio data) includes frequency components in the first audio data (or preprocessed first audio data) that are less than the frequency threshold. The high-frequency portion of the first audio data (or preprocessed first audio data) includes frequency components in the first audio data (or preprocessed first audio data) that are greater than the frequency threshold. In some embodiments, the audio data generation module 230 may determine, at least partially based on a frequency threshold, a third weight and a fourth weight for the low-frequency portion and the high-frequency portion (or preprocessed second audio data) of the second audio data (or preprocessed second audio data), respectively. The low-frequency portion of the second audio data (or preprocessed second audio data) includes frequency components in the second audio data (or preprocessed second audio data) that are less than the frequency threshold. The high-frequency portion of the second audio data (or preprocessed second audio data) includes frequency components in the second audio data (or preprocessed second audio data) that are greater than the frequency threshold. In some embodiments, the audio data generation module 230 can determine the third audio data by weighting the low-frequency and high-frequency portions of the first audio data (or preprocessed first audio data) and the low-frequency and high-frequency portions of the second audio data (or preprocessed second audio data) using a first weight, a second weight, a third weight, and a fourth weight, respectively.

[0050] In some embodiments, the audio data generation module 230 may determine, at least in part, a weight corresponding to the first audio data (or preprocessed first audio data) and a weight corresponding to the second audio data (or preprocessed second audio data) based on the first audio data (or preprocessed first audio data) and / or the second audio data (or preprocessed second audio data). The audio data generation module 230 may use the weights corresponding to the first audio data (or preprocessed first audio data) and the weights corresponding to the second audio data (or preprocessed second audio data) to weight the first audio data (or preprocessed first audio data) and the second audio data (or preprocessed second audio data) to determine the third audio data.

[0051] In some embodiments, the audio data generation module 230 may determine target audio data representing a user's speech based on third audio data, which has a higher fidelity than the first and second audio data. In some embodiments, the audio data generation module 230 may designate the third audio data as the target audio data. In some embodiments, the audio data generation module 230 may perform post-processing operations on the third audio data to obtain the target audio data. In some embodiments, the audio data generation module 230 may perform an inverse Fourier transform operation on the third audio data in the frequency domain to obtain the target audio data in the time domain. In some embodiments, the audio data generation module 230 may perform a noise reduction operation on the third audio data to obtain the target audio data. In some embodiments, the audio data generation module 230 may send a signal via network 150 to a client terminal (e.g., terminal 130), storage device 140, and / or any other storage device (not shown in the audio signal generation system 100). This signal may include the target audio data. The signal may also be configured to cause the client terminal to play the target audio data.

[0052] Storage module 240 can be configured to store data and / or instructions associated with audio signal generation system 100. For example, storage module 240 can store speech sample data, machine learning models, trained machine learning models and / or constructed filters, audio data acquired by bone conduction microphone 112 and / or air conduction microphone 114, etc. In some embodiments, storage module 240 can be the same as storage device 140 in the configuration.

[0053] It should be noted that the above description is for illustrative purposes only and is not intended to limit the scope of this application. Obviously, those skilled in the art can make various changes and modifications based on the description in this application. However, these changes and modifications will not depart from the scope of this application. For example, the storage module 240 may be omitted. As another example, the audio data generation module 230 and the storage module 240 may be integrated into a single module.

[0054] Figure 3 This is a block diagram illustrating an exemplary audio data generation module according to some embodiments of this application. For example... Figure 3 As shown, the audio data generation module 230 may include a frequency determination unit 310, a weight determination unit 320, and a combination unit 330. Each of the above sub-modules may be a hardware circuit designed to perform certain actions, for example, according to instructions stored in one or more storage media, and / or any combination of hardware circuitry and one or more storage media.

[0055] The frequency determination unit 310 can be configured to determine one or more frequency thresholds based at least in part on bone conduction audio data and / or air conduction audio data. In some embodiments, the frequency threshold may be a frequency point of the bone conduction audio data and / or air conduction audio data. In some embodiments, the frequency threshold may be different from the frequency point of the bone conduction audio data and / or air conduction audio data. In some embodiments, the frequency determination unit 310 may determine the frequency threshold based on a frequency response curve associated with the bone conduction audio data. The frequency response curve associated with the bone conduction audio data may include a frequency response value that varies with frequency. In some embodiments, the frequency determination unit 310 may determine one or more frequency thresholds based on the frequency response value of the frequency response curve associated with the bone conduction audio data. In some embodiments, the frequency determination unit 310 may determine one or more frequency thresholds based on the variation characteristics of the frequency response curve. In some embodiments, the frequency determination unit 310 may determine one or more frequency thresholds based on the frequency response curve associated with the reconstructed bone conduction audio data. In some embodiments, the frequency determination unit 310 may determine one or more frequency thresholds based on the noise level associated with at least a portion of the air conduction audio data. In some embodiments, the noise level may be represented by the signal-to-noise ratio (SNR) of the air conduction audio data. A higher SNR may indicate a lower noise level. The higher the signal-to-noise ratio associated with air conduction audio data, the higher the frequency threshold.

[0056] The weighting determination unit 320 can be configured to divide the bone conduction audio data and air conduction audio data into multiple segments based on one or more frequency thresholds. Each segment of the bone conduction audio data may correspond to a segment of the air conduction audio data. As used herein, a segment of air conduction audio data corresponding to a segment of bone conduction audio data may mean that two segments of the bone conduction audio data and air conduction audio data are defined by one or two identical frequency thresholds. In some embodiments, the count or number of frequency thresholds may be one, and the weighting determination unit 320 may divide the bone conduction audio data and air conduction audio data into two segments.

[0057] The weight determination unit 320 may also be configured to determine the weight of each segment among multiple segments of bone conduction audio data and air conduction audio data. In some embodiments, the weight of a specific segment of bone conduction audio data and the weight of a corresponding specific segment of air conduction audio data satisfy a certain condition such that the sum of the weights of the specific segments of bone conduction audio data and the corresponding specific segments of air conduction audio data equals 1. In some embodiments, the weight determination unit 320 may determine the weights of different segments of bone conduction audio data or air conduction audio data based on the SNR of the air conduction audio data.

[0058] The combining unit 330 can be configured to splice, fuse, and / or combine each segment of bone conduction audio data and air conduction audio data based on weights to generate spliced, fused, or combined audio data. In some embodiments, the combining unit 330 can determine the low-frequency portion of the bone conduction audio data and the high-frequency portion of the air conduction audio data based on a single frequency threshold. The combining unit 330 can splice and / or combine the low-frequency portion of the bone conduction audio data and the high-frequency portion of the air conduction audio data to generate spliced ​​audio data. The combining unit 330 can determine the low-frequency portion of the bone conduction audio data and the high-frequency portion of the air conduction audio data based on one or more filters. In some embodiments, the combining unit 330 can weight the low-frequency portion of the bone conduction audio data, the high-frequency portion of the bone conduction audio data, the low-frequency portion of the air conduction audio data, and the high-frequency portion of the air conduction audio data using a first weight, a second weight, a third weight, and a fourth weight, respectively, to determine the spliced, fused, or combined audio data. In some embodiments, the combining unit 330 can determine the fused or combined audio data by weighting the bone conduction audio data and the air conduction audio data, respectively.

[0059] It should be noted that the above description is for illustrative purposes only and is not intended to limit the scope of this application. Obviously, those skilled in the art can make various changes and modifications based on the description in this application. However, these changes and modifications will not depart from the scope of this application. For example, the audio data generation module 230 may also include an audio data partitioning submodule (…). Figure 3 (Not shown in the image). The audio data segmentation submodule can be configured to segment each bone conduction audio data and air conduction audio data into multiple segments based on one or more frequency thresholds. Alternatively, the weighting determination unit 320 and the combination unit 330 can be integrated into a single module.

[0060] Figure 4 This is a flowchart illustrating an exemplary process for generating an audio signal according to some embodiments of this application. In some embodiments, process 400 may be implemented as instructions (e.g., an application program) stored in storage device 140. Processing device 122 may execute the instructions, and when executing the instructions, processing device 122 may be configured to perform processing process 400. The operation of the process shown below is for illustrative purposes only. In some embodiments, process 400 may be accomplished using one or more additional operations not described, and / or without one or more operations discussed. Additionally, Figure 4 The order of operations of process 400 shown and described below is non-limiting.

[0061] In 410, processing device 122 (e.g., acquisition module 210) can acquire the first audio data collected by the bone conduction sensor. As used herein, a bone conduction sensor refers to any sensor (e.g., bone conduction microphone 112) that can collect vibrational signals conducted by the user's bone tissue (e.g., skull) while the user (or user) speaks, as described elsewhere in this application (e.g., Figure 1 (and its description). Vibration signals acquired by a bone conduction sensor can be converted into audio data (e.g., audio signals) by the bone conduction sensor or other devices (e.g., amplifiers, analog-to-digital converters (ADCs), etc.). The audio data acquired by the bone conduction sensor (e.g., first audio data) can also be referred to as bone conduction audio data. In some embodiments, the first audio data may include audio signals in the time domain, audio signals in the frequency domain, etc. The first audio data may include analog signals or digital signals. In some embodiments, the processing device 122 may acquire the first audio data in real time or periodically from the bone conduction sensor (e.g., bone conduction microphone 112), terminal 130, storage device 140, or any other storage device via network 150.

[0062] The first audio data can be represented by a superposition of multiple waves (e.g., sine waves, harmonics, etc.) with different frequencies and / or intensities (i.e., amplitudes). As used herein, a wave with a specific frequency can also be referred to as a frequency component with a specific frequency. In some embodiments, the frequency components included in the first audio data acquired by the bone conduction sensor can be in the frequency range of 0 Hz to 20 kHz, or 20 Hz to 10 kHz, or 20 Hz to 4000 Hz, or 20 Hz to 3000 Hz, or 800 Hz to 3500 Hz, or 800 Hz to 3000 Hz, or 1500 Hz to 3000 Hz. The first audio data can be acquired and / or generated by the bone conduction sensor when the user speaks. The first audio data can represent the content of the user's speech (i.e., the user's voice). For example, the first audio data can include acoustic features and / or semantic information that can reflect the content of the user's voice. The acoustic features of the first audio data can include duration-related features, energy-related features, fundamental frequency-related features, frequency spectrum-related features, phase spectrum-related features, etc. Duration-related features can also be referred to as duration features. Exemplary duration features may include speech rate, short-time average zero-crossing rate, etc. Features associated with energy may also be referred to as energy or amplitude features. Exemplary energy or amplitude features may include short-time average energy, short-time average amplitude, short-time energy gradient, average amplitude change rate, short-time maximum amplitude, etc. Features associated with the fundamental frequency may also be referred to as fundamental frequency features. Exemplary fundamental frequency features may include the fundamental frequency, the pitch of the fundamental frequency, the average fundamental frequency, the maximum fundamental frequency, the fundamental frequency range, etc. Exemplary features associated with the frequency spectrum may include formant features, linear predictive cepstral coefficients (LPCC), Mel frequency cepstral coefficients (MFCC), etc. Exemplary features associated with the phase spectrum may include instantaneous phase, initial phase, etc.

[0063] In some embodiments, first audio data can be acquired and / or generated by placing a bone conduction sensor on a part of the user's body and / or by bringing the bone conduction sensor into contact with the user's skin. Parts of the user's body in contact with the bone conduction sensor include, but are not limited to, the forehead, neck (e.g., throat), mastoid process, area around the ear, inner ear area, temple, face (e.g., area around the mouth, chin), and top of the head. For example, the bone conduction microphone 112 can be placed on and / or in contact with the tragus, auricle, inner ear canal, outer ear canal, etc. In some embodiments, the first audio data can vary depending on the part of the user's body in contact with the bone conduction sensor. For example, different parts of the user's body in contact with the bone conduction sensor can cause variations in the frequency characteristics (e.g., amplitude of frequency components), noise included in the first audio data, etc. For example, the signal strength of the first audio data acquired by a bone conduction sensor located on the neck is greater than the signal strength of the first audio data acquired by a bone conduction sensor located on the tragus. The signal strength of the first audio data acquired by a bone conduction sensor located at the tragus is greater than the signal strength of the first audio data acquired by a bone conduction sensor located at the ear canal. For example, bone conduction audio data acquired by a first bone conduction sensor located in the area surrounding the user's ear has more frequency components than bone conduction audio data simultaneously acquired by a second bone conduction sensor with the same configuration but located at the top of the user's head. In some embodiments, the first audio data may be acquired by a bone conduction sensor located at a part of the user's body applying a specific pressure within a certain range (e.g., 0N to 1N, or 0N to 0.8N, etc.) to that part. For example, the first audio data may be acquired by a bone conduction sensor located at the tragus of the user's body applying a specific pressure (e.g., 0 Newtons, or 0.2N, or 0.4N, or 0.8N, etc.) to that part. Differences in the pressure applied by the bone conduction sensors to the same body part may cause variations in the frequency components, acoustic characteristics (e.g., the amplitude of the frequency components), and noise in the first audio data acquired by the bone conduction sensors. For example, when the pressure increases from 0 N to 0.8 N, the signal strength of the first audio data gradually increases initially, then the rate of increase slows down, eventually reaching saturation. Further descriptions of the effects of placing the bone conduction sensor on bone conduction audio data at different body sites can be found elsewhere in this application (e.g., Figure 11 (and its description). Further description of the effect of different pressures applied by the bone conduction sensor to the user's body parts on bone conduction audio data can be found elsewhere in this application (e.g., Figure 12 (and its description).

[0064] In 420, processing device 122 (e.g., acquisition module 210) can acquire second audio data collected by the air conduction sensor. As used herein, an air conduction sensor can refer to any sensor capable of collecting vibration signals conducted through air while a user is speaking (e.g., air conduction microphone 114), as described elsewhere in this application (e.g., Figure 1 (and its description). The vibration signal acquired by the air conduction sensor can be converted into audio data (e.g., audio signal) by the air conduction sensor or other devices (e.g., amplifier, analog-to-digital converter (ADC) etc.). The audio data acquired by the air conduction sensor (e.g., second audio data) can also be referred to as air conduction audio data. In some embodiments, the second audio data may include an audio signal in the time domain, an audio signal in the frequency domain, etc. The second audio data may include analog signals or digital signals. In some embodiments, the processing device 122 may acquire the second audio data in real time or periodically from the air conduction sensor (e.g., air conduction microphone 114), terminal 130, storage device 140, or any other storage device via network 150. In some embodiments, the second audio data can be acquired by placing the air conduction sensor within a certain distance (e.g., 0cm, 1cm, 2cm, 5cm, 10cm, 20cm, etc.) from the user's mouth. In some embodiments, different distances between the air conduction sensor and the user's mouth may result in different acquired second audio data (e.g., the average amplitude of the second audio data).

[0065] The second audio data can be represented by a superposition of multiple waves (e.g., sine waves, harmonics, etc.) with different frequencies and / or intensities (i.e., amplitudes). In some embodiments, the frequency components included in the second audio data acquired by the air conduction sensor can be in the frequency range of 0 Hz to 20 kHz, or 20 Hz to 20 kHz, or 800 Hz to 10 kHz. The air conduction sensor can acquire and / or generate the second audio data when the user speaks. The second audio data can represent the content of the user's speech (i.e., the user's voice). For example, the second audio data includes acoustic features and / or semantic information that can reflect the content of the user's voice. The acoustic features of the second audio data can include features associated with duration, features associated with energy, features associated with fundamental frequency, features associated with frequency spectrum, features associated with phase spectrum, etc., as described in operation 410.

[0066] In some embodiments, the first audio data and the second audio data may represent the same speech of the same user through different frequency components. The first and second audio data representing the same speech of the same user may refer to the first audio data and the second audio data simultaneously acquired by a bone conduction sensor and an air conduction sensor, respectively, when the user speaks. The first audio data acquired by the bone conduction sensor may include a first frequency component. The second audio data may include a second frequency component. In some embodiments, the second frequency component includes at least a portion of the first frequency component. The semantic information included in the second audio data may be the same as or different from the semantic information included in the first audio data. The acoustic characteristics of the second audio data may be the same as or different from the acoustic characteristics of the first audio data. For example, the amplitude of a certain frequency component in the first audio data may be different from the amplitude of the same frequency component in the second audio data. As another example, the first audio data may have more frequency components below a certain frequency point (e.g., 2000Hz) or within a certain frequency range (e.g., 20Hz to 2000Hz) than the second audio data may have more frequency components below that frequency point (e.g., 2000Hz) or within that frequency range (e.g., 20Hz to 2000Hz). The frequency components in the first audio data that are above a certain frequency point (e.g., 3000Hz) or within a certain frequency range (e.g., 3000Hz to 20kHz) may be fewer than the frequency components in the second audio data that are above that frequency point (e.g., 3000Hz) or within that frequency range (e.g., 3000Hz to 20kHz). As used herein, the statement that the frequency components in the first audio data that are below a certain frequency point (e.g., 2000Hz) or within a certain frequency range (e.g., 20Hz to 2000Hz) are more than the frequency components in the second audio data that are below that frequency point (e.g., 2000Hz) or within that frequency range (e.g., 20Hz to 2000Hz) can mean that the count or number of frequency components in the first audio data that are below that frequency point (e.g., 2000Hz) or within that frequency range (e.g., 20Hz to 2000Hz) is greater than the count or number of frequency components in the second audio data that are below that frequency point (e.g., 2000Hz) or within that frequency range (e.g., 20Hz to 2000Hz).

[0067] In 430, processing device 122 (e.g., preprocessing module 220) can preprocess at least one of the first audio data or the second audio data. The preprocessed first audio data and the second audio data can also be referred to as preprocessed first audio data and preprocessed second audio data, respectively. Exemplary preprocessing operations may include domain transformation operations, signal calibration operations, audio reconstruction operations, speech enhancement operations, etc.

[0068] Domain transformation operations can be performed to convert the first audio data and / or the second audio data from the time domain to the frequency domain or vice versa. In some embodiments, the processing device 122 can perform the domain transformation operation by performing a Fourier transform or an inverse Fourier transform. In some embodiments, to perform the domain transformation operation, the processing device 122 can perform framing operations, windowing operations, etc., on the first audio data and / or the second audio data. For example, the first audio data can be divided into one or more speech frames. Each speech frame can be audio data including a duration segment (e.g., 5ms, 10ms, 15ms, 20ms, 25ms, etc.), within which the audio data of each frame can be considered approximately stable. A wave segmentation function can be used to perform a windowing operation on the speech frames to obtain processed speech frames. As used herein, the wave segmentation function can be referred to as a window function. Exemplary window functions may include the Hanning window, Hamming, Blackman-Harris window, etc. Finally, a Fourier transform operation can be used to convert the first audio data from the time domain to the frequency domain based on the processed speech frames.

[0069] Signal calibration operations can be used to unify the order of magnitude (e.g., amplitude) of the first and second audio data to eliminate differences in the order of magnitude between the first and / or second audio data caused by, for example, sensitivity differences between bone conduction sensors and air conduction sensors. In some embodiments, processing device 122 can perform a normalization operation on the first and / or second audio data to calibrate the first and / or second audio data, obtaining normalized first and / or normalized second audio data. For example, processing device 122 can determine the normalized first and / or normalized second audio data according to equation (1), as follows:

[0070]

[0071] Among them, S normalized Refers to the normalized first audio data (or normalized second audio data), S initial This refers to the first audio data (or the second audio data), |S max | can represent the maximum value of the absolute amplitude of the first audio data (or the second audio data).

[0072] Speech enhancement operations can be used to reduce noise or other irrelevant and unwanted information in audio data (e.g., first audio data and / or second audio data). Speech enhancement operations performed on the first audio data (or normalized first audio data) and / or the second audio data (or normalized second audio data) can use speech enhancement algorithms including spectral subtraction-based speech enhancement algorithms, wavelet analysis-based speech enhancement algorithms, Kalman filter-based speech enhancement algorithms, signal subspace-based speech enhancement algorithms, auditory masking effect-based speech enhancement algorithms, independent component analysis-based speech enhancement algorithms, neural network techniques, etc., or combinations thereof. In some embodiments, the speech enhancement operation may include a noise reduction operation. In some embodiments, the processing device 122 may perform a noise reduction operation on the second audio data (or normalized second audio data) to obtain denoised second audio data. In some embodiments, the normalized second audio data and / or the denoised second audio data may also be referred to as preprocessed second audio data. In some embodiments, the noise reduction operation may include using Wiener filters, spectral subtraction, adaptive algorithms, minimum mean square error (MMSE) estimation algorithms, etc., or any combination thereof.

[0073] Audio reconstruction operations can be used to enhance or increase the frequency components of initial bone conduction audio data (e.g., first audio data or normalized first audio data) above a certain frequency point (e.g., 2000Hz, 3000Hz) or within a frequency range (e.g., 2000Hz to 20kHz, 3000Hz to 20kHz) to improve the fidelity of the reconstructed bone conduction audio data relative to the initial bone conduction audio data (e.g., first audio data or normalized first audio data). The reconstructed bone conduction audio data can be similar to, close to, or identical to ideal air conduction audio data with little or no noise, and the reconstructed bone conduction audio data and the initial bone conduction audio data represent the same speech from the same user, the ideal air conduction audio data being acquired simultaneously by an air conduction sensor and a bone conduction sensor of the initial bone conduction audio data. The reconstructed bone conduction audio data can be equivalent to air conduction audio data, or referred to as equivalent air conduction audio data corresponding to the initial bone conduction audio data. As used herein, reconstructed bone conduction audio data that is similar to, close to, or identical to ideal air conduction audio data can refer to reconstructed bone conduction audio data whose similarity to ideal air conduction audio data can be greater than a certain threshold (e.g., 90%, 80%, 70%, etc.). Further descriptions of reconstructed bone conduction audio data, initial bone conduction audio data, and ideal air conduction audio data can be found elsewhere in this application (e.g., Figure 10 (and its description).

[0074] In some embodiments, the processing device 122 may reconstruct the first audio data using a trained machine learning model, a constructed filter, a harmonic correction model, sparse matrix techniques, or any combination thereof to generate reconstructed first audio data. In some embodiments, one of the following methods may be used to generate the reconstructed first audio data: a trained machine learning model, a constructed filter, a harmonic correction model, or sparse matrix techniques. In some embodiments, at least two of the following methods may be used to generate the reconstructed first audio data: a trained machine learning model, a constructed filter, a harmonic correction model, or sparse matrix techniques. For example, the processing device 122 may generate intermediate first audio data by reconstructing the first audio data using a trained machine learning model. The processing device 122 may generate reconstructed first audio data by reconstructing the intermediate first audio data using one of the following methods: a constructed filter, a harmonic correction model, or sparse matrix techniques. As another example, the processing device 122 may generate intermediate first audio data by reconstructing the first audio data using one of the following methods: a machine learning model, a constructed filter, a harmonic correction model, or sparse matrix techniques. Processing device 122 can reconstruct the first audio data using another method, such as a machine learning model, a constructed filter, a harmonic correction model, or sparse matrix techniques, to generate another intermediate first audio data. Processing device 122 can also generate reconstructed first audio data by averaging the intermediate first audio data and the other intermediate first audio data. Alternatively, processing device 122 can reconstruct the first audio data using two or more methods, such as a machine learning model, a constructed filter, a harmonic correction model, or sparse matrix techniques, to generate multiple intermediate first audio data. Processing device 122 can then generate reconstructed first audio data by averaging these multiple intermediate first audio data.

[0075] In some embodiments, the processing device 122 may reconstruct first audio data (or normalized first audio data) using a trained machine learning model to obtain reconstructed first audio data. Frequency components in the reconstructed first audio data that are above a certain frequency point (e.g., 2000Hz, 3000Hz) or within a certain frequency range (e.g., 2000Hz to 20kHz, 3000Hz to 20kHz, etc.) are increased relative to frequency components in the first audio data that are above that frequency point (e.g., 2000Hz, 3000Hz) or within that frequency range (e.g., 2000Hz to 20kHz, 3000Hz to 20kHz, etc.). The trained machine learning model may be constructed based on a deep learning model, a conventional machine learning model, or any combination thereof. Exemplary deep learning models may include convolutional neural network (CNN) models, recurrent neural network (RNN) models, long short-term memory network (LSTM) models, etc. Exemplary conventional machine learning models may include hidden Markov models (HMM), multilayer perceptron (MLP) models, etc.

[0076] In some embodiments, a primary machine learning model can be trained using multiple sets of training data to determine the trained machine learning model. Each set of training data may include bone conduction audio data and air conduction audio data. A set of training data may also be referred to as a speech sample. During the training of the primary machine learning model, the bone conduction audio data in the speech sample may be the input to the primary machine learning model, and the air conduction audio data in the speech sample corresponding to the bone conduction audio data may be the expected output of the primary machine learning model. The bone conduction audio data and air conduction audio data in the speech sample may represent the same speech and are simultaneously acquired by bone conduction sensors and air conduction sensors in a noise-free environment. As used herein, a noise-free environment may refer to an environment in which one or more noise assessment parameters (e.g., noise standard curve, statistical noise level, etc.) satisfy certain conditions, such as being less than a certain threshold. The trained machine learning model may be configured to provide a correspondence between bone conduction audio data (e.g., first audio data) and reconstructed bone conduction audio data (e.g., equivalent air conduction audio data). The trained machine learning model may reconstruct bone conduction audio data based on the correspondence. In some embodiments, bone conduction audio data from multiple sets of training data can be acquired by bone conduction sensors placed on the same part of the user's (e.g., the area around the ear) body. In some embodiments, the body part where the bone conduction sensor acquires bone conduction audio data for training a machine learning model is located is consistent with and / or the same as the body part where the bone conduction sensor acquires bone conduction audio data (e.g., first audio data) reconstructed using the trained machine learning model is located. For example, the body part where the bone conduction sensor acquires bone conduction audio data in each set of training data for training a machine learning model can be consistent with and / or the same as the body part where the bone conduction sensor acquires the first audio data is located. For another example, if the body part where the bone conduction sensor acquires the first audio data is the neck, the body part where the bone conduction sensor acquires bone conduction audio data for training the machine learning model is also the neck. The body part where the bone conduction sensor used to acquire multiple sets of training data is placed on the user's (e.g., the test subject) body affects the correspondence between bone conduction audio data (e.g., first audio data) and reconstructed bone conduction audio data (e.g., equivalent air conduction audio data). Therefore, reconstructing bone conduction audio data based on correspondences using a trained machine learning model can affect the reconstructed bone conduction audio data. Multiple sets of training data collected by bone conduction sensors located at different parts of the user's body can generate different correspondences between bone conduction audio data (e.g., first audio data) and reconstructed bone conduction audio data (e.g., equivalent air conduction audio data). For example, multiple bone conduction sensors with the same configuration can be located at different parts of the body, such as the mastoid process, temples, top of the head, external auditory canal, etc. Multiple bone conduction sensors can simultaneously collect bone conduction audio data generated when the user speaks.Multiple training sets can be formed based on bone conduction audio data acquired by multiple bone conduction sensors. Each training set may include multiple sets of training data acquired by one of the multiple bone conduction sensors and an air conduction sensor. Each set of training data may include bone conduction audio data and air conduction audio data representing the same speech. Each training set can be used to train a machine learning model to obtain a trained machine learning model. Multiple trained machine learning models can be obtained based on the multiple training sets. The multiple trained machine learning models can provide different correspondences between specific bone conduction audio data and reconstructed bone conduction audio data. For example, the same bone conduction audio data can be input into multiple trained machine learning models to generate different reconstructed bone conduction audio data. In some embodiments, the bone conduction audio data (e.g., frequency response curves, signal strength, acoustic features, etc.) acquired by bone conduction sensors with different configurations may be different. Therefore, the bone conduction sensor that acquires the bone conduction audio data used to train the machine learning model may be configured identically to the bone conduction sensor that acquires the bone conduction audio data (e.g., first audio data) to be reconstructed using the trained machine learning model. In some embodiments, different pressures applied to a part of the user's body result in different acquired bone conduction audio data (e.g., frequency response curves). Therefore, the pressure used to acquire bone conduction audio data for training a machine learning model can be the same as the pressure used to acquire bone conduction audio data reconstructed using the trained machine learning model (e.g., first audio data). Further description regarding the determination of the trained machine learning model and / or the reconstructed bone conduction audio data can be found in this document. Figure 5 And its description of bone conduction audio data.

[0077] In some embodiments, processing device 122 (e.g., preprocessing module 220) can reconstruct first audio data (or normalized first audio data) using a constructed filter to obtain reconstructed bone conduction audio data. The constructed filter can be configured to provide a relationship between specific air conduction audio data and specific bone conduction audio data corresponding to the specific air conduction audio data. As used herein, corresponding bone conduction audio data and air conduction audio data can refer to bone conduction audio data and air conduction audio data representing the same speech of the same user. Specific air conduction audio data can also be referred to as equivalent air conduction audio data corresponding to the specific bone conduction audio data or reconstructed bone conduction audio data. The air conduction audio data contains more frequency components above a certain frequency point (e.g., 2000Hz, 3000Hz) or within a certain frequency range (e.g., 2000Hz to 20kHz, 3000Hz to 20kHz, etc.) than the bone conduction audio data contains more frequency components above that frequency point (e.g., 2000Hz, 3000Hz) or within that frequency range (e.g., 2000Hz to 20kHz, 3000Hz to 20kHz, etc.). The processing device 122 can convert specific bone conduction audio data into specific air conduction audio data based on this relationship. For example, the processing device 122 can use a constructed filter to convert first audio data into reconstructed first audio data to obtain reconstructed first audio data. In some embodiments, the bone conduction audio data in a speech sample can be represented as d(n), and the corresponding air conduction audio data in the speech sample can be represented as s(n). The bone conduction audio data d(n) and the corresponding air conduction audio data s(n) can be determined based on the initial sound excitation signal e(n) through the bone conduction system and the air conduction system, respectively. The bone conduction system and the air conduction system can be equivalent to filter B and filter V, respectively. Then the constructed filter can be equivalent to filter H. Filter H can be determined according to the following equation (2):

[0078]

[0079] In some embodiments, long-time spectral techniques, for example, can be used to determine the constructed filter. For example, processing device 122 can determine the constructed filter according to equation (3) as shown below:

[0080]

[0081] in, This refers to the filter constructed in the frequency domain. This refers to the long-time spectrum expression corresponding to the air conduction audio data s(n). This refers to the long-time spectrum expression corresponding to the bone conduction audio data d(n). In some embodiments, the processing device 122 can acquire one or more sets of bone conduction audio data and air conduction audio data (also known as speech samples), the bone conduction audio data and air conduction audio data in each set being collected by the bone conduction sensor and air conduction sensor respectively when an operator (e.g., a tester) speaks in a noise-free environment. The processing device 122 can determine the constructed filter based on one or more sets of bone conduction audio data and air conduction audio data according to equation (3). For example, the processing device 122 can construct candidate filters based on the corresponding bone conduction audio data and air conduction audio data in each set according to equation (3). The processing device 122 can determine the constructed filter based on the candidate filters. In some embodiments, the processing device 122 can perform an inverse Fourier transform (IFT) (e.g., fast IFT) operation on the initial filter H(f) to obtain the constructed filter in the time domain.

[0082] In some embodiments, the body part where the bone conduction sensor that acquires bone conduction audio data for determining the constructed filter is located is the same as the body part where the bone conduction sensor that acquires bone conduction audio data for reconstruction using the constructed filter is located. For example, the body part where the bone conduction sensor that acquires bone conduction audio data for determining the constructed filter is located may be the same as the body part where the bone conduction sensor that acquires the first audio data is located. As another example, if the body part where the bone conduction sensor that acquires the first audio data is located is the neck, the body part where the bone conduction sensor that acquires bone conduction audio data for determining the constructed filter is also located is the neck. Multiple sets of training data acquired by bone conduction sensors located at different parts of the body can generate different filters. For example, a first set of bone conduction audio data and corresponding air conduction audio data can be acquired while the user is speaking, acquired by bone conduction sensors and air conduction sensors located at a first part of the user's body, respectively. A second set of bone conduction audio data and corresponding air conduction audio data can be acquired while the user is speaking, acquired by bone conduction sensors and air conduction sensors located at a second part of the user's body, respectively. A first filter can be determined based on the first set of bone conduction audio data and the corresponding air conduction audio data. The second filter can be determined based on the bone conduction audio data and the corresponding air conduction audio data from the second set. The first filter and the second filter are different; that is, the correspondence between the bone conduction audio data and the air conduction audio data provided by the first filter and the second filter is different.

[0083] In some embodiments, processing device 122 (e.g., preprocessing module 220) can reconstruct first audio data (or normalized first audio data) using a harmonic correction model to obtain reconstructed first audio data. The harmonic correction model can be configured to provide a relationship between the amplitude spectrum of a specific air conduction audio data and the amplitude spectrum of a specific bone conduction audio data corresponding to that specific air conduction audio data. As used herein, the specific air conduction audio data can also be referred to as equivalent air conduction audio data or reconstructed bone conduction audio data corresponding to the specific bone conduction audio data. The amplitude spectrum of the specific air conduction audio data can also be referred to as the corrected amplitude spectrum of the specific bone conduction audio data. Processing device 122 can determine the amplitude spectrum and phase spectrum of the first audio data (or normalized first audio data) in the frequency domain. Processing device 122 can use the harmonic correction model to correct the amplitude spectrum of the first audio data (or normalized first audio data) to obtain the corrected amplitude spectrum of the first audio data (or normalized first audio data). Then, processing device 122 can determine the reconstructed first audio data based on the corrected amplitude spectrum and the phase spectrum of the first audio data (or normalized first audio data). Further description of reconstructing the first audio data using the harmonic correction model can be found elsewhere in this application (e.g., Figure 6 (and its description).

[0084] In some embodiments, processing device 122 (e.g., preprocessing module 220) can reconstruct first audio data (or normalized first audio data) using sparse matrix techniques to obtain reconstructed first audio data. For example, processing device 122 can acquire a first transformation relation configured to convert a dictionary matrix of initial bone conduction audio data (e.g., first audio data) into a dictionary matrix of reconstructed bone conduction audio data (e.g., reconstructed first audio data) corresponding to the initial bone conduction audio data. Processing device 122 can acquire a second transformation relation configured to convert a sparse code matrix of initial bone conduction audio data into a sparse code matrix of reconstructed bone conduction audio data corresponding to the initial bone conduction audio data. Processing device 122 can use the first transformation relation to determine the dictionary matrix of the reconstructed first audio data based on the dictionary matrix of the first audio data. Processing device 122 can use the second transformation relation to determine the sparse code matrix of the reconstructed first audio data based on the sparse code matrix of the first audio data. Processing device 122 can determine the reconstructed first audio data based on the determined dictionary matrix and sparse code matrix of the reconstructed first audio data. In some embodiments, the first transformation relation and / or the second transformation relation can be default settings of the audio signal generation system 100. In some embodiments, the processing device 122 may determine a first transformation relationship and / or a second transformation relationship based on one or more sets of corresponding bone conduction audio data sets and air conduction audio data sets. Further description of reconstructing the first audio data using sparse matrix techniques can be found elsewhere in this application (e.g., Figure 7 (and its description).

[0085] In 440, processing device 122 (e.g., audio data generation module 230) can generate third audio data based on first audio data (or preprocessed first audio data) and second audio data (or preprocessed second audio data). The frequency components in the third audio data above a certain frequency point (or threshold) are increased relative to the frequency components above that frequency point (or threshold) in the first audio data (or preprocessed first audio data). In other words, the frequency components above that frequency point (or threshold) in the third audio data can be more numerous than the frequency components above that frequency point (or threshold) in the first audio data (or preprocessed first audio data). In some embodiments, the noise level associated with the third audio data can be lower than the noise level associated with the second audio data (or preprocessed second audio data). As used herein, an increase in frequency components in the third audio data above a certain frequency point (or threshold) relative to frequency components in the first audio data (or preprocessed first audio data) above that frequency point can refer to a greater count or number of waves (e.g., sine waves or harmonics) in the third audio data with frequencies above that frequency point than the count or number of waves (e.g., sine waves or harmonics) in the first audio data with frequencies above that frequency point. In some embodiments, the frequency point can be a constant in the range of 20 Hz to 20 kHz. For example, the frequency point can be 2000 Hz, 3000 Hz, 4000 Hz, 5000 Hz, 4000 Hz, etc. In some embodiments, the frequency point can be the frequency value of the frequency components in the third audio data and / or the first audio data.

[0086] In some embodiments, the processing device 122 may generate third audio data based on first audio data (or preprocessed first audio data) and second audio data (or preprocessed second audio data) according to one or more frequency thresholds. For example, the processing device 122 may determine one or more frequency thresholds at least in part based on the first audio data (or preprocessed first audio data) and / or the second audio data (or preprocessed second audio data). The processing device 122 may divide the first audio data (or preprocessed first audio data) and the second audio data (or preprocessed second audio data) into multiple segments according to one or more frequency thresholds. The processing device 122 may determine the weight of each segment in the multiple segments of the first audio data (or preprocessed first audio data) and the second audio data (or preprocessed second audio data). The processing device 122 may then determine the third audio data based on the weight of each segment in the multiple segments of the first audio data (or preprocessed first audio data) and the second audio data (or preprocessed second audio data).

[0087] In some embodiments, the processing device 122 may determine a single frequency threshold. The processing device 122 may concatenate first audio data (or pre-processed first audio data) and second audio data (or pre-processed second audio data) in the frequency domain based on the single frequency threshold to generate third audio data. For example, the processing device 122 may use a first filter to determine low-frequency components in the first audio data (or pre-processed first audio data) that include frequency components below the single frequency threshold. The processing device 122 may use a second filter to determine high-frequency components in the second audio data (or pre-processed second audio data) that include frequency components above the single frequency threshold. The processing device 122 may concatenate and / or combine the low-frequency components of the first audio data (or pre-processed first audio data) and the high-frequency components of the second audio data (or pre-processed second audio data) to generate third audio data. In some embodiments, the first filter may be a low-pass filter with a single frequency threshold as the cutoff frequency, which may allow frequency components in the first audio data below the single frequency threshold to pass through. The second filter may be a high-pass filter with a single frequency threshold as the cutoff frequency, which may allow frequency components in the second audio data above the single frequency threshold to pass through. In some embodiments, the processing device 122 may determine a single frequency threshold based at least in part on first audio data (or preprocessed first audio data) and / or second audio data (or preprocessed second audio data). Further description of determining the single frequency threshold can be found in [reference needed]. Figure 8 It was found in its description.

[0088] In some embodiments, the processing device 122 may determine, at least partially, a first weight and a second weight for the low-frequency portion and the high-frequency portion of the first audio data (or preprocessed first audio data) based on a single frequency threshold. The processing device 122 may also determine, at least partially, a third weight and a fourth weight for the low-frequency portion and the high-frequency portion (or preprocessed second audio data) of the second audio data (or preprocessed second audio data) based on a single frequency threshold. In some embodiments, the processing device 122 may use the first weight, the second weight, the third weight, and the fourth weight to weight the low-frequency portion (or preprocessed first audio data), the high-frequency portion (or preprocessed first audio data), the low-frequency portion (or preprocessed second audio data), and the high-frequency portion (or preprocessed second audio data) of the second audio data (or preprocessed second audio data), respectively, to determine the third audio data. Further description of determining the third audio data (or spliced ​​audio data) can be found in [the relevant documentation / reference]. Figure 8 It was found in its description.

[0089] In some embodiments, the processing device 122 may determine weights corresponding to the first audio data (or preprocessed first audio data) and the second audio data (or preprocessed second audio data), respectively, based at least in part on the first audio data (or preprocessed first audio data) and / or the second audio data (or preprocessed second audio data). The processing device 122 may use the weights corresponding to the first audio data (or preprocessed first audio data) and the weights corresponding to the second audio data (or preprocessed second audio data) to determine third audio data by weighting the first audio data (or preprocessed first audio data) and the second audio data (or preprocessed second audio data). Further description of determining the third audio data can be found elsewhere in this application (e.g., Figure 9 (and its description).

[0090] In 450, processing device 122 (e.g., audio data generation module 230) can determine target audio data representing a user's speech based on third audio data, the target audio data having a higher fidelity than the first and second audio data. The target audio data can represent the user's speech represented by the first and second audio data. As used herein, fidelity can be used to represent the similarity between output audio data (e.g., target audio data, first audio data, second audio data) and original input audio data (e.g., the user's speech). Fidelity can also represent the intelligibility of the output audio data (e.g., target audio data, first audio data, second audio data).

[0091] In some embodiments, the processing device 122 may designate third audio data as target audio data. In some embodiments, the processing device 122 may perform post-processing operations on the third audio data to obtain the target audio data. In some embodiments, the post-processing operations may include noise reduction operations, domain transformation operations (e.g., Fourier transform (FT) operations), or combinations thereof. In some embodiments, the noise reduction operation performed on the third audio data may include using a Wiener filter, spectral subtraction, an adaptive algorithm, a minimum mean square error (MMSE) estimation algorithm, or any combination thereof. In some embodiments, the noise reduction operation performed on the third audio data may be the same as or different from the noise reduction operation performed on the second audio data. For example, both the noise reduction operation performed on the second audio data and the noise reduction operation performed on the third audio data may use spectral subtraction. For another example, the noise reduction operation performed on the second audio data may use a Wiener filter, and the noise reduction operation performed on the third audio data may use spectral subtraction. In some embodiments, the processing device 122 may perform an inverse Fourier operation on the third audio data in the frequency domain to obtain the target audio data in the time domain.

[0092] In some embodiments, processing device 122 may transmit signals via network 150 to a client terminal (e.g., terminal 130), storage device 140, and / or any other storage device (not shown in the audio signal generation system 100). The signal may include target audio data. The signal may also be configured to instruct the client terminal to play the target audio data.

[0093] It should be noted that the foregoing is provided for illustrative purposes only and is not intended to limit the scope of this application. Various changes and modifications can be made by those skilled in the art based on the description in this application. However, these changes and modifications will not depart from the scope of this application. For example, operation 450 can be omitted. As another example, operations 410 and 420 can be integrated into a single operation.

[0094] Figure 5 This is a flowchart illustrating an exemplary process for reconstructing bone conduction audio data using a trained machine learning model, according to some embodiments of this application. The operation of the process shown below is for illustrative purposes only. In some embodiments, process 500 may be accomplished using one or more additional operations not described, and / or without the one or more operations discussed. Additionally, Figure 5 The order of operations of process 500 shown and described below is not limiting. In some embodiments, one or more operations of process 500 may be performed to achieve, for example... Figure 4 At least a portion of the described operation 430.

[0095] In 510, processing device 122 (e.g., acquisition module 210) can acquire bone conduction audio data. In some embodiments, bone conduction audio data may be raw audio data (e.g., first audio data) acquired by a bone conduction sensor when the user speaks. Figure 1 (and its description), as described elsewhere in this application. For example, a user's voice may be acquired by a bone conduction sensor (e.g., bone conduction microphone 112) to generate an electrical signal (e.g., an analog signal or a digital signal) (i.e., bone conduction audio data). The bone conduction sensor may transmit the electrical signal via network 150 to server 120, terminal 130, and / or storage device 140. In some embodiments, the bone conduction audio data includes acoustic features and / or semantic information that may reflect the content of the user's voice. Exemplary acoustic characteristics may include features associated with duration, features associated with energy, features associated with the fundamental frequency, features associated with the frequency spectrum, features associated with the phase spectrum, etc., as described elsewhere in this application (e.g., Figure 4 (and its description).

[0096] In step 520, processing device 122 (e.g., acquisition module 210) can obtain a trained machine learning model. The trained machine learning model can be provided by training a primary machine learning model using multiple sets of training data. In some embodiments, the trained machine learning model can be used to process specific bone conduction audio data to obtain processed bone conduction audio data. The processed bone conduction audio data can also be considered reconstructed bone conduction audio data. Frequency components of bone conduction audio data above a certain frequency threshold (e.g., 800Hz, 2000Hz, 3000Hz, 4000Hz, etc.) in the processed bone conduction audio data will increase relative to frequency components of bone conduction audio data above that frequency threshold or frequency point (e.g., 800Hz, 2000Hz, 3000Hz, 4000Hz, etc.) in the specific bone conduction audio data. Processed bone conduction audio data can be similar to or identical to ideal air conduction audio data with little or no noise, and the processed bone conduction audio data and unprocessed specific bone conduction audio data represent the same speech of the same user, wherein the ideal air conduction audio data is acquired by an air conduction sensor at the same time as the specific bone conduction audio data is acquired by the bone conduction sensor. As used herein, similarity or identity between processed bone conduction audio data and ideal air conduction audio data with little or no noise can mean that the similarity between the acoustic features of the processed bone conduction audio data and the acoustic features of the ideal air conduction audio data is greater than a certain threshold (e.g., 0.9, 0.8, 0.7, etc.). For example, in a noise-free environment, when a user speaks, bone conduction audio data and air conduction audio data are acquired simultaneously through bone conduction microphone 112 and air conduction microphone 114, respectively. Processed bone conduction audio data is generated by a trained machine learning model that processes the bone conduction audio data, and the processed bone conduction audio data has the same or similar acoustic features as the air conduction audio data acquired by the corresponding air conduction microphone 114. In some embodiments, the processing device 122 may obtain a trained machine learning model from the terminal 130, the storage device 140, or any other storage device.

[0097] In some embodiments, the primary machine learning model can be constructed based on deep learning models, traditional machine learning models, or any combination thereof. Deep learning models can include convolutional neural network (CNN) models, recurrent neural network (RNN) models, long short-term memory network (LSTM) models, or any combination thereof. Traditional machine learning models can include hidden Markov models (HMM), multilayer perceptron (MLP) models, or any combination thereof. In some embodiments, the primary machine learning model can include multiple layers, such as an input layer, multiple hidden layers, and an output layer. Multiple hidden layers can include one or more convolutional layers, one or more pooling layers, one or more batch normalization layers, one or more activation layers, one or more fully connected layers, loss function layers, etc. Each layer can include multiple nodes. In some embodiments, the primary machine learning model can be defined by at least two structural parameters and at least two learning parameters (or training parameters). The structural parameters can be changed by training the primary machine learning model using at least two sets of training data. Users can set and / or adjust the structural parameters before training the primary machine learning model. Exemplary structural parameters of a machine learning model may include the size of the layer kernels, the total number (or number of) layers, the number (or number of) nodes in each layer, the learning rate, batch size, stride, etc. For example, if a primary machine learning model includes a Long Short-Term Memory (LSTM) model, the LSM model may include an input layer with two nodes, four hidden layers, and an output layer with two nodes, each hidden layer comprising 30 nodes. The time-shifting step of the LSM model may be 65, and the learning rate may be 0.003. Exemplary learning parameters of a machine learning model may include connection weights between two connected nodes, node-associated bias vectors, etc. The connection weights between two connected nodes can be configured to represent the proportion of a node's output value as the input value of another connected node. The node-associated bias vector can be configured to control the output value of nodes deviating from the origin.

[0098] In some embodiments, a machine learning model training algorithm can be used to train a primary machine learning model using multiple sets of training data to determine the trained machine learning model. In some embodiments, one or more sets of training data can be acquired in a noise-free environment, such as in an anechoic chamber. A set of training data may include specific bone conduction audio data and corresponding specific air conduction audio data. The specific bone conduction audio data and corresponding specific air conduction audio data in a set of training data can be obtained from a specific user simultaneously via a bone conduction sensor (e.g., bone conduction microphone 112) and an air conduction sensor (e.g., air conduction microphone 114). In some embodiments, each set of training data in at least some of the multiple sets of training data may include specific bone conduction audio data and corresponding reconstructed bone conduction audio data, the reconstructed bone conduction audio data being generated by reconstructing the specific bone conduction audio data using one or more reconstruction techniques as described elsewhere in this application. Exemplary machine learning model training algorithms may include gradient descent algorithms, Newton's algorithm, quasi-Newton algorithms, Levenberg-Marquardt algorithms, conjugate gradient algorithms, etc., or combinations thereof. A trained machine learning model can be configured to provide a correspondence between bone conduction audio data (e.g., first audio data) and reconstructed bone conduction audio data (e.g., equivalent air conduction audio data). The trained machine learning model can reconstruct bone conduction audio data based on the correspondence. In some embodiments, bone conduction audio data from multiple sets of training data can be acquired by bone conduction sensors placed at the same location on the user's (e.g., test subject's) body (e.g., the area around the ear). In some embodiments, the body location of the bone conduction sensor acquiring the bone conduction audio data used to train the machine learning model can be the same as the body location of the bone conduction sensor acquiring the bone conduction sensor data (e.g., first audio data) to be reconstructed using the trained machine learning model. For example, the body location of the bone conduction sensor acquiring the bone conduction audio data in each set of training data used to train the machine learning model can be the same as the body location of the bone conduction sensor acquiring the first audio data. As another example, if the body location of the bone conduction sensor acquiring the first audio data is the neck, the body location of the bone conduction sensor acquiring the bone conduction audio data used to train the machine learning model is also the neck.

[0099] Bone conduction sensors used to collect multiple sets of training data are placed at different locations on the user's (e.g., test subject's) body, influencing the correspondence between bone conduction audio data (e.g., first audio data) and reconstructed bone conduction audio data (e.g., equivalent air conduction audio data). Therefore, reconstructing bone conduction audio data using a trained machine learning model affects the reconstructed bone conduction audio data generated based on the correspondence. Multiple sets of training data collected by bone conduction sensors located at different locations on the user's body can generate different correspondences between bone conduction audio data (e.g., first audio data) and reconstructed bone conduction audio data (e.g., equivalent air conduction audio data). For example, multiple bone conduction sensors with the same configuration can be located at different locations on the body, such as the mastoid process, temples, top of the head, external auditory canal, etc. Multiple bone conduction sensors can simultaneously collect bone conduction audio data generated when the user speaks. Multiple training sets can be formed based on the bone conduction audio data collected by multiple bone conduction sensors. Each training set in the multiple training sets may include multiple sets of training data collected by one of the multiple bone conduction sensors and an air conduction sensor. Each set of training data may include bone conduction audio data and air conduction audio data representing the same speech. Each training set in the multiple training sets can be used to train a machine learning model to obtain a trained machine learning model. Multiple trained machine learning models can be obtained based on multiple training sets. Multiple trained machine learning models can provide different correspondences between specific bone conduction audio data and reconstructed bone conduction audio data. For example, the same bone conduction audio data can be input into multiple trained machine learning models to generate different reconstructed bone conduction audio data. In some embodiments, bone conduction audio data (e.g., frequency response curves, signal strength, acoustic features, etc.) acquired by bone conduction sensors with different configurations can be different. Therefore, the bone conduction sensor used to acquire bone conduction audio data for training the machine learning model can be configured identically to the bone conduction sensor used to acquire bone conduction audio data (e.g., first audio data) to be reconstructed using the trained machine learning model. In some embodiments, applying different pressures (e.g., 0N to 1N or 0N to 0.8N) to a part of the user's body by the bone conduction sensor will result in different acquired bone conduction audio data (e.g., frequency response curves). Therefore, the pressure of acquiring bone conduction audio data for training a machine learning model can be the same as the pressure of acquiring reconstructed bone conduction audio data (e.g., first audio data) to be used with the trained machine learning model.

[0100] In some embodiments, the trained machine learning model can be obtained by performing at least two iterations to update one or more learning parameters of the primary machine learning model. For each of the at least two iterations, a specific set of training data can be input into the primary machine learning model. For example, specific bone conduction audio data from the specific training data set can be input into the input layer of the primary machine learning model, and specific air conduction audio data from the specific training data set can be input into the output layer of the primary machine learning model as the expected output of the primary machine learning model corresponding to the specific bone conduction audio data (i.e., the input). The primary machine learning model can extract one or more acoustic features (e.g., duration features, amplitude features, fundamental frequency features, etc.) from the specific bone conduction audio data and specific air conduction audio data in the specific training data set. Based on the extracted features, the primary machine learning model can determine a predicted output corresponding to the specific bone conduction audio data (i.e., the input). The predicted output corresponding to the specific bone conduction audio data is then compared with the expected output of the output layer (i.e., the specific air conduction audio data input) based on a cost function. The cost function of the primary machine learning model can be configured to evaluate the difference between the estimated value (e.g., the predicted output) of the primary machine learning model and the actual value (e.g., the expected output or the specific air conduction audio data input). If the value of the cost function exceeds a threshold in the current iteration, the learning parameters of the primary machine learning model can be adjusted and updated so that the value of the cost function (i.e., the difference between the predicted output and the specific aerodynamic audio data input) is less than the threshold. Therefore, in the next iteration, another set of training data can be input into the primary machine learning model to train it as described above. At least two iterations can then be performed to update the learning parameters of the primary machine learning model until a termination condition is met. The termination condition can indicate whether the primary machine learning model has been sufficiently trained. For example, the termination condition can be met if the value of the cost function associated with the primary machine learning model is minimum or less than a threshold (e.g., a constant). As another example, the termination condition can be met if the value of the cost function converges. The cost function can be considered to have converged if the change in the value of the cost function is less than a threshold (e.g., a constant) over two or more consecutive iterations. As yet another example, the termination condition can be met when a specified number of iterations are performed during training. The trained machine learning model can be determined based on the updated learning parameters. In some embodiments, the trained machine learning model can be sent to storage device 140 / storage module 240 or any other storage device for storage.

[0101] In step 530, processing device 122 (e.g., preprocessing module 220) can process bone conduction audio data using a trained machine learning model to obtain reconstructed bone conduction audio data. In some embodiments, processing device 122 can input bone conduction audio data into a trained machine learning model, which can then output processed bone conduction audio data. In some embodiments, processing device 122 can extract acoustic features from the bone conduction audio data and input the extracted acoustic features into a trained machine learning model. The trained machine learning model can output processed bone conduction audio data. The frequency components of bone conduction audio data above a frequency threshold or frequency point (e.g., 800Hz, 2000Hz, 3000Hz, etc.) in the processed bone conduction audio data are increased compared to the frequency components of bone conduction audio data above that frequency threshold or frequency point in the unprocessed bone conduction audio data. In some embodiments, processing device 122 can send the processed bone conduction audio data to a client terminal (e.g., terminal 130). The client terminal (e.g., terminal 130) can convert the processed bone conduction audio data into speech and play the speech to the user.

[0102] It should be noted that the foregoing is provided for illustrative purposes only and is not intended to limit the scope of this application. Various changes and modifications can be made by those skilled in the art based on the description herein. However, such changes and modifications will not depart from the scope of this application.

[0103] Figure 6 This is a flowchart illustrating an exemplary process for reconstructing bone conduction audio data based on a harmonic correction model, according to some embodiments of this application. The operation of the process shown below is for illustrative purposes only. In some embodiments, process 600 may be accomplished using one or more additional operations not described, and / or without the one or more operations discussed. Additionally, Figure 6 The order of operations of process 600 shown and described below is not limiting. In some embodiments, one or more operations of process 600 may be performed to achieve, as in combination Figure 4 At least a portion of the described operation 430.

[0104] In 610, processing device 122 (e.g., acquisition module 210) can acquire bone conduction audio data. In some embodiments, as described in conjunction with operation 410, the bone conduction audio data can be raw audio data (e.g., first audio data) acquired by a bone conduction sensor when the user speaks. For example, the user's voice can be acquired by a bone conduction sensor (e.g., bone conduction microphone 112) to generate an electrical signal (e.g., an analog or digital signal) (i.e., bone conduction audio data). In some embodiments, the bone conduction audio data can include multiple waves with different frequencies and amplitudes. Bone conduction audio data in the frequency domain can be represented as a matrix comprising multiple elements. Each of the multiple elements can represent the frequency and amplitude of the wave.

[0105] In 620, processing device 122 (e.g., preprocessing module 220) can determine the amplitude spectrum and phase spectrum of the bone conduction audio data. In some embodiments, processing device 122 can determine the amplitude spectrum and phase spectrum of the bone conduction audio data by performing a Fourier transform (FT) operation on the bone conduction audio data. Processing device 122 can determine the amplitude spectrum and phase spectrum of the bone conduction audio data in the frequency domain. For example, processing device 122 can utilize peak detection techniques, including but not limited to the Spectral Envelope Estimation Vocoder (SEEVOC) algorithm, to detect the peak values ​​of the waves in the bone conduction audio data. Processing device 122 can determine the amplitude spectrum and phase spectrum based on the peak values ​​of the waves. For example, the amplitude of the wave is half the distance between the peak and the trough.

[0106] In 630, processing device 122 (e.g., preprocessing module 220) can obtain a harmonic correction model. The harmonic correction model can be configured to provide a relationship between the amplitude spectrum of a specific air conduction audio data and the amplitude spectrum of a specific bone conduction audio data corresponding to that specific air conduction audio data. The amplitude spectrum of the specific air conduction audio data corresponding to the specific bone conduction audio data can be determined based on said relationship and the amplitude spectrum of the specific bone conduction audio data. As used herein, the specific air conduction audio data can also be referred to as equivalent air conduction audio data, bone conduction audio data, or reconstructed bone conduction audio data corresponding to the specific bone conduction audio data.

[0107] In some embodiments, the harmonic correction model may be the default setting of the audio signal generation system 100. In some embodiments, the processing device 122 may obtain the harmonic correction model from the storage device 140, the storage module 240, or any other storage device. In some embodiments, the harmonic correction model may be determined based on one or more sets of bone conduction audio data and corresponding air conduction audio data. The bone conduction audio data and corresponding air conduction audio data in each set may be simultaneously acquired by a bone conduction sensor and an air conduction sensor by an operator (e.g., a tester) speaking in a noise-free environment. The bone conduction sensor and air conduction sensor may be the same as or different from the bone conduction sensor used to acquire the first audio data and the air conduction sensor used to acquire the second audio data. In some embodiments, the harmonic correction model may be determined based on one or more sets of bone conduction audio data and corresponding air conduction audio data according to operations a1 to a3. In operation a1, processing device 122 can use peak detection techniques (e.g., the Spectral Envelope Estimation Vocoder Algorithm (SEEVOC)) to determine the amplitude spectrum of the bone conduction audio data in each group and the amplitude spectrum of the corresponding air conduction audio data in each group. In operation a2, processing device 122 can determine a candidate correction matrix based on the amplitude spectra of the bone conduction audio data and the corresponding air conduction audio data in each group. For example, processing device 122 can determine the candidate correction matrix based on the ratio of the amplitude spectrum of the air conduction audio data in each group to the amplitude spectrum of the corresponding bone-air conduction audio data. In operation a3, processing device 122 can determine a harmonic correction model based on the candidate correction matrix corresponding to each group of bone conduction audio data and the corresponding air conduction audio data in one or more groups. For example, processing device 122 can determine the average of the candidate correction matrices corresponding to one or more groups of bone conduction audio data and their corresponding air conduction audio data as the harmonic correction model.

[0108] In some embodiments, the body location of the bone conduction sensor that acquires bone conduction audio data for determining the harmonic correction model can be the same as and / or the same as the body location of the bone conduction sensor that acquires bone conduction audio data to be reconstructed using the harmonic correction model. For example, the body location of the bone conduction sensor that acquires bone conduction audio data for determining the harmonic correction model can be the same as the body location of the bone conduction sensor that acquires the first audio data. As another example, if the body location of the bone conduction sensor that acquires the first audio data is the neck, the body location of the bone conduction sensor that acquires bone conduction audio data for determining the harmonic correction model is also the neck. Multiple sets of data acquired by bone conduction sensors located at different parts of the user's body can generate different harmonic correction models. For example, a first set of bone conduction audio data and corresponding air conduction audio data can be acquired by a bone conduction sensor and an air conduction sensor located at a first part of the user's body while the user is speaking. A second set of bone conduction audio data and corresponding air conduction audio data can be acquired by a bone conduction sensor and an air conduction sensor located at a second part of the user's body while the user is speaking. A first harmonic correction model can be determined based on the first set of bone conduction audio data and the corresponding air conduction audio data. The second harmonic correction model can be determined based on the bone conduction audio data and the corresponding air conduction audio data from the second set. The first harmonic correction model and the second harmonic correction model are different. The correspondence between the amplitude spectrum of specific air conduction audio data provided by the first harmonic correction model and the amplitude spectrum of specific bone conduction audio data corresponding to specific air conduction audio data is different. The reconstructed bone conduction audio data obtained by reconstructing the same bone conduction audio data based on the first harmonic correction model and the second harmonic correction model are different.

[0109] In 640, processing device 122 (e.g., preprocessing module 220) can correct the amplitude spectrum of the bone conduction audio data to obtain a corrected amplitude spectrum of the bone conduction audio data. In some embodiments, the harmonic correction model may include a correction matrix that includes the amplitude spectrum of the bone conduction audio data (e.g., Figure 4 The weighting coefficients corresponding to each element in the amplitude spectrum of the first audio data described herein. As used herein, the elements in the amplitude spectrum may refer to the amplitude of a wave (i.e., a frequency component). The processing device 122 can process the correction matrix by comparing it with the bone conduction audio data (e.g., ... Figure 4 The amplitude spectrum of the first audio data described herein is multiplied to correct the bone conduction audio data (e.g., Figure 4 The amplitude spectrum of the first audio data (or normalized first audio data) described herein is used to obtain bone conduction audio data (e.g., Figure 4 The corrected amplitude spectrum of the first audio data described in the text.

[0110] In 650, processing device 122 (e.g., preprocessing module 220) can determine reconstructed bone conduction audio data based on the corrected amplitude spectrum and the phase spectrum of the bone conduction audio data. In some embodiments, processing device 122 can perform an inverse Fourier transform on the corrected amplitude spectrum and the phase spectrum of the bone conduction audio data to obtain the reconstructed bone conduction audio data.

[0111] It should be noted that the foregoing is provided for illustrative purposes only and is not intended to limit the scope of this application. Various changes and modifications can be made by those skilled in the art based on the description herein. However, such changes and modifications will not depart from the scope of this application.

[0112] Figure 7 This is a flowchart illustrating an exemplary process for reconstructing bone conduction audio data based on sparse matrix techniques according to some embodiments of this application. The operation of the process shown below is for illustrative purposes only. In some embodiments, process 700 may be accomplished using one or more additional operations not described, and / or without one or more operations discussed. Additionally, Figure 7 The order of operations of process 700 shown and described below is not limiting. In some embodiments, one or more operations of process 700 may be performed to achieve, as in combination Figure 4 At least a portion of the described operation 430.

[0113] In 710, processing device 122 (e.g., acquisition module 210) can acquire bone conduction audio data. In some embodiments, as described in conjunction with operation 410, the bone conduction audio data can be raw audio data (e.g., first audio data) acquired by a bone conduction sensor when the user speaks. For example, the user's voice can be acquired by a bone conduction sensor (e.g., bone conduction microphone 112) to generate an electrical signal (e.g., an analog signal or a digital signal) (i.e., bone conduction audio data). In some embodiments, the bone conduction audio data can include multiple waves with different frequencies and amplitudes. The bone conduction audio data in the frequency domain can be represented as a matrix X. Matrix X can be determined based on a dictionary matrix D and a sparse code matrix C. For example, the audio data can be determined according to equation (4):

[0114] X≈DC(4).

[0115] At 720, processing device 122 (e.g., preprocessing module 220) can obtain a first transformation relation for converting the dictionary matrix of bone conduction audio data into a dictionary matrix of reconstructed bone conduction audio data corresponding to the bone conduction audio data. In some embodiments, the first transformation relation may be the default setting of the audio signal generation system 100. In some embodiments, processing device 122 may obtain the first transformation relation from storage device 140, storage module 240, or any other storage device. In some embodiments, the first transformation relation may be determined based on one or more sets of bone conduction audio data and corresponding air conduction audio data (i.e., speech samples). The bone conduction audio data and corresponding air conduction audio data in each set may be simultaneously acquired by bone conduction sensors and air conduction sensors, respectively, in a noise-free environment while an operator (e.g., a test subject) speaks. For example, processing device 122 may determine the dictionary matrix of bone conduction audio data and the corresponding air conduction audio data dictionary matrix in each set of data as described in operation 740. The processing device 122 can divide the dictionary matrix of air conduction audio data in each set of data by the dictionary matrix of the corresponding bone conduction audio data to obtain candidate first transformation relationships for one or more sets of bone conduction audio data and corresponding air conduction audio data. In some embodiments, the processing device 122 can determine multiple candidate first transformation relationships based on multiple sets of bone conduction audio data and corresponding air conduction audio data. The processing device 122 can average the multiple candidate first transformation relationships to obtain a first transformation relationship. In some embodiments, the processing device 122 can determine one of the multiple candidate first transformation relationships as the first transformation relationship.

[0116] At 730, the processing device 122 (e.g., preprocessing module 220) can obtain a second transformation relation for converting the sparse code matrix of the bone conduction audio data into a sparse code matrix of the reconstructed bone conduction audio data corresponding to the bone conduction audio data. In some embodiments, the second transformation relation may be a default setting of the audio signal generation system 100. In some embodiments, the processing device 122 may obtain the second transformation relation from the storage device 140, the storage module 240, or any other storage device. In some embodiments, the second transformation relation may be determined based on one or more sets of bone conduction audio data and corresponding air conduction audio data. For example, the processing device 122 may determine the sparse code matrix of the bone conduction audio data and the sparse code matrix of the corresponding air conduction audio data in each of one or more sets of data, as described in operation 740. The processing device 122 may obtain candidate second transformation relations by dividing the sparse code matrix of the air conduction audio data by the sparse code matrix of the corresponding bone conduction audio data. In some embodiments, the processing device 122 may determine one or more candidate second transformation relations based on one or more sets of bone conduction audio data and corresponding air conduction audio data. The processing device 122 can average one or more candidate second transformation relationships to obtain a second transformation relationship. In some embodiments, the processing device 122 can determine one of the one or more candidate second transformation relationships as the second transformation relationship.

[0117] In some embodiments, the body location of the bone conduction sensor that acquires bone conduction audio data used to determine a first transformation relationship (and / or a second transformation relationship) can be the same as the body location of the bone conduction sensor that acquires bone conduction audio data to be reconstructed using the first transformation relationship (and / or the second transformation relationship). For example, the body location of the bone conduction sensor that acquires bone conduction audio data used to determine a first transformation relationship (and / or a second transformation relationship) can be the same as the body location of the bone conduction sensor that acquires the first audio data. As another example, if the body location of the bone conduction sensor that acquires the first audio data is the neck, the body location of the bone conduction sensor that acquires bone conduction audio data used to determine the first transformation relationship (and / or the second transformation relationship) is also the neck. Different bone conduction audio data acquired by bone conduction sensors located at different parts of the user's body can generate different first transformation relationships (and / or second transformation relationships). Reconstructing the same bone conduction audio data based on different first transformation relationships (and / or second transformation relationships) can yield different reconstructed bone conduction audio data.

[0118] In 740, the processing device 122 (e.g., preprocessing module 220) can be based on bone conduction audio data (e.g., Figure 4 The dictionary matrix of the first audio data (or normalized first audio data) described in the document is used to determine the reconstructed bone conduction audio data (e.g., ...) using a first transformation relation. Figure 4 The dictionary matrix of the reconstructed first audio data described herein. For example, processing device 122 can combine the first transformation relation (e.g., in matrix form) with the bone conduction audio data (e.g., Figure 4 Multiply the dictionary matrix of the first audio data (or normalized first audio data) to obtain the reconstructed bone conduction audio data (e.g., Figure 4 The processing device 122 can determine the dictionary matrix and / or sparse code matrix of the audio data (e.g., bone conduction audio data (e.g., the first audio data), bone conduction audio data of speech samples, and / or air conduction audio data) by performing at least two iterations. Before performing at least two iterations, the processing device 122 can initialize the dictionary matrix of the audio data (e.g., the first audio data) to obtain an initial dictionary matrix. For example, the processing device 122 can set each element in the initial dictionary matrix to 0 or 1. In each iteration, the processing device 122 can determine the estimated sparse code matrix of the audio data (e.g., the first audio data) based on the audio data (e.g., the first audio data) and the initial dictionary matrix using, for example, an orthogonal matching pursuit (OMP) algorithm. The processing device 122 can determine the estimated dictionary matrix based on the audio data (e.g., the first audio data) and the estimated sparse code matrix using, for example, a K-singular value decomposition (K-SVD) algorithm. The processing device 122 can determine the estimated audio data based on the estimated dictionary matrix and the estimated sparse code matrix according to equation (4). Processing device 122 can compare the estimated audio data with the audio data (e.g., the first audio data). If the difference between the estimated audio data generated in the current iteration and the audio data (e.g., the first audio data) exceeds a threshold, processing device 122 can update the initial dictionary matrix using the estimated dictionary matrix generated in the current iteration. Processing device 122 can execute the next iteration based on the updated initial dictionary matrix until the difference between the estimated audio data generated in the current iteration and the audio data (e.g., the first audio data) is less than the threshold. If the difference between the estimated audio data generated in the current iteration and the audio data is less than the threshold, processing device 122 can designate the estimated dictionary matrix and the estimated sparse code matrix generated in the current iteration as the dictionary matrix and / or sparse code matrix of the audio data (e.g., the first audio data).

[0119] In 750, processing device 122 (e.g., preprocessing module 220) can use a second transformation relationship based on bone conduction audio data (e.g., Figure 4 The sparse code matrix of the first audio data (or normalized first audio data) described in the text determines the reconstructed bone-guided audio data (e.g., Figure 4The sparse code matrix of the reconstructed first audio data described in [the previous section]. For example, processing device 122 can multiply a second transformation relation (e.g., a matrix) with the sparse code matrix of the bone conduction audio data to obtain the sparse code matrix of the reconstructed first audio data. The sparse code matrix of the first audio data can be determined as described in operation 740.

[0120] In 760, processing device 122 (e.g., preprocessing module 220) can determine the reconstructed bone conduction audio data (e.g., based on the dictionary matrix and sparse code matrix of the reconstructed bone conduction audio data) based on the dictionary matrix and sparse code matrix of the reconstructed bone conduction audio data. Figure 4 The reconstructed first audio data described in the text). The processing device 122 can determine the reconstructed bone conduction audio data based on the dictionary matrix and sparse code matrix determined by operations 740 and 750 according to equation (4).

[0121] It should be noted that the foregoing is provided for illustrative purposes only and is not intended to limit the scope of this application. Various changes and modifications can be made by those skilled in the art based on the description in this application. However, these changes and modifications will not depart from the scope of this application. For example, operations 720 and 730 can be integrated into a single operation.

[0122] Figure 8 This is a flowchart illustrating an exemplary process for generating audio data according to some embodiments of this application. The operation of the process shown below is for illustrative purposes only. In some embodiments, process 800 may be accomplished using one or more additional operations not described, and / or without one or more operations discussed. Additionally, Figure 8 The order of operations of process 800 shown and described below is not limiting. In some embodiments, one or more operations of process 800 may be performed to achieve, as in combination Figure 4 At least a portion of the described operation 440.

[0123] In 810, the processing device 122 (e.g., audio data generation module 230 or frequency determination unit 310) can determine one or more frequency thresholds based at least in part on bone conduction audio data and / or air conduction audio data. Bone conduction audio data (e.g., first audio data or preprocessed first audio data) and air conduction audio data (e.g., second audio data or preprocessed second audio data) can be simultaneously acquired by bone conduction sensors and air conduction sensors, respectively, while the user speaks. Further description of bone conduction audio data and air conduction audio data can be found elsewhere in this application (e.g., Figure 4 (and its description).

[0124] As described herein, a frequency threshold may also be referred to as a frequency point. In some embodiments, a frequency threshold may be a frequency value of a frequency component in bone conduction audio data and / or air conduction audio data. In some embodiments, a frequency threshold may differ from the frequency value of a frequency component in bone conduction audio data and / or air conduction audio data. In some embodiments, processing device 122 may determine a frequency threshold based on a frequency response curve associated with bone conduction audio data. The frequency response curve associated with bone conduction audio data may include a frequency response value that varies with frequency. In some embodiments, processing device 122 may determine a frequency threshold based on the frequency response value of a frequency response curve associated with bone conduction audio data. For example, processing device 122 may determine a frequency threshold within a certain frequency range (e.g., such as...). Figure 10 The maximum frequency (e.g., as shown in the frequency response curve m, within the range of 0-2000Hz) is the highest frequency. Figure 10 The frequency response curve m shown (within 2000Hz) is defined as the frequency threshold. Frequency response values ​​within this frequency range that are less than a certain threshold (e.g., ...) Figure 10 The frequency response curve m shown is approximately 80 dB. For example, processing device 122 can handle a certain frequency range (e.g., such as...). Figure 10 The minimum frequency (e.g., in the frequency response curve m of 4000Hz-20kHz) is shown. Figure 10 The frequency response curve m shown (4000Hz) is defined as the frequency threshold. Frequency response values ​​within this frequency range that are greater than a certain threshold (e.g., such as...) Figure 10 The frequency response curve m shown is approximately 90 dB. As another example, processing device 122 can determine the minimum and maximum frequencies within a frequency range as frequency thresholds, where the frequency response values ​​corresponding to frequencies within that range fall within a certain range. For example, such as... Figure 10As shown, the processing device 122 can determine one or more frequency thresholds based on the frequency response curve "m" of the bone conduction audio data. The processing device 122 can determine a frequency range (0-2000Hz) corresponding to frequency response values ​​less than a certain threshold (e.g., 70dB). The processing device 122 can determine the maximum frequency within this frequency range as the frequency threshold. In some embodiments, the processing device 122 can determine one or more frequency thresholds based on the variation characteristics of the frequency response curve. For example, the processing device 122 can determine the maximum and / or minimum frequency within a frequency range where the frequency response curve has a stable variation as the frequency threshold. As another example, the processing device 122 can determine the maximum and / or minimum frequency within a frequency range where the frequency response curve varies drastically as the frequency threshold. Yet another example, the frequency response curve m with a frequency range less than 800Hz varies substantially stably relative to a frequency range greater than 800Hz and less than 4000Hz. The processing device 122 can determine 800Hz and 4000Hz as frequency thresholds. In some embodiments, the processing device 122 can use one or more reconstruction techniques described elsewhere in this application (e.g., Figure 4 The processing device 122 reconstructs bone conduction audio data (as described above) to obtain reconstructed bone conduction audio data. The processing device 122 can determine a frequency response curve associated with the reconstructed bone conduction audio data. The processing device 122 can determine a frequency threshold based on the frequency response curve associated with the reconstructed bone conduction audio data, in a manner similar to or the same as described above for determining a frequency threshold based on the frequency response curve of bone conduction audio data.

[0125] In some embodiments, the processing device 122 may determine one or more frequency thresholds based on the noise level associated with at least a portion of the air conduction audio data. A higher noise level may result in a higher frequency threshold (e.g., a minimum frequency threshold). A lower noise level may result in a lower frequency threshold (e.g., a minimum frequency threshold). In some embodiments, the noise level associated with the air conduction audio data may be represented by the amount or energy of noise included in the air conduction audio data. The greater the amount or energy of noise included in the air conduction audio data, the higher the noise level. In some embodiments, the noise level may be represented by the signal-to-noise ratio (SNR) of the air conduction audio data. A higher SNR indicates a lower noise level. A higher SNR associated with the air conduction audio data results in a smaller threshold. For example, if the SNR is 0 dB, the frequency threshold may be 2000 Hz. If the SNR is 20 dB, the frequency threshold may be 4000 Hz. For example, the frequency threshold may be determined based on equation (5) as follows:

[0126]

[0127] Among them, F pointThese can represent frequency thresholds. F1, F2, and F3 can be values ​​within the range of 0-20kHz, where F1>F2>F3. A1 and A2 are constant values; for example, A1 can be 0 and A2 can be equal to 20.

[0128] Furthermore, the frequency threshold can be expressed by equation (6):

[0129]

[0130] In some embodiments, the processing device 122 may determine the signal-to-noise ratio of the air conduction audio data according to equation (6) as follows:

[0131]

[0132] Where n refers to the nth speech frame in the air conduction audio data. This refers to the energy of the pure audio data contained in air conduction audio data. This refers to the energy of the noise data contained in the air conduction audio data. In some embodiments, the processing device 122 can use a noise estimation algorithm to determine the noise data in the air conduction audio data, such as the Minimum Statistical (MS) algorithm, the Minimum Controlled Recursive Average (MCRA) algorithm, etc. The processing device 122 can determine the pure audio data in the air conduction audio data based on the noise data in the air conduction audio data. Then, the processing device 122 can determine the energy of the pure audio data in the air conduction audio data and the energy of the noise data in the air conduction audio data. In some embodiments, the processing device 122 can use a bone conduction sensor and an air conduction sensor to determine the noise data in the air conduction audio data. For example, the processing device 122 can determine reference audio data acquired by the air conduction sensor, where the bone conduction sensor did not acquire any signal at the same time that the air conduction sensor acquired the reference audio data, and the time when the air conduction sensor acquired the reference audio data is close to the time when the air conduction sensor acquired the air conduction audio data. As used herein, a time close to another time can mean that the difference between the two times is less than a certain threshold (e.g., 10 ms, 100 ms, 1 second, 2 seconds, 3 seconds, 4 seconds, etc.). The reference audio data can be equivalent to the noise data in the air conduction audio data. Then, the processing device 122 can determine the pure audio data in the air conduction audio data based on the noise data (i.e., the reference audio data) in the air conduction audio data. And the processing device 122 can determine the signal-to-noise ratio associated with the air conduction audio data according to equation (7).

[0133] In some embodiments, the processing device 122 can extract the energy of noise data from the air conduction audio data and determine the energy of the pure audio data based on the energy of the noise data and the total energy of the air conduction audio data. For example, the processing device 122 can subtract the energy of the noise data in the air conduction audio data from the total energy of the air conduction audio data to obtain the energy of the pure audio data in the air conduction audio data. The processing device 122 can determine the signal-to-noise ratio based on the energy of the pure audio data and the energy of the noise data according to equation (7).

[0134] In 820, processing device 122 (e.g., audio data generation module 230 or weight determination unit 320) can divide bone conduction audio data and air conduction audio data into multiple segments according to one or more frequency thresholds. In some embodiments, the bone conduction audio data and air conduction audio data are time-domain data, and processing device 122 can perform a domain transformation operation (e.g., a Fourier Transform operation) on the bone conduction audio data and air conduction audio data to convert them into the frequency domain. In some embodiments, the bone conduction audio data and air conduction audio data can be frequency-domain data. The bone conduction audio data and air conduction audio data in the frequency domain can each include a spectrum. The bone conduction audio data in the frequency domain can also be referred to as the bone conduction spectrum. The air conduction audio data in the frequency domain can also be referred to as the air conduction spectrum. Processing device 122 can divide the bone conduction spectrum and air conduction spectrum into multiple segments respectively. Each segment of the bone conduction audio data can correspond to a segment of the air conduction audio data. As used herein, a segment of air conduction audio data corresponding to a segment of bone conduction audio data can mean that the two segments of the bone conduction audio data and the air conduction audio data are defined by one or two identical frequency thresholds. For example, if a specific segment of bone conduction audio data is defined by frequency points 2000Hz and 4000Hz—in other words, if the specific segment of bone conduction audio data includes frequency components within the range of 2000Hz to 4000Hz—the corresponding segment of air conduction audio data can also be defined by frequency thresholds 2000Hz and 4000Hz. In other words, the segment of air conduction audio data that corresponds to the segment of bone conduction audio data defined by 2000Hz to 4000Hz includes frequency components within the range of 2000Hz to 4000Hz.

[0135] In some embodiments, the count or number of frequency thresholds can be one, and the processing device 122 can divide the bone conduction frequency spectrum and the air conduction frequency spectrum into two segments respectively. For example, one segment of the two segments of the bone conduction spectrum may include a portion of the bone conduction spectrum with frequency components lower than the frequency threshold, and the other segment of the two segments of the bone conduction frequency spectrum may include the remaining portion of the bone conduction spectrum with frequency components higher than the frequency threshold.

[0136] In step 830, the processing device 122 (e.g., audio data generation module 230 or weight determination unit 320) can determine the weight of each segment among multiple segments of bone conduction audio data and air conduction audio data, respectively. In some embodiments, the weight of a specific segment of bone conduction audio data and the weight of a corresponding specific segment of air conduction audio data can satisfy a criterion such that the sum of the weights of the specific segments of bone conduction audio data and air conduction audio data equals 1. For example, if the processing device 122 divides the bone conduction audio data and air conduction audio data into two segments based on a single frequency threshold, the weight of a segment of bone conduction audio data having a frequency component below the single frequency threshold (also referred to as the low-frequency portion of bone conduction audio data) can be equal to 1, or 0.9, or 0.8, etc. The weight of a segment of air conduction audio data having a frequency component below the single frequency threshold (also referred to as the low-frequency portion of air conduction audio data) can be correspondingly equal to 0, or 0.1, or 0.2, etc., corresponding to weights of 1, 0.9, or 0.8, etc., of the segment of bone conduction audio data, respectively. The weight of another segment in bone conduction audio data that has a frequency component above a single frequency threshold (also known as the high-frequency portion of bone conduction audio data) can be equal to 0, 0.1, or 0.2, etc. The weight of another segment in air conduction audio data that has a frequency component above a single frequency threshold (also known as the high-frequency portion of air conduction audio data) can be correspondingly equal to 1, 0.9, or 0.8, etc., corresponding to the weight data of another segment in bone conduction audio data of 0, 0.1, or 0.2, respectively.

[0137] In some embodiments, the processing device 122 can determine the weights of different segments of bone conduction audio data or air conduction audio data based on the signal-to-noise ratio (SNR) of the air conduction audio data. For example, the lower the SNR of the air conduction audio data, the greater the weight of a specific segment of the bone conduction audio data can be, and the lower the weight of the corresponding specific segment of the air conduction audio data can be.

[0138] In 840, processing device 122 (e.g., audio data generation module 230 or combination unit 330) can splice bone conduction audio data and air conduction audio data for each of a plurality of segments in the bone conduction audio data and air conduction audio data to generate spliced ​​audio data. The spliced ​​audio data can represent a user's speech, which has a higher fidelity than the bone conduction audio data and / or air conduction audio data. The splicing of bone conduction audio data and air conduction audio data can refer to selecting one or more portions of the frequency components of air conduction audio data and one or more portions of the frequency components of bone conduction audio data in the frequency domain according to one or more frequency thresholds, and generating audio data based on the selected portions of bone conduction audio data and air conduction audio data. As described herein, the frequency threshold can also be referred to as the frequency splicing point. In some embodiments, the selected portions of bone conduction audio data and / or air conduction audio data may include frequency components below the frequency threshold. In some embodiments, the selected portions of bone conduction audio data and / or air conduction audio data may include frequency components below the frequency threshold and above another frequency threshold. In some embodiments, the selected portions of bone conduction audio data and / or air conduction audio data may include frequency components above the frequency threshold.

[0139] In some embodiments, the processing device 122 may determine the spliced ​​audio data according to equation (8) as follows:

[0140]

[0141] in, Bone conduction audio data, Air conduction audio data, Including (a m1 a m2 , ..., a mN () refers to the weights of multiple segments of bone conduction audio data. Including (b) m1 b m2 , ..., b mN ) refers to the weights of multiple segments of the air-conducted audio data, (x m1 x m2 , ..., x mN (y) Multiple segments of bone conduction audio data, each segment including frequency components within a frequency range defined by a frequency threshold, (y m1 y m2 , ..., y mN () refers to multiple segments of air-conducted audio data, each segment comprising frequency components within a frequency range defined by a frequency threshold. For example, x m1 and y m1 This can refer to frequency components below 800Hz in bone conduction audio data and air conduction audio data, respectively. For example, xm2 and y m2 This can refer to the frequency components in the bone conduction audio data and air conduction audio data within the frequency range of 800Hz and 4000Hz, respectively. N can be a constant, such as 1, 2, 3, etc. mn(n=1,2,…N) It can be a constant in the range of 0 to 1. b mn(n=1,2,…N) It can be a constant in the range of 0 to 1. mn(n=1,2,…N) and b mn(n=1,2,…N) The sum equals 1. In some embodiments, N may equal 2. The processing device 122 may divide each bone conduction audio data and air conduction audio data into two segments based on a single frequency threshold. For example, the processing device 122 may determine the low-frequency portion and the high-frequency portion of the bone conduction audio data (or air conduction audio data) based on a single frequency threshold. The low-frequency portion of the bone conduction audio data (or air conduction audio data) may include frequency components in the bone conduction audio data (or air conduction audio data) below the single frequency threshold, and the high-frequency portion of the bone conduction audio data (or air conduction audio data) may include frequency components in the bone conduction audio data (or air conduction audio data) above the single frequency threshold. In some embodiments, the processing device 122 may determine the low-frequency portion and the high-frequency portion of the bone conduction audio data (or air conduction audio data) based on one or more filters. One or more filters may include low-pass filters, high-pass filters, band-pass filters, etc., or any combination thereof.

[0142] In some embodiments, the processing device 122 may determine, at least partially, a first weight and a second weight for the low-frequency portion and the high-frequency portion of the bone conduction audio data, respectively, based on a single frequency threshold. The processing device 122 may also determine, at least partially, a third weight and a fourth weight for the low-frequency portion and the high-frequency portion of the air conduction audio data, respectively, based on a single frequency threshold. In some embodiments, the first, second, third, and fourth weights may be determined based on the signal-to-noise ratio (SNR) of the air conduction audio data. For example, if the SNR of the air conduction audio data is greater than a threshold, the processing device 122 may determine that the first weight is less than the third weight, and / or the second weight is greater than the fourth weight. As another example, the processing device 122 may determine a plurality of SNR ranges, each corresponding to a fixed first, second, third, and fourth weight. The first and second weights may be the same or different, and the third and fourth weights may be the same or different. The sum of the first and third weights is 1, and the sum of the second and fourth weights is 1. The first, second, third, and / or fourth weights can be constant values ​​in the range of 0 to 1, such as 1, 0.9, 0.8, 0.7, 0.3, 0.4, 0.5, 0.6, 0.2, 0.1, 0, etc. In some embodiments, the processing device 122 can use the first, second, third, and fourth weights respectively to weight the low-frequency and high-frequency portions of the bone conduction audio data and the low-frequency and high-frequency portions of the air conduction audio data to determine the spliced ​​audio data. For example, the processing device 122 can use the first and third weights to perform a weighted summation of the low-frequency portions of the bone conduction audio data and the low-frequency portions of the air conduction audio data to determine the low-frequency portion of the spliced ​​audio data. The processing device 122 can use the second and fourth weights to perform a weighted summation of the high-frequency portions of the bone conduction audio data and the high-frequency portions of the air conduction audio data to determine the high-frequency portion of the spliced ​​audio data. The processing device 122 can combine the low-frequency portion and the high-frequency portion of the spliced ​​audio data to obtain the spliced ​​audio data.

[0143] In some embodiments, the first weight of the low-frequency portion of the bone conduction audio data can be equal to 1, and the second weight of the high-frequency portion of the bone conduction audio data can be equal to 0. The third weight of the low-frequency portion of the air conduction audio data can be equal to 0, and the fourth weight of the high-frequency portion of the air conduction audio data can be equal to 1. The spliced ​​audio data can be generated by concatenating the low-frequency portions of the bone conduction audio data and the high-frequency portions of the air conduction audio data. In some embodiments, the audio data generated after concatenating the bone conduction audio data and the air conduction audio data can vary depending on the single frequency threshold. For example, as... Figures 16 to 20 As shown, Figures 16 to 20It is a time-frequency diagram of spliced ​​audio data generated by splicing specific bone conduction audio data and specific air conduction audio data at frequency points of 2000Hz, 3000Hz and 4000Hz, respectively, according to some embodiments of this application. Figure 16 , 19 The noise levels in the spliced ​​audio data corresponding to 20 are different from each other. The larger the frequency splicing point, the less noise is in the spliced ​​audio data.

[0144] It should be noted that the foregoing is provided for illustrative purposes only and is not intended to limit the scope of this application. Various changes and modifications can be made by those skilled in the art based on the description herein. However, such changes and modifications will not depart from the scope of this application.

[0145] Figure 9 This is a flowchart illustrating an exemplary process for generating audio data according to some embodiments of this application. The operation of the process shown below is for illustrative purposes only. In some embodiments, process 900 may be accomplished using one or more additional operations not described, and / or without one or more operations discussed. Additionally, Figure 9 The order of operations of process 900 shown and described below is not limiting. In some embodiments, one or more operations of process 900 may be performed to achieve, as in combination Figure 4 At least a portion of the described operation 440.

[0146] In 910, the processing device 122 (e.g., audio data generation module 230 or weight determination unit 320) can determine a weight corresponding to the bone conduction audio data, at least in part, based on at least one of bone conduction audio data or air conduction audio data. In some embodiments, when a user speaks, bone conduction audio data and air conduction audio data can be simultaneously acquired by a bone conduction sensor and an air conduction sensor, respectively. The air conduction audio data and bone conduction audio data can represent the user's speech. Further description of bone conduction audio data and air conduction audio data is available in [the relevant section / document / etc.]. Figure 4 It was found in its description.

[0147] In some embodiments, the processing device 122 may determine the weights of the bone conduction audio data based on the signal-to-noise ratio of the air conduction audio data. Further description of determining the signal-to-noise ratio of the air conduction audio data can be found elsewhere in this application (e.g., Figure 8(and its description). The higher the signal-to-noise ratio of the air conduction audio data, the lower the weight of the bone conduction audio data. For example, if the signal-to-noise ratio of the air conduction audio data is greater than a predetermined threshold, the weight of the bone conduction audio data can be set to value A; if the signal-to-noise ratio of the air conduction audio data is less than the predetermined threshold, the weight of the bone conduction audio data can be set to value B, where A < B. For another example, the processing device 122 can determine the weight of the bone conduction audio data according to equation (9) as follows:

[0148]

[0149] Where a1>a2>a3. A1 and / or A2 can be the default settings of the audio signal generation system 100. Further, the processing device 122 can determine at least two signal-to-noise ratio (SNR) ranges, each corresponding to a weight value of the bone conduction audio data, for example, equation (10):

[0150]

[0151] Among them, W bone This refers to the weights corresponding to bone conduction audio data.

[0152] In operation 920, processing device 122 (e.g., audio data generation module 230 or weight determination unit 320) may determine weights corresponding to air conduction audio data, at least in part, based on at least one of bone conduction audio data or air conduction audio data. The method for determining the weights of the air conduction audio data may be similar to or the same as the method for determining the weights of the bone conduction audio data, as described in operation 910. For example, processing device 122 may determine the weights of the air conduction audio data based on the signal-to-noise ratio (SNR) of the air conduction audio data. Further description of determining the SNR of the air conduction audio data can be found elsewhere in this application (e.g., Figure 8 (and its description). The higher the signal-to-noise ratio (SNR) of the air conduction audio data, the higher its weight. For example, if the SNR of the air conduction audio data is greater than a predetermined threshold, the weight of the air conduction audio data can be set to value X; if the SNR of the air conduction audio data is less than the predetermined threshold, the weight of the air conduction audio data can be set to value Y, and X > Y. The weights of the bone conduction audio data and the air conduction audio data need to meet certain criteria such that the sum of the weights of the bone conduction audio data and the air conduction audio data equals 1. The processing device 122 can determine the weight of the air conduction audio data based on the weight of the bone conduction audio data. For example, the processing device 122 can determine the weight of the air conduction audio data base based on the difference between 1 and the weight of the bone conduction audio data.

[0153] At 930, the processing device 122 (e.g., audio data generation module 230 or combination unit 330) can use the weights of the bone conduction audio data and the air conduction audio data to perform a weighted summation of the bone conduction audio data and the air conduction audio data to determine the target audio data. The target audio data may represent the user's speech, which is the same as the speech represented by the bone conduction audio data and the air conduction audio data. In some embodiments, the processing device 122 can determine the target audio data according to equation (11) as follows:

[0154]

[0155] Among them, S air This refers to air conduction audio data, S bone This refers to bone conduction audio data, a1 refers to the weight of air conduction audio data, b1 refers to the weight of bone conduction audio data, and S refers to... out This refers to the target audio data. n and b n The standard is that the sum equals 1. For example, the target audio data can be determined according to equation (12) as follows:

[0156]

[0157] In some embodiments, the processing device 122 may send target audio data to a client terminal (e.g., terminal 130), storage device 140 and / or any other storage device (not shown in the audio signal generation system 100) via a network 150.

[0158] Example

[0159] The embodiments provided below are for illustrative purposes only and are not intended to limit the scope of this application.

[0160] Example 1: Frequency response curves of bone conduction audio data, reconstructed bone conduction audio data, and corresponding air conduction audio data.

[0161] For example Figure 10 As shown, curve "m" represents the frequency response curve of the bone conduction audio data, and curve "n" represents the frequency response curve of the corresponding air conduction audio data. The bone conduction and air conduction audio data represent the same speech of the user. Curve "m1" represents the frequency response curve of the reconstructed bone conduction audio data generated by reconstructing the bone conduction audio data using a trained machine learning model according to process 500. For example... Figure 10As shown, the frequency response curve "m1" is closer to the frequency response curve "n" than the frequency response curve "m". In other words, the reconstructed bone conduction audio data is closer to the air conduction audio data than the bone conduction audio data. Furthermore, the portion of the frequency response curve "m1" of the reconstructed bone conduction audio data below the frequency point (e.g., 2000Hz) is more similar to or closer to the frequency of the air conduction audio data.

[0162] Example 2: Frequency response curves of bone conduction audio data acquired by bone conduction sensors located at different parts of the user's body.

[0163] like Figure 11 As shown, curve "p" represents the frequency response curve of bone conduction audio data acquired by a first bone conduction sensor located in the user's neck. Curve "b" represents the frequency response curve of bone conduction audio data acquired by a second bone conduction sensor located at the mastoid process of the user's body. Curve "o" represents the frequency response curve of bone conduction audio data acquired by a third bone conduction sensor located in the user's ear canal (e.g., external auditory canal). In some embodiments, the second and third bone conduction sensors are configured identically to the first bone conduction sensor. The bone conduction audio data acquired by the first, second, and third bone conduction sensors represents the same speech of the same user and is acquired simultaneously by the first, second, and third bone conduction sensors. In some embodiments, the first, second, and third bone conduction sensors may employ different configurations. The frequency response curves of bone conduction audio data acquired by bone conduction sensors with different configurations at the same location may differ.

[0164] like Figure 11As shown, the frequency response curves "p", "b", and "o" are different from each other. In other words, the bone conduction audio data collected by the first, second, and third bone conduction sensors differs depending on the location of these sensors on the user's body. For example, the response value of the bone conduction audio data with frequency components below 800Hz collected by the first bone conduction sensor located in the user's neck is greater than the response value of the bone conduction audio data with frequency components below 800Hz collected by the second bone conduction sensor located at the user's mastoid process. The frequency response curve reflects the ability of the bone conduction sensor to convert sound energy into electrical signals. According to the frequency response curves "p", "b", and "o", the response values ​​of the bone conduction sensors in different parts of the body are greater in the frequency range of 0 to approximately 5000Hz than in the frequency range above approximately 5000Hz. The frequency response curves "p", "b", and "o" change relatively smoothly in the frequency range of 0 to approximately 2000Hz, and change drastically above 2000Hz. The sensor is located in different areas of the user's body and has a strong ability to pick up low-frequency signals or low-frequency components (e.g., 0-2000Hz or 0-5000HZ). That is, the signal energy collected by the bone conduction sensor is mainly concentrated in the low-frequency range.

[0165] Therefore, as Figure 11 As shown, a bone conduction device for acquiring and / or playing audio signals may include a bone conduction sensor for acquiring bone conduction audio signals. This bone conduction sensor can be positioned on one or more parts or locations of the user's body through the design of the bone conduction device structure. When designing the bone conduction device structure, the area of ​​the user's body where the bone conduction sensor is located can be determined based on one or more characteristics such as frequency response curves, signal strength, user comfort, aesthetics, and convenience. For example, a bone conduction device may include a bone conduction sensor for acquiring bone conduction audio signals. When a user wears a bone conduction device, the bone conduction sensor can be positioned on the user's tragus, ear canal, and / or in contact with the user's tragus and ear canal, resulting in relatively high signal strength of the audio signal acquired by the bone conduction sensor, while also being convenient and aesthetically pleasing to wear.

[0166] Example 3: An exemplary frequency response curve of bone conduction audio data collected by applying different pressures to the same area of ​​the user's body using a bone conduction sensor.

[0167] exist Figure 12In the diagram, curve "L1" represents the frequency response curve of bone conduction audio data acquired when the bone conduction sensor applies a pressure F1 of 0 N to the user's tragus. As used herein, the pressure applied by the bone conduction sensor to the user's body part can also be referred to as the clamping force of the bone conduction sensor or bone conduction device. Curve "L2" represents the frequency response curve of bone conduction audio data acquired when the bone conduction sensor applies a pressure F2 of 0.2 N to the user's tragus. Curve "L3" represents the frequency response curve of bone conduction audio data acquired when the bone conduction sensor applies a pressure F3 of 0.4 N to the user's tragus. Curve "L4" represents the frequency response curve of bone conduction audio data acquired when the bone conduction sensor applies a pressure F4 of 0.8 N to the user's tragus. Figure 12 In this study, the frequency response curves “L1” through “L4” are different from each other. In other words, the bone conduction audio data collected by applying different pressures to the same area of ​​the user's body through a bone conduction sensor are different.

[0168] When a bone conduction sensor applies varying pressure to a part of the user's body, the bone conduction audio data it collects can differ. For example, the signal strength of the bone conduction audio data collected by the sensor can vary with pressure. When the pressure increases from 0 N to 0.8 N, the signal strength initially increases gradually, then the rate of increase slows down, eventually reaching saturation. However, the greater the pressure applied by the bone conduction sensor to the user's body, the more uncomfortable the user will feel when wearing it. Therefore, according to... Figure 11 and 12 As shown, a bone conduction device for acquiring and / or playing audio signals may include a bone conduction sensor for acquiring bone conduction audio signals. This bone conduction sensor can be positioned on one or more parts or locations of the user's body through structural design of the bone conduction device, and the clamping force of the bone conduction device on that part of the user's body can be within a certain range when the user wears it. In the structural design of the bone conduction device, the area of ​​the bone conduction sensor on the user's body and / or the clamping force applied to that part of the user's body can be determined based on one or more characteristics such as frequency response curves, signal strength, and user comfort. For example, the bone conduction device may include a bone conduction sensor for acquiring bone conduction audio signals, whereby the bone conduction sensor contacts the user's tragus when the user wears the device, and the clamping force applied to the tragus is in the range of 0 to 0.8 N, such as 0.2 N, 0.4 N, 0.6 N, or 0.8 N. This allows for a higher signal strength of the acquired bone conduction signal, while the appropriate clamping force makes the user feel more comfortable when wearing it.

[0169] Example 4: An exemplary time-frequency plot of spliced ​​audio data.

[0170] Figure 13This is a time-frequency diagram of spliced ​​audio data generated by splicing bone conduction audio data and air conduction audio data according to some embodiments of this application. The bone conduction audio data and air conduction audio data represent the same speech of the same user. The air conduction audio data includes noise. Figure 14 This is a time-frequency diagram of spliced ​​audio data generated by splicing bone conduction audio data and preprocessed air conduction audio data according to some embodiments of this application. The preprocessed air conduction audio data is generated by using a Wiener filter to reduce noise in the air conduction audio data. Figure 15 This is a time-frequency diagram of spliced ​​audio data generated from bone conduction audio data and another preprocessed air conduction audio data according to some embodiments of this application. The other preprocessed audio data is generated by denoising the air conduction audio data using spectral subtraction techniques. Figures 13 to 15 The time-frequency diagram of the spliced ​​audio data shown is generated based on process 800 using the same 2000Hz frequency splicing point. For example... Figures 13 to 15 As shown, Figure 14 (For example, region M) and Figure 15 (For example, region N) shows the proportion of frequency components above 2000Hz in the spliced ​​audio data. Figure 13 (For example, region O) shows less noise in the spliced ​​audio data with frequency components above 2000Hz, indicating that the spliced ​​audio data generated based on the denoised air conduction audio data has higher fidelity than the spliced ​​audio data generated based on the undenoised air conduction audio data. Figure 14 The frequency components above 2000Hz in the spliced ​​audio data shown are... Figure 15 The spliced ​​audio data shown exhibits different frequency components above 2000Hz due to different noise reduction techniques applied to the air conduction audio data. For example... Figure 14 and 15 As shown, Figure 14 The spliced ​​audio data shown contains more frequency components above 2000Hz (e.g., region M) than... Figure 15 The frequency components above 2000Hz in the spliced ​​audio data shown (e.g., region N) have less noise.

[0171] Example 5: An exemplary time-frequency plot of spliced ​​audio data generated based on different frequency thresholds.

[0172] Figure 16 It is a time-frequency graph of bone conduction audio data. Figure 17 This is a time-frequency plot of air conduction audio data corresponding to bone conduction audio data. Bone conduction audio data (e.g., Figure 4 The first audio data described herein) and air conduction audio data (e.g., Figure 4 The second audio data described herein can be simultaneously collected by a bone conduction sensor and an air conduction sensor while the user speaks. Figures 18 to 20This is a time-frequency diagram of spliced ​​audio data generated according to some embodiments of this application, which splices bone conduction audio data and air conduction audio data based on frequency thresholds (frequency splicing points) of 2000Hz, 3000Hz, and 4000Hz respectively. (Comparison) Figures 18 to 20 The time-frequency diagram of the spliced ​​audio data shown is... Figure 17 The time-frequency plot of the air conduction audio data shown is as follows. Figure 18 , 19 The noise in the spliced ​​audio data in 20 is less than Figure 17 The air conduction audio data shown. The higher the frequency threshold, the less noise in the spliced ​​audio data. Figures 18 to 20 The time-frequency diagram of the spliced ​​audio data shown is... Figure 16 The time-frequency plots of the bone conduction audio data shown are compared with those of the bone conduction audio data. Figure 16 Frequency components below 2000Hz, 3000Hz, and 4000Hz. Figures 18 to 20 The frequency components below 2000Hz, 3000Hz and 4000Hz respectively increased.

[0173] It should be noted that the above descriptions of various embodiments are provided for illustrative purposes only and are not intended to limit the scope of this application. Those skilled in the art can make various changes and modifications based on the descriptions in this application. However, these changes and modifications will not depart from the scope of this application. Finally, it should be understood that the embodiments described in this application are only used to illustrate the principles of the embodiments of this application. Other variations may also fall within the scope of this application. Therefore, alternative configurations of the embodiments of this application are considered as examples and not limitations, and are regarded as consistent with the teachings of this application. Accordingly, the embodiments of this application are not limited to the embodiments explicitly described in this application.

Claims

1. An audio signal generation method, characterized by, include: Acquire the first audio data collected by the bone conduction sensor; Acquire second audio data collected by an air conduction sensor. The first audio data and the second audio data represent the user's voice. The first audio data and the second audio data are composed of different frequency components. The first audio data and the second audio data are divided into multiple segments based on one or more frequency thresholds, wherein each segment of the first audio data corresponds to a segment of the second audio data; The weight of each segment in the first audio data and the second audio data is determined based on the signal-to-noise ratio of the second audio data, wherein the weights of two corresponding segments in the first audio data and the second audio data are summed to 1. Based on weights, each segment of the first audio data and the second audio data is spliced, merged, and / or combined to generate the third audio data.

2. The method of claim 1, wherein, Based on the third audio data, target audio data representing the user's speech is determined, and the target audio data has a higher fidelity than the first audio data and the second audio data.

3. The method of claim 2, wherein, Post-processing operations are performed on the third audio data to obtain the target audio data. The post-processing operations include noise reduction operations, domain transformation operations, or combinations thereof.

4. The method of claim 1, wherein, The frequency threshold is determined by the following process, which includes: Determine the noise level associated with the second audio data; Based on the noise level associated with the second audio data, at least one of the one or more frequency thresholds is determined.

5. The method of claim 4, wherein, The noise level associated with the second audio data is represented by the signal-to-noise ratio (SNR) of the second audio data, and the SNR of the second audio data is determined by the following operations: The energy of the noise in the second audio data is determined using the bone conduction sensor and the air conduction sensor. Based on the energy of the noise in the second audio data, determine the energy of the pure audio data in the second audio data; The signal-to-noise ratio is determined based on the energy of the noise in the second audio data and the energy of the pure audio data in the second audio data.

6. The method of claim 4, wherein, The higher the noise level associated with the second audio data, the higher at least one of the one or more frequency thresholds.

7. The method of claim 1, wherein, At least one of the one or more frequency thresholds is determined based on the frequency response curve associated with the first audio data.

8. The method of claim 1, wherein, The first audio data and the second audio data are divided into multiple segments based on one or more frequency thresholds, including: The low-frequency portion of the first audio data is determined, wherein the low-frequency portion includes frequency components that are lower than one of the one or more frequency thresholds; Determine the high-frequency portion of the second audio data, the high-frequency portion including frequency components higher than one of the frequency thresholds.

9. An audio signal generation system, characterized by include: At least one processor; An executable instruction, which can be executed by the at least one processor, causes the system to perform the audio signal generation method as described in any one of claims 1-8.

10. A system for audio signal generation, characterized by include: The acquisition module is used to acquire first audio data collected by a bone conduction sensor and second audio data collected by a gas conduction sensor. The first audio data and the second audio data represent the user's voice, and the first audio data and the second audio data are composed of different frequency components. The weight determination unit is configured to divide the first audio data and the second audio data into multiple segments according to one or more frequency thresholds, and is configured to determine the weight of each segment in the multiple segments of the first audio data and the second audio data according to the signal-to-noise ratio of the second audio data, wherein each segment of the first audio data corresponds to a segment of the second audio data, and the weights of two corresponding segments in the first audio data and the second audio data are summed to 1. The combining unit is configured to splice, merge, and / or combine each segment of a plurality of segments of the first audio data and the second audio data based on weights to generate third audio data.

11. A non-transitory computer-readable medium, comprising: The medium stores computer instructions, which, when executed, perform the audio signal generation method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Voice enhancement method, device, equipment and storage medium

    CN109767783A

  • System and method for audio signal generation

    CN112581970A