Audio generation method and device and storage medium

By forming multiple candidate microphone arrays and selecting the optimal array based on sound field characteristics, the problems of low efficiency and low accuracy caused by fixed array configurations and mechanical adjustments of microphone arrays are solved, achieving more efficient and accurate audio generation.

CN121908190APending Publication Date: 2026-04-21TP-LINK
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TP-LINK
Filing Date
2026-01-14
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

The fixed array configuration and mechanical adjustment method of existing microphone arrays result in low audio generation efficiency and low accuracy, making them unsuitable for the needs of different scenarios.

Method used

By activating a target number of microphones to form multiple candidate microphone arrays, sound field features are extracted to generate sound field parameters, and the optimal microphone array is selected based on weighted combinations, replacing mechanical structure adjustment and improving adaptation accuracy and efficiency.

Benefits of technology

It improves the accuracy and efficiency of audio generation, avoids mechanical wear and redundant noise, and enables more objective array selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121908190A_ABST
    Figure CN121908190A_ABST
Patent Text Reader

Abstract

The invention provides an audio generation method and device and a storage medium, and the method comprises the steps: starting a target number of microphones which are used for forming a plurality of candidate microphone arrays through combination, and enabling the plurality of candidate microphone arrays to comprise an initial microphone array; performing audio sampling based on the initial microphone array to obtain first audio data, and performing sound field feature extraction on the first audio data to obtain sound field features; generating sound field parameters of different dimensions based on the sound field features, and generating a first weight combination based on the plurality of sound field parameters; fusing the plurality of sound field parameters based on the first weight combination to obtain a first score, and selecting a target microphone array from the initial microphone array and the reference candidate microphone array corresponding to the first weight combination based on the first score; and performing audio sampling based on the target microphone array to obtain second audio data, and generating a target audio based on the second audio data. Therefore, the audio generation efficiency and the accuracy of the generated audio can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet technology, and in particular to an audio generation method, an electronic device, and a computer-readable storage medium. Background Technology

[0002] Microphone arrays commonly used in related technologies are usually fixed arrays with fixed microphone spacing and arrangement, which limits their adaptability to different scenarios. Furthermore, while adaptive microphone technology in these technologies adjusts microphone positions within a small range while maintaining the original configuration, this adjustment relies on mechanical structures, resulting in slow response times. Long-term use leads to mechanical wear and tear, causing lower accuracy in the audio generated by the microphones. Therefore, the audio generation efficiency and accuracy of microphone array-based technologies in these technologies are both relatively low. Summary of the Invention

[0003] This application provides an audio generation method, an electronic device, and a computer-readable storage medium, which can improve the efficiency of audio generation and the accuracy of the generated audio.

[0004] The technical solution of this application embodiment is implemented as follows: This application provides an audio generation method, including: A target number of microphones are activated, which are used to combine to form a plurality of candidate microphone arrays, the plurality of candidate microphone arrays including an initial microphone array. Audio sampling is performed based on the initial microphone array to obtain first audio data, and sound field features are extracted from the first audio data to obtain sound field features. Based on the sound field features, sound field parameters of different dimensions are generated, and based on multiple sound field parameters, a first weight combination for fusing the multiple sound field parameters is generated; wherein, the weight combination has a corresponding relationship with the candidate microphone array; Based on the first weight combination, the multiple sound field parameters are fused to obtain a first score, and based on the first score, a target microphone array is selected from the reference candidate microphone arrays corresponding to the initial microphone array and the first weight combination. The first score is used to indicate the degree of adaptation of the reference candidate microphone array to the surrounding environment. Audio data is obtained by sampling the target microphone array to obtain second audio data, and target audio is generated based on the second audio data.

[0005] This application provides an audio generation apparatus, including: A startup module is used to start a target number of microphones, which are used to combine to form multiple candidate microphone arrays, including an initial microphone array. The sampling module is used to perform audio sampling based on the initial microphone array to obtain first audio data, and to extract sound field features from the first audio data to obtain sound field features. The first generation module is used to generate sound field parameters of different dimensions based on the sound field features, and to generate a first weight combination for fusing the multiple sound field parameters based on the multiple sound field parameters; wherein the weight combination has a corresponding relationship with the candidate microphone array; The fusion module is used to fuse the multiple sound field parameters based on the first weight combination to obtain a first score, and to select a target microphone array from the reference candidate microphone arrays corresponding to the initial microphone array and the first weight combination based on the first score; wherein, the first score is used to indicate the degree of adaptation of the reference candidate microphone array corresponding to the first weight combination to the surrounding environment. The second generation module performs audio sampling based on the target microphone array to obtain second audio data, and generates target audio based on the second audio data.

[0006] This application provides an electronic device, including: Memory is used to store executable instructions or computer programs. The processor, when executing computer-executable instructions or computer programs stored in the memory, implements the audio generation method provided in the embodiments of this application.

[0007] This application provides a computer-readable storage medium storing computer-executable instructions or computer programs, which, when executed by a processor, implement the audio generation method provided in this application.

[0008] The embodiments of this application have the following beneficial effects: By activating a target number of microphones to form multiple candidate microphone arrays (including the initial array), a diverse selection space for microphone arrays is provided for different sound fields, avoiding the limitations of fixed microphone arrays. Simultaneously, after generating sound field parameters based on the sound field features extracted from the initial microphone array, and then fusing these parameters through weighted combinations to generate a first score, the target microphone array is selected based on this first score. Here, the first score directly indicates the degree of adaptation of the candidate microphone array to the environment. The optimal target microphone array is selected from the candidate arrays based on the score, replacing methods that rely on mechanical structures to adjust microphone positions or manual selection. This selection process is more objective and quantitative, significantly improving adaptation accuracy and thus enhancing the accuracy of the audio generated subsequently based on the selected target microphone array. Furthermore, the target microphone array is selected from the candidate microphone arrays, rather than activating all microphones. This avoids unnecessary microphone activation while ensuring adaptability, reducing the computational power consumption for data acquisition, transmission, and processing, and avoiding additional noise introduced by redundant microphones. This not only improves audio generation efficiency but also enhances the accuracy of the generated audio. Attached Figure Description

[0009] Figure 1 This is a schematic diagram of the architecture of the audio generation system 100 provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application; Figure 3 This is a flowchart illustrating the audio generation method provided in an embodiment of this application; Figure 4 This is a schematic diagram of the target number of microphones provided in the embodiments of this application; Figure 5 This is a schematic diagram of the microphone-based signal processing process provided in an embodiment of this application; Figure 6 This is a schematic diagram of the process for generating sound field parameters of different dimensions based on sound field features, provided in an embodiment of this application. Detailed Implementation

[0010] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0011] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0012] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0013] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0014] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0015] 1) Client, also known as user terminal, refers to the program that provides local services to users in contrast to the server. Except for some applications that can only run locally, it is generally installed on the terminal and needs to work with the server. That is, there needs to be a corresponding server and service program on the network to provide the corresponding services. Thus, a specific communication connection needs to be established between the client and the server to ensure the normal operation of the application.

[0016] 2) A microphone, also known as a microphone or transducer, is an energy conversion device that converts sound signals into electrical signals.

[0017] See Figure 1 , Figure 1 This is a schematic diagram of the architecture of the audio generation system 100 provided in the embodiments of this application. The terminal (terminal 400 is shown as an example) is connected to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two, and data transmission is achieved using wireless or wired links.

[0018] In some embodiments, server 200 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminal 400 can be a smartphone, tablet, laptop, desktop computer, set-top box, smart voice interaction device, smart home appliance, virtual reality device, vehicle terminal, aircraft, portable music player, personal digital assistant, dedicated messaging device, portable gaming device, smart speaker, and smartwatch, but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited in this embodiment.

[0019] The following describes an electronic device that implements the audio generation method provided in the embodiments of this application, and an electronic device that implements the online form document processing method provided in the embodiments of this application. See also Figure 2 , Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device can be a server or a terminal. The electronic device is used as an example. Figure 1 Taking the server shown as an example, Figure 2 The illustrated electronic device includes at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. The various components in terminal 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to implement communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 2 The general labeled all buses as Bus System 440.

[0020] Processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor.

[0021] User interface 430 includes one or more output devices 431 that enable the display of media content, including one or more speakers and / or one or more visual displays. User interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0022] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices physically located away from the processor 410.

[0023] The memory 450 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 450 described in this application embodiment is intended to include any suitable type of memory.

[0024] In some embodiments, memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0025] Operating system 451 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks; The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc. Presentation module 453 is configured to enable the display of information (e.g., user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 (e.g., display screen, speaker, etc.) associated with user interface 430. The input processing module 454 is used to detect and translate one or more user inputs or interactions from one or more input devices 432.

[0026] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2An audio generation device 455 stored in memory 450 is shown. This device can be software in the form of programs and plugins, and includes the following software modules: a startup module 4551, a sampling module 4552, a first generation module 4553, a fusion module 4554, and a second generation module 4555. These modules are logically connected and can therefore be arbitrarily combined or further separated according to their implemented functions. The functions of each module will be described below.

[0027] In other embodiments, the apparatus provided in this application can be implemented in hardware. For example, the audio generation apparatus provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the audio generation method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0028] Based on the above description of the audio generation system and electronic device provided in the embodiments of this application, the audio generation method provided in the embodiments of this application is described below. In actual implementation, the audio generation method provided in the embodiments of this application can be implemented by a terminal or a server alone, or by a terminal and a server working together, so that... Figure 1 The following description uses the example of server 200 executing the audio generation method provided in this embodiment of the application independently. See also... Figure 3 , Figure 3 This is a flowchart illustrating the audio generation method provided in this application embodiment. Next, it will be discussed in conjunction with... Figure 3 The steps are shown and explained.

[0029] Step 101: Activate a target number of microphones, which are used to combine to form multiple candidate microphone arrays, including the initial microphone array.

[0030] In practice, the target number can be preset, and the microphone is the sound source. Based on this, the target number of microphones can be, for example, a 9x9 microphone hardware, such as... Figure 4 As shown, Figure 4 This is a schematic diagram of the target number of microphones provided in the embodiments of this application, based on Figure 4Each solid dot indicates a microphone.

[0031] It should be noted that the target number of microphones are used to form multiple candidate microphone arrays. That is, each candidate microphone array is formed by combining at least a portion of the target number of microphones. These multiple candidate microphone arrays include at least pre-defined linear arrays, circular arrays, rectangular arrays, cross arrays, etc., but this application does not limit the specific arrays. The microphones are uniformly encoded using a codec chip and processed by the main controller via I2S signals. Figure 5 As shown, Figure 5 This is a schematic diagram of the microphone-based signal processing process provided in the embodiments of this application.

[0032] In practical implementation, after activating the target number of microphones, the initial microphone array among multiple candidate microphone arrays is controlled to enter the sampling state, thereby performing audio sampling based on the initial microphone array. The initial microphone array can be pre-set, for example, a 16-microphone rectangular array; however, this embodiment does not limit its implementation.

[0033] Step 102: Audio sampling is performed based on the initial microphone array to obtain first audio data, and sound field features are extracted from the first audio data to obtain sound field features.

[0034] In practical implementation, for the process of obtaining the first audio data based on the initial microphone array, after the initial microphone array enters the sampling state, the main control unit issues a sampling command to start a continuous sampling process of a preset duration. The preset duration is determined by the minimum data volume requirement to be met for sound field feature extraction (e.g., 200ms, covering at least two speech frames or a stable background sound cycle). Afterwards, each microphone in the initial microphone array converts the sound signal into an analog electrical signal, which is then synchronously converted to digital by the codec chip, and noise reduction preprocessing (e.g., eliminating hardware noise floor) is performed. Encoding (such as I2S format encoding) ensures the time synchronization of data sampled by multiple microphones and avoids time delay deviations. Then, the encoded data is transmitted to the main control unit through the I2S bus to form the raw sampled data. Next, the main control optimizes the received raw sampled data (removing invalid sampling points with abnormal amplitude, such as signals that exceed the dynamic range of the microphone), and then integrates the optimized raw sampled data into structured data according to a preset format (such as PCM format), and marks the sampling timestamp, sampling rate and other metadata, finally forming the first audio data that can be directly used for sound field feature extraction.

[0035] In some embodiments, the process of extracting sound field features from the first audio data to obtain sound field features specifically includes: extracting spatial features from the first audio data to obtain sound field spatial features, and extracting temporal features from the first audio data to obtain sound field temporal features; and fusing the sound field spatial features and the sound field temporal features to obtain sound field features.

[0036] It should be noted that the process of extracting spatial features from the first audio data to obtain the sound field spatial features specifically includes extracting spatial features from the first audio data through a convolutional neural network to obtain the sound field spatial features. The sound field spatial features are used to indicate the location of the sound source and the noise distribution. The process of extracting temporal features from the first audio data to obtain the sound field temporal features specifically includes extracting temporal features from the first audio data through a recurrent neural network to obtain the sound field temporal features. The sound field temporal features are used to indicate the time decay of reverberation and the dynamic changes of the sound source.

[0037] It should be noted that the convolutional neural network adopts a lightweight design, containing two convolutional layers, a pooling layer, and a fully connected layer. The first convolutional layer is used to extract local spatial correlation features (such as the signal differences between adjacent microphones), while the second convolutional layer is used to capture global spatial pattern features (such as the spatial distribution of sound sources in the array). The pooling layer is used to reduce the dimensionality of the local spatial correlation features extracted by the first convolutional layer, so that the dimensionality-reduced features are input into the second convolutional layer. The fully connected layer is used to map the features output by the second convolutional layer into a fixed-dimensional vector (such as 64-dimensional), and finally outputs standardized sound field spatial features.

[0038] The recurrent neural network uses a Long Short-Term Memory (LSTM) or Gated Recurrent Unit (GUR) neural network, which includes a hidden layer and an output layer. The hidden layer captures the dependencies between time frames through a gating mechanism (such as the changes in signal amplitude between the previous and next frames and the pattern of spectrum migration), and focuses on learning the decay trend of reverberation (such as the decay rate of signal amplitude over time) and the dynamic changes of sound sources (such as the time nodes of sound source appearance / disappearance and the shift of direction over time). The output layer is used to encode the temporal features output by the hidden layer into a fixed-dimensional vector (such as 64-dimensional), which is the temporal feature of the sound field.

[0039] In actual implementation, after determining the spatial and temporal characteristics of the sound field, the spatial and temporal characteristics of the sound field are fused to obtain the sound field characteristics. The fusion of the spatial and temporal characteristics of the sound field can be, for example, by splicing or adding the spatial and temporal characteristics of the sound field. This application does not limit the specific method used in this embodiment.

[0040] Thus, by separating spatial features and temporal features for extraction, the "spatial distribution attributes" (such as sound source location and noise orientation) and "temporal dynamic attributes" (such as reverberation attenuation and signal fluctuation) of the sound field are captured respectively. This avoids the one-sidedness of single-dimensional features in describing the sound field, ensures that the sound field features can fully reflect the acoustic environment on site, and improves the comprehensiveness of the extracted sound field features.

[0041] Step 103: Based on the sound field features, generate sound field parameters of different dimensions, and based on multiple sound field parameters, generate a first weight combination for fusing multiple sound field parameters; wherein, the weight combination has a corresponding relationship with the candidate microphone array.

[0042] It should be noted that the multiple dimensions include the signal-to-noise ratio dimension, the directivity dimension, and the directivity gain dimension. Thus, the sound field parameters of different dimensions include the signal-to-noise ratio parameter of the signal-to-noise ratio dimension, the directivity parameter of the directivity dimension, and the directivity gain parameter of the directivity gain dimension. The process of generating sound field parameters of different dimensions based on sound field characteristics will only be described below, and will not be elaborated here in the embodiments of this application. There are multiple ways to generate a first weighted combination for fusing multiple sound field parameters. The following sections will explain the process of generating a first weighted combination for fusing multiple sound field parameters based on two different methods.

[0043] In some embodiments, the process of generating a first weight combination for fusing multiple sound field parameters based on multiple sound field parameters may involve obtaining a pre-built weight combination generation model, then inputting multiple sound field parameters into the weight combination generation model, and the model calculating and outputting the weight coefficients corresponding to each sound field parameter through forward propagation to obtain the first weight combination corresponding to the multiple sound field parameters.

[0044] It should be noted that the weight combination includes multiple weights, and different weights correspond to different sound field parameters. The weight combination generation model is trained based on a preset training dataset. The training dataset contains multiple sets of sound field parameter samples and corresponding optimal weight combination annotations. The model structure adopts a multilayer perceptron. The input layer dimension is consistent with the number of sound field parameters, and the output layer dimension is consistent with the number of sound field parameters. The output value is the weight coefficient corresponding to each sound field parameter, and the output layer is set with normalization constraints to ensure that the sum of all weight coefficients is 1.

[0045] In other embodiments, the process of generating a first weighted combination for fusing multiple sound field parameters based on multiple sound field parameters may include: identifying a target sound field scene corresponding to the surrounding environment based on multiple sound field parameters; obtaining a preset correspondence between sound field scenes and weighted combinations; and obtaining a first weighted combination corresponding to the target sound field scene based on the correspondence.

[0046] It should be noted that the process of identifying the target sound field scene corresponding to the surrounding environment based on multiple sound field parameters involves several steps. First, multiple sound field scenes (such as a solo speech scene, a roundtable meeting scene, a high-noise environment scene, and a high-reverberation scene) are acquired from a sound field scene type library. Each sound field scene is associated with a set of feature thresholds (e.g., for a solo speech scene, the directivity parameter ≥ 0.7, the signal-to-noise ratio parameter ≥ 40dB, and the reverberation intensity level ≤ 2). Second, for each sound field scene, multiple sound field parameters are matched with the feature thresholds of that sound field scene. Specifically, sound field parameters that meet the corresponding feature thresholds are acquired as matching parameters, and the ratio of the number of matching parameters to the total number of sound field parameters is obtained. This ratio is then used as the degree of matching between the multiple sound field parameters and the sound field scene. Finally, the sound field scene with the highest degree of matching is determined as the target sound field scene corresponding to the surrounding environment.

[0047] As for obtaining the correspondence between the preset sound field scene and the weight combination, the correspondence between the sound field scene and the weight combination is preset. Specifically, for each sound field scene in the type library, the appropriate weight combination is determined based on laboratory test data, and then the mapping relationship between the sound field scene and the appropriate weight combination is constructed. In this way, after determining the target sound field scene corresponding to the surrounding environment, the correspondence between the preset sound field scene and the weight combination is directly obtained, and the first weight combination corresponding to the target sound field scene is obtained based on the correspondence.

[0048] There is a correspondence between weight combinations and candidate microphone arrays. This correspondence is established through a pre-defined mapping table. The construction process of the mapping table includes: first, identifying the core features of all candidate microphone arrays, including array topology type (linear array, circular array, rectangular array, spiral array, etc.), number of microphones, array physical dimensions (such as length of linear array, radius of circular array), and microphone spacing; second, for each candidate microphone array, determining the appropriate weight combination based on laboratory training data and the core features of the candidate microphone array; and finally, storing the candidate microphone arrays and their corresponding weight combinations in the mapping table to form a one-to-one correspondence.

[0049] In practical implementation, as mentioned above, multiple dimensions are involved, including signal-to-noise ratio, directivity, and directivity gain. The process of generating sound field parameters of different dimensions based on sound field characteristics is described in [reference needed]. Figure 6 , Figure 6 This is a flowchart illustrating the process of generating sound field parameters of different dimensions based on sound field features, as provided in an embodiment of this application. Figure 6 The process of generating sound field parameters of different dimensions based on sound field characteristics is achieved through the following steps.

[0050] Step 1031: Based on the sound field characteristics, generate sound field information of the surrounding environment. The sound field information includes at least one of the following: the location of the main noise source, the reverberation intensity, and the sound source distribution heat map.

[0051] In practical implementation, for the process of generating sound field information of the surrounding environment based on sound field features, the first step is to perform dimensional analysis of the sound field features. The sound field features are high-dimensional vectors that integrate spatial and temporal features of the sound field. The vector is pre-divided into dimensions according to the following rules: the first part of the dimension corresponds to the location-related features of the main noise source (such as azimuth angle and amplitude difference dimension in spatial features), the second part of the dimension corresponds to the reverberation intensity-related features (such as attenuation coefficient and energy fluctuation dimension in temporal features), and the third part of the dimension corresponds to the sound source distribution heatmap-related features (such as multi-channel amplitude distribution and coherence dimension in spatial features), ensuring that each type of sound field information has corresponding feature dimensions to support it.

[0052] Then, after completing the dimensional analysis, generation operations are performed for different types of sound field information. Specifically, for the location of the main noise source, the feature values ​​of the corresponding dimension are extracted, and the feature values ​​are mapped to the actual azimuth angle (range 0°~360°) through a preset decoding formula, which is the location of the main noise source. The decoding formula is constructed based on the training dataset, which contains the correspondence between the feature values ​​and the measured azimuth angle of the main noise source. The location of the main noise source is used to clarify the spatial location of the main interference noise on site. For the reverberation intensity, the feature values ​​of the corresponding dimension are extracted, and the reverberation intensity level is divided according to the size of the feature value (e.g., level 1~5, the larger the feature value, the higher the level). The level division threshold is determined through laboratory testing. The reverberation intensity is used to quantify the reverberation degree of the surrounding environment. For the sound source distribution heatmap, the feature values ​​of the corresponding dimension are extracted, and the feature values ​​are mapped to a preset spatial grid (the grid matches the monitoring range of the microphone array). The value of the grid node represents the sound source energy density at that location. The higher the value, the more concentrated the sound source distribution. Finally, a visualized heatmap data is formed. The sound source distribution heatmap is used to visualize the spatial distribution of effective speech sounds.

[0053] Step 1032: Based on the sound field information, generate the signal-to-noise ratio parameter in the signal-to-noise ratio dimension, the directivity parameter in the directivity dimension, and the directivity gain parameter in the directivity gain dimension, and determine the signal-to-noise ratio parameter, the directivity parameter, and the directivity gain parameter as sound field parameters of different dimensions.

[0054] In practical implementation, the process of generating the signal-to-noise ratio (SNR) parameter based on sound field information can be as follows: Based on sound field information, identify the human voice signal and noise signal in the first audio data, and obtain the amplitude of the human voice signal and the amplitude of the noise signal; obtain the ratio of the amplitude of the human voice signal to the amplitude of the noise signal as the first ratio, and obtain the first adjustment coefficient corresponding to the SNR parameter; and determine the SNR parameter by multiplying the first ratio and the first adjustment coefficient.

[0055] It should be noted that, specifically, the process of identifying human voice signals and noise signals in the first audio data based on sound field information involves first performing signal region division, that is, using the location of the main noise source and the sound source distribution heatmap in the sound field information to determine the effective sound source region (the region where human voice signals are likely to be distributed) and the noise region (such as the energy sparse region of the heatmap). According to the spatial sampling characteristics of the initial microphone array, the first audio data is mapped to the corresponding region according to the direction of the signal source, thus achieving preliminary signal region separation.

[0056] After dividing the signal into regions, the acoustic features of the signals within each region are extracted, including time-domain features (short-time energy, zero-crossing rate) and frequency-domain features (Mel-frequency cepstral coefficients, spectral centroid). Based on the acoustic features of the signals in each effective sound source region, the corresponding signals are identified, and the identification results are used to indicate whether the corresponding signal is a human voice signal. The acoustic features of human voice signals are: short-time energy in the range of 0.01~0.1, zero-crossing rate in the range of 0.1~0.3, and spectral centroid concentrated in the range of 300~3400Hz. The acoustic features of noise signals are: small fluctuations in short-time energy, stable zero-crossing rate, and spectral centroid deviating from the human voice frequency band. Thus, based on the identification results, human voice signal regions are determined among multiple effective sound source regions, and the remaining effective sound source regions are identified as noise signal regions, achieving accurate identification of human voice and noise signals.

[0057] For the identified human voice signal and noise signal, the process of obtaining the amplitude of the human voice signal and the amplitude of the noise signal specifically involves calculating the average amplitude of all sampling points (such as sound sources) within the human voice signal area. This average is calculated as the ratio of the sum of the amplitudes of all sampling points within the human voice signal area to the number of sampling points, which is then used to determine the human voice signal amplitude. Simultaneously, the same average calculation is performed on all sampling points within the noise signal area. The process of determining all sampling points (such as sound sources) within the human voice signal area or noise signal area, and the process of locating the sound source, is described below as the process of locating multiple sound sources in the surrounding environment based on sound field information. This will not be elaborated upon in the embodiments of this application.

[0058] Then, the ratio of the amplitude of the human voice signal to the amplitude of the noise signal is obtained as the first ratio, and the first adjustment coefficient corresponding to the signal-to-noise ratio parameter is obtained; the product of the first ratio and the first adjustment coefficient is determined as the signal-to-noise ratio parameter; wherein, the first adjustment coefficient is preset, and combined with the reverberation intensity level correction, the higher the reverberation level, the lower the first adjustment coefficient, to avoid reverberation being misjudged as valid human voice; and the signal-to-noise ratio parameter is greater than 0 and less than 100, the higher the value, the more prominent the valid human voice.

[0059] In practical implementation, the process of generating directional parameters based on sound field information can be as follows: Based on sound field information, locate multiple sound sources in the surrounding environment and obtain the emission probability of each sound source; based on the emission probability of each sound source, identify the main sound source and other sound sources from the multiple sound sources; sum the emission probabilities of each other sound source to obtain the total emission probability of the other sound sources, and use the ratio of the emission probability of the main sound source to the total emission probability of the other sound sources as the second ratio; obtain the second adjustment coefficient corresponding to the directional parameter, and determine the directional parameter by multiplying the second ratio and the second adjustment coefficient.

[0060] It should be noted that the process of locating multiple sound sources in the surrounding environment based on sound field information specifically involves first extracting the sound source distribution heatmap and the azimuth data of the main noise source from the sound field information. The energy peak positions of the sound source distribution heatmap are used as candidate positions of the sound sources, and the number of energy peaks is used as the number of candidate sound sources. Secondly, a sound source localization algorithm (such as the controllable response power-phase transformation algorithm) is used, combined with the physical coordinates of the microphone array, to accurately locate each candidate position. The azimuth angle and distance of each candidate position are calculated to form the spatial coordinates of multiple sound sources. Then, the process of obtaining the emission probability of each sound source involves, specifically, extracting the signal segment corresponding to the spatial coordinate direction of the sound source from the first audio data for each located sound source, performing a short-time Fourier transform on the signal segment to obtain the spectrum distribution of the signal, identifying the peak probability in the spectrum distribution, i.e. the probability of the peak occurring, and determining the peak probability as the emission probability of the corresponding sound source.

[0061] Next, regarding the process of identifying the main sound source and other sound sources from multiple sound sources based on the sound emission probability of each sound source, specifically, based on the sound emission probability of each sound source, the sound source with the highest sound emission probability among multiple sound sources is determined as the main sound source, and the remaining sound sources other than the main sound source are determined as other sound sources.

[0062] Then, the emission probabilities of all other sound sources are summed to obtain the total emission probability of other sound sources. The ratio of the emission probability of the main sound source to the total emission probability of other sound sources is taken as the second ratio. The second adjustment coefficient corresponding to the directivity parameter is then obtained. The product of the second ratio and the second adjustment coefficient is determined as the directivity parameter. The second adjustment coefficient is used to indicate the spatial concentration of effective sound sources in the heat map. The more concentrated the concentration, the higher the coefficient. The directivity parameter is greater than 0 and less than 100. The higher the value, the more concentrated the sound sources.

[0063] In practical implementation, the process of generating directional gain parameters based on sound field information can be as follows: First, acquire the main sound sources in the surrounding environment and, based on the sound field information, identify the target distance between the main sound sources and the initial microphone array. The distance between the candidate microphone array and the sound source is used to influence the gain generated by the candidate microphone array for the corresponding sound source. Second, obtain the mapping relationship between the distance and gain corresponding to the initial microphone array, and based on this mapping relationship, obtain the maximum gain that the initial microphone array can generate, and the first gain at the target distance. Third, use the ratio of the first gain to the maximum gain as the third ratio, and obtain the third adjustment coefficient corresponding to the directional gain parameter. Finally, the product of the third ratio and the third adjustment coefficient is determined as the directional gain parameter.

[0064] It should be noted that the process of acquiring the main sound sources in the surrounding environment is as described above, and will not be repeated here in this embodiment. However, the process of identifying the target distance between the main sound source and the initial microphone array based on the sound field information specifically involves extracting the sound source distribution heatmap, the location of the main sound source, and the signal amplitude characteristics in the first audio data from the sound field information. Then, using the sound wave propagation attenuation model and combining it with the sensitivity parameters of the initial microphone array, a correlation equation between amplitude and distance is established. Next, based on the sound source distribution heatmap, the location of the main sound source, and the signal amplitude characteristics in the first audio data, the measured value of the signal amplitude in the direction of the main sound source is determined. Then, the measured value of the signal amplitude in the direction of the main sound source is substituted into the correlation equation to obtain the initial distance value.

[0065] It should be noted that, due to the amplitude attenuation of sound waves with increasing distance during propagation, the gain capability of the candidate microphone array must compensate for this attenuation. The greater the distance, the greater the gain the array needs to provide. At the same time, the topological characteristics of the candidate microphone array (such as array aperture and number of microphones) determine its upper limit of gain. At the same distance, the larger the array aperture and the more microphones, the higher the achievable gain. Therefore, distance is the core parameter affecting the gain of the candidate microphone array, and the mapping relationship between distance and gain is different for different candidate microphone arrays. To obtain the mapping relationship between distance and gain for the initial microphone array, a test environment is first set up. A physical array with the same topology as the initial microphone array is selected, different test distances are set, a standard sound source signal is played at each distance, the signal amplitude of the array output is collected, and the actual gain value (i.e., the ratio of output amplitude to input amplitude) is calculated, thereby establishing the mapping relationship between distance and gain.

[0066] Then, based on the mapping relationship, the maximum gain that the initial microphone array can produce and the first gain at the target distance are obtained. The ratio of the first gain to the maximum gain is then used as the third ratio, and the third adjustment coefficient corresponding to the directional gain parameter is obtained. Finally, the product of the third ratio and the third adjustment coefficient is determined as the directional gain parameter. Since reverberation weakens the gain effect, the third adjustment coefficient is equivalent to the reverberation correction coefficient. The higher the reverberation level, the lower the coefficient. The directional gain quantization parameter is greater than 0 and less than 100. The higher the value, the stronger the array's gain capability for the sound source at the current distance.

[0067] Step 104: Based on the first weight combination, multiple sound field parameters are fused to obtain a first score, and based on the first score, a target microphone array is selected from the reference candidate microphone arrays corresponding to the initial microphone array and the first weight combination; wherein, the first score is used to indicate the degree of adaptation of the reference candidate microphone array to the surrounding environment.

[0068] It should be noted that, as mentioned above, there is a correspondence between the weight combination and the candidate microphone array. Therefore, after determining the first weight combination, the reference candidate microphone array corresponding to the first weight combination is also determined. At the same time, after determining the first weight combination, multiple sound field parameters are fused, that is, multiple sound field parameters are weighted and summed to obtain the first score. Then, based on the first score, the target microphone array is selected from the reference candidate microphone array and the initial microphone array.

[0069] It should be noted that the first score is used to indicate the degree of adaptation of the reference candidate microphone array corresponding to the first weight combination to the surrounding environment. This means that the first score is directly proportional to the degree of adaptation of the corresponding reference candidate microphone array to the surrounding environment. In other words, the higher the first score, the higher the degree of adaptation of the corresponding reference candidate microphone array to the surrounding environment. Here, in addition to the first score, the scores obtained by weighted summation of multiple sound field parameters based on the weight combination all indicate the degree of adaptation of the reference candidate microphone array corresponding to the weight combination to the surrounding environment.

[0070] In practice, the process of selecting the target microphone array from the reference candidate microphone arrays corresponding to the initial microphone array and the first weight combination based on the first score can be as follows: obtain the second weight combination corresponding to the initial microphone array, and fuse multiple sound field parameters based on the second weight combination to obtain the second score; compare the first score and the second score to obtain the comparison result; when the comparison result indicates that the first score is greater than the second score, the reference candidate microphone array is determined as the target microphone array; when the comparison result indicates that the second score is greater than or equal to the first score, the initial microphone array is determined as the target microphone array.

[0071] It should be noted that, as mentioned earlier, there is a one-to-one correspondence between the candidate microphone arrays and the weight combinations. Therefore, after determining the initial microphone array, the second weight combination corresponding to the initial microphone array can be directly obtained. Based on the second weight combination, multiple sound field parameters are fused, i.e., weighted summed, to obtain the second score. Then, the first score and the second score are compared to obtain the comparison result. When the comparison result indicates that the first score is greater than the second score, the reference candidate microphone array is determined as the target microphone array; when the comparison result indicates that the second score is greater than or equal to the first score, the initial microphone array is determined as the target microphone array.

[0072] Step 105: Perform audio sampling based on the target microphone array to obtain second audio data, and generate target audio based on the second audio data.

[0073] In practical implementation, once the target microphone array is determined, if the target microphone array is the initial microphone array, audio sampling is directly performed based on the initial microphone array in the sampling state to obtain the second audio data. If the target microphone array is a reference candidate microphone array, the reference candidate microphone array is controlled to be in the sampling state, that is, the microphone array in the sampling state is switched from the initial microphone array to the reference candidate microphone array, thereby performing audio sampling based on the reference candidate microphone array in the sampling state to obtain the second audio data. The process of performing audio sampling based on the target microphone array to obtain the second audio data is as described above in the process of performing audio sampling based on the initial microphone array to obtain the first audio data, and will not be repeated here in the embodiments of this application. Then, the process of generating the target audio based on the second audio data specifically includes converting the second audio data to obtain the target audio; wherein, the conversion of the second audio data includes format conversion (if digital audio output is required, the target audio is digital audio) or digital-to-analog conversion (if analog audio output is required, the target audio is analog audio), which will not be elaborated in the embodiments of this application.

[0074] In some embodiments, after obtaining second audio data by sampling audio based on the target microphone array, a beamforming algorithm corresponding to the target microphone array can be obtained, and the parameters in the beamforming algorithm can be adjusted based on the second audio data to generate a target beamforming algorithm; the second audio data can be adjusted based on the target beamforming algorithm to obtain third audio data; thus, the process of generating target audio based on the second audio data can be that the target audio is generated based on the third audio data.

[0075] It should be noted that the process of obtaining the beamforming algorithm corresponding to the target microphone array involves the following steps: First, a pre-defined mapping table between microphone array topology and beamforming algorithms is constructed. This table stores the appropriate beamforming algorithms (such as delay-sum algorithm, minimum variance distortionless response algorithm, adaptive beamforming algorithm, etc.) for different array topologies (linear array, circular array, rectangular array, etc.) and the number of microphones. Second, based on the mapping table, the beamforming algorithm corresponding to the target microphone array is obtained. Beamforming is an array signal processing technique that adjusts the weighting coefficients (including amplitude and phase) of each unit in a sensor array (such as an antenna or microphone) to enhance signals in a specific direction and suppress interference from non-target directions. Its core principle is to use spatial filtering to control the phase relationship of signals so that signals in the desired direction produce constructive interference, while signals in other directions are suppressed through destructive interference, thereby forming a directional beam, such as a sound wave.

[0076] Then, regarding the process of adjusting the parameters in the beamforming algorithm based on the second audio data to generate the target beamforming algorithm, specifically, firstly, key features in the second audio data (such as effective sound source azimuth, noise distribution, signal amplitude attenuation characteristics, etc.) are identified, and based on these key features, the types of parameters to be adjusted in the beamforming algorithm (delay parameters, weighting coefficients, noise suppression thresholds, gain coefficients, etc.) are determined. Secondly, corresponding adjustment methods are obtained for different parameter types. For example, the delay parameter is adjusted according to the geometric relationship between the effective sound source azimuth and the array microphone coordinates (the larger the sound source azimuth angle, the larger the delay value of the far-end microphone), and the weighting coefficient is adjusted according to the signal amplitude attenuation characteristics (the more obvious the amplitude attenuation, the higher the weighting coefficient of the corresponding microphone). Then, based on the obtained adjustment methods and the feature values ​​of the key features in the second audio data, the corresponding parameters in the beamforming algorithm are adjusted to generate the adjusted beamforming algorithm, i.e., the target beamforming algorithm.

[0077] Next, based on the target beamforming algorithm, the second audio data is adjusted to obtain the third audio data. Specifically, firstly, the second audio data is split according to the microphone channel, and based on the target beamforming algorithm, the signal processing method for the audio data corresponding to each microphone channel is determined. Secondly, the audio data corresponding to each microphone channel is processed according to the signal processing method. For example, the adjusted delay parameter is applied to the audio data corresponding to each channel to achieve signal phase alignment, and then the effective sound source signal is superimposed and summed to enhance it. At the same time, the noise signal is filtered through the noise suppression threshold. After processing, the output data of each channel is integrated, and amplitude normalization (mapping the signal amplitude to a preset dynamic range) and smoothing processing (eliminating signal glitches) are performed to obtain the third audio data.

[0078] In some embodiments, after obtaining the second audio data by sampling the audio based on the target microphone array, the second audio data can be dynamically calibrated. Specifically, this involves combining large model training to compensate for channel differences. Specifically, the initial error (including phase / amplitude error) of each channel relative to the reference channel is calculated based on the Least Mean Squares (LMS) algorithm. The reference channel is usually the microphone channel at the center of the array or the channel with the best performance calibration. Then, the calculated initial error is corrected by combining the large model pre-training compensation parameters to obtain the error compensation parameters. Finally, the error compensation parameters are written into the database.

[0079] It should be noted that the process of calculating the initial error of each channel relative to the reference channel based on the LMS algorithm is as follows: First, a reference channel is selected, and the sampled data of the reference channel is used as the desired signal, while the sampled data of the other channels are used as the input signals. The weight vector of the LMS algorithm is initialized, and the mean square error between the input signal and the desired signal is minimized through iterative calculation. That is, in each iteration, the difference between the output after the input signal weight vector is adjusted and the desired signal is calculated, and the weight vector is updated according to the difference until the mean square error converges to a preset threshold. After the iteration is completed, the phase deviation value and amplitude deviation ratio of each channel relative to the reference channel are analyzed from the converged weight vector to form the initial error. Then, the process of correcting the calculated initial error by combining the pre-trained compensation parameters of the large model to obtain the error compensation parameters involves, specifically, extracting environmental data of the current sampling scene, including environmental parameters (such as temperature 25℃, reverberation time 0.3s), target microphone array hardware batch information, etc.; inputting the extracted environmental data into the pre-trained large model to match the corresponding pre-trained compensation parameters; fusing the matched pre-trained parameters with the initial error calculated by LMS, and obtaining the final error compensation parameters through weighted correction. The correction process specifically compensates for the error deviation of the LMS algorithm caused by sampling time and local environmental interference, while covering the channel inconsistency caused by hardware batch and environmental differences. The large model is pre-trained based on massive training data, which includes channel difference data (such as phase / amplitude deviation) of different acoustic environments (reverberation, temperature, humidity) and different batches of microphone arrays. In this way, the large model learns the correlation between channel differences and environment and hardware parameters, and outputs a channel difference compensation parameter library (including deviation correction coefficient and environment adaptation factor) covering multiple scenarios.

[0080] Finally, the process of writing the error compensation parameters involves selecting the writing location based on the application level of the error compensation parameters. If real-time hardware-level compensation is required, the error compensation parameters are written to the registers of the codec chip (the codec supports hardware-level phase adjustment and gain control). If algorithm-level compensation is required, the error compensation parameters are stored in the algorithm memory area of ​​the main control unit and embedded in processing links such as beamforming and audio enhancement. Thus, after obtaining the second audio data based on audio sampling of the target microphone array, the codec chip or the main control algorithm layer automatically calls the compensation parameters to apply compensation and phase shift compensation (e.g., increasing the phase of channel 3 by 0.1 rad) and amplitude gain compensation (e.g., multiplying the amplitude of channel 5 by 1.05) to the second audio data of each channel. This ensures that the output signal of each channel maintains phase / amplitude consistency with the reference channel, completing real-time calibration of channel differences.

[0081] It should be noted that, as mentioned above, after writing the error compensation parameters, compensation can be applied to the second audio data of each channel. After applying compensation to the second audio data of each channel, the target audio can be directly generated based on the compensated second audio data. Alternatively, the parameters in the beamforming algorithm can be adjusted based on the compensated second audio data to generate the target beamforming algorithm. Then, based on the target beamforming algorithm, the second audio data can be adjusted to obtain the third audio data, and the target audio can be generated based on the third audio data. This application does not limit the specific implementation of this method.

[0082] By applying the above embodiments of this application, multiple candidate microphone arrays (including the initial array) are formed by activating a target number of microphones, providing a diverse selection space for microphone arrays for different sound fields and avoiding the limitations of fixed microphone arrays. Simultaneously, after generating sound field parameters based on the sound field features extracted from the initial microphone array, and generating a first score by combining and fusing the sound field parameters through weighted combinations, the target microphone array is selected based on the first score. Here, the first score directly indicates the degree of adaptation of the candidate microphone array to the environment. The optimal target microphone array is selected from the candidate microphone arrays based on the score, replacing the method of adjusting the microphone position using mechanical structures or manual selection. The selection process is more objective and quantitative, significantly improving adaptation accuracy, thereby enhancing the accuracy of the audio generated subsequently based on the selected target microphone array. Furthermore, the target microphone array is selected from the candidate microphone arrays, rather than activating all microphones. This avoids unnecessary microphone activation while ensuring adaptability, reducing the computational power consumption for data acquisition, transmission, and processing, and avoiding additional noise introduced by redundant microphones. Thus, not only is the audio generation efficiency improved, but the accuracy of the generated audio is also enhanced.

[0083] The following description continues to illustrate the exemplary structure of the audio generation device 455 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 2 As shown, the software module stored in the audio generation device 455 in the memory 450 may include: The startup module 4551 is used to start a target number of microphones, which are used to combine to form multiple candidate microphone arrays, including an initial microphone array. The sampling module 4552 is used to perform audio sampling based on the initial microphone array to obtain first audio data, and to extract sound field features from the first audio data to obtain sound field features. The first generation module 4553 is used to generate sound field parameters of different dimensions based on the sound field features, and to generate a first weight combination for fusing the multiple sound field parameters based on the multiple sound field parameters; wherein the weight combination has a corresponding relationship with the candidate microphone array; The fusion module 4554 is used to fuse the multiple sound field parameters based on the first weight combination to obtain a first score, and to select a target microphone array from the reference candidate microphone arrays corresponding to the initial microphone array and the first weight combination based on the first score; wherein, the first score is used to indicate the degree of adaptation of the reference candidate microphone array to the surrounding environment. The second generation module 4555 performs audio sampling based on the target microphone array to obtain second audio data, and generates target audio based on the second audio data.

[0084] In some embodiments, the sampling module 4552 is further configured to extract spatial features from the first audio data to obtain sound field spatial features, and extract temporal features from the first audio data to obtain sound field temporal features; and fuse the sound field spatial features and the sound field temporal features to obtain the sound field features.

[0085] In some embodiments, the multiple dimensions include a signal-to-noise ratio dimension, a directivity dimension, and a directivity gain dimension. The first generation module 4553 is further configured to generate sound field information of the surrounding environment based on the sound field features. The sound field information includes at least one of the main noise source location, reverberation intensity, and sound source distribution heatmap. Based on the sound field information, the module generates a signal-to-noise ratio parameter of the signal-to-noise ratio dimension, a directivity parameter of the directivity dimension, and a directivity gain parameter of the directivity gain dimension, and determines the signal-to-noise ratio parameter, the directivity parameter, and the directivity gain parameter as the sound field parameters of the different dimensions.

[0086] In some embodiments, the first generation module 4553 is further configured to identify human voice signals and noise signals in the first audio data based on the sound field information, and obtain the amplitude of the human voice signal and the amplitude of the noise signal; obtain the ratio of the amplitude of the human voice signal to the amplitude of the noise signal as a first ratio, and obtain a first adjustment coefficient corresponding to the signal-to-noise ratio parameter; and determine the signal-to-noise ratio parameter by multiplying the first ratio and the first adjustment coefficient.

[0087] In some embodiments, the first generation module 4553 is further configured to: locate multiple sound sources in the surrounding environment based on the sound field information; obtain the emission probability of each sound source; identify the main sound source and other sound sources from the multiple sound sources based on the emission probability of each sound source; sum the emission probabilities of each other sound source to obtain the total emission probability of the other sound sources, and take the ratio of the emission probability of the main sound source to the total emission probability of the other sound sources as a second ratio; obtain the second adjustment coefficient corresponding to the directivity parameter, and determine the directivity parameter by multiplying the second ratio and the second adjustment coefficient.

[0088] In some embodiments, the first generation module 4553 is further configured to acquire the main sound source in the surrounding environment and, based on the sound field information, identify the target distance between the main sound source and the initial microphone array; wherein, the distance between the candidate microphone array and the sound source is used to affect the gain generated by the candidate microphone array for the corresponding sound source; acquire the mapping relationship between the distance and gain corresponding to the initial microphone array, and, based on the mapping relationship, acquire the maximum gain that the initial microphone array can generate and the first gain at the target distance; take the ratio of the first gain to the maximum gain as the third ratio, and acquire the third adjustment coefficient corresponding to the directional gain parameter; and determine the directional gain parameter by multiplying the third ratio and the third adjustment coefficient.

[0089] In some embodiments, the fusion module 4554 is further configured to obtain a second weight combination corresponding to the initial microphone array, and fuse the plurality of sound field parameters based on the second weight combination to obtain a second score; compare the first score and the second score to obtain a comparison result; when the comparison result indicates that the first score is greater than the second score, determine the reference candidate microphone array as the target microphone array; when the comparison result indicates that the second score is greater than or equal to the first score, determine the initial microphone array as the target microphone array.

[0090] In some embodiments, the first generation module 4553 is further configured to identify a target sound field scene corresponding to the surrounding environment based on multiple sound field parameters; obtain a preset correspondence between sound field scenes and weight combinations; and obtain the first weight combination corresponding to the target sound field scene based on the correspondence.

[0091] In some embodiments, the apparatus further includes an acquisition module, which is configured to acquire a beamforming algorithm corresponding to the target microphone array, and adjust the parameters in the beamforming algorithm based on the second audio data to generate a target beamforming algorithm; adjust the second audio data based on the target beamforming algorithm to obtain third audio data; the second generation module 4555 is further configured to generate target audio based on the third audio data.

[0092] This application provides a computer program product, which includes computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the audio generation method provided in this application.

[0093] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the audio generation method provided in this application. For example, ... Figure 3 The audio generation method is shown.

[0094] In some embodiments, the computer-readable storage medium may be a read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic surface memory, optical disk, or CD-ROM, etc.; or it may be a device that includes one or any combination of the above-mentioned memories.

[0095] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.

[0096] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).

[0097] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located in one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.

[0098] It should be noted that in this application embodiment, data such as text is involved. When this application embodiment is applied to a specific product or technology, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0099] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. An audio generation method, characterized in that, The method includes: A target number of microphones are activated, which are used to combine to form a plurality of candidate microphone arrays, the plurality of candidate microphone arrays including an initial microphone array. Audio sampling is performed based on the initial microphone array to obtain first audio data, and sound field features are extracted from the first audio data to obtain sound field features. Based on the sound field features, sound field parameters of different dimensions are generated, and based on multiple sound field parameters, a first weight combination for fusing the multiple sound field parameters is generated; wherein, the weight combination has a corresponding relationship with the candidate microphone array; Based on the first weight combination, the multiple sound field parameters are fused to obtain a first score, and based on the first score, a target microphone array is selected from the reference candidate microphone arrays corresponding to the initial microphone array and the first weight combination. The first score is used to indicate the degree of adaptation of the reference candidate microphone array to the surrounding environment. Audio data is obtained by sampling the target microphone array to obtain second audio data, and target audio is generated based on the second audio data.

2. The method according to claim 1, characterized in that, The step of extracting sound field features from the first audio data to obtain sound field features includes: Spatial features are extracted from the first audio data to obtain sound field spatial features, and temporal features are extracted from the first audio data to obtain sound field temporal features. The sound field spatial features are fused with the sound field temporal features to obtain the sound field features.

3. The method according to claim 1, characterized in that, Multiple dimensions include signal-to-noise ratio, directivity, and directivity gain. Based on the sound field characteristics, sound field parameters of different dimensions are generated, including: Based on the sound field characteristics, sound field information of the surrounding environment is generated, and the sound field information includes at least one of the following: the location of the main noise source, the reverberation intensity, and the sound source distribution heat map. Based on the sound field information, a signal-to-noise ratio parameter in the signal-to-noise ratio dimension, a directivity parameter in the directivity dimension, and a directivity gain parameter in the directivity gain dimension are generated, and the signal-to-noise ratio parameter, the directivity parameter, and the directivity gain parameter are determined as the sound field parameters of the different dimensions.

4. The method according to claim 3, characterized in that, Based on the sound field information, the signal-to-noise ratio (SNR) parameters of the SNR dimension are generated, including: Based on the sound field information, identify the human voice signal and noise signal in the first audio data, and obtain the amplitude of the human voice signal and the amplitude of the noise signal; The ratio of the amplitude of the human voice signal to the amplitude of the noise signal is obtained as a first ratio, and the first adjustment coefficient corresponding to the signal-to-noise ratio parameter is obtained. The product of the first ratio and the first adjustment coefficient is determined as the signal-to-noise ratio parameter.

5. The method according to claim 3, characterized in that, Based on the sound field information, the directional parameters of the directional dimension are generated, including: Based on the sound field information, multiple sound sources in the surrounding environment are located, and the emission probability of each sound source is obtained. Based on the emission probability of each of the sound sources, the main sound source and other sound sources are identified from the plurality of sound sources; The emission probabilities of each of the other sound sources are summed to obtain the total emission probability of the other sound sources, and the ratio of the emission probability of the main sound source to the total emission probability of the other sound sources is used as the second ratio. Obtain the second adjustment coefficient corresponding to the directional parameter, and multiply the second ratio by the second adjustment coefficient to determine the directional parameter.

6. The method according to claim 3, characterized in that, Based on the sound field information, the directional gain parameters of the directional gain dimension are generated, including: The main sound sources in the surrounding environment are acquired, and based on the sound field information, the target distance between the main sound sources and the initial microphone array is identified; wherein, the distance between the candidate microphone array and the sound source is used to affect the gain generated by the candidate microphone array for the corresponding sound source; Obtain the mapping relationship between distance and gain corresponding to the initial microphone array, and based on the mapping relationship, obtain the maximum gain that the initial microphone array can produce, and the first gain at the target distance; The ratio of the first gain to the maximum gain is used as the third ratio, and the third adjustment coefficient corresponding to the directional gain parameter is obtained. The product of the third ratio and the third adjustment coefficient is determined as the directional gain parameter.

7. The method according to claim 1, characterized in that, The step of selecting a target microphone array from the reference candidate microphone arrays corresponding to the initial microphone array and the first weight combination based on the first score includes: Obtain the second weight combination corresponding to the initial microphone array, and fuse the multiple sound field parameters based on the second weight combination to obtain the second score; The first score and the second score are compared to obtain the comparison result; When the comparison result indicates that the first score is greater than the second score, the reference candidate microphone array is determined as the target microphone array; When the comparison result indicates that the second score is greater than or equal to the first score, the initial microphone array is determined as the target microphone array.

8. The method according to claim 1, characterized in that, The step of generating a first weighted combination for fusing the multiple sound field parameters based on multiple sound field parameters includes: Based on multiple sound field parameters, the target sound field scene corresponding to the surrounding environment is identified; Obtain the correspondence between preset sound field scenes and weight combinations, and based on the correspondence, obtain the first weight combination corresponding to the target sound field scene.

9. The method according to claim 1, characterized in that, After obtaining the second audio data by performing audio sampling based on the target microphone array, the method further includes: Obtain the beamforming algorithm corresponding to the target microphone array, and adjust the parameters in the beamforming algorithm based on the second audio data to generate the target beamforming algorithm; Based on the target beamforming algorithm, the second audio data is adjusted to obtain the third audio data; The step of generating the target audio based on the second audio data includes: The target audio is generated based on the third audio data.

10. An electronic device, characterized in that, include: Memory is used to store executable instructions or computer programs. A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the audio generation method according to any one of claims 1 to 9.

11. A computer-readable storage medium, characterized in that, It stores computer-executable instructions or computer programs for causing a processor to execute, thereby implementing the audio generation method according to any one of claims 1 to 9.