A microphone optimization control method and system based on scene adaptation
By collecting and separating audio features through microphone arrays, matching sound source object types and predicting future audio features, it solves the problem that the existing technology is difficult to cope with complex noise scenarios in an open environment, and realizes real-time and effective noise reduction processing.
Patent Information
- Application Number
- CN202510127738.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-05
AI Technical Summary
The prior art is difficult to effectively deal with complex ambient noise scenarios in open environments, especially fixed noise reduction presets, which are difficult to adapt to uncertain sound sources in non-studio environments.
Ambient audio acquisition is performed through a microphone array, audio data is obtained in real time, and multiple sets of audio characteristics are obtained through differential separation. Based on these features, the sound source object type is matched, and future audio features are predicted to achieve real-time noise reduction processing.
It realizes effective processing of complex noise scenes in non-studio environments, can adapt to environmental changes in real time, improve noise reduction quality, and make up for the shortcomings of fixed noise reduction presets when dealing with uncommon noise events.
Smart Images

Figure CN119562180B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field related to microphone optimization control, and in particular to a microphone optimization control method and system based on scene adaptation. Background Art
[0002] When audio is collected and recorded in an open environment, the existence of uncertain sound sources in the environment will cause the recorded content to be mixed with a lot of non-target audio noise. For these noises, noise reduction processing is usually performed based on needs to make the required main sound more prominent and optimize the overall listening experience.
[0003] There are two commonly used noise reduction methods in the prior art. One is to optimize the noise reduction by post-processing. This type of method can ensure a high noise reduction quality, but it requires a long time for post-processing and is not time-sensitive. The other method uses a fixed noise reduction preset to suppress the ambient noise and optimize the collected content during collection. This method has a timeliness different from the aforementioned content and can be used in real-time video and audio content interaction scenarios. However, the fixed noise reduction preset is still difficult to cope with complex ambient sound scenes when actually used. Summary of the invention
[0004] The object of the present invention is to provide a microphone optimization control method and system based on scene adaptation to solve the problems raised in the above background technology.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A microphone optimization control method based on scene adaptation, comprising:
[0007] Based on the microphone array, the ambient audio is collected to obtain audio data in real time, and the audio data is differentially separated based on the microphone array to obtain multiple groups of audio features, wherein the differential separation is determined by the signal strength and sound source direction of multiple microphones in the array;
[0008] When the audio feature is separated, searching is performed based on the audio feature to obtain a matching feature object, wherein the feature object is used to characterize a sound source object type corresponding to the audio feature;
[0009] Performing audio statistics on feature objects through big data to fit and obtain multiple object audio features, wherein the multiple object audio features are used to characterize the audio and change characteristics of the feature object under different sound emission modes;
[0010] Sequentially fitting the audio features with the multiple groups of the object audio features based on the frame nodes, continuously screening and matching the matching object audio features, and calculating the comprehensive mean of the multiple object audio features within a preset time period to establish the estimated audio features within the prediction time period;
[0011] The time nodes waiting to be updated are pre-assigned noise reduction values according to the estimated audio features, and the audio data is subjected to scene noise reduction processing based on the pre-assigned noise reduction values when the time nodes are updated.
[0012] As a further solution of the present invention: it also includes a gain control management step:
[0013] When collecting ambient audio, the current real-time ambient volume is judged. If the ambient volume reaches the warning threshold, basic gain management is triggered to control the ambient volume within a safe threshold. The warning threshold includes an overly high threshold and an overly low threshold. The gain is used to achieve volume control through sensitive adjustment of the microphone.
[0014] Based on the estimated audio features, a pre-judgment of the ambient collection volume for a preset time period is performed. If the judgment result indicates that the ambient collection volume at a certain time node in the future reaches the warning threshold, the evaluation gain management is triggered accordingly to generate a pre-management signal. The pre-management signal is used to control the gain control to be induced and executed in advance when the actual collection time node reaches the corresponding predicted time node.
[0015] As a further solution of the present invention: the gain control management step further includes:
[0016] Matching multiple groups of audio features by pre-recorded subject audio features of a subject object to determine the recorded subject audio and the relative position information of the recording, wherein the subject object is the audio recording object during the recording process;
[0017] Real-time dynamic gain is performed based on the real-time volume of the recorded subject audio and the collected relative position information to control the recorded subject audio to be within a preset optimal range.
[0018] As a further solution of the present invention: it also includes a distortion management step, specifically including:
[0019] Audio distortion is judged based on the audio data obtained by real-time environment collection, and several distributed distortion nodes are obtained;
[0020] Matching the distortion node with the audio data after noise reduction and gain, obtaining the audio data before and after the distortion node, and performing natural language recognition on the audio data to obtain its language content;
[0021] Performing content retrieval on the language content before and after the distortion node through big data to fit and obtain the distorted language content corresponding to the distortion node;
[0022] A text-to-speech conversion model is obtained according to the main audio feature training, and the distorted language content is converted into speech based on the text-to-speech conversion model to correct the distorted content of the audio data.
[0023] As a further solution of the present invention: it also includes a scene twin optimization step:
[0024] Record the current location information in real time and synchronize it with the digital scene model;
[0025] When fitting to a feature object corresponding to an audio feature, the motion feature of the feature object is obtained through big data, and a twin trajectory is marked in the digital scene model based on the relative position information of the current audio feature, wherein the twin trajectory includes a motion direction and a motion time distribution;
[0026] When the audio features are collected and fitted again, the position relationship of the feature object relative to the collection device at consecutive time nodes is calculated based on the digital scene model;
[0027] Based on the positional relationship, accurate gain assignment calculation is performed on the object audio feature to obtain the gain assignment of each time node in a continuous time period, wherein the gain assignment includes a compensation coefficient for the volume change caused by the change in the relative positional relationship.
[0028] The embodiment of the present invention aims to provide a microphone optimization control system based on scene adaptation, comprising:
[0029] The acquisition and separation module is used to acquire ambient audio based on a microphone array, obtain audio data in real time, and perform differential separation on the audio data based on the microphone array to obtain multiple groups of audio features. The differential separation is determined by the signal strength and sound source direction of multiple microphones in the array.
[0030] A feature matching module, used for searching based on the audio feature when the audio feature is separated, and obtaining a matching feature object, wherein the feature object is used to characterize a sound source object type corresponding to the audio feature;
[0031] A feature statistics module is used to perform audio statistics on feature objects through big data to fit and obtain multiple object audio features, wherein the multiple object audio features are used to characterize the audio and change characteristics of the feature object under different sound emission modes;
[0032] A sequential prediction module, used to sequentially fit the audio features with the multiple groups of the object audio features based on the frame nodes, continuously screen and match the matching object audio features, and calculate the comprehensive mean of the multiple object audio features within a preset time period to establish the estimated audio features within the prediction time period;
[0033] The noise reduction gain module is used to pre-assign noise reduction values to time nodes waiting to be updated according to the estimated audio features, and perform scene noise reduction processing on the audio data based on the noise reduction pre-assignment when the time nodes are updated.
[0034] As a further solution of the present invention: it also includes a gain control management module:
[0035] The basic gain unit is used to judge the current real-time environment collection volume when collecting ambient audio. If the environment collection volume reaches the warning threshold, the basic gain management is triggered to control the environment collection volume to be within the safety threshold. The warning threshold includes too high and too low thresholds. The gain is used to realize the collection volume control through the sensitivity adjustment of the microphone;
[0036] The prediction gain unit is used to pre-judge the environment collection volume in a preset time period based on the estimated audio features. If the judgment result indicates that the environment collection volume at a certain time node in the future reaches the warning threshold, the evaluation gain management is triggered accordingly to generate a pre-management signal. The pre-management signal is used to control the induction and execution of gain control in advance when the actual collection time node reaches the corresponding prediction time node.
[0037] As a further solution of the present invention: the gain control management module further includes:
[0038] A subject matching unit, used to match multiple groups of audio features with pre-recorded subject audio features of a subject object to determine the recorded subject audio and the relative position information of the recording, wherein the subject object is the audio recording object during the recording process;
[0039] The gain maintenance unit is used to perform real-time dynamic gain based on the real-time volume of the recorded subject audio and the collected relative position information to control the recorded subject audio to be within a preset optimal range.
[0040] As a further solution of the present invention: it also includes a distortion management module:
[0041] A distortion judgment unit, used to judge audio distortion based on audio data acquired through real-time environment collection, and obtain a plurality of distributed distortion nodes;
[0042] A node mapping unit, used to match the distortion node with the audio data after noise reduction and gain, obtain the audio data before and after the distortion node, and perform natural language recognition on the audio data to obtain its language content;
[0043] A content association unit, used to perform content retrieval on the language content before and after the distorted node through big data, so as to fit and obtain the distorted language content corresponding to the distorted node;
[0044] The distortion correction unit is used to obtain a text-to-speech conversion model according to the main audio feature training, and perform speech conversion on the distorted language content based on the text-to-speech conversion model for correcting the distorted content of the audio data.
[0045] As a further solution of the present invention: it also includes a scene twin optimization module:
[0046] Twin synchronization unit, used to record the current position information in real time and perform twin synchronization within the digital scene model;
[0047] An object marking unit is used to obtain the motion characteristics of the feature object through big data when fitting to the feature object corresponding to the audio feature, and mark the twin trajectory in the digital scene model based on the relative position information of the current audio feature, wherein the twin trajectory includes the motion direction and the motion time distribution;
[0048] A relative evaluation unit, configured to calculate, based on the digital scene model, the positional relationship of the feature object relative to the acquisition device at consecutive time nodes when the audio feature is collected and fitted again;
[0049] The gain refinement unit is used to accurately calculate the gain assignment of the object audio feature based on the position relationship, and obtain the gain assignment of each time node in a continuous time period, wherein the gain assignment includes a compensation coefficient for the volume change caused by the change in relative position relationship.
[0050] Compared with the prior art, the beneficial effects of the present invention are: it is suitable for microphone sound collection control and optimization in non-recording studio environments, and matches and fits the characteristics and intensity possibilities of ambient sounds originating from the object in the future by judging the type of sound source object based on separated audio features, and performs preprocessing based on the possibility, matches targeted noise reduction assignment schemes in advance, and executes them at subsequent time nodes. Through such a control method, during the acquisition process, real-time noise reduction can be achieved to meet the standards for live call interaction, and more complex environmental noise scenarios can be coped with, thereby making up for the inefficiency of fixed noise reduction preset schemes for uncommon noise events. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1The present invention is a flowchart of a microphone optimization control method based on scene adaptation.
[0052] Figure 2 A flowchart of the distortion management step in a microphone optimization control method based on scene adaptation.
[0053] Figure 3 The block diagram of a microphone optimization control system based on scene adaptation. DETAILED DESCRIPTION
[0054] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0055] The specific implementation of the present invention is described in detail below in conjunction with specific embodiments.
[0056] like Figure 1 The microphone optimization control method based on scene adaptation provided by an embodiment of the present invention comprises the following steps:
[0057] S10, collecting ambient audio based on a microphone array, acquiring audio data in real time, and performing differential separation on the audio data based on the microphone array to obtain multiple groups of audio features, wherein the differential separation is determined by the signal strength and sound source direction of multiple microphones in the array;
[0058] S20, when the audio feature is separated, searching is performed based on the audio feature to obtain a matching feature object, where the feature object is used to characterize a sound source object type corresponding to the audio feature;
[0059] S30, performing audio statistics on the feature object through big data to fit and obtain multiple object audio features, wherein the multiple object audio features are respectively used to characterize the audio and change characteristics of the feature object in different sound emission modes;
[0060] S40, sequentially fitting the audio features with the multiple groups of the object audio features based on the frame nodes, continuously screening and matching the matching object audio features, and calculating the comprehensive mean of the multiple object audio features within a preset time period to establish the estimated audio features within the prediction time period;
[0061] S50, pre-assigning noise reduction values to time nodes waiting to be updated according to the estimated audio features, and performing scene noise reduction processing on the audio data based on the pre-assigned noise reduction values when the time nodes are updated.
[0062] In this embodiment, a microphone optimization control method based on scene adaptation is provided, which is suitable for microphone sound collection control and optimization in a non-recording studio environment. By judging the type of the sound source object based on the separated audio features, the characteristics and intensity possibilities of the ambient sound originating from the object in the future are matched and fitted, and preprocessing is performed based on the possibility, and a targeted noise reduction assignment scheme is matched in advance and executed at a subsequent time node. Through such a control method, during the acquisition process, real-time noise reduction is achieved to meet the interactive use standards of live calls, and more complex environmental noise scenes can be coped with, thereby compensating for the inefficiency of fixed noise reduction preset schemes for uncommon noise events. When audio acquisition and recording is performed in an open environment, due to the existence of uncertain sound source factors in the environment, the recorded content will be mixed with a lot of non-target audio noise. For these noises, there are two methods commonly used in the prior art. One is to optimize the noise reduction by post-processing, which can ensure a high noise reduction quality; the other method is to use a fixed noise reduction preset, which suppresses the environmental noise and optimizes the collected content by preset during acquisition. The method has a timeliness different from the aforementioned content, and can be used in real-time video and audio content interaction scenarios, but the fixed noise reduction preset is still difficult to cope with complex ambient sound scenes when actually used. The solution of this embodiment is: after collecting audio of continuous time frames, the array effect of multiple microphones can be used to determine the sound source and split the sound source channels in the audio data, and obtain multiple groups of audio features. For different audio features, the type of the sound source object (such as trains, sirens, animal calls, etc.) can be determined by retrieval matching, and these sound source objects usually have certain sound regularity characteristics. Therefore, by fitting and matching multiple object audio features of the sound source object type, the sound change in the future period can be accurately predicted. Based on the predicted sound change (i.e., the estimated audio feature), a series of estimated assignments for noise reduction can be generated (updated in real time as the recording collection process proceeds), and executed when the corresponding time node is reached, thereby realizing real-time and specific noise reduction execution management. Compared with the existing noise reduction solution, it has multiple advantages such as real-time, high efficiency and coping with complex environments.
[0063] As another preferred embodiment of the present invention, it also includes a gain control management step:
[0064] When collecting ambient audio, the current real-time ambient volume is judged. If the ambient volume reaches the warning threshold, basic gain management is triggered to control the ambient volume within a safe threshold. The warning threshold includes an overly high threshold and an overly low threshold. The gain is used to achieve volume control through sensitive adjustment of the microphone.
[0065] Based on the estimated audio features, a pre-judgment of the ambient collection volume for a preset time period is performed. If the judgment result indicates that the ambient collection volume at a certain time node in the future reaches the warning threshold, the evaluation gain management is triggered accordingly to generate a pre-management signal. The pre-management signal is used to control the gain control to be induced and executed in advance when the actual collection time node reaches the corresponding predicted time node.
[0066] In this embodiment, the step of automatic gain control (AGC) is added. The basis for achieving the purpose of effective recording and noise reduction management is that the microphone array can collect effective and recognizable audio data. In order to be suitable for environments with different sound intensities, the microphone is usually of controllable sensitivity. Therefore, when it is at a certain sensitivity, if high-intensity sound is generated in the external environment, it may cause the collected signal intensity to exceed the microphone's recognizable threshold, resulting in signal distortion. Therefore, automatic gain control is required. When the sound in the external environment begins to rise and touches the control threshold (the control threshold is much smaller than the distortion threshold), dynamic range compression is required to reduce the gain when the sound is too loud and avoid overload distortion. Similarly, when the sound is too small, the gain is increased.
[0067] As another preferred embodiment of the present invention, the gain control management step further includes:
[0068] Matching multiple groups of audio features by pre-recorded subject audio features of a subject object to determine the recorded subject audio and the relative position information of the recording, wherein the subject object is the audio recording object during the recording process;
[0069] Real-time dynamic gain is performed based on the real-time volume of the recorded subject audio and the collected relative position information to control the recorded subject audio to be within a preset optimal range.
[0070] In this embodiment, based on the previous embodiment, the frequent noise reduction and gain adjustment process will cause the volume of the main object captured by the recording to fluctuate, making the acquired audio unstable. Therefore, in order to output smooth audio content, it is necessary to perform variable gain on the audio data of the subject so that it is always maintained within an appropriate range to obtain the optimal auditory feedback effect.
[0071] like Figure 2 As shown, as another preferred embodiment of the present invention, it also includes a distortion management step, specifically including:
[0072] S61, performing audio distortion judgment based on the audio data obtained by real-time environment collection, and obtaining a plurality of distributed distortion nodes;
[0073] S62, matching the distortion node with the audio data after noise reduction and gain, obtaining the audio data before and after the distortion node, and performing natural language recognition on the audio data to obtain its language content;
[0074] S63, performing content retrieval on the language content before and after the distorted node through big data to fit and obtain the distorted language content corresponding to the distorted node;
[0075] S64, obtaining a text-to-speech conversion model according to the main audio feature training, and performing speech conversion on the distorted language content based on the text-to-speech conversion model for use in correcting the distorted content of the audio data.
[0076] In this embodiment, in the gain management of the aforementioned embodiment, its applicable scenario is still for controllable sound source objects. As for short, instantaneous high-peak environmental noise, it is easy to exceed the threshold because it directly reaches the highest peak at the same time of instantaneous occurrence, resulting in temporary distortion of the collected audio data, that is, the specific audio data content cannot be effectively identified, affecting the quality of the audio data. The method adopted here is the natural language big data recognition and matching method. By identifying the language content before and after the distorted node, the missing voice content at the node can be reversely judged, and the pronunciation can be simulated by AI training pronunciation to fill the distorted voice content. If it is used in recording scenarios such as live broadcasts for real-time interaction, this distortion repair process can be achieved by setting a short delay time, thereby further optimizing the audio interaction experience.
[0077] As another preferred embodiment of the present invention, the following step of scene twin optimization is also included:
[0078] Record the current location information in real time and synchronize it with the digital scene model;
[0079] When fitting to a feature object corresponding to an audio feature, the motion feature of the feature object is obtained through big data, and a twin trajectory is marked in the digital scene model based on the relative position information of the current audio feature, wherein the twin trajectory includes a motion direction and a motion time distribution;
[0080] When the audio features are collected and fitted again, the position relationship of the feature object relative to the collection device at consecutive time nodes is calculated based on the digital scene model;
[0081] Based on the positional relationship, accurate gain assignment calculation is performed on the object audio feature to obtain the gain assignment of each time node in a continuous time period, wherein the gain assignment includes a compensation coefficient for the volume change caused by the change in the relative positional relationship.
[0082] In this embodiment, it is suitable for relatively controllable and certain scenes, such as train stations, where the appearance of sound source objects is predictable and has a certain regularity. For example, in a train station, the running trajectory and sound emission pattern of the train are predictable. Therefore, by recording the motion trajectory of the collector, its relative position relationship with the motion trajectory of the train is determined, so when the sound of the train is detected and begins to approach, a series of gain assignments are generated to cover any position and time node of the train movement, so as to perform accurate and high-quality noise reduction on the audio data of the train.
[0083] like Figure 3 As shown, the present invention also provides a microphone optimization control system based on scene adaptation, which comprises:
[0084] The collection and separation module 100 is used to collect ambient audio based on a microphone array, obtain audio data in real time, and perform differential separation on the audio data based on the microphone array to obtain multiple groups of audio features. The differential separation is determined by the signal strength and sound source direction of multiple microphones in the array;
[0085] A feature matching module 200 is used to perform a search based on the audio feature when the audio feature is separated, and obtain a matching feature object, wherein the feature object is used to characterize a sound source object type corresponding to the audio feature;
[0086] A feature statistics module 300 is used to perform audio statistics on a feature object through big data to fit and obtain multiple object audio features, wherein the multiple object audio features are used to characterize the audio and change characteristics of the feature object under different sound emission modes;
[0087] A sequential prediction module 400 is used to sequentially fit the audio features with the multiple groups of the object audio features based on the frame nodes, continuously screen and match the matching object audio features, and calculate the comprehensive mean of the multiple object audio features within a preset time period to establish the estimated audio features within the prediction time period;
[0088] The noise reduction gain module 500 is used to pre-assign noise reduction values to time nodes waiting to be updated according to the estimated audio features, and perform scene noise reduction processing on audio data based on the noise reduction pre-assignment when the time nodes are updated.
[0089] As another preferred embodiment of the present invention, it also includes a gain control management module:
[0090] The basic gain unit is used to judge the current real-time environment collection volume when collecting ambient audio. If the environment collection volume reaches the warning threshold, the basic gain management is triggered to control the environment collection volume to be within the safety threshold. The warning threshold includes too high and too low thresholds. The gain is used to realize the collection volume control through the sensitivity adjustment of the microphone;
[0091] The prediction gain unit is used to pre-judge the environment collection volume in a preset time period based on the estimated audio features. If the judgment result indicates that the environment collection volume at a certain time node in the future reaches the warning threshold, the evaluation gain management is triggered accordingly to generate a pre-management signal. The pre-management signal is used to control the induction and execution of gain control in advance when the actual collection time node reaches the corresponding prediction time node.
[0092] As another preferred embodiment of the present invention, the gain control management module further includes:
[0093] A subject matching unit, used to match multiple groups of audio features with pre-recorded subject audio features of a subject object to determine the recorded subject audio and the relative position information of the recording, wherein the subject object is the audio recording object during the recording process;
[0094] The gain maintenance unit is used to perform real-time dynamic gain based on the real-time volume of the recorded subject audio and the collected relative position information to control the recorded subject audio to be within a preset optimal range.
[0095] As another preferred embodiment of the present invention, it also includes a distortion management module:
[0096] A distortion judgment unit, used to judge audio distortion based on audio data acquired through real-time environment collection, and obtain a plurality of distributed distortion nodes;
[0097] A node mapping unit, used to match the distortion node with the audio data after noise reduction and gain, obtain the audio data before and after the distortion node, and perform natural language recognition on the audio data to obtain its language content;
[0098] A content association unit, used to perform content retrieval on the language content before and after the distorted node through big data, so as to fit and obtain the distorted language content corresponding to the distorted node;
[0099] The distortion correction unit is used to obtain a text-to-speech conversion model according to the main audio feature training, and perform speech conversion on the distorted language content based on the text-to-speech conversion model for correcting the distorted content of the audio data.
[0100] As another preferred embodiment of the present invention, it also includes a scene twin optimization module:
[0101] Twin synchronization unit, used to record the current position information in real time and perform twin synchronization within the digital scene model;
[0102] An object marking unit is used to obtain the motion characteristics of the feature object through big data when fitting to the feature object corresponding to the audio feature, and mark the twin trajectory in the digital scene model based on the relative position information of the current audio feature, wherein the twin trajectory includes the motion direction and the motion time distribution;
[0103] A relative evaluation unit, configured to calculate, based on the digital scene model, the positional relationship of the feature object relative to the acquisition device at consecutive time nodes when the audio feature is collected and fitted again;
[0104] The gain refinement unit is used to accurately calculate the gain assignment of the object audio feature based on the position relationship, and obtain the gain assignment of each time node in a continuous time period, wherein the gain assignment includes a compensation coefficient for the volume change caused by the change in relative position relationship.
[0105] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0106] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the disclosure in the specification and examples. This application is intended to cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the claims.
[0107] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A microphone optimization control method based on scene adaptation, characterized in that: Include: Based on the microphone array, the ambient audio is collected to obtain audio data in real time, and the audio data is differentially separated based on the microphone array to obtain multiple groups of audio features, wherein the differential separation is determined by the signal strength and sound source direction of multiple microphones in the array; When the audio feature is separated, searching is performed based on the audio feature to obtain a matching feature object, wherein the feature object is used to characterize a sound source object type corresponding to the audio feature; Performing audio statistics on feature objects through big data to fit and obtain multiple object audio features, wherein the multiple object audio features are used to characterize the audio and change characteristics of the feature object under different sound emission modes; Sequentially fitting the audio features with the multiple groups of the object audio features based on the frame nodes, continuously screening and matching the matching object audio features, and calculating the comprehensive mean of the multiple object audio features within a preset time period to establish the estimated audio features within the prediction time period; The time nodes waiting to be updated are pre-assigned noise reduction values according to the estimated audio features, and the audio data is subjected to scene noise reduction processing based on the pre-assigned noise reduction values when the time nodes are updated.
2. The microphone optimization control method based on scene adaptation according to claim 1, characterized in that: Also included are the gain control management steps: When collecting ambient audio, the current real-time ambient volume is judged. If the ambient volume reaches the warning threshold, basic gain management is triggered to control the ambient volume within a safe threshold. The warning threshold includes an overly high threshold and an overly low threshold. The gain is used to achieve volume control through sensitive adjustment of the microphone. Based on the estimated audio features, a pre-judgment of the ambient collection volume for a preset time period is performed. If the judgment result indicates that the ambient collection volume at a certain time node in the future reaches the warning threshold, the evaluation gain management is triggered accordingly to generate a pre-management signal. The pre-management signal is used to control the gain control to be induced and executed in advance when the actual collection time node reaches the corresponding predicted time node.
3. The microphone optimization control method based on scene adaptation according to claim 2, characterized in that: The gain control management step further includes: Matching multiple groups of audio features by pre-recorded subject audio features of a subject object to determine the recorded subject audio and the relative position information of the recording, wherein the subject object is the audio recording object during the recording process; Real-time dynamic gain is performed based on the real-time volume of the recorded subject audio and the collected relative position information to control the recorded subject audio to be within a preset optimal range.
4. The microphone optimization control method based on scene adaptation according to claim 3, characterized in that: It also includes distortion management steps, including: Audio distortion is judged based on the audio data obtained by real-time environment collection, and several distributed distortion nodes are obtained; Matching the distortion node with the audio data after noise reduction and gain, obtaining the audio data before and after the distortion node, and performing natural language recognition on the audio data to obtain its language content; Performing content retrieval on the language content before and after the distortion node through big data to fit and obtain the distorted language content corresponding to the distortion node; A text-to-speech conversion model is obtained according to the main audio feature training, and the distorted language content is converted into speech based on the text-to-speech conversion model to correct the distorted content of the audio data.
5. The microphone optimization control method based on scene adaptation according to claim 2, characterized in that: It also includes scene twin optimization steps: Record the current location information in real time and synchronize it with the digital scene model; When fitting to a feature object corresponding to an audio feature, the motion feature of the feature object is obtained through big data, and a twin trajectory is marked in the digital scene model based on the relative position information of the current audio feature, wherein the twin trajectory includes a motion direction and a motion time distribution; When the audio features are collected and fitted again, the position relationship of the feature object relative to the collection device at consecutive time nodes is calculated based on the digital scene model; Based on the positional relationship, accurate gain assignment calculation is performed on the object audio feature to obtain the gain assignment of each time node in a continuous time period, wherein the gain assignment includes a compensation coefficient for the volume change caused by the change in the relative positional relationship.
6. A microphone optimization control system based on scene adaptation, characterized in that: Include: The acquisition and separation module is used to acquire ambient audio based on a microphone array, obtain audio data in real time, and perform differential separation on the audio data based on the microphone array to obtain multiple groups of audio features. The differential separation is determined by the signal strength and sound source direction of multiple microphones in the array. A feature matching module, used for searching based on the audio feature when the audio feature is separated, and obtaining a matching feature object, wherein the feature object is used to characterize a sound source object type corresponding to the audio feature; A feature statistics module is used to perform audio statistics on feature objects through big data to fit and obtain multiple object audio features, wherein the multiple object audio features are used to characterize the audio and change characteristics of the feature object under different sound emission modes; A sequential prediction module, used to sequentially fit the audio features with the multiple groups of the object audio features based on the frame nodes, continuously screen and match the matching object audio features, and calculate the comprehensive mean of the multiple object audio features within a preset time period to establish the estimated audio features within the prediction time period; The noise reduction gain module is used to pre-assign noise reduction values to time nodes waiting to be updated according to the estimated audio features, and perform scene noise reduction processing on the audio data based on the noise reduction pre-assignment when the time nodes are updated.
7. The microphone optimization control system based on scene adaptation according to claim 6, characterized in that: Also includes gain control management module: The basic gain unit is used to judge the current real-time environment collection volume when collecting ambient audio. If the environment collection volume reaches the warning threshold, the basic gain management is triggered to control the environment collection volume to be within the safety threshold. The warning threshold includes too high and too low thresholds. The gain is used to realize the collection volume control through the sensitivity adjustment of the microphone; The prediction gain unit is used to pre-judge the environment collection volume in a preset time period based on the estimated audio features. If the judgment result indicates that the environment collection volume at a certain time node in the future reaches the warning threshold, the evaluation gain management is triggered accordingly to generate a pre-management signal. The pre-management signal is used to control the induction and execution of gain control in advance when the actual collection time node reaches the corresponding prediction time node.
8. The microphone optimization control system based on scene adaptation according to claim 7, characterized in that: The gain control management module also includes: A subject matching unit, used to match multiple groups of audio features with pre-recorded subject audio features of a subject object to determine the recorded subject audio and the relative position information of the recording, wherein the subject object is the audio recording object during the recording process; The gain maintenance unit is used to perform real-time dynamic gain based on the real-time volume of the recorded subject audio and the collected relative position information to control the recorded subject audio to be within a preset optimal range.
9. The microphone optimization control system based on scene adaptation according to claim 8, characterized in that: Also includes distortion management modules: A distortion judgment unit, used to judge audio distortion based on audio data acquired through real-time environment collection, and obtain a plurality of distributed distortion nodes; A node mapping unit, used to match the distortion node with the audio data after noise reduction and gain, obtain the audio data before and after the distortion node, and perform natural language recognition on the audio data to obtain its language content; A content association unit, used to perform content retrieval on the language content before and after the distorted node through big data, so as to fit and obtain the distorted language content corresponding to the distorted node; The distortion correction unit is used to obtain a text-to-speech conversion model according to the main audio feature training, and perform speech conversion on the distorted language content based on the text-to-speech conversion model for correcting the distorted content of the audio data.
10. The microphone optimization control system based on scene adaptation according to claim 7, characterized in that: It also includes scene twin optimization modules: Twin synchronization unit, used to record the current position information in real time and perform twin synchronization within the digital scene model; An object marking unit is used to obtain the motion characteristics of the feature object through big data when fitting to the feature object corresponding to the audio feature, and mark the twin trajectory in the digital scene model based on the relative position information of the current audio feature, wherein the twin trajectory includes the motion direction and the motion time distribution; A relative evaluation unit, configured to calculate, based on the digital scene model, the positional relationship of the feature object relative to the acquisition device at consecutive time nodes when the audio feature is collected and fitted again; The gain refinement unit is used to accurately calculate the gain assignment of the object audio feature based on the position relationship, and obtain the gain assignment of each time node in a continuous time period, wherein the gain assignment includes a compensation coefficient for the volume change caused by the change in relative position relationship.
Citation Information
Patent Citations
Bluetooth earphone audio intelligent regulation and control method and system based on environmental noise
CN118338175A
Audio processing method, processing system, medium and program product
CN118841022A