Audio optimization method, device, system and medium
By receiving and integrating candidate audio from group receiving devices and utilizing neural network model scoring and scene recognition technology, the problem of unclear sound distinction in existing audio optimization methods is solved, thereby improving the auditory effect and user experience of live audio.
Patent Information
- Application Number
- CN202011032128.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-27
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2040-09-27
AI Technical Summary
Existing audio optimization methods cannot effectively distinguish and optimize the sounds that the audience is interested in and those that they are not interested in in live broadcast scenarios, resulting in a poor listening experience.
By receiving candidate audios sent by the receiving group devices, a candidate audio corresponding to each sound type is determined, and audio fusion is performed. The audio quality is optimized using pre-trained neural network model scoring and scene recognition technology.
The audio's auditory effect is enhanced, allowing viewers to clearly hear the sounds they are interested in and improving their listening experience.
Smart Images

Figure CN114360484B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence terminal technology, and in particular to an audio optimization method, device, system and medium. Background Art
[0002] As the number of users watching live broadcasts increases, their expectations for live broadcast quality are also increasing. They demand not only clear visuals but also clear audio. This is because the live broadcast includes not only the host's voice but also sounds from the surrounding scene. For example, let's say the host is broadcasting from the beach. The beach scene typically includes the sounds of waves, wind, seagulls, construction work, and advertisements from beachfront stores. The live broadcast audio heard by viewers includes both the host's voice and the sounds of the beach scene. Some of these sounds are desirable, while others are unwanted noise. For example, construction work or advertisements from beachfront stores can appear in the live broadcast audio. These unwanted noises can make the host's voice, waves, wind, and seagulls unclear, resulting in a poor listening experience.
[0003] Existing audio optimization methods treat the sounds of beach scenes as a whole and reduce them as noise, achieving an audio optimization effect that emphasizes the host's voice. However, this processing method also eliminates the sounds of waves and wind, preventing users from experiencing or perceiving the beach scene through the live broadcast audio, resulting in a poor listening experience. Summary of the Invention
[0004] The purpose of this application is to provide an audio optimization method, device, system and medium to improve the audio auditory effect and user experience.
[0005] In a first aspect, the present application provides an audio optimization method, which includes: receiving candidate audio sent by devices in a sound receiving group, the sound receiving group including at least one device, the candidate audio being audio of at least one sound type in a target scene, the target scene being a scene corresponding to the environment in which the devices in the sound receiving group are located, and each of the sound types corresponds to at least one candidate audio; determining a candidate audio corresponding to each of the sound types from a candidate audio set, and performing audio fusion on the determined candidate audio to obtain audio after audio fusion, the candidate audio set including the candidate audio sent by the devices in the sound receiving group.
[0006] In the above technical solution, the candidate audio sent by devices in the audio receiving group is audio of at least one sound type in the target scene. From the candidate audio set, one candidate audio corresponding to each sound type is determined and audio fusion is performed. The fused audio is the audio of the current environment, including at least one sound type in the corresponding scene. This fused audio allows the user to experience or perceive the environment, improving the audio listening effect and user experience.
[0007] In one possible design, at least one of the sound types corresponding to the target scene may include a human voice type and / or an environmental sound type. In the above technical solution, the sound type can be a human voice type or an environmental sound type, which can optimize the human voice and environmental sound in the environment, so that the human voice and environmental sound in the audio after audio fusion can both have a good auditory effect.
[0008] In one possible design, the target scene can also be determined before the candidate audio is received from the devices in the receiving group. In the above technical solution, the audio of the environment can be flexibly optimized by the scene corresponding to the environment where the devices in the receiving group are located, thereby improving the optimization effect.
[0009] In one possible design, when determining the target scene, keywords for the environment can be first obtained; then, based on the correspondence between the scene and the keyword, the scene corresponding to the keyword for the environment is determined. In the above technical solution, based on the correspondence between the scene and the keyword, the scene corresponding to the keyword for the environment is determined, thereby automatically determining the scene of the environment in which the sound receiving group is located, and improving the efficiency of determining the scene corresponding to the environment.
[0010] In one possible design, when obtaining keywords for the environment, at least one of the following methods may be used: determining keywords for the environment; or receiving keywords sent by at least one device in the sound receiving group. In the above technical solution, the keywords for the environment may be obtained by determining or receiving. If the keywords for the environment are obtained by determining keywords, the time required to determine the scene corresponding to the environment can be shortened. If the keywords for the environment are obtained by receiving keywords, the computing resources used can be reduced.
[0011] In one possible design, when determining the keywords of the environment, at least one of the following methods can be used to determine the keywords of the environment: performing semantic recognition on the collected audio to be processed to determine the keywords of the audio to be processed; performing image recognition on the collected image of the environment to determine the keywords corresponding to the image; and determining the keywords corresponding to the geographic location information based on the geographic location information of the environment.
[0012] In the above technical solution, the accuracy of determining the scene corresponding to the environment can be improved by performing semantic recognition on the processed audio, image recognition on the image of the environment, and determining the keywords of the environment based on the geographical location information of the environment.
[0013] In one possible design, after determining the target scene and before receiving the candidate audio sent by the devices in the receiving group, the target scene can also be notified to the devices in the receiving group. In the above technical solution, the target scene is notified to the devices in the receiving group, and the devices in the receiving group can process the collected audio to be processed according to the notified scene to obtain candidate audio, thereby reducing the computing resources consumed by the devices in the receiving group to determine the target scene, or avoiding the waste of resources caused by the devices in the receiving group repeatedly determining the target scene.
[0014] In one possible design, when determining a candidate audio corresponding to each of the sound types from the candidate audio set, the quality scores of the candidate audios in the candidate audio set are first obtained; then, the candidate audio corresponding to each of the sound types whose quality scores meet the preset conditions is selected as the candidate audio corresponding to each of the sound types. In the above technical solution, the quality scores of the candidate audios are used to select the candidate audios that meet the preset conditions from the candidate audios corresponding to each sound type for audio fusion, so that the audio after audio fusion includes the audio with better audio quality in each sound type, thereby improving the auditory effect of the audio.
[0015] In one possible design, the quality scores of the candidate audios in the candidate audio set are obtained by at least one of the following methods: determining the quality scores of the candidate audios in the candidate audio set; receiving the quality scores of the candidate audios sent by the devices in the sound receiving group. In the above technical solution, the quality scores of the candidate audios can be obtained by determining or receiving. If the quality scores of the candidate audios in the candidate audio set are obtained by determining the quality scores, the audio optimization time can be shortened. If the quality scores of the candidate audios in the candidate audio set are obtained by receiving the quality scores, the computing resources can be reduced.
[0016] In one possible design, a pre-trained neural network model is used to determine the quality scores of candidate audios in the candidate audio set. In the above technical solution, the quality scores determined by the pre-trained neural network model can reflect the auditory effects of the candidate audios and help improve the auditory effects of the fused audio.
[0017] In one possible design, a preset condition can be set based on the actual application scenario, for example, the preset condition is the highest quality score, or the highest quality score and the quality score is greater than a scoring threshold. In the above technical solution, the preset condition is the highest quality score, so that the sound of each sound type in the audio after the audio fusion is the sound of the highest audio quality. The preset condition is the highest quality score and the quality score is greater than the scoring threshold, so that the sound of each sound type in the audio after the audio fusion is the sound of the highest audio quality and provides a good listening experience.
[0018] In one possible design, candidate audio obtained by segmenting the collected audio to be processed can also be added to the candidate audio set. In the above technical solution, the audio in the candidate audio set can also be candidate audio obtained by segmenting the collected audio to be processed, which helps improve the efficiency of determining a candidate audio corresponding to each audio and video type.
[0019] In one possible design, after determining a candidate audio corresponding to each of the sound types, indication information can be sent to at least one device in the sound receiving group, and the indication information is used to indicate the sound type corresponding to the candidate audio that the device needs to send subsequently. In the above technical solution, by sending indication information to indicate the sound type corresponding to the candidate audio that the devices in the sound receiving group need to send subsequently, while ensuring the auditory effect of the audio after audio fusion, it can reduce the resources occupied by the candidate audio sent by the devices in the sound receiving group, shorten the processing time for determining a candidate audio corresponding to each of the sound types, and shorten the audio optimization processing time.
[0020] In a second aspect, the present application provides an audio optimization method, which includes: collecting audio to be processed; segmenting the audio to be processed to obtain at least one candidate audio; the target scene corresponding to the current environment includes at least one sound type, and the at least one candidate audio corresponds to one of the at least one sound type, and each obtained candidate audio corresponds to a different sound type; and then sending the at least one candidate audio to the first device.
[0021] In the above technical solution, the audio to be processed is segmented to obtain at least one candidate audio of the target scene corresponding to the current environment, and the candidate audio corresponds to one of the at least one sound type included in the target scene, so that the first device can receive the candidate audio corresponding to the at least one sound type.
[0022] In one possible design, the sound type included in the target scene may include a human voice type and / or an environmental sound type. In the above technical solution, the sound type may be a human voice type or an environmental sound type, so that the human voice and environmental sound in the current environment can be obtained from the audio to be processed.
[0023] In one possible design, before segmenting the audio to be processed, the target scene notified by the first device may also be received. In the above technical solution, the collected audio to be processed can be processed according to the notified scene to obtain candidate audio, thereby reducing the computing resources consumed in determining the target scene or avoiding the waste of resources caused by repeated determination of the target scene.
[0024] In one possible design, the target scene is determined before the audio to be processed is segmented. In the above technical solution, the scene corresponding to the current environment is determined, and candidate audio of the sound type contained in the target scene corresponding to the current environment can be obtained flexibly by segmenting the audio to be processed.
[0025] In one possible design, in the process of determining the target scene, the keyword of the current environment is determined; based on the correspondence between the scene and the keyword, the scene corresponding to the keyword of the current environment is determined. In the above technical solution, based on the correspondence between the scene and the keyword, the scene corresponding to the keyword of the current environment is determined, thereby automatically determining the scene of the environment where the sound receiving group is located, and improving the efficiency of determining the scene corresponding to the environment.
[0026] In one possible design, keywords for the environment can be determined using at least one of the following methods: semantic recognition of the collected audio to be processed to determine keywords for the audio to be processed; image recognition of the collected images of the environment to determine keywords corresponding to the images based on the names of the identified objects; determining keywords corresponding to the geographic location information based on the geographic location information of the environment; and using the determined keywords as keywords for the current environment. In the above technical solution, keywords for the environment can be determined by performing semantic recognition on the audio to be processed, performing image recognition on images of the environment, and determining keywords based on the geographic location information of the environment, thereby improving the accuracy of determining the scene corresponding to the environment.
[0027] In one possible design, when the audio to be processed is segmented to obtain candidate audio, if the at least one sound type includes a first environmental sound type, the audio to be processed is segmented to obtain audio without human voice, and the audio without human voice is segmented to obtain audio of the first environmental sound type; if the at least one sound type includes a human voice type, the audio to be processed is segmented to obtain audio of the human voice type.
[0028] In the above technical solution, if the sound types included in the target scene include the first ambient sound type, the audio to be processed is segmented to obtain non-voice audio, and the non-voice audio is segmented to obtain audio of the first ambient sound type, thereby improving the audio quality of the first ambient sound type. If the sound types included in the target scene include human voices, the audio to be processed is segmented to obtain audio of the human voice type, allowing the first device to obtain candidate audio of the human voice type.
[0029] In one possible design, before the audio to be processed is segmented, indication information sent by the first device can be received, and the indication information is used to indicate the sound type corresponding to the sent candidate audio. The audio to be processed is then segmented to obtain candidate audio corresponding to the sound type indicated in the indication information. In the above technical solution, the audio to be processed is segmented according to the sound type in the received indication information, thereby reducing the number of candidate audios sent and the resources occupied. There is no need to segment candidate audios of the sound type not included in the indication information from the audio to be processed, thereby reducing resource consumption.
[0030] In one possible design, after segmenting the audio to be processed to obtain at least one candidate audio, a quality score for each of the at least one candidate audio can be determined; and the quality score for each candidate audio can then be sent to the first device. In the above technical solution, the quality score for each determined candidate audio is sent to the first device, reducing the consumption of computing resources on the first device.
[0031] In one possible design, a pre-trained neural network model is used to determine the quality scores of candidate audios in the candidate audio set. In the above technical solution, the quality scores determined by the pre-trained neural network model can reflect the auditory effects of the candidate audios and help improve the auditory effects of the fused audio.
[0032] In a third aspect, the present application provides a method for establishing a sound receiving group, the method comprising: sending a first ultrasonic signal for establishing a sound receiving group; receiving a second ultrasonic signal, the second ultrasonic signal carrying information confirming joining the sound receiving group; and determining, based on the second ultrasonic signal, the device sending the second ultrasonic signal as a device in the sound receiving group, wherein the sound receiving group is a group composed of devices that can collect sounds in the current environment.
[0033] In the above technical solution, the poor penetrability of ultrasonic signals is utilized. By sending a first ultrasonic signal for establishing a receiving group, devices within a certain range can receive the first ultrasonic signal, and the second ultrasonic signal feedback is used to confirm the information of joining the receiving group. The device sending the second ultrasonic signal is determined as a device in the receiving group, and there is no need to screen the device sending the second ultrasonic signal or determine the device location.
[0034] In one possible design, the first ultrasonic signal is a signal of a first preset frequency.
[0035] In one possible design, the first ultrasonic signal carries a request to establish a receiving group.
[0036] In one possible design, the request to establish a reception group includes an identifier and / or an Internet Protocol address of the device that sends the first ultrasonic signal.
[0037] In one possible design, the second ultrasonic signal is a signal of a second preset frequency.
[0038] In one possible design, the information confirming joining the receiving group includes the identifier and / or Internet Protocol address of the device sending the second ultrasonic information signal.
[0039] In one possible design, after determining the device sending the second ultrasonic signal as a device in the sound receiving group, the method further includes: communicating with the devices in the sound receiving group based on a transmission control protocol.
[0040] In a fourth aspect, the present application provides a method for establishing a sound receiving group, the method comprising: receiving a third ultrasonic signal sent by the first device for establishing a sound receiving group; and sending a fourth ultrasonic signal, the fourth ultrasonic signal carrying information confirming joining the sound receiving group.
[0041] In this technical solution, the poor penetrability of ultrasonic signals is exploited. By receiving a third ultrasonic signal used to establish a sound reception group, it is possible to determine which devices within a predetermined range or space are being grouped. A fourth ultrasonic signal, carrying confirmation of joining the sound reception group, can be used to notify the device establishing the sound reception group. Furthermore, there is no need to determine the location of the device establishing the sound reception group, nor to determine whether it can join the sound reception group.
[0042] In one possible design, the third ultrasonic signal is a signal of a third preset frequency.
[0043] In one possible design, the third ultrasonic signal carries a request to establish a receiving group.
[0044] In one possible design, the request to establish a sound receiving group includes an identifier and / or an Internet Protocol address of the device that sends the third ultrasonic signal.
[0045] In one possible design, the fourth ultrasonic signal is a signal of a fourth preset frequency.
[0046] In one possible design, the information confirming joining the receiving group includes the identifier and / or Internet Protocol address of the device that sends the fourth ultrasonic information signal.
[0047] In one possible design, after sending the fourth ultrasonic signal, the method further includes:
[0048] Communicate with the first device based on the Transmission Control Protocol.
[0049] In the fifth aspect, the present application also provides an electronic device, which includes: a receiving module for receiving candidate audio sent by devices in a sound receiving group, the sound receiving group includes at least one device, the candidate audio is audio of at least one sound type in a target scene, the target scene is a scene corresponding to the environment where the devices in the sound receiving group are located, and each of the sound types corresponds to at least one candidate audio; a processing module for determining a candidate audio corresponding to each of the sound types from a candidate audio set, the candidate audio set includes candidate audio sent by devices in the sound receiving group; an audio fusion module for performing audio fusion on the determined candidate audio to obtain the audio after audio fusion.
[0050] In the sixth aspect, the present application also provides an electronic device, which includes: an acquisition module for acquiring audio to be processed; a processing module for segmenting the audio to be processed to obtain at least one candidate audio, each of the at least one candidate audio corresponding to one of the at least one sound types included in the target scene, and the target scene is the scene corresponding to the current environment; a sending module for sending the at least one candidate audio to the first device.
[0051] In the seventh aspect, the present application also provides an electronic device, which includes: one or more processors coupled to a memory, wherein the memory stores one or more computer programs; the one or more computer programs include instructions, and when the instructions are executed by the device, the device executes any possible design technical solution in the first aspect above.
[0052] In the eighth aspect, the present application also provides an electronic device, which includes: one or more processors coupled to a memory, wherein the memory stores one or more computer programs; the one or more computer programs include instructions, and when the instructions are executed by the device, the device executes any possible design technical solution in the above-mentioned second aspect.
[0053] In the ninth aspect, the present application also provides an electronic device, which includes: one or more processors coupled to a memory, wherein the memory stores one or more computer programs; the one or more computer programs include instructions, and when the instructions are executed by the device, the device executes any possible design technical solution in the third aspect above.
[0054] In the tenth aspect, the present application also provides an electronic device, which includes: one or more processors coupled to a memory, wherein the memory stores one or more computer programs; the one or more computer programs include instructions, and when the instructions are executed by the device, the device executes any possible design technical solution in the above-mentioned fourth aspect.
[0055] In the eleventh aspect, the present application also provides an electronic device (also referred to as a terminal, such as a mobile phone), which includes modules / units for executing the first aspect or any possible design method of the first aspect; these modules / units can be implemented through hardware, or corresponding software can be implemented through hardware.
[0056] In the twelfth aspect, the present application also provides an electronic device (also referred to as a terminal, such as a mobile phone), which includes modules / units for executing the second aspect or any possible design method of the second aspect; these modules / units can be implemented through hardware, or corresponding software implementations can be executed through hardware.
[0057] In a thirteenth aspect, the present application further provides an audio optimization system, which includes the apparatus in the seventh aspect and the apparatus in the eighth aspect. Alternatively, the system includes the electronic device in the eleventh aspect and the electronic device in the twelfth aspect.
[0058] In the fourteenth aspect, an embodiment of the present application further provides a chip, which is coupled with a memory in an electronic device to execute the technical solution of the first aspect of the embodiment of the present application and any possible design of the first aspect, or the technical solution of the second aspect and any possible design of the second aspect; in the embodiment of the present application, "coupling" refers to the direct or indirect combination of two components with each other.
[0059] In the fifteenth aspect, a computer-readable storage medium of an embodiment of the present application includes a computer program. When the computer program runs on an electronic device, the electronic device executes the technical solution of the first aspect of the embodiment of the present application and any possible design of the first aspect, or the technical solution of the second aspect and any possible design of the second aspect.
[0060] In the sixteenth aspect, a computer program product in an embodiment of the present application, when the computer program product is run on an electronic device, enables the electronic device to execute the technical solution of the first aspect of the embodiment of the present application and any possible design of the first aspect, or the technical solution of the second aspect and any possible design of the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application;
[0062] Figure 2 A schematic diagram of the software structure of the electronic device provided in an embodiment of the present application;
[0063] Figure 3 A schematic diagram of an AI engine functional module provided in an embodiment of the present application;
[0064] Figure 4 A schematic diagram of calling an AI engine provided in an embodiment of the present application;
[0065] Figure 5 A schematic diagram of an outdoor live broadcast scene provided in an embodiment of the present application;
[0066] Figure 6 A schematic diagram of establishing a radio group according to an embodiment of the present application;
[0067] Figure 7 A schematic diagram of an outdoor live broadcast scene provided in an embodiment of the present application;
[0068] Figure 8 A schematic diagram of establishing a radio group according to an embodiment of the present application;
[0069] Figure 9 A schematic diagram of establishing a radio group according to an embodiment of the present application;
[0070] Figure 10 A schematic diagram of a scene image provided in an embodiment of the present application;
[0071] Figure 11 A schematic diagram of an interface provided in an embodiment of the present application;
[0072] Figure 12A flowchart of an audio segmentation method provided in an embodiment of the present application;
[0073] Figure 13 A schematic diagram of a convolutional neural network training process provided in an embodiment of the present application;
[0074] Figure 14 A schematic diagram of an audio fusion method provided in an embodiment of the present application;
[0075] Figure 15 A schematic diagram of an audio fusion method provided in an embodiment of the present application;
[0076] Figure 16 A flowchart of an audio optimization method provided in an embodiment of the present application;
[0077] Figure 17 A flowchart of an audio optimization method provided in an embodiment of the present application;
[0078] Figure 18 A schematic structural diagram of an electronic device provided in an embodiment of the present application;
[0079] Figure 19 A schematic structural diagram of another electronic device provided in an embodiment of the present application;
[0080] Figure 20 A schematic structural diagram of another electronic device provided in an embodiment of the present application;
[0081] Figure 21 This is a structural diagram of another electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0082] The technical solutions in the embodiments of the present application will be described in detail below in conjunction with the drawings in the following embodiments of the present application.
[0083] Currently, in live webcasts or television broadcasts, the audio heard by viewers includes the voice of the host or presenter, as well as sounds from the live scene. To improve audio quality and provide a better listening experience, noise reduction is typically performed on the audio collected during the live broadcast. Existing noise reduction methods typically emphasize the host or presenter's voice, specifically the human voice, while treating all sounds from the live scene as noise for noise reduction or denoising.
[0084] However, especially in live webcast scenarios, in addition to being interested in the host's voice, the audience is also interested in some sounds in the scene. For example, in a live webcast scenario at the seaside, the audience is also interested in the sound of waves, wind, and seagulls in the seaside scene, but is not interested in the construction sounds and the promotional sounds of seaside shops in the seaside scene. If the sounds that the audience wants to hear, the host's voice, the sound of waves, the sound of wind, and the sound of seagulls can be clearly extracted, the audience can hear the sounds they want to hear. The audience will feel or perceive the live broadcast scene through audio, which will greatly enhance the auditory experience. To this end, the embodiment of the present application provides an audio optimization method that can be applied to live webcast scenarios, TV live broadcast scenarios, film and television work recording scenarios, music work recording scenarios, and the like.
[0085] When applied to film and television productions, audio recording is generally done through two methods: live recording and post-dubbing. During live recording, the recording environment must be as quiet as possible to ensure the clarity of the actors' voices. Especially for non-modern dramas, the recorded audio must avoid sounds from modern life. For example, when recording a Tang Dynasty drama, the recorded audio must avoid ringtones, vibrations, and car horns. Therefore, to ensure that the audio of a film or television production is consistent with the recorded content and to avoid any garbled audio, post-dubbing is often used. However, post-dubbing not only requires voice actors to dub the lines but also requires extensive effort to recreate the ambient sounds as closely as possible. Therefore, post-dubbing is labor-intensive and challenging.
[0086] Based on this, the embodiments of this application provide an audio optimization method and electronic device that can be applied during the recording of a film or television work to clearly extract the voices of actors and the sounds in scenes corresponding to the content of the film or television work, ensuring that the sounds heard by the audience are consistent with the scenes reflected in the content of the film or television work, thereby avoiding "bloopers". The method and electronic device are based on the same inventive concept. Since the principles of the method and device are similar, the implementation of the device and method can refer to each other, and the repeated parts will not be repeated.
[0087] If applied to a music recording scenario, for example, a concert recording, there are usually multiple sound recording devices at the recording site. The singer's voice is mainly picked up by the sound recording device carried by the singer, such as a microphone or a handset, as the main audio. There are other sound recording devices near the stage to collect the sound of the concert as auxiliary audio. The director or a dedicated post-production staff member decides at which moment to select which auxiliary audio to merge with the main audio in order to restore the atmosphere of the concert as much as possible. Therefore, an experienced director or staff member is required to make the fusion decision, which not only requires a lot of manpower but also takes a long time to process. The audio optimization method provided by this application, when applied to the music recording site, can determine the best singer's voice and the live sound from the live sounds collected by multiple sound recording devices, so that the audience can feel or perceive the atmosphere of the music recording scene through the sound.
[0088] The terms used in the following embodiments are only for the purpose of describing specific embodiments and are not intended to be limiting of the present application. As used in the specification of the present application and the appended claims, the singular expressions "a", "an", "said", "above", "the" and "this" are intended to also include expressions such as "one or more", unless there is a clear contrary indication in the context. It should also be understood that in the embodiments of the present application, "one or more" refers to one, two or more; "and / or" describes the association relationship of associated objects, indicating that three relationships may exist; for example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the related objects before and after are in an "or" relationship.
[0089] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0090] The at least one involved in the embodiments of the present application includes one or more; wherein, more means greater than or equal to two. In addition, it should be understood that in the description of this application, the words "first" and "second" are only used for the purpose of distinguishing descriptions and cannot be understood as indicating or implying relative importance or order.
[0091] Below, some of the terms involved in the embodiments of the present application are explained to facilitate understanding by those skilled in the art.
[0092] The embodiments of the present application can be implemented by any intelligent electronic device, and various applications (applications, apps) can be installed in the electronic device. The application involved here can be referred to as application, which is a software program that can realize one or more specific functions. Generally, multiple applications can be installed in an electronic device, such as live broadcast applications, video applications, audio applications, etc. Among them, live broadcast applications, for example, can be installed. Etc. Video applications, for example, can include Etc. Audio applications, for example, may include The applications mentioned in the following embodiments may be applications installed on the electronic device when it leaves the factory, or may be applications downloaded from the Internet or obtained from other electronic devices by the user during the use of the electronic device.
[0093] The audio optimization method provided in the embodiment of the present application can be implemented collaboratively by a first device and at least one second device. The first device can be any electronic device, for example, a terminal device with a display screen, such as a mobile phone, a tablet computer, a laptop computer, etc. For another example, a terminal device without a display screen, such as a smart box with a processor. The first device is used to control the formation of a device group, control the second device to process the locally collected audio, and control the second device to send the processed audio to the first device. Exemplarily, if the first device is implemented as a smart box, the first device can also be connected to a display device to provide an interactive interface to the user through the connected external display device.
[0094] The second device can be a terminal device with a microphone, such as a smart speaker, a smart microphone, a Bluetooth headset, a mobile phone, a tablet computer, or a laptop computer. The second device can also be a smart home device that can receive voice control, such as a smart refrigerator, a smart sweeping robot, etc. The second device can convert the sound signal collected by the microphone into an audio signal. The first device and the second device can be equipped with Or other operating systems.
[0095] For example, Figure 1 1 shows a schematic structural diagram of the electronic device 100. Figure 1As shown, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0096] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors. The controller may serve as the nerve center and command center of the electronic device 100. The controller may generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution. The processor 110 may also include memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a high-speed cache memory. This memory may store instructions or data that have just been used or are being recycled by the processor 110. If the processor 110 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids duplicate accesses, reduces the processor 110's waiting time, and thus improves system efficiency.
[0097] The USB interface 130 is an interface that complies with USB standards and specifications, and may be a Mini USB interface, a Micro USB interface, a USB Type-C interface, or the like. The USB interface 130 can be used to connect a charger to charge the electronic device 100, and can also be used to transfer data between the electronic device 100 and peripheral devices. The charging management module 140 is used to receive charging input from the charger. The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, and provides power to the processor 110, the internal memory 121, the external memory, the display 194, the camera 193, and the wireless communication module 160.
[0098] The wireless communication functionality of electronic device 100 can be implemented using antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, a modem processor, and a baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.
[0099] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the electronic device 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.
[0100] The wireless communication module 160 can provide wireless communication solutions applied to the electronic device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be transmitted from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2. It should be noted that the wireless communication module 160 can transmit information such as scene information determined by the processor 110, candidate audio corresponding to each sound type, and quality scores of the candidate audio corresponding to each sound type.
[0101] In some embodiments, the antenna 1 of the electronic device 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the electronic device 100 can communicate with the network and other devices through wireless communication technology. The wireless communication technology may include global system for mobile communications (GSM), general packet radio service (GPRS), long term evolution (LTE), new radio (NR) wireless communication technology, wireless communication technology emerging in future communication development, BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS may include global positioning system (GPS), global navigation satellite system (GLONASS), Beidou navigation satellite system (BDS), quasi-zenith satellite system (QZSS) and / or satellite-based augmentation system (SBAS). In some embodiments, the electronic device 100 can receive satellite signals from various satellite navigation systems using GNSS wireless communication technology, and use satellite positioning technology to determine geographic location information such as the location or coordinates of the electronic device 100 based on the received satellite signals. Keywords corresponding to the geographic location information of the electronic device 100 can be used to determine the scene of the environment in which the electronic device 100 is located.
[0102] The camera 193 is used to capture still images or videos. The camera 193 may include a front camera and a rear camera. Generally, the camera 193 may include a photosensitive element such as a lens group and an image sensor, wherein the lens group includes a plurality of lenses (convex lenses or concave lenses) for collecting light signals reflected by the object to be photographed and transmitting the collected light signals to the image sensor. The image sensor generates an original image of the object to be photographed based on the light signal. After the camera 193 collects the original image, it can send the original image to the processor 110. In addition, Figure 1 The camera 193 shown may include 1-N cameras. The processor 110 may perform image recognition on the original image captured by the camera 193, identify objects contained in the original image, and use the names of the identified objects as keywords corresponding to the image.
[0103] The display screen 194 is used to display the display interface of the application, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active matrix organic light-emitting diode or an active matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), Miniled, MicroLed, Micro-oLed, a quantum dot light emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include 1 or N display screens 194, where N is a positive integer greater than 1. In an embodiment of the present application, the display screen 194 can be used to display multiple application interfaces simultaneously.
[0104] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, and the software code of at least one application, etc. The data storage area can store data generated during the use of the electronic device 100 (such as collected audio data, various types of segmented audio data, scene type data, ambient sound type data, etc.). In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0105] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. For example, files such as pictures and videos can be stored on the external memory card.
[0106] The electronic device 100 can implement audio functions such as music playback and recording through the audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor. The application processor can be used to collect sound signals and convert them into audio signals or audio data, and the processor 110 can also be used to collect sound signals and convert them into audio signals or audio data. The application processor or processor 110 can control the speaker 170A to play a sound signal of a preset frequency, where the preset frequency can be higher than a sound frequency audible to humans, such as an ultrasonic sound signal of a preset frequency.
[0107] Pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A can be provided on display screen 194. Gyroscope sensor 180B can be used to determine the motion posture of electronic device 100. In some embodiments, gyroscope sensor 180B can be used to determine the angular velocity of electronic device 100 around three axes (i.e., x, y, and z axes).
[0108] The gyroscope sensor 180B can be used for anti-shake shooting. The air pressure sensor 180C is used to measure air pressure. In some embodiments, the electronic device 100 calculates the altitude through the air pressure value measured by the air pressure sensor 180C to assist in positioning and navigation. The magnetic sensor 180D includes a Hall sensor. The electronic device 100 can use the magnetic sensor 180D to detect the opening and closing of the flip leather case. In some embodiments, when the electronic device 100 is a flip phone, the electronic device 100 can detect the opening and closing of the flip according to the magnetic sensor 180D. Then, according to the detected opening and closing state of the leather case or the opening and closing state of the flip, features such as automatic unlocking of the flip can be set. The acceleration sensor 180E can detect the magnitude of the acceleration of the electronic device 100 in various directions (generally three axes). When the electronic device 100 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the posture of the electronic device, and is applied to applications such as horizontal and vertical screen switching and pedometers. In some embodiments, the electronic device 100 identifies the motion posture (stationary posture or motion posture) of the electronic device 100 through the parameters output by the gyro sensor 180B and the parameters output by the acceleration sensor 180E.
[0109] The distance sensor 180F is used to measure the distance. The electronic device 100 can measure the distance through infrared or laser. In some embodiments, when shooting a scene, the electronic device 100 can use the distance sensor 180F to measure the distance to achieve fast focusing. The proximity light sensor 180G may include, for example, a light emitting diode (LED) and a light detector, such as a photodiode. The light emitting diode can be an infrared light emitting diode. The electronic device 100 emits infrared light outward through the light emitting diode. The electronic device 100 uses the photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the electronic device 100. When insufficient reflected light is detected, the electronic device 100 can determine that there is no object near the electronic device 100. The electronic device 100 can use the proximity light sensor 180G to detect that the user holds the electronic device 100 close to the ear to talk, so as to automatically turn off the screen to save power. The proximity light sensor 180G can also be used in leather case mode and pocket mode to automatically unlock and lock the screen.
[0110] The ambient light sensor 180L is used to sense ambient light brightness. The electronic device 100 can adaptively adjust the brightness of the display screen 194 based on the perceived ambient light brightness. The ambient light sensor 180L can also be used to automatically adjust the white balance when taking photos. The ambient light sensor 180L can also work with the proximity light sensor 180G to detect whether the electronic device 100 is in a pocket to prevent accidental touches. The fingerprint sensor 180H is used to collect fingerprints. The electronic device 100 can use the collected fingerprint characteristics to implement fingerprint unlocking, access application locks, fingerprint photography, fingerprint call answering, etc.
[0111] The temperature sensor 180J is used to detect temperature. In some embodiments, the electronic device 100 uses the temperature detected by the temperature sensor 180J to execute a temperature processing strategy. For example, when the temperature reported by the temperature sensor 180J exceeds a threshold, the electronic device 100 reduces the performance of the processor located near the temperature sensor 180J to reduce power consumption and implement thermal protection. In other embodiments, when the temperature is lower than another threshold, the electronic device 100 heats the battery 142 to prevent the electronic device 100 from shutting down abnormally due to low temperature. In other embodiments, when the temperature is lower than another threshold, the electronic device 100 boosts the output voltage of the battery 142 to prevent abnormal shutdown due to low temperature.
[0112] The touch sensor 180K is also called a "touch panel." The touch sensor 180K can be disposed on the display screen 194. The touch sensor 180K and the display screen 194 form a touch screen, also called a "touch screen." The touch sensor 180K is used to detect touch operations applied thereto or in the vicinity thereof. The touch sensor can transmit the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operations can be provided via the display screen 194. In other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device 100, in a location different from that of the display screen 194.
[0113] Bone conduction sensor 180M can acquire vibration signals. In some embodiments, bone conduction sensor 180M can acquire vibration signals from vibrating bones in the human body. Bone conduction sensor 180M can also contact the human pulse to receive blood pressure signals.
[0114] The ultrasonic sensor 180N can receive ultrasonic signals. The ultrasonic sensor 180N can also transmit ultrasonic signals. In some embodiments, the ultrasonic sensor 180N can receive ultrasonic signals used to establish a sound receiving group. The ultrasonic sensor 180N can transmit an ultrasonic signal that carries a confirmation of joining the sound receiving group. The ultrasonic signal received by the ultrasonic sensor 180N can be decoded by the processor 110. The ultrasonic sensor 180N can also transmit the ultrasonic signal modulated by the processor 110. The processor 110 can control the ultrasonic sensor 180N to receive ultrasonic signals in the environment. The processor 110 can also control the ultrasonic sensor 180N to transmit ultrasonic signals.
[0115] The buttons 190 include a power button, a volume button, etc. The button 190 can be a mechanical button. It can also be a touch button. The electronic device 100 can receive button input and generate key signal input related to the user settings and function control of the electronic device 100. The motor 191 can generate a vibration prompt. The motor 191 can be used for incoming call vibration prompts, and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playback, etc.) can correspond to different vibration feedback effects. The indicator 192 can be an indicator light, which can be used to indicate the charging status, power changes, and can also be used to indicate messages, missed calls, notifications, etc. The SIM card interface 195 is used to connect the SIM card. The SIM card can be inserted into the SIM card interface 195 or pulled out from the SIM card interface 195 to achieve contact and separation with the electronic device 100.
[0116] It is understandable that Figure 1The components shown do not constitute a specific limitation on the mobile phone. The mobile phone may also include more or fewer components than shown in the figure, or combine some components, or split some components, or arrange the components differently. Figure 1 The electronic device shown is used as an example for introduction.
[0117] Figure 2 FIG1 shows a software structure block diagram of an electronic device provided by an embodiment of the present application. Figure 2 As shown, the software structure of the electronic device can be a layered architecture. For example, the software can be divided into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers: from top to bottom, the application layer, the application framework layer (framework, FWK), the Android runtime (Android runtime) and system library, and the kernel layer.
[0118] The application layer can include a series of application packages. Figure 2 As shown, the application layer may include camera, settings, skin module, user interface (UI), third-party applications, etc. Among them, third-party applications may include WeChat, QQ, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc.
[0119] The application framework layer provides application programming interface (API) and programming framework for the application layer. The application framework layer may include some predefined functions. Figure 3 As shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, an audio optimization control function manager, and the like.
[0120] The window manager manages windowed applications. It can determine the display size, determine whether a status bar is present, lock the screen, and take screenshots. Content providers store and retrieve data and make it accessible to applications. This data can include video, images, audio, incoming and outgoing calls, browsing history and bookmarks, and phone books.
[0121] The view system includes visual controls, such as those for displaying text and images. The view system is used to build applications. A display interface can consist of one or more views. For example, a display interface containing a text notification icon might include a view for displaying text and a view for displaying images.
[0122] The phone manager is used to provide communication functions of the electronic device 100, such as management of call status (including answering, hanging up, etc.).
[0123] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.
[0124] The Android runtime includes the core library and the virtual machine. The Android runtime is responsible for scheduling and management of the Android system.
[0125] The core library consists of two parts: one containing the Java language's callable functions and the other the Android core library. The application layer and the application framework layer run in a virtual machine. The virtual machine executes the Java files in the application layer and application framework layer as binary files. The virtual machine manages object lifecycles, stack management, thread management, security and exception management, and garbage collection.
[0126] The system library can include multiple functional modules, such as surface manager, media libraries, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), AI engine, etc.
[0127] The surface manager is used to manage the display subsystem and provide fusion of 2D and 3D layers for multiple applications.
[0128] The media library supports playback and recording of a variety of common audio and video formats, as well as static image files. The media library can support a variety of audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0129] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0130] A 2D graphics engine is a drawing engine for 2D drawings.
[0131] The AI engine can be used to implement the audio optimization method provided in the embodiment of this application. Figure 3As shown, the AI engine may include a fusion sound receiving module, an audio segmentation module, and an audio scoring module. The AI engine may also include an audio fusion module, and the AI engine may also include a control display module. Among them, the fusion sound receiving module is used to realize the establishment of a sound receiving group (sound receiving device group) and the establishment of a communication connection and data transmission function with the devices in the sound receiving group; the audio segmentation module is used to realize scene recognition and the function of segmenting the collected audio to be processed according to the ambient sound type corresponding to the scene, and obtaining at least one candidate audio. The candidate audio can be audio corresponding to the human voice type or audio corresponding to the ambient sound type; the audio scoring module is used to realize the function of determining the quality score of the candidate audio; the audio fusion module is used to realize the function of audio fusion of the candidate audio determined from the candidate audio set. It should be noted that the module can also refer to capabilities or services. The AI engine can be understood as an encapsulated code or function library, and specific functions can be realized by calling the interface of the code or function library. For example, audio segmentation function, etc.
[0132] In an embodiment of the present application, the program code of the audio optimization method can be encapsulated in the form of a software development kit (SDK). The SDK may include but is not limited to header files, library files, and dynamic link libraries. When developing or updating an APP, the SDK can be embedded, and the audio optimization method can be implemented when the APP is running. It should also be noted that when the SDK provided by this application is embedded during the development or update of the APP, it may also be necessary to obtain permission to use the SDK. In the case where the SDK provided by this application is embedded in a third-party application, the set API is used during the SDK instantiation process to call the AI engine in the system library (it can also be a software module with similar functions, which can be given other names) to implement the audio optimization method provided by this application. Such as Figure 4 The program code of the audio optimization method shown can be encapsulated in a toolkit in the AI engine. In the SDK instantiation for audio optimization in third-party applications, the toolkit in the AI engine can be called to implement the audio optimization method provided in this application.
[0133] In some embodiments, a user triggers the launch of an app through an electronic device, which calls the AI engine through the SDK embedded in the app. Methods for launching an app include, but are not limited to, button launch, icon launch, or voice launch, and this application does not limit these methods.
[0134] In some embodiments, the user can also modify the system settings to set the audio collection mode in the electronic device system settings to multi-device audio collection mode. When it is recognized that the user is using the electronic device to collect audio, the AI engine is triggered to call.
[0135] Exemplarily, the AI engine provided in the embodiment of the present application can be stored in the memory of the processor 110, or integrated into the processor 110. The processor 110 can run the audio optimization method provided in the present application to achieve the audio fusion effect of eliminating noise and scene-independent sounds.
[0136] The kernel layer is the layer between hardware and software. The kernel layer includes at least display driver, camera driver, audio driver, and sensor driver.
[0137] The hardware layer can include various sensors, such as accelerometers, gyroscopes, touch sensors, etc.
[0138] The following embodiments of the present application will be combined with the accompanying drawings and application scenarios to introduce the audio optimization method provided by the embodiments of the present application in detail. The following embodiments can be used in an electronic device 100 having the above hardware structure, or in an electronic device having Figure 2 The software structure shown is implemented in an electronic device.
[0139] The following uses a live broadcast scenario as an example to illustrate the audio optimization method provided in the embodiment of the present application.
[0140] Typically, a live streaming scenario may include multiple cameras and receivers in different locations. Live streaming scenarios can also be categorized into various types. Based on the location, these scenarios can be divided into indoor, outdoor, or outdoor scenarios. Based on the content, these scenarios can be divided into shopping, gaming, and more.
[0141] Taking the beach outdoor live broadcast as an example, the beach outdoor live broadcast scene can be seen in Figure 5 The audio optimization method provided by the present application is implemented collaboratively by the first device, the second device a, and the second device b. The first device is a smart phone, and the second device can be a smart phone (e.g. Figure 5 The second device a) shown in FIG. 1 may also be a smart speaker (eg Figure 5 The second device b) is shown in FIG.
[0142] like Figure 5As shown, the first device can be held by the host or placed elsewhere in the scene. The first device can also collect sound signals in the scene. Second device a is placed near the shore, and second device b is placed on a table displaying items. Within the same scene, the different positions of the sound receiving devices can result in different sound signal quality or listening experience. Generally, the closer the listening position is to the sound source, the clearer the sound will be. Therefore, it is expected that second device a will capture the sound of waves best, while second device b will capture the sound of displayed items best. The sounds collected by each device also contain a significant amount of noise. For users watching the live broadcast, sounds other than the host's voice, the sound of displayed items, and the sound of waves are considered noise, such as music playing in a shop near the beach. The host's voice, the sound of displayed items, and the sound of waves can be considered sounds of interest to the user. Similarly, in an outdoor grassland live broadcast, the host's voice, the sound of displayed items, the neighing of horses, and the sound of horses galloping are sounds of interest to the user. In a live game broadcast scenario, the host's voice and game sound effects are sounds that users are interested in. The audio optimization method provided in the embodiments of this application can extract the host's voice and sounds that users are interested in in the scene, and perform audio fusion to eliminate sounds that users are not interested in in the scene, highlight the host's voice and sounds that users are interested in in the scene, and improve the user's auditory experience.
[0143] The AI engine in the first device (which may also be a software module with similar functions, and may be given other names) may include a fusion sound receiving module, an audio segmentation module, an audio scoring module, an audio fusion module, and may also include a control display module. The AI engine in the second device may be the same as the AI engine in the first device. The AI engine in the second device may also be different from the AI engine in the first device, and the AI engine in the second device may include a fusion sound receiving module, an audio segmentation module, and an audio scoring module. It should be noted that the fusion sound receiving module in the AI engine has the ability to form or establish a sound receiving group. The audio segmentation module has the ability to recognize or determine scenes, segment the audio to be processed, and obtain audio corresponding to each sound type. The audio scoring module has the ability to determine the quality score of the audio. The audio fusion module has the ability to fuse audio. The audio optimization method provided in this application can be implemented by realizing these capabilities.
[0144] The following describes the audio optimization method involved in this application, using mobile phones as an example. The audio optimization method involved in this application may include establishing a sound collection group, determining a scene, and processing an audio signal. The following describes the process of establishing a sound collection group, determining a scene, and processing an audio signal separately.
[0145] A1. The process of establishing a radio group includes:
[0146] In an exemplary embodiment, the process of establishing a sound collection group can be as follows: Figure 6 shown.
[0147] S601: A first device sends a first ultrasonic signal for establishing a sound receiving group.
[0148] As an example, after the AI engine in the first device system library is called, it can send an ultrasonic signal transmission control signal to the audio driver in the kernel layer, so that the audio driver drives the speaker to transmit ultrasonic signals. It can also send an ultrasonic signal transmission control signal to the sensor driver in the kernel layer, so that the sensor driver drives the ultrasonic sensor to transmit ultrasonic signals.
[0149] For example, Figure 7 As shown, in actual application scenarios, the frequency of the ultrasonic wave sent by the first device can be any frequency between 20,000 Hz and 21,739 Hz. Since ultrasonic waves are mechanical vibration waves, compared with sound waves that can be heard by humans, they have a strong propagation direction and can be gathered into a directional narrow beam. And when ultrasonic waves propagate in a medium, the ultrasonic intensity will gradually weaken as the propagation distance increases. Therefore, the first device adjusts the frequency of the first ultrasonic signal so that electronic devices within a preset range can receive the first ultrasonic signal sent by the first device.
[0150] For example, in an outdoor scenario, a first ultrasonic signal emitted by a first device can be received by a second device within a preset range, where the preset range can be a circle or a sector with the first device as the center and the preset distance as the radius.
[0151] For another example, in an indoor scenario, the first ultrasonic signal emitted by the first device can be received by the second device within a preset range, where the preset range can refer to the second device in the same space as the first device, and the same space can refer to a room. By taking advantage of the fact that ultrasonic signals cannot penetrate the walls between rooms, the second device in the same room as the first device can receive the first ultrasonic signal, while the second device in a different room from the first device cannot receive the first ultrasonic signal. Figure 8 As shown, a first device, a second device a, and a second device b are in a first room, and a second device c and a second device d are in a second room. The first ultrasonic signal emitted by the first device cannot penetrate the wall between the first and second rooms, allowing second devices a and b in the first room to receive the first ultrasonic signal, while second devices c and d in the second room cannot receive the first ultrasonic signal emitted by the first device. The first device forms or establishes a sound reception group with the devices in the first room by transmitting the first ultrasonic signal.
[0152] In one implementation, the first device may send a first ultrasonic signal through an ultrasonic transmitting device, and the first ultrasonic signal may carry a request for establishing a sound receiving group.
[0153] Exemplarily, the first ultrasonic signal sent by the first device may be an ultrasonic signal of a first frequency, and the first frequency may be used to indicate that the first ultrasonic signal carries a request to establish a sound reception group.
[0154] In one example, the ultrasonic transmitting device may be implemented as a speaker, and the speaker transmits or plays the first ultrasonic signal.
[0155] In one implementation, a first device may generate a first ultrasonic signal using an ultrasonic modulation circuit. The first device may use encoding technology to cause the generated first ultrasonic signal to carry a request to establish a reception group. The request to establish a reception group may be first preset information, and decoding the first ultrasonic signal may reveal that the information carried in the first ultrasonic signal is the request to establish a reception group.
[0156] In addition, the first ultrasonic signal may also carry the identifier of the first device and / or the Internet Protocol (IP) address of the first device, so as to facilitate the second device to identify the first device as a trusted device and prevent resources from being maliciously occupied by other devices.
[0157] In a specific embodiment, the first device may continuously transmit the first ultrasonic signal or may transmit the first ultrasonic signal at a preset period. If the first device does not receive ultrasonic signals from other devices within a period of time after transmitting the first ultrasonic signal, it may be determined that there are no devices within the preset range that can be used to establish a reception group. The first device may display a prompt message to inform the user that there are no devices that can be used to establish a reception group, or prompt the user to switch to a non-multi-device reception mode.
[0158] The second device receives the first ultrasonic signal for establishing a sound reception group. The second device may be an electronic device within a preset range of the first device. In one example, the preset range may be a circle or sector with the first device as the center and a preset distance as the radius. In another example, the preset range may refer to a second device in the same space as the first device, and the same space may refer to a room. It should be noted that if there are multiple second devices within the preset range, each second device may receive the first ultrasonic signal for establishing a sound reception group.
[0159] As an example, after the AI engine in the second device's system library is called, it can send an ultrasonic signal reception control signal to the kernel layer's audio driver, causing the audio driver to drive the microphone to receive ultrasonic signals and send the received ultrasonic signals to the second device's processor. The processor's signal codec decodes the received ultrasonic signals and determines whether the received ultrasonic signals carry a group reception request. Alternatively, the ultrasonic reception control signal can be sent to the kernel layer's sensor driver, causing the sensor driver to drive the ultrasonic sensor to receive, collect, or detect ultrasonic signals.
[0160] In one implementation, an ultrasonic receiving device in a second device can receive a first ultrasonic signal emitted by a first device. In one example, the ultrasonic receiving device can be implemented as a microphone, and the second device can receive or collect the first ultrasonic signal via the microphone. In one example, the ultrasonic receiving device can be implemented as an ultrasonic sensor, and the second device can receive, collect, or detect the first ultrasonic signal via the ultrasonic sensor.
[0161] Since the second device can receive multiple signals or multiple sound wave signals in the environment, the second device can judge the received signal and confirm that the received signal is the first ultrasonic signal used to establish the sound receiving group.
[0162] In one implementation, the second device may determine the frequency of the first ultrasonic signal. In one example, the number of vibrations per second of the first ultrasonic signal may be determined as the frequency of the first ultrasonic signal. If the second device determines that the frequency of the first ultrasonic signal is the first frequency, the second device may determine that the received signal is the first ultrasonic signal used to establish the sound reception group.
[0163] In another implementation, the second device may decode the received signal using an ultrasonic demodulation circuit. If the decoded information is first preset information, the received signal may be a first ultrasonic signal for establishing a sound reception group. The first preset information may be a request to establish a sound reception group.
[0164] In another implementation, the second device decodes the first ultrasonic signal. If the decoded information includes a device identifier and / or an IP address, the device identifier and IP address can be compared with a device list locally stored on the second device to determine that the first device is a trusted device.
[0165] If the second device determines that the received signal is not the first ultrasonic signal for establishing the sound reception group, it may wait to receive other ultrasonic signals and determine whether the subsequently received signal is the first ultrasonic signal for establishing the sound reception group.
[0166] S602: The second device sends a second ultrasonic signal, where the second ultrasonic signal carries information confirming joining the audio receiving group.
[0167] As an example, after the AI engine in the system library of the second device is called, it can send an ultrasonic signal transmission control signal to the audio driver in the kernel layer, so that the audio driver drives the speaker to transmit ultrasonic signals. It can also send an ultrasonic signal transmission control signal to the sensor driver in the kernel layer, so that the sensor driver drives the ultrasonic sensor to transmit ultrasonic signals.
[0168] In one implementation, the second device may transmit a second ultrasonic signal via an ultrasonic transmitting device. The second ultrasonic signal may carry information confirming joining the audio receiving group.
[0169] For example, the second ultrasonic signal sent by the second device may be an ultrasonic signal of a second frequency, wherein the second frequency may be used to indicate that the second ultrasonic signal carries information confirming joining the audio reception group.
[0170] In one example, the ultrasonic transmitting device may be implemented as a speaker, and the speaker transmits or plays the second ultrasonic signal.
[0171] In one example, the ultrasonic transmitting device may be implemented as an ultrasonic sensor, and the ultrasonic sensor transmits the second ultrasonic signal.
[0172] In another implementation, the second device may generate a second ultrasonic signal using an ultrasonic modulation circuit. The second device may use encoding technology to cause the generated second ultrasonic signal to carry information confirming joining the receiving group. The information confirming joining the receiving group may be second preset information, and decoding the second ultrasonic signal may reveal that the information carried by the second ultrasonic signal is the information confirming joining the receiving group.
[0173] In addition, the second ultrasonic signal may also include a second device identifier and / or an IP address of the second device. Alternatively, the confirmation information for joining the reception group may include the second device identifier and / or the IP address of the second device. This facilitates the first device to identify the second device in the reception group or establish a communication connection with the second device.
[0174] S603: The first device determines the device that sends the second ultrasonic signal as a device in the sound receiving group.
[0175] As an example, after the AI engine in the system library of the first device is called, the ultrasonic signal reception control signal can also be sent to the audio driver of the kernel layer, so that the audio driver drives the microphone to receive the ultrasonic signal, and sends the received ultrasonic signal to the processor of the first device. The signal codec of the processor decodes the received ultrasonic signal to determine whether the received ultrasonic signal carries confirmation information for joining the receiving group.
[0176] In one implementation, an ultrasonic receiving device in a first device may receive a second ultrasonic signal emitted by a second device. In one example, the ultrasonic receiving device may be implemented as a microphone, and the first device may receive or collect the first ultrasonic signal via the microphone. In one example, the ultrasonic receiving device may be implemented as an ultrasonic sensor, and the first device may receive, collect, or detect the first ultrasonic signal via the ultrasonic sensor.
[0177] In another implementation, the first device may determine the frequency of the second ultrasonic signal. In one example, the number of vibrations per second of the second ultrasonic signal may be determined as the frequency of the second ultrasonic signal. If the first device determines that the frequency of the second ultrasonic signal is the second frequency, it may determine that the second ultrasonic signal carries information confirming joining the receiving group.
[0178] In another implementation, the first device may decode the received second ultrasonic signal using an ultrasonic demodulation circuit. If the decoded information is the second preset information, it may be determined that the information carried by the second ultrasonic signal is confirmation of joining the audio reception group.
[0179] In another implementation, the first device decodes the second ultrasonic signal. If the decoded information includes a device identifier and / or an IP address, the first device determines that the second device corresponding to the device identifier agrees to join the receiving group and that the second device has a receiving function.
[0180] If the first device determines that the second ultrasonic signal does not carry the information confirming joining the sound receiving group, it may wait to receive another ultrasonic signal, and may determine whether the newly received ultrasonic signal carries the information confirming joining the sound receiving group.
[0181] S604: The first device communicates with the second device based on the Transmission Control Protocol (TCP).
[0182] In practical applications, the first device and the second device may be connected to the same network, such as the same WiFi network. A three-way handshake or a four-way handshake process implemented using a TCP connection may be used to establish a trust relationship between the first device and the second device. TCP-based communication can provide reliable data transmission between the first device and the second device. For example, this can include data such as the scene of the environment in which the sound receiving group is located, audio signals, or audio quality rating data.
[0183] In one implementation, the first device uses the IP address carried in the received second ultrasonic signal to initiate a three-way handshake process or a four-way handshake process for establishing a TCP connection.
[0184] The first device and the second device complete the three-way handshake process or the four-way handshake process of establishing a TCP connection, indicating that the first device and the second device have established a sound reception group.
[0185] Furthermore, the first device may add the identification information of the second device that has successfully established a TCP connection with the first device to the sound reception group list or database. Optionally, the first device may also add the identification information of the first device to the sound reception group list or database.
[0186] In one implementation, Figure 9 A schematic diagram of the process of a first device establishing a radio group is shown. During the instantiation of the fusion radio module in the AI engine in the first device, after sending a first ultrasonic signal for establishing a radio group, the device that sent the second ultrasonic signal, such as the second device a and the second device b, is added to the radio group list or database based on the receipt of a second ultrasonic signal carrying information confirming joining the radio group. The first device can transmit data with the second device a and the second device b in the radio group through TCP communication (or other communication methods).
[0187] Optionally, if the first device triggers the establishment of a radio group again, the device identification information in the newly established radio group list or database can overwrite the radio group list or database established last time, or the device identification information in the previously established radio group list or database can be cleared before the radio group list or database is established this time.
[0188] After establishing a sound collection group, the devices in the sound collection group can use their microphones to collect sound signals from the environment and process them into collected audio. Each device can collect audio signals from the scene synchronously, and the collection period (or frequency) and collection duration can also be the same.
[0189] The sound signals in the environment can be sound waves, and the processed audio can be electrical signals. For ease of description, the audio collected by each device is actually the audio generated by processing the collected sound signals. The audio segmentation module of the AI engine in each device can then determine the scene corresponding to the environment in which the sound collection group is located.
[0190] A2, the process of determining the scenario includes:
[0191] Scene determination, also known as scene recognition, is the process of identifying the scene corresponding to the environment in which the sound is received. It is also the process of determining the scene type corresponding to the environment in which the sound is received from multiple scene types. Scene determination can be performed using one or more methods, such as semantic recognition, image recognition, and geographic location.
[0192] In one implementation, taking semantic recognition as an example, after a sound collection group is established, the devices in the sound collection group use semantic recognition technology to identify semantic information in the collected audio, identify keywords contained in the audio, or identify keywords that represent intent in the audio.
[0193] For example, in the audio clip "Let me share with you about spring in nature," devices in the listening group can use semantic recognition technology to identify the keywords "share," "nature," and "spring" that reflect the intended meaning. Based on the correspondence between scenes and keywords, the identified keywords in the audio clip can be used to determine that the scene corresponding to the listening group's environment is an outdoor live broadcast.
[0194] For example, in an audio clip titled "Let me show you how to lead teammates to rank up," devices in the listening group can use voice recognition technology to identify the keywords "show," "teammates," and "rank up" that reflect the intended meaning. Based on the mapping between scenarios and keywords, the identified keywords in the audio clip can be used to determine that the scenario corresponding to the listening group's environment is a live game broadcast.
[0195] Exemplarily, the semantic recognition method may be automatic speech recognition (ASR).
[0196] The correspondence between the scene type and the keyword is pre-established or set, and the devices in the sound receiving group may pre-establish a database or a data table to store the pre-set correspondence between the scene and the keyword.
[0197] For example, a scene (or scene type) may correspond to multiple keywords, and a keyword may also correspond to multiple scenes. The devices in the sound receiving group may determine the scene corresponding to the combination of multiple keywords by matching.
[0198] For example, a keyword may have a one-to-one correspondence with a scene, and the devices in the sound receiving group may determine a unique corresponding scene based on a keyword.
[0199] For example, after the AI engine in the system library of a device in a sound collection group is called, it sends an audio collection control signal to the kernel-layer audio driver, which drives the microphone to collect sound signals, converts the collected sound signals into audio, and feeds it back to the AI engine. The AI engine uses semantic recognition technology to identify keywords in the audio and, based on the corresponding relationship between the scene and the keyword, determines the scene corresponding to the keyword.
[0200] In another implementation, taking image recognition as an example, after a sound collection group is established, the devices in the sound collection group recognize the image captured by the camera and determine the keyword corresponding to the image through image recognition.
[0201] For example, the devices in the audio reception group use an intelligent recognition algorithm based on deep learning to detect objects in the image and determine the names of the recognized objects as keywords corresponding to the image. Then, based on the correspondence between the scene and the keywords, and the keywords determined by image recognition, the scene type of the scene in which the audio reception group is located is determined. For example, Figure 5 The device in the scene shown, the image captured by the camera is Figure 10 The image shown in FIG. Using image recognition method to identify Figure 10 The image shown can identify the objects included in the image and determine that the names of the identified objects are "beach", "coconut tree", and "sea". The keywords determined by image recognition are "beach", "coconut tree", and "sea". Based on a database that pre-establishes the correspondence between scenes and keywords, the scene corresponding to the environment of the radio group can be determined to be an outdoor seaside live broadcast scene based on the keywords "beach", "coconut tree", and "sea".
[0202] For example, after the AI engine in the system library of a device in the audio reception group is invoked, it sends an image acquisition control signal to the camera driver in the kernel layer, instructing the camera driver to capture an image and feed it back to the AI engine. The AI engine then uses imaging technology to identify objects in the image and uses the names of the identified objects as keywords. Based on the correspondence between the scene and the keyword, the scene type corresponding to the keyword is determined.
[0203] In another implementation, taking geolocation as an example, after establishing a listening group, the devices in the group use satellite positioning technology or algorithms to determine the geographic location information of the devices in the group. For example, the geographic location information may include the name of the location, such as a certain bathing beach in Beidaihe. It may also include latitude and longitude coordinates, or coordinates in a geodetic coordinate system. Based on the geographic location information, the devices in the group determine the keywords corresponding to the geographic location information.
[0204] For example, the geographic location information is "Beidaihe bathing beach", and the corresponding keyword is "bathing beach". Based on the correspondence between scenes and keywords, the keyword "bathing beach" can be used to determine that the scene corresponding to the environment in which the audio receiving group is located is an outdoor seaside live broadcast scene.
[0205] As an example, after the AI engine in the system library of the device in the receiving group is called, it obtains the geographic location information from the content provider, determines the keyword corresponding to the geographic location name in the geographic location information, and determines the scene corresponding to the keyword based on the correspondence between the scene and the keyword.
[0206] In another implementation, devices in a receiving group can perform scene recognition by combining any two or three of the following methods: semantic recognition, image recognition, and geographic location. Taking the combination of semantic recognition, image recognition, and geographic location as an example, the devices in the receiving group can more accurately and finely determine the scene corresponding to the receiving group's environment, such as an indoor game livestream, based on the correspondence between scenes and keywords, using keywords from semantic recognition, keywords from image recognition, and keywords corresponding to geographic location names.
[0207] In another implementation, the first device in the sound receiving group may also provide the user with a scene selection interface, such as Figure 11 As shown, by identifying the menu option or button selected by the user in the scene selection interface, the scene type corresponding to the identified menu option or button is used as the scene corresponding to the environment of the radio group. For example, a webcast scene, a TV live broadcast scene, a film and television recording scene, and a music recording scene can be buttons belonging to the same level of menu. If the user selects the webcast scene button, the next level menu corresponding to the webcast scene can be displayed to the user, for example, including buttons such as indoor live broadcast scene, outdoor live broadcast scene, game live broadcast scene, and shopping live broadcast scene. The scenes in the next level menu are more specific, and the scene division granularity is smaller. In actual application scenarios, scenes can be set according to business needs, application needs, and other aspects.
[0208] As an example, after the AI engine in the system library of the device in the receiving group is called, it will provide the user with a scene selection interface and send it to the display driver of the kernel layer, so that the display driver will drive the display screen to display the scene selection interface, and according to the scene selection instruction triggered by the user, the scene corresponding to the scene selection instruction will be used as the scene corresponding to the environment of the receiving group.
[0209] In some embodiments, each device in the sound receiving group can perform scene recognition by combining any one or more of semantic recognition, image recognition, and geographic location positioning.
[0210] In some embodiments, in order to reduce the computing power of the devices in the sound receiving group, it is not necessary for each device in the sound receiving group to perform the scene determination process. The scenes are determined collaboratively by the devices in the sound receiving group. The collaborative scene determination is described in the following example.
[0211] Example 1: Scene recognition is performed by a device in the sound receiving group.
[0212] For example, the first device may perform scene recognition using one or more of the three methods described in the above embodiments: semantic recognition, image recognition, and geographic location. The first device may establish a TCP connection with the second devices in the sound receiving group and send the determined scene to all second devices in the sound receiving group. This eliminates the need for the second devices in the sound receiving group to perform scene recognition, reducing computing resource consumption.
[0213] For another example, the target second device in the sound receiving group uses one or more of the three methods of semantic recognition, image recognition, and geographic location positioning in the above embodiments to perform scene recognition. The target second device can be any second device in the sound receiving group. The target second device sends the determined scene to the first device through the TCP connection established with the first device. The first device then sends the scene to all second devices in the sound receiving group except the target second device. The first device is not required to perform scene recognition, which reduces the load on computing resources in the first device.
[0214] Example 2: Scene recognition is performed by multiple devices in a sound receiving group.
[0215] The first and second devices in the audio reception group have different hardware structures, and therefore can implement different scene recognition methods. For example, if the second device is a smart speaker, since smart speakers do not have cameras or video modules, scene recognition cannot be performed through image recognition. Therefore, other electronic devices in the audio reception group that can implement image recognition technology can perform scene recognition through image recognition.
[0216] In one implementation, the devices in the sound receiving group send the determined keywords to the first device. The first device then determines the scene corresponding to the environment in which the sound receiving group is located based on the keywords determined by the devices in the sound receiving group and the corresponding relationship between the scenes and the keywords. The devices in the sound receiving group can perform scene recognition using one or more of the following methods: semantic recognition, image recognition, and geographic location.
[0217] For example, device A in the audio collection group can use semantic recognition to identify keywords in the collected audio and image recognition to identify keywords in the collected images. Device B in the audio collection group can use satellite positioning technology to determine the keywords corresponding to the geographical location information.
[0218] If device A is the first device and device B is the second device, device B sends the determined keyword to device A via a TCP connection. Device A determines the scene corresponding to the environment of the audio reception group based on the locally determined keyword and the received keyword. Device A sends the determined scene to the devices in the audio reception group, or to all second devices in the audio reception group except device B.
[0219] For example, device A in the audio collection group can use semantic recognition to identify keywords in the collected audio. Device B in the audio collection group can use image recognition to identify keywords in the collected images. Device C in the audio collection group can use satellite positioning technology to determine the geographic location information corresponding to the keyword.
[0220] If device C is the first device, and devices A and B are the second devices, device A sends the determined keyword to device C via a TCP connection. Device B also sends the determined keyword to device C via a TCP connection. Device C determines the scene corresponding to the environment of the sound reception group based on the locally determined keyword and the received keyword. Device C sends the determined scene to the devices in the sound reception group, or to all second devices in the sound reception group except devices A and B.
[0221] For example, device A in a sound collection group can use semantic recognition to identify keywords in the collected audio. Devices B, C, and D in the sound collection group can use image recognition to identify keywords in the collected images. Device E in the sound collection group can use keywords corresponding to geographic location information determined by satellite positioning technology.
[0222] If device A is the first device and devices B, C, D, and E are the second devices, devices B, C, D, and E each send their determined keywords to device A via a TCP connection. Device A determines the scene corresponding to the environment of the audio reception group based on the locally determined keywords and the received keywords. Device A then sends the determined scene to all devices in the audio reception group, or to all second devices in the audio reception group except devices B, C, D, and E.
[0223] In one implementation, each device in the sound receiving group sends the determined keyword to the first device. The first device then determines the scene type of the scene in which the sound receiving group is located based on the keywords determined by the devices in the sound receiving group and the correspondence between the scene and the keyword. The devices in the sound receiving group can perform scene recognition using one or more of the following methods: semantic recognition, image recognition, and geographic location.
[0224] In another implementation, the devices in the sound collection group can complete the above-mentioned scene determination process within a preset number of collection cycles, for example, within the first two collection cycles, without having to perform the scene determination process in each collection cycle.
[0225] After the sound collection group determines the scene, each device can combine the audio of at least one sound type in the determined scene to specifically extract audio of various sound types in the scene from the collected audio to be processed.
[0226] A3, the audio segmentation process includes:
[0227] Each device in the sound receiving group collects audio signals from the scene, and uses the AI engine audio segmentation module in each device to determine the scene. Based on the correspondence between the scene and the sound type, the sound type included in the scene corresponding to the environment in which the sound receiving device is located is determined. The collected audio to be processed is segmented to obtain audio corresponding to the sound type included in the scene. If the sound type included in the scene includes a human voice type, the audio to be processed is segmented to obtain human voice audio. If the sound type included in the scene includes at least one environmental sound type, the audio to be processed is segmented to obtain audio without human voice, and the audio without human voice is segmented to obtain audio of at least one environmental sound type.
[0228] The correspondence between scenes and sound types is pre-established, and the devices in the sound receiving group can use a database to store the pre-established correspondence between scene types and sound types. A scene has a correspondence with at least one sound type. The sound type can be a human voice type and / or an environmental sound type. The database also stores audio features corresponding to each environmental sound type. Type can also refer to category. Human voice type can also include female voice type, male voice type, children's voice type, etc. Environmental sound can refer to non-human sounds in the environment and can be classified according to the object that produces the sound.
[0229] For example, the ambient sound types included in the outdoor beach live broadcast scene may be the sound of waves, the sound of wind, etc. For another example, the ambient sound types included in the outdoor grassland live broadcast scene may be the sound of horses, the sound of horses running, etc.
[0230] In actual application scenarios, the audio collected by each device for processing includes both ambient sound and noise. If someone is speaking or making sounds in the scene, the collected audio for processing may also include human voices. Because each device is located in a different location within the scene, the types of sounds collected by each device may not be completely consistent.
[0231] If at least one of the sound types included in the scene includes an ambient sound type, each device can first identify whether the audio to be processed contains human voices. If it is determined that the audio to be processed contains human voices, each device can segment the collected audio to obtain audio without human voices. The audio without human voices can then be segmented to obtain audio corresponding to the ambient sound type.
[0232] If at least one sound type included in the scene includes a human voice type, each device can first identify whether the audio to be processed contains a human voice. If it is determined that the audio to be processed contains a human voice audio, each device will segment the collected audio to be processed to obtain audio corresponding to the human voice type.
[0233] like Figure 12 A process for implementing audio segmentation based on at least one sound type included in a scene is shown, comprising the following steps:
[0234] S1201: Each device performs frame processing on the audio to be processed.
[0235] Each device performs frame processing on the audio to be processed. For the convenience of description, each frame signal of the collected audio to be processed after the frame processing is recorded as a first signal.
[0236] S1202: Each device obtains the Mel-frequency cepstrum feature of the first signal of each frame.
[0237] At present, the Mel-frequency cepstrum feature has an excellent performance in extracting human voice features. Therefore, each device can determine the Mel-frequency cepstrum feature of the first signal of each frame. It should be understood that the embodiment of the present application only uses the Mel-frequency cepstrum feature as an example for illustration, and does not limit the expression method of the sound feature. In the specific implementation, other expressions can also be used to represent the sound feature. For example, in the field of human voice recognition, if there is a better sound feature representation method in the future, the sound feature used for human voice segmentation in the embodiment of the present application can be replaced with a better sound feature representation method. Each device can obtain the Mel-frequency cepstrum feature of the first signal of each frame by determining the Mel-frequency cepstrum feature of the first signal of each frame. Each device can also obtain the Mel-frequency cepstrum feature of the first signal of each frame by receiving the Mel-frequency cepstrum feature of the first signal of each frame determined by an external device.
[0238] S1203: Each device clusters the Mel-frequency cepstrum features of the first signal of each frame.
[0239] Each device uses clustering technology to cluster the Mel-frequency cepstrum features of the first signal of each frame to determine a clustering result for each frame. The clustering results can be human voice or non-human voice. Based on the clustering results, human voice audio is determined based on the first signal of each frame that is classified as human voice.
[0240] S1204: Each device generates human voice audio according to the clustering result of each frame of the first signal.
[0241] In one implementation, each device can sort the clustered signals into human voice categories based on the start and end times of the first signal in each frame. If the end time of the previous sequence signal and the start time of the next sequence signal are inconsistent between two adjacent sequence signals, a silent signal can be inserted between the previous and next sequence signals. The processed signal is then recorded as human voice audio, ensuring that the audio duration of the human voice audio matches the duration of the audio to be processed.
[0242] S1205 : Each device generates voice-free audio according to the clustering result of the first signal of each frame.
[0243] In one implementation, each device replaces the voice-identified signal in the audio being processed with a silent signal based on the start and end times of the voice-identified signal clustered as a human voice. This removes the human voice from the processed audio, and the processed signal is recorded as the non-voice audio. It should be noted that the non-voice audio typically also includes ambient sounds and noise in the scene.
[0244] The audio to be processed is identified. If it is determined that the audio to be processed contains human voice audio, the audio to be processed is segmented to obtain non-voice audio. The non-voice audio can then be identified or matched to determine whether it contains audio of a first ambient sound type, where the first ambient sound type is any of at least one sound type corresponding to the scene. When identifying or matching whether the non-voice audio contains an ambient sound type included in the scene, the ambient sound types included in the scene can be identified or matched one by one.
[0245] If it is determined that the audio to be processed does not include human voice audio, it can be identified or determined from the audio to be processed whether it includes audio of the first environmental sound type included in the scene.
[0246] It should be noted that the process of identifying whether audio without human voice contains audio of the type of ambient sound included in the scene is the same as the process of identifying whether audio to be processed that does not include human voice audio contains audio of the type of ambient sound included in the scene, and will not be repeated here.
[0247] S1206: Each device performs frame processing on the audio without human voice.
[0248] For the convenience of description, each frame signal obtained after the unvoiced audio is frame-processed is recorded as the second signal below.
[0249] S1207: Each device determines the audio features of each frame of the second signal.
[0250] In one implementation, each device determines the audio features of each frame of the second signal and may extract time-domain features from the time-domain signal of each frame of the second signal as the audio features. Alternatively, the device may extract frequency-domain features from the frequency-domain signal of each frame of the second signal as the audio features. Alternatively, the device may extract time-frequency domain features from the time-frequency domain signal of each frame of the second signal as the audio features.
[0251] For example, for each frame signal, each device can combine or splice any two or three of the extracted time domain features, frequency domain features, and time-frequency domain features as the audio features of the frame signal.
[0252] S1208: Determine the ambient sound type corresponding to each frame of the second signal based on the characteristics of the ambient sound type included in the scene.
[0253] Each device matches the audio features of each frame of the second signal with the audio features of the ambient sound type included in the scene, and determines the ambient sound type that matches the audio features of the frame signal.
[0254] For example, each device determines the similarity between the audio features of the frame signal and the audio features of the ambient sound type included in the scene. If the similarity with the audio features of a certain ambient sound type is greater than or equal to a preset threshold, such as 0.8, it is determined that the audio features of the frame signal match the ambient sound type. If the similarity with the audio features of a certain ambient sound type is less than the preset threshold, it is determined that the audio features of the frame signal do not match the ambient audio. The ambient sound type that matches the audio features of the frame signal is determined as the ambient sound type of the frame signal.
[0255] The audio features of the second signal of each frame are matched with the audio features of the ambient sound type included in the scene, and are not matched with the audio features of other ambient sound types (audio of ambient sound types not included in the scene), thereby reducing the computing consumption of each device and shortening the processing time for segmenting the ambient audio.
[0256] S1209 , generating audio of each ambient sound type according to the ambient sound type corresponding to the second signal of each frame.
[0257] In one implementation, for multiple frames of signals belonging to the same ambient sound type, the multiple frames of signals belonging to the same ambient sound type can be sorted according to the start and end time of the second signal of each frame. If the end time of the previous sequence signal of two adjacent sequence signals is inconsistent with the start time of the next sequence signal, a silent signal can be inserted between the previous sequence signal and the next sequence signal, and the processed signal is recorded as audio corresponding to the ambient sound type, such as a certain type of ambient audio, to obtain audio of the ambient sound type. Using this method, the audio duration of each type of ambient audio can be made the same as the duration of the audio to be processed.
[0258] For example, the determined scene is an outdoor beach live broadcast scene type, and the sound types included in the outdoor beach live broadcast scene include environmental sound types such as ocean wave sound type and wind sound type. The unvoiced audio is framed and the audio features of the second signal of each frame are extracted. The audio features of the second signal of each frame are matched with the audio features of the ocean wave sound type and the audio features of the wind sound type stored in the database to determine the environmental sound type corresponding to the second signal of each frame. The second signals of each frame that match the audio features of the ocean wave sound type are spliced together to obtain ocean wave sound type audio. The second signals of each frame that match the audio features of the wind sound type are spliced together to obtain wind sound type audio.
[0259] If the audio features of the second signal of each frame do not match the audio features of the wind sound type stored in the database, but the audio features of the second signal match the audio features of the ocean wave sound type stored in the database, the ocean wave audio is generated using the second signal of each frame that matches the audio features of the ocean wave sound type.
[0260] Each device in the audio collection group segments the collected audio to be processed, obtaining audio of the sound types included in the scene as candidate audio. Each device can use the audio scoring module in the AI engine to determine the quality score of the candidate audio.
[0261] As an example, after the AI engine in the electronic device system library is called, the audio acquisition control signal can be sent to the audio driver of the kernel layer, so that the audio driver drives the microphone to collect sound signals, converts the collected sound signals into audio, and feeds it back to the AI engine. The AI engine extracts the human voice features from the audio, segments the audio, and obtains human voice audio and non-voice audio. Based on the ambient sound type included in the sound type included in the determined scene, the AI engine extracts the ambient audio features from the non-voice audio, matches them with the audio features corresponding to the ambient sound type included in the sound type included in the scene, segments the non-voice audio based on the matching results, and obtains audio corresponding to the ambient sound type.
[0262] A4, the audio scoring process includes:
[0263] Each device in the audio collection group segments the audio to be processed based on at least one sound type identified in the scene. The audio scoring module in the AI engine then assigns a quality score to the candidate audio. The candidate audio can be either human voice or ambient sound.
[0264] First, each device determines a score of the third signal of each frame of the candidate audio.
[0265] The candidate audio is framed and each frame signal after the frame processing is recorded as a third signal of each frame.
[0266] In one implementation, the candidate audio is framed and the log-Mel coefficient of the third signal of each frame is determined. The log-Mel coefficient of the third signal of each frame is input into a pre-trained convolutional neural network model, the log-Mel coefficient of the third signal of each frame is scored, and the score of the third signal of each frame is output.
[0267] For example, after the framing process, the third signal of each frame can be pre-processed by level calibration and filtering, and then the pre-processed third signal of each frame can be subjected to fast Fourier transform, and the logarithmic Mel coefficient of the third signal of each frame can be determined from the spectrum information after the fast Fourier transform.
[0268] Among them, in the process of training the convolutional neural network model, such as Figure 13 The following is a flow chart showing an exemplary process of training a convolutional neural network model, including the following steps:
[0269] S1301, collect samples.
[0270] The collected samples are used to train the convolutional neural network model. The training samples can be audio samples from different scenarios. The collected samples include audio signals without human voices, that is, audio signals containing pure ambient sounds.
[0271] S1302, label samples.
[0272] Labeling samples also means setting labels or tagging them. The audio quality score is marked as the label of the sample. In one example, the audio score of each audio sample is determined by manual labeling, and the mean opinion score (MOS) evaluation method can be used. For example, the absolute category rating (ACR) or the degradation category rating (DCR) in the mean opinion score evaluation method is used.
[0273] For example, the training samples can be labeled according to evaluation criteria such as clarity, delay time, and fluency, as shown in Table 1, which shows the corresponding relationship between the evaluation criteria and the MOS value.
[0274] Table 1
[0275]
[0276] S1303, training model.
[0277] The neural network model is trained using samples and their labels as input and sample labels as output. During the training process, if the accuracy of the output sample labels exceeds a preset accuracy threshold, the training process can be terminated.
[0278] S1304: Generate a convolutional neural network model.
[0279] The trained convolutional neural network model can be used to determine the score of each frame signal. Each device can use the generated convolutional neural network model to determine the score of each frame third signal.
[0280] Each device determines a quality score of the candidate audio according to the score of the third signal of each frame of the candidate audio.
[0281] In one implementation, the electronic device may input the third signal of each frame of any candidate audio and the score of the frame signal into a pre-trained long short-term memory artificial neural network model, predict the audio quality of the candidate audio, and output the audio score of the candidate audio.
[0282] As an example, after the AI engine in the electronic device system library is called, it can send audio collection control signals to the core layer audio driver, which drives the microphone to collect sound signals, converts the collected sound signals into audio, and feeds it back to the AI engine. The AI engine segments the collected audio to be processed, obtains candidate audio, and uses convolutional neural network models and long short-term memory artificial neural network models to determine the quality scores of the candidate audio.
[0283] A5, the audio fusion process includes:
[0284] Each device in the audio collection group can determine the quality score of the candidate audio obtained locally by the device. The first device can perform audio fusion in any of the following two ways, but is not limited to the following.
[0285] Method 1:
[0286] Each device in the audio collection group segments the audio to be processed, collected during each sampling period, to obtain candidate audio and determine the quality score of the candidate audio. Each second device in the audio collection group sends the obtained candidate audio and the quality score of the candidate audio to the first device, which then performs audio fusion.
[0287] In one example, Figure 14 A Figure 5 The implementation process of audio fusion of collected audio in the sound receiving group in . The first device performs audio segmentation on the locally collected audio to be processed 1 to obtain human voice audio 1 and non-voice audio 1. The first device performs audio segmentation on the non-voice audio 1 to obtain three types of ambient sound audio, namely the first type of ambient audio 11, the second type of ambient audio 12, and the third type of ambient audio 13. The first device determines the quality scores of the human voice audio 1 and each type of ambient audio, and determines that the score of the human voice audio 1 is A1, the score of the first type of ambient audio 11 is B1, the score of the second type of ambient audio 12 is C1, and the score of the third type of ambient audio 13 is D1. Similarly, the second device a segments the audio to be processed 2 to obtain human voice audio 2, the first type of ambient audio 21, and the second type of ambient audio 22, and determines that the score of the human voice audio 2 is A2, the score of the first type of ambient audio 21 is B2, and the score of the second type of ambient audio 22 is C2. The third device b segments the audio 3 to be processed to obtain human voice audio 3, first-class environmental audio 31, and second-class environmental audio 32, and determines that the score of human voice audio 3 is A3, the score of first-class environmental audio 31 is B3, and the score of second-class environmental audio 32 is C3.
[0288] The process of the first device dividing the audio to be processed 1, the process of the second device a dividing the audio to be processed 2, and the process of the second device b dividing the audio to be processed 3 can be specifically referred to. Figure 12 The method shown will not be repeated here.
[0289] In one implementation, when the first device performs audio fusion, the first device can select the audio with the highest audio quality score from all the candidate audios corresponding to the human voice type in the sound receiving group as a candidate audio corresponding to the human voice type. For the first environmental sound type included in the sound types included in the scene, the first device can select the audio with the highest audio quality score from all the candidate audios corresponding to the first environmental sound type as a candidate audio corresponding to the first environmental sound type, wherein the first environmental sound type is any environmental sound type included in the sound types included in the scene. The first device performs audio fusion on a candidate audio corresponding to each determined sound type. The fused audio is not only noise-free, but also has high-quality human voices and contains environmental scene sounds.
[0290] For example, Figure 14 In the example, the first device selects human voice audio 1, which has the highest score, as the human voice audio to be fused, based on score A1 of human voice audio 1, score A2 of human voice audio 2, and score A3 of human voice audio 3. Based on score B1 of first type ambient audio 11 and score B2 of first type ambient audio 21, the first device selects first type ambient audio 11, which has the highest score, as the ambient audio to be fused with the first type ambient audio. Based on score C1 of second type ambient audio 12 and score C2 of second type ambient audio 22, the first device selects second type ambient audio 22, which has the highest score, as the ambient audio to be fused with the second type ambient audio. Based on score D1 of third type ambient audio 12 and score D3 of third type ambient audio 32, the first device selects third type ambient audio 33, which has the highest score, as the ambient audio to be fused with the third type ambient audio. The first device performs audio fusion on the human voice audio to be fused and the ambient audio to be fused of the various types of ambient audio, namely, human voice audio 1, first type ambient audio 11, second type ambient audio 22, and third type ambient audio 33.
[0291] In another implementation, the first device may perform audio fusion on the human voice audio to be fused and the ambient audio to be fused, both of which have audio scores greater than or equal to a fusion audio score threshold. The fused audio is not only noise-free but also contains high-quality human voice and ambient sounds, allowing the user to clearly hear both the human voice and the ambient sounds. The fusion audio score threshold may be pre-set.
[0292] For example, if Figure 14The quality score A1 of the human voice audio 1, the quality score B1 of the first type of environmental audio 11, and the quality score D3 of the third type of environmental audio 33 shown in the figure are all greater than the fusion audio score threshold, while the quality score C2 of the second type of environmental audio 22 is less than the fusion audio score threshold. In this case, the human voice audio 1, the first type of environmental audio 11, and the third type of environmental audio 33 are fused to obtain the fused audio.
[0293] Method 2:
[0294] In order to reduce the computing resources consumed by the AI engine in each second device for audio segmentation and audio scoring, and shorten the processing time. The first device can use the candidate audio and audio quality scores sent by each device in the sound receiving group in a preset number of acquisition cycles to determine the device with better audio quality corresponding to the first sound type, and notify the device to send the candidate audio corresponding to the first sound type to the first device, where the first sound type can be any one of the at least one sound type included in the scene corresponding to the environment in which the sound receiving group is located. That is, the first device determines that the second device sends a specified or fixed type of audio to the first device. In one example, the first device can indicate the sound type corresponding to the candidate audio that the devices in the sound receiving group need to send subsequently by sending an indication message.
[0295] Upon receiving the instruction information from the first device, the second device may segment the audio to be processed according to the sound type in the instruction information, obtain candidate audio corresponding to the sound type indicated in the instruction information, and send the candidate audio to the first device. It should be noted that the second device may not perform audio scoring on the obtained candidate audio.
[0296] For example, Figure 15 Shown in Figure 5 The data interaction process of each device in the sound receiving group in the scene during the audio fusion process.
[0297] In each cycle before and including the nth cycle, each second device in the audio reception group sends the candidate audio and quality score to the first device. It should be noted that the cycle can be an audio collection cycle or a data transmission cycle. The present invention does not specifically limit this.
[0298] The first device may add the received candidate audio to the candidate audio set, or may add the candidate audio obtained by segmenting the audio to be processed by the first device to the candidate audio combination. The first device determines a candidate audio corresponding to each sound type from the candidate audio set. For example, the audio with the highest quality score among all candidate audios corresponding to the first sound type is determined as the candidate audio corresponding to the first sound type, and determines the device to which the candidate audio corresponding to the first sound type belongs. The first sound type may be any sound type of at least one sound type included in the scene corresponding to the environment in which the sound receiving group is located.
[0299] The first device performs audio fusion on a candidate audio corresponding to each determined sound type.
[0300] In the nth cycle, the first device can determine by statistical means, based on the quality scores of the candidate audios corresponding to the human voice type obtained by each device in the sound receiving group and the quality scores of the candidate audios corresponding to each environmental sound type, that one candidate audio with better quality among all the candidate audios corresponding to the human voice type belongs to the first device, one candidate audio with better quality among all the candidate audios corresponding to the first type of environmental sound type belongs to the first device, one candidate audio with better quality among all the candidate audios corresponding to the second type of environmental sound type belongs to the second device a, and one candidate audio with better quality among all the candidate audios corresponding to the third type of environmental sound type belongs to the second device b.
[0301] For example, the statistical method can be to calculate the weighted average of the quality scores of all candidate audios corresponding to the human voice type of each device in the sound receiving group in the first n cycles, and the weighted average of the quality scores of all candidate audios corresponding to the first environmental sound type of each device. The higher the weighted average, the better the audio quality can be considered, wherein the first environmental sound type is any environmental sound type of at least one sound type included in the scene corresponding to the environment in which the sound receiving group is located.
[0302] For example, the statistical method can be to calculate the frequency of the first sound type of each device in the first n cycles being determined by the first device as a candidate audio corresponding to the first sound type, wherein the first sound type is any sound type of at least one sound type included in the scene corresponding to the environment in which the sound receiving group is located. The higher the frequency, the better the audio quality can be considered.
[0303] The first device notifies the second device a to send candidate audio of the second type of ambient sound in each cycle after the nth cycle, and notifies the second device b to send candidate audio of the third type of ambient sound in each cycle after the nth cycle. In one example, the first device can notify the devices in the sound receiving group by sending indication information, where the indication information is used to indicate the sound type corresponding to the candidate audio sent by the device to the first device in subsequent cycles.
[0304] In each cycle after the nth cycle, the first device segments the collected audio signal to be processed, obtains audio corresponding to the human voice type and audio corresponding to the first type of environmental sound type, and performs audio fusion with the received audio corresponding to the second type of environmental sound type and the received audio corresponding to the third type of environmental sound type.
[0305] In one implementation, when a first device determines the designated sound type sent by each second device to the first device, if there is a second device in the sound collection group and the first device does not determine the designated sound type corresponding to this second device, this may indicate that the audio quality corresponding to the sound type contained in the scene of the audio to be processed collected by the second device is poor and unsuitable for audio fusion. The first device may notify the second device to stop collecting audio or stop sending any candidate audio and scores to the first device.
[0306] If the sound source in the scene is a mobile sound source, each sound receiving device in the sound receiving group divides the collected signal to be processed, and the audio quality of the obtained candidate audio is not stable, the first device can use method one to perform audio fusion.
[0307] If the sound source in the scene is fixed, each receiving device in the receiving group segments the collected signal to be processed, and the audio quality of the obtained candidate audio is relatively stable. The first device can then perform audio fusion using method 2. For example, the sound source in an outdoor live broadcast scene is usually a mobile sound source, while the sound source in a game live broadcast scene is usually a fixed sound source.
[0308] In another implementation, the first device provides a sound source type selection interface to the user, identifies the sound source type selected by the user, and performs audio fusion of human voice audio and various environmental audios according to the audio fusion method corresponding to the sound source type selected by the user.
[0309] In another implementation, the first device determines the sound source type corresponding to the determined scene based on the correspondence between the scene and the sound source type, and performs audio fusion according to the audio fusion method corresponding to the determined sound source type. The correspondence between the scene and the sound source type can be pre-established and stored in a database or data table. Each scene corresponds to a sound source type. For example, the sound source type corresponding to the game live broadcast scene is a fixed sound source. The first device can use method 2 to perform audio fusion of human voice audio and various types of environmental audio.
[0310] As an example, after the AI engine in the system library of the first device is called, it can determine a candidate audio corresponding to each sound type in at least one sound type included in the scene based on the candidate audio set and the quality score of each candidate audio in the candidate audio set, and perform audio fusion on the determined candidates, so that the audio after audio fusion includes audio corresponding to each sound type with better audio quality, thereby improving the user's auditory experience and enabling the user to feel or perceive the scene in the audio through the audio.
[0311] Based on the above embodiments, the present application also provides an audio optimization method, see Figure 16 As shown, the method may include the following steps:
[0312] S1601: Receive candidate audio sent by devices in a receiving group, where the receiving group includes at least one device. The candidate audio is audio of at least one sound type in a target scene, where the target scene is a scene corresponding to the environment where the devices in the receiving group are located, and each sound type corresponds to at least one candidate audio.
[0313] In one possible design, the sound type includes a human voice type and / or an environmental sound type.
[0314] In an embodiment of the present application, before receiving the candidate audio sent by the device in the sound receiving group, the method further includes: determining a target scene.
[0315] In one possible design, determining the target scene includes: obtaining keywords of the environment; and determining the scene corresponding to the keywords of the environment based on the correspondence between the scene and the keywords.
[0316] In an embodiment of the present application, obtaining keywords of the environment includes: obtaining keywords of the environment in at least one of the following ways: determining keywords of the environment; and receiving keywords sent by at least one device in a sound receiving group.
[0317] Furthermore, determining the keywords of the environment includes: determining the keywords of the environment by at least one of the following methods: performing semantic recognition on the collected audio to be processed to determine the keywords of the audio to be processed; performing image recognition on the collected image of the environment to determine the keywords corresponding to the image; and determining the keywords corresponding to the geographic location information based on the geographic location information of the environment.
[0318] In one possible design, after determining the target scene and before receiving the candidate audio sent by the devices in the sound receiving group, the method also includes: notifying the devices in the sound receiving group of the target scene.
[0319] In one possible design, before receiving the candidate audio sent by a device in the receiving group, the method also includes: sending a first ultrasonic signal for establishing a receiving group; receiving a second ultrasonic signal, the second ultrasonic signal carrying information confirming joining the receiving group; and determining, based on the second ultrasonic signal, the device sending the second ultrasonic signal as a device in the receiving group.
[0320] S1602: Determine a candidate audio corresponding to each sound type from the candidate audio set, where the candidate audio set includes candidate audio sent by devices in the sound receiving group.
[0321] In one possible design, a candidate audio corresponding to each sound type is determined from a candidate audio set, including: obtaining quality scores of the candidate audios in the candidate audio set; and selecting a candidate audio whose quality score meets preset conditions from the candidate audios corresponding to each sound type as a candidate audio corresponding to each sound type.
[0322] Furthermore, obtaining the quality scores of the candidate audios in the candidate audio set includes: obtaining the quality scores of the candidate audios in the candidate audio set by using at least one of the following methods: determining the quality scores of the candidate audios in the candidate audio set; and receiving the quality scores of the candidate audios sent by the devices in the receiving group.
[0323] Furthermore, the preset condition is that the quality score is the highest, or the quality score is the highest and the quality score is greater than a score threshold.
[0324] Furthermore, the candidate audio set also includes candidate audio obtained by segmenting the collected audio to be processed.
[0325] In one possible design, after determining a candidate audio corresponding to each sound type, the method further includes: sending indication information to at least one device in the sound receiving group, where the indication information is used to indicate the sound type corresponding to the candidate audio that the device needs to send subsequently.
[0326] S1603: Perform audio fusion on the determined candidate audios to obtain audio after audio fusion.
[0327] The present application also provides an audio optimization method, see Figure 17 As shown, the method may include the following steps:
[0328] S1701: Collect audio to be processed.
[0329] In one possible design, before collecting the audio to be processed, the method also includes: receiving a third ultrasonic signal sent by the first device for establishing a sound receiving group; and sending a fourth ultrasonic signal, which carries information confirming joining the sound receiving group.
[0330] S1702: Segment the audio to be processed to obtain at least one candidate audio, where each candidate audio in the at least one candidate audio corresponds to a sound type of at least one sound type included in a target scene, where the target scene is a scene corresponding to the current environment.
[0331] In one possible design, the sound type includes a human voice type and / or an environmental sound type.
[0332] In one possible design, before segmenting the audio to be processed, the method also includes: receiving a target scene notified by the first device.
[0333] In a possible design, before segmenting the audio to be processed, the method further includes: determining a target scene.
[0334] Furthermore, determining the target scene includes: determining a keyword of the current environment; and determining a scene corresponding to the keyword of the current environment based on a correspondence between the scene and the keyword.
[0335] Furthermore, determining the keywords of the current environment includes: determining the keywords of the current environment by at least one of the following methods: performing semantic recognition on the collected audio to be processed to determine the keywords of the audio to be processed; performing image recognition on the collected image of the current environment to determine the keywords corresponding to the image; and determining the keywords corresponding to the geographic location information based on the geographic location information of the current environment.
[0336] In one possible design, the audio to be processed is segmented, including: if at least one sound type includes a first environmental sound type, the audio to be processed is segmented to obtain audio without human voice, and the audio without human voice is segmented to obtain audio of the first environmental sound type; if at least one sound type includes a human voice type, the audio to be processed is segmented to obtain audio of the human voice type.
[0337] In one possible design, before segmenting the audio to be processed, the method further includes: receiving indication information sent by the first device, where the indication information is used to indicate a sound type corresponding to the sent candidate audio. Segmenting the audio to be processed to obtain at least one candidate audio includes: segmenting the audio to be processed to obtain candidate audio corresponding to the sound type indicated in the indication information.
[0338] S1703: Send at least one candidate audio to the first device.
[0339] In one possible design, after segmenting the audio to be processed and obtaining at least one candidate audio, the method further includes: determining a quality score of each candidate audio of the at least one candidate audio; and sending the quality score of each candidate audio to the first device.
[0340] To implement the various functions of the methods provided in the embodiments of the present application, the mobile terminal device may include hardware structures and / or software modules, and the aforementioned functions may be implemented in the form of hardware structures, software modules, or a combination of hardware structures and software modules. Whether a particular one of the aforementioned functions is implemented in the form of hardware structures, software modules, or a combination of hardware structures and software modules depends on the specific application and design constraints of the technical solution.
[0341] In the embodiments provided in the present application above, the methods provided in the embodiments of the present application are introduced from the perspective of an electronic device as an execution subject. In order to implement the various functions in the methods provided in the embodiments of the present application above, the electronic device may include a hardware structure and / or a software module, and implement the above functions in the form of a hardware structure, a software module, or a hardware structure plus a software module. Whether a function of the above functions is executed in the form of a hardware structure, a software module, or a hardware structure plus a software module depends on the specific application and design constraints of the technical solution.
[0342] like Figure 18 As shown, some other embodiments of the present application disclose a device. Figure 18 As shown, the apparatus 1800 includes: one or more processors 1801 ; a memory 1802 ; and one or more computer programs 1803 (not shown in the figure), and the above components can be connected via one or more communication buses 1804 .
[0343] One or more computer programs are stored in the memory 1802, and the one or more computer programs include instructions. The processor 1801 calls the instructions stored in the memory 1802, so that the electronic device 1800 can perform the following steps:
[0344] Receive candidate audio signals sent by devices in a sound receiving group, where the sound receiving group includes at least one device, the candidate audio signals being audio signals of at least one sound type in a target scene, where the target scene is a scene corresponding to an environment in which the devices in the sound receiving group are located, and each sound type corresponds to at least one candidate audio signal;
[0345] Determining, from a candidate audio set, a candidate audio corresponding to each of the sound types, the candidate audio set including candidate audio sent by devices in the sound receiving group;
[0346] Perform audio fusion on the determined candidate audios to obtain audio fusion audio.
[0347] In a possible implementation, the sound type includes a human voice type and / or an environmental sound type.
[0348] In a possible implementation, when the instruction is called and executed by the one or more processors, the electronic device performs the following steps before receiving the candidate audio sent by the device in the receiving group: determining the target scene.
[0349] Specifically, when the instruction is called and executed by the one or more processors, the electronic device determines the target scene and performs the following steps: obtaining keywords of the environment; and determining the scene corresponding to the keyword of the environment based on the correspondence between the scene and the keyword.
[0350] Specifically, when the instruction is called and executed by the one or more processors, the electronic device obtains the keyword of the environment and performs the following steps: obtaining the keyword of the environment in at least one of the following ways:
[0351] Determine a keyword of the environment; and receive the keyword sent by at least one device in the sound receiving group.
[0352] Specifically, when the instruction is called and executed by the one or more processors, the electronic device determines the keywords of the environment and performs the following steps: determining the keywords of the environment in at least one of the following ways: performing semantic recognition on the collected audio to be processed to determine the keywords of the audio to be processed; performing image recognition on the collected image of the environment to determine the keywords corresponding to the image; and determining the keywords corresponding to the geographic location information based on the geographic location information of the environment.
[0353] In one possible implementation, when the instruction is called and executed by the one or more processors, the electronic device performs the following steps after determining the target scene: notifying the devices in the sound receiving group of the target scene before receiving the candidate audio sent by the devices in the sound receiving group.
[0354] In one possible implementation, when the instruction is called and executed by the one or more processors, the electronic device determines a candidate audio corresponding to each of the sound types from the candidate audio set, and performs the following steps: obtaining quality scores of the candidate audios in the candidate audio set; and selecting a candidate audio whose quality score meets a preset condition from the candidate audios corresponding to each of the sound types as a candidate audio corresponding to each of the sound types.
[0355] Specifically, when the instruction is called and executed by the one or more processors, the electronic device obtains the quality score of the candidate audio in the candidate audio set and performs the following steps: obtaining the quality score of the candidate audio in the candidate audio set using at least one of the following methods: determining the quality score of the candidate audio in the candidate audio set; receiving the quality score of the candidate audio sent by the device in the receiving group.
[0356] In a possible implementation, the preset condition is that the quality score is the highest, or that the quality score is the highest and is greater than a score threshold.
[0357] In a possible implementation, the candidate audio set further includes candidate audio obtained by segmenting the collected audio to be processed.
[0358] In one possible implementation, when the instruction is called and executed by the one or more processors, the electronic device performs the following steps: after determining a candidate audio corresponding to each of the sound types, sending indication information to at least one device in the sound receiving group, wherein the indication information is used to indicate the sound type corresponding to the candidate audio that the device needs to send subsequently.
[0359] In one possible implementation, when the instruction is called and executed by the one or more processors, the electronic device performs the following steps before receiving the candidate audio sent by the device in the receiving group: sending a first ultrasonic signal for establishing a receiving group; receiving a second ultrasonic signal, which carries information confirming joining the receiving group; and determining, based on the second ultrasonic signal, the device sending the second ultrasonic signal as a device in the receiving group.
[0360] In a possible implementation, the apparatus 1800 is an electronic device or a part of an electronic device, for example, a chip or a chip system.
[0361] In the embodiment of the present application, the processor 1801 can be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiment of the present application. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the embodiment of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The software module can be located in the memory 1802, and the processor 1801 reads the program instructions in the memory 1802 and completes the steps of the above method in combination with its hardware.
[0362] In the embodiment of the present application, the memory 1802 may be a non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), or a volatile memory (volatile memory), such as RAM. The memory may also be any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory in the embodiment of the present application may also be a circuit or any other device that can implement a storage function, for storing instructions and / or data.
[0363] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0364] Based on the above embodiments, the present application further provides a computer storage medium, in which a computer program is stored. When the computer program is executed by a computer, the computer executes the audio optimization method provided in the above embodiments.
[0365] An embodiment of the present application also provides a computer program product, including instructions, which, when executed on a computer, enables the computer to execute the audio optimization method provided in the above embodiment.
[0366] like Figure 19 As shown, some other embodiments of the present application disclose a device. Figure 19As shown, the apparatus 1900 includes: one or more processors 1901 ; a memory 1902 ; and one or more computer programs 1903 (not shown in the figure), and the above components can be connected via one or more communication buses 1904 .
[0367] One or more computer programs are stored in the memory 1902, and the one or more computer programs include instructions; the processor 1901 calls the instructions stored in the memory 1902, so that the electronic device 1900 can perform the following steps: collect the audio to be processed; segment the audio to be processed to obtain at least one candidate audio, each candidate audio in the at least one candidate audio corresponds to one of the at least one sound type included in the target scene, and the target scene is the scene corresponding to the current environment; and send the at least one candidate audio to the first device.
[0368] In a possible implementation, the sound type includes a human voice type and / or an environmental sound type.
[0369] In a possible implementation, when the instruction is called and executed by the one or more processors, the electronic device performs the following steps before segmenting the audio to be processed: receiving the target scene notified by the first device.
[0370] In a possible implementation, when the instruction is called and executed by the one or more processors, the electronic device performs the following steps before segmenting the audio to be processed: determining the target scene.
[0371] Specifically, when the instruction is called and executed by the one or more processors, the electronic device determines the target scene and performs the following steps: determining the keywords of the current environment; and determining the scene corresponding to the keywords of the current environment based on the correspondence between the scene and the keyword.
[0372] Specifically, when the instruction is called and executed by the one or more processors, the electronic device determines the keywords of the current environment and performs the following steps: determining the keywords of the current environment in at least one of the following ways: performing semantic recognition on the collected audio to be processed to determine the keywords of the audio to be processed; performing image recognition on the collected image of the current environment to determine the keywords corresponding to the image; and determining the keywords corresponding to the geographic location information based on the geographic location information of the current environment.
[0373] In one possible implementation, when the instruction is called and executed by the one or more processors, the electronic device segments the audio to be processed and performs the following steps: if the at least one sound type includes a first environmental sound type, the audio to be processed is segmented to obtain audio without human voice, and the audio without human voice is segmented to obtain audio of the first environmental sound type; if the at least one sound type includes a human voice type, the audio to be processed is segmented to obtain audio of the human voice type.
[0374] In one possible implementation, when the instruction is called and executed by the one or more processors, the electronic device performs the following steps before segmenting the audio to be processed: receiving indication information sent by the first device, wherein the indication information is used to indicate the sound type corresponding to the sent candidate audio.
[0375] Specifically, when the instruction is called and executed by the one or more processors, the electronic device segments the audio to be processed to obtain at least one candidate audio, and performs the following steps: segmenting the audio to be processed to obtain candidate audio corresponding to the sound type indicated in the indication information.
[0376] In one possible implementation, when the instruction is called and executed by the one or more processors, the electronic device performs the following steps after segmenting the audio to be processed and obtaining at least one candidate audio: determining a quality score for each candidate audio of the at least one candidate audio; and sending the quality score of each candidate audio to the first device.
[0377] In one possible implementation, when the instruction is called and executed by the one or more processors, the electronic device performs the following steps before collecting the audio to be processed: receiving a third ultrasonic signal sent by the first device for establishing a sound receiving group; and sending a fourth ultrasonic signal, wherein the fourth ultrasonic signal carries information confirming joining the sound receiving group.
[0378] In a possible implementation, the apparatus 1900 is an electronic device or a part of an electronic device, for example, a chip or a chip system.
[0379] In the embodiment of the present application, the processor 1901 can be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiment of the present application. A general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the embodiment of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The software module can be located in the memory 1902, and the processor 1901 reads the program instructions in the memory 1902 and completes the steps of the above method in combination with its hardware.
[0380] In the embodiment of the present application, the memory 1902 may be a non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), or a volatile memory (volatile memory), such as RAM. The memory may also be any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory in the embodiment of the present application may also be a circuit or any other device that can implement a storage function, for storing instructions and / or data.
[0381] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0382] Based on the above embodiments, the present application further provides a computer storage medium, in which a computer program is stored. When the computer program is executed by a computer, the computer executes the audio optimization method provided in the above embodiments.
[0383] An embodiment of the present application also provides a computer program product, including instructions, which, when executed on a computer, enables the computer to execute the audio optimization method provided in the above embodiment.
[0384] The embodiments of the present application further provide a system, which may include apparatus 1800 or apparatus 1900. The system may also include an electronic device composed of apparatus 1800 or an electronic device composed of apparatus 1900. The system may also include an electronic device that executes the audio optimization method provided in the above embodiments.
[0385] like Figure 20As shown, some other embodiments of the present application disclose a device 2000, which includes: a receiving module 2001, used to receive candidate audio sent by devices in a sound receiving group, the sound receiving group includes at least one device, the candidate audio is audio of at least one sound type in a target scene, the target scene is a scene corresponding to the environment where the devices in the sound receiving group are located, and each of the sound types corresponds to at least one candidate audio; a processing module 2002, used to determine a candidate audio corresponding to each of the sound types from a candidate audio set, the candidate audio set includes candidate audio sent by the devices in the sound receiving group; an audio fusion module 2003, used to perform audio fusion on the determined candidate audio to obtain audio after audio fusion.
[0386] In a possible implementation, the sound type includes a human voice type and / or an environmental sound type.
[0387] In a possible implementation, the processing module 2002 is further configured to: determine the target scene.
[0388] In a possible implementation, the processing module 2002 is specifically configured to: obtain keywords of the environment; and determine the scene corresponding to the keyword of the environment according to a correspondence between the scene and the keyword.
[0389] In a possible implementation, the processing module 2002 is specifically configured to: obtain the keyword of the environment in at least one of the following ways: determining the keyword of the environment; and receiving the keyword sent by at least one device in the sound receiving group.
[0390] In a possible implementation, the processing module 2002 is specifically configured to: determine the keywords of the environment by using at least one of the following methods: performing semantic recognition on the collected audio to be processed to determine the keywords of the audio to be processed;
[0391] Perform image recognition on the collected image of the environment to determine keywords corresponding to the image; and determine keywords corresponding to the geographic location information based on the geographic location information of the environment.
[0392] In a possible implementation, the processing module 2002 is further configured to: before receiving the candidate audio sent by the device in the sound receiving group, notify the devices in the sound receiving group of the target scene.
[0393] In one possible implementation, the processing module 2002 is specifically used to: obtain quality scores of candidate audios in the candidate audio set; and select a candidate audio whose quality score meets preset conditions among the candidate audios corresponding to each of the sound types as a candidate audio corresponding to each of the sound types.
[0394] In one possible implementation, the processing module 2002 is specifically used to: obtain the quality score of the candidate audio in the candidate audio set using at least one of the following methods: determining the quality score of the candidate audio in the candidate audio set; receiving the quality score of the candidate audio sent by the device in the receiving group.
[0395] In a possible implementation, the preset condition is that the quality score is the highest, or that the quality score is the highest and is greater than a score threshold.
[0396] In a possible implementation, the candidate audio set further includes candidate audio obtained by segmenting the collected audio to be processed.
[0397] In one possible implementation, the processing module 2002 is further used to: after determining a candidate audio corresponding to each of the sound types, send indication information to at least one device in the sound receiving group, wherein the indication information is used to indicate the sound type corresponding to the candidate audio that the device needs to send subsequently.
[0398] In one possible embodiment, the processing module 2002 is also used to: send a first ultrasonic signal for establishing a receiving group; receive a second ultrasonic signal, which carries information confirming joining the receiving group; and determine, based on the second ultrasonic signal, the device that sends the second ultrasonic signal as a device in the receiving group.
[0399] like Figure 21 As shown, some other embodiments of the present application disclose a device 2100, which includes: an acquisition module 2101, used to acquire audio to be processed; a processing module 2102, used to segment the audio to be processed to obtain at least one candidate audio, each of the at least one candidate audio corresponds to a sound type of at least one sound type included in a target scene, and the target scene is a scene corresponding to the current environment; a sending module 2103, used to send the at least one candidate audio to a first device.
[0400] In a possible implementation, the sound type includes a human voice type and / or an environmental sound type.
[0401] In a possible implementation, the processing module 2102 is further configured to: receive the target scenario notified by the first device.
[0402] In a possible implementation, the processing module 2102 is further configured to: determine the target scene.
[0403] In a possible implementation, the processing module 2102 is specifically configured to: determine keywords of the current environment;
[0404] According to the correspondence between scenes and keywords, the scene corresponding to the keyword of the current environment is determined.
[0405] In one possible implementation, the processing module 2102 is specifically used to: determine the keywords of the current environment using at least one of the following methods: performing semantic recognition on the collected audio to be processed to determine the keywords of the audio to be processed; performing image recognition on the collected image of the current environment to determine the keywords corresponding to the image; and determining the keywords corresponding to the geographic location information based on the geographic location information of the current environment.
[0406] In one possible implementation, the processing module 2102 is specifically used to: if the at least one sound type includes a first environmental sound type, segment the audio to be processed to obtain audio without human voice, and segment the audio without human voice to obtain audio of the first environmental sound type; if the at least one sound type includes a human voice type, segment the audio to be processed to obtain audio of the human voice type.
[0407] In a possible implementation, the processing module 2102 is further configured to: receive indication information sent by the first device, where the indication information is used to indicate a sound type corresponding to the sent candidate audio.
[0408] In a possible implementation, the processing module 2102 is specifically configured to segment the audio to be processed to obtain candidate audio corresponding to the sound type indicated in the indication information.
[0409] In a possible implementation, the processing module 2102 is further configured to: determine a quality score of each candidate audio of the at least one candidate audio; and send the quality score of each candidate audio to the first device.
[0410] In a possible implementation, the processing module 2102 is further configured to: receive a third ultrasonic signal sent by the first device for establishing a sound receiving group; and send a fourth ultrasonic signal, wherein the fourth ultrasonic signal carries information confirming joining the sound receiving group.
[0411] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0412] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0413] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0414] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0415] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
Claims
1. An audio optimization method, characterized in that: Applied to a first device, the method includes: receiving candidate audio signals sent by devices in a sound receiving group, where the sound receiving group is controlled and established by the first device, the devices in the sound receiving group are within a preset range of the first device, and can receive a first ultrasonic signal sent by the first device for establishing the sound receiving group; The sound receiving group includes multiple devices, the candidate audio is audio of at least one sound type in a target scene, the target scene is a scene corresponding to the environment where the devices in the sound receiving group are located, and each sound type corresponds to at least one candidate audio; Determining, from a candidate audio set, a candidate audio corresponding to each of the sound types, the candidate audio set including candidate audio sent by devices in the sound receiving group; The step of determining a candidate audio corresponding to each of the sound types from the candidate audio set includes: Obtaining quality scores of candidate audios in the candidate audio set; Selecting a candidate audio whose quality score meets a preset condition from the candidate audios corresponding to each sound type as a candidate audio corresponding to each sound type; Performing audio fusion on the determined candidate audios to obtain audio fusion; The method further comprises: Instruction information is sent to at least one device in the sound receiving group, where the instruction information is used to indicate the sound type corresponding to the candidate audio that the device needs to send subsequently.
2. The method according to claim 1, wherein The sound type includes a human voice type and / or an environmental sound type.
3. The method according to claim 1 or 2, wherein: Before receiving the candidate audio sent by the device in the audio reception group, the method further includes: The target scene is determined.
4. The method according to claim 3, wherein The determining of the target scene includes: Obtaining keywords for the environment; According to the correspondence between scenes and keywords, the scene corresponding to the keyword of the environment is determined.
5. The method according to claim 4, wherein The obtaining of keywords of the environment includes: At least one of the following methods is used to obtain keywords for the environment: Identify keywords for the environment; Receive a keyword sent by at least one device in the sound receiving group.
6. The method according to claim 5, wherein The keywords for determining the environment include: The keywords of the environment are determined by at least one of the following methods: Performing semantic recognition on the collected audio to be processed to determine keywords of the audio to be processed; Performing image recognition on the collected image of the environment to determine keywords corresponding to the image; According to the geographical location information of the environment, a keyword corresponding to the geographical location information is determined.
7. The method according to any one of claims 3 to 6, wherein: After determining the target scene and before receiving candidate audio sent by a device in the sound receiving group, the method further includes: Notify devices in the sound receiving group of the target scene.
8. The method according to claim 1, wherein The obtaining of quality scores of candidate audios in the candidate audio set includes: The quality scores of the candidate audios in the candidate audio set are obtained by using at least one of the following methods: Determining quality scores of candidate audios in the candidate audio set; Receive quality scores of candidate audios sent by devices in the audio receiving group.
9. The method according to claim 1 or 8, wherein: The preset condition is that the quality score is the highest, or the quality score is the highest and the quality score is greater than a score threshold.
10. The method according to any one of claims 1 to 9, characterized in that The candidate audio set also includes candidate audio obtained by segmenting the collected audio to be processed.
11. The method according to any one of claims 1 to 10, characterized in that Before receiving the candidate audio sent by the device in the audio reception group, the method further includes: Sending a first ultrasonic signal for establishing a sound receiving group; receiving a second ultrasonic signal, the second ultrasonic signal carrying information confirming joining the sound receiving group; According to the second ultrasonic signal, a device sending the second ultrasonic signal is determined as a device in the sound receiving group.
12. An audio optimization method, characterized in that: Applied to the second device, the method includes: receiving a third ultrasonic signal sent by the first device for establishing a sound receiving group, wherein the second device is within a preset range of the first device; Sending a fourth ultrasonic signal, wherein the fourth ultrasonic signal carries information confirming joining the sound receiving group; Collect audio to be processed; Segmenting the audio to be processed to obtain at least one candidate audio, each of the at least one candidate audio corresponding to one of at least one sound type included in a target scene, where the target scene is a scene corresponding to a current environment; sending the at least one candidate audio to the first device; The method further comprises: receiving indication information, where the indication information is used to indicate a sound type corresponding to the candidate audio that the second device needs to send subsequently; The newly collected audio to be processed is segmented to obtain candidate audio corresponding to the sound type indicated in the indication information, and sent to the first device.
13. The method according to claim 12, wherein: The sound type includes a human voice type and / or an environmental sound type.
14. The method according to claim 12 or 13, wherein: Before segmenting the audio to be processed, the method further includes: Receive the target scenario notified by the first device.
15. The method according to claim 12 or 13, wherein: Before segmenting the audio to be processed, the method further includes: The target scene is determined.
16. The method according to claim 15, wherein The determining of the target scene includes: determining keywords of the current environment; According to the correspondence between scenes and keywords, the scene corresponding to the keyword of the current environment is determined.
17. The method according to claim 16, wherein The keyword for determining the current environment includes: The keyword of the current environment is determined by at least one of the following methods: Performing semantic recognition on the collected audio to be processed to determine keywords of the audio to be processed; Performing image recognition on the collected image of the current environment to determine keywords corresponding to the image; According to the geographical location information of the current environment, a keyword corresponding to the geographical location information is determined.
18. The method according to any one of claims 12 to 17, wherein: The step of segmenting the audio to be processed includes: If the at least one sound type includes a first environmental sound type, segmenting the audio to be processed to obtain audio without human voices, and segmenting the audio without human voices to obtain audio of the first environmental sound type; If the at least one sound type includes a human voice type, the audio to be processed is segmented to obtain audio of the human voice type.
19. The method according to any one of claims 12 to 18, wherein: After segmenting the audio to be processed to obtain at least one candidate audio, the method further includes: determining a quality score for each candidate audio of the at least one candidate audio; The quality score of each candidate audio is sent to the first device.
20. An electronic device, characterized in that: include: One or more processors coupled to a memory, wherein the memory stores one or more computer programs; the one or more computer programs include instructions that, when executed by the device, cause the device to perform the audio optimization method described in any one of claims 1-11.
21. An electronic device, characterized in that: include: One or more processors coupled to a memory, wherein the memory stores one or more computer programs; the one or more computer programs include instructions that, when executed by the device, cause the device to perform the audio optimization method described in any one of claims 12-19.
22. The device according to claim 20 or 21, characterized in that The device is an electronic device or a part of an electronic device.
23. An audio optimization system, characterized in that: Comprising the apparatus of claim 20 or 21.
24. A computer storage medium, characterized in that The method comprises computer instructions, which, when executed on an electronic device, enable the electronic device to execute the audio optimization method according to any one of claims 1 to 19.
25. A computer program product, characterized in that When the computer program product is run on a computer, the computer is enabled to perform the audio optimization method according to any one of claims 1 to 19.
Citation Information
Patent Citations
Audio processing method and device
CN106782625A
Generating an audio signal from multiple microphones based on uncorrelated noise detection
CN110121889A