Audio processing method, readable storage medium, program product and electronic device

By acquiring scene images during audio recording, identifying and matching the target spatial scene, and processing the audio using simulation and optimization conversion parameters, the problem of insufficient spatial sense in audio recording by electronic devices is solved, achieving a more realistic spatial audio effect and immersive experience.

CN120475317BActive Publication Date: 2026-04-17HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HONOR DEVICE CO LTD
Filing Date
2024-09-05
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

When recording audio, existing electronic devices suffer from insufficient spatial sense due to microphone limitations. The spatial audio effects in post-processing do not match the real scene, affecting the user's immersive experience.

Method used

By acquiring scene images during audio recording, the spatial scene is determined, and the target scene is matched from the preset spatial scene. The audio is then spatially rendered using simulated conversion parameters and optimized conversion parameters to generate spatial audio that is more in line with the real scene.

Benefits of technology

It improves the spatial effect of audio, allowing users to more realistically feel the propagation of audio in the recording scene and enhance the immersive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120475317B_ABST
    Figure CN120475317B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of terminals, in particular to an audio processing method, a readable storage medium, a program product and an electronic device. The audio processing method is applied to the electronic device. The electronic device can collect an image and audio of a current space scene through a camera and a microphone array respectively, simulate the current space scene based on the image, perform spatial rendering on the audio in a propagation mode of the audio in the simulated space scene, so that the spatial effect of the generated spatial audio matches the space when the audio is recorded, and the sense of reality of the spatial audio is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminal technology, and in particular to an audio processing method, a readable storage medium, a program product, and an electronic device. Background Technology

[0002] Currently, many electronic devices typically support mono or stereo recording. Stereo recording, for example, uses a multi-microphone array to capture the sound differences between the left and right channels, thereby simulating the width (the spatial range of sound in the left-right direction) and depth (the spatial range of sound in the front-back direction) of sound to some extent, adding a sense of space to the audio. However, due to the limitations of the microphones in electronic devices (e.g., limitations in the number, layout, sensitivity, signal processing, and noise reduction capabilities of microphones), the sense of space in the audio recorded by electronic devices may not be obvious.

[0003] To enhance the spatial sense of audio played by electronic devices, post-processing techniques such as reverb, echo, and 3D rendering can be applied to the audio before playback. However, the parameters used in these post-processing steps (e.g., adjusting reverb, echo, and 3D rendering effects) are typically set based on experience. Therefore, post-processed audio cannot replicate the spatial sense of the audio in its original, real-world recording environment. Summary of the Invention

[0004] This application provides an audio processing method, a readable storage medium, a program product, and an electronic device, so as to enable the electronic device to record spatial audio with spatial effects corresponding to the spatial scene at the time of audio recording.

[0005] In a first aspect, embodiments of this application provide an audio processing method applied in an electronic device, comprising: acquiring a first audio and a scene image of the spatial scene in which the first audio was recorded; determining a first spatial scene corresponding to the first audio based on the scene image; determining a second spatial scene matching the first spatial scene from a plurality of preset spatial scenes; and processing the first audio based on a first audio adjustment parameter corresponding to the second spatial scene from pre-stored audio adjustment parameters to obtain a second audio.

[0006] In some embodiments of this application, the first audio may also be referred to as the original audio, the second audio may also be referred to as the target spatial audio, and the second spatial scene may also be referred to as the target spatial scene. The electronic device stores multiple preset spatial scenes and audio adjustment parameters corresponding to the multiple preset spatial scenes. The electronic device can determine the first spatial scene when recording the first audio through a scene image, and determine the second spatial scene matching the first spatial scene from the multiple preset spatial scenes. Then, based on the first audio adjustment parameters corresponding to the second spatial scene, it performs spatial rendering on the first audio, thereby rendering the first audio into a second audio with spatial effects.

[0007] It's understandable that the first audio adjustment parameters correspond to the second spatial scene. Therefore, the spatial effect of the second audio is similar to its propagation effect in the second spatial scene (the location of each sound source in the second audio within the second spatial scene, and the echoes, reverberations, and other sound effects produced by each sound source). Since the second spatial scene matches the first spatial scene, the spatial effect of the second audio better reflects the propagation process of audio in the first spatial scene, providing listeners with a better immersive experience within the first spatial scene.

[0008] In one possible implementation of the first aspect described above, the audio adjustment parameters are used to convert audio recorded in the corresponding spatial scene into spatial audio.

[0009] In some embodiments of this application, the audio adjustment parameters are matched with a preset spatial scene. The preset spatial scene can be determined first, and then the propagation process of sound within the preset spatial scene can be simulated to obtain the corresponding audio adjustment parameters. The audio adjustment parameters can be obtained based on the principles of sound propagation (or sound propagation models) and optimization model training, as detailed in the description below.

[0010] In one possible implementation of the first aspect above, determining the first spatial scene corresponding to the first audio based on the scene image includes: determining the first spatial information of the first spatial scene and the first sound source object in the scene image corresponding to the sound in the first audio based on the scene image and the first audio; wherein, the first spatial information includes the spatial size of the first spatial scene, the first sound source object and the sound transmission object in the first spatial scene, the position of each first sound source object and each sound transmission object, and the shape and acoustic characteristics of each sound transmission object.

[0011] For example, in some embodiments of this application, an electronic device can construct a first spatial scene based on a scene image using a corresponding model. For instance, the first spatial scene includes first spatial information and the first sound source object of the first audio in the scene image. It is understood that, based on the scene image, only objects such as people, animals, buildings, and vehicles in the first spatial scene can be identified, but it is impossible to determine which object emitted the sound. Therefore, it is also necessary to determine the first sound source object based on the first audio. For example, the first audio can be processed using sound source separation technology to obtain the audio corresponding to each sound source. Then, the audio corresponding to each sound source can be processed using sound source localization technology to determine the approximate location of each sound source. By fusing the approximate location of each sound source with the constructed first spatial scene, the first sound source object in the first spatial scene can be determined. That is, the first spatial scene includes the position of each first sound source object, as well as the position, shape, and acoustic characteristics of the sound-transmitting object, in order to better determine the corresponding audio adjustment parameters based on the first spatial scene.

[0012] In one possible implementation of the first aspect above, determining the second spatial scene that matches the first spatial scene from multiple preset spatial scenes includes: determining the spatial scene in which the spatial information and sound source object in the multiple preset spatial scenes match the first spatial information and the first sound source object in the first spatial scene as the second spatial scene.

[0013] For example, in some embodiments of this application, the spatial information and sound source objects of the preset spatial scene are matched with the first spatial information and first sound source objects of the first space. For instance, the number of sound source objects in the preset spatial scene is the same as the number of sound source objects in the first spatial scene. The similarity between the position of the sound source object in the preset spatial scene and the position of the first sound source object in the first spatial scene exceeds a threshold. Furthermore, the similarity between the spatial size, position, shape, acoustic characteristics, etc., of the preset spatial scene and the first spatial information of the first spatial scene exceeds a threshold. In this way, audio adjustment parameters that better match the first spatial scene can be selected to process the first audio.

[0014] In one possible implementation of the first aspect above, the first audio adjustment parameter includes a first analog conversion parameter and a first optimized conversion parameter; the first analog conversion parameter is used to synthesize the sounds emitted by multiple sound source objects in the second spatial scene to obtain the synthesized sound of each sound at a first position in the second spatial scene; the first optimized conversion parameter is used to convert the synthesized sound at the first position into spatial audio corresponding to the second spatial scene.

[0015] For example, in some embodiments of this application, the first simulation conversion parameter may be, for instance, based on the spatial information of a second spatial scene matching the first spatial scene, simulating the changes in the propagation of the sound of each sound source object in the second spatial scene to a first position in the second spatial scene, and synthesizing the sounds emitted by each sound source object to obtain a synthesized sound. It is understood that the synthesized sound is obtained through simulation and may deviate from the propagation process of sound in real space. Therefore, the synthesized sound can be adjusted based on the first optimized conversion parameter to obtain the spatial audio of the second spatial scene. The first optimized conversion parameter is used to adjust the synthesized sound to a spatial audio that is closer to the real sample spatial audio, wherein the sample spatial audio is spatial audio recorded by professional spatial audio recording equipment in the real spatial scene corresponding to the second spatial scene.

[0016] In one possible implementation of the first aspect above, the above-mentioned processing of the first audio to obtain the second audio includes: performing synthesis processing on the first sounds emitted by each first sound source object in the first spatial scene based on the first analog conversion parameters to obtain the first analog recorded audio; and performing spatial rendering processing on the first analog recorded audio based on the first optimized conversion parameters to obtain the second audio.

[0017] For example, in some embodiments of this application, after determining a second spatial scene that matches the first spatial scene, the electronic device can perform spatial rendering processing on the first audio using the first analog conversion parameters and the first optimized conversion parameters corresponding to the second spatial scene. It can be understood that the spatial rendering processing can be a processing of the first audio based on the spatial effects produced by the propagation process of the sound emitted by the first sound source object corresponding to the first audio in the second spatial scene.

[0018] In one possible implementation of the first aspect above, the first audio also includes a second sound emitted by a sound source object outside the first spatial scene; and, performing spatial rendering processing on the first analog recorded audio based on the first optimized conversion parameters to obtain the second audio includes: performing first spatial rendering processing on the first analog recorded audio based on the first optimized conversion parameters to obtain the first spatial audio, and performing second spatial rendering processing on the second sound to obtain the second spatial audio; and synthesizing the first spatial audio and the second spatial audio to obtain the second audio.

[0019] For example, in some embodiments of this application, the first audio includes a first sound emitted by a first sound source object within a first spatial scene, and the first spatial audio can be obtained by performing first spatial rendering processing on the first sound using first analog conversion parameters and first optimized conversion parameters.

[0020] The first audio also includes sounds from sources outside the first spatial scene, meaning these sources are not captured in the scene image. Therefore, the approximate direction and / or location of the sound source can be determined using sound source localization technology, and the second sound can be processed using this direction and / or location for second spatial rendering. For example, the time it takes for the second sound to travel to the recording location is determined based on the sound source's location, along with the energy loss of the sound, to obtain the second spatial audio. In other words, the second spatial rendering process cannot be adjusted based on the first analog conversion parameters and the first optimized conversion parameters.

[0021] The second audio is formed by fusing the first and second spatial audio to avoid losing audio during the processing of the first audio, thus making the second audio richer.

[0022] In one possible implementation of the first aspect above, the acquisition of the first audio and the scene image of the spatial scene where the first audio is recorded includes: acquiring the first audio through the audio acquisition module of the electronic device, and acquiring the scene image of the spatial scene in the scene where the electronic device is located based on the image acquisition module of the electronic device; or, using the audio file in the first video file as the first audio, and using the image in the first video as the scene image; or, using the first audio file as the first audio, and using at least one first image file as the scene image.

[0023] For example, in some embodiments of this application, the electronic device can process the first audio during recording to obtain a second audio. For instance, the electronic device can acquire the first audio through an audio acquisition module (e.g., a microphone or microphone array) and capture scene images of the current spatial scene through an image acquisition module (e.g., a camera) in order to process the first audio.

[0024] In some embodiments, the electronic device may also perform spatial rendering processing on the first audio in the first video file. It is understood that the first video file includes the first audio and a scene image of the spatial scene corresponding to the first audio.

[0025] In some embodiments, the electronic device may further process the first audio and determine the first spatial scene based on a first image file of the spatial scene corresponding to the first audio.

[0026] Secondly, this application provides an electronic device comprising: a memory for storing instructions; and at least one processor for executing the instructions to cause the device to implement the methods provided in the first aspect and any possible implementation of the first aspect. The beneficial effects achievable in the second aspect can be referred to the beneficial effects of the methods provided in any embodiment of the first aspect, and will not be repeated here.

[0027] In one possible implementation of the second aspect above, the electronic device further includes: an image acquisition module for acquiring scene images; and an audio acquisition module for acquiring first audio.

[0028] Thirdly, this application provides a computer-readable storage medium storing instructions that, when executed by a device, cause a computer to implement the methods provided in the first aspect and any possible implementation of the first aspect. The beneficial effects achievable in this third aspect can be referenced to the beneficial effects of the methods provided in any embodiment of the first aspect, and will not be repeated here.

[0029] Fourthly, this application provides a computer program product that, when run on a device, enables the device to implement the methods provided in the first aspect and any possible implementation of the first aspect. The beneficial effects achievable in the fourth aspect can be found in the beneficial effects of the methods provided in any embodiment of the first aspect, and will not be repeated here. Attached Figure Description

[0030] Figure 1 A schematic diagram is shown of a method for recording video in a concert setting using electronic devices;

[0031] Figure 2A According to some embodiments of this application, a flowchart of an implementation of spatial processing of raw audio is shown;

[0032] Figure 2B According to some embodiments of this application, a system schematic diagram of an electronic device is shown;

[0033] Figure 2C According to some embodiments of this application, schematic diagrams of audio emitted by various sound source objects are shown;

[0034] Figure 3A According to some embodiments of this application, a flowchart of an implementation of recording spatial video with an electronic device is shown;

[0035] Figure 3B According to some embodiments of this application, a flowchart of an implementation of an electronic device recording spatial audio is shown;

[0036] Figure 4 According to some embodiments of this application, a schematic diagram of identifying a first spatial scene is shown;

[0037] Figure 5A According to some embodiments of this application, a schematic diagram of the sound propagation process in space is shown;

[0038] Figure 5BAccording to some embodiments of this application, a flowchart of an implementation of training and optimizing transformation parameters is shown;

[0039] Figure 6 A schematic diagram of the structure of an electronic device is shown according to an embodiment of this application. Detailed Implementation

[0040] The illustrative embodiments of this application include, but are not limited to, audio processing methods, readable storage media, program products, and electronic devices.

[0041] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0042] For ease of understanding, some of the terms and related technologies used in this application are explained below.

[0043] 1. Spatial audio: Spatial audio, also known as immersive audio, 3D audio, or 360-degree surround sound, uses specific encoding and processing techniques to enable audio signals to simulate the position, direction, and trajectory of sound in three-dimensional space.

[0044] 2. Microphone Directivity: Microphone directivity, also known as microphone polarity, refers to a microphone's ability to pick up sound from different directions. Based on its directivity, microphones can be classified as follows:

[0045] (1) Omnidirectional microphones have the same sensitivity to sound from all directions and can pick up sound from all directions evenly. They are suitable for scenarios that require wide sound pickup, such as musicals, plays, speeches, and broadcasts.

[0046] (2) Unidirectional microphone,

[0047] Cardioid microphones are highly sensitive to sound coming directly in front of them and effectively isolate sound from other directions, especially from behind. They are suitable for picking up specific vocals or instruments on stage or in a recording studio.

[0048] Supercardioid microphones have a narrower pickup area than cardioid microphones, making them more directional and able to more effectively block out ambient noise. They are suitable for performance venues that require isolation or for targeted instrument recording.

[0049] Bidirectional microphone: also known as a figure-eight microphone, it has the highest sensitivity to sound in the 0° and 180° directions and can pick up sound from these two directions equally, but it is not sensitive to sound sources on the left and right sides. It is suitable for recording environments that need to capture two sound sources, such as face-to-face interviews or instrumental duets.

[0050] 3. 3D Rendering: 3D audio rendering involves virtually placing sound sources in a 3D space, including the listener's position in all directions (front, back, left, right, up, and down). Through specific algorithms and technologies, these sound sources can simulate the sense of location and distance of sound in natural space.

[0051] 4. Acoustic characteristics of materials: This mainly involves their properties regarding the propagation, reflection, absorption, and scattering of sound waves. For example, the speed at which sound propagates in a material, as well as the material's absorption rate and reflectivity.

[0052] The technical solution of this application is described below with reference to the accompanying drawings.

[0053] As mentioned in the background, electronic devices can perform post-processing on recorded audio using appropriate equipment or software to enhance its spatial sense, such as adding reverb, echo, and 3D rendering. However, the adjustment parameters for audio post-processing (such as those for adjusting reverb, echo, and 3D rendering effects) are generally set based on experience and do not match the propagation process of audio in the real spatial scene (the scene at the time of audio recording), resulting in an unrealistic sense of spatiality in the audio.

[0054] The following describes a scenario where electronic devices record audio.

[0055] For example, Figure 1 This diagram illustrates a method of recording video in a concert setting using electronic devices.

[0056] It should be noted that this application does not limit the specific form of the electronic device. The electronic device can be a mobile phone, camera, laptop, tablet, desktop computer, large-screen device, wearable device (e.g., watch, smart glasses, helmet), augmented reality (AR) / virtual reality (VR) device, personal digital assistant (PDA), etc. The following uses a mobile phone as an example of electronic device 100.

[0057] Reference Figure 1The concert venue A100 includes a stage A110, walls A121 and A122 surrounding the stage A110. Speakers A111 and A112 are mounted on the stage A110, speaker A114 is mounted on wall A121, and speaker A113 is mounted on wall A122. Audio from the stage A110 can be played through speakers A111 and A112, while background music from the concert venue A100 can be played through speakers A113 and A114. User A130 records the concert on stage A110 using electronic device 100.

[0058] For example, taking the audio played by speaker A111 as an example, the audio emitted by speaker A111 at the same time can propagate in all directions. For instance, speaker A111 emits audio B01, audio B02, and audio B03 propagating in different directions at time T0. Audio B01 is the audio that reaches user A130 directly, audio B02 is the audio that first propagates towards wall A121 and then reflects back to user A130, and audio B03 is the audio that first propagates towards wall A122 and then reflects back to user A130. For example, user A130 hears audio B01 at time T1, audio B02 at time T2, and audio B03 at time T3. Since T1 < T2 < T3, audio B01, which user A130 hears first, can be considered the initial sound, and audio B02 and audio B03 are the echoes.

[0059] Understandably, after hearing the initial sound (audio B01), user A130 can hear echoes (e.g., audio B02, audio B03) reflected from various surfaces (e.g., walls A121 and A122) in concert scene A100. Furthermore, the propagation path of the reflected echoes is typically longer than that of the initial sound, and the intensity of the echoes attenuates along the propagation path. Additionally, user A130 can also receive reverberation noise (the phenomenon where sound waves continue to linger for a period of time after the sound source stops emitting sound due to multiple reflections and absorptions by obstacles such as walls during indoor propagation) mixed with the initial sound and subsequent echoes. Generally, user A130 can often determine the source of the sound from the initial sound. Subsequent echoes and / or reverberation often provide user A130 with additional information about concert scene A100. For example, echoes and / or reverberation can convey that sound propagates along many different paths within concert scene A100. Based on audio B01, user A130 can determine that speaker A111 is located to the northwest of user A130. Based on audio B02 and B02, user A130 can determine the approximate locations of walls A121 and A122 in concert scene A100.

[0060] For example, in some embodiments of this application, user A130 can record video of concert scene A100 through electronic device 100. For instance, on the recording interface 10 of electronic device 100, user A130 can select recording control 11 to record video of concert scene A100.

[0061] However, in scenarios where electronic devices are equipped with only a single microphone, the limited directionality and sensitivity of a single microphone typically prevent it from capturing all sounds in the space. Consequently, the video recorded by the electronic device only captures sound from certain directions. Therefore, when the electronic device plays the recorded video, the lack of sound from certain directions during audio recording prevents the user from perceiving the audio from all directions and thus from experiencing the effect of audio propagation in space.

[0062] For example, in some cases, electronic device 100 records audio C01 using a single microphone, and the recording software of electronic device 100 does not perform spatial audio rendering or other processing on audio C01. When electronic device 100 plays audio C01, the listener can hear the sound source as either electronic device 100 or a sound output device (such as headphones). The listener can hear the music and cheers of the audience in concert scene A100 based on audio C01, but cannot determine the location of other sound sources such as speaker A111 in concert scene A100 based on audio C01. This results in the user not being able to perceive the spatial environment of concert scene A100 through audio C01 (such as the approximate location of walls A121 and A122 in concert scene A100), and the user does not have an immersive experience.

[0063] To improve the spatial sense of audio recorded by electronic devices, some devices are equipped with multiple microphones (such as microphone arrays) to record sound from different directions, thereby enhancing the spatial information carried in the recorded audio. However, even with microphone arrays, the spatial sense of the recorded audio remains insufficient due to limitations in the number, layout, sensitivity, signal processing, and noise reduction capabilities of the microphones.

[0064] For example, in some cases, electronic device 100 can record audio CO2 using a multi-microphone array, and the recording software of electronic device 100 can perform spatial audio rendering and other processing on audio CO2. Alternatively, the debugging personnel can process audio CO2 in post-production using other software or devices, such as performing spatial audio rendering to add effects like echo and reverb, thereby giving audio CO2 a good sense of space. Therefore, when electronic device 100 plays audio CO2, the listener may experience a spatial effect to some extent. For example, the listener can feel the sound coming from all directions, enhancing the sense of auditory immersion and making audio CO2 more natural.

[0065] However, the spatial audio rendering processing of audio CO2 by electronic devices is based on experience or preset processing parameters, adjusting parameters such as echo and reverberation of audio CO2 to give audio CO2 a spatial effect. However, spatial audio rendering processing methods based on experience or preset processing parameters may not match the recording scenario of audio CO2 (such as concert scenario A100), resulting in the obtained spatial audio not being able to reproduce the actual recording scenario of audio CO2.

[0066] In conclusion, when electronic devices perform spatial audio rendering on recorded audio using empirical or preset parameters, the resulting spatial audio may not accurately reflect the recording context. This prevents users from experiencing the recording context through the audio played on the electronic device, thus negatively impacting the user experience.

[0067] It should be noted that for a spatial scene with known spatial information (such as the size of the space, the materials, locations, shapes, and acoustic characteristics of the facilities within the space, and the distribution of sound sources in the space), the sound produced by mixing the sounds emitted by each sound source at any location can be calculated based on the principles of sound propagation (such as the loss of sound intensity during propagation, and the reflection and absorption of sound by various surfaces).

[0068] Based on this, in some embodiments, the conversion relationship between the mixed sound emitted by various sound sources at a certain location in a spatial scene with known spatial information and the spatial audio at that location can be determined in the following way: First, sample spatial audio can be obtained by recording sample sounds played by various sound sources in the real spatial scene at a recording location in the real spatial scene using professional spatial audio recording equipment. This can reflect the spatial information in the real spatial scene (e.g., space size, materials, locations, shapes, and acoustic characteristics of facilities within the space, and the distribution of sound sources in the space).

[0069] Then, a simulated spatial scene corresponding to the real spatial scene can be established, where the positions of the simulated sound sources relative to the simulated recording position in the simulated space are the same as the positions of each sound source relative to the recording position in the real spatial scene. Based on the established simulated spatial scene and a sound propagation model (e.g., a model simulating the sound propagation process), the electronic device can obtain the sample sound from each simulated sound source propagating to the simulated recording position, and based on the fusion of the simulated sound source audios, obtain the simulated recorded audio of the sample sound, and record the analog conversion parameters from sample sound to simulated recorded audio.

[0070] Secondly, electronic devices can use sample space audio and simulated recorded audio to train the optimization model, so that the optimization model can convert simulated recorded audio into sample space audio and record the optimization conversion parameters from simulated recorded audio to sample space audio.

[0071] Based on the aforementioned analog and optimized conversion parameters, if a user records stereo audio using an electronic device in the aforementioned real recording scenario or a scenario similar to the aforementioned real recording scenario (e.g., a spatial scenario with similar spatial information to the real recording scenario), the electronic device can first separate the stereo audio into sound components corresponding to each sound source, and then convert the sound components corresponding to each sound source into analog recorded audio based on the analog conversion parameters. Then, the electronic device converts the analog recorded audio into spatial audio based on the optimized conversion parameters. Since the spatial scenario where the stereo sound was recorded matches the real recording scenario corresponding to the optimized conversion parameters, the converted spatial audio can reflect the spatial effect of the recorded stereo sound propagating in the corresponding spatial scenario.

[0072] Based on the above, to improve the spatial effect of audio, this application proposes an audio processing method that can pre-store analog conversion parameters and optimized conversion parameters corresponding to multiple preset spatial scenes. An electronic device can identify the first spatial scene in which the electronic device acquired the original audio (as an example of the first audio) by using an image of the spatial scene in which it acquired the original audio (as an example of a scene image). Then, the electronic device can acquire a target spatial scene (as an example of a second spatial scene) that matches the first spatial scene from among the multiple preset spatial scenes, and perform spatial audio processing on the original audio acquired by the electronic device based on the audio adjustment parameters corresponding to the target spatial scene (e.g., the analog conversion parameters and optimized conversion parameters corresponding to the target spatial scene) to obtain the target spatial audio.

[0073] Through the above method, the electronic device can adjust the original audio to target spatial audio based on the target spatial scene corresponding to the first spatial scene, thereby providing corresponding spatial information to the original audio and improving the realism and layering of the original audio. Furthermore, the spatial sense of this target spatial audio better matches the first spatial scene in which the electronic device recorded the original audio.

[0074] In some embodiments, the analog conversion parameters can be determined based on the principles of sound propagation. For example, the intensity loss of sound during propagation can be calculated based on the distance between the sound source and the receiving location; the propagation speed of sound in various directions can be determined based on the propagation medium in space; and the reflectivity and absorption rate of the materials of the facilities in space can be determined, thereby calculating parameters such as the intensity, echo, and reverberation of the sound upon reaching the receiving location. It can be understood that the analog conversion parameters can be parameters used to calculate the state change of the sound emitted by the sound source to the state of the sound received at the receiving location. Exemplarily, in some embodiments, the analog conversion parameters can be determined using models that simulate sound propagation; these models will be described in detail below and will not be elaborated upon here.

[0075] The optimized transformation parameters can be obtained through model training. For example, simulated recorded audio and sample space audio can be used as training sample pairs to train the corresponding optimization model. This allows the trained optimization model to adjust the simulated audio to more closely resemble the sample space audio propagating in real space. Specifically, the simulated recorded audio samples can be input into the optimization model. The transformation parameters in the optimization model process the simulated recorded audio and output the predicted spatial audio. Then, a loss function is calculated based on the difference between the predicted spatial audio and the sample space audio. The transformation parameters of the optimization model are then adjusted based on the loss function. When the loss function converges, or the optimization model achieves the expected result, or the number of training iterations reaches the preset number, training ends, and the transformation parameters in the optimization model are output. These transformation parameters are the optimized transformation parameters. Details will be described below and will not be elaborated upon here.

[0076] In some embodiments of this application, the preset spatial scene refers to a spatial scene with different spatial information. The preset spatial scene and the corresponding audio adjustment parameters can be stored in an electronic device, or in a server or other device. The electronic device can directly obtain the adjustment parameters corresponding to the target spatial scene from the preset spatial scene, or obtain the audio adjustment parameters corresponding to the target spatial scene from the server or other devices.

[0077] In some embodiments of this application, obtaining a target spatial scene that matches the first spatial scene involves, for example, obtaining a target spatial scene with similar spatial information from a preset spatial scene based on the spatial information of the first spatial scene. For example, obtaining a spatial scene from a preset spatial scene that is similar to the first spatial scene in terms of spatial size, material, location, shape, acoustic characteristics of facilities within the space, and distribution of sound sources in the space as the target spatial scene.

[0078] The following describes the spatial processing of the original audio in the embodiments of this application.

[0079] For example, Figure 2A According to some embodiments of this application, a flowchart of an implementation of spatial processing of raw audio is shown.

[0080] It should be noted that this application does not limit the specific form of the electronic device. The electronic device can be a mobile phone, laptop, tablet, large-screen device, wearable device (e.g., watch, smart glasses, helmet), desktop computer, AR / VR device, PDA, etc. It is understood that the executing entity for each of the following processes is an electronic device, and no limitation is made on the executing entity when describing each process.

[0081] like Figure 2A As shown, the process includes:

[0082] S201, by acquiring an image of the spatial scene where the original audio was acquired, identifies the first spatial scene in which the electronic device acquired the original audio.

[0083] For example, in some embodiments of this application, the electronic device also captures images of the spatial scene where the original audio is located during the recording process. For instance, in a concert scene, a user can record video of the concert scene using the electronic device, thereby capturing the original audio of the concert scene as well as images of the spatial scene of the concert.

[0084] In other embodiments, the electronic device may also obtain the original audio and the spatial scene in which the original audio was acquired through other means. For example, the electronic device may acquire video from another device, the video including the original audio and an image of the spatial scene corresponding to the original audio. Alternatively, the electronic device may also obtain the original audio and an image of the spatial scene of the original audio.

[0085] In some embodiments of this application, the original audio can be a recording file or audio from a video file. The image of the spatial scene in which the original audio is recorded can be an image from a video file, an image of the spatial scene corresponding to the original audio captured before video recording, or an image of the spatial scene corresponding to the original audio captured before recording the original audio.

[0086] After acquiring the original audio and an image of the spatial scene where the original audio was recorded, the electronic device can identify the first spatial scene corresponding to the original audio.

[0087] For example, Figure 2B According to some embodiments of this application, a system schematic diagram of an electronic device is shown.

[0088] like Figure 2B The electronic device system shown includes a large-scale image analysis model, a sound source separation module, a sound source localization module, an acoustic generation module, a spatial matching module, a simulation module, and an optimization module.

[0089] In some embodiments of this application, the large-scale image analysis model may be, for example, a neural radiation fields for talking head synthesis (NeRF) model, used to output spatial information of a first spatial scene based on the input image.

[0090] For example, in some embodiments of this application, the first spatial scene includes corresponding spatial information. Figure 1 Taking a concert scene A100 as an example, the first spatial scene includes a stage A110, walls A121 and A122 surrounding the stage A110. Speakers A111 and A112 are located on the stage A110. It is understood that in some cases, the photographer may only want to record images of the stage A110; therefore, in the first spatial scene, only partial images of walls A121 and A122 may be recorded. Furthermore, images of speakers A114 on wall A121 and A113 on wall A122 may not exist in the first spatial scene. That is, in some cases, the images acquired by the electronic device may only include a portion of the first spatial scene containing the original audio. In other embodiments, the photographer may also record all images of the spatial scene containing the original audio. The embodiments of this application do not limit the first spatial scene acquired by the electronic device.

[0091] Exemplary examples, in some embodiments of this application, the electronic device can input acquired images into a NeRF model to determine a first spatial scene and spatial information of the first spatial scene. For example, the input data to the NeRF model includes:

[0092] Three-dimensional spatial coordinates represent a location point in a first spatial scene within an image, typically denoted by (x, y, z). For example, these three-dimensional spatial coordinates are not obtained directly from the image, but rather by extracting feature points from the image and matching these feature points across different images. Common feature point extraction algorithms include scale-invariant feature transform (SIFT) and speedup robust features (SURF). Feature matching is the process of associating corresponding feature points across different images. Using the matched feature points, methods such as triangulation can be used to calculate the three-dimensional coordinates of an object. Triangulation is based on the correspondence of feature points from multiple viewpoints, using geometric calculations to determine the position of the feature points in three-dimensional space.

[0093] The line-of-sight direction represents the directional relationship between the line of sight and a point in the scene, usually denoted by (θ, φ), where θ is the angle between the line-of-sight direction and the y-axis, and φ is the angle between the line-of-sight direction and the x-axis. These two angles uniquely determine the direction of a ray. Think of it this way: when an electronic device captures an image, each pixel on the image plane corresponds to a ray that originates from the optical center of the camera (the center of the camera lens), passes through that pixel, and points to a point in the scene. The direction of this ray is the line-of-sight direction.

[0094] The input data for the NeRF model includes three-dimensional spatial coordinates (x, y, z) along a given line-of-sight direction (θ, φ). The output data can be a vector consisting of four quantities (r, g, b, σ).

[0095] Here, (r, g, b) represents the color values ​​of the three RGB channels, ranging from [0, 1]. This indicates the RGB color value emitted from the input position (x, y, z) when viewed along the direction (θ, φ).

[0096] σ is the density value, ranging from [0, +∞). σ represents the volumetric density at the input position (x, y, z), which can be understood as the opacity or the density of clouds in space.

[0097] It can be understood that the NeRF model reconstructs a 3D representation of a first-dimensional scene from images viewed from multiple perspectives. This 3D representation includes not only the geometry of the first-dimensional scene but also the color and texture information of the surfaces of objects within the scene. Within the NeRF framework, the 3D scene is represented as a continuous radiation field, where each point (in 3D space) has a corresponding density and color value.

[0098] After the 3D representation is reconstructed, the density field can be analyzed to determine the object's location and shape. Regions with higher density values ​​typically correspond to the object's surface because light is more easily absorbed or scattered in these areas. By analyzing the distribution of the density field, the object's location and boundaries can be roughly determined.

[0099] In some embodiments, to more accurately determine the position and shape of an object, it is often necessary to extract the object's surface from the density field. This can be achieved by setting a density threshold or using more advanced algorithms (such as the marching cubes algorithm). After surface extraction, one or more three-dimensional meshes are obtained, representing the geometry of the object in the first spatial scene.

[0100] Once the geometry of the object is determined, the distance from any point on the object to the shooting point (i.e., the shooting position of the electronic device) can be calculated.

[0101] It is understandable that the NeRF model can simulate spatial information in the first spatial scene, such as the size of the space, and the materials, locations, and shapes of the facilities within it. Based on the materials of the facilities, the acoustic properties of those materials can be determined. In other words, based on images of the concert scene A100 captured by electronic devices, the NeRF model can determine that the first spatial scene includes the stage A110, walls A121 and A122 surrounding the stage A110, and speakers A111 and A112 on the stage A110. It can also determine the acoustic characteristics of the stage A110, walls A121 and A122, and speakers A111 and A112.

[0102] In some embodiments of this application, the sound source objects in the first spatial scene are also matched with the original audio. For example, in some embodiments of this application, after acquiring the original audio data, the electronic device can process the original audio based on sound source separation technology to obtain the audio emitted by each sound source in the first spatial scene.

[0103] For example, the raw audio can be input into a sound source separation module. This module uses sound source separation technology to separate and extract different sound sources from the raw audio, allowing the user to hear only the sound sources of interest. Sound source separation technology analyzes the characteristics of different sound sources in the sound signal (such as arrival time, frequency, intensity, etc.) and uses algorithms to separate these sound sources.

[0104] For example, sound source separation techniques can include blind source separation methods: separating different source signals by processing the mixed signal without knowing the source signals. Common methods include independent component analysis and principal component analysis.

[0105] Deep learning-based methods utilize deep neural networks to model and analyze mixed audio signals, thereby separating sound sources. The processing power of deep learning enables it to effectively separate different sound sources, and even process highly correlated audio signals.

[0106] For example, taking the concert venue A100 as an example, in which, Figure 2C According to some embodiments of this application, schematic diagrams of audio emitted by various sound source objects are shown.

[0107] Reference Figure 2C By using audio source separation technology, the audio frequencies of speakers A111, A112, A113, and A114 can be determined.

[0108] After the sound source is isolated, it can be located to determine its position.

[0109] For example, in some embodiments of this application, an electronic device may be configured with multiple microphones, which are arranged in an array according to a certain geometric shape (such as linear, circular, spherical, etc.). Each microphone in the microphone array will collect sound signals in a first spatial scene. These signals include direct sound signals from various sound source objects, interference signals, and environmental noise, etc.

[0110] After obtaining the audio from each sound source, the corresponding audio from each sound source can be input into the sound source localization module. For example, the sound source localization module can preprocess the audio from each sound source to remove noise and improve signal quality. Preprocessing steps may include filtering, noise reduction, gain control, etc. After preprocessing the audio from each sound source, feature parameters for sound source localization can be extracted from the preprocessed audio. In some embodiments, the feature parameters for sound source localization may include time difference (the time difference between the arrival of the same sound source at different microphones), angle of arrival, intensity difference (the intensity difference between the arrival of the same sound source at different microphones), etc. These parameters can reflect the propagation characteristics of the sound signal in space and the relative positional relationship between the sound source and the microphone. After extracting the feature parameters for sound source localization, the feature parameters can be processed based on a localization algorithm. The localization algorithm determines the position of the sound source in three-dimensional space based on these feature parameters and the geometric information of the microphone array. For example, commonly used localization algorithms include methods based on maximum likelihood estimation, methods based on beamforming, and methods based on subspace partitioning.

[0111] For example, refer to Figure 1Taking the concert scene A100 as an example, the sound source localization technology can determine that the audio of speaker A111 (as an example of a sound source) comes from the front left, the audio of speaker A112 comes from the front right, the audio of speaker A113 comes from the rear right, and the audio of speaker A114 comes from the rear left.

[0112] After determining the location of each sound source, the location of each sound source object in the simulated first spatial scene can be determined.

[0113] For example, the spatial information of the first spatial scene, along with the audio and location information of each sound source, is input into the acoustic generation module. The acoustic generation module can then fuse the locations of the sound sources with the spatial information of the first spatial scene to simulate the first spatial scene. This first spatial scene includes the locations and audio of each sound source object. In other words, the acoustic generation module can fuse the locations of the sound sources with the spatial information of the simulated first spatial scene to determine the individual sound source objects within the simulated scene. For instance, if the first spatial scene is a concert scene A100, then based on the location (front left) of the audio from speaker A111, the audio from speaker A111 can be matched with the audio from speaker A111 in the first spatial scene. Similarly, based on the location (front right) of the audio from speaker A112, the audio from speaker A112 can be matched with the audio from speaker A112.

[0114] It is understood that in some embodiments of this application, the original audio may also include audio from other sound sources, but these sound sources are not present in the simulated first spatial scene. For example, the original audio may include audio from speaker A113 and audio from speaker A114. However, speakers A113 and A114 are located behind the photographer, and the photographer did not capture images of speakers A113 and A114 when filming stage A110. In other words, no object in the simulated first spatial scene matches the audio from speakers A113 and A114. Therefore, in subsequent processing, the audio from speakers A113 and A114 can be adjusted based on the positions of speakers A113 and A114 determined by their audio.

[0115] S202, Obtain the target spatial scene that matches the first spatial scene from multiple preset spatial scenes.

[0116] For example, in some embodiments of this application, the electronic device stores multiple preset space scenarios.

[0117] Among them, the preset spatial scenes are obtained by simulation based on real spatial scenes, and each preset spatial scene includes corresponding audio adjustment parameters (including analog conversion parameters and optimized conversion parameters).

[0118] For example, refer to Figure 2B After simulating a first spatial scene, the electronic device can input the simulated first spatial scene into a spatial matching module. The spatial matching module can then obtain a target spatial scene from a preset set of spatial scenes that matches the simulated first spatial scene. The spatial parameters of the target spatial scene are the same as or similar to those of the first spatial scene. For example, the size of the target spatial scene, the materials, locations, shapes, and acoustic characteristics of the facilities within it are similar to those of the first spatial scene. Furthermore, the number and locations of sound source objects in the target spatial scene are also similar to those in the first spatial scene.

[0119] For example, in some embodiments of this application, the target space scene may be similar to the concert scene A100.

[0120] S203, based on the audio adjustment parameters corresponding to the target spatial scene, perform spatial audio processing on the original audio to obtain the target spatial audio.

[0121] For example, in some embodiments of this application, after the electronic device determines the target spatial scene, it can adjust the original audio based on the audio adjustment parameters corresponding to the target spatial scene so that the original audio has a corresponding sense of space.

[0122] For example, the target spatial scene and the original audio (audio from each sound source object) can be input into the simulation module. The simulation module can adjust the original audio based on the analog conversion parameters corresponding to the target spatial scene, thereby simulating the propagation process of the original audio in the target spatial scene and obtaining the simulated audio. Then, the optimization module can optimize and adjust the simulated audio based on the optimization conversion parameters corresponding to the target spatial scene to obtain the target spatial audio.

[0123] It is understood that, in some embodiments, the simulation module and the optimization module can also be configured as an adjustment module, which includes simulation conversion parameters and optimization conversion parameters corresponding to the target space scene. The adjustment module can directly adjust the input raw audio to target space audio.

[0124] It is understandable that the target spatial scene corresponding to the audio adjustment parameters used to adjust the original audio is similar to the first spatial scene when the original audio data was recorded. Therefore, adjusting the original audio based on the audio adjustment parameters of the target spatial scene results in spatial audio that better matches the spatial effect corresponding to the first spatial scene.

[0125] In other embodiments, for audio emitted by a sound source not within the simulated first spatial scene, the audio emitted by that sound source can be adjusted based on the sound source location determined from the corresponding audio source, so that the audio has a sense of space. For example, refer to... Figure 1 In the simulated first spatial scene, speakers A113 and A114 are not included. Therefore, when adjusting the audio emitted by speakers A113 and A114, the audio of speakers A113 and A114 is rendered in 3D based solely on the positions determined by their audio. Then, the audio of speakers A113 and A114 is blended with other audio in the original audio to obtain complete spatial audio.

[0126] Using the above method, the first spatial scene of the original audio can be determined based on the image information of the scene. Then, the original audio is adjusted according to the audio adjustment parameters corresponding to the target spatial scene that is similar to the first spatial scene to obtain the corresponding spatial audio, so that the spatial effect of the spatial audio matches the first spatial scene corresponding to the original audio.

[0127] The process of recording spatial video and spatial audio in this application is described below.

[0128] For example, Figure 3A According to some embodiments of this application, a flowchart of an implementation of an electronic device recording spatial video is shown.

[0129] It should be noted that the executing entities for each of the following processes are all electronic devices, and no restrictions are placed on the executing entities for each process when describing them.

[0130] like Figure 3A As shown, the process includes:

[0131] S301, spatial video recording operation detected.

[0132] For example, in some embodiments of this application, a user can record video using an electronic device. For instance, a user can use the camera of an electronic device to film a concert scene A100 and record audio of the concert scene A100 using a multi-microphone array.

[0133] S302 uses a multi-microphone array to pick up audio data.

[0134] For example, in some embodiments of this application, when a user is recording audio and video using an electronic device, they can use a multi-microphone array to pick up sound and record audio data, thereby obtaining the original audio. For example, in concert scenario A100, the original audio includes audio data corresponding to speakers A111, A112, A113, and A114.

[0135] S303, acquires image data recorded by the camera.

[0136] For example, in some embodiments of this application, the electronic device can record video using a camera, and then the electronic device can obtain corresponding image data from the recorded video. For instance, taking concert scene A100 as an example, after the electronic device records the video of concert scene A100, the electronic device can obtain image data of different perspectives corresponding to concert scene A100 from the corresponding video.

[0137] It is understood that in some embodiments of this application, the processes of the electronic device acquiring image data and acquiring audio data do not have a specific order. That is, the processes S302 and S303 can be executed synchronously, or S303 can be executed first, followed by S302. The embodiments of this application do not limit the execution order of the processes S303 and S302.

[0138] S304 analyzes image data based on a large image analysis model.

[0139] For example, in some embodiments of this application, after an electronic device acquires image data recorded by a camera, it can analyze the image data using a corresponding large-scale image analysis model. For instance, in some embodiments of this application, the large-scale image analysis model may be a NeRF model.

[0140] S305, acquire object information and environmental information from image data, where object information includes the number of objects, spatial orientation, and distance, and environmental information includes the geometric environment of the spatial scene and the acoustic properties of the materials.

[0141] For example, in some embodiments of this application, image data acquired by an electronic device can be analyzed according to a large image analysis model (e.g., a NeRF model) to determine a first spatial scene based on the image data. It is understood that the first spatial scene includes the spatial orientation, distance, and environmental information of various objects. The objects are people, animals, plants, buildings, etc., within the first spatial scene. Environmental information may include, for example, the size of the first spatial scene, the shape of the materials of various facilities within the first spatial scene, and their acoustic characteristics.

[0142] For example, the method of obtaining object information and environmental information by analyzing image data through the NeRF model can be referred to the process of S201 in the embodiment of Figure 2.

[0143] S306. Based on object information, environmental information and audio data, determine the sound source object, and perform coordinate transformation of the sound source object to obtain a sound field centered on the sound source object.

[0144] For example, in some embodiments of this application, after obtaining object information and environmental information in the first spatial scene, the electronic device can determine the sound source object in the first spatial scene based on the audio data and perform coordinate transformation of the sound source object.

[0145] For example, an electronic device can separate the audio from each sound source in the audio data and determine the location information of each sound source. Then, it fuses the location information of each sound source with the object information analyzed by the electronic device based on a large image analysis model to determine the sound source object emitting the audio, and performs coordinate transformation on the sound source object to obtain a sound field centered on the sound source object. The process of determining each sound source object can be exemplarily described in the flow of S201 in the embodiment of Figure 2.

[0146] In some embodiments of this application, the coordinate transformation of the sound source object is, for example, to determine the position of the photographer relative to the sound source object with the sound source object as the origin of the coordinate system, so as to determine the propagation path of the sound emitted by the sound source object to the photographer.

[0147] S307, establish a simulated space environment based on environmental information, and determine the target space environment from the preset space environment based on the simulated space environment.

[0148] For example, in some embodiments of this application, after the electronic device obtains environmental information for recording audio and video based on image data, it can establish a simulated spatial environment based on the environmental information. The simulated spatial environment is a spatial environment simulated based on the environmental information of the real spatial environment when the electronic device records the video. It can be understood that the simulated spatial environment is the same as or similar to the real spatial environment when the electronic device records the video in terms of size, acoustic characteristics of facilities within the space, shape, etc.

[0149] It is understood that the simulated spatial environment can be the simulated first spatial scene in the embodiment of Figure 2. The electronic device can determine a target spatial environment similar to the simulated spatial environment from a preset spatial environment based on the simulated spatial environment. For example, the process of determining the target spatial environment can refer to the flow of S202 in the embodiment of Figure 2.

[0150] S308 adjusts audio data to obtain spatial audio based on audio adjustment parameters of the target spatial environment.

[0151] For example, in some embodiments of this application, after the electronic device determines the target spatial scene, it can adjust the original audio based on the audio adjustment parameters corresponding to the target spatial scene to obtain spatial audio. For example, the process of obtaining spatial audio can be referred to the flow of S203 in the embodiment of Figure 2.

[0152] For example, in some embodiments of this application, the processes of S306 to S308 can be processed by an acoustic generation module in an electronic device.

[0153] S309, record image.

[0154] For example, in some embodiments of this application, the electronic device can continue to record images via a camera to obtain video data.

[0155] S310 obtains spatial video based on recorded images and target spatial audio.

[0156] For example, in some embodiments of this application, the electronic device adjusts the recorded audio data into spatial audio and then fuses the spatial audio with image data to obtain spatial video. Here, spatial video is video with spatial audio.

[0157] Figure 3B According to some embodiments of this application, a flowchart illustrating an implementation of an electronic device recording spatial audio is shown.

[0158] It should be noted that the executing entities for each of the following processes are all electronic devices, and no restrictions are placed on the executing entities for each process when describing them.

[0159] like Figure 3B As shown, the process includes:

[0160] S31, Spatial audio recording operation detected.

[0161] For example, in some embodiments of this application, a user can directly record spatial audio using an electronic device. Upon detecting the spatial audio recording operation, the electronic device can activate a multi-microphone array for sound pickup and a camera to record images of the current environment.

[0162] S32 acquires audio data based on a multi-microphone array for sound pickup.

[0163] S33, acquire image data recorded by the camera.

[0164] S34 analyzes image data based on a large image analysis model.

[0165] S35, acquire object information and environmental information from image data, where object information includes the number of objects, spatial orientation, and distance, and environmental information includes the geometric environment of the spatial scene and the acoustic properties of the materials.

[0166] S36. Based on object information, environmental information and audio data, determine the sound source object, and perform coordinate transformation of the sound source object to obtain a sound field centered on the sound source object.

[0167] S37, establish a simulated space environment based on environmental information, and determine the target space environment from the preset space environment based on the simulated space environment.

[0168] S38, based on the audio adjustment parameters of the target spatial environment, adjusts the audio data to obtain spatial audio.

[0169] The process of S32 to S38 can be referred to as an example. Figure 3A The process described in steps S302 to S308 is as follows. It can be understood that, compared to recording spatial video, recording spatial audio does not require recording images, and the recorded images are then fused with the spatial audio.

[0170] In summary, electronic devices can directly record spatial video and audio using cameras and microphone arrays. Therefore, there is no need for post-production processing such as 3D rendering of the recorded audio data. Furthermore, the spatial effects of the recorded audio more closely match the spatial scenes in the video, providing users with a more realistic sense of space.

[0171] In other embodiments, when recording spatial video, the electronic device can also store files corresponding to spatial audio, so that users can extract the spatial audio separately for appropriate processing, etc.

[0172] The process of identifying the first spatial scene is described below.

[0173] For example, Figure 4 According to some embodiments of this application, a schematic diagram of identifying a first spatial scene is shown.

[0174] For example, in some embodiments of this application, the process of determining a first spatial scene based on image data from the recording of the original audio is described using the NeRF model as an example.

[0175] like Figure 4 As shown, after acquiring the audio data corresponding to the original audio and the image data of the spatial scene when the original audio was recorded, the electronic device can process the audio data and the image data respectively.

[0176] For example, the process of processing an image involves inputting the image into a NeRF model to obtain the density of each spatial point in the image, thereby constructing the spatial environment of the first spatial scene in the image based on the density.

[0177] For example, the input data for the NeRF model includes the three-dimensional spatial coordinates (x, y, z) and the viewing direction (θ, φ) of the image. These three-dimensional spatial coordinates are obtained by extracting feature points from images of the first spatial scene from multiple angles, matching these feature points across different images, and then using appropriate algorithms to calculate the position of the feature points in three-dimensional space based on the correspondence between feature points from multiple viewpoints. In other words, the three-dimensional spatial coordinates (x, y, z) are obtained through image analysis. The viewing direction (θ, φ) represents the directional relationship between the viewing direction and a point in the scene. After inputting the three-dimensional spatial coordinates (x, y, z) and the viewing direction (θ, φ) into the NeRF model, the color (r, g, b) and density σ data of each spatial coordinate in the first spatial scene can be obtained. The position and shape of objects are determined by analyzing the density field. Regions with higher density values ​​usually correspond to object surfaces because light is more easily absorbed or scattered in these areas. By analyzing the distribution of the density field, the position and boundaries of objects can be roughly determined, thereby simulating the spatial information of the first spatial scene.

[0178] The processing of audio data is, for example:

[0179] The audio is processed using sound source separation technology to obtain the audio emitted by each sound source in the first spatial scene.

[0180] After determining the audio emitted by each sound source, the sound source can be processed based on sound source localization technology to obtain the corresponding position (X, Y) and orientation (ψ) of each sound source.

[0181] Then, the position (X, Y) and orientation (ψ) of each sound source can be fused with the spatial information of the simulated first spatial scene to determine the object (sound source object) corresponding to each sound source in the first spatial scene.

[0182] After identifying each sound source object in the first spatial scene, the sound source objects can be fused with their corresponding audio data to simulate the first spatial scene.

[0183] It is understood that the above method can simulate the first spatial scene based on the audio data corresponding to the original audio recorded by the electronic device and the image data of the first spatial scene at the time of recording the original audio. This allows for subsequent adjustments to the original audio based on the first spatial scene to obtain spatial audio.

[0184] The following describes the process of obtaining and optimizing simulation conversion parameters.

[0185] For example, Figure 5A According to some embodiments of this application, a schematic diagram of the sound propagation process in space is shown.

[0186] In some embodiments of this application, the process of sound propagation is described using an empty room 500 as an example. Figure 5A As shown, room 500 includes a sound source 501 and a recording device 502. The straight-line distance between the sound source 501 and the recording device 502 is d1. Sound N1 emitted by the sound source 501 travels in a straight line to the recording device 502, with a propagation distance of d1. Sound N2 first travels to point M on the wall 503, and then reflects off point M back to the recording device 502, with a propagation distance of d2 + d3. It can be understood that sound N1 and sound N2 are the sounds of sound N emitted by the sound source 501 at the same time, under different propagation paths.

[0187] In some embodiments of this application, the sound N can be, for example, 50 dB. The intensity of the sound N1 received by the recording device 502 can then be calculated using a corresponding formula, for example, to be 40.5 dB, and the sound propagation time can be, for example, 0.0088 s. That is, the sound N1 loses 9.5 dB over the distance d1 it travels.

[0188] Similarly, the intensity of the sound N2 received by the recording device 502 is, for example, 29.2 dB, and the propagation time is 0.0159 s. That is to say, the sound N loses 20.8 dB after passing through the propagation paths d2 and d3 and after being reflected by the wall 503.

[0189] In this context, sound N2 is the echo of sound N1, and the time difference between sound N1 and sound N2 received by recording device 502 is, for example, 0.0071s. By measuring the time it takes for sound N to reach recording device 502 under different propagation paths, and the intensity of sound N when it reaches recording device 502, analog audio recording can be simulated.

[0190] For example, the above-described sound propagation process is just an example. In the process of simulating sound propagation, it is also necessary to consider the propagation medium of the sound, the frequency of the sound, the scattering of materials and other parameters. The embodiments of this application do not limit the process of simulating sound propagation.

[0191] It can be understood that the above calculations can simulate the conversion process from the sound emitted by the sound source 501 to the sound received by the recording device 502, thereby determining the corresponding analog conversion parameters. By storing the analog conversion parameters and the corresponding spatial scene in the electronic device, the electronic device can adjust the corresponding original audio using the analog conversion parameters.

[0192] For example, in some embodiments, simulation conversion parameters can be determined using models that simulate sound propagation. One such model is the sound wave equation model, which is a mathematical model describing the propagation of sound waves in a medium. Based on the wave equation, it can simulate phenomena such as the propagation and reflection of sound waves in three-dimensional space.

[0193] A ray model is a simplified model of sound propagation that treats sound waves as a series of straight lines (sound ray) originating from a sound source. These sound rays undergo reflection, refraction, or diffraction when they encounter obstacles. Ray models are very useful for predicting the propagation path and distribution of sound in large spaces such as theaters and concert halls. By simulating the propagation path of sound ray, metrics such as sound pressure level and sound intelligibility at different locations can be evaluated.

[0194] Geometric acoustic models, an extension of ray sound models, consider the wave-like and diffraction effects of sound waves during propagation. By constructing a geometric model of a room and simulating phenomena such as reflection, refraction, and diffraction of sound waves within the room, geometric acoustic models calculate the sound field distribution at various points within the room.

[0195] Digital sound synthesis models simulate the generation and propagation of sound using computer programs. These models utilize mathematical models (such as wave equations and acoustic propagation equations) to describe the vibration and propagation characteristics of sound, and generate sound signals through computer algorithms. Digital sound synthesis models have wide applications in audio production and sound effects processing, and can generate various complex sound effects, such as echoes and reverberation.

[0196] Next, we will continue to introduce the process of determining the optimization conversion parameters.

[0197] For example, Figure 5B According to some embodiments of this application, a flowchart of an implementation of training and optimizing transformation parameters is shown.

[0198] For example, Figure 5B The execution entities of each process can be electronic devices such as servers and computers. When introducing each process, no limitation is made on the execution entity of each process.

[0199] like Figure 5B As shown, the process includes:

[0200] S51, input simulated recorded audio and sample space audio into the optimization model.

[0201] For example, in some embodiments of this application, simulated recorded audio and sample spatial audio can be used as training sample pairs input into the optimization model. Based on the optimization transformation parameters in the optimization model, the simulated recorded audio is adjusted to a target spatial audio that approximates the sample spatial audio. It can be understood that sample spatial audio is audio obtained by recording sample sounds played from various sound sources in a real spatial scene at recording locations in a real spatial scene using professional spatial audio recording equipment. Sample spatial audio can reflect the spatial information in a real spatial scene.

[0202] S52, based on the optimized transformation parameters, adjusts the analog recorded audio to obtain the predicted spatial audio.

[0203] For example, in some embodiments of this application, simulated recorded audio is adjusted by optimizing transformation parameters in an optimization model to obtain predicted spatial audio.

[0204] S53, determine the loss function based on the difference between the predicted spatial audio and the sample spatial audio.

[0205] For example, in some embodiments of this application, the predicted spatial audio obtained after optimizing the transformation parameters of the simulated recorded audio may differ from the sample spatial audio. Therefore, a corresponding loss function can be determined based on the difference between the predicted spatial audio and the sample spatial audio. In other words, the difference between the predicted spatial audio and the sample spatial audio can be determined using the loss function.

[0206] S54, determine whether the predicted spatial audio meets the preset conditions.

[0207] For example, in some embodiments of this application, the preset condition may be the convergence of the loss function, the achievement of the expected effect of the optimized model, or the achievement of a preset number of training iterations.

[0208] If the judgment result is yes, then execute S56 and output the optimized conversion parameters.

[0209] If the judgment result is negative, then execute S55 to update and optimize the conversion parameters.

[0210] S55, updated and optimized conversion parameters.

[0211] For example, in some embodiments of this application, if the preset conditions are not met, such as the loss function not converging, the optimized transformation parameters can be updated, and then S52 can be executed again to adjust the simulated recorded audio based on the optimized transformation parameters to obtain the predicted spatial audio, until the predicted spatial audio meets the preset conditions. In some embodiments, the optimized transformation parameters can be updated according to the loss function so that the predicted spatial audio can quickly meet the preset conditions.

[0212] S56 outputs optimized conversion parameters.

[0213] For example, in some embodiments of this application, if the predicted spatial audio meets preset conditions, it indicates that the predicted spatial audio is closer to the sample spatial audio. That is, if the current optimized transformation parameters can adjust the simulated recorded audio to be closer to the sample spatial audio, then the optimized transformation parameters can be output.

[0214] It is understandable that the optimized transformation parameters can be obtained through the training process of the optimization model described above. These optimized transformation parameters are then matched with the spatial scene corresponding to the sample audio, and the optimized transformation parameters and the corresponding spatial scene are configured in the electronic device. When the first spatial scene corresponding to the original audio recorded by the electronic device matches the spatial scene corresponding to the sample space, the original audio can be adjusted based on these optimized transformation parameters to obtain the corresponding spatial audio.

[0215] The following section uses a mobile phone as an example to provide a detailed description of the electronic devices involved in some embodiments of the present invention.

[0216] Figure 6 A schematic diagram of the structure of an electronic device is shown according to an embodiment of this application.

[0217] Electronic device 100 may include processor 110, external memory interface 120, internal memory 121, universal serial bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, button 190, motor 191, indicator 192, camera 193, display screen 194, and subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0218] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0219] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors.

[0220] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0221] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0222] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0223] The charging management module 140 is used to receive charging input from the charger.

[0224] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.

[0225] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0226] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization.

[0227] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.

[0228] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through an audio device (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.

[0229] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0230] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a glass cover 10 and a display panel 20. The display panel 20 can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini-LED, a Micro-LED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1.

[0231] Camera 193 is used to capture still images or videos. In some embodiments of this application, when electronic device 100 records raw audio, camera 193 can be turned on to record the spatial scene corresponding to the raw audio for subsequent processing of the raw audio. For example, a user can record a portion of the current scene using camera 193 of electronic device 100, for example, referring to... Figure 1In concert scene A100, the user can record the scene on stage A110 using the camera 193 of the electronic device 100. In other embodiments, to enhance the spatial sense of the spatial audio generated by the electronic device 100, the user can also record the spatial scene from various angles using the camera 193 of the electronic device 100.

[0232] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.

[0233] Internal memory 121 can be used to store computer executable program code, which includes instructions. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of electronic device 100 by running instructions stored in internal memory 121 and / or instructions stored in memory located in the processor.

[0234] Electronic device 100 can implement audio functions such as music playback and recording through an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, and an application processor. For example, in some embodiments of this application, electronic device 100 may include multiple microphones 170C, arranged in an array according to a certain geometric shape (e.g., linear, circular, spherical, etc.) to capture sound from different directions. This allows the audio module 170 of electronic device 100 to determine each sound source and its corresponding location based on the audio recorded by the microphones 170C. It is understood that in embodiments of this application, the audio module 170 of electronic device 100 can also be used to process the audio data recorded by the microphones 170C to analyze the audio corresponding to each sound source and the location of each sound source.

[0235] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to make contact with and separate from the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 simultaneously. The multiple cards can be of the same or different types. The SIM card interface 195 is also compatible with different types of SIM cards. The SIM card interface 195 is also compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to realize functions such as calls and data communication. In some embodiments, the electronic device 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.

[0236] This application also provides a program product that, when executed on an electronic device, enables the electronic device to implement the methods provided in the foregoing embodiments.

[0237] This application also provides a readable storage medium storing one or more programs, which, when executed by an electronic device, enable the electronic device to implement the methods provided in the foregoing embodiments.

[0238] The various embodiments of the mechanisms disclosed in this application can be implemented in hardware, software, firmware, or a combination of these implementation methods. Embodiments of this application can be implemented as computer programs or program code executable on a programmable system, the programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0239] Program code can be applied to input instructions to execute the functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, the processing system includes any system having a processor such as, for example, a digital signal processor, a microcontroller, an application-specific integrated circuit, or a microprocessor.

[0240] The program code can be implemented using a high-level procedural language or an object-oriented programming language to communicate with the processing system. Assembly language or machine language can also be used when needed. In fact, the mechanisms described in this application are not limited to any particular programming language. In either case, the language can be a compiled language or an interpreted language.

[0241] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried on or stored thereon on one or more transient or non-transitory machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, the instructions may be distributed via a network or through other computer-readable media. Therefore, machine-readable media can include any mechanism for storing or transmitting information in a machine-readable (e.g., computer-readable) form, including but not limited to floppy disks, optical disks, CD-ROMs, compact disc-read-only memory (CD-ROMs), magneto-optical disks, read-only memory (ROM), random-access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic cards or optical cards, flash memory, or tangible machine-readable storage for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in the form of electrical, optical, acoustic, or other forms of propagation signals. Therefore, machine-readable media includes any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a machine-readable (e.g., computer-readable) form.

[0242] In the accompanying drawings, some structural or methodological features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be necessary. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Furthermore, the inclusion of structural or methodological features in a particular figure does not imply that such features are required in all embodiments, and in some embodiments, these features may be omitted or may be combined with other features.

[0243] It should be noted that all units / modules mentioned in the device embodiments of this application are logical units / modules. Physically, a logical unit / module can be a physical unit / module, a part of a physical unit / module, or a combination of multiple physical units / modules. The physical implementation of these logical units / modules themselves is not the most important factor; the combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed in this application. Furthermore, to highlight the innovative aspects of this application, the above-described device embodiments of this application have not introduced units / modules that are not closely related to solving the technical problems proposed in this application. This does not mean that the above-described device embodiments do not contain other units / modules.

[0244] It should be noted that in the examples and description of this patent, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0245] Although this application has been illustrated and described with reference to certain preferred embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made thereto without departing from the scope of this application.

Claims

1. An audio processing method, applied in an electronic device, characterized in that, include: Acquire the first audio and the scene image of the spatial scene where the first audio was recorded; Based on the scene image, a first spatial scene corresponding to the first audio is determined. The first spatial scene includes: first spatial information of the first spatial scene and a first sound source object in the first spatial scene. A second spatial scene matching the first spatial scene is determined from multiple preset spatial scenes. The matching of the first spatial scene with the second spatial scene includes: the spatial information and the sound source object of the second spatial scene are respectively matched with the first spatial information and the first sound source object in the first spatial scene. Based on the first audio adjustment parameter corresponding to the second spatial scene from the pre-stored audio adjustment parameters, the first audio is processed to obtain the second audio, wherein the first audio adjustment parameter is determined based on the spatial information of the second spatial scene and the sound source object.

2. The method according to claim 1, characterized in that, The audio adjustment parameters are used to convert audio recorded in the corresponding spatial scene into spatial audio.

3. The method according to claim 1, characterized in that, The first spatial information includes the size of the first spatial scene, the sound transmission object in the first spatial scene, the relative positional relationship between the first sound source object and the sound transmission object, and the shape and acoustic characteristics of the sound transmission object.

4. The method according to claim 1, characterized in that, The first audio adjustment parameters include a first analog conversion parameter and a first optimized conversion parameter; The first analog conversion parameter is used to synthesize the sounds emitted by multiple sound source objects in the second spatial scene to obtain the synthesized sound of each sound at a first position in the second spatial scene; The first optimized conversion parameters are used to convert the synthesized sound at the first position into spatial audio corresponding to the second spatial scene.

5. The method according to claim 4, characterized in that, The process of processing the first audio to obtain the second audio includes: Based on the first analog conversion parameters, the first sounds emitted by each of the first sound source objects in the first spatial scene are synthesized to obtain the first analog recorded audio. Based on the first optimized conversion parameters, the first simulated recorded audio is spatially rendered to obtain the second audio.

6. The method according to claim 5, characterized in that, The first audio also includes a second sound emitted by a sound source object outside the first spatial scene; and the step of performing spatial rendering processing on the first simulated recorded audio based on the first optimized conversion parameters to obtain the second audio includes: Based on the first optimized conversion parameters, the first analog recorded audio is subjected to a first spatial rendering process to obtain a first spatial audio, and the second sound is subjected to a second spatial rendering process to obtain a second spatial audio; The first spatial audio and the second spatial audio are synthesized to obtain the second audio.

7. The method according to claim 1, characterized in that, The acquisition of the first audio and the scene image of the spatial scene where the first audio was recorded includes: The first audio is acquired through the audio acquisition module of the electronic device, and a scene image of the spatial scene in the scene where the electronic device is located is acquired through the image acquisition module of the electronic device. Alternatively, the audio file in the first video file can be used as the first audio, and the image in the first video can be used as the scene image; Alternatively, a first audio file may be used as the first audio, and at least one first image file may be used as the scene image.

8. An electronic device, characterized in that, include: Memory, used to store instructions; At least one processor is configured to execute the instructions to cause the electronic device to implement the method of any one of claims 1 to 7.

9. The electronic device according to claim 8, characterized in that, The electronic device also includes: An image acquisition module is used to acquire images of the scene; An audio acquisition module is used to acquire the first audio.

10. A computer-readable storage medium, characterized in that, The readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the method of any one of claims 1 to 7.

11. A computer program product, characterized in that, When the computer program product is run on the device, it causes the device to perform the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Video recording implementation method and device and electronic equipment

    CN112165590A

  • Voice processing method and device based on scene recognition, medium and system

    CN113129917A