Method for playing audio in cooperation with video playing and communication system
Patent Information
- Application Number
- CN202111552635.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-17
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2041-12-17
AI Technical Summary
视频播放设备在播放视频时虽然同步播放音频,但用户往往难以辨别该音频中不同发声对象的声音的方位
[0097]可以理解地,上述第四方面提供的电子设备、第五方面提供的通信系统、第六方面提供的计算机可读存储介质、第七方面提供的计算机程序产品、第八方面提供的芯片均用于执行本申请实施例所提供的方法。因此,其所能达到的有益效果可参考对应方法中的有益效果,此处不再赘述。
Smart Images

Figure CN116266874B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal technology, and in particular to a method and communication system for co-playing audio during video playback. Background Technology
[0002] Sounds produced in nature, such as human voices, thunder, and train sounds, are all stereo. When people hear these sounds, they can perceive not only the loudness, pitch, and timbre, but also the direction of the sound. Although video playback devices play audio simultaneously with the video, users often find it difficult to distinguish the location of different sounds within the audio. Consequently, during video playback, users cannot quickly (or even be able to) immerse themselves in the scene, cannot experience the feeling of being there, and have a limited user experience. Summary of the Invention
[0003] To address the aforementioned technical problems, this application provides a method and communication system for collaborative audio playback during video playback. The technical solution provided in this application, through the positions of multiple audio playback devices, the user's position, and the video being played, enables the multiple audio playback devices to simulate the locations of different sound-emitting objects in the video's live environment. This allows the user to perceive the locations of different sound-emitting objects during video playback, making them feel as if they are in the live environment of the video, increasing the user's immersion and playability, and improving the user experience.
[0004] In a first aspect, this application provides a method for coordinating audio playback during video playback. This method can be applied to an electronic device. The electronic device is capable of communicating with M audio playback devices. The electronic device may include a display screen. M is a positive integer greater than or equal to 2. The electronic device can acquire a first video. The first video contains a first image group and a first audio within a first time period. Based on the first image group and the first audio, the electronic device can determine that the first video contains a first sound source and a first background sound within the first time period, and separate a first audio component of the first sound source and a second audio component of the first background sound from the first audio. The electronic device can send a first message to the first audio playback device among the M audio playback devices. The first message contains the first audio component and a first playback parameter, and the first message is used to instruct the first audio playback device to play the first audio component with the first playback parameter. The electronic device can send a second message to the second audio playback device among the M audio playback devices. The second message contains the second audio component and a second playback parameter, and the second message is used to instruct the second audio playback device to play the second audio component with the second playback parameter. During the playback of the first video in the first time period, the electronic device can display images from the first image group.
[0005] In this way, during the playback of the first video, the first audio playback device simulates the effect of the sound source's voice in terms of sound location, matching the position of the first sound source in the three-dimensional space represented by the video. The observer can immerse themselves in the scene presented in the video and more realistically perceive the location of different sound sources. This increases the observer's sense of immersion and engagement with the video, enhancing their overall experience.
[0006] According to the first aspect, in some embodiments, based on the first image group and the first audio, the electronic device further determines that the first video contains a second sound-emitting object within a first time period, and separates a third audio component of the second sound-emitting object from the first audio. The electronic device may send a third message to a third audio playback device among M audio playback devices, the third message containing the third audio component and third playback parameters, the third message being used to instruct the third audio playback device to play the third audio component with the third playback parameters.
[0007] The first audio playback device and the second audio playback device mentioned above can be different audio playback devices.
[0008] In this way, when the first video contains multiple sound-producing objects, the electronic device can select an audio playback device to simulate the sound of different objects. The effect of each audio device simulating the sound's location matches the simulated sound object's position in the three-dimensional space represented by the video. This can increase the viewer's immersion and engagement with the video, enhancing their overall experience.
[0009] The first video may contain more sound-producing objects within the first time period, not limited to the first and second sound-producing objects mentioned above. Understandably, the methods for separating the audio components of other sound-producing objects, determining the audio playback device for playing the audio components of other sound-producing objects, and determining the playback time can all refer to the processing methods for the first and second sound-producing objects mentioned above. These will not be elaborated upon here.
[0010] It should be noted that the content included in the first time period of the aforementioned first video can be a portion of the first video's duration. For example, the first video is 1 minute long. The aforementioned first time period can be the period from 0 to 5 seconds of the first video. Alternatively, the content included in the first time period of the aforementioned first video can be the entire duration of the first video. For example, the first video is 1 minute long. The aforementioned first time period can be the entire period of the first video from beginning to end.
[0011] According to the first aspect, or any implementation of the first aspect above, in some embodiments, the first audio playback device may be obtained based on the position of all or part of the M audio playback devices in the first scene relative to the observer viewing the electronic device, the position of the first sound-emitting object in the second scene relative to the virtual camera of the first video, and determining the position of the virtual camera in the second scene as a position corresponding to the position of the observer in the first scene.
[0012] In this way, the electronic device can determine the position of the first sound-emitting object in the first scene and select an audio playback device that is close to the position of the first sound-emitting object in the first scene to simulate the sound emitted by the first sound-emitting object. Specifically, the closer an audio device is to a sound-emitting object, the easier it is for the observer to perceive the sound as emanating from that location by adjusting the playback parameters of the audio device to play the audio component of the sound-emitting object. If the positions of the observer, the audio playback device, or the first sound-emitting object change in the first scene, the electronic device can reselect an audio playback device to simulate the sound emitted by the first sound-emitting object.
[0013] According to the first aspect, or any implementation of the first aspect above, in some embodiments, the first playback parameter is obtained based on the position of the first audio playback device relative to the observer of the viewing electronic device in the first scene, the position of the first sound-emitting object relative to the virtual camera of the first video in the second scene, and determining the position of the virtual camera in the second scene as a position corresponding to the position of the observer in the first scene.
[0014] In this way, when the first audio playback device plays the first audio component with the first playback parameters, the observer perceives the sound as originating from the location of the first sound-emitting object within the first scene. This allows the observer to immerse themselves in the scene presented in the video and more realistically perceive the location of different sound-emitting objects. This increases the observer's sense of immersion and engagement with the video, enhancing their overall experience.
[0015] According to the first aspect, or any implementation of the first aspect above, in some embodiments, the third audio playback device is obtained based on the positions of all or part of the M audio playback devices in the first scene relative to the observer viewing the electronic device, the position of the second sound-emitting object in the second scene relative to the virtual camera of the first video, and determining the position of the virtual camera in the second scene as a position corresponding to the position of the observer in the first scene.
[0016] In this way, the electronic device can determine the position of the second sound-producing object in the first scene and select an audio playback device that is close to the position of the second sound-producing object in the first scene to simulate the sound of the second sound-producing object.
[0017] According to the first aspect, or any implementation of the first aspect above, in some embodiments, the third playback parameter is obtained based on the position of the third audio playback device relative to the observer of the viewing electronic device in the first scene, the position of the second sound-emitting object relative to the virtual camera of the first video in the second scene, and determining the position of the virtual camera in the second scene as a position corresponding to the position of the observer in the first scene.
[0018] In this way, the third audio playback device plays the third audio component with third playback parameters, making the sound perceived by the observer as originating from the location of the second sound object in the first scene. This can increase the observer's immersion and playability in watching the video, thus enhancing the observer's overall experience.
[0019] According to the first aspect, or any implementation of the first aspect above, in some embodiments, the position corresponding to the observer's position in the first scene may be the same as the observer's position in the first scene.
[0020] In this way, during the video playback, the observer can feel the process of the voice-speaking object in the video making a sound from different directions, following the perspective of the virtual camera.
[0021] According to the first aspect, or any implementation of the first aspect above, in some embodiments, the observer is located at a first position in the first scene, and the first audio playback device is located at a second position in the first scene. After determining the position of the virtual camera in the second scene as the position corresponding to the observer's position in the first scene, the electronic device can obtain the third position of the first sound-emitting object in the first scene based on the first position and the position of the first sound-emitting object relative to the virtual camera in the second scene. Wherein, with the first position as the vertex of the included angle, the first position, the second position, and the third position form a first included angle, which is the smallest among the included angles formed by the first position as the vertex and the positions of the first position, the third position, and any one of the M audio playback devices.
[0022] According to the first aspect, or any implementation of the first aspect above, in some embodiments, the second audio playback device is one of the M audio playback devices that does not play the audio components of the sound-producing object contained in the first video during the first time period; or, the second audio playback device is one of the M audio playback devices that plays the fewest audio components of the sound-producing object contained in the first video during the first time period.
[0023] In this way, electronic devices can prioritize selecting idle audio playback devices to play the second audio component of the first background sound. This reduces the likelihood of a single audio playback device playing too many audio components simultaneously. This allows for more efficient use of the M audio playback devices, achieving a better stereo effect and helping the observer perceive the location of different sound sources during video playback.
[0024] According to the first aspect, or any implementation of the first aspect above, in some embodiments, the first image group includes one or more image frames, the first playback parameters include a first playback time and a first sound intensity, and the second playback parameters include a second playback time and a second sound intensity.
[0025] Adjusting the playback time of the audio component of the sound source on the audio playback device can alter the observer's perception of the sound's direction. Adjusting the sound intensity of the audio component can alter the observer's perception of the sound's distance. This allows the audio playback device to play the audio component of the sound source, making the sound perceived by the observer as originating from a secondary sound source within the primary scene. This helps the observer immerse themselves in the video's context and perceive the location of different sound sources.
[0026] According to the first aspect, or any implementation of the first aspect above, in some embodiments, the first playback time and the second playback time are within a first time period. For example, the first time period is 0-5 seconds of the first video. Then, within 0-5 seconds of the first video playback, the first audio playback device plays the first audio component, and the second audio playback device plays the second audio component. That is, both the first playback time and the second playback time are within the 0-5 second time period of the first video. In other embodiments, because it is necessary to adjust the observer's resolution of the direction of the sound source, the electronic device instructs the audio playback device to adjust the playback time of the audio component of the sound source. The adjustment range of the playback time is typically on the order of milliseconds. Therefore, the first playback time and / or the second playback time may not be within the first time period.
[0027] According to the first aspect, or any implementation of the first aspect above, in some embodiments, the electronic device may send a fourth message to a fourth audio playback device among M audio playback devices. The fourth message includes a second audio component and a fourth playback parameter. The fourth message is used to instruct the fourth audio playback device to play the second audio component with the fourth playback parameter.
[0028] The second audio playback device can be located on the first side of the electronic device, and the fourth audio playback device can be located on the second side of the second electronic device. The first side and the second side are two sides divided by the orientation of the electronic device's display screen.
[0029] In this way, both the second and fourth audio playback devices play the second audio component of the first background sound, which can better form stereo sound in the video playback environment and help the observer perceive the process of the sound-producing object from different directions.
[0030] According to the first aspect, or any implementation thereof, in some embodiments, there are multiple observers viewing the electronic device. The electronic device can obtain a first position based on the positions of the multiple observers in a first scene, the first position representing the position of the observers viewing the electronic device. For example, the first position could be the center of the positions of the multiple observers.
[0031] In this way, during video playback, the electronic device can adjust the playback parameters of the corresponding sound-producing objects on each audio device and the audio components of those objects based on the positions of multiple observers. Each observer can immerse themselves in the scene of the sound-producing objects in the video, experiencing the different sounds emanating from their own location. Furthermore, even if an observer moves while watching the video, the electronic device can still adjust the playback parameters of the corresponding sound-producing objects on each audio device and the audio components of those objects in real time. This allows the observer to still experience the process of the sound-producing objects in the video emanating from their own location, following the perspective of the virtual camera, even while moving.
[0032] According to the first aspect, or any implementation of the first aspect above, in some embodiments, the first scene is the scene of an observer watching an electronic device, and the second scene is the scene presented by the first video. That is, the first scene can be equivalent to a scene in a real three-dimensional space. The second scene can be equivalent to a scene in a virtual three-dimensional space.
[0033] According to the first aspect, or any implementation of the first aspect above, in some embodiments, the first video includes a second image group and a second audio during a second time period. Based on the second image group and the second audio, the electronic device can determine that the first video includes a first sound source and a second background sound during the second time period, and separate a fourth audio component of the first sound source and a fifth audio component of the second background sound from the second audio. The electronic device can send a fifth message to a first audio playback device, the fifth message including the fourth audio component and a fifth playback parameter, the fifth message being used to instruct the first audio playback device to play the fourth audio component with the fifth playback parameter. The electronic device sends a sixth message to a second audio playback device, the sixth message including the fifth audio component and a sixth playback parameter, the sixth message being used to instruct the second audio playback device to play the fifth audio component with the sixth playback parameter. While the first video is playing in the second time period, the electronic device can display images from the second image group.
[0034] It should be noted that in some embodiments, the first background sound and the second background sound described above may be the same. In other embodiments, the first background sound and the second background sound described above are different.
[0035] In this way, the electronic device can analyze the first video segment by segment, detecting in real time changes in the positions of the speaker, the observer, and the audio playback device during the playback of each time segment of the first video. The electronic device can promptly adjust the audio playback device and playback parameters for the speaker's audio components when one or more of these positions change. This method helps the observer immerse themselves in the video's scene throughout playback, experiencing the process of the speaker emitting sound from different locations within the video, following the perspective of a virtual camera.
[0036] According to the first aspect, or any implementation thereof, in some embodiments, the electronic device can obtain the second playback time of the second audio component of the first background sound played by the second audio playback device based on the first time period of the first video. The electronic device can calculate the second sound pressure level of the second audio component and obtain the second sound intensity of the second audio component played by the second audio playback device based on the relationship between sound pressure level and sound intensity.
[0037] According to the first aspect, or any implementation of the first aspect above, in some embodiments, the electronic device can obtain the third playback time of the first audio component of the first sound-emitting object based on the first time period of the first video. It is understood that the third playback time is the original playback time of the first audio component in the first video. The electronic device can determine the first time difference based on the first position of the observer viewing the electronic device in the first scene, the second position of the first audio playback device in the first scene, and the third position of the first sound-emitting object in the first scene. The third position is obtained based on the position of the first sound-emitting object in the second scene relative to the virtual camera of the first video, and by determining the position corresponding to the observer's position in the first scene as the position of the virtual camera in the second scene, wherein, with the first position as the vertex of the included angle, the smaller the included angle formed by the first position, the second position, and the third position, the smaller the first time difference. The electronic device obtains the first playback time of the first audio component played by the first audio playback device based on the sum of the third playback time and the first time difference. The electronic device can determine the first sound pressure level based on the second sound pressure level, the first position, the second position, and the third position, wherein, the closer the second position is to the first position than the third position, the smaller the first sound pressure level, and vice versa. The electronic device can obtain the first sound intensity of the first audio component played by the first audio playback device based on the first sound pressure level and the relationship between sound pressure level and sound intensity.
[0038] According to the first aspect, or any implementation thereof, in some embodiments, the fourth playback parameter includes a fourth playback time and a third sound intensity. The fourth playback parameter is the playback parameter for the second audio component of the first background sound played by the fourth audio playback device. The electronic device can obtain the fourth playback time based on the second playback time and the positions of the second and fourth audio playback devices in the first scene relative to an observer viewing the electronic device, wherein the closer the fourth audio playback device is to the observer than the second audio playback device, the later the fourth playback time is compared to the second playback time. The electronic device can obtain the third sound intensity based on the second sound intensity and the positions of the second and fourth audio playback devices in the first scene relative to an observer viewing the electronic device, wherein the closer the fourth audio playback device is to the observer than the second audio playback device, the lower the third sound intensity is compared to the second sound intensity.
[0039] In this way, when multiple audio playback devices play the second audio component of the first background sound, the first background sound played by these multiple audio playback devices can reach the observer simultaneously, and the sound intensity reaching the observer is the same.
[0040] In some embodiments, an audio playback device can simultaneously play audio components of multiple sound-producing objects. An audio playback device can also simultaneously play audio components of the sound-producing objects and a second audio component of a first background sound. The audio playback device can play the corresponding audio components according to playback parameters sent by the electronic device.
[0041] According to the first aspect, or any implementation of the first aspect above, in some embodiments, during the playback of the first video, the position of an observer viewing the electronic device changes in the first scene. For example, during the playback of the first video from a first time period to a second time period, the observer moves from a first position to a fourth position. While the observer is at the fourth position, the content of the second time period in the first video is playing. The second time period in the first video includes a second image group and a second audio. The first and second time periods can be adjacent time periods in the first video. Based on the second image group and the second audio, the electronic device determines that the first video contains a first sound-emitting object within the second time period, and separates a fourth audio component of the first sound-emitting object from the second audio. The electronic device can send a seventh message to a fifth audio playback device among M audio playback devices. The seventh message includes the fourth audio component and a seventh playback parameter, and the seventh message instructs the fifth audio playback device to play the fourth audio component with the seventh playback parameter. While the first video is playing in the second time period, the electronic device can display images from the second image group.
[0042] The seventh playback parameter can be obtained based on the position of the fifth audio playback device relative to the fourth position in the first scene, the position of the first sound-emitting object relative to the virtual camera of the first video in the second scene, and the position corresponding to the fourth position as the position of the virtual camera in the second scene.
[0043] The aforementioned fifth audio playback device is obtained based on the positions of all or part of the M audio playback devices in the first scene relative to the fourth position, the position of the first sound-emitting object in the second scene relative to the virtual camera of the first video, and the position corresponding to the fourth position being determined as the position of the virtual camera in the second scene.
[0044] In this way, when the observer moves during video playback, the electronic device can adjust the audio playback device and playback parameters of the simulated sound-producing object. Even after changing position, the observer can still immerse themselves in the scene presented in the video, following the perspective of the virtual camera and experiencing the different methods of sound production. This method can increase the observer's sense of immersion and playability when watching videos, thus enhancing the overall viewing experience.
[0045] According to the first aspect, or any implementation of the first aspect above, in some embodiments, the first video includes a third image group and a third audio during a third time period. The third time period follows the aforementioned first time period. Based on the third image group and the third audio, the electronic device can determine that the first video includes a first sound-emitting object during the third time period, and separate a sixth audio component of the first sound-emitting object from the third audio. The position of the first sound-emitting object in the first scene changes from a third position to a fifth position. The second position is the position of the first sound-emitting object during the first time period of the first video, and the fifth position is the position of the first sound-emitting object during the third time period of the first video. The electronic device can send an eighth message to the sixth audio playback device among M audio playback devices. The eighth message includes the sixth audio component and an eighth playback parameter, and the eighth message is used to instruct the sixth audio playback device to play the sixth audio component with the eighth playback parameter. While the first video is playing during the third time period, the electronic device can display the images in the third image group.
[0046] The aforementioned eighth playback parameter can be obtained based on the position of the sixth audio playback device relative to the observer of the viewing electronic device in the first scene, the fifth position, and the position corresponding to the observer's position in the first scene as the position of the virtual camera in the second scene.
[0047] The aforementioned sixth audio playback device is determined based on the positions of all or part of the M audio playback devices in the first scene relative to the observer viewing the electronic device, the fifth position, and the position corresponding to the observer's position in the first scene as the position of the virtual camera in the second scene.
[0048] In this way, when the position of the sound-producing object changes during video playback, the electronic device can adjust the audio playback device and playback parameters to simulate the sound. The observer can immerse themselves in the scene presented in the video and perceive the change in the sound-producing object's position. This method can increase the observer's sense of immersion and engagement when watching the video, thus enhancing the overall viewing experience.
[0049] Secondly, this application provides a method for collaborative audio playback during video playback. This method can be applied to a communication system including an electronic device and M audio playback devices. The electronic device includes a display screen, and M is a positive integer greater than or equal to 2. The electronic device can acquire a first video, which includes a first image group and a first audio signal within a first time period. Based on the first image group and the first audio signal, the electronic device can determine that the first video contains a first sound source and a first background sound within the first time period, and separate a first audio component of the first sound source and a second audio component of the first background sound from the first audio signal. The electronic device can send a first message to the first audio playback device among the M audio playback devices. The first message includes the first audio component and a first playback parameter, and the first message instructs the first audio playback device to play the first audio component with the first playback parameter. The electronic device can send a second message to the second audio playback device among the M audio playback devices. The second message includes the second audio component and a second playback parameter, and the second message instructs the second audio playback device to play the second audio component with the second playback parameter. While the first video is playing in a first time period, the electronic device can display images in the first image group, the first audio playback device can synchronously play the first audio component according to the first message and the first playback parameters, and the second audio playback device can synchronously play the second audio component according to the second message and the second playback parameters.
[0050] In this way, during the playback of the first video, the first audio playback device simulates the effect of the sound source's voice in terms of sound location, matching the position of the first sound source in the three-dimensional space represented by the video. The observer can immerse themselves in the scene presented in the video and more realistically perceive the location of different sound sources. This increases the observer's sense of immersion and engagement with the video, enhancing their overall experience.
[0051] According to the second aspect, in some embodiments, based on the first image group and the first audio, the electronic device further determines that the first video contains a second sound-emitting object within a first time period, and separates a third audio component of the second sound-emitting object from the first audio. The electronic device may send a third message to a third audio playback device among M audio playback devices. The third message contains the third audio component and third playback parameters, and the third message is used to instruct the third audio playback device to play the third audio component with the third playback parameters. While the first video is playing in the first time period, the third audio playback device can synchronously play the third audio component with the third playback parameters according to the third message.
[0052] In this way, when the first video contains multiple sound-producing objects, the electronic device can select an audio playback device to simulate the sound of different objects. The effect of each audio device simulating the sound's location matches the simulated sound object's position in the three-dimensional space represented by the video. This can increase the viewer's immersion and engagement with the video, enhancing their overall experience.
[0053] According to the second aspect, or any implementation of the second aspect above, in some embodiments, the first audio playback device may be obtained based on the position of all or part of the M audio playback devices in the first scene relative to the observer viewing the electronic device, the position of the first sound-emitting object in the second scene relative to the virtual camera of the first video, and determining the position of the virtual camera in the second scene as a position corresponding to the position of the observer in the first scene.
[0054] In this way, the electronic device can determine the position of the first sound-producing object in the first scene and select an audio playback device that is close to the position of the first sound-producing object in the first scene to simulate the sound produced by the first sound-producing object. If the position of one or more of the observer, audio playback device, and first sound-producing object in the first scene changes, the electronic device can reselect an audio playback device to simulate the sound produced by the first sound-producing object.
[0055] According to the second aspect, or any implementation of the second aspect above, in some embodiments, the first playback parameter may be obtained based on the position of the first audio playback device relative to the observer of the viewing electronic device in the first scene, the position of the first sound-emitting object relative to the virtual camera of the first video in the second scene, and determining the position of the virtual camera in the second scene as a position corresponding to the position of the observer in the first scene.
[0056] In this way, when the first audio playback device plays the first audio component with the first playback parameters, the observer perceives the sound as originating from the location of the first sound-emitting object within the first scene. This allows the observer to immerse themselves in the scene presented in the video and more realistically perceive the location of different sound-emitting objects. This increases the observer's sense of immersion and engagement with the video, enhancing their overall experience.
[0057] According to the second aspect, or any implementation of the second aspect above, in some embodiments, the third audio playback device is obtained based on the positions of all or part of the M audio playback devices in the first scene relative to the observer viewing the electronic device, the position of the second sound-emitting object in the second scene relative to the virtual camera of the first video, and determining the position of the virtual camera in the second scene as a position corresponding to the position of the observer in the first scene.
[0058] According to the second aspect, or any implementation of the second aspect above, in some embodiments, the third playback parameter is obtained based on the position of the third audio playback device relative to the observer of the viewing electronic device in the first scene, the position of the second sound-emitting object relative to the virtual camera of the first video in the second scene, and determining the position of the virtual camera in the second scene as a position corresponding to the position of the observer in the first scene.
[0059] The first video may contain more sound-producing objects within the first time period, not limited to the first and second sound-producing objects mentioned above. Understandably, the methods for separating the audio components of other sound-producing objects, determining the audio playback device for playing the audio components of other sound-producing objects, and determining the playback time can all refer to the processing methods for the first and second sound-producing objects mentioned above. These will not be elaborated upon here.
[0060] According to the second aspect, or any implementation of the second aspect above, in some embodiments, the position corresponding to the observer's position in the first scene may be the same as the observer's position in the first scene.
[0061] In this way, during the video playback, the observer can feel the process of the voice-speaking object in the video making a sound from different directions, following the perspective of the virtual camera.
[0062] According to the second aspect, or any implementation thereof, in some embodiments, the observer is located at a first position in the first scene, and the first audio playback device is located at a second position in the first scene. After determining the position of the virtual camera in the second scene as the position corresponding to the observer's position in the first scene, the electronic device can obtain the third position of the first sound-emitting object in the first scene based on the first position and the position of the first sound-emitting object relative to the virtual camera in the second scene. Wherein, with the first position as the vertex of the included angle, the first position, the second position, and the third position form a first included angle, which is the smallest among the included angles formed by the first position as the vertex and the positions of the first position, the third position, and any one of the M audio playback devices.
[0063] According to the second aspect, or any implementation of the second aspect above, in some embodiments, the second audio playback device is one of the M audio playback devices that does not play the audio components of the sound-producing object contained in the first video during the first time period; or, the second audio playback device is one of the M audio playback devices that plays the fewest audio components of the sound-producing object contained in the first video during the first time period.
[0064] In this way, electronic devices can prioritize selecting idle audio playback devices to play the second audio component of the first background sound. This reduces the likelihood of a single audio playback device playing too many audio components simultaneously. This allows for more efficient use of the M audio playback devices, achieving a better stereo effect and helping the observer perceive the location of different sound sources during video playback.
[0065] According to the second aspect, or any implementation of the second aspect above, in some embodiments, the first image group includes one or more image frames, the first playback parameters include a first playback time and a first sound intensity, and the second playback parameters include a second playback time and a second sound intensity.
[0066] According to the second aspect, or any implementation thereof, in some embodiments, the first playback time and the second playback time are within a first time period. In other embodiments, because it is necessary to adjust the observer's resolution of the direction of sound emission from the sound-emitting object, the electronic device instructs the audio playback device to adjust the playback time of the audio component of the sound-emitting object. The adjustment range of the playback time is typically on the order of milliseconds. Therefore, the first playback time and / or the second playback time may not be within the first time period.
[0067] According to the second aspect, or any implementation thereof, in some embodiments, the electronic device may send a fourth message to a fourth audio playback device among M audio playback devices. The fourth message includes a second audio component and fourth playback parameters, and is used to instruct the fourth audio playback device to play the second audio component using the fourth playback parameters. During the playback of the first video in a first time period, the fourth audio playback device may synchronously play the second audio component according to the fourth message and the fourth playback parameters.
[0068] The second audio playback device can be located on the first side of the electronic device, and the fourth audio playback device can be located on the second side of the second electronic device. The first side and the second side are two sides divided by the orientation of the electronic device's display screen.
[0069] In this way, both the second and fourth audio playback devices play the second audio component of the first background sound, which can better form stereo sound in the video playback environment and help the observer perceive the process of the sound-producing object from different directions.
[0070] According to the second aspect, or any implementation thereof, in some embodiments, there are multiple observers viewing the electronic device. The electronic device can obtain a first position based on the positions of the multiple observers in a first scene, the first position representing the position of the observers viewing the electronic device. For example, the first position could be the center of the positions of the multiple observers.
[0071] In this way, during video playback, the electronic device can adjust the playback parameters of the corresponding sound-producing objects on each audio device and the audio components of those objects based on the positions of multiple observers. Each observer can immerse themselves in the scene of the sound-producing objects in the video, experiencing the different sounds emanating from their own location. Furthermore, even if an observer moves while watching the video, the electronic device can still adjust the playback parameters of the corresponding sound-producing objects on each audio device and the audio components of those objects in real time. This allows the observer to still experience the process of the sound-producing objects in the video emanating from their own location, following the perspective of the virtual camera, even while moving.
[0072] According to the second aspect, or any implementation of the second aspect above, in some embodiments, the first scenario is a scenario where an observer watches an electronic device, and the second scenario is a scenario where a first video is presented.
[0073] According to the second aspect, or any implementation thereof, in some embodiments, the first video includes a second image group and a second audio during a second time period. Based on the second image group and the second audio, the electronic device can determine that the first video includes a first sound source and a second background sound during the second time period, and separate a fourth audio component of the first sound source and a fifth audio component of the second background sound from the second audio. The electronic device can send a fifth message to a first audio playback device, the fifth message including the fourth audio component and a fifth playback parameter, the fifth message being used to instruct the first audio playback device to play the fourth audio component with the fifth playback parameter. The electronic device sends a sixth message to a second audio playback device, the sixth message including the fifth audio component and a sixth playback parameter, the sixth message being used to instruct the second audio playback device to play the fifth audio component with the sixth playback parameter. While the first video is playing in the second time period, the electronic device can display images in the second image group, the first audio playback device can synchronously play the fourth audio component with the fifth playback parameter according to the fifth message, and the second audio playback device can synchronously play the fifth audio component with the sixth playback parameter according to the sixth message.
[0074] It should be noted that in some embodiments, the first background sound and the second background sound described above may be the same. In other embodiments, the first background sound and the second background sound described above are different.
[0075] In this way, the electronic device can analyze the first video segment by segment, detecting in real time changes in the positions of the speaker, the observer, and the audio playback device during the playback of each time segment of the first video. The electronic device can promptly adjust the audio playback device and playback parameters for the speaker's audio components when one or more of these positions change. This method helps the observer immerse themselves in the video's scene throughout playback, experiencing the process of the speaker emitting sound from different locations within the video, following the perspective of a virtual camera.
[0076] Thirdly, this application provides a method for coordinating audio playback during video playback, applicable to an electronic device. The electronic device is capable of communicating with M audio playback devices. The electronic device may include a display screen. M is a positive integer greater than or equal to 2. The electronic device can acquire a first video, which contains a first image group and first audio within a first time period. Based on the first image group and the first audio, the electronic device can determine that the first video contains a first sound-emitting object and a second sound-emitting object within the first time period, and separate a first audio component of the first sound-emitting object and a third audio component of the second sound-emitting object from the first audio. The electronic device can send a first message to the first audio playback device among the M audio playback devices, the first message containing the first audio component and a first playback parameter, the first message instructing the first audio playback device to play the first audio component with the first playback parameter. The electronic device can send a third message to the third audio playback device among the M audio playback devices, the third message containing the third audio component and a third playback parameter, the third message instructing the third audio playback device to play the third audio component with the third playback parameter. While the first video is playing in the first time period, the electronic device can display images from the first image group.
[0077] In this way, during the playback of the first video, the sound location of the audio played by each audio device matches the position of the simulated sound-producing object in the three-dimensional space presented in the first video. Observers can immerse themselves in the scene presented in the first video and more realistically perceive the location of different sound-producing objects. This increases the observer's sense of immersion and engagement with the video, enhancing their overall experience.
[0078] In conjunction with the third aspect, in some embodiments, the aforementioned first audio playback device may be obtained based on the positions of all or part of the M audio playback devices in the first scene relative to the observer viewing the electronic device, the position of the first sound-emitting object in the second scene relative to the virtual camera of the first video, and determining the position of the virtual camera in the second scene as a position corresponding to the position of the observer in the first scene.
[0079] In this way, the electronic device can determine the position of the first sound-producing object in the first scene and select an audio playback device that is close to the position of the first sound-producing object in the first scene to simulate the sound produced by the first sound-producing object. If the position of one or more of the observer, audio playback device, and first sound-producing object in the first scene changes, the electronic device can reselect an audio playback device to simulate the sound produced by the first sound-producing object.
[0080] In conjunction with the third aspect, or any of the above-mentioned third aspects, in some embodiments, the first playback parameter is obtained based on the position of the first audio playback device relative to the observer of the viewing electronic device in the first scene, the position of the first sound-emitting object relative to the virtual camera of the first video in the second scene, and determining the position of the virtual camera in the second scene as corresponding to the position of the observer in the first scene.
[0081] In this way, when the first audio playback device plays the first audio component with the first playback parameters, the observer perceives the sound as originating from the location of the first sound-emitting object within the first scene. This allows the observer to immerse themselves in the scene presented in the video and more realistically perceive the location of different sound-emitting objects. This increases the observer's sense of immersion and engagement with the video, enhancing their overall experience.
[0082] In conjunction with the third aspect, or any of the above-mentioned third aspects, in some embodiments, the third audio playback device is obtained based on the positions of all or part of the M audio playback devices in the first scene relative to the observer viewing the electronic device, the position of the second sound-emitting object in the second scene relative to the virtual camera of the first video, and determining the position of the virtual camera in the second scene as a position corresponding to the position of the observer in the first scene.
[0083] In conjunction with the third aspect, or any of the above-mentioned third aspects, in some embodiments, the third playback parameter is obtained based on the position of the third audio playback device relative to the observer of the viewing electronic device in the first scene, the position of the second sound-emitting object relative to the virtual camera of the first video in the second scene, and determining the position of the virtual camera in the second scene as corresponding to the position of the observer in the first scene.
[0084] In this way, the third audio playback device plays the third audio component with third playback parameters, making the sound perceived by the observer as originating from the location of the second sound object in the first scene. This can increase the observer's immersion and playability in watching the video, thus enhancing the observer's overall experience.
[0085] The first video may contain more sound-producing objects within the first time period, not limited to the first and second sound-producing objects mentioned above. Understandably, the methods for separating the audio components of other sound-producing objects, determining the audio playback device for playing the audio components of other sound-producing objects, and determining the playback time can all refer to the processing methods for the first and second sound-producing objects mentioned above. These will not be elaborated upon here.
[0086] In conjunction with the third aspect, or any of the implementations of the third aspect above, in some embodiments, the position corresponding to the observer's position in the first scene may be the same as the observer's position in the first scene.
[0087] In this way, during the video playback, the observer can feel the process of the voice-speaking object in the video making a sound from different directions, following the perspective of the virtual camera.
[0088] In conjunction with the third aspect, or any implementation thereof, in some embodiments, the first image group includes one or more image frames, and the playback parameters include playback time and sound intensity. For example, the first playback parameters play a first playback time and a first sound intensity.
[0089] In conjunction with the third aspect, or any of the implementations of the third aspect above, in some embodiments, the first scenario is a scenario where an observer watches an electronic device, and the second scenario is a scenario where the first video is presented.
[0090] In conjunction with the third aspect, or any implementation thereof, in some embodiments, there are multiple observers viewing the electronic device. The electronic device can determine a first position based on the positions of the multiple observers in a first scene, and this first position represents the position of the observer viewing the electronic device. For example, the first position could be the center of the positions of these multiple observers.
[0091] In this way, during video playback, the electronic device can adjust the playback parameters of the corresponding sound-producing objects on each audio device and the audio components of those objects based on the positions of multiple observers. Each observer can immerse themselves in the scene of the sound-producing objects in the video, experiencing the different sounds emanating from their own location. Furthermore, even if an observer moves while watching the video, the electronic device can still adjust the playback parameters of the corresponding sound-producing objects on each audio device and the audio components of those objects in real time. This allows the observer to still experience the process of the sound-producing objects in the video emanating from their own location, following the perspective of the virtual camera, even while moving.
[0092] Fourthly, this application provides an electronic device. The electronic device includes a memory and a processor. The memory can be used to store a computer program. The processor can be used to invoke the computer program, causing the electronic device to execute any possible implementation method as described in the first or third aspect.
[0093] Fifthly, this application provides a communication system, characterized in that the communication system includes an electronic device and M audio playback devices. The electronic device includes a display screen, where M is a positive integer greater than or equal to 2, and the M audio playback devices include a first audio playback device and a second audio playback device. The electronic device can be used to execute any possible implementation method as described in the first or third aspect. The first audio playback device can be used to synchronously play a first audio component with first playback parameters during the playback of a first video in a first time period. The second audio playback device can be used to synchronously play a second audio component with second playback parameters during the playback of the first video in the first time period.
[0094] Sixthly, this application provides a computer-readable storage medium storing a computer program. When the computer program is run on an electronic device, it causes the electronic device to perform any of the possible implementations of the first or third aspect.
[0095] In a seventh aspect, this application provides a computer program product that may include computer instructions that, when executed on an electronic device, cause the electronic device to perform any possible implementation method as described in the first or third aspect.
[0096] Eighthly, this application provides a chip for use in an electronic device, the chip including one or more processors, the processors being configured to invoke computer instructions to cause the electronic device to perform any of the possible implementation methods of the first or third aspect.
[0097] Understandably, the electronic device provided in the fourth aspect, the communication system provided in the fifth aspect, the computer-readable storage medium provided in the sixth aspect, the computer program product provided in the seventh aspect, and the chip provided in the eighth aspect are all used to execute the methods provided in the embodiments of this application. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods, and will not be repeated here. Attached Figure Description
[0098] Figure 1 This is a schematic diagram illustrating a scenario of coordinated audio playback during video playback, provided in an embodiment of this application.
[0099] Figures 2A to 2D These are schematic diagrams illustrating the principles of sound source location identification provided in the embodiments of this application;
[0100] Figure 3A This is a schematic diagram of the structure of an electronic device 100 provided in an embodiment of this application;
[0101] Figure 3B This is a schematic diagram of the structure of an audio playback device 200 provided in an embodiment of this application;
[0102] Figure 4 This is a flowchart illustrating a method for co-playing audio during video playback, as provided in an embodiment of this application.
[0103] Figure 5A This is a flowchart of a method for determining a sound-producing object and its audio components in a video, provided in an embodiment of this application.
[0104] Figure 5B This is a schematic diagram of an image recognition method provided in an embodiment of this application;
[0105] Figure 5C This is a schematic diagram of a method for separating audio components provided in an embodiment of this application;
[0106] Figure 5D This is a flowchart of a method for determining a sound-producing object and its audio components in a video, provided in an embodiment of this application.
[0107] Figure 5E This is a flowchart of another method for determining the sound-producing object and the audio components of the sound-producing object in a video, provided in an embodiment of this application.
[0108] Figure 5F This is a flowchart of another method for determining the sound-producing object and the audio components of the sound-producing object in a video, provided in an embodiment of this application.
[0109] Figure 6A This is a schematic diagram illustrating a method for determining the location of a sound-emitting object and the location of a virtual camera, provided in an embodiment of this application.
[0110] Figure 6B and Figure 6C This is a schematic diagram showing the positional relationship between some sound-emitting objects and virtual cameras provided in the embodiments of this application;
[0111] Figure 6D This is a flowchart illustrating a method for obtaining the location of an audio playback device according to an embodiment of this application;
[0112] Figure 6E This is a schematic diagram illustrating the positional relationship between an observer, an electronic device 100, and an audio playback device, provided in an embodiment of this application.
[0113] Figure 6FThis is a schematic diagram illustrating the positional relationship between an observer, an electronic device 100, an audio playback device, and a sound-emitting object, provided in an embodiment of this application.
[0114] Figures 7A to 7C These are schematic diagrams illustrating the determination of the sound-producing object and playback parameters corresponding to an audio playback device, as provided in embodiments of this application.
[0115] Figure 8A and Figure 8B This is a schematic diagram showing the positional relationships of other observers, electronic devices 100, audio playback devices, and sound-emitting objects provided in the embodiments of this application;
[0116] Figure 9A This is another schematic diagram showing the positional relationship between the observer, electronic device 100, and audio playback device provided in an embodiment of this application;
[0117] Figure 9B This is a schematic diagram showing the positional relationship between another sound-emitting object and a virtual camera provided in an embodiment of this application;
[0118] Figure 9C This is another schematic diagram showing the positional relationship between the observer, electronic device 100, audio playback device, and sound-emitting object provided in an embodiment of this application;
[0119] Figure 9D and Figure 9E This is a schematic diagram illustrating a scenario of a method for a user to select audio to play concurrently during video playback, as provided in an embodiment of this application.
[0120] Figures 10-12 These are schematic diagrams of the structures of some communication systems provided in the embodiments of this application. Detailed Implementation
[0121] The technical solutions of the embodiments of this application are described below with reference to the accompanying drawings. In the description of the embodiments of this application, the terminology used in the following embodiments is for the purpose of describing specific embodiments only and is not intended to limit the application. As used in the specification and appended claims of this application, the singular expressions "a," "the," "the," "the," and "this" are intended to also include expressions such as "one or more," unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of this application, "at least one" and "one or more" refer to one or more (including two). The term "and / or" is used to describe the relationship between related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.
[0122] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized. The term "connection" includes direct connections and indirect connections, unless otherwise stated. "First" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated.
[0123] In the embodiments of this application, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of the words "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.
[0124] This section first introduces a typical scenario of collaborative audio playback during video playback, as provided in an embodiment of this application.
[0125] like Figure 1 As shown, the device for playing video may include electronic device 100 and one or more audio playback devices. For example, the audio playback devices may include audio playback devices 200, 201, 202, and 203. Audio playback devices 200, 201, 202, and 203 can all establish a communication connection with electronic device 100. The aforementioned communication connection may include a wired communication connection, or a wireless communication connection such as a Bluetooth communication connection, a Wi-Fi communication connection, or a ZigBee communication connection. This application embodiment does not limit the specific method of establishing the aforementioned communication connection.
[0126] For example, the electronic device 100 may be a device with a display screen (e.g., a television, a monitor, etc.) that can be used to display images from a video. In some embodiments, the electronic device 100 has an audio output device (e.g., a speaker), and the electronic device 100 can also be used to play audio from the video.
[0127] For example, the audio playback devices 200-203 can be devices with audio output devices (such as speakers), amplifiers, etc., and can be used to play audio from the video in conjunction with the video playback. In some embodiments, the audio playback devices 200-203 can be distributed on both sides (e.g., the left and right sides) of the electronic device 100. Optionally, the audio playback devices 200-203 can be devices of the same model or different models.
[0128] In one possible implementation, during video playback, electronic device 100 can instruct one or more of audio playback devices 200-203 to play the audio in the video in coordination, thereby providing the user with a better listening experience.
[0129] For example, Figure 1 The observer shown can be a user watching the video. During the video viewing process, the user can browse the video screen through electronic device 100 and listen to the audio in the video through audio playback devices 200-203.
[0130] In the scenario where audio is played in conjunction with video playback, the user can perceive the loudness, pitch, and timbre of the sound produced by the audio playback devices 200-203 using their ears. Using both ears, the user can also distinguish the direction of the sound source and determine its location. For example, the location of the sound source could be the location of one or more audio playback devices among audio playback devices 200-203.
[0131] Users are able to pinpoint sound locations primarily by sensing at least one of the following: the time difference, phase difference, sound level difference, and timbre difference between the two ears. The time difference represents the difference in the order in which the sound reaches the two ears. The phase difference represents the difference in the phase of the sound waves heard by the two ears. The sound level difference represents the difference in the intensity of the sound heard by the two ears. The timbre difference represents the difference in the timbre of the sound heard by the two ears.
[0132] The following describes a method for users to distinguish the location of a sound source by the difference in sound heard by both ears.
[0133] For example, Figure 2A This illustration depicts a scenario where sound reaches a user, as provided in an embodiment of this application. Figure 2AAs shown, the sound source is to the user's lower left (or to the user's right front). The distance between the sound source and the user's right ear is less than the distance between the sound source and the user's left ear. The user's right ear hears the sound produced by the sound source before their left ear. Since sound energy decreases as it travels further in the air, the sound intensity heard by the user's right ear is greater than the sound intensity heard by their left ear. The user can determine that the sound source is to their right based on one or more of the time difference and sound level difference mentioned above.
[0134] When a sound source produces sound that reaches a user's ears, there may be a time difference between the two ears, and the sound waves may also have a phase difference.
[0135] For example, Figure 2B A phase diagram of a signal provided in an embodiment of this application is shown. Figure 2B The phase of a sound wave is illustrated using a sine wave as an example. Phase is a measure of change in a signal waveform. Phase can be measured in degrees (°). When a signal waveform changes periodically, one cycle of the signal waveform is 360°. For example... Figure 2B As shown, an ideal sine wave, within a complete cycle, can have its phase change from 0° to 360°, experiencing a peak and a trough before returning to its original position. If the initial phase of the sine wave is 0°, then the wave starts vibrating in the positive direction, with a peak phase of 90° and a trough phase of 270°, a phase difference of 180°. Since sound from a sound source reaches a user's ears at different times, the phase of the sound wave may differ when it reaches the ears.
[0136] For example, Figure 2C and Figure 2D This application illustrates other scenarios where sound reaches the user, as provided in its embodiments. For example... Figure 2C As shown, the phase of the sound wave when it reaches the user's right ear is phase A1, and the phase when it reaches the user's left ear is phase A2. Phase A1 and phase A2 are different. The human ear can perceive this phase difference, and thus the user can determine the location of the sound source based on this phase difference. Of course, in Figure 2A In this system, the human ear can also perceive the aforementioned phase difference, and users can then determine the location of the sound source based on this phase difference.
[0137] Furthermore, when propagating sound waves encounter obstacles with geometric dimensions greater than or equal to their wavelength, a shielding effect occurs (also known as the masking effect). The higher the frequency of the sound wave, the shorter its wavelength. For example, a sound wave with a frequency of 20 Hz at room temperature has a wavelength of 17 meters (m), while a sound wave with a frequency of 200 Hz has a wavelength of 1.7 meters. High-frequency sound waves are blocked by obstacles during propagation, making it difficult for them to continue. Low-frequency sound waves, however, can diffract when encountering obstacles, allowing them to bypass them and continue their propagation.
[0138] like Figure 2A As shown, the sound from the sound source reaches the user's right ear without interference from any obstacles. The user's right ear can hear all sound waves (such as high-frequency and low-frequency sound waves). However, the sound from the sound source is interfered with by the user's head as it reaches the user's left ear. Some high-frequency sound waves may be blocked and unable to reach the user's left ear. The user's left ear cannot hear these blocked high-frequency sound waves. Therefore, the frequencies of the sound heard by the user's left and right ears differ, meaning the timbre of the sound heard by both ears is different.
[0139] like Figure 2D As shown, because the shape of the auricle is not symmetrical, the propagation process of sound into the ear canal differs depending on whether the sound source is in front of or behind the user. Sound from a sound source in front of the user can be reflected by the auricle and enter the ear canal directly. Sound from a sound source behind the user needs to bypass the auricle to enter the ear canal. It can be seen that sound from a sound source behind the user is blocked by the auricle during propagation. Therefore, some high-frequency sound waves from the rear may not be able to enter the user's ear canal. The same sound source will produce a different timbre depending on whether it is in front of or behind the user. The auditory cortex of the brain can compare the timbre of the heard sound with previously known signals to determine whether the sound source is in front of or behind the user.
[0140] As can be seen from the above, through the time difference, phase difference, sound level difference, and timbre difference, users can distinguish the specific location of the sound source they hear.
[0141] In the scenario where audio is played in tandem with video playback, electronic device 100 can display images from the video on a screen, while audio playback devices 200-203 can synchronously play audio from the video. During video viewing, the user hears sounds from the audio playback devices (i.e., the sound source is the aforementioned audio playback devices). However, the video being played may contain multiple voices (for example, if the video is a recording of a conversation between multiple people, then the voices in the video include multiple individuals). The sounds of these multiple voices are transmitted to the user by the audio playback devices during video playback. It is difficult for the user to determine the location of these multiple voices based on the sounds they hear. In other words, the user cannot quickly immerse themselves in the video's scene, cannot perceive different voices emanating from their different locations, and the user experience is limited.
[0142] To provide users with a better listening and viewing experience, this application provides a method for co-playing audio during video playback. In this method, an electronic device 100 analyzes the images and audio contained in the video to be played, determining the sound-producing object in the video and its audio components. The electronic device 100 determines the position of the sound-producing object and a virtual camera based on the images contained in the video. The position of the virtual camera represents the position of the camera during the actual shooting of the video to be played. The electronic device 100 determines the positions of the observer and the audio playback device. Using the position of the virtual camera as the observer's position, the electronic device 100 determines the positional relationship between the sound-producing object in the video and the observer. Based on the positional relationship between the audio playback device, the sound-producing object, and the observer, the electronic device 100 instructs the audio playback device to simulate the sound of the sound-producing object. Specifically, the electronic device 100 can instruct the audio playback device to adjust playback parameters (such as playback time, sound intensity, etc.), and use the adjusted playback parameters to play the audio components of the simulated sound source of the audio playback device, so that when the user distinguishes the sound of different sound sources during video playback, the location of the sound source emitted by the audio playback device can be determined as the location of the simulated sound source of the audio playback device.
[0143] In other words, users can immerse themselves in the scene where the voice is being spoken in the video, experiencing the different voices coming from different directions within themselves. This method increases the user's sense of immersion and engagement when watching videos, thus enhancing the overall user experience.
[0144] In some embodiments, the electronic device 100 may be a display device with image display functionality. The device used in the method of co-playing audio during video playback to analyze the images and audio contained in the video and instruct the audio playback device to simulate the sound of a sound-producing object may be other electronic devices (e.g., video analysis devices and control devices). The video analysis device may be used to analyze the images and audio contained in the video. The control device may be used to instruct the electronic device 100 to display images in the video and to instruct the audio playback device to simulate the sound of a sound-producing object. That is, the communication system to which the method of co-playing audio during video playback is applied may include the electronic device 100 and one or more audio playback devices. Optionally, the communication system may also include the video analysis device and control device. This application embodiment does not limit the devices included in the communication system.
[0145] The following embodiments of this application will be specifically described using a communication system consisting of an electronic device 100 and an audio playback device as an example.
[0146] To facilitate understanding, some concepts involved in this application are introduced below.
[0147] 1. Video
[0148] In this application, "video" can refer to multimedia data containing images and audio. In some embodiments, the video may be captured by a device equipped with a camera, such as a camera, mobile phone, tablet computer, laptop computer, or television. In some embodiments, the video may also be obtained by synthesizing multiple frames of images. This application does not limit the method of video generation. Video file formats may include avi, mp4, mov, wmv, etc. This application does not limit the video file format. Video playback devices (such as mobile phones, televisions, etc.) have a display screen and an audio output device. The video playback device can play the images contained in the video at a preset frame rate, such as 24 frames per second, 30 frames per second, 60 frames per second, etc. The aforementioned 24 frames per second can mean that the video playback device continuously displays 24 frames of images on the display screen per second.
[0149] The images and audio in a video have a temporal correspondence. When a video playback device plays a video based on this temporal correspondence, it can synchronize the video image and sound. Specifically, the display time of an image is the same as the playback time of its corresponding audio. Alternatively, the interval between the display time of an image and the playback time of its corresponding audio is less than a preset time interval. This preset time interval can be the maximum time difference between the video image and sound that the user does not perceive as being out of sync, for example, 100 milliseconds. In other words, if the time interval between the display time of an image and the playback time of its corresponding audio does not exceed 100 milliseconds, the video image and sound are perceived as synchronized by the user. For example, if a user sees someone speaking on a video screen and hears that same person speaking simultaneously, then the video image and sound can be considered synchronized.
[0150] 2. Virtual camera
[0151] Images in a video can be two-dimensional. Objects contained within images in a video can reside in a three-dimensional space. These objects can be located in different positions within that three-dimensional space. A virtual camera can be a camera that captures objects in three-dimensional space as the objects contained in the images of the aforementioned video. In other words, images in a video can be considered as being captured by a virtual camera.
[0152] By performing 3D reconstruction on a 2D image, the objects contained in the image and the position of the virtual camera in 3D space can be obtained. This 3D reconstruction can construct the 3D space of the video recording, facilitating the determination of the positional relationship between the user watching the video and the sound-producing objects in the video. This allows the user to experience an immersive experience, feeling as if different sound-producing objects in the video are emitting sounds from their different locations, thus placing the user in the scene.
[0153] The implementation method of the above-mentioned three-dimensional reconstruction will be introduced in subsequent embodiments, and will not be elaborated here.
[0154] 3. The target of the speech
[0155] The sound-producing object can refer to any object capable of emitting sound. For example, a person, or animals such as cats and dogs, or vehicles such as trains and cars, or natural features such as waterfalls, rain, thunder, and the sea. This application does not limit the type of sound-producing object in its embodiments.
[0156] 4. Playback parameters
[0157] Playback parameters represent the parameters used by an audio playback device to play audio. An audio playback device is any device used to play audio. Playback parameters can include playback time and sound intensity. Sound intensity is also known as volume. When the playback parameters of an audio playback device change, the sound effects produced by the device can also change. Sound effects can include the user's subjective auditory perception of sound, such as loudness, pitch, and timbre. Sound effects can also include objective physical quantities of sound, such as the phase of a sound wave and sound pressure level.
[0158] For example, when an audio playback device increases the intensity of the audio being played, the louder the sound heard by the user, the greater the sound pressure at the user's ears. When sound waves propagate through the air, the density of air particles changes with the sound wave, and the pressure at that point changes accordingly. This change in pressure is called sound pressure. In other words, sound pressure can represent the change in pressure caused by the vibration of sound waves as they travel through a medium. The magnitude of sound pressure can be represented by sound pressure level (SPL). Sound pressure and sound intensity have the following relationship:
[0159] p 2 =I*ρ*C (1)
[0160] The value of 'p' above represents sound pressure. The value of 'I' above represents sound intensity. The value of 'ρ' above represents the density of the medium. The value of 'C' above represents the speed of sound. In air, the value of 'C' can be 340 m / s. It can be seen that, for the same medium density, the greater the sound intensity, the greater the sound pressure.
[0161] In some embodiments, two audio playback devices can play different audio components of an audio segment. For example, audio playback device A1 plays audio component B1 of an audio segment. Audio playback device A2 plays audio component B2 of the same audio segment. Audio playback device A1 does not change its playback parameters while playing audio component B1. When audio playback device A2 adjusts the sound intensity of audio component B2, the user's perception of the distance between themselves and the sound source of audio component B2 (i.e., audio playback device A2) can change. For example, when audio playback device A2 increases the sound intensity of audio component B2, the user may perceive that the distance between themselves and audio playback device A2 is decreasing.
[0162] When audio playback device A2 adjusts the playback time of audio component B2, a time difference will occur between audio components B1 and B2 during playback. For example, the audio at time point Tc of audio component B1 in the aforementioned audio segment should ideally be played simultaneously with the audio at time point Tc of audio component B2 in the same audio segment. However, due to the adjustment of the playback time of audio component B2 by audio playback device A2, the audio at time point Tc of audio component B1 in the aforementioned audio segment is no longer played simultaneously with the audio at time point Tc of audio component B2 in the same audio segment, resulting in a certain time difference. Therefore, the phase of the aforementioned audio segment reaching the user's ears may change, thereby altering the user's perception of the direction of the sound source of audio component B2 (i.e., audio playback device A2) relative to themselves. In other words, by changing the playback time of audio components in an audio segment, creating a time difference between the playback times of different audio components, it is possible to change the phase of the sound waves.
[0163] This application does not limit the type of playback parameters. For example, the playback parameters may also include the frequency of sound waves. Adjusting the frequency of the sound waves can change the pitch, timbre, and other sound effects heard by the user.
[0164] The following is a schematic diagram of the structure of an electronic device 100 involved in this application.
[0165] For example, such as Figure 3A As shown, the electronic device 100 may include a processor 110, internal memory 121, universal serial bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, button 190, motor 191, indicator 192, camera 193, display screen 194, etc.
[0166] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0167] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0168] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of instruction fetching and execution.
[0169] USB port 130 is a USB standard compliant interface. USB port 130 can be used to connect a charger to charge electronic device 100, and can also be used for data transfer between electronic device 100 and peripheral devices. It can also be used to connect headphones for audio playback.
[0170] The charging management module 140 receives charging input from the charger. While charging the battery 142, the charging management module 140 can also supply power to the electronic device through the power management module 141.
[0171] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to power the processor 110, internal memory 121, external memory, display 194, camera 193, and wireless communication module 160, etc.
[0172] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.
[0173] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals.
[0174] The mobile communication module 150 can provide wireless communication solutions, including 2G / 3G / 4G / 5G, for use on electronic devices 100.
[0175] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, etc.
[0176] In some embodiments, the electronic device 100 may further include a millimeter-wave radar module. This millimeter-wave radar module can transmit millimeter-wave radar signals and receive reflected millimeter-wave radar signals. Based on the transmitted millimeter-wave type signal and the received reflected millimeter-wave type signal, the millimeter-wave radar module can obtain a difference frequency signal and use this difference frequency signal to determine the position of a target (such as an object, human body, etc.).
[0177] In some embodiments, the electronic device 100 may further include an ultrawideband (UWB) module. The UWB module can provide a wireless communication solution based on UWB technology applied to the electronic device 100. For example, the UWB module can function as a UWB base station. The UWB base station can be used to locate UWB tags. Specifically, the distance between the UWB tag and the UWB base station can be obtained by detecting the UWB signal and combining it with certain positioning algorithms to calculate the duration of the UWB signal's flight through the air. This duration multiplied by the transmission rate of the UWB signal in the air (e.g., the speed of light).
[0178] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU performs mathematical and geometric calculations and is used for graphics rendering. The display screen 194 is used to display images, videos, etc. In some embodiments, electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1. Electronic device 100 can implement shooting functions through an ISP, a camera 193, a video codec, a GPU, the display screen 194, and the application processor. The camera 193 is used to capture still images or videos. In some embodiments, electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0179] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0180] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs can enable intelligent cognitive applications in electronic devices, such as image recognition, 3D image reconstruction, facial recognition, speech recognition, and text understanding.
[0181] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121.
[0182] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.
[0183] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.
[0184] The loudspeaker 170A, also known as a "loudspeaker", is used to convert audio electrical signals into sound signals.
[0185] The receiver 170B, also known as the "earpiece", is used to convert audio electrical signals into sound signals.
[0186] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. Electronic device 100 may include at least one microphone 170C. In some embodiments, electronic device 100 may include two microphones 170C, which, in addition to acquiring sound signals, can also perform noise reduction. In other embodiments, electronic device 100 may also include three, four, or more microphones 170C, enabling sound signal acquisition, noise reduction, sound source identification, and directional recording, among other functions.
[0187] The 170D headphone jack is used to connect wired headphones.
[0188] The sensor module 180 may include one or more of the following: pressure sensor, gyroscope sensor, barometric pressure sensor, magnetic sensor, accelerometer, distance sensor, proximity sensor, infrared sensor, fingerprint sensor, temperature sensor, touch sensor, ambient light sensor, and bone conduction sensor.
[0189] Buttons 190 include a power button, volume buttons, etc. Electronic device 100 can receive button input and generate key signal inputs related to user settings and function control. Motor 191 can generate vibration alerts. Indicator 192 can be an indicator light, used to indicate charging status, battery level changes, and also to indicate messages, missed calls, notifications, etc.
[0190] Figure 3A The electronic device 100 shown may be equipped with Alternatively, it can be a portable electronic device with other operating systems, such as a mobile phone, tablet computer, laptop computer, etc., or a non-portable electronic device such as a laptop computer or desktop computer with a touch-sensitive surface or touch panel. This application does not limit the type of electronic device 100.
[0191] The following is a schematic diagram of the structure of an audio playback device 200 involved in this application.
[0192] For example, such as Figure 3B As shown, the audio playback device 200 may include a communication device 210, a memory 211, a processor 212, a microphone 213, and a speaker 214 coupled via a bus. Wherein:
[0193] The communication device 210 can be used to establish a communication connection between the audio playback device 200 and other electronic devices (such as electronic device 100). For example, the audio playback device 200 can receive audio components, playback parameters, and playback instructions sent by the electronic device 100 through the communication device 210. The playback instructions can be used to instruct the audio playback device 200 to play the audio components according to the playback parameters.
[0194] The memory 211 can be used to store various software programs and / or multiple sets of instructions. The memory 211 can also store a communication program, which can be used to communicate with devices such as the electronic device 100. In some embodiments, the memory 211 can also store a video analysis program. This video analysis program can be used to analyze the images and audio contained in a video to obtain information such as the sound-producing object in the video, the audio components of the sound-producing object, and the position of the sound-producing object. The memory 211 can also store a playback parameter determination program. This playback parameter determination program can be used to determine the playback parameters used by an audio playback device to play the audio component of a sound-producing object.
[0195] The processor 212 can be used to read and execute programs stored in the memory 211. Examples include communication programs, video analysis programs, playback parameter determination programs, etc. In other words, the audio playback device 200 can determine the sound-producing object in the video and its audio components. The audio playback device 200 can determine which audio playback device will play the audio components of a sound-producing object, and the playback parameters of the audio playback device when simulating the sound-producing object's sound.
[0196] Microphone 213 can be used to collect sound signals and convert them into electrical signals. Audio playback device 200 may include one or more microphones. In some embodiments, audio playback device 200 can use microphones to collect sound signals to identify the location of a sound source. Not limited to microphone 213, audio playback device 200 may also include other types of audio input devices.
[0197] Speaker 214 can be used to convert audio electrical signals into sound signals. That is, audio playback device 200 can play audio from a video through speaker 214. Not limited to speaker 214, audio playback device 200 may also include other types of audio output devices.
[0198] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the audio playback device 200. In other embodiments of this application, the audio playback device 200 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware. The structures of other audio playback devices (such as audio playback devices 201-203) involved in the embodiments of this application can be referred to. Figure 3B The structure shown is not described in detail here.
[0199] The following describes in detail a method for co-playing audio during video playback, provided by an embodiment of this application.
[0200] Please refer to Figure 4 , Figure 4A flowchart illustrating a method for collaborative audio playback during video playback, provided in an embodiment of this application, is shown as an example. This method can be applied to a communication system including an electronic device 100, an audio playback device 200, and an audio playback device 201. It is not limited to audio playback devices 200 and 201; the communication system may also include more audio playback devices. This embodiment specifically uses two audio playback devices as an example for illustration.
[0201] The electronic device 100 establishes a communication connection with the audio playback devices 200 and 201. This communication connection may include wired communication connections, as well as wireless communication connections such as Bluetooth, Wi-Fi, and ZigBee. This application embodiment does not limit the specific method of establishing the above communication connection.
[0202] This method may include steps S411 to S416. Wherein:
[0203] An observer can perform a first operation. This first operation can be used to play a first video. The observer can refer to a user watching the first video (i.e., watching electronic device 100). In some embodiments, the first operation can be applied to electronic device 100. For example, electronic device 100 has a video playback control for playing the first video. The first operation can be an operation on that video playback control. In other embodiments, the first operation can also be applied to other electronic devices besides electronic device 100 (such as a remote control, speaker, mobile phone, etc. for controlling electronic device 100). After receiving the first operation, the other electronic device can send a video playback command to electronic device 100. This video playback command can be used to instruct electronic device 100 to play the first video. The embodiments of this application do not limit the specific implementation of the first operation.
[0204] In some embodiments, the first video can be divided into multiple different time segments according to the order of playback time. For example, a first time segment, a second time segment, a third time segment, etc. (the specific number of time segments is not limited). The duration of a time segment can be, for example, 4 seconds, 5 seconds, etc. This application embodiment does not limit the value of the duration of a time segment. The electronic device 100 can analyze the first video segment by segment according to the order of time segments. In addition to the sound emitted by the specific sound-emitting object, the audio in the first video usually also includes background noise. The background noise can include sounds other than the sound of the sound-emitting object in the environment during the recording of the first video, machine noise of the video recording device, etc. The background noise does not have a corresponding sound-emitting object. In one possible implementation method, the electronic device 100 can determine the audio components of different sound-emitting objects in an audio segment. After determining the audio components of all sound-emitting objects in an audio segment, the electronic device 100 can determine the remaining audio components after excluding the audio components of the sound-emitting objects as the audio components of the background noise.
[0205] For example, the first video contains a first group of images and a first audio track within a first time period. The electronic device 100 can determine, based on the first group of images and the first audio track, that the first video contains a first sound-emitting object and a first background sound within the first time period, and separate a first audio component of the first sound-emitting object and a second audio component of the first background sound from the first audio track. The first video contains a second group of images and the first audio track within a second time period. The electronic device 100 can determine, based on the second group of images and the second audio track, that the first video contains a fifth audio component of the first sound-emitting object and the second background sound within the first time period. The first background sound and the second background sound can be the same. Alternatively, the first background sound and the second background sound can be different.
[0206] In some embodiments, the first video may also contain only one time period, namely the first time period.
[0207] To facilitate understanding and make the description clearer, this embodiment and subsequent embodiments of this application specifically use the first audio component of the first sound-emitting object and the second audio component of the first background sound contained in the first video within the first time period as examples. Those skilled in the art should understand that the processing method for the audio components of the sound-emitting object and background sound contained in other time periods of the first video (e.g., the second time period) is the same as the processing method for the first and second audio components described above. This application does not elaborate on the processing of the audio components of the sound-emitting object and background sound contained in other time periods of the first video.
[0208] S411, Electronic device 100 receives the first operation and acquires the first video. The data in the first video within a first time period includes a first image group and a first audio. The images in the first image group present the action of a first sound-emitting object making a sound, and the first audio includes a first audio component of the first sound-emitting object and a second audio component of a first background sound.
[0209] Optionally, the first video may be a video stored locally on the electronic device 100. In response to a first operation of playing the first video, the electronic device 100 may retrieve the data of the first video from its local memory.
[0210] Optionally, the first video may also be a video stored on a cloud server. In response to the first operation of playing the first video, the electronic device 100 may request the data of the first video from the cloud server. This application embodiment does not limit the method by which the electronic device 100 obtains the first video.
[0211] In one possible implementation, the electronic device 100 can analyze the first video in real time during playback to determine the sound-emitting object, its audio components, and its location. The location of the sound-emitting object can be used to determine the audio playback device for playing its audio components. Specifically, the electronic device 100 can select one audio playback device from multiple options based on the sound-emitting object's location, the location of the audio playback device, and the observer's location to play the sound-emitting object's audio components, and determine the playback parameters for playing those audio components.
[0212] The method by which the electronic device 100 obtains the position of the sound-emitting object, the position of the audio playback device, and the position of the observer will be described in detail in subsequent embodiments. It will not be elaborated here.
[0213] Understandably, the position of the sound-producing object in the first video may change, and the observer's position may also change during the playback of the first video. Real-time analysis of the first video to determine the audio components and playback parameters played by the audio playback device can enable the audio playback device to better simulate the sound-producing process of the sound-producing object, thereby improving the user's sense of immersion when watching the video.
[0214] The first image group mentioned above may contain one or more frames of images.
[0215] The method by which the aforementioned electronic device 100 determines the sound-emitting object in the video and the audio components of the sound-emitting object will be described in detail in subsequent embodiments. It will not be elaborated here.
[0216] S412, the electronic device 100 sends a first message to the audio playback device 200. The first message includes a first audio component and a first playback parameter. The first message is used to instruct the first audio component to be played with the first playback parameter.
[0217] In one possible implementation, the electronic device 100 can determine the correspondence between the audio playback device 200 and the first sound-emitting object based on the position of the first sound-emitting object, the position of the audio playback device 200, and the position of the observer. The position of the first sound-emitting object can be its location within the first time period. The position of the observer can be their location when the video is played within the first time period.
[0218] Once it is determined that the audio playback device 200 corresponds to the first sound-emitting object, the electronic device 100 can instruct the audio playback device 200 to play the first audio component of the first sound-emitting object. The first playback parameters may include a first playback time and a first sound intensity. The aforementioned first playback parameters may be determined based on the first audio, the position of the first sound-emitting object, the position of the audio playback device 200, and the position of the observer.
[0219] In this context, both the observer and the audio playback device 200 can be located in a first scene, which can represent the scene where the first video is viewed (i.e., a real scene). The observer's position can be within the first scene. Similarly, the audio playback device 200's position can also be within the first scene. Both the first sound-emitting object and the virtual camera of the first video can be located in a second scene, which can represent the scene depicted in the first video (i.e., a virtual scene). Based on the positional relationship between the first sound-emitting object and the virtual camera in the second scene, and by determining the position of the virtual camera to correspond to the observer's position in the first scene, the positional relationship between the first sound-emitting object and the observer in the first scene can be obtained.
[0220] The method by which electronic device 100 determines the correspondence between the audio playback device and the sound-emitting object, and the method for determining the playback parameters, will be described in detail in subsequent embodiments. It will not be elaborated here.
[0221] S413, the electronic device 100 sends a second message to the audio playback device 201. The second message contains a second audio component and a second playback parameter. The second message is used to instruct the second audio component to be played with the second playback parameter.
[0222] In one possible implementation, the electronic device 100 can select an idle audio playback device from among multiple audio playback devices to play the second audio component of the first background sound. For example, during the video playback within the aforementioned first time period, the audio playback device 201 does not play the audio component of the sound source (i.e., the audio playback device 201 is in an idle state). The electronic device 100 can instruct the audio playback device 201 to play the second audio component. The second playback parameters may include a second playback time and a second sound intensity. The aforementioned second playback parameters can be determined based on the second audio component.
[0223] In another possible implementation, the electronic device 100 can select at least one audio playback device from the audio playback devices located on one side (e.g., the left side) and at least one audio playback device from the audio playback devices located on the other side (e.g., the right side) to play the second audio component. The one side and the other side of the electronic device 100 can be the side and the other side of the orientation of the display screen of the electronic device 100. The electronic device 100 can preferentially select an idle audio playback device to play the second audio component of the first background sound. Having audio playback devices on both sides of the electronic device 100 playing the second audio component of the first background sound can better create stereo sound in the video playback environment, helping the user perceive the process of sound emanating from different directions.
[0224] When multiple audio playback devices are playing the second audio component of the first background sound, one of these devices can be a reference audio playback device. This reference device can play the second audio component according to the aforementioned second playback parameters. The playback parameters for the other audio playback devices playing the first audio component 2 can be determined based on the second playback parameters, the position of the reference audio playback device, the positions of the other audio playback devices, and the position of the observer. The first background sound played by the multiple audio playback devices can reach the observer simultaneously, and the sound intensity of the first background sound is the same when it reaches the observer.
[0225] The method by which the electronic device 100 determines the playback parameters for playing the second audio component will be described in detail in subsequent embodiments. It will not be elaborated upon here.
[0226] S414, Electronic device 100 displays images in the first image group.
[0227] S415, the audio playback device 200 synchronously plays the first audio component according to the first playback parameters.
[0228] S416, Audio playback device 201 synchronously plays the second audio component according to the second playback parameters.
[0229] The first audio component and the second audio component are both audio from the first video within a first time period. The playback of the first audio component by audio playback device 200 and the playback of the second audio component by audio playback device 201 are synchronized with the display of images in the first image group by electronic device 100. For example, the first image group includes images depicting a first sound-emitting object (such as an image of person 1 speaking). During the display of the images depicting the first sound-emitting object, audio playback device 200 plays the first audio component. In other words, the video image and the audio within the video are synchronized.
[0230] From the above Figure 4 As shown in the method, different audio playback devices can simulate different sound sources during video playback. The effect of each audio playback device simulating the sound source's location is matched with the simulated location of the sound source in the three-dimensional space depicted in the video. This allows users to more realistically experience being present in the scene, feeling as if they are witnessing the sound source in different locations within the video. This method increases the user's immersion and engagement with the video, enhancing the overall user experience.
[0231] The following describes in detail an implementation method for determining the sound-producing object contained in a video and the audio components of the sound-producing object, provided by an embodiment of this application.
[0232] Please refer to Figure 5A , Figure 5A An exemplary flowchart illustrates a method for determining a sound-producing object and its audio components within a video. This method can be applied to an electronic device 100. Specifically, it is illustrated here using the determination of a sound-producing object and its audio components within a first time period in a first video.
[0233] The method may include steps S510 to S530. Wherein:
[0234] S510. Perform image recognition on the images in the first image group to identify the objects in the images.
[0235] The electronic device 100 can perform image recognition on the images contained in the first image group of the first video to determine the objects contained in the images in the first image group. Optionally, the electronic device 100 can determine potential sound-producing objects contained in the images in the first image group. The aforementioned potential sound-producing objects can represent objects that may emit sound. The aforementioned first image group can contain multiple frames of images. The multiple frames of images contained in the first image group can be multiple consecutive frames of images within a certain period of time in the first video.
[0236] For example, the first image group contains Figure 5B The image shown here. Figure 5BThe image shown is used as an example to illustrate the method of identifying images within images.
[0237] 100 pairs of electronic devices Figure 5B Image recognition of the image 610 shown can identify that the image 610 contains the following objects: person Ha611, person Hb612, cat 613, trash can 614, car 615, and tree 616. In some embodiments, person Ha611 and person Hb612 may represent two different people. The electronic device 100 can identify person Ha611, person Hb612, cat 613, and car 615 as potential sound-producing objects. Since trash can 614 and tree 616 typically do not produce sound, the electronic device 100 can determine that trash can 614 and tree 616 are not potential sound-producing objects.
[0238] Specifically, the electronic device 100 can utilize an image recognition model to identify objects contained in the images of the first image group. This image recognition model can be retrieved by the electronic device 100 from its local memory or from a server (e.g., a cloud server). The image recognition model can be a pre-trained model based on a neural network. When an image frame is input into the image recognition model, the model can output information such as the category, quantity, and location of objects contained in that image frame. This application does not limit the type of image recognition model or the training method of the image recognition model.
[0239] The electronic device 100 may store a category table of potential sound-producing objects. When an object contained in an image of a first image group is identified, the electronic device 100 can determine the potential sound-producing object contained in the image of the first image group according to the aforementioned category table of potential sound-producing objects.
[0240] Optionally, the image recognition model described above can be trained to recognize objects of a specified type. These specified types of objects can be the potential sound-producing objects described above. The electronic device 100 can use this image recognition model to determine the potential sound-producing objects contained in the images of the first image group.
[0241] In some embodiments, the electronic device 100 can also identify the gender of people in an image, such as men and women. Since the voiceprint characteristics of men and women are significantly different, the electronic device 100 can associate people in an image with audio components in the audio based on gender. For example, in image 610, person Ha611 is a woman, and person Hb612 is a man. If it is determined that the audio component played during the display time of image 610 is a female voice, the electronic device 100 can associate the female voice audio component with person Ha611 in image 610. If it is determined that the audio component played during the display time of image 610 is a male voice, the electronic device 100 can associate the male voice audio component with person Hb611 in image 610.
[0242] The embodiments of this application do not limit the method by which the electronic device 100 identifies objects contained in an image.
[0243] S520. Identify the sound-producing objects in the first audio and separate the audio components belonging to different sound-producing objects in the first audio.
[0244] The electronic device 100 can also perform audio recognition on the first audio in the first video to determine the audio components of different sound-producing objects in the first audio.
[0245] Specifically, the electronic device 100 can utilize an audio recognition model to identify audio components of different sound types in a first audio file. These sound types can include human speech, cat meows, dog barks, car sounds, rain sounds, thunder sounds, etc. The audio recognition model can be retrieved by the electronic device 100 from its local memory or from a server (e.g., a cloud server). The image recognition model can be a pre-trained model based on a neural network. The audio recognition model can be trained to recognize audio components of a specified sound type. When an audio segment is input into the audio recognition model, the model can input one or more audio components of different sound types contained in that audio segment.
[0246] An audio segment can be composed of one or more audio components. The time-domain signal of an audio segment can be obtained by adding the time-domain signals of the audio components that make up the audio segment. An audio component can represent the sound signal produced by a sound-producing object emitting sound over a period of time.
[0247] Through the above audio recognition, the electronic device 100 can obtain the audio components of different types of sound-producing objects in audio A.
[0248] In some embodiments, there may be multiple voices of the same type in an audio clip. For example, the first audio clip contains audio components produced by multiple people speaking. The voiceprint features of different voices are different. The aforementioned voiceprint features can represent the sound wave spectrum carried by the sound signal. The electronic device 100 can perform voiceprint feature recognition on the audio components corresponding to the voices in the first audio clip to distinguish the audio components of different people.
[0249] This section describes in detail an implementation method for separating audio components of different people, provided by an embodiment of this application.
[0250] Specifically, the electronic device 100 can slice audio A using a sliding window method to obtain multiple speech segments of audio A. The length of the window used for slicing can be any length, such as 0.5 seconds or 1 second. This embodiment does not limit this. The electronic device 100 can extract the voiceprint features of each speech segment. The extraction of voiceprint features can be the extraction of mel-frequency cepstral coefficients (MFCC) features. The electronic device 100 can detect the similarity of the voiceprint features of each speech segment to determine the number of speakers contained in audio A. The method for detecting the similarity of voiceprint features can include a similarity judgment method based on the Bayesian information criterion (BIC), a similarity judgment method based on the Akaike information criterion (AIC), etc. The electronic device 100 can cluster the speech segments according to the voiceprint features of each speech segment to obtain speech segments of different speakers in audio A. The clustering methods described above may include: clustering using a Gaussian mixture model (GMM), clustering using a support vector machine (SVM), clustering using a k-means clustering algorithm (k-means), clustering using a neural network-based clustering algorithm, etc. The electronic device 100 can splice together speech segments from a speaker to obtain the audio components of that speaker. Optionally, the electronic device 100 can also combine features such as the energy, zero-crossing rate, and formants of the speech segments to optimize the above clustering results, improving the accuracy of classifying speech segments from different speakers. This application embodiment does not limit the implementation method for separating the audio components of different speakers in the above-described audio.
[0251] As can be seen from the above implementation method, the electronic device 100 can separate audio components of different sound types from the first audio, and separate audio components of different people from the audio components of human voices. In this way, the electronic device 100 can determine the audio components of different voice-producing objects in the first audio.
[0252] From the above Figure 4 As shown in the embodiments, the first audio contains not only the audio components of the sound-producing objects but also a second audio component of the first background sound. After determining the audio components of all sound-producing objects in the first audio, the electronic device 100 can determine the remaining audio components in the first audio, excluding the audio components of the sound-producing objects, as the second audio component of the first background sound. The electronic device 100 can also extract the second audio component of the first background sound from the first audio through other implementation methods. This application embodiment does not limit this.
[0253] In some embodiments, the type of the sound-emitting object may only include people. The electronic device 100 can directly use the above method to separate the audio components of different people in the first audio, and determine the other audio components besides the audio components of the people as the second audio components of the first background sound.
[0254] For example, such as Figure 5C As shown, f(t) can represent the time-domain signal of the first audio. The electronic device 100 analyzes the first audio and determines that it contains two sound-producing objects: a first sound-producing object and a second sound-producing object. The electronic device 100 can separate the first audio into the audio component of the first sound-producing object, the audio component of the second sound-producing object, and the second audio component of the first background sound. Here, f1(t) can represent the time-domain signal of the audio component of the first sound-producing object. f2(t) can represent the time-domain signal of the audio component of the second sound-producing object. fn(t) can represent the time-domain signal of the second audio component of the first background sound. Where f(t) = f1(t) + f2(t) + fn(t).
[0255] The electronic device 100 can also determine the sound pressure of different sound-emitting objects. The sound pressure of a sound-emitting object can be obtained by integrating the time-domain signal of the audio component of that sound-emitting object. For example, the sound pressure p1 of the first sound-emitting object can be obtained by integrating f1(t). The sound pressure p2 of the second sound-emitting object can be obtained by integrating f2(t). The sound pressure pn of the first background sound can be obtained by integrating fn(t).
[0256] S530. By combining the objects in the image and the audio components in the first audio, the sound-emitting objects and the audio components of the sound-emitting objects contained in the first video within the first time period are determined.
[0257] After step S510, the electronic device 100 can determine one or more objects contained in the images of the first image group. After step S520, the electronic device 100 can determine the audio components contained in the first audio. Further, the electronic device 100 needs to associate the objects in the images with the audio components in the first audio to determine which object in the image is emitting sound, and which audio component is produced by the sound emitted by one object in the image. Thus, while the image displayed by the electronic device 100 presents a sounding object emitting sound, the corresponding audio playback device can play the audio component of that sounding object. This allows the video and sound to be synchronized in the user's perception.
[0258] This section describes in detail an implementation method for associating objects contained in an image with audio components, provided by an embodiment of this application.
[0259] Please refer to Figure 5D , Figure 5D An exemplary flowchart of a method for associating objects contained in an image with audio components is shown. This method can be applied to an electronic device 100. Specifically, the example described here is associating objects contained in an image in a first image group with audio components in a first audio file.
[0260] An audio component is generated by which object in the image makes a sound.
[0261] The method may include steps S531 to S538. Wherein:
[0262] S531. Determine whether the first sound-producing object indicated by the first audio component of the first audio exists in the image Ga1 displayed during the playback time of the first audio component.
[0263] Typically, during the process of a voice-speaking object speaking in a video, the video frame can contain an image showing the voice-speaking object speaking, and the audio in the video can contain the audio component of the voice-speaking object.
[0264] The playback time of the first audio component can be a period of time. Image Ga1 can be an image displayed by the electronic device 100 during the playback of the first audio component. Image Ga1 can contain one or more frames. Image Ga1 is an image in the first group of images mentioned above.
[0265] For example, if the first audio component indicates a person as the first source of sound, the electronic device 100 can determine whether the image Ga1 contains a person. If the first audio component indicates a cat as the first source of sound, the electronic device 100 can determine whether the image Ga1 contains a cat.
[0266] In one possible implementation, if the image Ga1 contains the aforementioned first sound-producing object, the electronic device 100 can directly associate the first sound-producing object in the image Ga1 with the first audio component. That is, the first audio component is generated by the sound produced by the first sound-producing object in the image Ga1. For example, the first audio component indicates that the first sound-producing object is a cat. The image Ga1 contains objects including cats. The electronic device 100 can associate the cat in the image Ga1 with the first audio component.
[0267] Given that the first video may be a video of a conversation between multiple people, the image Ga1 may contain multiple people, and the first audio may contain multiple audio components of human voices. The electronic device 100 needs to identify which person in the image is associated with a particular audio component of a human voice.
[0268] In another possible implementation, if the first sound-emitting object is present in the image Ga1, the electronic device 100 can perform the following step S532 to determine whether the first sound-emitting object is a person.
[0269] S532. If the first voice-emitting object exists in image Ga1, determine whether the first voice-emitting object is a person.
[0270] If the first sound source is a person, the electronic device 100 can perform the following step S533 to determine whether the image Ga1 contains multiple people.
[0271] If the first sound source is not a person, the electronic device 100 can perform the following step S535 to directly associate the first sound source in the image Ga1 with the first audio component.
[0272] S533. If the first voice is a person, determine whether the image Ga1 contains multiple people.
[0273] Specifically, the electronic device 100 can determine whether there are multiple figures in a single frame of image Ga1. If there are multiple figures in a single frame, the electronic device 100 may have difficulty directly determining which figure in the image is the first speaker. The electronic device 100 can perform the following step S534 to determine which figure in image Ga1 is speaking.
[0274] If image Ga1 contains only one person, electronic device 100 can directly determine that the first sound source is this one person contained in image Ga1. That is, the first audio component is produced by the sound of this one person. Then, electronic device 100 can perform the following step S535.
[0275] S534. If image Ga1 contains multiple people, perform face recognition on the multiple people in image Ga1, determine the facial movements of these multiple people, and identify person 1 among these multiple people as the one making a sound. Person 1 is the first person making a sound.
[0276] Specifically, the electronic device 100 can identify the facial regions of an image Ga1 containing multiple people to determine which person in the image Ga1 is speaking. Facial movements when speaking typically differ from facial movements when not speaking. The electronic device 100 can extract facial features of different people in the image Ga1 and, based on these features, identify the facial movements (such as mouth movements) of different faces to determine which person is speaking. The electronic device 100 can associate the person identified as speaking in the image Ga1 with the audio component (such as a first audio component) of the voice that is played at the same time as the display time of the image Ga1. It can be understood that when a person is identified as speaking in an image from a first video, the audio component of the voice played in the first video at the display time of that image is the audio component produced by that person speaking in the image.
[0277] In this way, the electronic device 100 can associate different people in an image with the audio components of different human voices in an audio file.
[0278] Steps S532 to S534 described above are optional. In some embodiments, if it is determined that the first sound-emitting object exists in the image Ga1, the electronic device 100 can directly execute step S535 below.
[0279] S535. Associate the first sound-emitting object in the image Ga1 with the first audio component. The first sound-emitting object is a sound-emitting object in the first video within the first time period.
[0280] Associating the first sound-producing object in image Ga1 with the first audio component indicates that the first audio component is produced by the sound produced by the first sound-producing object in image Ga1. That is, the first audio component is the audio component of the first sound-producing object in image Ga1.
[0281] S536. If the first sound-emitting object does not exist in image Ga1, determine whether the first sound-emitting object is contained in image Gb before the first time period in the first video.
[0282] In some embodiments, the first sound-emitting object does not exist in image Ga1, but it may exist in images prior to a first time period in the first video, such as image Gb. Image Gb may contain one or more frames.
[0283] If the image Gb contains a first sound-emitting object, the electronic device 100 may perform the following step S537 to associate the first sound-emitting object in the image Gb with a first audio component.
[0284] If the image Gb does not contain the first sound-emitting object, the electronic device 100 may perform the following step S538 to determine the first audio component as the second audio component of the first background sound in the first video within the first time period.
[0285] S537. If the image Gb contains a first sound object, associate the first sound object in the image Gb with a first audio component. The first sound object is a sound object of the first video in the first time period.
[0286] For example, when the first sound-emitting object is a cat, if image Ga1 does not contain a cat, but image Gb contains a cat, the electronic device 100 can associate the cat in image Gb with the first audio component. The sound-emitting objects in the first video during the first time period include the aforementioned cat. When it is necessary to determine the location of the sound-emitting object, the electronic device 100 can use image Gb to determine the location of the sound-emitting object.
[0287] This document exemplifies a method for associating audio components of a human voice with objects in an image, as provided in an embodiment of this application.
[0288] In one possible implementation, when the electronic device 100 matches the sound source in the audio with an object in the image, associating an audio component of a voice with a person in the image, the electronic device 100 can associate the voiceprint features of the audio component with the facial features of that person in the image. If, during subsequent video playback, an audio component with the same voiceprint features appears in the audio, but the corresponding image does not contain that person, the electronic device 100 can still determine the person associated with that audio component based on the facial features associated with the voiceprint features.
[0289] Understandably, a video may contain images and audio of multiple people conversing. One of these people may appear in the video for a period of time while speaking, and not appear in the video for another period, but the audio from both periods contains audio components generated by that person speaking. When the electronic device 100 associates the audio component of a person's voice with a person in an image at the same time, it can associate the voiceprint features of the audio component with the facial features of that person in the image. Therefore, during subsequent playback of the video, regardless of whether the person's face appears in the image, the electronic device 100 can determine that the audio component, which has voiceprint features associated with that person's facial features, is the audio component generated by that person speaking when separating it from the audio.
[0290] For example, the electronic device 100 can determine whether other audio components with the same voiceprint characteristics as the first audio component, such as audio component Fa3, have been associated with an object in the image. If audio component Fa3 has been associated with object 1 in image Gc, the electronic device 100 can associate the first audio component with object 1, which has been associated with audio component Fa3. That is, the first audio component is the audio component generated by object 1 emitting sound. It can be seen that even if object 1 does not exist in image Ga, the electronic device 100 can still determine the association between the first audio component and object 1. Object 1 is a sound-emitting object in the first video within the first time period. The position of object 1 can be determined based on the image Gc.
[0291] Steps S536 and S537 described above are optional. In some embodiments, if it is determined that the first sound-emitting object does not exist in the image Ga1, the electronic device 100 may directly execute step S538 below.
[0292] S538. If the image Gb does not contain the first sound-emitting object, the first audio component is determined as the second audio component of the first background sound in the first video within the first time period.
[0293] It is understandable that the aforementioned image Ga1 and the images displayed before Ga1 in the first video do not contain the first sound-emitting object, indicating that the first audio component failed to be associated with the sound-emitting object in the first video. The electronic device 100 can determine the first audio component as the second audio component of the first background sound in the first video within a first time period. Specifically, the electronic device 100 can mix the first audio component with the already separated second audio component of the first background sound (e.g., add the time-domain signals of the audio components). For example, the electronic device 100 separates the audio component of a dog barking sound from the first audio, but does not identify a dog from the images contained in the first image group. Therefore, the electronic device 100 can determine the audio component of the dog barking sound as an audio component contained in the second audio component of the first background sound.
[0294] Furthermore, if an object in the images contained in the first image group cannot be associated with an audio component in the first audio, it can be considered that the object did not produce sound during the display time of the images contained in the first image group, and is not a sound-producing object during that display time. For example, the electronic device 100 recognizes the presence of a tree in the images contained in the first image group, but fails to separate the audio component of the tree from the first audio. In this case, the tree in the image is not a sound-producing object.
[0295] The first video may include more sound-producing objects within the first time period. The method by which the electronic device 100 determines the audio components of these sound-producing objects can refer to the method described above for determining the first audio component of the first sound-producing object. It will not be repeated here.
[0296] The audio components of the aforementioned sound-producing objects can have temporal attributes. These temporal attributes can be determined by the position of the audio component on the video's timeline. The temporal attribute of an audio component can be used to determine its playback time. For example, if an audio component is the audio component of the first time segment on the video's timeline, then the audio playback device can play this audio component within the first time segment of the video playback. By utilizing the temporal attributes of the audio components, the electronic device 100 can improve the synchronization between audio playback and video playback, reducing situations where audio components from different sound-producing objects are ahead or behind during playback.
[0297] From the above Figure 5A As shown in the method, the electronic device 100 can analyze the images and audio in the first video separately to determine the objects in the images and the sound-producing objects in the audio. The electronic device 100 can determine whether the images and audio corresponding to the same time in the first video contain the same object. The images and audio corresponding to the same time in the first video can represent images displayed and audio played at the same time. If the images and audio corresponding to the same time in the first video contain the same object, such as object 1, the electronic device 100 can associate the audio component of object 1 in the audio with the image of object 1. For example, if the images and audio corresponding to the same time in the first video contain a car and an audio component of a car, the electronic device 100 can determine the audio component of the car in the image.
[0298] The following describes in detail another method for determining the sound-producing object contained in a video and the audio components of the sound-producing object, provided by an embodiment of this application.
[0299] Please refer to Figure 5E , Figure 5E An exemplary flowchart illustrates another method for determining the sound-producing object and its audio components within a video. This method can be applied to an electronic device 100. Specifically, it is illustrated here using the determination of the sound-producing object and its audio components within a first time period in a first video.
[0300] The method may include steps S551 to S554. Wherein:
[0301] S551. Identify the sound-producing objects in the first audio and separate the audio components belonging to different sound-producing objects in the first audio.
[0302] Step S551 can be referred to the above. Figure 5A Step S520 in the process will not be described in detail here.
[0303] S552. Identify whether the images in the first image group contain the first sounding object among the aforementioned sounding objects.
[0304] In one possible implementation, the electronic device 100 can identify whether an image in the first image group contains a corresponding sound-producing object based on the sound-producing object in the first audio. For example, the electronic device 100 separates the audio component of a cat's meow from the first audio. Based on this audio component of the cat's meow, the electronic device 100 can identify whether a cat exists in the images of the first image group.
[0305] S553. If the images in the first image group contain a first sound-emitting object, associate the first audio component in the first audio with the first sound-emitting object in the image. The first sound-emitting object is a sound-emitting object in the first video within the first time period.
[0306] The first image group and the first audio are data from the first video within the same time period (i.e., the first time period). If the first audio contains a first audio component of the first sound-emitting object, and the images in the first image group contain the first sound-emitting object, then the electronic device 100 can determine that the first audio component is the audio component generated by the first sound-emitting object emitting sound within the first time period.
[0307] In some embodiments, the electronic device 100 determines that the first sound-emitting object is a person based on a first audio component of the first sound-emitting object. The images in the first image group contain multiple people. The electronic device 100 can determine the first sound-emitting object as described above. Figure 5D Steps S534 and S535 show how to determine which of the aforementioned multiple characters is the first vocal subject, and then associate the determined character with audio component 1.
[0308] S554. If the images in the first image group do not contain the first sound-emitting object, the first audio component in the first audio is determined as the second audio component of the first background sound in the first video within the first time period.
[0309] The fact that the images in the first image group do not contain the first sound-producing object indicates that the first audio component representing the sound-producing object can be associated with the object contained in the images in the first image group. The electronic device 100 can determine the first audio component of the first sound-producing object in the first audio as the second audio component of the first background sound in the first video within a first time period. The electronic device 100 can then mix the first audio component with the separated second audio component of the first background sound (e.g., add the time-domain signals of the audio components).
[0310] Optionally, if the images in the first image group do not contain the first sound-emitting object, the electronic device 100 can further determine whether the images of the first video before the first time period contain the first sound-emitting object. If the images of the first video before the first time period contain the first sound-emitting object, the electronic device 100 can associate the first audio component of the first sound-emitting object with the first sound-emitting object contained in the images of the first video before the first time period. For specific implementation methods, please refer to the foregoing. Figure 5D Step S537 is shown. It will not be described again here.
[0311] From the above Figure 5E As shown in the method, the electronic device 100 can analyze the audio in the first video to determine the sound-producing object and its audio components. Based on the sound-producing object determined from the audio, the electronic device 100 can identify whether the sound-producing object exists in an image played at the same time as the audio in the first video. If it exists, the electronic device 100 can associate the sound-producing object in the image with the corresponding audio components in the audio.
[0312] As can be seen, in the above implementation method, the electronic device 100 can perform image recognition on the video image based solely on the sound-producing object in the audio, without needing to recognize all objects in the image. This saves the computing resources of the electronic device 100 and improves the efficiency of determining the sound-producing object in the video using the matching relationship between the image and audio.
[0313] The following describes in detail another method for determining the sound-producing object contained in a video and the audio components of the sound-producing object, provided by an embodiment of this application.
[0314] Please refer to Figure 5F , Figure 5F An exemplary flowchart illustrates another method for determining the sound-producing object and its audio components within a video. This method can be applied to an electronic device 100. Specifically, it is illustrated here using the determination of the sound-producing object and its audio components within a first time period in a first video.
[0315] The method may include steps S561 to S564. Wherein:
[0316] S561. Perform image recognition on the images in the first image group to identify the objects in the images.
[0317] Step S561 can be referred to the above. Figure 5A Step S510 in the process will not be described in detail here.
[0318] S562. Identify whether the audio component of object 1 in the above-mentioned objects exists in the first audio.
[0319] In one possible implementation, the electronic device 100 can identify whether the first audio contains a corresponding object based on the objects contained in the images in the first image group. For example, the electronic device 100 identifies that the images in the first image group contain a cat. Based on the cat in the image, the electronic device 100 can identify whether there is an audio component of a cat's meow in the first audio. Specifically, the electronic device 100 can use an audio recognition model for cat meows to determine whether there is an audio component of a cat's meow in the first audio. If there is an audio component of a cat's meow in the first audio, the electronic device 100 can use the aforementioned audio recognition model for cat meows to separate the audio component of the cat's meow from the first audio.
[0320] S563. If there is an audio component of object 1 in the first audio, associate object 1 in the image with the audio component of object 1 in the first audio, and object 1 is a sounding object in the first video within the first time period.
[0321] For example, if the object 1 is a cat and the first audio contains an audio component of a cat's meow, the electronic device 100 can associate the cat contained in the images in the first image group with the audio component of the cat's meow in the first audio.
[0322] In some embodiments, if the electronic device 100 identifies that the images in the first image group contain multiple people, the electronic device 100 can further determine which of these multiple people is speaking based on the image. Specific methods can be found in the foregoing. Figure 5D Step S534 of the method shown. The electronic device 100 can identify whether there is a human voice audio component in the audio played during the display time of the image showing the person speaking, using the speaking person as the source of sound in the image. If a human voice audio component is present, the electronic device 100 can associate the person in the image with the human voice audio component. Wherein, according to the aforementioned step S534, if a person in the image who is not speaking is determined not to be a source of sound during the image's display time, the electronic device 100 does not need to identify the audio component in the audio played during the image's display time using the aforementioned non-speaking person.
[0323] In other words, when the object 1 is a person and the images in the first image group contain multiple people, the electronic device 100 can identify the person speaking from the images in the first image group alone, and recognize the audio component of the human voice associated with the person speaking in the audio.
[0324] S564. If the audio component of object 1 is not present in the first audio, determine that object 1 in the image did not make a sound during the first time period of the first video.
[0325] For example, electronic device 100 identifies that an image in the first image group contains a dog. Electronic device 100 can use an audio recognition model of dog barking to determine whether there is an audio component of dog barking in the first audio. If there is no audio component of dog barking in the first audio, electronic device 100 can determine that the dog contained in the image Ga did not bark during the first time period of the first video.
[0326] After traversing all objects contained in the images of the first image group according to steps S562-S564 above, the electronic device 100 can associate one or more objects contained in the images of the first image group with audio components in the first audio. In this way, the electronic device 100 can determine the sound-producing object and its audio component within the first time period of the first video. Furthermore, the electronic device 100 can determine the audio components in the first audio other than the audio components of all sound-producing objects as the second audio component of the first background sound.
[0327] From the above Figure 5F As shown in the method, electronic device 100 can analyze images in a first video to determine objects present in the images. Based on the objects determined from the images, electronic device 100 can identify whether the objects in the images exist in audio played at the same time as the images in the first video. If they exist, electronic device 100 can associate the audio components in the audio with the corresponding objects in the images.
[0328] As can be seen, in the above implementation method, the electronic device 100 can perform audio recognition based solely on objects in the images contained in the video, without needing to identify all sound-producing objects in the audio and separate the audio components of all sound-producing objects from the audio. This can save the computing resources of the electronic device 100 and improve the efficiency of using the matching relationship between images and audio to determine the sound-producing objects in the video.
[0329] In addition to separating audio components (such as the first audio component of the first sound source and the second audio component of the first background sound) from the first audio contained in the first video within the first time period, the electronic device 100 also needs to determine which audio playback device will play the audio component. Specifically, the electronic device 100 can determine the correspondence between audio playback devices and sound sources. An audio playback device can play the audio component corresponding to its own sound source. The electronic device 100 can also select one or more audio playback devices from all available audio playback devices to play the second audio component of the first background sound.
[0330] In one possible implementation, the electronic device 100 can determine the correspondence between the audio playback device and the sound-producing object based on the position of the sound-producing object, the position of the audio playback device, and the position of the observer.
[0331] The following describes in detail an implementation method for determining the position of a sound-emitting object in a video, provided by an embodiment of this application.
[0332] Please refer to Figure 6A , Figure 6A An exemplary diagram illustrates a method for determining the location of a sound-emitting object in a video. Specifically, this explanation focuses on determining the location of a sound-emitting object within a first time period in a first video.
[0333] The electronic device 100 can input the images from the first image group into the three-dimensional object reconstruction model to obtain the position of the sound-emitting object in the image and the position of the virtual camera that captured the image. The images input into the three-dimensional object reconstruction model may include one or more frames. The position of the sound-emitting object and the position of the virtual camera are the same as the position of the sound-emitting object and the position of the virtual camera in the first video within the first time period.
[0334] The aforementioned 3D object reconstruction model can be used to perform 3D reconstruction of the images in the first image group. Specifically, capturing an object in 3D space yields a 2D image. According to the pinhole imaging theorem, the position of an object in 3D space corresponds to its position in the 2D image. Therefore, the aforementioned 3D image reconstruction can be described as using a 2D image of a 3D scene or object as the base data, processing this base data to obtain the 3D data of the scene or object, thereby generating a three-dimensional scene or object.
[0335] The aforementioned 3D object reconstruction model can be a neural network-based model, capable of reconstructing a 3D model from a 2D image. The 3D object reconstruction model can be trained to reconstruct a 3D object from an image of a specified type, along with the virtual camera that captured the image. One or more frames of an image are input into the 3D object reconstruction model, which can then input the positions of one or more objects in the 3D space, as well as the position of the virtual camera in the 3D space. This application does not limit the training method for the aforementioned 3D object reconstruction model.
[0336] The aforementioned three-dimensional object reconstruction model can be obtained by the electronic device 100 from its local memory or from a server (e.g., a cloud server). This application does not limit this.
[0337] In one possible implementation, the electronic device 100 can input a frame of image contained in the first image group into the three-dimensional object reconstruction model to obtain the position of the sound-emitting object in that frame and the position of the virtual camera that captured that frame. The first image group may contain multiple frames. The electronic device 100 can input the multiple frames contained in the first image group into the three-dimensional object reconstruction model frame by frame in chronological order of display time to determine the position of the sound-emitting object in each frame and the position of the virtual camera that captured the corresponding frame. This improves the accuracy of locating the sound-emitting object and the virtual camera in the video and reduces errors caused by changes in the position of the sound-emitting object or the virtual camera. Optionally, the electronic device 100 can also select one frame from the multiple frames contained in the first image group every preset number of frames. The electronic device 100 can input the selected image into the three-dimensional object reconstruction model to determine the position of the sound-emitting object in the selected image and the position of the virtual camera that captured the selected image. It is understood that the frame rate of the video can be, for example, 24 frames per second, 60 frames per second, etc. The positions of the sound-emitting object and the virtual camera in the video typically remain unchanged for a short period (e.g., within two consecutive frames or five consecutive frames). The electronic device 100 determining the positions of the sound-emitting object and the virtual camera at preset intervals does not cause significant errors. This saves the computing resources of the electronic device 100 and reduces the computational demands on it.
[0338] Electronic device 100 uses a 3D object reconstruction model to determine the position of the sound-emitting object in a frame of an image and the position of the virtual camera that captured that frame of the image, as follows: Figure 6B As shown.
[0339] For example, please refer to Figure 6B The coordinate system Ow-Xw-Yw-Zw can be the world coordinate system. The 3D object reconstruction model can use this world coordinate system as the reference coordinate system to describe the position of the sound-emitting object and the position of the virtual camera. This application does not limit the method for establishing the coordinate system Ow-Xw-Yw-Zw. The coordinate system Ow-Xw-Yw-Zw can be a left-handed coordinate system or a right-handed coordinate system. This application uses a left-handed coordinate system as an example for explanation.
[0340] For example, the aforementioned frame image may contain four sound-emitting objects. In the coordinate system Ow-Xw-Yw-Zw, the position of the first sound-emitting object can be (x_a, y_a, z_a). The position of the second sound-emitting object can be (x_b, y_b, z_b). The position of the third sound-emitting object can be (x_c, y_c, z_c). The position of the fourth sound-emitting object can be (x_d, y_d, z_d). The position of the virtual camera can be (x_e, y_e, z_e).
[0341] In one possible implementation, the electronic device 100 can project the positional relationship between the sound-emitting object and the virtual camera from a three-dimensional coordinate system to a two-dimensional coordinate system. For example, the electronic device 100 can project along the Zw axis of the aforementioned coordinate system Ow-Xw-Yw-Zw, and establish a coordinate system with the virtual camera's position as the origin (0,0). Figure 6C The coordinate system shown is Xc-Oc-Yc. The Yc axis of the coordinate system Xc-Oc-Yc can be a straight line parallel to the optical axis of the virtual camera, and the Xc axis can be a straight line perpendicular to the Yc axis. This application does not limit the method for determining the Yt axis and Xt axis in the coordinate system Xc-Oc-Yc.
[0342] Based on the principle of projecting from a three-dimensional coordinate system to a two-dimensional coordinate system, the electronic device 100 can determine the position of the aforementioned sound-emitting object in the coordinate system Xc-Oc-Yc. For example, in the coordinate system Xc-Oc-Yc, the position of the first sound-emitting object can be (x5, y5). The position of the second sound-emitting object can be (x6, y6). The position of the third sound-emitting object can be (x7, y7). The position of the fourth sound-emitting object can be (x8, y8).
[0343] It can be seen that the above Figure 6B and Figure 6C The position of the sound-emitting object shown is its position in the second scene (i.e., the virtual scene), and the position of the virtual camera is its position in the second scene.
[0344] Both the location of the aforementioned sound-emitting object and the location of the virtual camera can have temporal attributes. The temporal attributes of the sound-emitting object's location and the virtual camera's location can be determined by the position of the image at that location on the video's timeline. The temporal attribute of the virtual camera's location can be used to determine the duration of the virtual camera's position at that location. The temporal attribute of the sound-emitting object's location can be used to determine the duration of the sound-emitting object's position at that location. It is understood that if the temporal attribute of the sound-emitting object's location matches the temporal attribute of the sound-emitting object's audio component, that audio component can be the audio component generated by the sound-emitting object speaking at that location. Alternatively, if the audio component of the sound-emitting object is associated with a sound-emitting object in the image, that audio component can be the audio component generated by the sound-emitting object speaking at a location determined based on the image.
[0345] Because the positions of the virtual camera and the sound-producing object may change, the electronic device 100 can adjust the audio playback device simulating the sound-producing object's voice in real time, taking into account the aforementioned time attributes. This allows the user to more realistically experience the process of the sound-producing object in the video making a sound from different positions, following the perspective of the virtual camera.
[0346] From the above Figures 6A to 6C As can be seen from the embodiments shown, the electronic device 100 can determine the position of the virtual camera and the position of the sound-emitting object in the first video in real time.
[0347] The following describes in detail an implementation method for obtaining the location of an audio playback device provided by an embodiment of this application.
[0348] In one possible implementation, the electronic device 100 can determine the location of the audio playback device using ultrasonic positioning. Both the electronic device 100 and the audio playback device can have an audio output device (such as a speaker) and an audio input device (such as a microphone). The audio output device can emit ultrasonic waves. The audio input device can receive ultrasonic waves.
[0349] It's understandable that the frequency of ultrasound exceeds the highest threshold of human hearing, 20,000 Hz. This means humans cannot perceive ultrasound in the environment. Therefore, the use of ultrasound for positioning by electronic device 100 will not affect the user's ability to hear other sounds in the environment.
[0350] Specifically, electronic device 100 can acquire the direction and distance of multiple audio playback devices relative to itself. Here, we will use acquiring the position of audio playback device 200 as an example for explanation.
[0351] Please refer to Figure 6D , Figure 6D An exemplary flowchart of a method for obtaining the location of an audio playback device is shown.
[0352] The method may include steps S611 to S615. Wherein:
[0353] S611, the electronic device 100 sends an ultrasonic wave transmission command to the audio playback device 200, the ultrasonic wave transmission command being used to instruct the audio playback device 200 to transmit ultrasonic waves.
[0354] Electronic device 100 can send an ultrasonic wave transmission command to audio playback device 200 via its communication connection with audio playback device 200. This ultrasonic wave transmission command can be used to instruct audio playback device 200 to emit ultrasonic waves.
[0355] S612, Audio playback device 200 emits ultrasonic waves.
[0356] When the ultrasonic wave transmission command is received, the audio playback device 200 can transmit ultrasonic waves.
[0357] S613. Electronic device 100 receives ultrasonic waves from audio playback device 200 and determines the direction of audio playback device 200 based on the ultrasonic waves.
[0358] Electronic device 100 can receive ultrasonic waves from audio playback device 200. The audio input device of electronic device 100 may include a microphone array. A microphone array can be understood as multiple microphones distributed according to a specified rule (such as three rows of three columns, five rows of five columns, etc.). Electronic device 100 can obtain the direction of the ultrasonic waves by the time difference between the ultrasonic waves received by the multiple microphones in the microphone array. The direction of the ultrasonic waves is the direction of audio playback device 200. Thus, electronic device 100 can obtain the direction of audio playback device 200 relative to itself. For the specific implementation process of obtaining the direction of the audio playback device through the microphone array, please refer to Chinese invention patent application No. 202011556351.2. It will not be repeated here. It should be noted that all content related to positioning in Chinese invention patent application No. 202011556351.2 is incorporated into this application and is within the scope of this application.
[0359] S614, Electronic device 100 emits ultrasonic waves.
[0360] S615. The electronic device 100 obtains the distance between the audio playback device 200 and the audio playback device 200 based on the time difference between the time of emitting the differential sound wave and the time of receiving the reflected ultrasonic wave in the direction of the audio playback device 200.
[0361] Electronic device 100 can emit ultrasonic waves and receive the reflected ultrasonic waves. Electronic device 100 can determine the distance between itself and audio playback device 200 by the time ts of emitting the ultrasonic wave and the time tr of receiving the ultrasonic wave reflected from the direction of audio playback device 200. The distance between audio playback device 200 and electronic device 100 can be C*(tr-ts) / 2, where C is the speed of ultrasonic wave propagation. In air, C can be 340 meters per second.
[0362] Understandably, during the process of the electronic device 100 emitting ultrasonic waves, the audio playback device 200 can stop emitting ultrasonic waves. The ultrasonic waves received by the electronic device 100 are the ultrasonic waves after the ultrasonic waves it emitted are reflected. This can avoid interference from other devices emitting ultrasonic waves to the ultrasonic positioning of the electronic device 100. In one possible implementation, the electronic device 100 can send a command to the audio playback device 200 to stop emitting ultrasonic waves before emitting ultrasonic waves. Upon receiving the command to stop emitting ultrasonic waves, the audio playback device 200 can stop emitting ultrasonic waves. In another possible implementation, after receiving the ultrasonic wave emission command in step S611, the audio playback device 200 can emit ultrasonic waves within a preset duration (e.g., 1 second, 2 seconds, etc.). When it stops emitting ultrasonic waves, the audio playback device 200 can send an ultrasonic wave stop message to the electronic device 100. The electronic device 100 can then resume emitting ultrasonic waves after receiving the ultrasonic wave stop message. This application does not limit the method for implementing the audio playback device 200 to stop emitting ultrasonic waves.
[0363] Similarly, electronic device 100 can be achieved through the above... Figure 6D The ultrasonic positioning method shown obtains the direction and distance of other audio playback devices besides audio playback device 200 relative to itself.
[0364] In another possible implementation, electronic device 100 can emit ultrasonic waves before an audio playback device is placed and receive the reflected ultrasonic waves. Then, electronic device 100 can emit ultrasonic waves again after the audio playback device is placed and receive the reflected ultrasonic waves. Electronic device 100 can determine the location of the audio playback device by comparing the reflected ultrasonic waves received before and after its placement.
[0365] In another possible implementation, the electronic device 100 may have an image acquisition device (such as a camera). The electronic device 100 can acquire an image containing the audio playback device. The electronic device 100 can perform image recognition on the image to obtain the model number of the audio playback device, thereby obtaining the actual size of the audio playback device. Based on the position of the audio playback device in the image and the ratio of the actual size of the audio playback device to its size in the image, the electronic device 100 can determine the location of the audio playback device.
[0366] Optionally, the electronic device 100 can also capture images of the audio playback device at different focal lengths. The electronic device 100 can identify the location of the audio playback device by recognizing images captured by the audio playback device at different focal lengths. Furthermore, after acquiring the image, the electronic device 100 can also instruct other devices (such as video analysis devices) to perform image recognition. This application embodiment does not limit this aspect.
[0367] Not limited to the positioning technologies mentioned in the above embodiments, the electronic device 100 can also obtain the location of each audio playback device through positioning technologies such as millimeter-wave radar positioning and UWB positioning. The specific implementation process of locating audio playback devices using millimeter-wave radar positioning and UWB positioning can be found in Chinese invention patent application No. 202111243798.9. It will not be repeated here. It should be noted that all content related to positioning in Chinese invention patent application No. 202111243798.9 is incorporated into this application and is within the scope of this application.
[0368] The following describes in detail the implementation method for obtaining the position of the observer provided in the embodiments of this application.
[0369] In one possible implementation, the electronic device 100 can obtain the observer's position using ultrasonic positioning. Specifically, the electronic device 100 can emit ultrasonic waves and receive the reflected ultrasonic waves. The electronic device 100 can identify a target with a human-like shape (such as a standing human-like shape, a sitting human-like shape, etc.) from the reflected ultrasonic waves. This target can be considered the observer. The electronic device 100 can obtain the direction and position of the target relative to the electronic device 100 based on the ultrasonic waves corresponding to the aforementioned human-like target in the reflected ultrasonic waves. In this way, the electronic device 100 can obtain the observer's position.
[0370] In another possible implementation, electronic device 100 can acquire an image containing the photographer. Electronic device 100 can identify the positional relationship between the photographer and one or more audio playback devices in the image. Furthermore, electronic device 100 can combine the positions of the one or more audio playback devices to obtain the observer's position. The specific implementation process of locating the observer using millimeter-wave radar and UWB positioning can be found in Chinese invention patent application number 202111243798.9. It will not be repeated here. It should be noted that all content related to positioning in Chinese invention patent application number 202111243798.9 is incorporated into this application and is within the scope of this application.
[0371] The number of observers mentioned above can be one or more.
[0372] This application does not limit the method by which the electronic device 100 obtains the observer's location. For example, the electronic device 100 can also obtain the observer's location through positioning technologies such as millimeter-wave radar positioning, UWB positioning, and infrared positioning.
[0373] In some implementations, once the direction and distance of the audio playback device and the observer relative to the electronic device 100 are obtained, the electronic device 100 can establish a coordinate system with the observer's location as the origin to determine the positional relationship between the electronic device 100, each audio playback device, and the observer.
[0374] Please refer to Figure 6E , Figure 6E This example illustrates the positions of the audio playback devices and electronic devices 100 in a coordinate system established with the observer's position as the origin. Four audio playback devices are used as an example for illustration.
[0375] like Figure 6E As shown, the origin (0,0) of the coordinate system Xt-Ot-Yt represents the observer's position. The Yt axis can be a straight line parallel to the orientation of the display screen of the electronic device 100, and the Xt axis can be a straight line perpendicular to the Yt axis. This application does not limit the method for determining the Yt axis and Xt axis in the coordinate system Xt-Ot-Yt.
[0376] It should be noted that, for ease of explanation and for simplicity, a two-dimensional coordinate system is used as an example here. Those skilled in the art will understand that a three-dimensional coordinate system is similar, and will not be described in detail here.
[0377] Electronic device 100 can determine its own position and the positions of each audio playback device in the coordinate system Xt-Ot-Yt based on the direction and distance of each audio playback device and the observer relative to itself. For example, in the coordinate system Xt-Ot-Yt, the coordinates of electronic device 100 are (0, y0). The coordinates of audio playback device 200 are (x1, y1). The coordinates of audio playback device 201 are (x2, y2). The coordinates of audio playback device 202 are (x3, y3). The coordinates of audio playback device 203 are (x4, y4).
[0378] It can be seen that, Figure 6E The observer's position shown is the observer's position in the first scene (i.e., the real scene), and the audio playback device's position is the audio playback device's position in the first scene.
[0379] In some embodiments, the electronic device 100 can determine whether all audio playback devices are located on its side. If it is determined that all audio playback devices are located on one side of the electronic device 100, the electronic device 100 can prompt the user to adjust the position of the audio playback devices so that audio playback devices are present on both sides of the electronic device 100, and then re-acquire the position of the audio playback devices.
[0380] Typically, an observer faces the display screen of electronic device 100 when watching a video. To enhance the observer's immersion in the video, multiple audio playback devices can play audio from the video such that the sound of some objects reaches the observer from their left side, while the sound of others reaches the observer from their right side. Therefore, with audio playback devices distributed on both the left and right sides of electronic device 100, the audio playback devices can better simulate the sound production process of the objects in the video.
[0381] In some embodiments, based on the positions of the audio playback devices and the observer obtained from the foregoing embodiments, the electronic device 100 can determine whether all audio playback devices are located on its side. For example, the electronic device 100 can compare... Figure 6EIn the coordinate system shown, the coordinates of audio playback devices 200-203 and electronic device 100 are represented on the Xt axis. If the Xt axis coordinates of audio playback devices 200-203 are all less than or all greater than the Xt axis coordinates of electronic device 100, then electronic device 100 can determine that audio playback devices 200-203 are all located on one side of electronic device 100. If it is determined that all audio playback devices are located on one side of electronic device 100, electronic device 100 can prompt the user to adjust the position of the audio playback devices. For example, the user can move some of the audio playback devices to the side of electronic device 100 where no audio playback devices are located.
[0382] The electronic device 100 can reacquire the location of the audio playback device. Optionally, the electronic device 100 can reacquire only the location of the audio playback device whose location has changed.
[0383] This application embodiment does not limit the time for the electronic device 100 to acquire the location of each audio playback device.
[0384] For example, when a user configures electronic device 100 and one or more audio playback devices at home, electronic device 100 can obtain the location of the aforementioned audio playback devices. As another example, when receiving an operation to play a video, electronic device 100 can obtain the location of the audio playback devices and the observer before playing the video. During video playback, electronic device 100 can periodically or intermittently obtain the location of the audio playback devices and the observer. As yet another example, when receiving an operation to play a video, electronic device 100 can provide the user with video playback options: normal playback mode and stereo playback mode. The aforementioned normal playback mode can mean that electronic device 100 does not distinguish the audio components of different sound-producing objects in the video and distributes them to different audio playback devices for playback. During a video playback process, all audio playback devices can play the same audio. The aforementioned stereo playback mode can mean that electronic device 100 distributes the audio components of different sound-producing objects to different audio playback devices for playback according to the method of coordinated audio playback in video playback provided in this application. When receiving an operation from the user to select the aforementioned stereo playback mode, electronic device 100 can obtain the location of the audio playback devices and the observer before playing the video.
[0385] The following describes in detail an implementation method for determining the positional relationship between a sound-emitting object, an audio playback device, and an observer, provided by an embodiment of this application.
[0386] In one possible implementation, electronic device 100 can determine the position of a virtual camera in a second scene based on the observer's position in the first scene, thus obtaining the positional relationship between the sound-emitting object, the audio playback device, and the observer in the first scene. For example, electronic device 100 can... Figure 6C The coordinate system Xc-Oc-Yc shown is Figure 6E The coordinate system Xt-Ot-Yt shown coincides, thus obtaining Figure 6F The virtual camera and the sound-emitting object are shown in the coordinate system Xt-Ot-Yt.
[0387] like Figure 6F As shown, the origin of the coordinate system Xt-Ot-Yt can be the observer's position in the first scene, which is also the position of the virtual camera in the second scene. The positions of the first to fourth sound-emitting objects and the audio playback devices 200 to 203 in this two-dimensional coordinate system can be referred to the description in the foregoing embodiment. It is not limited to... Figure 6F The first video may contain more or fewer sound-producing objects and audio playback devices, as shown. The audio playback devices used to play the audio in the first video may be more or fewer in number.
[0388] As can be seen from the above embodiments, the position of the sound-emitting object in the first scene can be obtained based on the positional relationship between the sound-emitting object and the virtual camera in the second scene, and by determining the position of the observer in the first scene as the position of the virtual camera in the second scene.
[0389] It is understood that the positions of the audio playback devices and the sound-emitting objects illustrated in the above embodiments are merely illustrative examples and should not be construed as limiting this application.
[0390] The positional relationship between the sound-emitting object, the audio playback device, and the observer mentioned in the embodiments of this application can refer to the positional relationship between the sound-emitting object, the audio playback device, and the observer in the first scene.
[0391] In the process of determining the positional relationship between the sound source, audio playback device, and observer, setting the observer's position in the first scene as the position of the virtual camera in the second scene allows the user to stand in the virtual camera's perspective, change with the camera's position, and experience the three-dimensional scene presented by the video screen in an immersive way, feeling different sound sources making sounds from different directions.
[0392] The electronic device 100 is not limited to determining the position of the virtual camera in the second scene as the position of the observer in the first scene. It can also determine the position of the virtual camera in the second scene as a position that has a mapping relationship with the observer's position in the first scene. For example, the position that has a mapping relationship with the observer's position in the first scene could be a position 1 meter away from the observer's position in the first direction. This application embodiment does not specifically limit the above mapping relationship.
[0393] This application embodiment does not limit the time at which the electronic device 100 acquires the positions of the audio playback devices, the observer, the sound source, and the virtual camera. For example, the electronic device 100 can acquire the positions of each audio playback device when it first establishes a communication connection with each audio playback device. When a new audio playback device is added, the electronic device 100 can establish a communication connection with this audio playback device and acquire its position. Since the positions of the audio playback devices change relatively infrequently, the electronic device 100 can store the positions of each audio playback device. When the electronic device 100 detects a change in the position of an audio playback device, it can reacquire the position of the audio playback device and update it in memory. When the position of the audio playback device is needed during video playback, the electronic device 100 can retrieve the position of the audio playback device from memory. In this way, the electronic device 100 does not need to calculate the position of the audio playback device every time it determines the audio component played by the audio playback device during video playback. The electronic device 100 can calculate the positions of the observer, the sound source in the video, and the virtual camera in real time during video playback.
[0394] Once the positional relationship between the sound-producing object, the audio playback device, and the observer is obtained, the electronic device 100 can determine the correspondence between the audio playback device and the sound-producing object.
[0395] The following describes in detail an implementation method for determining the correspondence between an audio playback device and a sound-producing object, provided by an embodiment of this application.
[0396] The closer an audio playback device is to a sound-producing object, the better it simulates that object's sound. This is understandable, as humans can use their two ears to pinpoint the location of a sound source. The closer an audio playback device is to a sound-producing object, the easier it is for the user to perceive the sound as originating from that object's location by adjusting the playback parameters. In other words, simulating a sound-producing object is equivalent to the audio playback device being physically located at the object's location while playing its audio components.
[0397] Please refer to Figure 7A , Figure 7A An exemplary flowchart illustrates a method for determining the correspondence between an audio playback device and a sound-producing object. The explanation specifically focuses on determining the audio playback device corresponding to the first sound-producing object.
[0398] This method may include steps S711 to S717. Wherein:
[0399] S711. Based on the angle between the audio playback device and the observer's line, and between the first sound source and the observer's line L1, select the audio playback device on the line with the smallest angle to line L1.
[0400] S712. Determine whether there is only one audio playback device selected in step S711.
[0401] S713. If only one audio playback device is selected in step S711, then the audio playback device selected in step S711 is instructed to play the first audio component of the first sound-producing object. The selected audio playback device corresponds to the first sound-producing object.
[0402] S714. If there are multiple audio playback devices selected in step S711, then select the audio playback device with the smallest distance from the first sound source from the audio playback devices selected in step S711.
[0403] S715. Determine whether there is only one audio playback device selected in step S714.
[0404] S716. If only one audio playback device is selected in step S714, then the audio playback device selected in step S714 is instructed to play the first audio component of the first sound-producing object. The selected audio playback device corresponds to the first sound-producing object.
[0405] In steps S717 and S714, multiple audio playback devices are selected. One idle audio playback device is selected from the audio playback devices selected in step S714, or one audio playback device with the fewest corresponding sound-producing objects is selected. The audio playback device selected in step S717 is instructed to play the first audio component of the first sound-producing object. The audio playback device selected in step S717 corresponds to the first sound-producing object.
[0406] The aforementioned idle audio playback device can refer to an audio playback device that does not play other audio components during the playback time of the first audio component of the first sound-emitting object.
[0407] Steps S712 to S716 described above are optional. In some embodiments, when there are multiple lines with the smallest angle to line L1, the electronic device 100 can select multiple audio playback devices. The electronic device 100 can select one audio playback device from these multiple audio playback devices according to step S717 described above. Alternatively, step S717 described above is also optional. When the electronic device 100 selects multiple audio playback devices in step S711 described above, the electronic device 100 can arbitrarily select one audio playback device from these multiple electronic devices to play the first audio component of the first sound source.
[0408] In some embodiments, when the electronic device 100 selects a plurality of audio playback devices in step S714 above, the electronic device 100 may arbitrarily select one of the plurality of electronic devices to play the first audio component of the first sound source.
[0409] In some embodiments, the electronic device 100 can directly select the audio playback device corresponding to the first sound-emitting object based on the distance between the first sound-emitting object and each audio playback device. Specifically, the electronic device 100 can calculate the distance between the first sound-emitting object and each audio playback device. The electronic device 100 can select the audio playback device with the shortest distance to the first sound-emitting object to play the first audio component of the first sound-emitting object.
[0410] Here Figure 7B The diagram shown illustrates the method for determining the audio playback device corresponding to the first sound-producing object.
[0411] Figure 7B An example is shown of four audio playback devices (audio playback device 200, audio playback device 201, audio playback device 202, and audio playback device 203), a sound-producing object (first sound-producing object), and an electronic device 100. Figure 7B The coordinate system Xt-Ot-Yt shown can be referenced from the previous section. Figure 6FThe introduction is as follows. It is not limited to the four audio playback devices and one sound-producing object mentioned above; there can be more sound-producing objects in a video, and more or fewer audio playback devices used to play the audio.
[0412] like Figure 7B As shown, the angle between the first sound-emitting object and the observer's line L1, and between the audio playback device 200 and the observer's line L2, is θ1. θ1 is smaller than the angle between line L2, the audio playback device 201 and the observer's line L3. θ1 is smaller than the angle between line L2, the audio playback device 202 and the observer's line L4. θ1 is smaller than the angle between line L2, the audio playback device 203 and the observer's line L5. The electronic device 100 can correspond the first sound-emitting object to the audio playback device 200. The audio playback device 200 can play the audio component corresponding to its own sound-emitting object.
[0413] For example, electronic device 100 can instruct audio playback device 200 to play the first sound source when it is in a certain position. Figure 7B The image shows the audio component at the indicated position. Specifically, an image where the audio contains the first audio component of the first sound-producing object, and the display time is the same as the playback time of that audio component, indicates that the first sound-producing object is located at... Figure 7B In the case shown, the audio playback device 200 can play the audio component. If the audio contains a first audio component of the first sound-emitting object, and the image displayed at the same time as the playback time of that audio component does not contain the first sound-emitting object, the electronic device 100 can determine the position of the first sound-emitting object based on the image from which the first sound-emitting object most recently appeared on the video screen. If the image from which the first sound-emitting object most recently appeared on the video screen indicates that the first sound-emitting object is located... Figure 7B At the position shown, the audio playback device 200 can play the first audio component of the first sound source.
[0414] In some embodiments, during video playback, the position of a sound-producing object in a second scene may change. Therefore, the correspondence between the sound-producing object and the audio playback device can change as the position of the sound-producing object changes. For example, when the position of the first sound-producing object changes, the electronic device 100 can redetermine the positional relationship between the first sound-producing object, the audio playback device, and the observer. If the position of the first sound-producing object after the change is closest to the position of the audio playback device 201, the electronic device 100 can instruct the audio playback device 201 to play the audio component of the first sound-producing object when it is in the changed position. In other words, the electronic device 100 can detect whether the position of the sound-producing object has changed in real time, thereby adjusting the audio playback device used to play the audio component of that sound-producing object in real time.
[0415] In some embodiments, during video playback, the observer's position in the first scene may change. The correspondence between the sound-emitting object and the audio playback device can change with the observer's position. For example, when the observer's position changes, the electronic device 100 can redetermine the positional relationship between the first sound-emitting object, the audio playback device, and the observer. If the observer's position changes, the electronic device 100 can, according to the aforementioned... Figure 7A The method shown determines that the audio playback device corresponding to the first sound-producing object changes to audio playback device 202. Electronic device 100 can instruct audio playback device 202 to play the first audio component of the first sound-producing object after the observer's position changes. In other words, electronic device 100 can detect whether the observer's position has changed in real time, and thus adjust the audio playback device used to play the audio component of the sound-producing object in real time.
[0416] In some embodiments, the electronic device 100 can also detect in real time whether the position of the audio playback device has changed. A change in the position of the audio playback device will affect the positional relationship between the audio playback device and the sound-producing object. After detecting a change in the position of the audio playback device, the electronic device 100 can re-determine the correspondence between the sound-producing object and the audio playback device based on the position of the audio playback device, the position of the sound-producing object, and the position of the observer.
[0417] In addition to determining the correspondence between the sound source and the audio playback device, the electronic device 100 can also select an audio playback device for playing the second audio component of the first background sound.
[0418] In one possible implementation, the electronic device 100 can select one audio playback device from a plurality of audio playback devices to play the second audio component of the first background sound. The audio playback device playing the second audio component of the first background sound can be an audio playback device in an idle state.
[0419] Preferably, the electronic device 100 can select at least one audio playback device from audio playback devices located on one side (e.g., the left side) and at least one audio playback device from audio playback devices located on the other side (e.g., the right side) of the electronic device 100 to play the second audio component of the first background sound. The one side and the other side of the electronic device 100 can be one side and the other side in the direction in which the display screen of the electronic device 100 faces. The electronic device 100 can preferentially select an idle audio playback device to play the second audio component of the first background sound. Having audio playback devices on both sides of the electronic device 100 playing the second audio component of the first background sound can better create stereo sound in the video playback environment, helping the user perceive the process of sound emanating from different directions.
[0420] Understandably, when the position of the sound-producing object changes and / or the observer's position changes, the electronic device 100 can adjust the audio playback device used to simulate the sound produced by the sound-producing object. That is, the audio playback device idle during video playback may change. The electronic device 100 may preferentially select an idle audio playback device to play the second audio component of the first background sound. Therefore, during video playback, the audio playback device used to play the second audio component of the first background sound can change.
[0421] If all audio playback devices have corresponding sound sources, the electronic device 100 can select the audio playback device that requires the fewest simulated sound sources to play the second audio component of the first background sound. Specifically, the audio playback device can play the audio component of the sound source according to the playback parameters of the audio component of the sound source, and play the second audio component of the first background sound according to the playback parameters of the second audio component of the first background sound.
[0422] Once it is determined which audio playback device will play the audio component, the electronic device 100 can further determine the playback parameters for the audio playback device. In this way, the electronic device 100 can instruct the audio playback device to adjust the playback parameters to play the corresponding audio component, so that when a user distinguishes the sounds of different sound sources during video playback, the location of the sound source emitted by the audio playback device can be identified as the location of the sound source simulated by that audio playback device.
[0423] The following describes in detail an implementation method for determining playback parameters provided by an embodiment of this application.
[0424] In some embodiments, the electronic device 100 may instruct the audio playback device to adjust the playback time and sound intensity so that when the audio playback device plays the audio component of the sound source corresponding to the audio playback device, the user can perceive that the sound is emitted from the location of the sound source corresponding to the audio playback device.
[0425] As can be seen from the aforementioned method of users distinguishing the location of a sound source by using both ears, users can distinguish the location of a sound source by the time difference, phase difference, sound level difference, and timbre difference of the sound reaching both ears.
[0426] The aforementioned phase difference is related to the time difference. In one possible implementation, the electronic device 100 can adjust the playback time of the audio components of different sound-producing objects in the audio played by each audio playback device, based on the playback time of the second audio component of the first background sound, to change the phase of the audio reaching the observer's ears. That is, adjusting the time difference between the audio component of a sound-producing object played by each audio playback device and the second audio component of the first background sound can adjust the observer's ability to distinguish the direction of sound from that sound-producing object. The playback time of the second audio component of the first background sound, which is used as a reference, can be determined based on the time attributes of the second audio component of the first background sound in the video.
[0427] The aforementioned sound level difference can represent the difference in sound intensity. Sound intensity gradually decreases during propagation. Sound intensity and sound pressure are positively correlated. The lower the sound intensity, the lower the sound pressure. The farther the sound source is from the observer, the lower the sound pressure reaching the user's ears at the same sound intensity. The electronic device 100 can adjust the sound intensity of the audio components of different sound-producing objects in the audio played by each audio playback device, based on the sound intensity of the second audio component of the first background sound, to change the sound pressure reaching the observer's ears for each audio component. That is, adjusting the sound intensity of the audio component of a sound-producing object played by each audio playback device can adjust the observer's ability to distinguish the distance of that sound-producing object. The electronic device 100 can determine the sound intensity of the audio component with the aforementioned sound intensity as a reference using formula (1) from the aforementioned embodiment. The sound pressure used to calculate the sound intensity can be obtained by integrating the time-domain signal of the audio component.
[0428] Here it is still Figure 7B The method for determining the playback parameters of an audio playback device by an electronic device 100 is illustrated using the positional relationship shown as an example.
[0429] like Figure 7B As shown, the distance between the audio playback device 200 and the observer is D1. The distance between the first sound-emitting object and the observer is D2.
[0430] 1. Determine the sound intensity of the audio components played by the audio playback device.
[0431] Electronic device 100 can determine the sound pressure p1′ of the first audio component of the first sound-emitting object at the audio playback device 1 according to the following formula (2).
[0432] p1′=pn*D1 / D2 (2)
[0433] Where pn can represent the sound pressure of the second audio component of the first background sound.
[0434] The electronic device 100 can use the above-described p1′ and the aforementioned formula (1) to determine the sound intensity of the first audio component of the first sound source played by the audio playback device 200. For example, the sound intensity can be I1′=p1′. 2 / (ρ*C). That is, the audio playback device 200 can play the audio at sound intensity I1′ from the location of the first sound source. Figure 7B The audio components at the indicated positions.
[0435] As can be seen from the above method, the electronic device 100 can adjust the sound intensity of the audio components of each sound source by using the sound intensity of the second audio component of the first background sound as a reference.
[0436] In some embodiments, the electronic device 100 may select an audio playback device to play the second audio component of the first background sound. Using the sound intensity of the second audio component of the first background sound as a reference, the electronic device 100 can determine the sound intensity of the second audio component of the first background sound using the aforementioned pn and formula (1).
[0437] In other embodiments, the electronic device 100 may select at least two audio playback devices to play the second audio component of the first background sound. The electronic device 100 may designate one of the multiple audio playback devices playing the first background sound as a reference audio playback device. The reference audio playback device may play the audio component at the sound intensity of the second audio component of the first background sound. The electronic device 100 may determine the sound intensity of the second audio component of the first background sound played by other audio playback devices based on the distances of the reference audio playback device and the other audio playback devices playing the second audio component of the first background sound relative to the observer. Since sound intensity attenuates during propagation, the farther the audio playback device is from the observer, the greater the sound intensity of the second audio component of the first background sound played by that audio playback device can be. This ensures that the sound intensity of the first background sound is the same when it reaches the observer from different audio playback devices.
[0438] For example, electronic device 100 can select audio playback device 201 and audio playback device 202 to play the second audio component of the first background sound. The distance between audio playback device 201 and the observer is D3. The distance between audio playback device 202 and the observer is D4. Audio playback device 201 is the aforementioned reference audio playback device. The sound intensity of the second audio component of the first background sound played by audio playback device 201 can be In. In = pn 2 / (ρ*C). Electronic device 100 can determine the sound pressure pn3 of the second audio component of the first background sound at the audio playback device 202 according to the following formula (3).
[0439] pn3=pn*D4 / D3 (3)
[0440] Based on pn3 and the aforementioned formula (1), the electronic device 100 can determine the sound intensity of the second audio component of the first background sound played by the audio playback device 202 as In3. Wherein3 = pn3 2 / (ρ*C).
[0441] This application embodiment does not limit the sound intensity used as a reference. In addition to the sound intensity of the second audio component of the first background sound described above, the electronic device 100 can also use the sound intensity of the audio component of a sound-producing object in the video as a reference to determine the sound intensity of the audio components of other sound-producing objects played by the audio playback device.
[0442] It should be noted that the audio component whose sound intensity is used as a reference can be a different audio component determined from the audio within a certain time period in the video. For example, electronic device 100 separates multiple audio components of sound-producing objects and a second audio component of a first background sound from the audio within a first time period. Electronic device 100 can use the sound intensity of the second audio component of the first background sound within the first time period as a reference to determine the sound intensity of the multiple audio components of sound-producing objects played by the audio playback device within the first time period.
[0443] 2. Determine the playback time of the audio components played by the audio playback device.
[0444] Electronic device 100 can determine the playback time T′ of the first audio component of the first sound source played by audio playback device 200 according to the following formulas (4-1) and (4-2).
[0445] T′=T1+ΔT (4-1)
[0446]
[0447] Where T1 can be the playback time of the first audio component determined based on the time attribute of the first audio component of the first sound-emitting object in the video. That is, T1 is the original playback time of the first audio component of the first sound-emitting object before the above playback time adjustment. C represents the speed of sound, and the value can be 340 meters per second.
[0448] D1 can represent the straight-line distance from the first audio component of the first sound-producing object to the observer when played by the audio playback device 200. The electronic device 100 adjusts the playback time of the first audio component of the first sound-producing object, thereby adjusting the user's perception of the sound-producing object. This is equivalent to... Figure 7BThe audio playback device 200 is moved along the tangent line at the location of the audio playback device 200 shown.
[0449] Specifically, determining the time difference for adjusting the playback time of the first audio component of the first sound-emitting object as ΔT above can be equivalent to moving the audio playback device 200 to... Figure 7B Position S1 is shown. The aforementioned ΔT can be the sound traveling along a straight line from... Figure 7B The given time is the time required for the sound to propagate from position S1 to position S2. Position S1 can be the intersection of the tangent line of the location of the audio playback device 200 and the straight line between the first sound source and the observer. That is, the direction of position S1 relative to the observer is the same as the direction of the first sound source relative to the observer. Position S2 can be the location on the straight line between the first sound source and the observer, at a distance D1 from the observer. The distance between positions S1 and S2 is D0. It can be seen that, with D1 remaining constant, the smaller θ1 is, the smaller the value of D0, and the smaller the value of ΔT. A smaller θ1 indicates that the direction of the audio playback device 200 relative to the observer is closer to the direction of the first sound source relative to the observer. Therefore, the smaller θ1 is, the smaller the adjustment amount of the playback time can be for the audio playback device 200 to perceive the sound as originating from the direction of the first sound source. For example, when θ1 is 0 (i.e., the audio playback device 200, the first sound-producing object, and the observer are in the same direction), the adjustment amount of the playback time of the first audio component of the first sound-producing object by the audio playback device 200 can be 0 (i.e., ΔT is 0).
[0450] Understandably, when ΔT is greater than 0, the playback time of the audio component lags behind T1 by ΔT. When ΔT is less than 0, the playback time of the audio component advances T1 by ΔT.
[0451] For example, the first audio component of the first sound source is originally scheduled to start playing at 1 minute 30 seconds and 10 milliseconds of the video. The distance D1 between the audio playback device 200 and the observer is 8 meters. The aforementioned θ1 is 60°. The electronic device 100 can determine that the aforementioned time difference ΔT is 23.5 milliseconds. Therefore, the electronic device 100 can instruct the audio playback device 200 to start playing the first audio component of the first sound source at 1 minute 30 seconds and 33.5 milliseconds of the video.
[0452] As can be seen, the adjustment range of the playback time mentioned above is typically on the order of milliseconds. During video playback, the time difference between the playback of audio and the display of the corresponding image is usually within 100 milliseconds, which is generally acceptable. That is, users will not perceive any asynchrony between the video and audio. In the video playback environment, the distance between the audio playback device and the user will generally not cause the time difference for adjusting the playback time to exceed 100 milliseconds. Therefore, adjusting the playback time of the audio playback device to play the corresponding audio component will not cause the user to perceive any asynchrony between the video and audio.
[0453] In some embodiments, the electronic device 100 may select an audio playback device to play the second audio component of the first background sound. The electronic device 100 may determine the playback time of the second audio component of the first background sound based on the time attribute of the second audio component of the first background sound in the video. The second audio component of the first background sound is the audio component of a first time segment on the timeline of the video, and the electronic device 100 may instruct the audio playback device to play the second audio component of the first background sound within the first time segment of the video playback.
[0454] In other embodiments, the electronic device 100 may select at least two audio playback devices to play the audio component of the first background sound. The electronic device 100 may designate one of the multiple audio playback devices playing the first background sound as a reference audio playback device. The playback time of the reference audio playback device playing the second audio component of the first background sound can be determined based on the time attributes of the second audio component of the first background sound in the video. The electronic device 100 can determine the playback time of the second audio component of the first background sound played by other audio playback devices based on the distances of the reference audio playback device and other audio playback devices playing the second audio component of the first background sound relative to the observer. Since the greater the distance between the audio playback device and the observer, the longer it takes for the sound generated by the audio playback device to reach the observer, the earlier the playback time of the second audio component of the first background sound can be. This allows the first background sound played by different audio playback devices to reach the observer simultaneously.
[0455] For example, electronic device 100 may select audio playback device 201 and audio playback device 202 to play the second audio component of the first background sound. The distance between audio playback device 201 and the observer is D3. The distance between audio playback device 202 and the observer is D4. Audio playback device 201 is the aforementioned reference audio playback device. The playback time of audio playback device 201 playing the second audio component of the first background sound can be T2. T2 can be determined based on the time attribute of the second audio component of the first background sound in the video. Electronic device 100 can determine the time difference ΔT2 between audio playback device 201 and audio playback device 202 playing the second audio component of the first background sound according to the following formula (5).
[0456] ΔT2=(D3-D4) / C (5)
[0457] Where C is the speed of sound, and its value can be 340 meters per second.
[0458] The electronic device 100 can determine the playback time T3 of the second audio component of the first background sound played by the audio playback device 202 based on T2 and ΔT2. Wherein, T3 = T2 + ΔT2.
[0459] In addition to the playback time and sound intensity mentioned above, audio playback devices can also adjust other playback parameters, such as the frequency of sound waves.
[0460] Understandably, the playback parameters for an audio playback device to play the audio components of a sound source can be determined based on the positional relationship between the sound source, the video playback device, and the observer. Therefore, when one or more of the positions of the sound source, the audio playback device, and the observer change, the positional relationship between them all changes. The electronic device can then redetermine the playback parameters for the audio components of the sound source based on this changed positional relationship. Specifically, when a change in the position of one or more of the sound source, the audio playback device, and the observer results in a change in the audio playback device playing the audio components of a sound source, the electronic device 100 can determine the playback parameters for that specific audio component based on the changed positional relationship between the audio playback device, the sound source, and the observer.
[0461] This application does not limit the method for implementing the playback of an audio component of a sound-producing object by an audio playback device, enabling an observer to determine the location of the sound source as the location of the sound-producing object. Besides the method described above, which involves playing an audio component of a sound-producing object through a single audio playback device and adjusting the playback time and sound intensity of the audio component, it is also possible to play an audio component of a sound-producing object through multiple audio playback devices and adjust the playback time and sound intensity of the audio component played by these multiple audio playback devices, so that an observer can determine the location of the sound source as the location of the sound-producing object.
[0462] The following describes another method for determining playback parameters provided in the embodiments of this application.
[0463] Please refer to Figure 7C , Figure 7C An exemplary diagram illustrates another method for adjusting playback parameters of an audio playback device.
[0464] This explanation uses audio playback devices 200 and 201 as an example to illustrate simulating the sound of a first sound-producing object. However, it is not limited to two audio playback devices; more audio playback devices can be used to simulate the sound of a single sound-producing object. Figure 7C The positional relationships of the observer, audio playback device 200, audio playback device 201, and the first sound-emitting object shown are merely illustrative and should not be construed as limiting this application.
[0465] In one possible implementation, electronic device 100 can adjust the playback time of the first audio component of the first sound-producing object played by audio playback device 200 according to the method of the foregoing embodiments, such that audio playback device 200 is equivalent to emitting sound at the location of audio playback device 200'. The direction vector of audio playback device 200' relative to the observer's right ear is a1. The direction vector of audio playback device 200' relative to the observer's left ear is a2. Similarly, electronic device 100 can adjust the playback time of the first audio component of the first sound-producing object played by audio playback device 201 according to the method of the foregoing embodiments, such that audio playback device 201 is equivalent to emitting sound at the location of audio playback device 201'. The direction vector of audio playback device 201' relative to the observer's right ear is b1. The direction vector of audio playback device 201' relative to the observer's left ear is b2. The direction vector of the first sound-producing object relative to the observer's right ear is c1. The direction vector of the first sound-producing object relative to the observer's left ear is c2.
[0466] In this case, the direction vector obtained by a1+b1 has the same direction as c1. The direction vector obtained by a2+b2 has the same direction as c2. Therefore, when audio playback devices 200 and 201 simulate the sound emitted by the first sound-producing object, they can adjust the phase of the first audio component of the first sound-producing object reaching the observer's ears. By observing the phase difference between the sound heard by the left and right ears, the observer can determine the direction of the sound as the direction of the first sound-producing object.
[0467] Understandably, if the actual position of audio playback device 200 is the position of audio playback device 200', and the actual position of audio playback device 201 is the position of audio playback device 201', electronic device 100 does not need to adjust the playback time of the first audio component of the first sound source played by audio playback device 200 and audio playback device 201.
[0468] The method for determining the sound intensity of the first audio component of the first sound source played by the audio playback device 200 and the audio playback device 201 can refer to the aforementioned embodiments, and will not be repeated here.
[0469] It should be noted that the audio playback devices 200 and 201 mentioned above can be the two audio playback devices that are closest to the location of the first sound-producing object among all audio playback devices. In scenarios where multiple audio playback devices simulate the sound of a single sound-producing object, the method for selecting these multiple audio playback devices is not limited in this embodiment.
[0470] In another possible implementation, the electronic device 100 may store a playback parameter determination model (ΔT1, ΔI) = g(s1, s2, s3, s4, s5). Here, ΔT1 can represent the time difference between the playback times of the first audio component of the first sound-emitting object played by the audio playback device 200 and the audio playback device 201. ΔI can represent the intensity difference between the sound intensities of the first sound-emitting object played by the audio playback device 200 and the audio playback device 201. If the first audio component of the first sound-emitting object played by the audio playback device 200 is taken as a reference, the playback time of the audio component played by the audio playback device 200 can be determined based on the time attribute of the audio component in the video, for example, the playback time is Tb. The sound intensity of the audio component played by the audio playback device 200 can be determined based on the sound pressure of the audio component and the aforementioned formula (1), for example, the sound intensity is Ib. Therefore, the playback time of the first audio component of the first sound-emitting object played by the audio playback device 201 can be Tb + ΔT1. The sound intensity of the first audio component played by the audio playback device 201 can be Ib+ΔI.
[0471] The position of s1 can represent the position of the observer's left ear. The position of s2 can represent the position of the observer's right ear. The position of s3 can represent the position of audio playback device 1. The position of s4 can represent the position of audio playback device 2. The position of s5 can represent the position of the first sound-emitting object. The above positions can be in a two-dimensional coordinate system or a three-dimensional coordinate system.
[0472] The aforementioned playback parameter determination model can be a neural network-based model, or a head-related transfer function (HRTF)-based model, etc. This application does not limit the specific type of the playback parameter determination model.
[0473] Understandably, when two audio playback devices play the same audio component, the sound intensity of that component played by the two devices will affect not only the observer's ability to distinguish the distance between the sound source and themselves, but also their ability to distinguish the direction of the sound source relative to themselves. For example, if the sound intensity of the audio component played by audio playback device 200 is greater than that played by audio playback device 201, the observer will audibly perceive the sound source as being more towards the direction of audio playback device 200.
[0474] Furthermore, the distances between the two audio playback devices and the observer (more specifically, the observer's left and right ears) may differ. A time difference between the two audio playback devices playing the same audio component can alter the phase difference of the audio component reaching the user's ears, thereby changing the user's perception of the sound source's direction relative to themselves.
[0475] Optionally, the above playback parameter determination model can also be (ΔT1, ΔI) = h(s0, s3, s4, s5). Here, s0 can represent the observer's position. That is, the above playback parameter determination model can directly use the distances between each audio playback device and the sound-emitting object and the observer to determine the playback parameters, without specifying the exact distances between each audio playback device and the sound-emitting object and the observer's left ear, or between each audio playback device and the sound-emitting object and the observer's right ear. This simplifies the calculation process and saves the computing resources of the electronic device 100.
[0476] As can be seen, electronic device 100 can instruct multiple audio playback devices to simulate the sound of a sound-producing object, and adjust the playback time and sound intensity of the audio component of the sound-producing object played by these multiple audio playback devices according to the aforementioned playback parameters. The time difference and phase difference between the arrival of an audio component of a sound-producing object at the user's ears in the above method allow the user to match the location of the sound source of the audio component with the location of the sound-producing object. Matching the location of the sound source of the audio component with the location of the sound-producing object can include: the location of the sound source of the audio component being the same as the location of the sound-producing object, or the location of the sound source of the audio component being within a preset range of the location of the sound-producing object (e.g., within a radius of 1 meter).
[0477] In the subsequent embodiments of this application, the example of an audio component of a sound-producing object being played by an audio playback device will be used for illustration.
[0478] In some embodiments, the electronic device 100 detects multiple observers. The electronic device 100 can calculate the center of the positions of these multiple observers to obtain the center position. For example, the electronic device 100 detects two observers, and the positions of these two observers are (x11, y11) and (x21, y21), respectively. The electronic device 100 can calculate the center position of these two observers as ((x11+x21) / 2, (y11+y21) / 2). The specific calculation method of the center position is not limited in the embodiments of this application.
[0479] Electronic device 100 can determine the positional relationship between each audio playback device and the central location. For example, electronic device 100 can establish a system with the central location as the origin. Figure 6E The electronic device 100 establishes a coordinate system and determines the position of each audio playback device within that system. Furthermore, the electronic device 100 can use the position of the virtual camera as the center position to determine the positional relationship between each sound-producing object in the video and this center position. In this way, the electronic device 100 can determine the sound-producing object corresponding to each audio playback device, as well as the playback parameters for the audio components of the sound-producing object played by the audio playback device.
[0480] If the positions of the aforementioned observers change, the electronic device 100 can redetermine the center positions of the observers based on the changed positions, thereby updating the positional relationships between the center positions, the sound-emitting object, and the audio playback devices. In other words, the electronic device 100 can detect the positions of the observers in real time and adjust the playback parameters of the sound-emitting object corresponding to each audio playback device and the audio components of the sound-emitting object played by the audio playback devices accordingly after the observer positions change.
[0481] For example, if one of the multiple observers leaves, the electronic device 100 can determine the aforementioned center position based on the positions of the remaining observers, and adjust the playback parameters of the sound-producing object corresponding to each audio playback device and the audio components of the sound-producing object played by the audio playback devices. This allows observers who are still watching the video to have a better sense of immersion.
[0482] As can be seen, during video playback, the electronic device 100 can adjust the playback parameters of the sound-producing objects and their audio components on each audio playback device based on the positions of multiple observers. Each observer can immerse themselves in the scene of the sound-producing objects in the video, experiencing the different sound-producing objects emitting sound from their own different locations. Furthermore, even if an observer moves while watching the video, the electronic device 100 can still adjust the playback parameters of the sound-producing objects and their audio components on each audio playback device in real time. This allows the observer to still experience the process of the sound-producing objects in the video emitting sound from their own different locations, following the perspective of the virtual camera, even while moving.
[0483] In some embodiments, the electronic device 100 can determine whether the range of change in the observer's position is less than a preset range. If the range of change in the observer's position is less than the preset range, the electronic device 100 may not change the playback parameters of the sound-producing object corresponding to the audio playback device and the audio component of the sound-producing object played by the audio playback device. The preset range may be, for example, a range within a radius of 1 meter. This application embodiment does not limit the value of the preset range.
[0484] Understandably, when the observer's position changes within a small range, the audio playback device can still provide a good sense of immersion for the observer by playing the audio component corresponding to the sound source according to the playback parameters determined before the observer's position changes. Therefore, if the electronic device 100 does not adjust the corresponding sound source and playback parameters of the audio playback device when the observer's position changes within a small range, it can save the computing resources of the electronic device 100 and reduce the requirements on the computing power of the electronic device 100.
[0485] In some embodiments, after determining that the observer's position has changed, the electronic device 100 can further determine whether the observer returns to the position before the position change within a preset time. Specifically, the electronic device 100 determines that the observer has been at position 1 for a period of time and then moves from position 1 to another position. Position 1 can be the position before the position change. The observer returning to the position before the position change within the preset time can include: the electronic device 100 determining again that the observer is at position 1 within the preset time, or the electronic device 100 determining within the preset time that the observer is within a preset range of position 1 (e.g., a radius of 1 meter).
[0486] If it is determined that the observer returns to the position before the position change within a preset time, the electronic device 100 may not change the sound source and playback parameters of the audio playback device. The preset time may be, for example, 1 minute, 3 minutes, etc. This application embodiment does not limit the value of the preset time.
[0487] Understandably, if an observer's position changes but returns to its previous position within a short period, the observer's absence may be temporary (e.g., getting up to get a drink of water or to use the restroom). Because the observer returns to its original position quickly, the electronic device 100 does not need to adjust the sound source and playback parameters of the audio playback device. This saves the computing resources of the electronic device 100 and reduces the computational demands on it.
[0488] The electronic device 100 can send the audio component of the corresponding sound source, along with playback parameters for playing that audio component, to the audio playback device based on the correspondence between the audio playback device and the sound source. The electronic device 100 can also send the second audio component of the first background sound, along with playback parameters, to the audio playback device used to play the second audio component of the first background sound. After receiving the playback parameters and the audio component, the audio playback device can play the audio component according to the instructions of the electronic device 100 using the received playback parameters.
[0489] Here Figure 8A The positional relationship shown is used as an example to illustrate the process of electronic device 100 instructing audio playback device to play audio components.
[0490] For example, electronic device 100 identifies that the first video contains sound-emitting objects E1 and E2 within a first time period. Electronic device 100 separates the audio component BE1 of sound-emitting object E1, the audio component BE2 of sound-emitting object E2, and the second audio component BE3 of the first background sound from the first audio within the first time period of the first video. Electronic device 100 can also determine the positions of sound-emitting objects E1 and E2 based on a first group of images within the first time period of the first video. The positions of sound-emitting objects E1 and E2 can be determined as follows: Figure 8A As shown.
[0491] Based on the positions of the sound-emitting object E1, the sound-emitting object E2, the observer, and the positions of the audio playback devices 200 to 203, the electronic device 100 can determine that the position of the audio playback device 200 is closest to that of the sound-emitting object E1, and the position of the audio playback device 203 is closest to that of the sound-emitting object E2.
[0492] Furthermore, the electronic device 100 can determine that the playback parameter of the audio playback device 200 playing audio component BE1 is R1, and the playback parameter of the audio playback device 203 playing audio component BE2 is R2.
[0493] Electronic device 100 can instruct audio playback devices 201 and 202 to play the second audio component BE3 of the first background sound. Specifically, electronic device 100 can determine that the playback parameter for audio playback device 201 playing audio component BE3 is R3, and the playback parameter for audio playback device 202 playing audio component BE3 is R4.
[0494] Electronic device 100 can send audio component BE1, playback parameter R1, and a first playback instruction message to audio playback device 200. The first playback instruction message can be used to instruct audio playback device 200 to play audio component BE1 with playback parameter R1. Upon receiving the audio component BE1, playback parameter R1, and the first playback instruction message, audio playback device 200 can play audio component BE1 with playback parameter R1 according to the first playback instruction message.
[0495] Electronic device 100 can send audio component BE2, playback parameter R2, and a second playback instruction message to audio playback device 203. The second playback instruction message can be used to instruct audio playback device 203 to play audio component BE2 with playback parameter R2. Upon receiving the audio component BE2, playback parameter R2, and the second playback instruction message, audio playback device 203 can play audio component BE2 with playback parameter R2 according to the second playback instruction message.
[0496] Electronic device 100 can send audio component BE3, playback parameter R3, and playback instruction message 3 to audio playback device 201. The playback instruction message 3 can be used to instruct audio playback device 201 to play audio component BE3 with playback parameter R3. When receiving the audio component BE3, playback parameter R3, and playback instruction message 3, audio playback device 201 can play audio component BE3 with playback parameter R3 according to playback instruction message 3.
[0497] Electronic device 100 can send audio component BE3, playback parameter R4, and playback instruction message 4 to audio playback device 202. The playback instruction message 4 can be used to instruct audio playback device 202 to play audio component BE3 with playback parameter R4. Upon receiving the audio component BE3, playback parameter R4, and playback instruction message 4, audio playback device 202 can play audio component BE3 with playback parameter R4 according to the playback instruction message 4.
[0498] As can be seen from the aforementioned embodiment of determining playback parameters, when the audio playback device 200 plays the audio component BE1 with playback parameter R1, the observer can auditorily distinguish the first sound source. Figure 8A The sound originates from the location of the first sound-producing object. The audio playback device 203 plays the audio component BE2 with playback parameter R2, allowing the observer to audibly distinguish the location of the second sound-producing object. Figure 8A The sound is emitted from the location of the second sound-emitting object shown. In other words, the effect of the audio played by each audio playback device in terms of sound location matches the position of the simulated sound-emitting object in the three-dimensional space represented by the audio playback device in the video frame. This allows the observer to have a better sense of immersion when watching the video.
[0499] In some embodiments, an audio playback device may correspond to multiple sound-producing objects. That is, an audio playback device may play audio components of multiple sound-producing objects.
[0500] For example, such as Figure 8B As shown, among the audio playback devices 200-203, audio playback device 200 is the closest to the location of sound-producing objects E1, E2, and E3. Electronic device 100 can instruct audio playback device 200 to play the audio components of sound-producing objects E1, E2, and E3. Specifically, audio playback device 200 can play the audio component of sound-producing object E1 using playback parameters. Similarly, audio playback device 200 can play the audio component of sound-producing object E2 using playback parameters.
[0501] Depend on Figure 8B As can be seen, the audio playback device 202 is located on one side of the electronic device 100, while the audio playback devices 201 and 203 are located on the other side. When the audio playback device 200 plays the audio components of sound source E1, sound source E2, and sound source E3, the electronic device 100 can instruct the audio playback device 202 to play the second audio component of the first background sound, and instruct the audio playback devices 201 and / or 203 to play the second audio component of the first background sound.
[0502] In some embodiments, an audio playback device may play not only the audio components of one or more sound-producing objects, but also a second audio component of a first background sound.
[0503] In the foregoing embodiments, the electronic device 100 can determine the sound-producing object and playback parameters corresponding to the audio playback device based on the positional relationship between the observer, the audio playback device, and the sound-producing object in a two-dimensional coordinate system. However, the positions of the observer, the audio playback device, and the sound-producing object may not actually be on the same horizontal plane. Differences in height will also affect the effect of each audio playback device simulating the sound-producing object's sound location.
[0504] The following describes another method for determining the sound-producing object and playback parameters corresponding to an audio playback device, provided by an embodiment of this application.
[0505] In some embodiments, the electronic device 100 can determine the positional relationship between the observer, the audio playback device, and the sound-producing object in a three-dimensional coordinate system, and determine the sound-producing object corresponding to the audio playback device and the playback parameters based on the positional relationship. This can improve the accuracy of the audio playback device simulating the sound-producing object's sound, making the effect of each audio playback device simulating the sound-producing object's sound location in the three-dimensional space represented by the audio playback device more closely match the position of the sound-producing object simulated by the audio playback device in the three-dimensional space represented by the video image.
[0506] Specifically, the electronic device 100 can acquire the positions of the observer and the audio playback device in three-dimensional space using three-dimensional ultrasound imaging. For example, the electronic device 100 can establish a [structure / structure] with the observer's position as the origin. Figure 9A The three-dimensional coordinate system Ot-Xt-Yt-Zt is shown, and the positions of electronic device 100 and each audio playback device in this three-dimensional coordinate system are determined. The Xt axis and Yt axis of the three-dimensional coordinate system Ot-Xt-Yt-Zt can be referred to the aforementioned... Figure 6EThe Zt axis can be a straight line perpendicular to the Xt-Ot-Yt plane. Here, we use two audio playback devices as an example. The electronic device 100 is located at (0, y0, z0). The audio playback device 200 is located at (x1, y1, z1). The audio playback device 201 is located at (x2, y2, z2). The positions of the electronic device 100 and the audio playback devices described above are merely illustrative and should not be construed as limiting this application.
[0507] In addition to the three-dimensional ultrasound imaging methods described above, the electronic device 100 can also acquire the positions of the observer, display device, and audio playback device in three-dimensional space through other methods.
[0508] As can be seen from the foregoing embodiments, the electronic device 100 can perform three-dimensional reconstruction of images in a video to determine the positions of the sound-emitting object and the virtual camera in the video. The positions of the sound-emitting object and the virtual camera can be determined with reference to the aforementioned... Figure 6B The description shown is provided. Electronic device 100 can be based on... Figure 6B Based on the positional relationships shown, a three-dimensional coordinate system is established with the virtual camera's position as the origin. For example, the aforementioned three-dimensional coordinate system with the virtual camera's position as the origin can be referenced... Figure 9B The diagram shows a three-dimensional coordinate system Oc-Xc-Yc-Zc. The Xc and Yc axes of this three-dimensional coordinate system can be referenced from the previous section. Figure 6C The Zc axis can be a straight line perpendicular to the Xc-Oc-Yc plane. Here, we use two sound-emitting objects as an example. The position of the first sound-emitting object is (x5, y5, z5). The position of the second sound-emitting object is (x6, y6, z6). The positions of the sound-emitting objects described above are merely illustrative and should not be construed as limiting this application.
[0509] Electronic device 100 can use the observer's position as the position of the virtual camera to obtain the positional relationship between the sound-emitting object and the observer. For example, the control device can... Figure 9B The three-dimensional coordinate system Oc-Xc-Yc-Zc shown is Figure 9A The three-dimensional coordinate system Ot-Xt-Yt-Zt shown coincides, resulting in Figure 9C The virtual camera and the sound-emitting object are shown in the three-dimensional coordinate system Ot-Xt-Yt-Zt.
[0510] Electronic device 100 can Figure 9CIn the three-dimensional coordinate system shown, data such as the distance between each sound-producing object and the observer, the distance between each audio component and the observer, and the angle between the line containing the sound-producing object and the observer, and the line containing the audio playback device and the observer, are calculated. Based on the data calculated in the three-dimensional coordinate system, the electronic device 100 can determine the sound-producing object and playback parameters corresponding to the audio playback device. The principle and specific method by which the electronic device 100 determines the sound-producing object and playback parameters corresponding to the audio playback device can be found in the description of calculations based on the positional relationship between the observer, the audio playback device, and the sound-producing object in the two-dimensional coordinate system in the foregoing embodiments.
[0511] In one possible implementation, the electronic device 100 can determine whether to use positional relationships in a two-dimensional coordinate system or a three-dimensional coordinate system to determine the sound-producing object and playback parameters corresponding to the audio playback device, based on its own computing power. Understandably, the complexity of calculations using positional relationships in a three-dimensional coordinate system is higher than that using positional relationships in a two-dimensional coordinate system. Therefore, when its computing power is high, the electronic device 100 can use positional relationships in a three-dimensional coordinate system to determine the sound-producing object and playback parameters corresponding to the audio playback device.
[0512] For example, electronic device 100 may have a dedicated processing module for determining the sound-producing object and playback parameters corresponding to an audio playback device. This processing module has strong computational capabilities. Electronic device 100 can perform calculations using the positional relationships of the observer, audio playback device, and sound-producing object in a three-dimensional coordinate system. Alternatively, electronic device 100 may use a general-purpose processing module (such as a CPU) to determine the sound-producing object and playback parameters corresponding to an audio playback device. When electronic device 100 has many applications running, and the computational resources available for determining the sound-producing object and playback parameters are limited, electronic device 100 can perform calculations using the positional relationships of the observer, audio playback device, and sound-producing object in a two-dimensional coordinate system.
[0513] In another possible implementation, the electronic device 100 can determine the sound-producing object and playback parameters of the audio playback device based on the user's choice, using positional relationships in a two-dimensional coordinate system or in a three-dimensional coordinate system.
[0514] For example, such as Figure 9D As shown, the electronic device 100 can display a video playback interface 910. The video playback interface 910 can display a video image. The video playback interface 910 may include a stereo precision option 911. This stereo precision option 911 can be used by the user to select the precision of the sound produced by the simulated sound source on the audio playback device.
[0515] In response to the operation of stereo precision option 911, electronic device 100 can display Figure 9E The option box 912 is shown. This option box 912 may contain a low-precision option 912A and a high-precision option 912B. When an operation to select low-precision option 912A is detected, the electronic device 100 can determine the sound-producing object and playback parameters corresponding to the audio playback device using the positional relationship between the observer, the audio playback device, and the sound-producing object in a two-dimensional coordinate system. When an operation to select high-precision option 912B is detected, the electronic device 100 can determine the sound-producing object and playback parameters corresponding to the audio playback device using the positional relationship between the observer, the audio playback device, and the sound-producing object in a three-dimensional coordinate system.
[0516] Optionally, when the operation of selecting the high-precision option 912B is detected, the electronic device 100 may prompt the user that selecting the high-precision option 912B requires more computing resources and the device will consume more power.
[0517] Please refer to Figure 10 , Figure 10 An exemplary schematic diagram of a communication system 1000 provided in an embodiment of this application is shown.
[0518] like Figure 10 As shown, the communication system 1000 may include an electronic device 100 and one or more audio playback devices (such as audio playback device 200, audio playback device 201, etc.). A communication connection is established between the electronic device 100 and each audio playback device.
[0519] The electronic device 100 may include a control unit 1001, a video analysis unit 1002, and a display unit 1003. The control unit 1001 can be used to acquire the position of the observer and the positions of each audio playback device. The video analysis unit 1002 can be used to analyze the video to determine the sound-emitting object in the video, the audio components of the sound-emitting object, the position of the sound-emitting object, and the position of the virtual camera. The control unit 1001 can also be used to determine the correspondence between the sound-emitting object and the audio playback devices, as well as the playback parameters. The implementation methods for the control unit 1001 and the video analysis unit 1002 to determine the corresponding information can be referred to the foregoing embodiments. Further details are omitted here.
[0520] The control unit 1001 and video analysis unit 1002 described above can be integrated on the same processor, such as a CPU, in the electronic device 100. Alternatively, the control unit 1001 and video analysis unit 1002 can also be integrated on different processors in the electronic device 100. For example, the control unit 1001 is integrated on the CPU, and the video analysis unit 1002 is integrated on the NPU.
[0521] The display unit 1003 described above can be used to display video images. The display unit 1003 may include a display screen.
[0522] The electronic device 100 may include more units than just the control unit 1001, video analysis unit 1002, and display unit 1003 described above. For example, it may include a communication unit, an audio output unit, and an audio input unit. The audio output unit may include a speaker. The audio input unit may include a microphone.
[0523] During video playback, electronic device 100 can display images from the video and instruct one or more audio playback devices to play audio components contained in the video's audio. These audio playback devices can simulate the sound production process of objects in the video, giving the viewer an immersive experience. The communication system 1000 enhances the user's sense of immersion and engagement when watching videos, improving the overall user experience.
[0524] Please refer to Figure 11 , Figure 11 An exemplary schematic diagram of a communication system 1100 provided in an embodiment of this application is shown.
[0525] like Figure 11 As shown, the communication system 1100 may include a control device 1101, a video analysis device 1102, one or more audio playback devices (such as audio playback device 1104, audio playback device 1105, etc.), and a display device 1103. The control device 1101 establishes a communication connection with the video analysis device 1102. The control device 1101 also establishes communication connections with each audio playback device and the display device 1103.
[0526] The aforementioned control device 1101 can be used to obtain the position of the observer and the positions of each audio playback device. The control device 1101 can send video and instructions for analyzing the video to the video analysis device 1102 through its communication connection with the video analysis device 1102.
[0527] The aforementioned video analysis device 1102 can be used to analyze video to determine the sound-emitting object, its audio components, its location, and the location of the virtual camera. Upon receiving video and video analysis instructions from the control device 1101, the video analysis device 1102 can analyze the video to determine the sound-emitting object, its audio components, its location, and the location of the virtual camera. The video analysis device 1102 can also send the sound-emitting object, its audio components, its location, and the location of the virtual camera to the control device 1101.
[0528] The aforementioned control device 1101 can also be used to determine the correspondence between the sound-emitting object and the audio playback device, as well as the playback parameters, using the sound-emitting object in the video, the audio components of the sound-emitting object, the position of the sound-emitting object, and the position of the virtual camera. The control device 1101 can send the audio components and playback parameters to the audio playback device and instruct the audio playback device to play the audio components according to the playback parameters. The control device 1101 can also send images from the video to the display device 1103 and instruct the display device 1103 to display the images from the video.
[0529] During video playback, display device 1103 can display images from the video. One or more audio playback devices can play audio components contained in the video's audio. These audio playback devices can simulate the sound production process of objects in the video, giving viewers an immersive experience. The communication system 1100 enhances the user's sense of immersion and engagement when watching videos, improving the overall user experience.
[0530] Please refer to Figure 12 , Figure 12 An exemplary schematic diagram of a communication system 1200 provided in an embodiment of this application is shown.
[0531] like Figure 12 As shown, the communication system 1200 may include one or more audio playback devices (such as audio playback device 1210, audio playback device 1220, audio playback device 1230, etc.) and a display device 1240.
[0532] The audio playback device 1210 may include a control unit 1211, a video analysis unit 1212, and an audio output unit 1213. The control unit 1211 can be referred to the foregoing. Figure 10 The control unit 1001 in the communication system 1000 shown. The video analysis unit 1212 can be referred to the above. Figure 10 The video analysis unit 1002 in the communication system 1000 shown is not described in detail here. The audio output unit 1213 may include a speaker and can be used to convert audio into sound signals. The audio playback device 1210 can establish a communication connection with other audio playback devices. The audio playback device 1210 can establish a communication connection with the display device 1240.
[0533] The audio playback device 1210 may include more units than those described above, such as the control unit 1211, video analysis unit 1212, and audio output unit 1213. For example, a communication unit and an audio input unit.
[0534] As can be seen, the audio playback device 1210 can acquire the observer's position, the positions of each audio playback device, and analyze the video to determine the sound-producing object in the video, the audio components of the sound-producing object, the position of the sound-producing object, and the position of the virtual camera. The audio playback device 1210 can also determine the audio components played by each audio playback device, as well as the playback parameters.
[0535] During video playback, audio playback device 1210 can play its own audio components and instruct other audio playback devices to play corresponding audio components. Audio playback device 1210 can also instruct display device 1240 to display images from the video. The aforementioned one or more audio playback devices can simulate the sound production process of objects in the video, giving viewers an immersive experience. The aforementioned communication system 1200 can enhance the user's sense of immersion and engagement when watching videos, thereby improving the user experience.
[0536] It should be noted that, without causing contradictions or conflicts, any feature in any embodiment of this application, or any part of any feature, can be freely combined, and the combined technical solution is also within the scope of this application.
[0537] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A method for co-playing audio during video playback, characterized in that, The method is applied to an electronic device capable of communicating with M audio playback devices, wherein the electronic device and the M audio playback devices are independent devices, the electronic device includes a display screen, and M is a positive integer greater than or equal to 2. The method includes: The electronic device acquires a first video, which includes a first group of images and a first audio within a first time period. Based on the first image group and the first audio, the electronic device determines that the first video contains a first sound-emitting object and a first background sound within the first time period, and separates a first audio component of the first sound-emitting object and a second audio component of the first background sound from the first audio. The electronic device sends a first message to the first audio playback device among the M audio playback devices. The first message includes the first audio component and the first playback parameters. The first message is used to instruct the first audio playback device to play the first audio component with the first playback parameters. The first audio playback device is obtained based on the positions of all or part of the M audio playback devices in the first scene relative to the observer viewing the electronic device, the position of the first sound source relative to the virtual camera of the first video in the second scene, and determining the position of the virtual camera in the second scene as corresponding to the position of the observer in the first scene. The electronic device sends a second message to the second audio playback device among the M audio playback devices. The second message contains the second audio component and the second playback parameter. The second message is used to instruct the second audio playback device to play the second audio component with the second playback parameter. While the first video is playing during the first time period, the electronic device displays images from the first image group.
2. The method according to claim 1, characterized in that, The method further includes: Based on the first image group and the first audio, the electronic device determines that the first video contains a second sound-emitting object within the first time period, and separates a third audio component of the second sound-emitting object from the first audio. The electronic device sends a third message to the third audio playback device among the M audio playback devices. The third message contains the third audio component and the third playback parameters. The third message is used to instruct the third audio playback device to play the third audio component with the third playback parameters.
3. The method according to claim 1, characterized in that, The first playback parameter is obtained based on the position of the first audio playback device relative to the observer viewing the electronic device in the first scene, the position of the first sound-emitting object relative to the virtual camera of the first video in the second scene, and determining the position of the virtual camera in the second scene to correspond to the position of the observer in the first scene.
4. The method according to claim 2, characterized in that, The third audio playback device is obtained based on the position of all or part of the M audio playback devices in the first scene relative to the observer watching the electronic device, the position of the second sound-emitting object in the second scene relative to the virtual camera of the first video, and determining the position of the virtual camera in the second scene as corresponding to the position of the observer in the first scene.
5. The method according to claim 2, characterized in that, The third playback parameter is obtained based on the position of the third audio playback device relative to the observer viewing the electronic device in the first scene, the position of the second sound-emitting object relative to the virtual camera of the first video in the second scene, and determining the position of the virtual camera in the second scene to correspond to the position of the observer in the first scene.
6. The method according to any one of claims 1-5, characterized in that, The position corresponding to the observer's position in the first scene includes: The same position as the observer in the first scene.
7. The method according to any one of claims 1-5, characterized in that, The observer is located at a first position in the first scene, and the first audio playback device is located at a second position in the first scene; After determining the position of the virtual camera in the second scene to correspond to the position of the observer in the first scene, the method further includes: The electronic device obtains a third position of the first voice-emitting object in the first scene based on the first position and the position of the first voice-emitting object relative to the virtual camera in the second scene; Wherein, with the first position as the vertex of the included angle, the first position, the second position, and the third position form a first included angle, and the first included angle is the smallest among the included angles formed by the first position as the vertex, the first position, the third position, and the position of any one of the M audio playback devices (all or some of them).
8. The method according to claim 1, characterized in that, The second audio playback device is one of the M audio playback devices that does not play the audio components of the sound-producing object contained in the first video during the first time period; or, the second audio playback device is one of the M audio playback devices that plays the fewest audio components of the sound-producing object contained in the first video during the first time period.
9. The method according to claim 1, characterized in that, The first image group includes one or more image frames, the first playback parameters include a first playback time and a first sound intensity, and the second playback parameters include a second playback time and a second sound intensity.
10. The method according to claim 9, characterized in that, The first playback time and the second playback time are within the first time period.
11. The method according to claim 1, characterized in that, The method further includes: The electronic device sends a fourth message to the fourth audio playback device among the M audio playback devices. The fourth message contains the second audio component and a fourth playback parameter. The fourth message is used to instruct the fourth audio playback device to play the second audio component with the fourth playback parameter.
12. The method according to claim 11, characterized in that, The second audio playback device is located on the first side of the electronic device, and the fourth audio playback device is located on the second side of the electronic device. The first side and the second side are two sides divided by the orientation of the display screen of the electronic device.
13. The method according to claim 1, characterized in that, The method further includes the following: Multiple observers are viewing the electronic device. The electronic device obtains a first position based on the positions of the plurality of observers in the first scene, and the first position is used to represent the position of the observers viewing the electronic device.
14. The method according to claim 1, characterized in that, The first scenario is the scenario where the observer is watching the electronic device, and the second scenario is the scenario where the first video is presented.
15. The method according to claim 1, characterized in that, The first video includes a second group of images and a second audio during a second time period, and the method further includes: Based on the second image group and the second audio, the electronic device determines that the first video contains the first sound source and the second background sound during the second time period, and separates the fourth audio component of the first sound source and the fifth audio component of the second background sound from the second audio. The electronic device sends a fifth message to the first audio playback device. The fifth message includes the fourth audio component and a fifth playback parameter. The fifth message is used to instruct the first audio playback device to play the fourth audio component with the fifth playback parameter. The electronic device sends a sixth message to the second audio playback device. The sixth message includes the fifth audio component and a sixth playback parameter. The sixth message is used to instruct the second audio playback device to play the fifth audio component with the sixth playback parameter. While the first video is playing in the second time period, the electronic device displays images from the second image group.
16. A method for co-playing audio during video playback, characterized in that, The method is applied to a communication system comprising an electronic device and M audio playback devices, wherein the electronic device and the M audio playback devices are independent devices, the electronic device includes a display screen, and M is a positive integer greater than or equal to 2. The method includes: The electronic device acquires a first video, which includes a first group of images and a first audio within a first time period. Based on the first image group and the first audio, the electronic device determines that the first video contains a first sound-emitting object and a first background sound within the first time period, and separates a first audio component of the first sound-emitting object and a second audio component of the first background sound from the first audio. The electronic device sends a first message to the first audio playback device among the M audio playback devices. The first message includes the first audio component and the first playback parameters. The first message is used to instruct the first audio playback device to play the first audio component with the first playback parameters. The first audio playback device is obtained based on the positions of all or part of the M audio playback devices in the first scene relative to the observer viewing the electronic device, the position of the first sound source relative to the virtual camera of the first video in the second scene, and determining the position of the virtual camera in the second scene as corresponding to the position of the observer in the first scene. The electronic device sends a second message to the second audio playback device among the M audio playback devices. The second message contains the second audio component and the second playback parameter. The second message is used to instruct the second audio playback device to play the second audio component with the second playback parameter. While the first video is playing during the first time period, the electronic device displays images from the first image group, the first audio playback device synchronously plays the first audio component according to the first message and the first playback parameters, and the second audio playback device synchronously plays the second audio component according to the second message and the second playback parameters.
17. The method according to claim 16, characterized in that, The method further includes: Based on the first image group and the first audio, the electronic device determines that the first video contains a second sound-emitting object within the first time period, and separates a third audio component of the second sound-emitting object from the first audio. The electronic device sends a third message to the third audio playback device among the M audio playback devices. The third message includes the third audio component and the third playback parameters. The third message is used to instruct the third audio playback device to play the third audio component with the third playback parameters. While the first video is playing during the first time period, the third audio playback device synchronously plays the third audio component according to the third message and the third playback parameters.
18. The method according to claim 16, characterized in that, The first playback parameter is obtained based on the position of the first audio playback device relative to the observer viewing the electronic device in the first scene, the position of the first sound-emitting object relative to the virtual camera of the first video in the second scene, and determining the position of the virtual camera in the second scene to correspond to the position of the observer in the first scene.
19. The method according to claim 16, characterized in that, The method further includes: The electronic device sends a fourth message to the fourth audio playback device among the M audio playback devices. The fourth message contains the second audio component and a fourth playback parameter. The fourth message is used to instruct the fourth audio playback device to play the second audio component with the fourth playback parameter. While the first video is playing during the first time period, the fourth audio playback device synchronously plays the second audio component according to the fourth message and the fourth playback parameters.
20. The method according to claim 16, characterized in that, The first video includes a second group of images and a second audio during a second time period, and the method further includes: Based on the second image group and the second audio, the electronic device determines that the first video contains the first sound source and the second background sound during the second time period, and separates the fourth audio component of the first sound source and the fifth audio component of the second background sound from the second audio. The electronic device sends a fifth message to the first audio playback device. The fifth message includes the fourth audio component and a fifth playback parameter. The fifth message is used to instruct the first audio playback device to play the fourth audio component with the fifth playback parameter. The electronic device sends a sixth message to the second audio playback device. The sixth message includes the fifth audio component and a sixth playback parameter. The sixth message is used to instruct the second audio playback device to play the fifth audio component with the sixth playback parameter. While the first video is playing in the second time period, the electronic device displays the images in the second image group, the first audio playback device synchronously plays the fourth audio component according to the fifth message and the fifth playback parameters, and the second audio playback device synchronously plays the fifth audio component according to the sixth message and the sixth playback parameters.
21. An electronic device, characterized in that, The electronic device includes a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to invoke the computer program to cause the electronic device to perform the method according to any one of claims 1-15.
22. A communication system, characterized in that, The communication system includes an electronic device and M audio playback devices. The electronic device includes a display screen, where M is a positive integer greater than or equal to 2. The M audio playback devices include a first audio playback device and a second audio playback device. The electronic device is configured to perform the method as described in any one of claims 1-15. The first audio playback device is configured to synchronously play a first audio component with first playback parameters while the first video is playing during the first time period. The second audio playback device is configured to synchronously play a second audio component with second playback parameters while the first video is playing during the first time period.
23. A computer-readable storage medium storing a computer program, characterized in that, When the computer program runs on an electronic device, the electronic device performs the method according to any one of claims 1-15.
24. A computer program product, characterized in that, The computer program product includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method described in any one of claims 1-15.
Citation Information
Patent Citations
Wake-up identification method, audio device and audio device group
CN114121024A
Automatic control method and system based on human body perception, and first electronic equipment
CN116033331A
Method, apparatus and device for achieving co-location of voices and images and medium
CN109194999A
Video sound processor, and video sound processing method, and program
JP2011071686A
Information processing system, computer-readable non-transitory storage medium having stored therein information processing program, information processing control method, and information processing apparatus
US20140112505A1