Rendering method and related devices
The rendering method addresses the challenge of spatial audio rendering for single sound objects by determining sound source positions and performing spatial rendering, resulting in an immersive stereo sound effect for users.
Patent Information
- Application Number
- JP2023565286
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-04-29
- Filing Date
- 2022-04-18
- Publication Date
- 2025-06-30
- Estimated Expiration
- 2042-04-18
AI Technical Summary
Current audio and video playback devices struggle to effectively render spatial audio for single sound objects within dual-channel/multi-channel audio tracks, failing to provide an immersive stereo surround sound experience.
A rendering method that obtains a single-object audio track from a multimedia file, determines the sound source position based on reference information, and performs spatial rendering to enhance the stereo spatial perception and provide an immersive stereo sound effect.
The method improves the stereo spatial perception of single-object audio tracks, resulting in an immersive stereo sound effect that enhances the user's audio experience.
Smart Images

Figure 0007700266000114 
Figure 0007700266000115 
Figure 0007700266000116
Abstract
Description
Technical Field
[0001] This application claims priority to Chinese Patent Application No. CN202110477321.0, titled "RENDERING METHOD AND RELATED DEVICE", filed with the State Intellectual Property Office of China on April 29, 2021, the entire content of which is incorporated herein by reference.
[0002] This application relates to the field of audio applications, and in particular, to a rendering method and related devices.
Background Art
[0003] As audio and video playback technologies become increasingly mature, people's requirements for the playback effects of audio and video playback devices are also increasing.
[0004] Currently, in order to enable users to experience a realistic stereo surround sound effect when playing audio and video, audio and video playback devices can process the audio and video data to be played by using processing technologies such as the head related transfer function (HRTF).
[0005] However, a large amount of audio and video data on the Internet (such as music or movies and TV works) is in dual-channel / multi-channel audio tracks. How to perform spatial rendering on a single sound object within an audio track is an urgent problem to be solved.
Summary of the Invention
Means for Solving the Problems
[0006] Embodiments of the present application provide a rendering method for improving the stereo spatial perception of a first single-object audio track corresponding to a first sound object in a multimedia file and providing an immersive stereo sound effect to a user.
[0007] A first aspect of an embodiment of the present application provides a rendering method. This method can be applied to scenarios such as the production of music or movies and television works. This method can be implemented by a rendering device or by components of a rendering device (such as a processor, a chip, a chip system, etc.). This method includes the steps of obtaining a first single-object audio track based on a multimedia file, where the first single-object audio track corresponds to a first sound object; determining a first sound source position of the first sound object based on reference information, where the reference information includes reference position information and / or media information of the multimedia file, and the reference position information indicates the first sound source position; and performing spatial rendering on the first single-object audio track based on the first sound source position to obtain a rendered first single-object audio track.
[0008] In this embodiment of the present application, the first single-object audio track is obtained based on a multimedia file, the first single-object audio track corresponds to a first sound object, the first sound source position of the first sound object is determined based on reference information, and spatial rendering is performed on the first single-object audio track based on the first sound source position to obtain a rendered first single-object audio track. The stereo spatial perception of the first single-object audio track corresponding to the first sound object in the multimedia file can be improved, and as a result, an immersive stereo sound effect is provided to the user.
[0009] Optionally, in a possible implementation of the first aspect, the media information in the foregoing steps needs to include at least one of the text to be displayed in the multimedia file, the image to be displayed in the multimedia file, the musical characteristics of the music to be played in the multimedia file, and the sound source type corresponding to the first sound object.
[0010] In this possible implementation, when the media information includes musical characteristics, the rendering device can perform orientation and dynamics settings on the extracted specific sound object based on the musical characteristics of the music. As a result, the audio track corresponding to the sound object becomes more natural in 3D rendering, and the artistic aspects are better reflected. When the media information includes text, images, etc., a 3D immersion feeling is obtained through rendering in a headset or loudspeaker environment. In this way, the sound and the image can be synchronized, and as a result, the user can obtain an optimal sound effect experience. In addition, when the media information includes video, the sound object in the video is tracked, and the audio track corresponding to the sound object throughout the video is rendered. This can also be applied to professional mixing post-production to improve the work efficiency of the mixing engineer.
[0011] Optionally, in a possible implementation of the first aspect, the reference position information in the foregoing steps includes the first position information of the sensor or the second position information selected by the user.
[0012] In this possible implementation, when the reference position information includes the first position information of the sensor, the user can perform real-time or subsequent dynamic rendering on the selected sound object based on the orientation or position given by the sensor. In this way, the control can assign specific spatial orientations and specific movements to the sound object so that the generation of interaction between the user and the audio is implemented to provide the user with a new experience. When the reference position information includes the second position information selected by the user, the user can control the selected sound object by using the drag method in the interface and perform real-time or subsequent dynamic rendering. Thereby, specific spatial orientations and specific movements can be assigned to the sound object so that the generation of interaction between the user and the audio is implemented to provide the user with a new experience. Also, when the user does not have a sensor, the sound image of the sound object may be further edited.
[0013] Optionally, in a possible implementation of the first aspect, the foregoing step is a step of determining the type of the playback device, the playback device being configured to play a target audio track, the target audio track being obtained based on the rendered first single object audio track, and the step of performing spatial rendering on the first single object audio track based on the first sound source position includes performing spatial rendering on the first single object audio track based on the first sound source position and the type of the playback device.
[0014] In this possible implementation, when spatial rendering is performed on an audio track, the type of playback device is taken into consideration. Different playback device types can correspond to different spatial rendering methods so that the spatial effect of subsequently playing the first single-object audio track rendered by the playback device is more realistic and accurate.
[0015] Optionally, in a possible implementation of the first aspect, the reference information in the foregoing steps includes media information. When the media information includes an image and the image includes a first sound object, the step of determining the first sound source position of the first sound object based on the reference information is a step of determining third position information of the first sound object in the image, where the third position information includes the two-dimensional coordinates and depth of the first sound object in the image, and a step of obtaining the first sound source position based on the third position information.
[0016] In this possible implementation, after the coordinates of the sound object and the single-object audio track are extracted with reference to the multimodal features of audio, video, and images, a 3D immersive feeling is obtained through rendering in a headset or loudspeaker environment. In this way, the audio and the image can be synchronized, and as a result, the user obtains an optimal sound effect experience. In addition, the technology of tracking and rendering the object audio in the entire video after a sound object is selected can also be applied to specialized mixing post-production to improve the working efficiency of the mixing engineer. The single-object audio track of the audio in the video is separated, and the sound objects in the video image are analyzed and tracked to obtain the movement information of the sound objects, and real-time or subsequent dynamic rendering is performed on the selected sound objects. In this way, the video image is matched with the sound source direction of the audio, and as a result, the user experience is improved.
[0017] Optionally, in a possible implementation of the first aspect, the reference information in the foregoing steps includes media information. When the media information includes the musical characteristics of the music that needs to be played within the multimedia file, the step of determining the first sound source position of the first sound object based on the reference information is a step of determining the first sound source position based on the association relationship and the musical characteristics, where the association relationship indicates the association between the musical characteristics and the first sound source position, and the step includes.
[0018] In this possible implementation, based on the musical characteristics of the music, orientation and dynamics settings are implemented for the extracted specific sound object, and as a result, the 3D rendering becomes more natural and the artistic elements are better reflected.
[0019] Optionally, in a possible implementation of the first aspect, the media information in the foregoing steps includes media information. When the media information includes the text that needs to be displayed within the multimedia file and the text includes position text related to a position, the step of determining the first sound source position of the first sound object based on the reference information includes the step of identifying the position text and the step of determining the first sound source position based on the position text.
[0020] In this possible implementation, the position text related to the position is identified, and a 3D immersion feeling is obtained through rendering in a headset or loudspeaker environment. In this way, a sense of space corresponding to the position text is achieved, and as a result, the user obtains an optimal sound effect experience.
[0021] Optionally, in a possible implementation of the first aspect, the reference information in the foregoing steps includes reference position information. When the reference position information includes first position information, before the step of determining the first sound source position of the first sound object based on the reference information, the method further includes a step of obtaining the first position information, where the first position information includes the first attitude angle of the sensor and the distance between the sensor and the playback device. The step of determining the first sound source position of the first sound object based on the reference information includes a step of converting the first position information into the first sound source position.
[0022] In this possible implementation, the user can perform real-time or subsequent dynamic rendering on the selected sound object based on the orientation provided by the sensor (i.e., the first attitude angle). In this case, the sensor is similar to a laser pointer, and the position pointed by the laser is the sound source position. In this way, the control can assign specific spatial orientations and specific movements to the sound object so that the generation of interaction between the user and the audio is implemented to provide the user with a new experience.
[0023] Optionally, in a possible implementation of the first aspect, the reference information in the foregoing steps includes reference position information. When the reference position information includes first position information, before the step of determining the first sound source position of the first sound object based on the reference information, the method further includes a step of obtaining the first position information, where the first position information includes the second attitude angle of the sensor and the acceleration of the sensor. The step of determining the first sound source position of the first sound object based on the reference information includes a step of converting the first position information into the first sound source position.
[0024] In this possible implementation, the user can control the sound object by using the actual position information of the sensor as the sound source position and perform real-time or subsequent dynamic rendering. In this way, the motion track of the sound object can be easily and fully controlled by the user, and as a result, the flexibility of editing is greatly improved.
[0025] Optionally, in a possible implementation of the first aspect, the reference information in the foregoing steps includes reference position information. When the reference position information includes second position information, before the step of determining the first sound source position of the first sound object based on the reference information, the method further includes providing a spherical view for the user to select, where the center of the circle of the spherical view is the user's position and the radius of the spherical view is the distance between the user's position and the playback device, and obtaining the second position information selected by the user within the spherical view. The step of determining the first sound source position of the first sound object based on the reference information includes converting the second position information into the first sound source position.
[0026] In this possible implementation, the user can use the spherical view to select the second position information (e.g., through operations such as tapping, dragging, sliding) to control the selected sound object and perform real-time or subsequent dynamic rendering, so that specific spatial orientation and specific movement can be assigned to the sound object to implement the generation of interaction between the user and the audio and provide a new experience for the user. In addition, when the user does not have a sensor, the sound image of the sound object can be further edited.
[0027] Optionally, in a possible implementation of the first aspect, the aforementioned step of obtaining the first single-object audio track based on the multimedia file is a step of separating the first single-object audio track from the original audio track in the multimedia file, where the original audio track is obtained by combining at least the first single-object audio track and the second single-object audio track, and the second single-object audio track corresponds to a second sound object, and the step is included.
[0028] In this possible implementation, when the original audio track is obtained by combining at least the first single-object audio track and the second single-object audio track, the first single-object audio track is separated, and as a result, spatial rendering can be performed on a specific sound object within the audio track. This can improve the user's audio editing ability and can be applied to the production of objects in music or movies and TV works. In this way, the user's controllability and reproducibility for music are improved.
[0029] Optionally, in a possible implementation of the first aspect, the aforementioned step of separating the first single-object audio track from the original audio track in the multimedia file includes the step of separating the first single-object audio track from the original audio track by using a trained separation network.
[0030] In this possible implementation, when the original audio track is obtained by combining at least a first single-object audio track and a second single-object audio track, the first single-object audio track is separated using a separation network, and as a result, spatial rendering can be performed on specific sound objects within the original audio track. This can improve the user's audio editing ability and can be applied to the object production of music or movies and TV works. In this way, the user's controllability and reproducibility for music are improved.
[0031] Optionally, in a possible implementation of the first aspect, the trained separation network in the foregoing steps is obtained by training the separation network using training data as the input of the separation network and using a value of a loss function smaller than a first threshold as a target. The training data includes a training audio track, and the training audio track is obtained by combining at least an initial third single-object audio track and an initial fourth single-object audio track. The initial third single-object audio track corresponds to a third sound object, the initial fourth single-object audio track corresponds to a fourth sound object, the third sound object and the first sound object have the same type, and the second sound object and the fourth sound object have the same type. The output of the separation network includes a third single-object audio track obtained through separation. The loss function indicates the difference between the third single-object audio track obtained through separation and the initial third single-object audio track.
[0032] In this possible implementation, the separation network is trained to reduce the value of the loss function, that is, to continuously reduce the difference between the third single-object audio track output by the separation network and the initial third single-object audio track. In this way, the single-object audio track separated by using the separation network is more accurate.
[0033] Optionally, in a possible implementation of the first aspect, the aforementioned step of performing spatial rendering on the first single-object audio track based on the first sound source position and the type of the playback device, when the playback device is a headset,
Number
[0034]
Number
[0035] In this possible implementation, when the playback device is a headset, the technical problem of how to obtain the rendered first single-object audio track is solved.
[0036] Optionally, in a possible implementation of the first aspect, the aforementioned step of performing spatial rendering on the first single-object audio track based on the first sound source position and the type of the playback device includes, when the playback device is an N-channel loudspeaker device,
Number
Number
Number
[0037]
Number
[0038] In this possible implementation, when the playback device is a loudspeaker device, the technical problem of how to obtain the rendered first single object audio track is solved.
[0039] Optionally, in a possible implementation of the first aspect, the aforementioned steps further include the steps of obtaining a target audio track based on the rendered first single object audio track, the original audio track in the multimedia file, and the type of the playback device, and sending the target audio track to the playback device, where the playback device is configured to play the target audio track.
[0040] In this possible implementation, the target audio track can be obtained. This helps to store the rendered audio track, facilitates subsequent playback, and reduces repeated rendering operations.
[0041] Optionally, in a possible implementation of the first aspect, the aforementioned step of obtaining a target audio track based on the rendered first single object audio track, the original audio track in the multimedia file, and the type of the playback device, when the type of the playback device is a headset,
Number
[0042] i represents the left channel or the right channel,
Number
[0043] In this possible implementation, when the playback device is a headset, the technical problem of how to obtain the target audio track is solved. This helps to store the rendered audio track, facilitates subsequent playback, and reduces the repeated rendering operation.
[0044] Optionally, in a possible implementation of the first aspect, the foregoing step of obtaining a target audio track based on the rendered first single object audio track, the original audio track in the multimedia file, and the type of the playback device includes, when the type of the playback device is an N-channel loudspeaker device,
Number
Number
Number
[0045] i represents the i-th channel among a plurality of channels,
Number
Number
Number
[0046] In this possible implementation, when the playback device is a loudspeaker device, the technical problem of how to obtain the target audio track is solved. This helps to store the rendered audio track, facilitates subsequent playback, and reduces the repeated rendering operation.
[0047] Optionally, in a possible implementation of the first aspect, the musical features in the foregoing steps include at least one of musical structure, musical emotion, and singing mode.
[0048] Optionally, in a possible implementation of the first aspect, the aforementioned steps further include separating a second single-object audio track from the multimedia file, determining a second sound source position of the second sound object based on the reference information, and performing spatial rendering on the second single-object audio track based on the second sound source position to obtain the rendered second single-object audio track.
[0049] In this possible implementation, at least two single-object audio tracks may be separated from the multimedia file, and corresponding spatial rendering is performed. This enhances the user's ability to edit specific sound objects in audio and can be applied to object production in music or movies and TV works. In this way, the controllability and reproducibility of music for the user are improved.
[0050] A second aspect of the embodiments of the present application provides a rendering method. This method may be applied to scenarios such as the production of music or movies and TV works, may be implemented by a rendering device, or may be implemented by components of a rendering device (such as a processor, a chip, a chip system, etc.). The method includes steps of obtaining a multimedia file, obtaining a first single-object audio track based on the multimedia file, where the first single-object audio track corresponds to a first sound object, displaying a user interface, where the user interface includes rendering mode options, determining an automatic rendering mode or an interactive rendering mode from the rendering mode options in response to a first operation of a user in the user interface, when the automatic rendering mode is determined, obtaining the rendered first single-object audio track in a preset mode, or when the interactive rendering mode is determined, obtaining reference position information in response to a second operation of the user, determining a first sound source position of the first sound object based on the reference position information, and rendering the first single-object audio track based on the first sound source position to obtain the rendered first single-object audio track.
[0051] In this embodiment of the present application, the rendering device determines an automatic rendering method or an interactive rendering method from rendering method options based on the user's first action. In one aspect, the rendering device can automatically obtain a first single-object audio track that has been rendered based on the user's first action. In another aspect, the spatial rendering of the audio track corresponding to the first sound object in the multimedia file can be implemented through the interaction between the rendering device and the user so that an immersive stereo sound effect is provided to the user.
[0052] Optionally, in a possible implementation of the second aspect, the pre-set method in the foregoing steps includes the steps of obtaining media information of the multimedia file, determining a first sound source position of the first sound object based on the media information, and rendering a first single-object audio track based on the first sound source position to obtain the rendered first single-object audio track.
[0053] In this possible implementation, Optionally, in a possible implementation of the second aspect, the media information in the foregoing steps includes at least one of text that needs to be displayed in the multimedia file, an image that needs to be displayed in the multimedia file, musical characteristics of music that needs to be played in the multimedia file, and a sound source type corresponding to the first sound object.
[0054] In this possible implementation, the rendering device determines the multimedia file to be processed through the interaction between the rendering device and the user, and as a result, the controllability and reproducibility of the user with respect to the music in the multimedia file are enhanced.
[0055] Optionally, in a possible implementation of the second aspect, the reference position information in the foregoing steps includes the first position information of the sensor or the second position information selected by the user.
[0056] In this possible implementation, when spatial rendering is performed on an audio track, the type of the playback device is determined based on the user's operation. Different playback device types can correspond to different spatial rendering methods so that the spatial effect of continuously playing the audio track rendered by the playback device is more realistic and accurate.
[0057] Optionally, in a possible implementation of the second aspect, when the media information includes an image and the image includes a first sound object, the foregoing steps of determining the first sound source position of the first sound object based on the media information are the steps of presenting the image and determining the third position information of the first sound object in the image, where the third position information includes the two-dimensional coordinates and depth of the first sound object in the image, and the step of obtaining the first sound source position based on the third position information.
[0058] In this possible implementation, the rendering device can automatically present the image, determine the sound object in the image, obtain the third position information of the sound object, and then obtain the first sound source position. In this way, the rendering device can automatically identify the multimedia file. When the multimedia file includes an image and the image includes a first sound object, the rendering device can automatically obtain the rendered first single object audio track. After the coordinates of the sound object and the single object audio track are automatically extracted, a 3D immersion feeling is obtained through rendering in a headset or loudspeaker environment. In this way, the sound and the image can be synchronized, and as a result, the user can obtain an optimal sound effect experience.
[0059] Optionally, in a possible implementation of the second aspect, the foregoing step of determining the third position information of the first sound object in the image includes the step of determining the third position information of the first sound object in response to a third action performed by the user on the image.
[0060] In this possible implementation, the user can select the first sound object from a plurality of sound objects in the presented image, that is, can select the rendered first single object audio track corresponding to the first sound object. The coordinates of the sound object and the single object audio track are extracted based on the user action, and the 3D immersion feeling is obtained through rendering in a headset or loudspeaker environment. In this way, the sound and the image can be synchronized, and as a result, the user can obtain an optimal sound effect experience.
[0061] Optionally, in a possible implementation of the second aspect, when the media information includes the musical characteristics of music that needs to be played in the multimedia file, the step of determining the first sound source position of the first sound object based on the media information includes the step of identifying the musical characteristics and the step of determining the first sound source position based on the association relationship and the musical characteristics, where the association relationship indicates the association between the musical characteristics and the first sound source position.
[0062] In this possible implementation, based on the musical characteristics of the music, orientation and dynamics settings are performed on the extracted specific sound object, and as a result, the 3D rendering becomes more natural and the artistic aspects are better reflected.
[0063] Optionally, in a possible implementation of the second aspect, when the media information includes text that needs to be displayed within a multimedia file and the text includes position text related to a position, the step of determining a first sound source position of a first sound object based on the media information includes the step of identifying the position text and the step of determining the first sound source position based on the position text.
[0064] In this possible implementation, position text related to a position is identified, and a 3D immersion feeling is obtained through rendering in a headset or loudspeaker environment. In this way, a sense of space corresponding to the position text is achieved, and as a result, the user obtains an optimal sound effect experience.
[0065] Optionally, in a possible implementation of the second aspect, when the reference position information includes the first position information, the step of obtaining the reference position information in response to a second action of the user includes the step of obtaining the first position information in response to a second action performed on the sensor by the user, and the first position information includes a first attitude angle of the sensor and a distance between the sensor and the playback device. The step of determining a first sound source position of a first sound object based on the reference position information includes the step of converting the first position information into the first sound source position.
[0066] In this possible implementation, the user can perform real-time or subsequent dynamic rendering on the selected sound object based on the orientation provided by the sensor (i.e., the first attitude angle). In this case, the sensor is similar to a laser pointer, and the position pointed to by the laser is the sound source position. In this way, the control can assign specific spatial orientations and specific movements to the sound object so that the generation of interaction between the user and the audio is implemented to provide a new experience for the user.
[0067] Optionally, in a possible implementation of the second aspect, when the reference position information includes the first position information, the step of obtaining the reference position information in response to the second action of the user includes the step of obtaining the first position information in response to the second action performed by the user on the sensor, and the first position information includes the second attitude angle of the sensor and the acceleration of the sensor. The step of determining the first sound source position of the first sound object based on the reference position information includes the step of converting the first position information into the first sound source position.
[0068] In this possible implementation, the user can control the sound object by using the actual position information of the sensor as the sound source position and perform real-time or subsequent dynamic rendering. In this way, the motion track of the sound object can be easily and completely controlled by the user, and as a result, the flexibility of editing is greatly improved.
[0069] Optionally, in a possible implementation of the second aspect, when the reference position information includes the second position information, the step of obtaining the reference position information in response to the second action of the user includes the step of presenting a spherical view, where the center of the circle of the spherical view is the position of the user and the radius of the spherical view is the distance between the position of the user and the playback device, and the step of determining the second position information within the spherical view in response to the second action of the user. The step of determining the first sound source position of the first sound object based on the reference position information includes the step of converting the second position information into the first sound source position.
[0070] In this possible implementation, the user can select second position information by using a spherical view (e.g., through operations such as tapping, dragging, sliding, etc.) to control the selected sound object and perform real-time or subsequent dynamic rendering, and specific spatial orientation and specific movement can be assigned to the sound object so that interaction generation between the user and the audio is implemented to provide a new experience for the user. In addition, when the user does not have a sensor, the sound image of the sound object can be further edited.
[0071] Optionally, in a possible implementation of the second aspect, the aforementioned step of obtaining a multimedia file includes determining a multimedia file from at least one stored multimedia file in response to a fourth action of the user.
[0072] In this possible implementation, the multimedia file can be determined from at least one stored multimedia file based on the user's selection to implement the rendering and production of a first single-object audio track corresponding to a first sound object in the multimedia file selected by the user. This improves the user experience.
[0073] Optionally, in a possible implementation of the second aspect, the user interface further includes playback device type options. The method further includes determining the type of the playback device from the playback device type options in response to a fifth action of the user. The step of rendering a first single-object audio track based on a first sound source position to obtain the rendered first single-object audio track includes rendering a first single-object audio track based on the first sound source position and type to obtain the rendered first single-object audio track.
[0074] In this possible implementation form, a rendering method suitable for the playback device used by the user is selected based on the type of the playback device used by the user. As a result, the rendering effect of the playback device is improved, and the 3D rendering becomes more natural.
[0075] Optionally, in a possible implementation form of the second aspect, the above-mentioned step of obtaining the first single-object audio track based on the multimedia file is a step of separating the first single-object audio track from the original audio track in the multimedia file, where the original audio track is obtained by combining at least the first single-object audio track and the second single-object audio track, and the second single-object audio track corresponds to a second sound object, and the step is included.
[0076] In this possible implementation form, when the original audio track is obtained by combining at least the first single-object audio track and the second single-object audio track, the first single-object audio track is separated, and as a result, spatial rendering can be performed on a specific sound object in the audio track. This can improve the user's audio editing ability and can be applied to the object production of music or movies and TV works. In this way, the user's controllability and reproducibility for music are improved.
[0077] In this possible implementation form, the first single-object audio track can be separated from the multimedia file in order to implement the rendering of the single-object audio track corresponding to a specific sound object in the multimedia file. This helps the user to perform audio production and improves the user experience.
[0078] In this possible implementation, the user can perform real-time or subsequent dynamic rendering on the selected sound object based on the orientation or position provided by the sensor. In this way, the control can assign specific spatial orientations and specific movements to the sound object so that the generation of interaction between the user and the audio is implemented to provide the user with a new experience.
[0079] In this possible implementation, the user can perform real-time or subsequent dynamic rendering on the selected sound object based on the orientation (i.e., the first attitude angle) provided by the sensor. In this case, the sensor is similar to a laser pointer, and the position indicated by the laser becomes the sound source position. In this way, the control can assign specific spatial orientations and specific movements to the sound object so that the generation of interaction between the user and the audio is implemented to provide the user with a new experience.
[0080] In this possible implementation, the sound object is controlled by using the actual position information of the sensor as the sound source position, and real-time or subsequent dynamic rendering is performed. In this way, the motion track of the sound object can be easily and completely controlled by the user, and as a result, the flexibility of editing is greatly improved.
[0081] In this possible implementation, the user can control the selected sound object by using the drag method in the interface and perform real-time or subsequent dynamic rendering. Thereby, specific spatial orientations and specific movements can be assigned to the sound object so that the generation of interaction between the user and the audio is implemented to provide the user with a new experience. Also, when the user does not have a sensor, the sound image of the sound object may be further edited.
[0082] In this possible implementation form, the rendering device can perform orientation and dynamics settings for the extracted specific sound objects based on the musical characteristics of the music. As a result, the audio track corresponding to the sound object becomes more natural in 3D rendering, and the artistic elements are better reflected.
[0083] In this possible implementation form, a 3D immersion feeling is obtained through rendering in a headset or loudspeaker environment. In this way, the audio and the image can be synchronized, and as a result, the user can obtain an optimal sound effect experience.
[0084] In this possible implementation form, after determining the sound object, the rendering device can automatically track the sound object in the video and render the audio track corresponding to the sound object throughout the video. This can also be applied to specialized mixing post-production to improve the work efficiency of the mixing engineer.
[0085] In this possible implementation form, the rendering device can determine the sound object in the image based on the user's fourth action, track the sound object in the image, and render the audio track corresponding to the sound object. In this way, the control can assign specific spatial orientations and specific movements to the sound object so that the generation of interaction between the user and the audio is implemented to provide the user with a new experience.
[0086] Optionally, in a possible implementation form of the second aspect, the musical characteristics in the foregoing steps include at least one of musical structure, musical emotion, and singing mode.
[0087] Optionally, in a possible implementation of the second aspect, the aforementioned steps further include separating a second single-object audio track from the original audio track, determining a second sound source position of the second sound object based on reference information, and performing spatial rendering on the second single-object audio track based on the second sound source position to obtain the rendered second single-object audio track.
[0088] In this possible implementation, at least two single-object audio tracks may be separated from the original audio track, and corresponding spatial rendering is performed. This can enhance the user's ability to edit specific sound objects in audio and can be applied to object production in music or movies and TV works. In this way, the controllability and reproducibility of music for the user are improved.
[0089] The third aspect of the present application provides a rendering device. The rendering device can be applied to scenarios such as music or movie production and TV works. The rendering device obtains a first single-object audio track based on a multimedia file, where the first single-object audio track is configured to correspond to a first sound object, and an acquisition unit determines a first sound source position of the first sound object based on reference information, where the reference information includes reference position information and / or media information of the multimedia file, and the reference position information is configured to indicate the first sound source position, and a determination unit and a rendering unit configured to perform spatial rendering on the first single-object audio track based on the first sound source position to obtain the rendered first single-object audio track.
[0090] Optionally, in a possible implementation of the third aspect, the media information includes at least one of text that needs to be displayed in the multimedia file, an image that needs to be displayed in the multimedia file, musical features of music that needs to be played in the multimedia file, and a sound source type corresponding to the first sound object.
[0091] Optionally, in a possible implementation of the third aspect, the reference position information includes the first position information of the sensor or the second position information selected by the user.
[0092] Optionally, in a possible implementation of the third aspect, the determination unit is further configured to determine the type of the playback device, the playback device is configured to play the target audio track, and the target audio track is obtained based on the rendered first single object audio track. The rendering unit is specifically configured to perform spatial rendering on the first single object audio track based on the first sound source position and the type of the playback device.
[0093] Optionally, in a possible implementation of the third aspect, the reference information includes the media information. When the media information includes an image and the image includes the first sound object, the determination unit is specifically configured to determine the third position information of the first sound object in the image, and the third position information includes the two-dimensional coordinates and depth of the first sound object in the image. The determination unit is specifically configured to obtain the first sound source position based on the third position information.
[0094] Optionally, in a possible implementation of the third aspect, the reference information includes the media information. When the media information includes the musical features of music that need to be played in the multimedia file, the determination unit is specifically configured to determine the first sound source position based on the association relationship and the musical features, and the association relationship indicates the association between the musical features and the first sound source position.
[0095] Optionally, in a possible implementation form of the third aspect, the media information includes media information. When the media information includes text that needs to be displayed in a multimedia file and the text includes position text related to a position, the determination unit is specifically configured to identify the position text. The determination unit is specifically configured to determine a first sound source position based on the position text.
[0096] Optionally, in a possible implementation form of the third aspect, the reference information includes reference position information. When the reference position information includes first position information, the acquisition unit is further configured to acquire the first position information, and the first position information includes a first attitude angle of the sensor and a distance between the sensor and the playback device. The determination unit is specifically configured to convert the first position information into a first sound source position.
[0097] Optionally, in a possible implementation form of the third aspect, the reference information includes reference position information. When the reference position information includes first position information, the acquisition unit is further configured to acquire the first position information, and the first position information includes a second attitude angle of the sensor and the acceleration of the sensor. The determination unit is specifically configured to convert the first position information into a first sound source position.
[0098] Optionally, in a possible implementation form of the third aspect, the reference information includes reference position information. When the reference position information includes second position information, the rendering device further includes a providing unit configured to provide a spherical view for the user to select, the center of the circle of the spherical view is the position of the user, and the radius of the spherical view is the distance between the position of the user and the playback device. The acquisition unit is further configured to acquire second position information selected by the user within the spherical view. The determination unit is specifically configured to convert the second position information into a first sound source position.
[0099] Optionally, in a possible implementation of the third aspect, the acquisition unit is specifically configured to separate a first single object audio track from the original audio track in the multimedia file, and the original audio track is obtained by combining at least a first single object audio track and a second single object audio track, and the second single object audio track corresponds to a second sound object.
[0100] Optionally, in a possible implementation of the third aspect, the acquisition unit is specifically configured to separate a first single object audio track from the original audio track by using a trained separation network.
[0101] Optionally, in a possible implementation of the third aspect, the trained separation network is obtained by training the separation network by using training data as the input of the separation network and using a value of a loss function smaller than a first threshold as the target. The training data includes training audio tracks, and the training audio tracks are obtained by combining at least an initial third single object audio track and an initial fourth single object audio track. The initial third single object audio track corresponds to a third sound object, the initial fourth single object audio track corresponds to a fourth sound object, the third sound object and the first sound object have the same type, and the second sound object and the fourth sound object have the same type. The output of the separation network includes a third single object audio track obtained through separation. The loss function indicates the difference between the third single object audio track obtained through separation and the initial third single object audio track.
[0102] Optionally, in a possible implementation of the third aspect, when the playback device is a headset, the acquisition unit is specifically configured to acquire the rendered first single-object audio track in accordance with [Number] .
[0103] [Number] represents the rendered first single-object audio track, S represents the sound object of the multimedia file, the sound object includes the first sound object, i represents the left channel or the right channel, and a s (t) represents the adjustment coefficient of the first sound object at instant t, and h i,s (t) represents the head-related transfer function (HRTF) filter coefficient of the left channel or the right channel corresponding to the first sound object at instant t. The HRTF filter coefficient is related to the first sound source position, and o s (t) represents the first single-object audio track at instant t, and τ represents the integral term.
[0104] Optionally, in a possible implementation of the third aspect, when the playback device is N loudspeaker devices, the acquisition unit is specifically configured to acquire the rendered first single-object audio track in accordance with [Number] , [Number] where [Number] .
[0105] [Number] represents the rendered first single-object audio track, i represents the i-th channel among a plurality of channels, S represents the sound object of the multimedia file, the sound object includes the first sound object, and a s (t) represents the adjustment coefficient of the first sound object at instant t, and g s (t) represents the translation coefficient of the first sound object at instant t, and o s (t) represents the first single-object audio track at instant t, and λ i represents the azimuth angle obtained when the calibrator calibrates the i-th loudspeaker device, and Φ i represents the tilt angle obtained when the calibrator calibrates the i-th loudspeaker device, and r i represents the distance between the i-th loudspeaker device and the calibrator, N is a positive integer, i is a positive integer, i ≤ N, and the first sound source position is within the tetrahedron formed by N loudspeaker devices.
[0106] Optionally, in a possible implementation of the third aspect, the acquisition unit is further configured to acquire a target audio track based on the rendered first single-object audio track and the original audio track in the multimedia file. The rendering device further includes a transmission unit configured to transmit the target audio track to the playback device, and the playback device is configured to play the target audio track.
[0107] Optionally, in a possible implementation of the third aspect, when the playback device is a headset, the acquisition unit [Number] is specifically configured to acquire the target audio track according to
[0108] i represents the left or right channel,
Number
Number
Number
[0109] Optionally, in a possible implementation of the third aspect, when the playback device is an N-channel loudspeaker device, the acquisition unit [Number] is particularly configured to obtain a target audio track according to [Number] and is [Number] as follows.
[0110] i represents the i-th channel among a plurality of channels, [Number] represents the target audio track at instant t, X i (t) represents the original audio track at instant t, [Number] represents the first single object audio track not being rendered at instant t, [Number] represents the first single object audio track that has been rendered, a s (t) represents the adjustment coefficient of the first sound object at instant t, g s (t) represents the translation coefficient of the first sound object at instant t, g i,s (t) is g s (t) represents the i-th row in g s(t) represents the first single-object audio track at instant t, S1 represents the sound object that needs to be replaced in the original audio track. When the first sound object replaces the sound object in the original audio track, S1 represents the null set. S2 represents the sound object added to the target audio track compared to the original audio track. When the first sound object is a copy of the sound object in the original audio track, S2 represents the null set. S1 and / or S2 represent the sound objects of the multimedia file. The sound object includes the first sound object, λ i represents the azimuth angle obtained when the calibrator calibrates the i-th loudspeaker device, Φ i represents the tilt angle obtained when the calibrator calibrates the i-th loudspeaker device, r i represents the distance between the i-th loudspeaker device and the calibrator. N is a positive integer, i is a positive integer, and i ≤ N. The first sound source position is within the tetrahedron formed by N loudspeaker devices.
[0111] The fourth aspect of this application provides a rendering device. The rendering device can be applied to scenarios such as music or movie production and TV works. The rendering device is an acquisition unit configured to acquire a multimedia file, wherein the acquisition unit is further configured to acquire a first single-object audio track based on the multimedia file, and the first single-object audio track is further configured to correspond to the first sound object, the acquisition unit; a display unit configured to display a user interface, where the user interface includes rendering mode options, the display unit; A determination unit configured to determine an automatic rendering mode or an interactive rendering mode from rendering mode options in response to a first operation of a user in a user interface, and The acquisition unit is further configured to acquire a first single object audio track rendered in a preset manner when the determination unit determines the automatic rendering mode, or The acquisition unit, when the determination unit determines the interactive rendering mode, acquires reference position information in response to a second operation of the user, determines a first sound source position of a first sound object based on the reference position information, and is further configured to render a first single object audio track based on the first sound source position to acquire the rendered first single object audio track.
[0112] Optionally, in a possible implementation form of the fourth aspect, the preset manner includes that the acquisition unit is further configured to acquire media information of a multimedia file, the determination unit is further configured to determine a first sound source position of a first sound object based on the media information, and the acquisition unit is further configured to render a first single object audio track based on the first sound source position to acquire the rendered first single object audio track.
[0113] Optionally, in a possible implementation form of the fourth aspect, the media information includes at least one of text that needs to be displayed in the multimedia file, an image that needs to be displayed in the multimedia file, musical characteristics of music that needs to be played in the multimedia file, and a sound source type corresponding to the first sound object.
[0114] Optionally, in a possible implementation form of the fourth aspect, the reference position information includes first position information of a sensor or second position information selected by the user.
[0115] Optionally, in a possible implementation of the fourth aspect, when the media information includes an image and the image includes a first sound object, the determination unit is specifically configured to present the image. The determination unit is specifically configured to determine third position information of the first sound object in the image, and the third position information includes the two-dimensional coordinates and depth of the first sound object in the image. The determination unit is specifically configured to obtain a first sound source position based on the third position information.
[0116] Optionally, in a possible implementation of the fourth aspect, the determination unit is specifically configured to determine third position information of the first sound object in response to a third operation performed on the image by the user.
[0117] Optionally, in a possible implementation of the fourth aspect, when the media information includes the musical characteristics of music that needs to be played in a multimedia file, the determination unit is specifically configured to identify the musical characteristics.
[0118] The determination unit is specifically configured to determine a first sound source position based on an association relationship and musical characteristics, and the association relationship indicates an association between the musical characteristics and the first sound source position.
[0119] Optionally, in a possible implementation of the fourth aspect, when the media information includes text that needs to be displayed in a multimedia file and the text includes position-related position text, the determination unit is specifically configured to identify the position text. The determination unit is specifically configured to determine a first sound source position based on the position text.
[0120] Optionally, in a possible implementation of the fourth aspect, when the reference position information includes the first position information, the determination unit is specifically configured to obtain the first position information in response to a second operation performed on the sensor by the user, and the first position information includes the first attitude angle of the sensor and the distance between the sensor and the playback device. The determination unit is specifically configured to convert the first position information into the first sound source position.
[0121] Optionally, in a possible implementation of the fourth aspect, when the reference position information includes the first position information, the determination unit is specifically configured to obtain the first position information in response to a second operation performed on the sensor by the user, and the first position information includes the second attitude angle of the sensor and the acceleration of the sensor. The determination unit is specifically configured to convert the first position information into the first sound source position.
[0122] Optionally, in a possible implementation of the fourth aspect, when the reference position information includes the second position information, the determination unit is specifically configured to present a spherical view, the center of the circle of the spherical view is the position of the user, and the radius of the spherical view is the distance between the position of the user and the playback device. The determination unit is specifically configured to determine the second position information within the spherical view in response to the second operation of the user. The determination unit is specifically configured to convert the second position information into the first sound source position.
[0123] Optionally, in a possible implementation of the fourth aspect, the acquisition unit is specifically configured to determine a multimedia file from at least one stored multimedia file in response to a fourth operation of the user.
[0124] Optionally, in a possible implementation of the fourth aspect, the user interface further includes playback device type options. The determination unit is further configured to determine the type of the playback device from the playback device type options in response to a fifth action of the user. The acquisition unit is specifically configured to render a first single object audio track based on a first sound source position and type in order to acquire the rendered first single object audio track.
[0125] Optionally, in a possible implementation of the fourth aspect, the acquisition unit is specifically configured to separate a first single object audio track from an original audio track in a multimedia file, where the original audio track is obtained by combining at least a first single object audio track and a second single object audio track, and the second single object audio track corresponds to a second sound object.
[0126] Optionally, in a possible implementation of the fourth aspect, the musical feature includes at least one of a musical structure, a musical emotion, and a singing mode.
[0127] Optionally, in a possible implementation of the fourth aspect, the acquisition unit further separates a second single object audio track from the multimedia file, determines a second sound source position of the second sound object, and performs spatial rendering on the second single object audio track based on the second sound source position to acquire the rendered second single object audio track.
[0128] The fifth aspect of this application provides a rendering device. The rendering device implements the method in any one of the first aspect or a possible implementation of the first aspect, or implements the method in any one of the second aspect or a possible implementation of the second aspect.
[0129] A sixth aspect of the present application provides a rendering device including a processor. The processor is coupled to a memory. The memory is configured to store a program or instructions. When the program or instructions are executed by the processor, the rendering device is enabled to implement the method in any one of the first aspect or possible implementations of the first aspect, or the rendering device is enabled to implement the method in any one of the second aspect or possible implementations of the second aspect.
[0130] A seventh aspect of the present application provides a computer-readable medium. The computer-readable medium stores a computer program or instructions. When the computer program or instructions are implemented on a computer, the computer is enabled to implement the method in any one of the first aspect or possible implementations of the first aspect, or the computer is enabled to implement the method in any one of the second aspect or possible implementations of the second aspect.
[0131] An eighth aspect of the present application provides a computer program product. When the computer program product is executed on a computer, the computer is enabled to implement the method of any one of the first aspect or possible implementations of the first aspect, or the computer is enabled to implement the method of any one of the second aspect or possible implementations of the second aspect.
[0132] Regarding the technical effects brought about by any one of the third aspect, the fifth aspect, the sixth aspect, the seventh aspect, or the eighth aspect, or possible implementations of the third aspect, the fifth aspect, the sixth aspect, the seventh aspect, or the eighth aspect, please refer to the technical effects brought about by the first aspect or different possible implementations of the first aspect. Details are not repeated herein.
[0133] Regarding the technical effects brought about by any one of the fourth aspect, the fifth aspect, the sixth aspect, the seventh aspect, or the eighth aspect, or the possible implementation forms of the fourth aspect, the fifth aspect, the sixth aspect, the seventh aspect, or the eighth aspect, please refer to the technical effects brought about by the second aspect or different possible implementation forms of the second aspect. Details are not repeated in this specification.
[0134] From the above technical solutions, it can be seen that the embodiments of the present application have the following advantages: The first single-object audio track is obtained based on a multimedia file, the first single-object audio track corresponds to the first sound object, the first sound source position of the first sound object is determined based on reference information, and in order to obtain the rendered first single-object audio track, spatial rendering is performed on the first single-object audio track based on the first sound source position. The stereo spatial sense of the first single-object audio track corresponding to the first sound object in the multimedia file can be improved, and as a result, an immersive stereo sound effect is provided to the user.
[0135] To more clearly explain the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings for explaining the embodiments or the prior art will be briefly described below. The accompanying drawings in the following description only show some embodiments of the present invention, and it is obvious that those skilled in the art can further derive other drawings from these accompanying drawings without creative efforts.
Brief Description of the Drawings
[0136]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Figure 34
Figure 35
Figure 36
Figure 37
Figure 38
Figure 39
Figure 40
Figure 41
Figure 42
Figure 43
Figure 44
Figure 45
Figure 46
Figure 47
Figure 48
Figure 49
Figure 50
Figure 51
Figure 52
Figure 53
Figure 54
Figure 55
Figure 56
Figure 57
Figure 58
Figure 59
Figure 60
Figure 61
Figure 62
Mode for Carrying Out the Invention
[0137] Embodiments of the present application provide a rendering method for improving the stereo spatial perception of a first single - object audio track corresponding to a first sound object in a multimedia file and providing an immersive stereo sound effect to a user.
[0138] Hereinafter, with reference to the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be described. It is obvious that the described embodiments are only a part, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0139] For ease of understanding, the mainly related terms and concepts in the embodiments of this application will be first described below.
[0140] 1. Neural Network A neural network may include neurons. A neuron may be an arithmetic unit that takes X s and a section of 1 as inputs. The output of the arithmetic unit may be as follows:
Equation
[0141] where s = 1, 2, …, n, and n is a natural number greater than 1. W s represents the weight of X s . b represents the bias of the neuron. f represents the activation functions of the neuron used to introduce non-linear features into the neural network and convert the input signal in the neuron into an output signal. The output signal of the activation function can be used as the input of the next convolutional layer. The activation function may be a sigmoid function. A neural network is a network formed by connecting a large number of single neurons to each other. Specifically, the output of a neuron can be the input of another neuron. The input of each neuron can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field may be a region containing a plurality of neurons.
[0142] 2. Deep Neural Network A deep neural network (DNN), also referred to as a multi-layer neural network, can be understood as a neural network having multiple hidden layers. In this specification, there is no special criterion for "a plurality of". A DNN is divided based on the positions of different layers, and the neural network of a DNN can be divided into three types: an input layer, a hidden layer, and an output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the middle layers are hidden layers. The layers are fully connected. Specifically, any neuron in the i-th layer is necessarily connected to any neuron in the (i + 1)-th layer. Naturally, a deep neural network may not include a hidden layer. This is not particularly limited in this specification.
[0143] The operations in each layer of a deep neural network are described by the mathematical formula
Number
Number
Number
[0144] 3. Convolutional Network A convolutional neural network (CNN) is a deep neural network with a convolutional structure. The convolutional network includes a feature extractor that contains convolutional layers and sampling sub-layers. The feature extractor can be regarded as a filter. The convolution process can be regarded as performing convolution on an input image or a convolutional feature map by using the same trainable filter. A convolutional layer is a neuron layer in a convolutional network that performs a convolution process on an input signal. At the convolutional layer of the convolutional network, one neuron may be connected to only a part of the neurons in the adjacent layer. One convolutional layer usually includes a plurality of feature planes, and each feature plane may include a plurality of neurons arranged in a rectangular shape. Neurons in the same feature plane share weights, and the weights shared herein are convolutional kernels. Weight sharing can be understood as the fact that the image information extraction method is independent of position. The principle suggested herein is that the statistical information of a part of an image is the same as that of another part. This means that the image information learned from one part can also be used in other parts. Therefore, the same image information obtained through learning can be used for all positions in the image. In the same convolutional layer, a plurality of convolutional kernels are used to extract different image information. Usually, the larger the number of convolutional kernels, the richer the image information reflected through the convolution operation.
[0145] The convolutional kernel can be initialized in the form of a random-sized matrix. In the process of training the convolutional network, the convolutional kernel can obtain appropriate weights through learning. In addition, the advantage directly brought about by weight sharing is that the connection between layers of the convolutional network is reduced, and the risk of overfitting is also reduced. In the embodiments of the present application, networks such as a separation network, an identification network, a detection network, and a deep estimation network may all be CNNs.
[0146] 4. Recurrent Neural Network (RNN) In a conventional neural network model, the layers are fully connected and the nodes between the layers are not connected. However, this general neural network cannot solve many problems. For example, the previous and next words in a sentence are not independent, and since the previous word is usually required to be used, the problem of predicting the next word in a sentence cannot be solved. A recurrent neural network (RNN) means that the current output of a sequence is related to the previous output. The specific expression form is that the network remembers the previous information, stores that information in the internal state of the network, and applies that information to the calculation of the current output.
[0147] 5. Loss Function In the process of training a deep neural network, since the output of the deep neural network is expected to be as close as possible to the actually expected value for the predicted value, the current predicted value of the network can be compared with the actually expected target value, and then the matrix vectors in each layer of the neural network are updated based on the difference between the current predicted value and the target value (usually, there is an initialization process before the first update, that is, the parameters are preconfigured for each layer of the neural network). For example, when the predicted value of the network is large, the matrix vector is adjusted to make the predicted value smaller, and the adjustment continues to be carried out until the neural network can predict the actually expected target value. Therefore, it is necessary to predefine "how to obtain the difference between the predicted value and the target value through comparison". This is the loss function or objective function. The loss function and the objective function are important formulas for measuring the difference between the predicted value and the target value. The loss function is used as an example. The larger the output value (loss) of the loss function, the greater the difference indicates. Therefore, the training of a deep neural network is a process of minimizing the loss as much as possible.
[0148] 6. Head Transfer Function Head Related Transfer Function (HRTF): Sound waves transmitted by a sound source reach the two ears after being scattered by the head, pinna, torso, etc. The physical process can be regarded as a linear time-invariant acoustic filtering system, and the characteristics of the process can be described by using the HRTF. In other words, the HRTF describes the process of transmitting sound waves from the sound source to the two ears. A clearer explanation is as follows: If the audio signal transmitted by the sound source is X, and the corresponding audio signal after the audio signal X is transmitted to a preset position is Y, then X * Z = Y (the convolution of X and Z is equal to Y), and Z represents the HRTF.
[0149] 7. Audio Track An audio track is a track for recording audio data. Each audio track has one or more attribute parameters. The attribute parameters include audio format, bit rate, dubbing language, sound effects, number of channels, volume, etc. When the audio data is multi-audio track data, two different audio tracks have at least one different attribute parameter, or at least one attribute parameter of two different audio tracks has different values. An audio track can be a single audio track or a multi-audio track (also called a mixed audio track). A single audio track can correspond to one or more sound objects, and a multi-audio track includes at least two single audio tracks. Generally, one single object audio track corresponds to one sound object.
[0150] 8. Short-Time Fourier Transform The central concept of the short-time Fourier transform (STFT), also known as the short-term Fourier transform, is "windowing". Specifically, the entire time-domain process is divided into a number of small processes of equal length, each small process is approximately stable, and then the fast Fourier transform (FFT) is performed on each small process.
[0151] Hereinafter, the system architecture provided in the embodiments of the present application will be described.
[0152] Refer to FIG. 1. One embodiment of the present invention provides a system architecture 100. As shown in the system architecture 100, the data collection device 160 is configured to collect training data. In this embodiment of the present application, the training data includes multimedia files, and the multimedia files include original audio tracks, and the original audio tracks correspond to at least one sound object. The training data is stored in the database 130. The training device 120 obtains the target model / rule 101 through training based on the training data maintained in the database 130. Hereinafter, using Embodiment 1, how the training device 120 obtains the target model / rule 101 based on the training data will be described in more detail. The target model / rule 101 can be used to implement the rendering method provided in the embodiments of the present application. The target model / rule 101 has multiple cases. In the case of the target model / rule 101 (when the target model / rule 101 is the first model), by inputting the multimedia file into the target model / rule 101, a first single-object audio track corresponding to the first sound object can be obtained. In another case of the target model / rule 101 (when the target model / rule 101 is the second model), the first single-object audio track corresponding to the first sound object can be obtained by inputting the multimedia file into the target model / rule 101 after the relevant preprocessing. The target model / rule 101 in this embodiment of the present application can specifically include a separation network, and can further include an identification network, a detection network, a deep estimation network, etc. This is not particularly limited herein. In this embodiment provided in the present application, the separation network is obtained through training by using the training data. It should be noted that in actual use, the training data maintained in the database 130 is not necessarily all collected by the data collection device 160, or may be received from another device.The training device 120 does not necessarily have to train the target model / rule 101 based entirely on the training data maintained in the database 130, or it may also obtain the training data from the cloud or another location and perform model training. It should be further noted that the foregoing description should not be construed as a limitation to this embodiment of the present application.
[0153] The target model / rule 101 obtained through the training by the training device 120 can be applied to different systems or devices, for example, the execution device 110 shown in FIG. 1. The execution device 110 may be a terminal device, such as a mobile phone terminal, a tablet computer, a laptop computer, augmented reality (AR) / virtual reality (VR), or an in-vehicle terminal device, or it may be a server, a cloud, etc. In FIG. 1, the execution device 110 is configured to include an I / O interface 112 configured to exchange data with an external device. The user can input data to the I / O interface 112 by using the client device 140. In this embodiment of the present application, the input data may include a multimedia file, may be input by the user, or may be uploaded by the user by using an audio device, or may of course be from a database. This is not particularly limited in this specification.
[0154] The preprocessing module 113 is configured to perform preprocessing based on the multimedia file received by the I / O interface 112. In this embodiment of the present application, the preprocessing module 113 may be configured to perform a short-time Fourier transform process on the audio track in the multimedia file to obtain a spectrogram.
[0155] In the process where the execution device 110 preprocesses the input data, or in the process where the calculation module 111 of the execution device 110 performs related processes such as calculation, the execution device 110 can call data, code, etc. in the data storage system 150 for the corresponding process, and can store the data, instructions, etc. obtained through the corresponding process in the data storage system 150.
[0156] Finally, the I / O interface 112 returns the processing result, for example, the aforementioned acquired first single object audio track corresponding to the first sound object, to the client device 140 in order to provide the processing result to the user.
[0157] It should be noted that the training device 120 can generate corresponding target models / rules 101 for different targets or different tasks based on different training data. The corresponding target models / rules 101 can be used to implement the aforementioned targets or complete the aforementioned tasks to provide the necessary results to the user.
[0158] In the case shown in FIG. 1, the user can manually provide input data to the interface provided by the I / O interface 112. In another case, the client device 140 can automatically send the input data to the I / O interface 112. If it is required that the client device 140 needs to obtain authorization from the user to automatically send the input data, the user may set the corresponding permission on the client device 140. The user can view the results output by the execution device 110 on the client device 140. Specifically, it may be presented in forms such as display, sound, action, etc. The client device 140 can collect the input data input to the I / O interface 112 shown in the figure and the output result output from the I / O interface 112 as new sample data, and can alternatively be used as a data collection end to store the new sample data in the database 130. Of course, the client device 140 may alternatively not perform the collection. Instead, the I / O interface 112 directly stores the input data input to the I / O interface 112 shown in the figure and the output result output from the I / O interface 112 in the database 130 as new sample data.
[0159] Note that FIG. 1 is only a schematic diagram of the system architecture according to an embodiment of the present invention. The positional relationship among the devices, components, modules, etc. shown in the figure does not constitute a limitation. For example, in FIG. 1, the data storage system 150 is an external memory for the execution device 110. In another case, the data storage system 150 may alternatively be disposed within the execution device 110.
[0160] As shown in FIG. 1, the target model / rule 101 is obtained through training by the training device 120. The target model / rule 101 can be the separation network in this embodiment of the present application. In particular, in the network provided in this embodiment of the present application, the separation network can be a convolutional network or a recurrent neural network.
[0161] Since CNN is a general neural network, hereinafter, the structure of CNN will be focused on in detail with reference to FIG. 2. As described in the explanation of the above basic concept, the convolutional network is a deep neural network having a convolutional structure and is a deep learning architecture. In the deep learning architecture, multilayer learning is performed at different abstraction levels by using a machine learning algorithm. As a deep learning architecture, CNN is a feed-forward artificial neural network. Neurons in the feed-forward artificial neural network can respond to the image input to the feed-forward artificial neural network.
[0162] As shown in FIG. 2, the convolutional network (CNN) 100 may include an input layer 110, a convolutional layer / pooling layer 120, and a neural network layer 130. The pooling layer is optional.
[0163] Convolutional layer / pooling layer 120 Convolutional layer As shown in FIG. 2, for example, the convolutional layer / pooling layer 120 may include layers 121 to 126. In one implementation, layer 121 is a convolutional layer, layer 122 is a pooling layer, layer 123 is a convolutional layer, layer 124 is a pooling layer, layer 125 is a convolutional layer, and layer 126 is a pooling layer. In another implementation, layers 121 and 122 are convolutional layers, layer 123 is a pooling layer, layers 124 and 125 are convolutional layers, and layer 126 is a pooling layer. That is, the output of the convolutional layer can be used as the input of the subsequent pooling layer or as the input of another convolutional layer to continue the convolutional operation.
[0164] The convolutional layer 121 is used as an example. The convolutional layer 121 may include a plurality of convolutional operators. The convolutional operator is also referred to as a kernel. In image processing, the convolutional operator functions as a filter that extracts specific information from the input image matrix. The convolutional operator may essentially be a weight matrix, and the weight matrix is usually predefined. In the process of performing a convolution operation on an image, the weight matrix is usually used to process pixels at the granularity level of one pixel (or two pixels depending on the value of the stride) in the horizontal direction on the input image in order to extract specific features from the image. The size of the weight matrix needs to be associated with the size of the image. Note that the depth dimension of the weight matrix is the same as the depth dimension of the input image. In the process of performing the convolution operation, the weight matrix extends across the entire depth of the input image. Therefore, by performing a convolution with a single weight matrix, a convolution output of a single depth dimension is generated. However, in most cases, instead of a single weight matrix, a plurality of weight matrices having the same dimensions are used. The outputs of the weight matrices are stacked to form the depth dimension of the convolved image. Different weight matrices can be used to extract different features of the image. For example, one weight matrix is used to extract edge information of the image, another weight matrix is used to extract a specific color of the image, and yet another weight matrix is used to blur unnecessary noise in the image. Since the plurality of weight matrices have the same dimensions, the feature maps extracted using the plurality of weight matrices having the same dimensions also have the same dimensions. Then, the plurality of extracted feature maps having the same dimensions are combined to form the output of the convolution operation.
[0165] The weight values within the weight matrix need to be obtained through large-scale training in actual applications. In order to assist the convolutional network 100 in performing correct predictions, a weight matrix formed by using the weight values obtained through training is used to extract information from the input picture.
[0166] When the convolutional network 100 includes a plurality of convolutional layers, a large number of general features are usually extracted in the initial convolutional layer (for example, convolutional layer 121). The general features may also be referred to as low-level features. As the depth of the convolutional network 100 increases, the features extracted in later convolutional layers (for example, convolutional layer 126) are more complex and are, for example, high-level semantic features. Features with higher semantics are more applicable to the problem to be solved.
[0167] Pooling layer Since the number of training parameters usually needs to be reduced, the pooling layer usually needs to be periodically introduced after the convolutional layer. Specifically, for the 120 layers 121 to 126 shown in FIG. 2, one convolutional layer may be followed by one pooling layer, or a plurality of convolutional layers may be followed by one or more pooling layers. In the image processing process, the pooling layer is only used to reduce the spatial size of the image. The pooling layer may include an average pooling operator and / or a max pooling operator to perform sampling on the input image to obtain an image with a small size. The average pooling operator can calculate the pixel values within a specific range in the image to generate an average value. The max pooling operator can be used to select the pixel with the maximum value within a specific range as the max pooling result. In addition, similar to the case where the size of the weight matrix in the convolutional layer needs to be associated with the size of the image, the operator in the pooling layer also needs to be associated with the size of the image. The size of the image output after processing in the pooling layer may be smaller than the size of the image input to the pooling layer. Each pixel in the image output from the pooling layer represents the average value or the maximum value of the corresponding sub-region of the image input to the pooling layer.
[0168] Neural network layer 130 After the processing performed in the convolutional layer / pooling layer 120, the convolutional network 100 is not ready to output the required output information. As described above, in the convolutional layer / pooling layer 120, only features are extracted and the parameters resulting from the input image are reduced. However, in order to generate the final output information (the required class information or other relevant information), the convolutional network 100 needs to use the neural network layer 130 to generate the output of one required class or the output of a group of required classes. Therefore, the neural network layer 130 may include a plurality of hidden layers (131, 132, …, 13n in FIG. 2) and an output layer 140. The parameters included in the plurality of hidden layers can be obtained through pre-training based on the relevant training data of a specific task type. For example, the task type may include multi-audio track separation, image recognition, image classification, super-resolution image reconstruction, and the like.
[0169] The plurality of hidden layers within the neural network layer 130 are followed by the output layer 140, that is, the last layer of the entire convolutional network 100. The output layer 140 has a loss function similar to categorical cross-entropy, and the loss function is specifically used to calculate the prediction error. When the forward propagation of the entire convolutional network 100 (for example, the propagation from layer 110 to layer 140 in FIG. 2 is the forward propagation) is completed, the backward propagation (for example, the propagation from layer 140 to layer 110 in FIG. 2 is the backward propagation) is started to update the weight values and biases of the above-described layers, and reduce the error between the loss of the convolutional network 100 and the result output by the convolutional network 100 through the output layer and the ideal result.
[0170] It should be noted that the convolutional network 100 shown in FIG. 2 is only used as an example of a convolutional network. In a specific application, the convolutional network may alternatively exist in the form of another network model, for example, a network model in which a plurality of convolutional layers / pooling layers are in parallel as shown in FIG. 3, and the extracted features are all input into the neural network layer 130 for processing.
[0171] Hereinafter, the hardware structure of the chip provided in the embodiments of the present application will be described.
[0172] FIG. 4 shows the hardware structure of a chip provided in an embodiment of the present invention. The chip includes a neural network processing unit 40. The chip may be disposed in the execution device 110 shown in FIG. 1 to complete the computing operation of the computing module 111. The chip may alternatively be disposed in the training device 120 shown in FIG. 1 to complete the training operation of the training device 120 and output the target model / rule 101. All algorithms of the layers in the convolutional network shown in FIG. 2 may be implemented in the chip shown in FIG. 4.
[0173] The neural network processing unit 40 may be any processor suitable for large-scale exclusive OR operation processing, for example, a neural-network processing unit (NPU), a tensor processing unit (TPU), a graphics processing unit (GPU), etc. NPU is used as an example. The neural network processing unit NPU 40 is mounted on a host central processing unit (host CPU) as a coprocessor. The host CPU assigns tasks. The core part of the NPU is the arithmetic circuit 403, and the controller 404 controls the arithmetic circuit 403 to fetch data in the memory (weight memory or input memory) and perform arithmetic operations.
[0174] In some implementation forms, the arithmetic circuit 403 includes a plurality of processing engines (PEs) internally. In some implementation forms, the arithmetic circuit 403 is a two-dimensional systolic array. The arithmetic circuit 403 may alternatively be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementation forms, the arithmetic circuit 403 is a general-purpose matrix processor.
[0175] For example, it is assumed that there are an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit fetches the data corresponding to the matrix B from the weight memory 402 and buffers the data to each PE in the arithmetic circuit. The arithmetic circuit fetches the data of the matrix A from the input memory 401, performs a matrix operation on the data of the matrix B and the matrix A, and stores the partial result or the final result of the obtained matrix in the accumulator 408.
[0176] The vector calculation unit 407 can perform further processing on the output of the arithmetic circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. For example, the vector calculation unit 407 may be configured to perform network computing such as Pooling, Batch Normalization, Local Response Normalization, etc. in the non-convolutional / non-FC layers in the neural network.
[0177] In some implementation forms, the vector calculation unit 407 can store the processed output vector in the integrated buffer 406. For example, the vector calculation unit 407 may apply a non-linear function to the output of the arithmetic circuit 403 (for example, a vector of accumulated values) to generate activation values. In some implementation forms, the vector calculation unit 407 generates both normalized values, combined values, or both normalized values and combined values. In some implementation forms, the processed output vector can be used as an activation input for the arithmetic circuit 403, for example, to be used in subsequent layers in a neural network.
[0178] The integrated memory 406 is configured to store input data and output data.
[0179] The direct memory access controller (DMAC) 405 transfers input data in the external memory to the input memory 401 and / or the integrated memory 406, stores the weight data in the external memory in the weight memory 402, and stores the data in the integrated memory 506 in the external memory.
[0180] The bus interface unit (BIU) 410 is configured to implement the interaction between the host CPU, the DMAC, and the instruction fetch buffer 409 via the bus.
[0181] The instruction fetch buffer 409 connected to the controller 404 is configured to store the instructions used by the controller 404.
[0182] The controller 404 is configured to call the instructions buffered in the instruction fetch buffer 409 to control the working process of the operation accelerator.
[0183] Generally, the integrated memory 406, the input memory 401, the weight memory 402, and the instruction fetch buffer 409 are on-chip memories. The external memory is the memory outside the NPU. The external memory can be a double data rate synchronous dynamic random access memory (DDR SDRAM), a high bandwidth memory (HBM), or another readable and writable memory.
[0184] Note that the operations of each layer in the convolutional network shown in FIG. 2 or FIG. 3 may be performed by the arithmetic circuit 403 or the vector calculation unit 407.
[0185] Hereinafter, with reference to the accompanying drawings, the separation network training method and the rendering method in the embodiments of the present application will be described in detail.
[0186] First, referring to FIG. 5, the separation network training method in the embodiments of the present application will be described in detail. The method shown in FIG. 5 may be executed by a training device for a separation network. The training device for a separation network may be a cloud service device, or a terminal device, such as a computer, a server, etc., having sufficient operating capabilities to implement the separation network training method, or a system including a cloud service device and a terminal device. For example, the training method may be implemented by the training device 120 in FIG. 1 or the neural network processing unit 40 in FIG. 4.
[0187] Optionally, the training method may be processed by the CPU, or may be processed by both the CPU and the GPU, or the GPU may not be used, but another processor suitable for neural network calculation is used. This is not limited in the present application.
[0188] In this embodiment of the present application, it can be understood that when there are a plurality of sound objects corresponding to the original audio track in the multimedia file, separation can be performed on the original audio track by using a separation network to obtain at least one single object audio track. Of course, when the original audio track in the multimedia file corresponds to only one sound object, the original audio track is a single object audio track, and separation does not need to be performed by using a separation network.
[0189] The training method may include step 501 and step 502. Hereinafter, step 501 and step 502 will be described in detail.
[0190] Step 501: Obtain training data.
[0191] In this embodiment of the present application, the training data is obtained by combining at least an initial third single object audio track and an initial fourth single object audio track. Alternatively, it can be understood that the training data includes a multi-audio track obtained by combining single object audio tracks corresponding to at least two sound objects. The initial third single object audio track corresponds to the third sound object, and the initial fourth single object audio track corresponds to the fourth sound object. In addition, the training data may further include an image that matches the original audio track. The training data may alternatively be a multimedia file, and the multimedia file includes the aforementioned multi-audio track. In addition to the audio track, the multimedia file may further include a video track, a text track (or a track called a subtitle screen track), etc. This is not particularly limited in this specification.
[0192] The audio track (such as the original audio track, the first single object audio track, etc.) in this embodiment of the present application may include an audio track generated by a sound object (or referred to as a sound emission object), such as a human voice track, an instrument track (e.g., a drum track, a piano track, a trumpet track, etc.), the sound of an airplane, etc. The specific sound object corresponding to the audio track is not limited herein.
[0193] In this embodiment of the present application, the training data may be obtained by directly recording the sound emitted by the sound object, or may be obtained by the user inputting audio information and video information, or may be received from a capture device. In actual applications, the training data may be obtained by other methods. The method of obtaining the training data is not particularly limited herein.
[0194] Step 502: To obtain the trained separation network, use the training data as the input of the separation network and use the value of the loss function less than the first threshold as the target to train the separation network.
[0195] The separation network in this embodiment of the present application may be referred to as a separation neural network, or may be referred to as a separation model, or may be referred to as a separation neural network model. This is not particularly limited herein.
[0196] The loss function indicates the difference between the third single object audio track obtained through separation and the initial third single object audio track.
[0197] In this case, the separation network is trained to reduce the value of the loss function, that is, to continuously reduce the difference between the third single-object audio track output by the separation network and the initial third single-object audio track. The training process can be understood as a separation task. The loss function can be understood as a loss function corresponding to the separation task. The output of the separation network (at least one single-object audio track) is a single-object audio track corresponding to at least one sound object in the input (audio track). The third sound object and the first sound object have the same type, and the second sound object and the fourth sound object have the same type. For example, both the first sound object and the third sound object are human voices, but the first sound object may be user A, and the second sound object may be user B. In other words, the third single-object audio track and the first single-object audio track are audio tracks corresponding to voices emitted by different people. In this embodiment of the present application, the third sound object and the first sound object may be two sound objects of the same type, or one sound object of the same type. This is not particularly limited herein.
[0198] Optionally, the training data input to the separation network includes the original audio track corresponding to at least two sound objects. The separation network may output a single-object audio track corresponding to one of the at least two sound objects, or can output single-object audio tracks corresponding to each of the at least two sound objects.
[0199] For example, a multimedia file includes an audio track corresponding to human voice, an audio track corresponding to a piano, and an audio track corresponding to the sound of a vehicle. After separation is performed on the multimedia file by using a separation network, one single-object audio track (for example, a single-object audio track corresponding to human voice), two single-object audio tracks (for example, a single-object audio track corresponding to human voice and a single-object audio track corresponding to the sound of a vehicle), or three single-object audio tracks can be obtained.
[0200] In a possible implementation form, the separation network is shown in FIG. 6, and the separation network includes one-dimensional convolution and a residual structure. By adding the residual structure, the gradient transfer efficiency can be improved. Of course, the separation network may further include activation, pooling, etc. The specific structure of the separation network is not limited in this specification. In the case of the separation network shown in FIG. 6, a signal source (that is, a signal corresponding to the audio track in the multimedia file) is used as an input, and conversion is performed through multiple convolutions and deconvolutions to output an object signal (a single audio track corresponding to a sound object). Also, a recurrent neural network module may be added to improve the temporal correlation, or different output layers may be connected to improve the relationship between high-dimensional features and low-dimensional features.
[0201] In another possible implementation, the separation network is shown in FIG. 7. Before the signal source is input into the separation network, the signal source may first be preprocessed. For example, in order to obtain a spectrogram, STFT mapping processing is performed on the signal source. For the amplitude spectrum in the spectrogram, conversion is performed by two-dimensional convolution and inverse convolution to obtain a mask spectrum (the spectrum obtained by screening). The mask spectrum and the amplitude spectrum are combined to obtain a target amplitude spectrum. Then, the target amplitude spectrum is multiplied by the phase spectrum to obtain a target spectrogram. In order to obtain an object signal (a single audio track corresponding to a sound object), inverse short-time Fourier transform (iSTFT) mapping is performed on the target spectrogram. Different output layers may be connected to improve the relationship between high-dimensional features and low-dimensional features, a residual structure may be added to improve the gradient transfer efficiency, and a recurrent neural network module may be added to improve the time-series correlation.
[0202] The input of FIG. 6 may be understood as a one-dimensional time-domain signal, and the input of FIG. 7 is a two-dimensional spectrogram signal.
[0203] The two separation models described above are merely examples. In actual applications, there are other possible structures. The input of the separation model may be a time-domain signal, the output of the separation model may be a time-domain signal, the input of the separation model may be a time-frequency domain signal, and the output of the separation model may be a time-frequency domain signal. The structure, input, or output of the separation model is not particularly limited in this specification.
[0204] Optionally, before being input into the separation network, the multi-audio tracks in the multimedia file may first be further identified by using an identification network to identify the number of audio tracks included in the multi-audio track and the object type (e.g., human voice, drum sound, etc.), so that the training period of the separation network can be reduced. Of course, the separation network may alternatively include a multi-audio track identification sub-network. This is not particularly limited herein. The input of the identification network may be a time-domain signal, and the output of the identification network may be a class probability. It is equivalent that a time-domain signal is input into the identification network, the obtained object is a class probability, and the class whose probability exceeds the threshold is selected as the classification class. The object in this specification can be understood as a sound object.
[0205] For example, the input in the identification network is a multimedia file obtained by combining the audio corresponding to vehicle A and the audio corresponding to vehicle B. The multimedia file is input into the identification network, and the identification network can output the vehicle class. Of course, if the training data is sufficiently comprehensive, the identification network can also identify a specific vehicle type. This corresponds to a more fine-grained identification. The identification network is set based on actual requirements. This is not particularly limited herein.
[0206] It should be noted that in the training process, another training method may be used instead of the aforementioned training method. This is not limited in this specification.
[0207] FIG. 8 shows another system architecture according to the present application. The system architecture includes an input module, a function module, a database module, and an output module. Each module will be described in detail below.
[0208] 1. Input Module The input module includes a database option sub-module, a sensor information acquisition sub-module, a user interface input sub-module, and a file input sub-module. The above four sub-modules can be understood as four input methods.
[0209] The database option sub-module is configured to perform spatial rendering in a rendering method stored in the database and selected by the user.
[0210] The sensor information acquisition sub-module is configured to specify the spatial position of a specific sound object by using a sensor (which may be a sensor within the rendering device or another sensor device, and is not particularly limited herein). In this way, the user can select the position of a specific sound object.
[0211] The user interface input sub-module is configured to determine the spatial position of a specific sound object in response to the user's operation on the user interface. Optionally, the user may control the spatial position of a specific sound object by means such as tapping or dragging.
[0212] The file input sub-module is configured to track a specific sound object based on image information or character information (such as lyrics, subtitles, etc.), and determine the spatial position of the specific sound object based on the tracked position of the specific sound object.
[0213] 2. Function Module The function module includes a signal transmission sub-module, an object identification sub-module, a calibration sub-module, an object tracking sub-module, an orientation calculation sub-module, an object separation sub-module, and a rendering sub-module.
[0214] The signal transmission sub-module is configured to receive and transmit information. Specifically, the signal transmission sub-module can receive the input information of the input module and output the feedback information to another module. For example, the feedback information includes information such as the position change information of a specific sound object, a single object audio track obtained by separation, etc. Of course, the signal transmission sub-module may be further configured to feedback the identified object information to the user via a user interface (UI), etc. This is not particularly limited in this specification.
[0215] The object identification sub-module is configured to identify all object information of the multi-audio track information transmitted by the input module and received by the signal transmission sub-module. The object referred to in this specification is a sound object (also referred to as a sound emission object), for example, a human voice, a drum sound, an airplane sound, etc. Optionally, the object identification sub-module may be an identification sub-network within the identification network or separation network described in the embodiment shown in FIG. 5.
[0216] The calibration sub-module is configured to calibrate the initial state of the playback device. For example, when the playback device is a headset, the calibration sub-module is configured to calibrate the headset, or when the playback device is a loudspeaker device, the calibration sub-module is configured to calibrate the loudspeaker device. In the case of headset calibration, the initial state of the sensor (the relationship between the sensor device and the playback device is described in FIG. 9) can be regarded as the front-facing direction by default, and the calibration is then performed based on the front-facing direction. Alternatively, the actual position of the sensor placed by the user may be obtained to ensure that the front-facing direction of the audio-visual content is the front-facing direction of the headset. For the calibration of the loudspeaker device, the coordinate positions of each loudspeaker device are first obtained (this may be obtained through interaction by using the sensors of the user terminal, and the corresponding description will be provided later in FIG. 9). After calibrating the playback device, the calibration sub-module transmits information about the calibrated playback device to the database module via the signal transmission sub-module.
[0217] The object tracking sub-module is configured to track the motion track of a specific sound object. The specific sound object can be a text or an image in a multi-modal file (e.g., audio information and corresponding video information, audio information and corresponding character information, etc.) that is displayed. Optionally, the object tracking sub-module can be further configured to render the motion track on the audio side. In addition, the object tracking sub-module can further include a target identification network and a depth estimation network. The target identification network is configured to identify a specific sound object that needs to be tracked, and the depth estimation network is configured to obtain the relative coordinates of a specific sound object in the image (a detailed description will be provided in subsequent embodiments). As a result, the object tracking sub-module renders the orientation and motion track of the audio corresponding to the specific sound object based on the relative coordinates.
[0218] The orientation calculation sub-module is configured to convert the information obtained by the input module (e.g., sensor information, input information of the UI interface, file information, etc.) into orientation information (which can also be referred to as the sound source position). There are corresponding conversion methods for different information, and the specific conversion process will be described in detail in subsequent embodiments.
[0219] The object separation sub-module is configured to separate at least one single object audio track from a multimedia file (or multimedia information) or multi-audio track information, for example, to extract a separate human voice track (i.e., an audio file containing only human voices) from a piece of music. The object separation sub-module may be the separation network in the embodiment shown in FIG. 5. Further, the structure of the object separation sub-module may be the structure shown in FIG. 6 or FIG. 7. This is not particularly limited in this specification.
[0220] The rendering sub-module is configured to obtain the sound source position acquired by the orientation calculation sub-module and perform spatial rendering on the sound source position. Further, based on the playback device selected according to the input information of the UI in the input module, the corresponding rendering method can be determined. For different playback devices, the rendering method is different. The rendering process will be described in detail in subsequent embodiments.
[0221] 3. Database Module The database module includes a database selection sub-module, a rendering rule editing sub-module, and a rendering rule sharing sub-module.
[0222] The database selection sub-module is configured to store the rendering rules. The rendering rules may be default rendering rules that convert dual-channel / multi-channel audio tracks into a three-dimensional (3D) spatial sense provided by the system during system initialization, or rendering rules stored by the user. Optionally, different objects may correspond to the same rendering rule, or different objects may correspond to different rendering rules.
[0223] The rendering rule editing sub-module is configured to re-edit the stored rendering rules. Optionally, the stored rendering rules may be the rendering rules stored in the database selection sub-module, or newly input rendering rules. This is not particularly limited in this specification.
[0224] The rendering rule sharing sub-module is configured to upload rendering rules to the cloud and / or download specific rendering rules from the rendering rule database in the cloud. For example, the rendering rule sharing module can upload the rendering rules customized by the user to the cloud and share the rendering rules with another user. The user can select a rendering rule that matches the multi-audio track information shared and played by another user from the rendering rule database stored in the cloud, and download the rendering rule to the database on the terminal side as a data file for the audio 3D rendering rule.
[0225] 4. Output module The output module is configured to play the rendered single-object audio track or the target audio track (obtained based on the original audio track and the rendered single-object audio track) by using a playback device.
[0226] First, the application scenario to which the rendering method provided in the embodiments of this application is applied will be described.
[0227] Referring to FIG. 9. The application scenario includes a control device 901, a sensor device 902, and a playback device 903.
[0228] The playback device 903 in this embodiment of this application may be a loudspeaker device, or may be a headset (for example, in-ear earphones, headphones, etc.), or may be a large screen (for example, a projection screen), etc. This is not particularly limited in this specification.
[0229] The control device 901, the sensor device 902, and the sensor device 902 and the playback device 903 may be connected by a wired method, a wireless fidelity (WIFI) method, a mobile data network method, or another connection method. This is not particularly limited in this specification.
[0230] The control device 901 in this embodiment of the present application is a terminal device configured to provide services to a user. The terminal device may include a head mount display (HMD) device. The head mount display device may be a combination of a virtual reality (VR) box, an integrated VR headset, a personal computer (PC) VR, an augmented reality (AR) device, a mixed reality (MR) device, and the like. Alternatively, the terminal device may include a cellular phone, a smart phone, a personal digital assistant (PDA), a tablet computer, a laptop computer, a personal computer (PC), an in-vehicle terminal, and the like. This is not particularly limited in this specification.
[0231] The sensor device 902 in this embodiment of the present application is a device configured to sense orientation and / or position, and may be a laser pointer, a cellular phone, a smart watch, a smart band, a device having an inertial measurement unit (IMU), a device having a simultaneous localization and mapping (SLAM) sensor, and the like. This is not particularly limited in this specification.
[0232] The playback device 903 in this embodiment of the present application is a device configured to play audio or video, and may be a loudspeaker device (for example, a speaker, or a terminal device having an audio or video playback function), or an in-ear monitoring device (for example, in-ear headphones, headphones, an AR device, a VR device, etc.). This is not particularly limited in this specification.
[0233] In the application scenario shown in FIG. 9, it can be understood that for each device, one or more such devices may exist. For example, a plurality of loudspeaker devices may exist. For each device, the quantity is not particularly limited in this specification.
[0234] In this embodiment of the present application, the control device, the sensor device, and the playback device may be three devices, or two devices, or one device. This is not particularly limited in this specification.
[0235] In a possible implementation, the control device and the sensor device in the application scenario shown in FIG. 9 are the same device. For example, the control device and the sensor device are the same mobile phone, and the playback device is a headset. In another example, the control device and the sensor device are the same mobile phone, and the playback device is a loudspeaker device (which may also be referred to as a loudspeaker device system, and the loudspeaker device system includes one or more loudspeaker devices).
[0236] In another possible implementation, the control device and the playback device in the application scenario shown in FIG. 9 are the same device. For example, the control device and the playback device are the same computer. In another example, the control device and the playback device are the same large screen.
[0237] In another possible implementation, in the application scenario shown in FIG. 9, the control device, the sensor device, and the playback device are the same device. For example, the control device, the sensor device, and the playback device are the same tablet computer.
[0238] Hereinafter, with reference to the foregoing application scenarios and the accompanying drawings, the rendering method in the embodiments of the present application will be described in detail.
[0239] FIG. 10 shows an embodiment of the rendering method provided in an embodiment of the present application. This method can be implemented by a rendering device or by components of a rendering device (such as a processor, a chip, a chip system, etc.). This embodiment includes steps 1001 to 1004.
[0240] In this embodiment of the present application, the rendering device can have the functions of the control device in FIG. 9, the functions of the sensor device in FIG. 9, and / or the functions of the playback device in FIG. 9. This is not particularly limited in this specification. The following uses an example in which the rendering device is a control device (such as a notebook computer), the sensor device is a device having an IMU (such as a mobile phone), and the playback device is a loudspeaker device (such as a speaker) to explain the rendering method.
[0241] The sensor described in the embodiments of the present application may be a sensor within the rendering device or a sensor within a device other than the rendering device (such as the aforementioned sensor device). This is not particularly limited in this specification.
[0242] Step 1001: Calibrate the playback device. This step is optional.
[0243] Optionally, before the playback device plays the rendered audio track, the playback device may first be calibrated. The calibration aims to improve the realism of the spatial effect of the rendered audio track.
[0244] In this embodiment of the present application, there may be multiple ways to calibrate the playback device. The following describes the process of calibrating the playback device by only using the example where the playback device is a loudspeaker device. FIG. 11 shows the playback device calibration method provided in this embodiment. This method includes steps 1 to 5.
[0245] Optionally, before calibration, the mobile phone held by the user establishes a connection to the loudspeaker device. The connection method is the same as the connection method between the sensor device and the playback device in the embodiment shown in FIG. 9. Details are not repeated herein.
[0246] Step 1: Determine the type of the playback device.
[0247] In this embodiment of the present application, the rendering device may determine the type of the playback device based on the user's operation, or may adaptively detect the type of the playback device, or may determine the type of the playback device based on the default setting, or may determine the type of the playback device in another way. This is not particularly limited herein.
[0248] For example, when the rendering device determines the type of the playback device based on the user's operation, the rendering device can display the interface shown in FIG. 12. The interface includes a playback device type selection icon. In addition, the interface may further include an input file selection icon, a rendering method selection (i.e., reference information option) icon, a calibration icon, a sound hunter icon, an object bar, volume, progress, and a spherical view (or a three-dimensional view). As shown in FIG. 13, the user can tap the "playback device type selection icon" 101. As shown in FIG. 14, the rendering device displays a drop-down list in response to the tap operation. The drop-down list may include "loudspeaker device option" and "headset option". Further, the user can tap the "loudspeaker device option" 102 to determine that the type of the playback device is a loudspeaker device. As shown in FIG. 15, in the interface displayed by the rendering device, "playback device type selection" can be replaced with "loudspeaker device" to prompt the user that the current type of the playback device is a loudspeaker device. It can also be understood that the rendering device displays the interface shown in FIG. 12, the rendering device receives the user's fifth operation (i.e., the tap operations shown in FIGS. 13 and 14), and in response to the fifth operation, the rendering device selects a loudspeaker device as the type of the playback device from the playback device type options.
[0249] In addition, since this method is used to calibrate the playback device, as shown in FIG. 16, the user can further tap the "Calibration Icon" 103. As shown in FIG. 17, the rendering device displays a drop-down list in response to the tap operation. The drop-down list may include "Default Option" and "Manual Calibration Option". Further, the user can tap the "Manual Calibration Option" 104 to determine that the calibration method is automatic calibration. Automatic calibration can be understood as calibrating the playback device by the user using a mobile phone (i.e., the sensor device).
[0250] In FIG. 14, only an example where the drop-down list of the "Playback Device Type Selection Icon" includes the "Loudspeaker Device Option" and the "Headset Option" is used. In actual use, the drop-down list may further include options for specific headset types, such as headphones, in-ear earphones, wired headsets, Bluetooth headsets, etc. This is not particularly limited in this specification.
[0251] In FIG. 17, only an example where the drop-down list of the "Calibration Icon" includes the "Default Option" and the "Manual Calibration Option" is used. In actual use, the drop-down list may further include other types of options. This is not particularly limited in this specification.
[0252] Step 2: Determine the test audio.
[0253] The test audio in this embodiment of the present application may be a test signal (e.g., pink noise) specified by default, or a single-object audio track corresponding to human speech separated from a piece of music (i.e., the multimedia file is a piece of music) by using the separation network in the embodiment shown in FIG. 5, or may be audio corresponding to another single-object audio track in the piece of music, or may be audio including only a single-object audio track. This is not particularly limited herein.
[0254] For example, the user can select the test audio by tapping the "input file selection icon" in the interface displayed by the rendering device.
[0255] Step 3: Obtain the attitude angle of the mobile phone and the distance between the sensor and the loudspeaker device.
[0256] After the test sound source is determined, the loudspeaker device plays the test audio in sequence, and the user holds the sensor device (e.g., a mobile phone) and points to the loudspeaker device playing the test audio. After the mobile phone is stably placed, the current orientation of the mobile phone and the signal energy of the received test audio are recorded, and the distance between the mobile phone and the loudspeaker device is calculated according to the following formula 1. The same applies when there are multiple loudspeaker devices. Details are not repeated herein. The fact that the mobile phone is stably placed can be understood as that the dispersion of the orientation of the mobile phone is less than a threshold value (e.g., 5 degrees) within a certain period (e.g., 200 milliseconds).
[0257] Optionally, if the playback device is two loudspeaker devices, the first loudspeaker device plays the test audio first, and the user holds the mobile phone so as to point to the first loudspeaker device. After the first loudspeaker device is calibrated, the user holds the mobile phone to point to the second loudspeaker device and performs calibration.
[0258] In this embodiment of the present application, the orientation of the mobile phone can be the posture angle of the mobile phone. The posture angle can include the azimuth angle and the tilt angle (or the angle called the tilt angle), or the posture angle can include the azimuth angle, the tilt angle, and the pitch angle. The azimuth angle represents the angle around the z-axis, the tilt angle represents the angle around the y-axis, and the pitch angle represents the angle around the x-axis. The relationship between the orientation of the mobile phone and the x-axis, y-axis, and z-axis can be shown in FIG. 18.
[0259] For example, the above-mentioned example is still used. The playback device is two loudspeaker devices. The first loudspeaker device plays the test audio first. The user holds the mobile phone so as to point to the first loudspeaker device and records the current orientation of the mobile phone and the signal energy of the received test audio. Next, the second loudspeaker device plays the test audio. The user holds the mobile phone so as to point to the second loudspeaker and records the current orientation of the mobile phone and the signal energy of the received test audio.
[0260] Furthermore, in the process of calibrating the loudspeaker device, the rendering device can display the interface shown in FIG. 19. The right side of the interface shows a spherical view. The calibrated loudspeaker device and the loudspeaker device being calibrated can be displayed in the spherical view. Additionally, an uncalibrated loudspeaker device (not shown) may be further displayed. This is not particularly limited in this specification. The center of the circle of the spherical view is the position of the user (which may be understood as the position where the user holds the mobile phone, and since the user holds the mobile phone, the position of the mobile phone is close to the position of the user). The radius may be the distance between the position of the user (or the position of the mobile phone) and the loudspeaker device, or may be set to a default value (for example, 1 meter), etc. This is not particularly limited in this specification.
[0261] For ease of understanding, FIG. 20 is an effect diagram of an example where the user holds the mobile phone facing the loudspeaker device.
[0262] In this embodiment of the present application, there are N loudspeaker devices, where N is a positive integer. The i-th loudspeaker device is one of the N loudspeaker devices, and i is a positive integer, and i ≤ N. In all the formulas in this embodiment of the present application, the i-th loudspeaker device is used as an example for calculation, and the calculation for another loudspeaker device is the same as the calculation for the i-th loudspeaker device.
[0263] Equation 1 used to calibrate the i-th loudspeaker device can be as follows: Equation 1:
Number
[0264] x(t) represents the energy of the test signal received by the mobile phone at instant t, X(t) represents the energy of the test signal reproduced by the loudspeaker device at instant t, t is a positive number, r i represents the distance between the mobile phone and the i-th loudspeaker device (since the user holds the mobile phone, the distance can be understood as the distance between the user and the i-th loudspeaker device), r s represents the normalized distance, and the normalized distance may be understood as a coefficient and is used to convert the ratio of X(t) to x(t) into a distance. The coefficient may be specified based on the actual situation of the loudspeaker device, r s The specific value of is not limited in this specification.
[0265] In addition, when there are multiple loudspeaker devices, the test signal is reproduced sequentially, the mobile phone faces the loudspeaker device, and the distance is obtained by using Equation 1.
[0266] It can be understood that Equation 1 is an example. In actual applications, Equation 1 may alternatively be in another form, such as removal, etc. This is not particularly limited in this specification.
[0267] Step 4: Based on the attitude angle and the distance, determine the position information of the loudspeaker device.
[0268] In Step 3, the mobile phone records the attitude angle of the mobile phone with respect to each loudspeaker device and calculates the distance between the mobile phone and each loudspeaker device using Equation 1. Of course, the mobile phone may alternatively transmit the measured distance to the rendering device, and the rendering device calculates the distance between the mobile phone and each loudspeaker device by using Equation 1. This is not particularly limited in this specification.
[0269] After obtaining the posture angle of the mobile phone and the distance between the mobile phone and the loudspeaker device, the rendering device can convert the posture angle of the mobile phone and the distance between the mobile phone and the loudspeaker device into the position information of the loudspeaker device in the spherical coordinate system by using Equation 2. The position information includes the azimuth angle, the tilt angle, and the distance (specifically, the distance between the sensor device and the playback device). When there are multiple loudspeaker devices in the loudspeaker device system, it is the same to determine the position information of another loudspeaker device. Details are not repeated herein.
[0270] Equation 2 can be as follows: Equation 2:
Number
[0271] λ(t) represents the azimuth angle of the i-th loudspeaker device in the spherical coordinate system at instant t, Φ(t) represents the tilt angle of the i-th loudspeaker device in the spherical coordinate system at instant t, and d(t) represents the distance between the mobile phone and the i-th loudspeaker device. Ω(t)[0] represents the azimuth angle of the mobile phone at instant t (i.e., the rotation angle of the mobile phone around the z-axis), Ω(t)[1] represents the pitch angle of the mobile phone at instant t (i.e., the rotation angle of the mobile phone around the x-axis), and r i represents the distance calculated by using Equation 1, sign represents a positive or negative value. When Ω(t)[1] is positive, sign is positive, or when Ω(t)[1] is negative, sign is negative. %360 is used to adjust the angle range to 0 degrees to 360 degrees. For example, when the angle of Ω(t)[0] is -80 degrees, Ω(t)[0]%360 represents -80 + 360 = 280 degrees.
[0272] It can be understood that Equation 2 is merely an example. In actual applications, Equation 2 may be in another form. This is not particularly limited herein.
[0273] For example, after the user calibrates the loudspeaker device, the rendering device can display the interface shown in FIG. 21. The interface can display a "calibrated icon", and the position of the calibrated playback device can be displayed in a spherical view on the right side.
[0274] Since the playback device is calibrated, the problem of calibrating irregular loudspeaker devices can be solved. In this way, the user can obtain the spatial position of each loudspeaker device in subsequent operations, and as a result, the position required for the single object audio track is accurately rendered, and the realism of the spatial effect of rendering the audio track is improved.
[0275] Step 1002: Obtain a first single object audio track based on the multimedia file.
[0276] In this embodiment of the present application, the rendering device may obtain the multimedia file by directly recording the audio created by the first sound object, or may obtain the multimedia file transmitted by another device. For example, receive the multimedia file transmitted by a capture device (such as a camera, a recorder, a mobile phone, etc.). In actual applications, the multimedia file may be obtained by another method. The specific method of obtaining the multimedia file is not limited in this specification.
[0277] The multimedia file in this embodiment of the present application may specifically be audio information, for example, stereo audio information, multi-channel audio information, etc. Alternatively, the multimedia file may specifically be multi-modal information. For example, the multi-modal information may be video information, image information corresponding to audio information, character information, etc. In addition to the audio track, it can also be understood that the multimedia file may further include a video track, a text track (or a track called a subtitle screen track), etc. This is not particularly limited herein.
[0278] In addition, the multimedia file may include a first single-object audio track or may include the original audio track. The original audio track is obtained by combining at least two single-object audio tracks. This is not particularly limited herein. The original audio track may be a single audio track or a multi-audio track. This is not particularly limited herein. The original audio track may include an audio track generated by a sound object (or a sound emission object), such as a human voice track, an instrument track (for example, a drum track, a piano track, a trumpet track, etc.), the sound of an airplane, etc. The specific type of the sound object corresponding to the original audio track is not limited in this specification.
[0279] According to multiple cases of the original audio track in the multimedia file, the processing method of this step may be different and will be described separately below.
[0280] In the first method, the audio track in the multimedia file is a single-object audio track.
[0281] In this case, the rendering device can directly obtain the first single-object audio track from the multimedia file.
[0282] In the second method, the audio track in the multimedia file is a multi-object audio track.
[0283] In this case, the original audio track in the multimedia file can be understood as corresponding to a plurality of sound objects. Optionally, in addition to the first sound object, the original audio track further corresponds to a second sound object. In other words, the original audio track is obtained by combining at least a first single-object audio track and a second single-object audio track. The first single-object audio track corresponds to the first sound object, and the second single-object audio track corresponds to the second sound object.
[0284] In this case, the rendering device can separate the first single-object audio track from the original audio track, or can separate the first single-object audio track and the second single-object audio track from the original audio track. This is not particularly limited herein.
[0285] Optionally, the rendering device can separate a first single-object audio track from the original audio track using the separation network in the embodiment shown in FIG. 5. Additionally, the rendering device can alternatively separate a first single-object audio track and a second single-object audio track from the original audio track by using the separation network. This is not particularly limited herein. Different outputs depend on different ways of training the separation network. For details, refer to the description in the embodiment shown in FIG. 5. Details are not repeated herein.
[0286] Optionally, after determining a multimedia file, the rendering device can identify the sound objects of the original audio track in the multimedia file by using an identification network or a separation network. For example, the sound objects included in the original audio track include a first sound object and a second sound object. The rendering device may randomly select one of the sound objects as the first sound object, or determine the first sound object based on the user's selection. Further, after determining the first sound object, the rendering device may obtain a first single-object audio track by using the separation network. Of course, after determining the multimedia file, the rendering device can first obtain the sound objects by using the identification network, and then obtain the single-object audio tracks of the sound objects by using the separation network. Alternatively, the sound objects included in the multimedia file and the single-object audio tracks corresponding to the sound objects can be directly obtained by using the identification network and / or the separation network. This is not particularly limited herein.
[0287] For example, the foregoing example is still used. After the playback device is calibrated, the rendering device can display the interface shown in FIG. 21 or the interface shown in FIG. 22. The user can select a multimedia file by tapping the "Input File Selection Icon" 105. For example, the multimedia file in this specification is "Dream it possible.wav". The rendering device receives the user's fourth operation, and in response to the fourth operation, the rendering device can also be understood to select "Dream it possible.wav" (i.e., the target file) as the multimedia file from at least one multimedia file stored in the storage area. The storage area may be a storage area within the rendering device or a storage area within an external device (e.g., a USB flash drive). This is not particularly limited in this specification. After the user selects a multimedia file, the rendering device can display the interface shown in FIG. 23. In the interface, "Input File Selection" may be replaced with "Dream it possible.wav" to prompt the user that the current multimedia file is "Dream it possible.wav". In addition, by using the identification network and / or separation network in the embodiment shown in FIG. 4, the rendering device can identify the sound objects in "Dream it possible.wav" and separate the single object audio tracks corresponding to each sound object. For example, the rendering device identifies that the sound objects included in "Dream it possible.wav" are a person, a piano, a violin, and a guitar.As shown in FIG. 23, the interface displayed by the rendering device may further include an object bar, and icons such as a "human voice icon", a "piano icon", a "violin icon", and a "guitar icon" may be displayed within the object bar for the user to select a sound object to be rendered. Optionally, a "combination icon" may be further displayed within the object bar, and the user may stop selecting a sound object by tapping the "combination icon".
[0288] Furthermore, as shown in FIG. 24, the user can determine that the audio track to be rendered is a single object audio track corresponding to human voice by tapping the "human voice icon" 106. The rendering device identifies "Dream it possible.wav", obtains the interface displayed by the rendering device and shown in FIG. 24, the rendering device receives the user's tap operation, and in response to the tap operation, the rendering device selects the first icon (i.e., the "human voice icon" 106) within the interface. As a result, it can also be understood that the rendering device determines that the first single object audio track is human voice.
[0289] It can be understood that the example where the type of the playback device shown in FIGS. 22 to 24 is a loudspeaker device is simply used. Of course, the user may select a headset as the type of the playback device. The example where a loudspeaker device is selected by the user as the type of the playback device during calibration is simply used for the following description.
[0290] In addition, the rendering device can further duplicate one or more single object audio tracks in the original audio track. For example, as shown in FIG. 25, the user may further duplicate the "human voice icon" in the object bar to obtain a "human voice 2 icon", and the single object audio track corresponding to the human voice 2 is the same as the single object audio track corresponding to the human voice. The duplication method may be that the user double-taps the "human voice icon", or double-taps the human voice in the spherical view. This is not particularly limited herein. After the user obtains the "human voice 2 icon" by the duplication method, it may be considered by default that the user loses the control permission for the human voice and can start the control of the human voice 2. Optionally, after the user obtains the human voice 2 by the duplication method, the first sound source position of the human voice may be further displayed in the spherical view. Of course, the user may duplicate the sound object or delete the sound object.
[0291] Step 1003: Determine the first sound source position of the first sound object based on the reference information.
[0292] In this embodiment of the present application, when the original audio track of the multimedia file includes a plurality of sound objects, the sound source position of one sound object can be determined based on the reference information, or a plurality of sound source positions corresponding to the plurality of sound objects can be determined based on the reference information. This is not particularly limited herein.
[0293] For example, the foregoing example is still used. The rendering device determines that the first sound object is a human voice, and the rendering device can display the interface shown in FIG. 26. The user can tap the "rendering method selection icon" 107 to select reference information, and the reference information is used to determine the first sound source position of the first sound object. As shown in FIG. 27, the rendering device can display a drop-down list in response to the user's first operation (i.e., the foregoing tap operation). The drop-down list can include "automatic rendering option" and "interactive rendering option". The "interactive rendering option" corresponds to the reference position information, and the "automatic rendering option" corresponds to the media information.
[0294] It can also be understood that the rendering method includes an automatic rendering method and an interactive rendering method. The automatic rendering method means that the rendering device automatically obtains a first single-object audio track rendered based on the media information in the multimedia file. The interactive rendering method means that the first single-object audio track rendered is obtained through the interaction between the user and the rendering device. In other words, when the automatic rendering method is determined, the rendering device can obtain the first single-object audio track rendered in a preset manner, or when the interactive rendering method is determined, the rendering device obtains reference position information in response to the user's second operation, determines the first sound source position of the first sound object based on the reference position information, renders the first single-object audio track based on the first sound source position, and obtains the rendered first single-object audio track. The preset method includes steps of obtaining the media information of the multimedia file, determining the first sound source position of the first sound object based on the media information, and rendering the first single-object audio track based on the first sound source position to obtain the rendered first single-object audio track.
[0295] In addition, the sound source positions (the first sound source position and the second sound source position) in this embodiment of the present application may be fixed positions at a certain moment, or may be multiple positions (for example, a motion track) during a certain period. This is not particularly limited in this specification.
[0296] In this embodiment of the present application, the reference information has multiple cases. The cases will be described separately below.
[0297] In the first case, the reference information includes reference position information.
[0298] The reference position information in this embodiment of the present application indicates the sound source position of the first sound object. The reference position information may be the first position information of the sensor device, or may be the second position information selected by the user, etc. This is not particularly limited in this specification.
[0299] In this embodiment of the present application, the reference position information has multiple cases. Hereinafter, the cases will be described separately.
[0300] 1. The reference position information is the first position information of a sensor device (hereinafter referred to as a sensor).
[0301] For example, the aforementioned example is still used. As shown in FIG. 27, further, the user can tap the "Interactive Rendering Option" 108 to determine that the rendering method is interactive rendering. The rendering device can display a drop-down list in response to the tap operation. The drop-down list may include "Orientation Control Option", "Position Control Option", and "Interface Control Option".
[0302] In this embodiment of the present application, the first position information has multiple cases. Hereinafter, the cases will be described separately.
[0303] 1.1. The first position information includes the first attitude angle of the sensor.
[0304] Similar to how a loudspeaker device was previously calibrated by using the orientation of a sensor, a user can adjust the orientation of a handheld sensor device (e.g., a mobile phone) through a second operation (e.g., up, down, left, and right translations) to determine the first sound source position of a first single-object audio track. The rendering device may receive the first pose angle of the mobile phone and obtain the first sound source position of the first single-object audio track according to Equation 3 below. It can also be understood that the first sound source position includes the azimuth angle, the tilt angle, and the distance between the loudspeaker device and the mobile phone.
[0305] Furthermore, the user can further determine the second sound source position of a second single-object audio track by adjusting the orientation of the handheld mobile phone again. The rendering device may receive the first pose angle (including the azimuth angle and the tilt angle) of the mobile phone and obtain the second sound source position of the second single-object audio track according to Equation 3 below. It can also be understood that the second sound source position includes the azimuth angle, the tilt angle, and the distance between the loudspeaker device and the mobile phone.
[0306] Optionally, if no connection is established between the mobile phone and the rendering device, the rendering device may send reminder information to the user, and the reminder information is used to remind the user to connect the mobile phone to the rendering device. Of course, the mobile phone and the rendering device may alternatively be the same mobile phone. In this case, the reminder information need not be sent.
[0307] Equation 3 may be as follows: Equation 3:
Number
[0308] λ(t) represents the azimuth angle of the i-th loudspeaker device in the spherical coordinate system at instant t, Φ(t) represents the elevation angle of the i-th loudspeaker device in the spherical coordinate system at instant t, and d(t) represents the distance between the mobile phone and the i-th loudspeaker device at instant t. Ω(t)[0] represents the azimuth angle of the mobile phone at instant t (i.e., the rotation angle of the mobile phone around the z-axis), and Ω(t)[1] represents the elevation angle of the mobile phone at instant t (i.e., the rotation angle of the mobile phone around the y-axis). d(t) represents the distance between the mobile phone and the i-th loudspeaker device at instant t. The distance may be the distance calculated by using Equation 1 during calibration, or may be set to a default value (e.g., 1 meter), and d(t) may be adjusted based on requirements. sign represents a positive or negative value. When Ω(t)[1] is positive, sign is positive, or when Ω(t)[1] is negative, sign is negative. %360 is used to adjust the angle range to 0 degrees to 360 degrees. For example, if the angle of Ω(t)[0] is -80 degrees, Ω(t)[0]%360 represents -80 + 360 = 280 degrees.
[0309] It can be understood that Equation 3 is merely an example. In actual applications, Equation 3 may be in another form. This is not particularly limited in this specification.
[0310] For example, the aforementioned example is still used. As shown in FIG. 28, the rendering device can display a drop-down list. The drop-down list may include "orientation control options", "position control options", and "interface control options". The user can tap on the "orientation control option" 109 to determine that the rendering method is orientation control in interactive rendering. In addition, after the user selects orientation control, the rendering device can display the interface shown in FIG. 29. In the interface, "rendering method selection" may be replaced with "orientation control rendering" to prompt the user that the current rendering method is orientation control. In this case, the user can adjust the orientation of the mobile phone. When the user adjusts the orientation of the mobile phone, as shown in FIG. 30, a dashed line may be displayed in a spherical view on the display interface of the rendering device, and the dashed line represents the current orientation of the mobile phone. In this way, the user can intuitively view the orientation of the mobile phone in a spherical view to help the user determine the first sound source position. After the orientation of the mobile phone is stabilized (refer to the aforementioned description regarding the mobile phone being stably placed, details of which are not described again herein), the current first posture angle of the mobile phone is determined. Also, the first sound source position is obtained using Equation 3. Also, if the user's position does not change compared to the time of calibration, the distance between the mobile phone and the loudspeaker device obtained at the time of calibration may be used as d(t) in Equation 3. In this way, the user determines the first sound source position of the first sound object based on the first posture angle. Alternatively, the first sound source position of the first sound object is understood as the first sound source position of the first single object audio track. Further, the rendering device can display the interface shown in FIG. 31. The spherical view within the interface includes the first position information of the sensor corresponding to the first sound source position (i.e., the first position information of the mobile phone).
[0311] Also, the above example is an example of determining the first sound source position. Further, the user can determine the second sound source position of the second single object audio track. For example, as shown in FIG. 32, the user may determine that the second sound object is a violin by tapping the "violin icon" 110. The rendering device monitors the posture angle of the mobile phone and uses Equation 3 to determine the second sound source position. As shown in FIG. 32, the spherical view in the display interface of the rendering device can display the currently determined first sound source position of the first sound object (person) and the currently determined second sound source position of the second sound object (violin).
[0312] In this way, the user can perform real-time or subsequent dynamic rendering on the selected sound object based on the orientation provided by the sensor (i.e., the first posture angle). In this case, the sensor is similar to a laser pointer, and the position pointed to by the laser is the sound source position. In this way, the control can assign specific spatial orientations and specific movements to the sound object so that the generation of interaction between the user and the audio is implemented to provide the user with a new experience.
[0313] 1.2. The first position information includes the second posture angle and acceleration of the sensor.
[0314] The user can control the position of a sensor device (e.g., a mobile phone) through a second operation to determine a first sound source position. It can also be understood that the rendering device can receive the second posture angles (including azimuth angle, tilt angle, and pitch angle) and acceleration of the mobile phone, and obtain the first sound source position according to the following equations 4 and 5. The first sound source position includes the azimuth angle, tilt angle, and the distance between the loudspeaker device and the mobile phone. Specifically, the second posture angles and acceleration of the mobile phone are first converted into the coordinates of the mobile phone in a spatial orthogonal coordinate system by using equation 4, and then the coordinates of the mobile phone in the spatial orthogonal coordinate system are converted into the coordinates of the mobile phone in a spherical coordinate system, that is, the first sound source position, by using equation 5.
[0315] Equations 4 and 5 can be as follows: Equation 4:
Number
Number
Number
[0316] x(t), y(t), and z(t) represent the position information of the mobile phone in a spatial orthogonal coordinate system at instant t, g represents the acceleration due to gravity, a(t) represents the acceleration of the mobile phone at instant t, Ω(t)[0] represents the azimuth angle of the mobile phone at instant t (i.e., the rotation angle of the mobile phone around the z-axis), Ω(t)[1] represents the pitch angle of the mobile phone at instant t (i.e., the rotation angle of the mobile phone around the x-axis), Ω(t)[2] represents the tilt angle of the mobile phone at instant t (i.e., the rotation angle of the mobile phone around the y-axis), λ(t) represents the azimuth angle of the i-th loudspeaker device at instant t, Φ(t) represents the tilt angle of the i-th loudspeaker device at instant t, and d(t) represents the distance between the i-th loudspeaker device and the mobile phone at instant t.
[0317] It can be understood that Equations 4 and 5 are merely examples. In actual applications, Equations 4 and 5 may alternatively be in other forms respectively. This is not particularly limited in this specification.
[0318] For example, the foregoing example is still used. When the rendering device displays the interface shown in FIG. 27, after the user determines that the rendering method is interactive rendering, the rendering device can display a drop-down list. The drop-down list may include "orientation control options", "position control options", and "interface control options". The user can tap on the "orientation control options". As shown in FIG. 33, the user can tap on the "position control option" 111 to determine that the rendering method is position control in interactive rendering. In addition, after the user selects position control, in the display interface of the rendering device, "rendering method selection" may be replaced with "position control rendering" to prompt the user that the current rendering method is position control. In this case, the user can adjust the position of the mobile phone. After the position of the mobile phone stabilizes (refer to the foregoing description regarding the mobile phone being stably placed, the details of which will not be described again herein), the current second attitude angle and the current acceleration of the mobile phone are determined. Also, the first sound source position is obtained using Equations 4 and 5. In this way, the user determines the first sound source position of the first sound object based on the second attitude angle and the acceleration. Alternatively, the first sound source position of the first sound object is understood as the first sound source position of the first single object audio track. Further, during the process in which the user adjusts the mobile phone or after the position of the mobile phone stabilizes, the rendering device can display the interface shown in FIG. 34. The spherical view within the interface includes the first position information of the sensor corresponding to the first sound source position (i.e., the first position information of the mobile phone). In this way, the user can intuitively view the position of the mobile phone in the spherical view to assist the user in determining the first sound source position.When the interface of the rendering device displays the first position information in a spherical view in the process of the user adjusting the position of the mobile phone, the first sound source position can change in real time based on the position change of the mobile phone.
[0319] Also, the above example is an example of determining the first sound source position. Further, the user can determine the second sound source position of the second single object audio track. The method of determining the second sound source position is the same as the method of determining the first sound source position. Details are not repeated in this specification.
[0320] It can be understood that the above two ways of the first position information are merely examples. In actual applications, the first position information may alternatively have other cases. This is not particularly limited in this specification.
[0321] In this way, the single object audio track corresponding to the sound object in the audio is separated, the sound object is controlled by using the actual position information of the sensor as the sound source position, and real-time or subsequent dynamic rendering is performed. In this way, the motion track of the sound object can be easily and completely controlled, and as a result, the flexibility of editing is greatly improved.
[0322] 2. The reference position information is the second position information selected by the user.
[0323] The rendering device can provide a spherical view for the user to select the second position information. The center of the sphere of the spherical view is the user's position, and the radius of the spherical view is the distance between the user's position and the loudspeaker device. The rendering device acquires the second position information selected by the user within the spherical view and converts the second position information into the first sound source position. It can also be understood that the rendering device acquires the second position information of the point selected by the user within the spherical view and converts the second position information of the point into the first sound source position. The second position information includes the two-dimensional coordinates and depth (i.e., the distance between the tangent plane and the center of the sphere) of the point selected by the user on the tangent plane within the spherical view.
[0324] For example, the aforementioned example is still used. When the rendering device displays the interface shown in FIG. 27, after the user determines that the rendering method is interactive rendering, the rendering device can display a drop-down list. The drop-down list may include "orientation control options", "position control options", and "interface control options". The user can tap on the "orientation control options". As shown in FIG. 35, the user can tap on the "interface control option" 112 to determine that the rendering method is interface control in interactive rendering. In addition, after the user selects interface control, the rendering device can display the interface shown in FIG. 36. In the interface, "rendering method selection" may be replaced with "interface control rendering" to prompt the user that the current rendering method is interface control.
[0325] In this embodiment of the present application, the second position information has multiple cases. The cases will be described separately below.
[0326] 2.1. The second position information is acquired based on the user's selection on the vertical tangent plane.
[0327] The rendering device acquires the two-dimensional coordinates of a point selected by a user on a vertical plane, and the distance (hereinafter referred to as depth) between the vertical plane on which the point is located and the circle center, and converts the two-dimensional coordinates and the depth into a first sound source position according to the following formula 6. The first sound source position includes the azimuth angle, the tilt angle, and the distance between the loudspeaker device and the mobile phone.
[0328] For example, the foregoing example is still used. Further, when the vertical plane is jumped by default, as shown in FIG. 37, the spherical view, the vertical plane, and the depth control bar can be displayed on the right side of the interface of the rendering device. The depth control bar is used to adjust the distance between the vertical plane and the sphere center. The user can tap a point (x, y) on the horizontal plane (as shown by 114). Correspondingly, the position of the point in the spherical coordinate system is displayed at the upper right corner of the spherical view. In addition, when the horizontal tangent plane is jumped by default, the user can tap a meridian (as shown by 113 in FIG. 37) in the spherical view. In this case, the interface displays the vertical plane in the interface shown in FIG. 37. Of course, the user may alternatively adjust the distance between the vertical plane and the sphere center through a slide operation (as shown by 115 in FIG. 37). The second position information includes the two-dimensional coordinates (x, y) of the point and the depth r. The first sound source position is obtained using formula 6.
[0329] Formula 6 can be as follows: Formula 6:
Number
Number
Number
[0330] x represents the horizontal coordinate of the point selected by the user on the vertical plane, y represents the vertical coordinate of the point selected by the user on the vertical plane, r represents the depth, λ represents the azimuth angle of the i-th loudspeaker device, Φ represents the tilt angle of the i-th loudspeaker device, and d represents the distance between the i-th loudspeaker device and the mobile phone (which can also be understood as the distance between the i-th loudspeaker device and the user). %360 is used to adjust the angle range from 0 degrees to 360 degrees. For example,
Number
Number
[0331] It can be understood that Equation 6 is merely an example. In actual applications, Equation 6 may be in another form. This is not particularly limited in this specification.
[0332] 2.2. The second position information is obtained based on the user's selection on the horizontal plane.
[0333] The rendering device obtains the two-dimensional coordinates of the point selected by the user on the horizontal plane and the distance between the horizontal plane where the point is located and the center of the circle (hereinafter referred to as the depth), and converts the two-dimensional coordinates and the depth into the first sound source position according to the following Equation 7. The first sound source position includes the azimuth angle, the tilt angle, and the distance between the loudspeaker device and the mobile phone.
[0334] For example, the foregoing example is still used. Further, when the horizontal tangent plane is jumped by default, as shown in FIG. 38, the spherical view, the horizontal tangent plane, and the depth control bar can be displayed on the right side of the interface of the rendering device. The depth control bar is used to adjust the distance between the horizontal tangent plane and the center of the sphere. The user can tap a point (x, y) on the horizontal plane (as shown by 117). Correspondingly, the position of the point in the spherical coordinate system is displayed at the upper right corner of the spherical view. In addition, when the vertical tangent plane is jumped by default, the user can tap the latitude (as shown by 116 in FIG. 38) within the spherical view. In this case, the interface displays the horizontal tangent plane in the interface shown in FIG. 38. Naturally, the user may alternatively adjust the distance between the horizontal tangent plane and the center of the sphere through a slide operation (as shown by 118 in FIG. 38). The second position information includes the two-dimensional coordinates (x, y) of the point and the depth r. The first sound source position is obtained using Equation 7.
[0335] Equation 7 can be as follows: Equation 7:
Number
Number
Number
[0336] x represents the horizontal coordinate of the point selected by the user on the vertical plane, y represents the vertical coordinate of the point selected by the user on the vertical plane, r represents the depth, λ represents the azimuth angle of the i-th loudspeaker device, Φ represents the tilt angle of the i-th loudspeaker device, and d represents the distance between the i-th loudspeaker device and the mobile phone (which can also be understood as the distance between the i-th loudspeaker device and the user). %360 is used to adjust the angle range from 0 degrees to 360 degrees. For example,
Number
Number
[0337] it can be understood that Equation 7 is merely an example. In actual applications, Equation 7 may be in another form. This is not particularly limited in this specification.
[0338] It can be understood that the above two methods of reference position information are merely examples. In actual applications, the reference position information may alternatively have other cases. This is not particularly limited in this specification.
[0339] In this way, the user can use the spherical view (for example, through a second operation such as tap, drag, slide, etc.) to select the second position information in order to control the selected sound object and perform real-time or subsequent dynamic rendering, and assign a specific spatial orientation and specific movement to the sound object so as to implement the interaction generation between the user and the audio and provide a new experience for the user. In addition, when the user does not have a sensor, the sound image of the sound object can be further edited.
[0340] In the second case, the reference information includes the media information of the multimedia file.
[0341] The media information in this embodiment of the present application includes at least one of text that needs to be displayed in the multimedia file, an image that needs to be displayed in the multimedia file, the musical characteristics of the music in the multimedia file, the sound source type corresponding to the first sound object, etc. This is not particularly limited in this specification.
[0342] In addition, determining the first sound source position of the first sound object based on the musical characteristics of the music in the multimedia file or the sound source type corresponding to the first sound object can be understood as automatic 3D remixing. Determining the first sound source position of the first sound object based on the position text that needs to be displayed in the multimedia file or the image that needs to be displayed in the multimedia file can be understood as multimodal remixing. Explanations will be provided separately below.
[0343] For example, the aforementioned example is still used. The rendering device determines that the first sound object is human speech, and the rendering device can display the interface shown in FIG. 26. The user can tap the "rendering method selection icon" 107 to select a rendering method, and the rendering method is used to determine the first sound source position of the first sound object. As shown in FIG. 39, the rendering device can display a drop-down list in response to the tap operation. The drop-down list can include "automatic rendering option" and "interactive rendering option". The "interactive rendering option" corresponds to the reference position information, and the "automatic rendering option" corresponds to the media information. Further, as shown in FIG. 39, the user can tap the "automatic rendering option" 119 to determine that the rendering method is automatic rendering.
[0344] 1. Automatic 3D remixing For example, as shown in FIG. 39, the user can tap the "automatic rendering option" 119. The rendering device can display the drop-down list shown in FIG. 40 in response to the tap operation. The drop-down list can include "automatic 3D remixing option" and "multimodal remixing option". Further, the user can tap the "automatic 3D remixing option" 120. In addition, after the user selects automatic 3D remixing, the rendering device can display the interface shown in FIG. 41. In the interface, "rendering method selection" may be replaced with "automatic 3D remixing" to prompt the user that the current rendering method is automatic 3D remixing.
[0345] The following describes multiple cases of automatic 3D remixing.
[0346] 1.1. The media information includes the musical characteristics of the music within the multimedia file.
[0347] The musical characteristics in this embodiment of the present application can be at least one of musical structure, musical emotion, singing mode, etc. The musical structure may include a prelude, the human voice in the prelude, a verse, a bridge, a refrain, etc. The musical emotion includes a sense of happiness, a tragic sense, a sense of panic, etc., and the singing mode includes a solo, a chorus, an accompaniment, etc.
[0348] After determining the multimedia file, the rendering device can analyze the musical characteristics in the audio track (which can also be understood as audio, music, etc.) of the multimedia file. Of course, the musical characteristics may alternatively be identified in a manual manner or a neural network manner. This is not particularly limited in this specification. After the musical characteristics are identified, the first sound source position corresponding to the musical characteristics may be determined based on a pre-set association relationship. The association relationship is the relationship between the musical characteristics and the first sound source position.
[0349] For example, the above example is still used. The rendering device determines that the first sound source position is around, and the rendering device can display the interface shown in FIG. 41. The spherical view within the interface displays the motion track of the first sound source position.
[0350] As described above, the musical structure may generally include at least one of a prelude, the human voice in the prelude, a verse, a bridge, or a refrain. Hereinafter, for the purpose of exemplary explanation, the analysis of the music structure is used.
[0351] Optionally, the human voice and instrument sounds in the music can be separated manually or by a neural network method. This is not particularly limited in this specification. After the human voice is separated, the music can be segmented by determining the dispersion of the mute segments and the pitch of the human voice. The specific steps include that when the mute segment of the human voice is greater than a threshold value (e.g., 2 seconds), the segment is considered to have ended. Based on this, the large segments of the music are divided. If there is no human voice in the first large segment, that large segment is determined to be the prelude of the instrument. If there is a human voice in the first large segment, the first large segment is determined to be the prelude of the human voice. The large intermediate mute segment is determined to be the bridge. Further, the center frequency of each large segment containing the human voice (referred to as the large human voice segment) is calculated according to the following formula 8, and the dispersion of the center frequency at all instants within the large human voice segment is calculated. The large human voice segments are sorted based on the dispersion. The large human voice segments ranked in the first 50% of the dispersion are marked as the refrain, and the large human voice segments ranked in the last 50% of the dispersion are marked as the verse. That is, the musical characteristics of the music are determined based on the frequency variation. In subsequent renderings, for different large segments, the sound source position or the motion track of the sound source position may be determined based on a pre-set association relationship, and then different large segments of the music are rendered.
[0352] For example, if the musical feature is an intro, the first sound source position is determined to be orbiting (or understood to be surrounding) above the user. First, the multi-channel is downmixed to a mono-channel (e.g., an average is calculated), and then the entire human voice is set to orbit around the head throughout the intro phase. The speed at each instant is determined based on the value of the human voice energy (represented by RMS or variance). Higher energy indicates a higher rotational speed. If the musical feature is a panic, the first sound source position is determined to be sudden right and sudden left. If the musical feature is a chorus, the human voice in the left channel and the human voice in the right channel can be expanded and enlarged to increase the delay. The number of instruments in each period is determined. If there is an instrument solo, the instrument is enabled to circulate based on the energy within the solo time segment.
[0353] Equation 8 can be as follows: Equation 8:
Number
[0354] f c represents the center frequency of the large human voice segment per second, N represents the number of large segments, N is a positive integer, and 0 < n < N - 1. f(n) represents the frequency domain obtained by performing a Fourier transform on the time domain waveform corresponding to the large segment, and x(n) represents the energy corresponding to the frequency.
[0355] It can be understood that Equation 8 is merely an example. In actual applications, Equation 8 may be in another form. This is not particularly limited in this specification.
[0356] In this way, based on the musical features of the music, orientation and dynamics settings are implemented for the extracted specific sound objects, as a result, the 3D rendering becomes more natural and the artistic aspects are better reflected.
[0357] 1.2. The media information includes a sound source type corresponding to the first sound object.
[0358] The sound source type in this embodiment of the present application may be a person or an instrument, or may be a drum sound, a piano sound, etc. In actual applications, the sound source type may be classified based on requirements. This is not particularly limited in this specification. Of course, the rendering device can identify the sound source type in a manual manner or a neural network manner. This is not particularly limited in this specification.
[0359] After the sound source type is identified, the first sound source position corresponding to the sound source type may be determined based on a pre-set association relationship, and the association relationship is the relationship between the sound source type and the first sound source position (this is the same as the aforementioned musical features, and the details will not be described again in this specification).
[0360] It can be understood that the two methods of the aforementioned automatic 3D remixing are merely examples. In actual applications, the automatic 3D remixing may have other cases. This is not particularly limited in this specification.
[0361] 2. Multimodal remixing For example, as shown in FIG. 42, the user can select a multimedia file by tapping the “input file selection icon” 121. In this specification, an example where the multimedia file is “car.mkv” is used. The rendering device receives the fourth operation of the user, and in response to the fourth operation, it can also be understood that the rendering device selects “car.mkv” (i.e., the target file) from the storage area as the multimedia file. The storage area may be a storage area within the rendering device or a storage area within an external device (e.g., a USB flash drive). This is not particularly limited in this specification. After the user selects the multimedia file, the rendering device can display the interface shown in FIG. 43. In the interface, “input file selection” may be replaced with “car.mkv” to prompt the user that the current multimedia file is car.mkv. In addition, by using the identification network and / or separation network in the embodiment shown in FIG. 4, the rendering device can identify the sound objects in “car.mkv” and separate the single-object audio tracks corresponding to each sound object. For example, the rendering device identifies that the sound objects included in “car.mkv” are the sounds of people, vehicles, and wind. As shown in FIGS. 43 and 44, the interfaces displayed by the rendering device may each further include an object bar, and icons such as a “human voice icon”, a “vehicle icon”, and a “wind sound icon” may be displayed in the object bar for the user to select the sound objects to be rendered.
[0362] In the following, multiple cases of multimodal remixing will be described.
[0363] 2.1. The media information includes the images that need to be displayed within the multimedia file.
[0364] Optionally, after the rendering device obtains a multimedia file (including an audio track of an image or an audio track of a video), the video can be split into frames of the image (there can be one or more frames of the image), the third position information of the first sound object is obtained based on the frames of the image, the first sound source position is obtained based on the third position information, and the third position information includes the two-dimensional coordinates and depth of the first sound object in the image.
[0365] Optionally, the specific steps for obtaining the first sound source position based on the third position information may include inputting the frames of the image into a detection network and obtaining the tracking box information (x0, y0, w0, h0) corresponding to the first sound object in the frames of the image. Of course, the frames of the image and the first sound object may alternatively be used as the input to the detection network, and the detection network outputs the tracking box information of the first sound object. The tracking box information includes the two-dimensional coordinates (x0, y0) of the corner points in the tracking box, and the height h0 and width w0 of the tracking box. The rendering device calculates the tracking box information (x0, y0, w0, h0) by using Equation 9 to obtain the coordinates (x c , y c ) of the center point within the tracking box, and then inputs the coordinates (x c , y c ) of the center point within the tracking box into a depth estimation network to obtain the relative depth
Number
Number
[0366] Equation 9 can be as follows: Equation 9: [Number] , and [Number] Equation 10: [Number] Equation 11: [Number] , [Number] , and [Number] Equation 12: λ i = x norm * θ x_max , Φ i = y norm * θ y_max , and [Number]
[0367] (x0, y0) represents the two-dimensional coordinates of the corner point (for example, the lower left corner point) within the tracking box, h0 represents the height of the tracking box, w0 represents the width of the tracking box, h1 represents the height of the image, and w1 represents the width w1 of the image. [Number] represents the relative depth of each point within the tracking box, where z c represents the average depth of all points within the tracking box. θ x_max represents the maximum horizontal angle of the playback device (when the playback device is an N - loudspeaker device, the N loudspeaker devices have the same playback device information), and θ y_max represents the maximum vertical angle of the playback device, and d y_max represents the maximum depth of the playback device, and λ i represents the azimuth angle of the i - th loudspeaker device, and Φ i represents the tilt angle of the i - th loudspeaker device, and r i represents the distance between the i - th loudspeaker device and the user.
[0368] It can be understood that Equations 9 to 12 are merely examples. In actual applications, Equations 9 to 12 may each have different forms. This is not particularly limited in this specification.
[0369] For example, as shown in FIG. 43, the user can tap on the "Multimodal Remixing Option" 122. In response to the tap operation, the rendering device can display the interface shown in FIG. 43. The right side of the interface includes the frame of the image of "car.mkv" (e.g., the first frame) and playback device information, and the playback device information includes the maximum horizontal angle, the maximum vertical angle, and the maximum depth. If the playback device is a headset, the user can input the playback device information. If the playback device is a loudspeaker device, the user may input the playback device information, or may directly use the calibration information obtained in the calibration stage as the playback device information. This is not particularly limited herein. Further, after the user selects the multimodal remixing option, the rendering device can display the interface shown in FIG. 44. In the interface, "Rendering Method Selection" may be replaced with "Multimodal Remixing" to prompt the user that the current rendering method is multimodal remixing.
[0370] When the media information includes an image that needs to be displayed within a multimedia file, there are multiple ways to determine the first sound object within the image. These ways are described separately below.
[0371] (1) The first sound object is determined through a tap by the user on the object bar.
[0372] Optionally, the rendering device may determine the first sound object based on a tap by the user on the object bar.
[0373] For example, as shown in FIG. 44, the user may determine that the sound object to be rendered is a vehicle by tapping on the "vehicle icon" 123. The rendering device displays a tracking box of the vehicle on the image of "car.mkv" on the right side, obtains third position information, and converts the third position information to the first sound source position using Equations 9 to 12. In addition, the interface further includes the coordinates (x0, y0) of the corner point at the lower left corner within the tracking box and the coordinates (x c , y c ). For example, in the loudspeaker device information, the maximum horizontal angle is 120 degrees, the maximum vertical angle is 60 degrees, and the maximum depth is 10 (the unit may be meters, decimeters, etc. and is not limited in this specification).
[0374] (2) The first sound object is determined by the user tapping on the image.
[0375] Optionally, the rendering device may use, as the first sound object, the sound object determined by the user by performing a third operation (e.g., tapping) on the image with respect to the image.
[0376] For example, as shown in FIG. 45, the user may determine the first sound object by tapping on the sound object (as indicated by 124) within the image.
[0377] (3) The first sound object is determined based on the default settings.
[0378] Optionally, the rendering device can identify the sound object by using the audio track corresponding to the image, track the default sound object or all sound objects within the image, and determine the third position information. The third position information includes the two-dimensional coordinates of the sound object within the image and the depth of the sound object within the image.
[0379] For example, the rendering device can specifically select the "combination" in the object bar by default, track all sound objects in the image, and separately determine the third position information of the first sound object.
[0380] It can be understood that the above-mentioned multiple methods for determining the first sound object in the image are merely examples. In actual applications, the first sound object in the image may alternatively be determined by another method. This is not particularly limited herein.
[0381] In this way, after the coordinates of the sound object and the single object audio track are extracted with reference to the multimodal features of audio, video, and images, a 3D immersive feeling is obtained through rendering in a headset or loudspeaker environment. In this way, the audio and the image can be synchronized, and as a result, the user can obtain an optimal sound effect experience. In addition, the technology of tracking and rendering the object audio in the entire video after the sound object is selected can also be applied to professional mixing post-production to improve the working efficiency of the mixing engineer. The single object audio track of the audio in the video is separated, and the sound object in the video image is analyzed and tracked to obtain the movement information of the sound object, and real-time or subsequent dynamic rendering is performed on the selected sound object. In this way, the video image is matched with the sound source direction of the audio, and as a result, the user experience is improved.
[0382] 2.2. The media information includes position text that needs to be displayed in the multimedia file.
[0383] In this case, the rendering device can determine a first sound source position based on the position text that needs to be displayed in the multimedia file, and the position text indicates the first sound source position.
[0384] Optionally, the position text may be understood as text having a meaning such as a position or an orientation. For example, the wind is blowing northward, blowing upward into the sky, blowing toward the shell, front, rear, left, right, etc. This is not particularly limited in this specification. Of course, the position text may specifically be lyrics, subtitles, an advertising slogan, etc. This is not particularly limited in this specification.
[0385] Optionally, the semantics of the displayed position text can be identified based on reinforcement learning or a neural network, and then the first sound source position is determined based on the semantics.
[0386] In this way, the position text related to the position is identified, and a 3D immersion feeling is obtained through rendering in a headset or loudspeaker environment. In this way, a sense of space corresponding to the position text is achieved, and as a result, the user obtains an optimal sound effect experience.
[0387] It can be understood that the above two methods of media information are merely examples. In actual applications, the media information may alternatively have other cases. This is not particularly limited in this specification.
[0388] In addition, in step 1003, how to determine the first sound source position based on the reference information will be described in multiple cases. In actual applications, the first sound source position may alternatively be determined in a combined manner. For example, after the first sound source position is determined using the sensor orientation, the motion track of the first sound source position is determined using musical characteristics. For example, as shown in FIG. 46, assume that the rendering device determines that based on the first attitude angle of the sensor, the sound source position of the human voice has become as shown on the right side of the interface in FIG. 46. Further, the user can determine the motion track of the human voice by tapping the "Circling Option" 125 in the menu on the right side of the "Human Voice Icon". It can also be understood that the first sound source position at a certain moment is first determined by using the sensor orientation, and then the motion track of the first sound source position is determined as a circle by using musical characteristics or according to pre-set rules. Accordingly, as shown in FIG. 46, the interface of the rendering device can display the motion track of the sound object.
[0389] Optionally, in the above-described process of determining the first sound source position, the user can control the distance at the first sound source position by controlling the volume button of the mobile phone or through taps, drags, slides, etc. on the spherical view.
[0390] Step 1004: Perform spatial rendering on the first single object audio track based on the first sound source position.
[0391] After determining the first sound source position, the rendering device can perform spatial rendering on the first single object audio track to obtain the rendered first single object audio track.
[0392] Optionally, the rendering device performs spatial rendering on the first single-object audio track based on the first sound source position to obtain the rendered first single-object audio track. Of course, the rendering device alternatively performs spatial rendering on the first single-object audio track based on the first sound source position and performs rendering on the second single-object audio track based on the second sound source position to obtain the rendered first single-object audio track and the rendered second single-object audio track.
[0393] Optionally, when spatial rendering is performed on a plurality of single-object audio tracks corresponding to a plurality of sound objects in the original audio track, the combination of the plurality of methods in step 1003 may be used in the method for determining the sound source position in this embodiment of the present application. This is not particularly limited herein.
[0394] For example, as shown in FIG. 32, the first sound object is a person, the second sound object is a violin, the method in the interactive rendering can be used for the first sound source position of the first single-object audio track corresponding to the first sound object, and the method in the automatic rendering can be used for the second sound source position of the second single-object audio track corresponding to the second sound object. The specific methods for determining the first sound source position and the second sound source position may be any two methods of the aforementioned step 1003. Of course, the specific methods for determining the first sound source position and the second sound source position may alternatively be the same method. This is not particularly limited herein.
[0395] In addition, in the attached drawings including the spherical view, the spherical view can further include a volume bar. The user can control the volume of the first single-object audio track by performing operations such as finger sliding, mouse dragging, and mouse wheel scrolling on the volume bar. This improves the real-time performance of rendering the audio track. For example, as shown in FIG. 47, the user may adjust the volume bar 126 to adjust the volume of the single-object audio track corresponding to the guitar.
[0396] The rendering method in this step may vary depending on different types of playback devices. It can also be understood that the method used by the rendering device to perform spatial rendering on the original audio track or the first single-object audio track based on the first sound source position and the type of the playback device varies depending on different types of playback devices and will be described separately below.
[0397] In the first case, the type of the playback device is a headset.
[0398] In this case, after determining the first sound source position, the rendering device can render the audio track based on the HRTF filter coefficient table according to Equation 13. The audio track can be the first single-object audio track, or the second single-object audio track, or the first single-object audio track and the second single-object audio track. This is not particularly limited herein. The HRTF filter coefficient table shows the association relationship between the sound source position and the coefficient. It can be understood that one sound source position corresponds to one HRTF filter coefficient.
[0399] Equation 13 can be as follows: Equation 13:
Number
[0400]
Number
[0401] It can be understood that Equation 13 is merely an example. In actual applications, Equation 13 may be in another form. This is not particularly limited in this specification.
[0402] In the second method, the type of the playback device is a loudspeaker device.
[0403] In this case, after determining the first sound source position, the rendering device can render the audio track according to Equation 14. The audio track can be the first single-object audio track, or the second single-object audio track, or the first single-object audio track and the second single-object audio track. This is not particularly limited in this specification.
[0404] Equation 14 may be as follows: Equation 14: [Number] , where [Number] , where [Number]
[0405] There may be N loudspeaker devices, [Number] represents the rendered first single object audio track, i represents the i-th channel among the plurality of channels, S represents the sound object of the multimedia file, the sound object includes the first sound object, a s (t) represents the adjustment coefficient of the first sound object at instant t, g s (t) represents the translation coefficient of the first sound object at instant t, o s (t) represents the first single object audio track at instant t, λ i represents the azimuth angle obtained when the calibrator (e.g., the aforementioned sensor device) calibrates the i-th loudspeaker device, Φ i represents the tilt angle obtained when the calibrator calibrates the i-th loudspeaker device, r i represents the distance between the i-th loudspeaker device and the calibrator, N is a positive integer, i is a positive integer, i ≤ N, and the first sound source position is within the tetrahedron formed by the N loudspeaker devices.
[0406] In addition, for spatial rendering of the original audio track, a single object audio track corresponding to the sound object in the original audio track may be rendered, for example, S1 in the above formula may be replaced. Alternatively, the single object audio track corresponding to the sound object in the original audio track may be rendered after being duplicated and added, for example, it is S2 in the above formula. Of course, it may alternatively be a combination of S1 and S2.
[0407] It can be understood that Equation 14 is merely an example. In actual applications, Equation 14 may be in another form. This is not particularly limited in this specification.
[0408] To facilitate the understanding of N, please refer to FIG. 48. The figure is a schematic diagram of the architecture of a loudspeaker device in a spherical coordinate system. When the sound source position of the sound object is within the tetrahedron formed by four loudspeaker devices, N = 4. When the sound source position of the sound object is on the surface of the region formed by three loudspeaker devices, N = 3. When the sound source position of the sound object is on the connection line between two loudspeaker devices, N = 2. When the sound source position of the sound object directly points to one loudspeaker device, N = 1. Since the point in FIG. 48 is on the connection line between loudspeaker device 1 and loudspeaker device 2, N shown in FIG. 48 is 2.
[0409] Step 1005: Obtain a target audio track based on the rendered first single object audio track.
[0410] The target audio track obtained in this step may vary depending on different types of playback devices. It can be understood that the method used by the rendering device to obtain the target audio track varies with different types of playback devices and will be described separately below.
[0411] In the first case, the type of the playback device is a headset.
[0412] In this case, after obtaining the rendered first single-object audio track and / or the rendered second single-object audio track, the rendering device can obtain the target audio track based on the rendered audio track according to Equation 15. The audio track can be the first single-object audio track, or the second single-object audio track, or both the first single-object audio track and the second single-object audio track. This is not particularly limited herein.
[0413] Equation 15 can be as follows: Equation 15:
Number
[0414] i represents the left channel or the right channel,
Number
Number
[0415] It can be understood that Formula 15 is merely an example. In actual applications, Formula 15 may be in another form. This is not particularly limited in this specification.
[0416] In the second method, the type of the playback device is a loudspeaker device.
[0417] In this case, after obtaining the rendered first single object audio track and / or the rendered second single object audio track, the rendering device can obtain the target audio track according to Formula 16 based on the rendered audio track. The audio track can be the first single object audio track, or the second single object audio track, or both the first single object audio track and the second single object audio track. This is not particularly limited in this specification.
[0418] Formula 16 can be as follows Formula 16:
Number
Number
Number
[0419] There may be N loudspeaker devices, and i represents the i-th channel among a plurality of channels.
Number
[0420] It can be understood that Formula 16 is merely an example. In actual applications, Formula 16 may be in another form. This is not particularly limited herein.
[0421] Of course, the new multimedia file can alternatively be generated based on the multimedia file and the target audio track. This is not particularly limited herein.
[0422] In addition, after the first single-object audio track is rendered, the user may upload to the database module corresponding to FIG. 8 the method of setting the sound source position by the user in the rendering process, and as a result, another user renders another audio track in the method of that setting. Of course, the user may alternatively download the setting method from the database module and modify the setting method so that the user performs spatial rendering on the audio track. In this way, modification of the rendering rules and sharing among different users are added. In this way, in the multimodal mode, repeated object identification and tracking for the same file can be avoided, and as a result, the overhead on the terminal side is reduced. In addition, the free creation of the user in the interactive mode can be shared with another user, and as a result, the application interaction is further enhanced.
[0423] For example, as shown in FIG. 49, the user can choose to synchronize the rendering rule file stored in the local database to another device of the user. As shown in FIG. 50, the user may choose to upload the rendering rule file stored in the local database to the cloud for sharing with another user, and another user may choose to download the corresponding rendering rule file from the cloud database to the terminal side.
[0424] The metadata files stored in the database are mainly used to render sound objects separated by the system or specified by the user in automatic mode, or to render sound objects that need to be automatically rendered according to the specified and stored rendering rules in hybrid mode by the user. The metadata files stored in the database may be pre-specified in the system as shown in Table 1.
[0425]
Table 1A
Table 1B
[0426] Sequence numbers 1 and 2 in Table 1 may alternatively be generated during production by the user when the user uses the interactive mode of the present invention, for example, sequence numbers 3 to 6 in Table 1, or may be stored after the system automatically identifies the motion track of a specified sound object in a video picture in the multimodal mode, for example, sequence number 7 in Table 1. The metadata file may be strongly related to the audio content in the multimedia file or the multimodal file content. For example, in Table 1, sequence number 3 represents the metadata file corresponding to audio file A1, and sequence number 4 represents the metadata file corresponding to audio file A2. Alternatively, the metadata file may be separated from the multimedia file. The user performs an interactive operation on object X in audio file A in the interactive mode and stores the motion track of object X as the corresponding metadata file (for example, sequence number 5 in Table 1 represents a free spiral ascending state). Next, when automatic rendering is used, the user can select, from the database module, the metadata file corresponding to the free spiral ascending state in order to render object Y in audio file B.
[0427] In a possible implementation, the rendering method provided in the embodiments of the present application includes steps 1001 to 1005. In another possible implementation, the rendering method provided in the embodiments of the present application includes steps 1002 to 1005. In another possible implementation, the rendering method provided in the embodiments of the present application includes steps 1001 to 1004. In another possible implementation, the rendering method provided in the embodiments of the present application includes steps 1002 to 1004. In addition, in the embodiments of the present application, the time series relationship between the steps shown in FIG. 10 is not limited. For example, step 1001 in the foregoing method may alternatively be performed after step 1002. Specifically, the playback device is calibrated after the audio track is acquired.
[0428] In an embodiment of the present application, the user can control the sound image and volume of a specific sound object in the audio, as well as the number of specific sound objects in the audio, by using a mobile phone sensor. The sound image and volume of a specific sound object, as well as the number of specific sound objects, are controlled through a drag in the mobile phone interface, and spatial rendering is performed on specific sound objects in the music according to automatic rules to improve the sense of space. The sound source position is automatically rendered through multimodal identification. According to the method of rendering a single sound object, a sound effect experience completely different from the conventional music / movie interaction mode is provided. Thereby, a new interaction mode for music appreciation is provided. Automatic 3D playback improves the sense of space of dual-channel music and improves the music listening level. In addition, separation is introduced into the designed interaction method to enhance the user's audio editing ability. The interaction method can be applied to the production of sound objects in music or movies and TV works to easily edit the movement information of specific sound objects. Furthermore, the controllability and reproducibility of the music for the user are increased, and as a result, the user experiences the pleasure of generating audio by the user and the ability to control specific sound objects.
[0429] In addition to the foregoing training method and rendering method, the present application further provides two specific application scenarios to which the foregoing rendering method is applied. The two specific application scenarios are described separately below.
[0430] The first application scenario is the "Sound Hunter" game scenario.
[0431] This scenario can also be understood as determining whether the sound source position indicated by the user matches the actual sound source position when the user indicates the sound source position, scoring the user's actions, and improving the user's entertainment experience.
[0432] For example, the foregoing example is still used. After the user determines a multimedia file and renders a single object audio track, as shown in FIG. 51, the user can tap the "Sound Hunter Icon" 126 to enter the sound hunter game scenario, the rendering device can display the interface shown in FIG. 51, and the user can decide to start the game by tapping the play button at the bottom of the interface. The playback device plays at least one single object audio track at an arbitrary position in a specific sequence. When the playback device plays the single object audio track of the piano, the user determines the sound source position based on hearing and holds the mobile phone to indicate the sound source position determined by the user. If the position indicated by the user's mobile phone coincides with (or the error is within a specific range) the actual sound source position of the single object audio track of the piano, the rendering device can display a prompt on the right side of the interface shown in FIG. 51: "You hit the first instrument, which took 5.45 seconds and beat 99.33% of people worldwide." In addition, after the user indicates the correct position of the sound object, the corresponding sound object in the object bar can change from red to green. Of course, if the user indicates an incorrect position of the sound source within a specific period, a failure may be displayed. As shown in FIGS. 52 and 53, after the first single object audio track is played and a preset period (for example, the time interval T in FIG. 54) has elapsed, the next single object audio track is played to continue the game. In addition, after the user indicates an incorrect position of the sound object, the corresponding sound object in the object bar can remain red. The rest can be inferred by analogy (as shown in FIG. 54). After the user presses the pause button at the bottom of the interface or after the single object audio track is played, the game is determined to end.Furthermore, after the game ends, if the user points to the correct position several times, the rendering device can display the interface shown in FIG. 53.
[0433] In this scenario, the user interacts with the playback system in real time to render the orientation of the object audio in the playback system in real time. The game is designed so that the user can obtain the ultimate "audio localization" experience. The present invention can be applied to home entertainment, AR and VR games, etc. Compared with the prior art where the "audio localization" technology is only for the whole piece of music, the present application provides a game that is implemented after the human voice and the sound of musical instruments are separated from the music.
[0434] The second application scenario is a multi-user interaction scenario.
[0435] This scenario can be understood as multiple users each controlling the sound source position of a specific sound object, and as a result, multiple users each rendering an audio track to increase entertainment and communication among multiple users. For example, the interaction scenario may specifically be a scenario where multiple people create a band online, a scenario where an anchor controls a symphony online, etc.
[0436] For example, a multimedia file is music reproduced by using a plurality of musical instruments. User A can select the multi - person interaction mode and invite User B to complete the production jointly. Each user can select a different musical instrument as an interactive sound object for control. After rendering is performed based on the rendering track provided by the user corresponding to the sound object, remixing is completed, and then the audio file obtained through remixing is sent to each participating user. The interaction modes selected by different users can be different. This is not particularly limited in this specification. For example, as shown in FIG. 55, User A selects an interactive mode in which interactive control is performed on the position of Object A by changing the orientation of the mobile phone used by User A, and User B selects an interactive mode in which interactive control is performed on the position of Object B by changing the orientation of the mobile phone used by User B. As shown in FIG. 56, the system (rendering device or cloud server) can send the audio file obtained through remixing to each user participating in the multi - person interaction application, and the positions of Object A and Object B in the audio file respectively correspond to the control of User A and the control of User B.
[0437] An example of the aforementioned specific interaction process between User A and User B is as follows: User A selects an input multimedia file, and the system identifies the object information in the input file and feeds back the object information to User A via the UI interface. User A selects a mode. If User A selects the multi-user interaction mode, User A sends a multi-user interaction request to the system and sends information about the specified invitees to the system. In response to the request, the system sends an interactive request to User B selected by User A. If the request is accepted, User B sends an acceptance command to the system to participate in the multi-user interaction application created by User A. User A and User B each select a sound object to be operated and control the sound object selected in the aforementioned rendering mode and the corresponding rendering rule file. The system separates the single-object audio track by using a separate network, renders the separated single-object audio track based on the rendering track provided by the user corresponding to the sound object, remixes the rendered single-object audio track to obtain the target audio track, and then sends the target audio track to each participating user.
[0438] In addition, the multi-person interaction mode may be the real-time online multi-person interaction described in the foregoing example, or may be an offline multi-person interaction. For example, the multimedia file selected by User A is a duet music including Singer A and Singer B. As shown in FIG. 57, User A may select an interaction mode to control the rendering effect of Singer A, and share the target audio track obtained through re-rendering with User B. User B may use the received target audio track shared by User A as an input file to control the rendering effect of Singer B. The interaction modes selected by different users may be the same or different. This is not particularly limited herein.
[0439] It can be understood that the foregoing multiple application scenarios are merely examples. In actual applications, there may be other application scenarios. This is not particularly limited herein.
[0440] In this scenario, real-time and non-real-time interactive rendering control involving multiple people is supported. A user can invite another user to jointly complete the re-rendering and production of different sound objects of a multimedia file, thereby improving the interaction experience and the fun of the application. The foregoing method is used to implement the rendering of a multimedia file by multiple people and to implement audio-visual control for different objects through multi-person cooperation.
[0441] The above describes the rendering method in the embodiments of the present application. The following describes the rendering device in the embodiments of the present application. Referring to FIG. 58. The embodiment of the rendering device in an embodiment of the present application is Obtain a first single-object audio track based on the multimedia file, where the first single-object audio track includes an acquisition unit 5801 configured to correspond to a first sound object, and Determine a first sound source position of the first sound object based on reference information, where the reference information includes reference position information and / or media information of the multimedia file, and the reference position information includes a determination unit 5802 configured to indicate the first sound source position, and A rendering unit 5803 configured to perform spatial rendering on the first single-object audio track based on the first sound source position to obtain the rendered first single-object audio track.
[0442] In this embodiment, the operations performed by the units in the rendering device are the same as those described in the embodiments shown in FIGS. 5 to 11. Details are not repeated herein.
[0443] In this embodiment, the acquisition unit 5801 obtains a first single-object audio track based on the multimedia file, the first single-object audio track corresponds to a first sound object, the determination unit 5802 determines a first sound source position of the first sound object based on reference information, and the rendering unit 5803 performs spatial rendering on the first single-object audio track based on the first sound source position to obtain the rendered first single-object audio track. The stereo spatial sense of the first single-object audio track corresponding to the first sound object in the multimedia file can be improved, and as a result, an immersive stereo sound effect is provided to the user.
[0444] Referring to FIG. 59. Another embodiment of the rendering device in an embodiment of the present application is Obtain a first single-object audio track based on a multimedia file, where the first single-object audio track includes an acquisition unit 5901 configured to correspond to a first sound object, and Determine a first sound source position of the first sound object based on reference information, where the reference information includes reference position information and / or media information of the multimedia file, and the reference position information is a determination unit 5902 configured to indicate the first sound source position, and A rendering unit 5903 configured to perform spatial rendering on the first single-object audio track based on the first sound source position to obtain the rendered first single-object audio track.
[0445] The rendering device in this embodiment Includes a providing unit 5904 configured to provide a spherical view for the user to select, where the center of the circle of the spherical view is the user's position, and the radius of the spherical view is the distance between the user's position and the playback device. Further includes a transmission unit 5905 configured to transmit a target audio track to a playback device, and the playback device is configured to play the target audio track.
[0446] In this embodiment, the operations performed by the units in the rendering device are the same as those described in the embodiments shown in FIGS. 5 to 11. Details are not repeated herein.
[0447] In this embodiment, the acquisition unit 5901 acquires a first single-object audio track based on a multimedia file, the first single-object audio track corresponds to a first sound object, the determination unit 5902 determines a first sound source position of the first sound object based on reference information, and the rendering unit 5903 performs spatial rendering on the first single-object audio track based on the first sound source position to obtain a rendered first single-object audio track. The stereo spatial perception of the first single-object audio track corresponding to the first sound object in the multimedia file can be improved, and as a result, an immersive stereo sound effect is provided to the user. In addition, a sound effect experience completely different from the conventional music / movie interaction mode is provided. Thereby, a new interaction mode for music appreciation is provided. Automatic 3D playback improves the spatial perception of dual-channel music and improves the music listening level. In addition, in order to enhance the user's audio editing ability, separation is introduced into the designed interaction method. The interaction method can be applied to the production of sound objects in music or movies and TV works in order to easily edit the movement information of specific sound objects. Furthermore, the user's controllability and reproducibility for music are increased, and as a result, the user experiences the pleasure of generating audio by the user and the ability to control specific sound objects.
[0448] Refer to FIG. 60. The embodiment of the rendering device in an embodiment of the present application is an acquisition unit 6001 configured to acquire a multimedia file, wherein the acquisition unit 6001 is further configured to acquire a first single-object audio track based on the multimedia file, and the first single-object audio track corresponds to a first sound object; A display unit 6002 configured to display a user interface, the user interface including a rendering method option, the display unit 6002, and A determination unit 6003 configured to determine an automatic rendering method or an interactive rendering method from the rendering method options in response to a first action of a user in the user interface, and When the determination unit determines the automatic rendering method, the acquisition unit 6001 is further configured to acquire a first single-object audio track rendered in a preset method, or When the determination unit determines the interactive rendering method, the acquisition unit 6001 acquires reference position information in response to a second action of the user, determines a first sound source position of a first sound object based on the reference position information, and is further configured to render a first single-object audio track based on the first sound source position to acquire the rendered first single-object audio track.
[0449] In this embodiment, the operations performed by the units in the rendering device are the same as those described in the embodiments shown in FIGS. 5 to 11. Details are not repeated herein.
[0450] In this embodiment, the determination unit 6003 determines an automatic rendering method or an interactive rendering method from the rendering method options based on the first action of the user. In one aspect, the acquisition unit 6001 can automatically acquire a first single-object audio track rendered based on the first action of the user. In another aspect, the spatial rendering of the audio track corresponding to the first sound object in the multimedia file can be implemented through the interaction between the rendering device and the user so that an immersive stereo sound effect is provided to the user.
[0451] FIG. 61 is a schematic diagram of the structure of another rendering device according to the present application. The rendering device may include a processor 6101, a memory 6102, and a communication interface 6103. The processor 6101, the memory 6102, and the communication interface 6103 are connected to each other via a line. The memory 6102 stores program instructions and data.
[0452] The memory 6102 stores program instructions and data corresponding to the steps implemented by the rendering device in the corresponding implementation forms shown in FIGS. 5 to 11.
[0453] The processor 6101 is configured to implement the steps implemented by the rendering device in any one of the embodiments shown in FIGS. 5 to 11.
[0454] The communication interface 6103 may be configured to receive and transmit data, and is configured to implement the steps related to acquisition, transmission, and reception in any one of the embodiments shown in FIGS. 5 to 11.
[0455] In one implementation form, the rendering device may include more or fewer components than the components shown in FIG. 61. This is merely an example for explanation and is not limited in the present application.
[0456] As shown in FIG. 62, an embodiment of the present application further provides a sensor device. For ease of explanation, only the parts related to the embodiments of the present application are shown. For specific technical details not disclosed, refer to the method part in the embodiments of the present application. The sensor device may be any terminal device such as a mobile phone or a tablet computer. For example, the sensor is a mobile phone.
[0457] FIG. 62 is a block diagram showing a partial structure of a sensor device (i.e., a mobile phone) according to an embodiment of the present technology. Refer to FIG. 62. The mobile phone includes components such as a radio frequency (RF) circuit 6210, a memory 6220, an input unit 6230, a display unit 6240, a sensor 6250, an audio circuit 6260, a wireless fidelity (WiFi) module 6270, a processor 6280, and a power supply 6290. Those skilled in the art can understand that the structure of the mobile phone shown in FIG. 62 does not constitute a limitation on the mobile phone, and the mobile phone may include more or fewer components than those shown in the figure, or some components may be combined, or different component arrangements may be used.
[0458] Hereinafter, with reference to FIG. 62, each component of the mobile phone will be described in detail.
[0459] The RF circuit 6210 receives and transmits signals in the information reception / transmission process or call processing. In particular, after receiving the downlink information of the base station, it transmits the downlink information to the processor 6280 for processing and transmits the related uplink data to the base station. The RF circuit 6210 may typically include, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. Further, the RF circuit 6210 may further communicate with the network and other devices via wireless communication. Any communication standard or protocol may be used for wireless communication, including, but not limited to, Global System for Mobile Communication (GSM) for mobile communication, General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0460] Memory 6220 can be configured to store software programs and modules. Processor 6280 performs various functional applications and data processing on the mobile phone by executing the software programs and modules stored in memory 6220. Memory 6220 may mainly include a program storage area and a data storage area. The program storage area may store an operating system, application programs required for at least one function (such as a voice playback function, an image playback function, etc.). The data storage area can store data created based on the use of the mobile phone (such as audio data, phone book, etc.). In addition, memory 6220 may include a high-speed random access memory and may further include a non-volatile memory such as at least one magnetic storage device, a flash memory, or another volatile solid-state storage device.
[0461] The input unit 6230 may be configured to receive the input digital or character information and generate button signal inputs related to user settings and function controls on the mobile phone. Specifically, the input unit 6230 may include a touch panel 6231 and another input device 6232. The touch panel 6231, also referred to as a touch screen, can collect touch operations (e.g., operations performed by the user on or near the touch panel 6231 using any suitable object or accessory such as a finger or a stylus) and drive the corresponding connection device based on a pre-set program. Optionally, the touch panel 6231 may include two parts, namely, a touch detection device and a touch controller. The touch detection device detects the touch position of the user, detects the signal generated by the touch operation, and transmits the signal to the touch controller. The touch controller can receive touch information from the touch detection device, convert the touch information into touch point coordinates, transmit the touch point coordinates to the processor 6280, receive and execute the command transmitted by the processor 6280. Also, the touch panel 6231 may be implemented in multiple ways such as the resistive film method, the capacitance method, the infrared method, the surface acoustic wave method, etc. In addition to the touch panel 6231, the input unit 6230 may further include another input device 6232. Specifically, the other input device 6232 may include one or more of, but not limited to, a physical keyboard, function buttons (such as volume control buttons, on / off buttons, etc.), a trackball, a mouse, a joystick, etc.
[0462] The display unit 6240 may be configured to display information input by or provided to the user, as well as various menus on the mobile phone. The display unit 6240 may include a display panel 6241. Optionally, the display panel 6241 may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like. Also, the touch panel 6231 may cover the display panel 6241. After detecting a touch operation on or near the touch panel 6231, the touch panel 6231 transfers the touch operation to the processor 6280 to determine the type of touch event. Then, the processor 6280 provides a corresponding visual output on the display panel 6241 based on the type of touch event. In FIG. 62, the touch panel 6231 and the display panel 6241 function as two independent components for implementing the input and input functions of the mobile phone. However, in some embodiments, the touch panel 6231 and the display panel 6241 may be integrated to implement the input and output functions of the mobile phone.
[0463] The mobile phone may further include at least one sensor 6250, such as an optical sensor, a motion sensor, and another sensor. Specifically, the optical sensor includes an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel 6241 based on the brightness of the ambient light, and the proximity sensor can turn off the display panel 6241 and / or the backlight when the mobile phone moves close to the ear. As a type of motion sensor, the acceleration sensor may detect the values of acceleration in all directions (usually three axes), may also detect the value and direction of gravity when the mobile phone is stationary, and is used in applications for identifying the posture of the mobile phone (such as switching between landscape mode and portrait mode, related games, or attitude calibration of a magnetometer), functions related to vibration identification (such as a pedometer or a knock), etc. Another sensor such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, an IMU sensor, a SLAM sensor, etc. may be further configured in the mobile phone. Details are not described in this specification.
[0464] The audio circuit 6260, the speaker 6262, and the microphone 6262 can provide an audio interface between the user and the mobile phone. The audio circuit 6260 can convert the received audio data into an electrical signal and transmit the electrical signal to the speaker 6262, and the speaker 6262 converts the electrical signal into an audio signal for output. Also, the microphone 6262 converts the collected audio signal into an electrical signal. The audio circuit 6260 receives the electrical signal, converts the electrical signal into audio data, and then outputs the audio data to the processor 6280 for processing, and through the RF circuit 6210, sends the audio data to, for example, another mobile phone or outputs the audio data to the memory 6220 for further processing.
[0465] WiFi is a short-range wireless transmission technology. By using the WiFi module 6270, a mobile phone can help a user receive and send emails, browse web pages, access streaming media, etc., and provide the user with wireless broadband Internet access. Although Figure 62 shows the WiFi module 6270, it can be understood that the WiFi module 6270 is not an essential component of the mobile phone.
[0466] The processor 6280 is the control center of the mobile phone connected to various parts of the mobile phone through various interfaces and lines. It implements or executes software programs and / or modules stored in the memory 6220, and calls the data stored in the memory 6220 to implement various functions and data processing of the mobile phone and perform overall monitoring of the mobile phone. Optionally, the processor 6280 may include one or more processing units. Preferably, an application processor and a modem processor may be integrated into the processor 6280. The application processor mainly processes the operating system, user interface, application programs, etc. The modem processor mainly processes wireless communication. It can be understood that the modem processor may not be alternatively integrated into the processor 6280.
[0467] The mobile phone further includes a power source 6290 (e.g., a battery) that supplies power to the components. Preferably, the power source can be logically connected to the processor 6280 by using a power management system to implement functions such as charge management, discharge management, and power consumption management.
[0468] Although not shown in the figure, the mobile phone may further include a camera, a Bluetooth module, etc. Details are not described in this specification.
[0469] In this embodiment of the present application, the processor 6280 included in the mobile phone can implement the functions in the embodiments shown in FIGS. 5 to 11. Details are not repeated herein.
[0470] For convenience and brevity of description, for the detailed operation processes of the above-mentioned systems, devices, and units, reference may be made by those skilled in the art to the corresponding processes in the embodiments of the above-mentioned methods. Details are not repeated herein.
[0471] It should be understood that in the multiple embodiments provided in the present application, the disclosed systems, devices, and methods may be implemented in other ways. For example, the above-described device embodiments are merely examples. For example, the division into units is merely a logical function division. At the actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not implemented. In addition, the mutual coupling, direct coupling, or communication connection shown or discussed may be implemented through some interfaces. The indirect coupling or communication connection between devices or units may be implemented in electronic, mechanical, or other forms.
[0472] The units described as separate parts may or may not be physically separate, and the parts shown as units may or may not be physical units. They may be located at one position or dispersed on multiple network units. Some or all of the units may be selected based on actual requirements to achieve the objectives of the embodiment solutions.
[0473] In addition, the functional units in the embodiments of the present application may be integrated into one processing unit, each of the units may exist physically alone, or two or more units may be integrated into one unit. The integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0474] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, the integrated unit may be stored in a computer-readable storage medium. Based on such understanding, the essence of the technical solution of this application, or the part that contributes to the prior art, or all or part of the technical solution, may be implemented in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for instructing a computer device (which may be a personal computer, a server, a network device, etc.) to implement all or part of the steps of the method in the embodiments of this application. The aforementioned storage medium includes any medium that can store program code, such as a USB flash drive, a removable hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, etc.
[0475] In the description, claims, and accompanying drawings of this application, terms such as "first", "second", etc. are intended to distinguish similar objects, but do not necessarily indicate a specific order or sequence. Such terms are interchangeable in appropriate situations, and it should be understood that this is merely a method of distinguishing objects having the same attributes in the embodiments of this application. In addition, the terms "include", "have", and any other variations are meant to cover non-exclusive inclusion, so that a process, method, system, product, or device that includes a series of units is not necessarily limited to those units, and may include other units not explicitly listed or not specific to such a process, method, system, product, or device.
Description of Reference Numerals
[0476] 40 Neural Network Processing Unit 100 System Architecture, Convolutional Network 101 Target model / rule, playback device type selection icon 102 Loudspeaker device option 103 Calibration icon 104 Manual calibration option 105 Input file selection icon 106 Human voice icon 107 Rendering method selection icon 108 Interactive rendering option 109 Orientation control option 110 Execution device, input layer, violin icon 111 Calculation module, position control option 112 I / O interface, interface control option 113 Preprocessing module 119 Automatic rendering option 120 Training device, convolutional layer / pooling layer, automatic 3D remixing option 121 Layer, input file selection icon 122 Layer, multimodal remixing option 123 Layer, vehicle icon 124 Layer 125 Layer, orbiting option 126 Layer, sound hunter icon 130 Database, neural network layer 131 Hidden layer 1 132 Hidden layer 2 13n Hidden layer n 140 Client device, output layer 150 Data storage system 160 Data collection device 401 Input memory 402 Weight memory 403 Arithmetic circuit 404 Controller 405 Direct memory access controller 406 Integrated memory 407 Vector Calculation Unit 408 Accumulator 409 Instruction Fetch Buffer 410 Bus Interface Unit 901 Control Device 902 Sensor Device 903 Reproduction Device 5801 Acquisition Unit 5802 Decision Unit 5803 Rendering Unit 5901 Acquisition Unit 5902 Decision Unit 5903 Rendering Unit 5904 Provision Unit 5905 Transmission Unit 6001 Acquisition Unit 6002 Display Unit 6003 Decision Unit 6101 Processor 6102 Memory 6103 Communication Interface 6210 RF Circuit 6220 Memory 6230 Input Unit 6231 Touch Panel 6232 Another Input Device 6240 Display Unit 6241 Display Panel 6250 Sensor 6260 Audio Circuit 6261 Loudspeaker 6262 Microphone 6270 WiFi Module 6280 Processor 6290 Power Supply
Claims
1. A rendering method, comprising: obtaining a first single-object audio track based on a multimedia file, wherein the first single-object audio track corresponds to a first sound object; obtaining first position information including the attitude angle of a sensor and a distance between the sensor and a playback device, the distance being determined based on a ratio of the energy of a test signal received by the sensor to the energy of the test signal played by the playback device; determining a first sound source position of the first sound object based on the first position information; performing spatial rendering on the first single-object audio track based on the first sound source position to obtain a rendered first single-object audio track.
2. The method according to claim 1, wherein the multimedia file includes media information, and the media information includes at least one of text that needs to be displayed in the multimedia file, an image that needs to be displayed in the multimedia file, musical characteristics of music that needs to be played in the multimedia file, and a sound source type corresponding to the first sound object.
3. The method further includes: determining the type of the playback device, wherein the playback device is configured to play a target audio track, and the target audio track is obtained based on the rendered first single-object audio track; the step of performing spatial rendering on the first single-object audio track based on the first sound source position includes: performing spatial rendering on the first single-object audio track based on the first sound source position and the type of the playback device.
4. The step of obtaining the first single-object audio track based on the multimedia file includes: The method according to claim 1, comprising the step of separating the first single object audio track from the original audio track in the multimedia file, wherein the original audio track is obtained by combining at least the first single object audio track and a second single object audio track, and the second single object audio track corresponds to a second sound object.
5. The step of separating the first single object audio track from the original audio track in the multimedia file The method according to claim 4, comprising the step of separating the first single object audio track from the original audio track by using a trained separation network.
6. The trained separation network is obtained by training the separation network by using training data as an input to the separation network and using a value of a loss function less than a first threshold as a target, the training data includes a training audio track, the training audio track is obtained by combining at least an initial third single object audio track and an initial fourth single object audio track, the initial third single object audio track corresponds to a third sound object, the initial fourth single object audio track corresponds to a fourth sound object, the third sound object and the first sound object have the same type, the second sound object and the fourth sound object have the same type, and the output of the separation network includes a third single object audio track obtained through separation, The method according to claim 5, wherein the loss function indicates a difference between the third single object audio track obtained through separation and the initial third single object audio track.
7. The step of performing spatial rendering on the first single-object audio track based on the first sound source position and the type of the playback device is when the playback device is a headset, 【Number 1】 according to, a step of obtaining the rendered first single-object audio track, 【Number 2】 represents the rendered first single object audio track, S represents the sound object of the multimedia file, the sound object includes the first sound object, i represents the left channel or the right channel, and a s (t) represents the adjustment coefficient of the first sound object at instant t, and h i,s (t) represents the head-related transfer function (HRTF) filter coefficient of the left channel or the right channel corresponding to the first sound object at instant t, the HRTF filter coefficient is related to the first sound source position, and o s (t) represents the first single object audio track at instant t, τ represents the integral term, a method according to claim 3, including steps. **Claim 8** The step of performing spatial rendering on the first single-object audio track based on the first sound source position and the type of the playback device is when the playback device is an N-channel loudspeaker device, 【Mathematics 3】 according to, a step of obtaining the rendered first single-object audio track, 【Number 4】 wherein 【Number 5】 wherein 【Number 6】 represents the rendered first single object audio track, i represents the i-th channel among a plurality of channels, S represents the sound object of the multimedia file, the sound object includes the first sound object, a s (t) represents the adjustment coefficient of the first sound object at instant t, g s (t) represents the translation coefficient of the first sound object at instant t, o s (t) represents the first single object audio track at instant t, λ i represents the azimuth angle obtained when the calibrator calibrates the i-th loudspeaker device, Φ i represents the tilt angle obtained when the calibrator calibrates the i-th loudspeaker device, r i represents the distance between the i-th loudspeaker device and the calibrator, N is a positive integer, i is a positive integer, i ≦ N, and the first sound source position is within the tetrahedron formed by the N loudspeaker devices. The method according to claim 3, comprising the step of **Claim 9** The method further includes a step of obtaining the target audio track based on the rendered first single-object audio track, the original audio track in the multimedia file, and the type of the playback device; and a step of transmitting the target audio track to the playback device, wherein the playback device is configured to play the target audio track. The method according to claim 3. **Claim 10** The step of obtaining the target audio track based on the rendered first single-object audio track, the original audio track in the multimedia file, and the type of the playback device is when the type of the playback device is a headset, 【Number 7】 according to, a step of obtaining the target audio track, where i represents the left channel or the right channel, 【Number 8】 represents the target audio track at instant t, and X i (t) represents the original audio track at the instant t, and 【Number 9】 represents the first single-object audio track not being rendered at the instant t, 【Number 10】 represents the rendered first single object audio track, a s (t) represents the adjustment coefficient of the first sound object at the instant t, h i,s (t) represents the head-related transfer function (HRTF) filter coefficient of the left or right channel corresponding to the first sound object at the instant t, and the HRTF filter coefficient is related to the first sound source position, o s (t) represents the first single object audio track at the instant t, τ represents the integral term, S 1 represents the sound object that needs to be replaced in the original audio track. When the first sound object replaces the sound object in the original audio track, S 1 represents the null set, S 2 represents the sound object added to the target audio track compared to the original audio track. When the first sound object is a copy of the sound object in the original audio track, S 2 represents the null set, S 1 and / or S 2 represents the sound object of the multimedia file, the sound object includes the first sound object, includes steps, the method according to claim 9. **Claim 11** The step of obtaining the target audio track based on the rendered first single-object audio track, the original audio track in the multimedia file, and the type of the playback device is when the type of the playback device is an N-channel loudspeaker device, 【Number 11】 according to, a step of obtaining the target audio track, 【Number 12】 wherein 【Number 13】 and i represents the i-th channel among a plurality of channels, 【Number 14】 represents the target audio track at instant t, and X i (t) represents the original audio track at the instant t, 【Number 15】 represents the first single-object audio track that has not been rendered at the moment t, 【Number 16】 represents the rendered first single object audio track, a s (t) represents the adjustment coefficient of the first sound object at the instant t, g s (t) represents the translation coefficient of the first sound object at the instant t, g i,s (t) is g s (t) represents the i-th row in g s (t) represents the first single object audio track at the instant t, S 1 represents the sound object that needs to be replaced in the original audio track. When the first sound object replaces the sound object in the original audio track, S 1 represents the null set, S 2 represents the sound object added to the target audio track compared to the original audio track. When the first sound object is a copy of the sound object in the original audio track, S 2 represents the null set, S 1 and / or S 2 represents the sound object of the multimedia file. The sound object includes the first sound object, λ i represents the azimuth angle obtained when the calibrator calibrates the i-th loudspeaker device, Φ i represents the tilt angle obtained when the calibrator calibrates the i-th loudspeaker device, r i represents the distance between the i-th loudspeaker device and the calibrator. N is a positive integer, i is a positive integer, i ≦ N, and the first sound source position is within the tetrahedron formed by the N loudspeaker devices. The method according to claim 9, comprising the steps.
12. A rendering device, configured to obtain a first single-object audio track based on a multimedia file, the first single-object audio track corresponding to a first sound object and including an attitude angle of a sensor and a distance between the sensor and a playback device, the distance being determined based on a ratio of the energy of a test signal received by the sensor to the energy of the test signal reproduced by the playback device, and configured to obtain first position information including the distance; an acquisition unit; a determination unit configured to determine a first sound source position of the first sound object based on the first position information; a rendering unit configured to perform spatial rendering on the first single-object audio track based on the first sound source position to obtain the rendered first single-object audio track. A rendering device comprising:
13. The multimedia file includes media information, the media information including at least one of text that needs to be displayed in the multimedia file, an image that needs to be displayed in the multimedia file, musical features of music that needs to be played in the multimedia file, and a sound source type corresponding to the first sound object. The rendering device according to claim 12.
14. The determination unit is further configured to determine a type of the playback device, the playback device being configured to play a target audio track, the target audio track being obtained based on the rendered first single-object audio track, The rendering unit is particularly configured to perform spatial rendering on the first single-object audio track based on the first sound source position and the type of the playback device. The rendering device according to claim 12.
15. The acquisition unit separates the first single object audio track from the original audio track in the multimedia file, the original audio track being obtained by combining at least the first single object audio track and a second single object audio track, the second single object audio track being specifically configured to correspond to a second sound object, the rendering device according to claim 12.
16. The acquisition unit is specifically configured to separate the first single object audio track from the original audio track by using a trained separation network, the rendering device according to claim 15.
17. The trained separation network is obtained by training the separation network by using training data as an input to the separation network and using a value of a loss function less than a first threshold as a target, the training data including a training audio track, the training audio track being obtained by combining at least an initial third single object audio track and an initial fourth single object audio track, the initial third single object audio track corresponding to a third sound object, the initial fourth single object audio track corresponding to a fourth sound object, the third sound object and the first sound object having the same type, the second sound object and the fourth sound object having the same type, the output of the separation network including a third single object audio track obtained through separation, the loss function indicating a difference between the third single object audio track obtained through separation and the initial third single object audio track, the rendering device according to claim 16.
18. When the playback device is a headset, the acquisition unit 【Number 17】 obtains the rendered first single object audio track according to 【Number 18】 represents the rendered first single object audio track, S represents the sound object of the multimedia file, the sound object includes the first sound object, i represents the left channel or the right channel, a s (t) represents the adjustment coefficient of the first sound object at instant t, h i,s (t) represents the head-related transfer function (HRTF) filter coefficient of the left channel or the right channel corresponding to the first sound object at instant t, the HRTF filter coefficient is related to the first sound source position, o s (t) represents the first single object audio track at instant t, τ represents an integral term, and is specifically configured as such. The rendering device according to claim 14.
19. When the playback device is N loudspeaker devices, the acquisition unit 【Number 19】 acquires the rendered first single object audio track according to 【Number 20】 and 【Number 21】 and 【Number 22】 represents the rendered first single object audio track, i represents the i-th channel among a plurality of channels, S represents the sound object of the multimedia file, the sound object includes the first sound object, a s (t) represents the adjustment coefficient of the first sound object at instant t, g s (t) represents the translation coefficient of the first sound object at instant t, o s (t) represents the first single object audio track at instant t, λ i represents the azimuth angle obtained when the calibrator calibrates the i-th loudspeaker device, Φ i represents the tilt angle obtained when the calibrator calibrates the i-th loudspeaker device, r i represents the distance between the i-th loudspeaker device and the calibrator, N is a positive integer, i is a positive integer, i ≦ N, and the first sound source position is particularly configured to be within the tetrahedron formed by the N loudspeaker devices. The rendering device according to claim 14.
20. The acquisition unit is further configured to acquire the target audio track based on the rendered first single object audio track and the original audio track in the multimedia file, The rendering device further includes a transmission unit configured to transmit the target audio track to the playback device, and the playback device is configured to play the target audio track. The rendering device according to claim 14.
21. When the playback device is a headset, the acquisition unit 【Number 23】 acquires the target audio track according to where i represents the left channel or the right channel, 【24 Points】 represents the target audio track at instant t, and X i (t) represents the original audio track at the instant t, and 【Number 25】 represents the first single object audio track that is not being rendered at the instant t, 【Number 26】 represents the rendered first single object audio track, a s (t) represents the adjustment coefficient of the first sound object at the instant t, h i,s (t) represents the head-related transfer function (HRTF) filter coefficient of the left or right channel corresponding to the first sound object at the instant t, and the HRTF filter coefficient is related to the first sound source position, o s (t) represents the first single object audio track at the instant t, τ represents the integral term, S 1 represents the sound object that needs to be replaced in the original audio track. When the first sound object replaces the sound object in the original audio track, S 1 represents the null set, S 2 represents the sound object added to the target audio track compared to the original audio track. When the first sound object is a duplicate of the sound object in the original audio track, S 2 represents the null set, S 1 and / or S 2 represents the sound object of the multimedia file, and the sound object is specifically configured to include the first sound object. The rendering device according to claim 20
22. When the playback device is N loudspeaker devices, the acquisition unit 【Number 27】 acquires the target audio track according to 【Number 28】 and 【No. 29】 and where i represents the i-th channel among a plurality of channels, 【30 numbers】 represents the target audio track at instant t, and X i (t) represents the original audio track at the instant t, and 【Number 31】 represents the first single object audio track that is not being rendered at the instant t, 【Number 32】 represents the rendered first single object audio track, a s (t) represents the adjustment coefficient of the first sound object at the instant t, g s (t) represents the translation coefficient of the first sound object at the instant t, g i,s (t) is g s represents the i-th row in g s (t) represents the first single object audio track at the instant t, S 1 represents the sound object that needs to be replaced in the original audio track. When the first sound object replaces the sound object in the original audio track, S 1 represents the null set, S 2 represents the sound object added to the target audio track compared to the original audio track. When the first sound object is a copy of the sound object in the original audio track, S 2 represents the null set, S 1 and / or S 2 represents the sound object of the multimedia file. The sound object includes the first sound object, λ i represents the azimuth angle obtained when the calibrator calibrates the i-th loudspeaker device, Φ i represents the tilt angle obtained when the calibrator calibrates the i-th loudspeaker device, r i represents the distance between the i-th loudspeaker device and the calibrator. N is a positive integer, i is a positive integer, i ≦ N, and the first sound source position is specifically configured to be within the tetrahedron formed by the N loudspeaker devices. The rendering device according to claim 20.
23. A rendering device comprising a processor, the processor being coupled to a memory, the memory being configured to store a program or instructions, and when the program or the instructions are executed by the processor, the rendering device is enabled to implement the method according to any one of claims 1 to 11.
24. A computer-readable storage medium, the computer-readable storage medium storing instructions, and when the instructions are executed on a computer, the computer is enabled to implement the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Music data processing method and device and computer storage medium
CN112037738A
Method and device for reproducing synthesized voice
JP2003099078A
Audio signal processing method and audio signal processing device using the same
JP2014522181A
Signal processing device, method, and program
JP2022017880A
Method and apparatus for improved matching of auditory space to visual space in video viewing applications
US20100328419A1