Audio data processing method and device, storage medium and electronic equipment

By obtaining track points and speaker gains in audio data processing, dynamically matching virtual speakers and physical devices, the problem of poor audio data playback effect on multi-channel devices is solved, and immersive sound field reconstruction and coordinated enhancement of audio and video spatial features is achieved.

CN120386508APending Publication Date: 2025-07-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510466670.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing audio data processing methods are difficult to adapt to differentiated sound field requirements on multi-channel devices, resulting in poor user immersion, and the voice barrage and main soundtrack are prone to conflict, affecting the playback effect.

Method used

By obtaining the audio and track points to be processed, initial audio data is generated, the location information of the track points is used to determine the virtual source position in the virtual playback scene, and mixing is performed based on the speaker gain, dynamically matching the virtual speakers and physical devices to realize multi-dimensional fusion and gain control of the audio data.

Benefits of technology

It realizes universal restoration of sound field and immersive sound field reconstruction in complex environments, effectively suppresses multi-channel superposition distortion, and improves the playback effect of audio data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386508A_ABST
    Figure CN120386508A_ABST
Patent Text Reader

Abstract

The invention discloses an audio data processing method and device, a storage medium and electronic equipment. The method comprises the steps that to-be-processed audio and at least one track point are acquired, initial audio data are generated based on the to-be-processed audio and the track point, loudspeaker gain corresponding to a target virtual loudspeaker is determined based on the initial audio data, and position information of the track point is used for determining a virtual source position in a virtual playing scene. Virtual loudspeakers are deployed in the virtual playing scene, the target virtual loudspeaker is at least one of the virtual loudspeakers, the target virtual loudspeaker represents the virtual loudspeaker determined based on the virtual source position, and the loudspeaker gain is used for simulating playing of the to-be-processed audio at the virtual source position; and performing sound mixing processing on the audio to be processed and the media information indicated by the source media information identifier based on the loudspeaker gain to generate target audio data. The technical problem that the playing effect of the audio data is poor is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computers, and in particular, to a method and apparatus for processing audio data, a storage medium, and an electronic device. Background Art

[0002] In the field of multimedia interaction, the processing method of audio data (for example, voice bullet screens) usually adopts an audio superposition scheme, that is, directly mixing and playing the user's voice with the main audio track of the video. However, such a scheme has significant defects. When voice bullet screens appear densely, it is easy to conflict with the key content of the main audio track (such as dialogue, background music).

[0003] In the related art, the mixing mode with fixed gain is difficult to meet the differentiated sound field requirements of multi-channel devices, and the user's immersion is poor. Due to the single generation and playback methods of audio data, there will be technical problems of poor playback effect of audio data.

[0004] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention

[0005] Embodiments of the present application provide a method and apparatus for processing audio data, a storage medium, and an electronic device, so as to at least solve the technical problem of poor playback effect of audio data.

[0006] According to one aspect of the embodiments of the present application, a method for processing audio data is provided, including: obtaining the audio to be processed and at least one trajectory point, generating initial audio data based on the audio to be processed and the trajectory point, where the initial audio data includes a source media information identifier corresponding to the audio to be processed and position information of the trajectory point; determining a speaker gain corresponding to a target virtual speaker based on the initial audio data, where the position information of the trajectory point is used to determine a virtual source position in a virtual playback scene, the virtual playback scene is deployed with virtual speakers, the target virtual speaker is at least one of the virtual speakers, the target virtual speaker represents a virtual speaker determined based on the virtual source position, and the speaker gain is used to simulate the audio to be processed playing at the virtual source position; performing a mixing process on the audio to be processed and the media information indicated by the source media information identifier based on the speaker gain to generate target audio data.

[0007] According to another aspect of the embodiments of the present application, there is also provided a processing device for audio data, including: an acquisition module, configured to acquire the audio to be processed and at least one track point, and generate initial audio data based on the audio to be processed and the track point, where the initial audio data includes a source media information identifier corresponding to the audio to be processed and position information of the track point; a determination module, configured to determine a speaker gain corresponding to a target virtual speaker based on the initial audio data, where the position information of the track point is used to determine a virtual source position in a virtual playback scenario, the virtual playback scenario is deployed with virtual speakers, the target virtual speaker is at least one of the virtual speakers, the target virtual speaker represents a virtual speaker determined based on the virtual source position, and the speaker gain is used to simulate the playback of the audio to be processed at the virtual source position; a generation module, configured to perform a mixing process on the audio to be processed and the media information indicated by the source media information identifier based on the speaker gain to generate target audio data.

[0008] In an exemplary embodiment, the device is configured to determine a speaker gain corresponding to a target virtual speaker based on the initial audio data in the following manner: determine the virtual source position based on the position information; determine the target virtual speaker and the speaker gain according to the virtual source position and the speaker positions corresponding to the respective virtual speakers in the virtual playback scenario, where the target virtual speaker is a speaker in the virtual playback scenario whose distance from the virtual source position satisfies a preset distance condition.

[0009] In an exemplary embodiment, the device is configured to determine the virtual source position based on the position information in the following manner: convert the position information into spatial track point coordinates in a spherical coordinate system, where the position information represents track point coordinates in a rectangular coordinate system; determine a target position vector based on the spatial track point coordinates, and determine the position indicated by the target position vector as the virtual source position, where the target position vector is used to indicate the spatial position of the corresponding track point in the virtual playback scenario.

[0010] In an exemplary embodiment, the device is configured to determine the target virtual speaker and the speaker gain according to the virtual source position and the speaker positions corresponding to the respective virtual speakers in the virtual playback scene in the following manner: obtain the distances between the respective virtual speakers in the virtual playback scene and the virtual source position; determine the speakers whose distances meet a preset distance condition as the target virtual speakers; determine the speaker position vectors corresponding to the target virtual speakers; and determine a gain coefficient based on the target position vector and the speaker position vectors, wherein the speaker position vectors, after being weighted and summed according to the gain coefficient, are the same as the target position vector, and the speaker gain includes the gain coefficient.

[0011] In an exemplary embodiment, the device is configured to obtain the audio to be processed and at least one trajectory point, and generate initial audio data based on the audio to be processed and the trajectory point in the following manner: obtain the audio to be processed and mark the audio timestamp; receive the trajectory points input by the user, where each trajectory point includes a spatial position parameter and associated time information; and bind the audio to be processed and the spatial position parameter based on the audio timestamp and the associated time information to generate the initial audio data.

[0012] In an exemplary embodiment, the device is configured to obtain the audio to be processed and mark the audio timestamp by at least one of the following methods: in the case where the media information indicated by the source media information identifier includes a recorded video, mark the audio timestamp based on the playback progress of the recorded video, where the audio timestamp indicates that the audio to be processed is set to be played at the corresponding playback progress; in the case where the media information indicated by the source media information identifier includes a live video, mark the audio timestamp based on the system time, where the audio timestamp indicates that the audio to be processed is set to be played at the corresponding system time.

[0013] In an exemplary embodiment, the device is configured to bind the audio to be processed and the spatial position parameter in at least one of the following manners based on the audio timestamp and the associated time information to generate the initial audio data: when the media information indicated by the source media information identifier includes a recorded video, bind the associated time information to the first playback progress of the recorded video, where the first playback progress indicates that when the recorded video is played to the first playback progress, the audio to be processed is played according to the speaker gain set at the corresponding trajectory point; when the media information indicated by the source media information identifier includes a live video, bind the associated time information to the second playback progress of the audio to be processed, where the second playback progress indicates that when the audio to be processed is played to the second playback progress, the audio to be processed is played according to the speaker gain set at the corresponding trajectory point.

[0014] In an exemplary embodiment, the device is configured to obtain the audio to be processed and at least one trajectory point, and generate initial audio data based on the audio to be processed and the trajectory point: in response to a first interaction operation performed on the application interface, start recording the audio to be processed; in response to the end of recording the audio to be processed, display a trajectory point generation interface, and display the at least one trajectory point in the virtual playback scene in the trajectory point generation interface.

[0015] In an exemplary embodiment, the device is configured to, in response to the end of recording the audio to be processed, display a trajectory point generation interface and display the at least one trajectory point in the virtual playback scene in the trajectory point generation interface in the following manner: in response to the end of recording the audio to be processed, determine the number of trajectory points based on the duration of the audio to be processed; display the trajectory point generation interface, and display the trajectory points in the virtual playback scene in the trajectory point generation interface according to the number of trajectory points.

[0016] In an exemplary embodiment, the apparatus is configured to determine a speaker gain corresponding to a target virtual speaker based on the initial audio data in the following manner: when the number of track points meets a preset number condition, determine a target position vector corresponding to each track point based on the position information corresponding to each track point, and determine the position indicated by each target position vector as the virtual source position corresponding to each track point, where the target position vector is used to indicate the spatial position of the corresponding track point in the virtual playback scene; obtain the distances between each virtual speaker and each virtual source position in the virtual playback scene; determine the speakers whose distances meet a preset distance condition as the target virtual speakers corresponding to each track point, and determine the speaker position vectors corresponding to each target virtual speaker corresponding to each track point; based on the target position vector and the speaker position vector, determine the gain coefficient of each target virtual speaker corresponding to each track point, where the speaker position vector is the same as the target position vector after being weighted and summed according to the gain coefficient, and the speaker gain includes the gain coefficient.

[0017] In an exemplary embodiment, the apparatus is further configured to: when the number of track points meets a preset number condition, obtain the playback time corresponding to each track point; increase track points by interpolation based on the playback time interval between a first track point and a second track point with adjacent playback times.

[0018] In an exemplary embodiment, the apparatus is configured to mix the audio to be processed and the media information indicated by the source media information identifier based on the speaker gain in the following manner to generate target audio data: obtain the information type corresponding to the media information indicated by the source media information identifier; determine the gain allocation ratio based on the information type; mix the audio to be processed and the media information indicated by the source media information identifier according to the gain allocation ratio to generate target audio data.

[0019] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium storing a computer program, where the computer program, when run by a processor, executes the above-mentioned audio data processing method.

[0020] According to another aspect of the embodiments of the present application, there is provided a computer program product including a computer program, which, when executed by a processor, implements the steps in the above-mentioned audio data processing method.

[0021] According to another aspect of the embodiments of the present application, an electronic device is further provided, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to execute the processing method of the audio data through the computer program.

[0022] In the embodiments of the present application, an audio track dynamic binding and spatial metadata fusion technology is adopted. By performing spatio-temporal correlation encoding on the audio to be processed and multi-dimensional track point information, initial audio data is generated, achieving the purpose of accurately anchoring the spatial attributes and content relevance of the sound source. By adopting a virtual sound field dynamic mapping and physical device adaptive matching technology, an acoustic model of the virtual playback scene is constructed by analyzing the position information of the track points, and the gain parameters are calculated in combination with the physical characteristics of the target virtual speaker, achieving the purpose of dynamically adapting the virtual source position to the physical playback device, thereby realizing the general technical effect of sound field restoration in complex environments. This method breaks through the limitation of fixed speaker layout on spatial audio through the intelligent matching of virtual source positions and physical devices.

[0023] Furthermore, a multi-source data spatialization mixing and dynamic gain control technology is adopted. Through a frequency band allocation and timing alignment strategy based on speaker gain, the audio to be processed and associated media information are fused in multiple dimensions, achieving the purpose of synergistically enhancing the spatial characteristics of sound and picture, thereby realizing the technical effect of immersive sound field reconstruction. This process precisely controls the energy distribution of each channel through a spatial audio rendering algorithm. While retaining the characteristics of the original audio content, elements such as voice bullet screens and environmental sound effects present azimuth movement trajectories that conform to physical laws. At the same time, dynamic level control effectively suppresses multi-channel superposition distortion, thereby solving the technical problem of poor playback effect of audio data. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0025] Figure 1 is a schematic diagram of an application environment of an optional method for processing audio data according to an embodiment of the present application;

[0026] Figure 2 is a schematic flowchart of an optional method for processing audio data according to an embodiment of the present application;

[0027] Figure 3 is a schematic diagram of an optional method for processing audio data according to an embodiment of the present application;

[0028] Figure 4 is a schematic diagram of another optional method for processing audio data according to an embodiment of the present application;

[0029] Figure 5 It is a schematic diagram of another optional method for processing audio data according to an embodiment of the present application;

[0030] Figure 6 It is a schematic diagram of another optional method for processing audio data according to an embodiment of the present application;

[0031] Figure 7 It is a schematic diagram of another optional method for processing audio data according to an embodiment of the present application;

[0032] Figure 8 It is a schematic diagram of another optional method for processing audio data according to an embodiment of the present application;

[0033] Figure 9 It is a schematic diagram of another optional method for processing audio data according to an embodiment of the present application;

[0034] Figure 10 It is a schematic diagram of another optional method for processing audio data according to an embodiment of the present application;

[0035] Figure 11 It is a schematic diagram of another optional method for processing audio data according to an embodiment of the present application;

[0036] Figure 12 It is a schematic diagram of another optional method for processing audio data according to an embodiment of the present application;

[0037] Figure 13 It is a schematic diagram of another optional method for processing audio data according to an embodiment of the present application;

[0038] Figure 14 It is a schematic diagram of another optional method for processing audio data according to an embodiment of the present application;

[0039] Figure 15 It is a schematic diagram of the structure of an optional audio data processing device according to an embodiment of the present application;

[0040] Figure 16 It is a schematic diagram of the structure of an optional audio data processing product according to an embodiment of the present application;

[0041] Figure 17 It is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present application. Detailed implementation manners

[0042] To enable those skilled in the art to better understand the solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0043] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0044] To more clearly understand the technical solutions provided by the embodiments of this application, the key terms related to the embodiments of this application will be introduced here first:

[0045] Voice barrage is a new form of interaction that combines voice interaction and real-time comment functions, mainly applied to scenarios such as videos and live broadcasts. Its core lies in that users send voice messages to replace traditional text barrages. Users can directly record voices (usually limited to 5 - 15 seconds) and send them instead of text. Tone and emotion are transmitted in real time through sound, which is especially suitable for the scene resonance of music / dance live broadcasts.

[0046] Voice barrages can be presented in the picture in forms including but not limited to the following:

[0047] Sound wave visualization: Dynamic sound wave patterns float in a specific area of the screen.

[0048] Avatar bubble: The sender's avatar slides with the voice progress bar.

[0049] Intelligent subtitles: Real-time text overlays are generated through ASR technology.

[0050] It should be noted that when there are multiple voice barrages and they intersect, a mixing special effect can be automatically triggered.

[0051] This application may include, but is not limited to, eliminating environmental noise through the RNN noise reduction algorithm, marking emotion tags by real-time analyzing pitch frequencies, and detecting illegal content through voiceprint feature comparison and / or semantic double detection.

[0052] In an exemplary embodiment, taking VR live broadcast as an example, 3D sound effect positioning can be supported in VR live broadcast, and the voice bullet screen can present a sense of orientation as the viewer's head moves, enhancing the immersive experience.

[0053] The following describes this application in conjunction with embodiments:

[0054] According to one aspect of the embodiments of this application, a method for processing audio data is provided. Optionally, in this embodiment, the above method for processing audio data can be applied to a hardware environment composed of a server 101 and a terminal device 103 as Figure 1 shown. As Figure 1 shown, the server 101 is connected to the terminal device 103 through a network and can be used to provide services for the terminal device or an application installed on the terminal device. The application can be a video application, an instant messaging application, a browser application, an educational application, a game application, etc. A database 105 can be set up on the server or independently of the server to provide data storage services. The above network can include, but is not limited to: a wired network, a wireless network. Among them, the wired network includes: a local area network, a metropolitan area network, and a wide area network, and the wireless network includes: Bluetooth, WIFI, and other networks that implement wireless communication. The terminal device 103 can be a terminal device configured with an application and can include, but is not limited to, at least one of the following: a mobile phone (such as an Android mobile phone, an iOS mobile phone, etc.), a laptop computer, a tablet computer, a handheld computer, a desktop computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal device, an aircraft, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a mixed reality (MR) terminal device, etc. computer devices. The above server can be a single server, a server cluster composed of multiple servers, or a cloud server.

[0055] Combined with Figure 1 shown, the above method for processing audio data can be executed by an electronic device. The electronic device can be a terminal device or a server. The above method for processing audio data can be implemented separately by the terminal device or the server, or jointly implemented by the terminal device and the server.

[0056] In an exemplary embodiment, through the collaborative operation of the terminal device and the server, this application realizes intelligent processing of full-link audio data, including but not limited to the following steps:

[0057] S1. The terminal device 103 acquires the original audio stream through the integrated acquisition module and performs frame segmentation processing and feature marking. The built-in preprocessing unit of the device preliminarily denoises and regularizes the format of the audio data to generate data packets that conform to the transmission specifications.

[0058] Among them, the terminal device 103 is built with a dedicated audio processing module, which effectively eliminates environmental noise interference and ensures the integrity of the original voice features. The intelligent coding technology is adopted to dynamically adapt to the network status, significantly reducing the amount of transmitted data while maintaining the voice clarity, and realizing a smooth real-time voice interaction experience.

[0059] S2. The application 107 establishes an intelligent transmission channel and dynamically selects the optimal coding scheme according to the network quality. The device identifier and time sequence mark are attached to the data packet, and it is pushed to the server 101 through the encrypted tunnel to ensure the data integrity and privacy security during the transmission process.

[0060] Among them, the server 101 can build a multi-level processing pipeline, and through the coordinated operation of voiceprint feature extraction and deep semantic analysis, it realizes high-precision voice content understanding. Combining heterogeneous computing resource scheduling technology, it intelligently allocates computing tasks to different processing units to balance the resource requirements of real-time processing and batch analysis.

[0061] S3. After receiving the data stream, the server 101 starts a multi-stage processing process, including format standardization, voiceprint feature extraction, and semantic analysis. Deploy an elastic computing resource pool, and automatically expand the processing nodes according to the data traffic to ensure the processing timeliness in high-concurrency scenarios.

[0062] S4. The database 105 implements a hierarchical storage strategy, storing the structured metadata and the time-sequenced audio segments separately. An intelligent index engine is established to support multi-dimensional joint queries based on voice content features, realizing the efficient organization and management of massive data.

[0063] Among them, the database system establishes a multi-modal storage structure, realizing the dual capabilities of fast response for hot data and efficient archiving for cold data. Through content label indexing and feature vector matching technology, it supports millisecond-level data retrieval under complex conditions, significantly improving the utilization rate of historical voice data.

[0064] S5. The server 101 constructs a data feedback loop, regularly extracts feature data from the database 105 to optimize the analysis model. The processing results are transmitted back to the terminal device 103 through compression and encapsulation technology, and the analysis conclusions are dynamically displayed on the visualization interface.

[0065] S6. The network transmission layer deploys an intelligent routing selection mechanism, automatically switching the optimal transmission path according to the real-time network status. An end-to-end quality monitoring system is established, and when the transmission quality drops, the data fidelity transmission mode is automatically enabled.

[0066] Among them, a dual-active data storage architecture and an intelligent disaster recovery mechanism are adopted to ensure the continuous stability of the data processing link. By combining a trusted computing environment and dynamic encryption technology, a full-process security protection system from terminal device collection to cloud storage is constructed.

[0067] This framework has been applied on a large scale in the field of intelligent interaction, significantly improving the audio data processing efficiency and system stability, and providing a technical foundation for building a highly reliable audio service system. Through modular design, it supports function expansion and can flexibly adapt to the personalized needs of different scenarios.

[0068] The above is only an example, and this embodiment is not specifically limited.

[0069] Optionally, as an alternative implementation, as Figure 2 shown, the above method for processing audio data includes:

[0070] S202, obtaining the audio to be processed and at least one trajectory point, and generating initial audio data based on the audio to be processed and the trajectory point, where the initial audio data includes the source media information identifier corresponding to the audio to be processed and the position information of the trajectory point;

[0071] Optionally, in the embodiments of the present application, the audio to be processed may include, but is not limited to, the original voice data recorded by the user through the terminal device, usually a mono audio stream that has not been spatialized. Specifically, it is the digital audio data formed by the analog signal collected by the microphone after analog-to-digital conversion and compression encoding, and may contain voice content, background sound or mixed sound effects. For example, its storage format can be a combination of the ADTS header and the original data encoded by AAC, and the compression parameters can include the sampling rate (such as 44.1 kHz or 48 kHz), the number of channels (mono or stereo), the bit depth (16 bit or 24 bit), etc. In addition, the recording duration of the audio to be processed can be restricted by business rules (for example, within 5 seconds), and it supports dynamic adjustment of preprocessing operations such as noise suppression and gain control.

[0072] It should be noted that there are various possibilities for the source and processing method of the audio to be processed. For example, the audio input device can include the built-in microphone of the mobile phone, the external directional microphone or the bone conduction sensor of the AR glasses; the audio format can be extended to MP3, WAV or FLAC, and the encoding standard can adopt CELT, Speex or EVS; the preprocessing link can integrate modules such as echo cancellation, voice enhancement or semantic analysis, for example, automatically filtering non-human voice noise or extracting keywords to generate subtitles through an AI model. The present application does not make specific limitations on this.

[0073] Exemplarily, Figure 3It is a schematic diagram of an optional method for processing audio data according to an embodiment of the present application. As Figure 3 shown, a voice barrage recording entry button is added to the playback page. In response to an interaction operation performed on the entry button, the voice barrage recording page is opened to start recording the above-mentioned audio to be processed.

[0074] Optionally, in the embodiment of the present application, the above-mentioned trajectory points may include, but are not limited to, dynamic position markers of user-defined voice barrages in a spatial coordinate system. Each trajectory point contains a timestamp, spatial coordinates, and motion parameters, which are used to describe the real-time azimuth change of the voice barrage during playback. For example, the spatial coordinates can adopt a rectangular coordinate system (x, y, z) or a spherical coordinate system to represent. The timestamp marks the starting moment when this position becomes effective, and the motion parameters can include uniform speed, acceleration, or a Bezier curve interpolation algorithm. The number and interval of trajectory points can be set according to business requirements (such as at most 5 points and the interval is not less than 500 ms), and it supports dragging and adding or path drawing through a graphical editor.

[0075] It should be noted that the definition and generation method of trajectory points can be flexibly adjusted according to the actual scenario. For example, the acquisition of position information can rely on manual editing (such as a user dragging path points in a 3D interface), automatic generation (such as automatically matching the surrounding trajectory of an explosion sound according to video content analysis), or external device input (such as real-time recording through the spatial positioning data of a VR handle); the reference point of the coordinate system can be set as the center of the user's head, the center of the screen, or the anchor point of the virtual environment; the motion trajectory interpolation algorithm can adopt linear interpolation, spline curves, or a physical engine to simulate parabolic motion. The present application does not make specific limitations on this.

[0076] Exemplarily, Figure 4 It is a schematic diagram of another optional method for processing audio data according to an embodiment of the present application. As Figure 4As shown, on the voice barrage recording page, the left half is for voice recording, allowing users to record sounds. The maximum recording time is only allowed to be 5 seconds. (The maximum recording duration can be changed according to business requirements). After the recording is completed, operations can be performed on the right-side sound orientation editor. It should be noted that track points can be added for the sound spatial playback position through the "Add Track" button. T is the start playback time of the current sound position, with the unit of milliseconds. X, Y, and Z are the coordinate values of the spatial coordinate system centered on the human head. After the addition is successful, the current specified position will be displayed. In addition, multiple positions can be added to the currently recorded barrage. For example, Track 2, Track 2, and Track 3. Among them, the tracks and track points mentioned in this application both represent the positions selected by the user in the virtual scene. The interval between each position needs to be no less than 500 ms. Each voice barrage shall not exceed 5 position coordinates. (Both of the above two parameters can be adjusted by the business according to the actual situation). And, multiple spatial coordinates can form a barrage movement trajectory. The barrage voice is moved between each position based on the smooth algorithm using the barrage movement trajectory. If there is only one spatial coordinate, it can be fixed at the specified position for playback.

[0077] Optionally, in the embodiment of the present application, the above initial audio data may include, but is not limited to, a structured data packet formed by binding the audio to be processed with track point information. This data packet usually contains audio content, spatial track metadata, and associated identifiers, such as audio encoding formats (such as AAC / Opus), metadata fields (such as start time, duration, coordinate system type), video source ID (for matching playback content), and user identity identifier. Its generation process may involve data encapsulation (such as binary serialization), checksum addition (such as CRC check), and encryption processing (such as AES algorithm) to ensure transmission integrity and security.

[0078] It should be noted that the encapsulation structure and transmission protocol of the initial audio data can be adapted to different business requirements. For example, the metadata part can add priority tags (such as barrages of VIP users are rendered first), emotion tags (such as "cheering" barrages automatically increase the volume), or copyright watermarks; the transmission protocol can choose HTTP / 2 streaming transmission, WebRTC peer-to-peer transmission, or blockchain distributed storage; the security mechanism can combine digital signature verification of the sender's identity or use differential privacy technology to protect user track data. The present application does not make specific limitations on this.

[0079] Optionally, in the embodiments of the present application, the above-mentioned source media information identifier may include, but is not limited to, unique index information for associating the voice barrage with the target video content. This identifier can be embodied as a video resource hash value, an internal platform number, or a timestamp combination code. For example, it can be a video ID assigned by the video platform, a shard sequence number (such as the 12th minute segment of the 3rd episode), or a user-defined tag (such as "science fiction movie climax segment"). Its function is to ensure the precise synchronization of the voice barrage with the video frame during playback and support content matching retrieval across devices and sessions.

[0080] It should be noted that the construction logic of the source media information identifier can be extended to multi-dimensional combinations. For example, the identifier can include a three-segment code of the video platform domain name, the uploader ID, and the upload time, or use content fingerprint technology to generate a unique identifier based on the key frame hash value. In the live broadcast scenario, the identifier can be synchronized with the NTP time server to generate a millisecond-level timestamp to ensure the precise alignment of the barrage with the real-time streaming media. The present application does not make specific limitations on this.

[0081] Optionally, in the embodiments of the present application, the above-mentioned position information may include, but is not limited to, a set of parameters describing the orientation, distance, and motion state of the voice barrage in three-dimensional space. Specifically, it can be decomposed into static coordinates (such as a fixed point x = 1.2, y = 0, z = 0.8), a dynamic path (such as a parabolic motion from coordinate A to B), or a relative orientation (such as 30 degrees and 2 meters in front of the listener's left). Its expression form can include Cartesian coordinate coefficient values, spherical coordinate angle-distance combinations, or polar coordinate parameters, and can be converted into multi-channel speaker gain parameters through a spatial coding algorithm (such as Ambisonics B-format).

[0082] It should be noted that the application scenarios of the position information can extend to diverse terminal devices. For example, in a car audio system, the position information can be mapped to the surround sound field of the driver's seat. In a home theater scenario, the speaker gain can be adaptively adjusted through a room calibration system (such as DiracLive). In a VR headset device, the position information can be fused with the head tracking data to achieve dynamic sound field updates. The present application does not make specific limitations on this.

[0083] S204. Determine the speaker gain corresponding to the target virtual speaker based on the initial audio data, where the position information of the trajectory point is used to determine the virtual source position in the virtual playback scene. The virtual playback scene is deployed with virtual speakers, the target virtual speaker is at least one of the virtual speakers, the target virtual speaker represents the virtual speaker determined based on the virtual source position, and the speaker gain is used to simulate the playback of the audio to be processed at the virtual source position;

[0084] Optionally, in an embodiment of the present application, the target virtual speaker may include but is not limited to an actual audio output device selected based on the matching rules between the virtual source position and the physical speaker layout. The selection basis may involve the relative orientation of the speaker and the virtual source position, the distance attenuation model, and the acoustic characteristics of the environment. For example, the three physical speakers closest to each other are determined as the main output nodes through spatial vector calculation, and the participation weights of the secondary speakers are dynamically adjusted according to the room reverberation parameters. In a specific scenario, the target virtual speaker can be the left front surround speaker of a home theater system, the center speaker of a car audio system, or the left and right earmuff-type sound units of a VR headset.

[0085] For example, Figure 5 is a schematic diagram of another optional audio data processing method according to an embodiment of the present application, such as Figure 5 As shown, taking the 5.1.4 speaker sound bed layout as an example, it includes the best listening position 1, left and right speakers 2, center speaker 3, woofer 4, left and right rear surround speakers 5, left and right front overhead speakers 6, and left and right rear overhead speakers 7. Figure 6 This is a schematic diagram of another optional method for processing audio data according to an embodiment of the present application. The calculation of the bullet screen position information is as follows: Figure 6 As shown, the three closest speakers can be determined as the target virtual speakers through space vector calculation.

[0086] also, Figure 7 is a schematic diagram of another optional method for processing audio data according to an embodiment of the present application, such as Figure 7 As shown, given the azimuth and elevation angles of the input object in the spherical coordinate system, the spherical coordinates are converted into a three-dimensional unit vector p = [p1 p2 p3] T .

[0087]

[0088] The target virtual speaker gain is calculated as follows: find the three speakers closest to the virtual source: m , l k , l n Map the 5.1.4 channel layout to the spherical space and obtain three three-dimensional vectors pointing to lm, lk, ln, L1 = [l 11 ,l 12 ,l 13 ] points to l m , L2=[l 21 ,l 22 ,l 23 ] points to l k , L3=[l 31 ,l 32 ,l33 Pointing to l n , such that:

[0089] p = g1l1 + g2l2 + g3l3;

[0090] Here, g = [g1, g2, g3], representing the gains of three speakers. If the sound is close to l m , the gain of g1 will approach 1, and the same applies to other positions. And the energies of the final speakers p1, p2, and p3 are obtained.

[0091] Among them, the above speaker gains may include, but are not limited to, a set of adjustment coefficients for controlling the output power of the target virtual speaker. This coefficient set is calculated through a spatial audio algorithm and is used to simulate the attenuation, diffraction, and azimuth perception effects of sound waves propagating from the virtual source position to the listener. For example, generating a binaural gain difference based on the head-related transfer function (HRTF), calculating the volume attenuation ratio according to the inverse square law of distance, or enhancing the sound pressure level in a specific direction through a beamforming algorithm. The gain parameters can be stored as a floating-point number array, and each element corresponds to the amplification factor of the target virtual speaker channel, and the value range is usually between 0.0 (mute) and 1.0 (maximum power).

[0092] It should be noted that the calculation method of the speaker gain can be designed diversely according to the sound field model and device capabilities. For example, the distance attenuation algorithm can adopt a spherical wave model, a plane wave model, or a mixed attenuation curve; the azimuth perception processing can choose an amplitude translation method based on the number of channels, a binaural rendering method based on HRTF, or a sound field reconstruction method based on higher-order Ambisonics; the dynamic adjustment mechanism can integrate automatic level control (to avoid sudden volume changes), environmental noise compensation (such as increasing the gain according to the real-time noise collected by the microphone), or emotional response adaptation (such as enhancing the low-frequency shock feeling according to the user's heart rate data). This application does not make specific limitations on this.

[0093] Optionally, in the embodiments of the present application, the above virtual playback scenario may include, but is not limited to, a computable sound field environment defined by a three-dimensional coordinate system and acoustic parameters. This environment includes the spatial layout of virtual speakers, the absorption coefficient of the reflecting surface material, and the air attenuation model. For example, the length, width, and height dimensions of a rectangular room, the scattering coefficient of the wall sound-absorbing cotton, and the influence parameters of temperature and humidity on the sound speed. Its function is to map the position information of the trajectory points into a sound wave propagation path that conforms to physical laws and support dynamic update of scene parameters to adapt to different playback environments, such as automatically adjusting the reverberation time and the direct sound ratio when switching from a closed meeting room to an open-air square.

[0094] It should be noted that the construction logic of the virtual playback scenario can be adapted to the requirements of different application scenarios. For example, the coordinate system type can be selected as the Cartesian coordinate system (suitable for regular rooms), the polar coordinate system (suitable for rotationally symmetric scenarios), or the grid voxel space (supporting complex obstacle modeling); the environmental acoustic parameters can be preset as the concert hall mode (long reverberation time), the recording studio mode (short reverberation time), or the custom mode (the user uploads the reflection coefficient of the wall material); the virtual speaker layout strategy can adopt a fixed standard configuration, dynamic density adjustment (automatically optimized according to the number of physical speakers), or the sound field hot spot distribution generated based on machine learning. This application does not make specific limitations on this.

[0095] Optionally, in the embodiments of this application, the above virtual source position may include, but is not limited to, the equivalent sound source coordinates mapped in the virtual playback scenario through the trajectory point position information. This coordinate can be appended with multi-dimensional attributes, such as the movement speed (affecting the Doppler shift), the radiation pattern (omnidirectional or directional), and the sound source size (point source or plane source). Its calculation process may involve coordinate system conversion (such as converting the user-defined world coordinates into local coordinates centered on the listener's head) and occlusion detection (such as judging the blocking degree of the virtual wall on the sound ray), so as to generate accurate spatial audio rendering parameters.

[0096] It should be noted that the attribute definition of the virtual source position supports multi-dimensional expansion. For example, the movement trajectory can be associated with the parabolic ballistic simulated by the physics engine, the random Brownian motion, or the free curve drawn by the user's gesture; the sound source radiation characteristics can be defined as the cardioid directivity (enhancing front-end pickup), the super-cardioid directivity (suppressing side and rear noise), or the shotgun directivity (long-distance focusing); the additional metadata can include emotion tags (such as "urgent warning sound" triggering gain increase), content classification (such as "human voice" automatically enabling voice enhancement), or copyright marking (restricting the replay of unauthorized devices). This application does not make specific limitations on this.

[0097] Optionally, in the embodiments of this application, the above virtual speaker may include, but is not limited to, the audio rendering reference points predefined or dynamically generated in the virtual playback scenario. Its function is similar to the digital twin of the physical speaker, including parameters such as position coordinates, frequency response curves, and maximum output sound pressure levels. The mapping relationship between the virtual speaker and the actual physical device can be established through automatic detection or manual configuration (such as the user inputting the number of speakers and the azimuth angle).

[0098] It should be noted that the selection strategy of the target virtual speaker can be extended in combination with hardware characteristics and service requirements. For example, the priority rule can be set to prefer speakers with a wider high-frequency response, top speakers that support Dolby Atmos encoding, or devices marked as the "main listening area" by the user; the dynamic adjustment mechanism can automatically switch to a standby node when a speaker failure is detected, or recalculate the optimal speaker combination according to the change in the listener's position (by following with a camera); the fault tolerance strategy can include smooth gain transition (to avoid audio breakage during switching), redundant output of multiple devices (to ensure the stability of sound image positioning), or a degradation mode (switching to stereo rendering when there are insufficient available speakers). This application does not make specific limitations in this regard.

[0099] S206, perform mixing processing on the audio to be processed and the media information indicated by the source media information identifier based on the speaker gain to generate target audio data.

[0100] Optionally, in the embodiments of this application, the above mixing processing may include, but is not limited to, an audio rendering process of spatially fusing the audio to be processed with associated media information. This process performs multi-channel energy distribution on the original audio signal through the speaker gain parameter, and combines the time synchronization mark and content association identifier in the media information to achieve sound field reconstruction and content matching. For example, after performing gain weighting on each frequency band of the audio to be processed, it is phase-aligned and superimposed with the background music track corresponding to the media information, and at the same time, the delay parameters of each channel are adjusted according to the position of the virtual sound source to simulate the spatial propagation effect. The specific implementation may include processing links such as dynamic range compression (to avoid signal clipping), multi-track electrical average balance (to balance the primary and secondary audio ratios), and surround sound encoding (such as Dolby Digital encoding).

[0101] It should be noted that the specific algorithms and processes of the mixing processing can be flexibly adjusted according to the rendering target and resource constraints. For example, the spatialization algorithm can select physically based ray tracing sound field simulation, machine learning-based environmental timbre transfer, or a simplified amplitude translation method; the multi-track fusion strategy can adopt automatic avoidance of primary and secondary audio (reduce the background music level through side-chain compression), dynamic frequency band masking elimination (weaken frequency band conflicts), or emotion-driven reverb stacking (increase the environmental atmosphere layer according to semantic analysis); the processing stage can be divided into pre-processing (noise reduction / normalization), real-time processing (low-latency convolutional reverb), and post-processing (loudness matching / limiting). This application does not make specific limitations in this regard.

[0102] Exemplarily, Figure 8 is a schematic diagram of another optional method for processing audio data according to the embodiments of this application, as Figure 8As shown, after recording the barrage, click Send to send the generated barrage to the server backend for storage. Services can charge membership fees based on the barrage. The uploaded barrage will be distributed to all users on the playback end. Users can click the barrage-style button to render and play it.

[0103] Optionally, in an embodiment of the present application, the above-mentioned media information may include but is not limited to a multimodal content data set bound to a source media information identifier. The information may include a video frame sequence, subtitle text, an interactive event timestamp, or environmental sound effect parameters, such as a key frame hash value of a video clip played synchronously with a voice barrage, scene lighting change data, or a vibration feedback parameter triggered by the user. Its function is to provide a contextual association basis for mixing processing, such as automatically adjusting the audio fade-in and fade-out rate according to the rhythm of the video action, or adapting the reverberation intensity parameters according to the emotional analysis results of the subtitle text, thereby achieving an immersive experience enhanced by the synergy of sound and picture.

[0104] It should be noted that there are multiple implementation paths for the source and fusion of media information. For example, the content type can be expanded to three-dimensional model motion data (driving the sound to move with the object), tactile feedback waveforms (to achieve sound-vibration synchronization effects) or real-time game engine events (such as explosion coordinates triggering corresponding directional sound effects); additional dimensions of metadata can include copyright licensing information (restricting decoding of unauthorized devices), barrier-free auxiliary information (adding a voice annotation layer for visually impaired users) or multi-language audio track indexes (switching dubbing versions according to the system language); application scenarios can be adapted to film and television secondary creation (users add commentary audio tracks to synchronize the original film), virtual meetings (integrating multi-person voices with shared whiteboard operation sounds) or immersive education (superposition of experimental operation sound effects and explanation voices). This application does not make specific restrictions on this.

[0105] Optionally, in an embodiment of the present application, the above-mentioned target audio data may include but is not limited to a terminal device playable audio stream formed after multi-dimensional mixing processing. Its data structure usually contains multi-channel audio samples, metadata tags and device adaptation instructions, such as 5.1-channel PCM data, head-following auxiliary information for headphone virtualization, or frequency band compression parameters. The generation process may involve format conversion (such as upmixing from mono to surround sound), metadata embedding (such as dynamic loudness normalization tags) and device compatibility packaging (such as packaging as MPEG-DASH streaming media segments) to ensure that the preset spatial sound field effect is achieved on different hardware devices.

[0106] It should be noted that the output form of the target audio data and the optimization strategy can be extended according to business requirements. For example, the encoding format can be adapted to low-bitrate scenarios (using Opus encoding), high-fidelity requirements (using FLAC lossless compression), or special formats for spatial audio (such as Ambisonics B-format); the transmission protocol can be designed as real-time streaming (transmitted through WebRTC data channels), offline packages (distributed as encrypted ZIP files), or edge computing node caching (pre-loaded according to the user's location); the rendering method can support traditional speaker array driving, frequency response optimization of bone conduction devices, or ultrasonic directional beam synthesis (to achieve a localizable private auditory space). This application does not make specific limitations in this regard.

[0107] In an exemplary embodiment, Figure 9 is a schematic diagram of another optional method for processing audio data according to an embodiment of the present application, as Figure 9 shown. Assuming that taking the example of a user creating a voice barrage on an online video platform, it includes but is not limited to the following processes:

[0108] S902: Voice barrage recording and trajectory binding:

[0109] The user records the audio to be processed with a duration of 3 seconds through a mobile terminal device, with a sampling rate of 48 kHz and a bit depth of 16 bit, using the AAC encoding format. During the recording process, the user draws a dynamic path containing 3 trajectory points in the three-dimensional preview window of the video playback interface: the starting point coordinates are (1.2, 0.5, -0.8), the middle point is generated through the Bezier curve interpolation algorithm, and the ending point coordinates are (-0.3, 0.2, 1.5). The time interval between trajectory points is 500 ms. The system binds the audio data with the trajectory point information to generate an initial audio data packet, where the source media information identifier is the video segment hash value (such as the SHA-256 digest value 0x7d3f…a91b), the trajectory point position information is stored in a rectangular coordinate system, and an additional timestamp sequence [0 ms, 500 ms, 1000 ms] is attached.

[0110] Among them, the recording parameters can be set flexibly. For example, the sampling rate can be set to 44.1 kHz or 96 kHz, and the bit depth can be selected as 24 bit; the trajectory definition method can be set flexibly. For example, the user can choose to input manually in the spherical coordinate system (r = 2.0 m, θ = 30°, ) or draw in real time through a gyroscope-controlled virtual laser pen; the data encapsulation method can be set flexibly. For example, the initial audio data packet can include a CRC32 check code (such as 0x8e5d2c49) or be encoded using the TLV (Tag-Length-Value) structure.

[0111] S904: Virtual sound field mapping and gain calculation:

[0112] After the system loads the initial audio data, it parses the position information of the track points and constructs a rectangular room model with dimensions of 5m × 4m × 3m in the virtual playback scene, setting the wall sound absorption coefficient to 0.4. According to the virtual source position (e.g., the coordinates at t = 500ms are 0.8, 0.6, -0.2), the distance attenuation model is used to calculate the gain of the target virtual speaker:

[0113] Exemplarily, the distance calculation includes but is not limited to: the Euclidean distance d between the listener position (0, 0, 0) and the virtual source is d = √(0.8² + 0.6² + 0.2²) = 1.02m; the gain distribution includes but is not limited to: based on the 1 / d 2 law, the base gain is 0.96 (1 / 1.02²), the left front speaker obtains an additional gain coefficient of 0.15 due to the azimuth angle of 35°, and the right rear speaker reduces the gain coefficient by 0.2 due to the occlusion detection result. Finally, a gain matrix [0.82, 0.91, 0.45, 0.33, 0.28, 0.17] containing 6 physical speakers is generated.

[0114] Among them, the virtual scene parameters can be flexibly set. For example, it can be replaced with a dome-shaped scene (radius 4m) or import a CAD building model; the gain algorithm can be flexibly set. For example, when using HRTF binaural rendering, a pair of headphone stereo gains (left 0.7, right 0.3) are generated; the selection of the target virtual speaker can be flexibly set. For example, when it is detected that the user is using a soundbar, it is automatically switched to match the virtual Dolby 5.1.2 layout.

[0115] S906: Multi-source mixing and data generation:

[0116] Mix the audio to be processed with the background music (sampling rate 48kHz, average loudness -18dBFS) associated with the source media information identifier, which specifically includes but is not limited to the following sub-steps:

[0117] S3-1, Gain application: Apply the speaker gain matrix to the audio to be processed, with a 3dB boost for the left front channel;

[0118] S3-2, Timing alignment: Synchronize the start point of the voice bullet screen with the video frame at 12 minutes and 35 seconds according to the video timeline, with the error controlled within ±20ms;

[0119] S3-3, Spatialization processing: Add early reflections (delay 15ms, attenuation 40%) and late reverberation (reverberation time 1.2s) based on the room model;

[0120] S3-4, Dynamic compression: Use a compression ratio of 4:1 to limit the peak level to avoid the total output exceeding -1dBFS. Finally, generate the target audio data containing 7.1 channels, encapsulated in the MP4 container format, with an audio bitrate of 256kbps.

[0121] Among them, the mixing strategy can be flexibly set. For example, in the game live broadcast scenario, side-chain compression can be used to dynamically avoid the voice barrage and the game gun sound effects. The media information fusion can be flexibly set. For example, the associated barrage text can be used to generate an SRT subtitle track to achieve lip-sync. The output format can be flexibly set. For example, when adapting to bone conduction devices, it is converted into mono + vibration waveform data, and the frequency response range is adjusted to 200Hz - 3kHz.

[0122] In this embodiment, through the dynamic binding of trajectory points and virtual sound field mapping, the voice barrage has three-dimensional spatial positioning characteristics. By using physical modeling gain calculation and multi-source mixing strategy, it is possible to restore the sound image movement trajectory under different hardware configurations. For example, in a 5.1 surround sound system, the sound effect of continuous sliding from the front left to the back right can be achieved. Combining with the media information synchronization mechanism, it ensures that the time alignment accuracy between the voice barrage added by the user and the video content and background audio track reaches the frame level. This method significantly improves the accuracy of the user's spatial orientation perception, supports the adaptation of dynamic trajectory complexity and multi-scene acoustic characteristics, and at the same time avoids the distortion problem caused by multi-channel superposition through compression limiting technology.

[0123] In another exemplary embodiment, taking the teacher-student voice interaction in a virtual reality teaching platform as an example, it includes but is not limited to the following processes:

[0124] Teaching voice collection and knowledge point positioning:

[0125] The user records the question voice in the virtual classroom environment as the audio to be processed, and the audio format uses the standard sampling rate and lossless compression coding. The system synchronously obtains the trajectory point information associated with the knowledge point, and the trajectory point is automatically generated at the specific area coordinates of the virtual scene according to the teaching content. For example, literary questions correspond to the coordinates of the ancient poetry exhibition hall, and physical questions are mapped to the coordinates of the mechanics laboratory. When the initial audio data packet is encapsulated, the source media information identifier is bound to the digital fingerprint of the current teaching chapter, and the trajectory point position information includes a dynamic orientation mark and a knowledge point classification label.

[0126] Among them, the audio input method can be flexibly set. For example, it supports near-field voice collection, far-field array microphone noise reduction, or historical recording import. The trajectory generation logic can be flexibly set. For example, the sound source distribution density can be automatically adjusted according to the knowledge point popularity, or the coordinates can be dynamically optimized according to the student's gaze direction. The data identifier construction can be flexibly set. For example, the course tree coding, semantic feature vector, or blockchain deposit hash value is used as the source identifier.

[0127] Teaching scene sound field modeling and device adaptation:

[0128] Analyze the trajectory point information in the initial data and construct a multi-region acoustic model in the virtual teaching scenario. This model includes the reflection surface parameters of the lecture hall, the difference in sound absorption coefficients in different subject areas, and the student seat position data. After determining the virtual source position according to the knowledge point coordinates, the system combines the spatial layout of physical speakers and calculates the gain coefficients of each device through the acoustic wave diffraction algorithm. For example, for the sound source point located at the chemistry experiment table, the mid-high frequency speakers surrounding the student seats are preferentially activated, while the output weight of the low-frequency speakers in the podium area is suppressed.

[0129] Among them, the acoustic field modeling method can be flexibly set. For example, geometric acoustic simulation, machine learning to predict reverberation parameters, or import BIM building information model can be used; the gain allocation strategy can be flexibly set. For example, in the group discussion mode, the adjacent speaker groups are enhanced, and in the exam mode, the directional sound beam is enabled to suppress interference; the device adaptation mechanism can be flexibly set. For example, when a mobile terminal device is detected to be connected, it automatically switches to the head-related transfer function rendering mode.

[0130] Multi-modal content fusion and spatialized output:

[0131] Mix the voice to be processed with the associated teaching resources, including the operation sound effects of 3D courseware, the writing sound of electronic blackboard writing, and the background environmental sound. During the processing, the main voice stream is enhanced in frequency band and strengthened in azimuth according to the gain coefficient. For example, in the biological dissection demonstration scenario, the teacher's explanation voice is focused on the operating table area, while the organ dissection sound effects are dispersed to the corresponding dissection azimuths. The final generated target audio data integrates spatial metadata, supporting adaptive dynamic range control and multi-device synchronous calibration.

[0132] Among them, the content fusion dimension can be flexibly set. For example, associated experimental instrument vibration feedback data can be used to generate tactile synchronization signals, or the azimuth prompt light effects of AR glasses can be superimposed; the spatialization processing can be flexibly set. For example, in the language learning scenario, the typical environmental reverberation characteristics of the corresponding country are added to the foreign teacher's voice; the output optimization can be flexibly set. For example, for hearing-impaired students, dedicated bone conduction frequency band compensation is added, or a sound field equalization version is generated for the group learning mode.

[0133] In this embodiment, through knowledge-point-driven sound source localization and sound field modeling in teaching scenarios, spatialized cognitive guidance for educational content is achieved. The dynamic matching mechanism between the virtual source position and physical devices enables the teacher's voice to exhibit different spatial characteristics in different subject areas. For example, the voice for literary appreciation has a warm tone of wooden structure in the antique academy scene, while the knowledge explanation simulates the acoustic vacuum effect of a space capsule. The multi-modal mixing strategy effectively enhances the immersion of knowledge transfer and solves the problem of single-dimensional voice interaction in traditional distance education. The system supports intelligent switching of sound field preset templates according to teaching content to adapt to the multi-mode teaching requirements such as group discussion, experimental operation, and exam invigilation. At the same time, the voice salience of key knowledge is enhanced through azimuth enhancement technology, significantly reducing the cognitive load of learners.

[0134] By adopting the technology of dynamic binding of audio tracks and fusion of spatial metadata, the initial audio data is generated by performing spatio-temporal correlation encoding on the audio to be processed and multi-dimensional track point information, achieving the purpose of accurately anchoring the spatial attributes and content relevance of the sound source. By adopting the technology of dynamic mapping of virtual sound fields and adaptive matching with physical devices, an acoustic model of the virtual playback scene is constructed by analyzing the position information of track points, and the gain parameters are calculated in combination with the physical characteristics of the target virtual speaker, achieving the dynamic adaptation of the virtual sound source position and the physical playback device, thus realizing the general technical effect of sound field restoration in complex environments. This method breaks through the limitation of fixed speaker layouts on spatial audio through the intelligent matching of virtual source positions and physical devices.

[0135] Furthermore, by adopting the technology of spatialized mixing of multi-source data and dynamic gain control, through the frequency band allocation and timing alignment strategy based on speaker gain, the audio to be processed and associated media information are fused in multiple dimensions, achieving the purpose of synergistically enhancing the spatial characteristics of sound and picture, thus realizing the technical effect of immersive sound field reconstruction. This process precisely controls the energy distribution of each channel through spatial audio rendering algorithms. While retaining the characteristics of the original audio content, elements such as voice bulletins and environmental sound effects present azimuth movement trajectories that conform to physical laws. At the same time, dynamic level control effectively suppresses multi-channel superposition distortion, thus solving the technical problem of poor playback effects of audio data.

[0136] As an optional solution, Figure 10 is a schematic diagram of another optional method for processing audio data according to an embodiment of the present application. As Figure 10 shown, determining the speaker gain corresponding to the target virtual speaker based on the initial audio data includes:

[0137] S204-1, determining the virtual source position based on the position information;

[0138] S204-2. Determine a target virtual speaker and speaker gain according to the virtual source position and the speaker positions corresponding to the respective virtual speakers in the virtual playback scenario, where the target virtual speaker is the speaker in the virtual playback scenario whose distance from the virtual source position satisfies a preset distance condition.

[0139] Optionally, in the embodiments of the present application, the above position information may include, but is not limited to, a set of absolute or relative positioning parameters describing the virtual source in a spatial coordinate system. This information may include numerical expressions of three-dimensional Cartesian coordinates, spherical coordinates, or polar coordinates, along with additional motion direction vectors, velocity vectors, or acceleration parameters. For example, in an indoor positioning scenario, the position information can be converted into coordinate values with the center of the room as the origin based on UWB ultra-wideband ranging data, or 6DoF pose data generated using the SLAM algorithm in an AR scenario. Its storage format may include binary floating-point arrays, JSON structured data, or byte streams of a custom protocol, and supports encrypted transmission and dynamic updates.

[0140] It should be noted that the expression form and acquisition method of the position information can be extended according to the application scenario. For example, the coordinate system reference can be set to the center of the listener's head (applicable to VR headsets), the origin of the room building structure (applicable to smart homes), or the earth's longitude and latitude (applicable to outdoor sound reinforcement systems); the dynamic following method can use an inertial measurement unit (IMU) for real-time positioning, camera vision recognition of marker points, or indoor positioning technology based on Wi-Fi RTT; data accuracy control can include noise reduction filtering processing (such as Kalman filtering), multi-sensor data fusion, or adaptive error compensation algorithms. The present application does not make specific limitations in this regard.

[0141] Optionally, in the embodiments of the present application, the above speaker position may include, but is not limited to, the spatial coordinate definition of a physical or virtual audio output device in a sound field model. This coordinate may include device installation height, horizontal deflection angle, and distance parameters from the listener reference point. For example, the left front speaker coordinates (1.5m, 0m, 2.1m) in a home theater system, and the elevation angle of 60 degrees of the top reflection speaker. Its measurement methods may include laser rangefinder calibration, an automatic room correction system (such as acoustic impulse response analysis), or user manual input configuration, and supports real-time update of coordinate data according to the device movement state.

[0142] It should be noted that the determination strategy of the speaker position can be adapted to diverse device types and environmental requirements. For example, the automatic detection of the physical speaker layout can read the device topology through the HDMI CEC protocol, match the acoustic fingerprints (such as sending sweep signals to analyze the response characteristics), or combine Bluetooth beacon positioning; the deployment of virtual speakers can be based on standard surround sound configuration templates (such as 5.1.4 Dolby Atmos), optimized layouts generated by machine learning, or user-defined sound field hotspots; the coordinate update mechanism can be designed for regular automatic calibration, remeasurement triggered by device displacement, or dynamic synchronization with the Building Information Model (BIM). This application does not make specific limitations in this regard.

[0143] Optionally, in the embodiments of this application, the above preset distance conditions may include, but are not limited to, a set of spatial correlation determination rules for screening target virtual speakers. This condition can be set based on the Euclidean distance threshold, spherical coverage area, or sound wave propagation path loss. For example, all speakers within 3 meters of the virtual source position are selected, or a group of speakers whose direct sound wave paths are not blocked by obstacles are selected. Its judgment logic can include the K-Nearest Neighbor algorithm (KNN), regional block spatial indexing (such as grid-based spatial hashing), or weighted screening based on the acoustic energy attenuation model, and supports dynamic adjustment of condition parameters to adapt to different sound field environments.

[0144] It should be noted that there are multiple possibilities for the setting dimensions and calculation methods of the preset distance conditions. For example, the distance metric can choose the Manhattan distance to simplify the operation complexity, consider the temperature compensation of the sound speed to correct the distance, or introduce the equivalent distance considering the reverberation energy ratio; the condition combination can include multi-level screening (first select speakers within 5 meters and then filter by azimuth angle), weighted selection based on priority (such as the center speaker has a higher weight), or dynamic conditions (automatically switch the theater mode / music mode threshold according to the current playback content type); the exception handling mechanism can include an alternative speaker queue, a downgraded rendering strategy (switch to stereo output when no device meets the conditions), or pre-computation optimization based on the prediction of the listener's position. This application does not make specific limitations in this regard.

[0145] In an exemplary embodiment, taking the spatial rendering of voice bullet screens in a video interaction platform as an example, it includes, but is not limited to, the following processes:

[0146] Virtual source position parsing and scene mapping:

[0147] The system reads the trajectory point position information in the initial audio data and maps it to the virtual playback scene based on the set coordinate system. The acoustic parameters of the virtual playback scene include the sound absorption coefficient, sound speed parameter, and spatial dimension definition. A continuous virtual source position sequence is generated through interpolation algorithms to ensure the smoothness of the movement trajectory. For example, the trajectory points can be mapped to a rectangular or dome-shaped virtual scene, the interpolation process supports linear or non-linear path generation, and the scene acoustic model can adapt to the reverberation characteristics of closed or open environments.

[0148] Among them, the coordinate system conversion can be flexibly set. For example, it supports the rectangular coordinate system, the spherical coordinate system, or other custom spatial encoding methods; the interpolation strategy can be flexibly set. For example, a high-order interpolation algorithm is used to generate a curved path, or a physical engine is combined to simulate a dynamic trajectory; the scene configuration can be flexibly set. For example, a standard acoustic template (such as a cinema, a concert hall) is selected according to the application requirements, or a custom 3D model is imported.

[0149] Target virtual speaker screening and gain calculation:

[0150] Traverse the virtual speakers deployed in the virtual playback scene and calculate the spatial relationship between each speaker and the virtual source position. Based on the preset distance condition, screen the target virtual speakers. For example, preferentially select the adjacent speakers or the devices within a specific azimuth range. The gain calculation process integrates the distance attenuation model and the azimuth perception algorithm. For example, a basic gain is assigned to the speakers that meet the conditions, and the weight coefficients are adjusted according to factors such as azimuth deviation and obstacle occlusion. Finally, a gain matrix containing multiple channels is generated to guide the audio rendering of physical devices.

[0151] Among them, the screening conditions can be flexibly set. For example, they can be dynamically adjusted based on the integrity of the direct sound path, the frequency response characteristics of the device, or the ambient noise level; the attenuation model can be flexibly set. For example, a spherical wave, a plane wave, or a mixed propagation model is used to calculate the energy attenuation; the weight adjustment can be flexibly set. For example, an acoustic field uniformity optimization algorithm is introduced or the gain correction based on the auditory perception model is used.

[0152] Multi-device gain allocation and real-time rendering:

[0153] Adapt the gain matrix to the physical speaker system and perform multi-dimensional audio processing. Perform frequency division processing on the audio to be processed, and allocate specific frequency bands to different speaker groups. For example, the center channel enhances the voice frequency band, and the surround channels focus on the environmental sound effects. Synchronously superimpose dynamic effects, including distance-based delay compensation, azimuth-related frequency response correction, and environmental reverberation rendering. A smooth transition mechanism is adopted during the rendering process to ensure the continuity of the sound image movement and support the automatic switching of the encoding strategy according to the device type.

[0154] Among them, the frequency division strategy can be flexibly set. For example, the frequency band range is dynamically divided, or the frequency response curve is preset according to the content type (voice / music); the effect superposition can be flexibly set. For example, the simulation of early reflected sound, the high-frequency attenuation of air absorption, or the dynamic range compression is integrated; the device adaptation can be flexibly set. For example, different rendering parameters are generated for headphone virtualization, soundbar, or panoramic sound system.

[0155] Through the embodiments of the present application, by using the dynamic matching mechanism between the virtual source position and the physical device, accurate sound image positioning of the voice barrage in the three-dimensional space is achieved. Based on the gain allocation strategy calculated according to the spatial relationship, the energy utilization rate of the multi-channel system is effectively improved, and the signal interference of the ineffective channels is reduced. The combination of the frequency division processing and the effect superimposing technology enhances the spatial atmosphere while retaining the voice clarity. The adaptive rendering scheme is compatible with a variety of playback devices, supports seamless switching from a simple stereo system to a panoramic sound system, and ensures a coherent sound field movement trajectory can be presented under different hardware environments. The dynamic transition mechanism eliminates the auditory discontinuity during the azimuth switching, making the movement trajectory of the voice barrage smooth and natural, and significantly enhancing the user's spatial immersion and content perception dimension in the interactive scenario.

[0156] As an optional solution, as Figure 10 shown, determining the virtual source position based on the position information includes:

[0157] S204-1-1, converting the position information into the spatial trajectory point coordinates in the spherical coordinate system, where the position information represents the trajectory point coordinates in the rectangular coordinate system;

[0158] S204-1-2, determining the target position vector based on the spatial trajectory point coordinates, and determining the position indicated by the target position vector as the virtual source position, where the target position vector is used to indicate the spatial position of the corresponding trajectory point in the virtual playback scenario.

[0159] Optionally, in the embodiments of the present application, the above position information may include, but is not limited to, a set of three-dimensional space positioning parameters describing the trajectory point in the rectangular coordinate system. This information usually includes the coordinate values of the x, y, and z axes, and is used to accurately represent the absolute or relative position of an object in the Cartesian space. For example, in a three-dimensional modeling scenario, the position information can be reflected as the coordinate value of the center point of an object (such as the coordinates of a certain trajectory point are x = 1.5, y = 0.8, z = -0.3), or a sequence of movement path coordinates composed of multiple discrete points. Its data form can include a floating-point array, a binary encoded stream, or a structured data field, and supports encrypted storage and dynamic update. The application scenarios of the position information cover virtual reality space positioning, robot motion trajectory planning, or the definition of the sound source coordinates in the sound field modeling, and its precision control can be achieved through coordinate normalization processing or error compensation algorithms.

[0160] It should be noted that the coordinate system conversion method can be flexibly selected according to the application scenario requirements. For example, the conversion algorithm can select the standard spherical coordinate system formula, the custom cylindrical coordinate system hybrid model, or the local coordinate system based on geographic information; the dynamic conversion strategy can adopt the incremental update algorithm in the real-time rendering scenario, or batch convert the historical trajectory data during offline processing; the error processing mechanism can include coordinate rounding rules (such as retaining two decimal places), outlier filtering (such as removing coordinates beyond the scene boundary), or using a sliding window for smoothing processing. The present application does not make specific limitations on this.

[0161] Optionally, in the embodiments of the present application, the above target position vector may include, but is not limited to, a mathematical representation quantity generated by coordinate system conversion for describing the azimuth of the virtual source space. This vector includes parameters such as modulus length, azimuth angle, and elevation angle, and is used to accurately describe the distance and direction of the sound source relative to the reference point in the spherical coordinate system. For example, when converting the rectangular coordinates (2.0, 1.5, 0.8) to spherical coordinates, the target position vector can be expressed as radius r = 2.7, azimuth angle θ = 36.8 degrees, and elevation angle φ = 16.3 degrees. Its calculation process may involve trigonometric function conversion, quaternion rotation operation, or a non-linear mapping algorithm based on a machine learning model. The role of the target position vector is to adapt to the input requirements of different spatial audio rendering algorithms (such as Ambisonics) and provide a unified spatial reference benchmark for multi-device collaboration.

[0162] It should be noted that the mapping strategy of the virtual source position can be adapted to different rendering engines and hardware devices. For example, simplified mapping rules (such as ignoring the height axis) are adopted in the lightweight scenario of mobile devices, and the superposition of reflection coefficients of multiple layers of materials is supported in professional acoustic simulation systems; the multi-source fusion mechanism can include priority weight assignment (such as retaining complete vector parameters for the main sound source), motion trajectory interpolation optimization, or parameter compression based on a perception model; the exception handling can be designed as default position fallback, adjacent device mapping, or triggering a user interaction calibration process. The present application does not make specific limitations on this.

[0163] Exemplarily, Figure 11 is a schematic diagram of another optional audio data processing method according to the embodiments of the present application. As Figure 11 shown, it may include, but is not limited to, converting the position information into spatial trajectory point coordinates in the spherical coordinate system in the following manner: The vector pointing from the origin to the sound source becomes the direction vector of the sound source. θ and φ represent the elevation angle and azimuth angle respectively. θ ∈ [0, π] is defined as the angle between the direction vector and the positive direction of the z-axis; φ ∈ [0, 2π] is the angle between the direction vector and the positive direction of the x-axis, and r represents the radial distance from the origin of the coordinates. The relationship between spherical coordinates (r, θ, φ) and rectangular coordinates (x, y, z) is:

[0164]

[0165] Among them, the bullet screen position information includes the pitch angle, azimuth angle and radius of the direction vector of the sound, the start time, and the bullet screen duration. Each voice bullet screen includes the content of the voice bullet screen (audio adts header + aac-encoded voice raw data), multiple bullet screen position information (for forming a trajectory), the bullet screen duration, the time when the voice bullet screen appears (for the appearance of the voice bullet screen), and the video id matched by the bullet screen (for matching the corresponding video source).

[0166] As an optional solution, as Figure 10 shown, determining the target virtual speaker and the speaker gain according to the virtual source position and the speaker positions corresponding to the respective virtual speakers in the virtual playback scene includes:

[0167] S204-2-1, obtaining the distances between the respective virtual speakers in the virtual playback scene and the virtual source position;

[0168] S204-2-2, determining the speakers whose distances meet the preset distance condition as the target virtual speakers;

[0169] S204-2-3, determining the speaker position vector corresponding to the target virtual speaker;

[0170] S204-2-4, determining the gain coefficient based on the target position vector and the speaker position vector, wherein the speaker position vector is the same as the target position vector after being weighted and summed according to the gain coefficient, and the speaker gain includes the gain coefficient.

[0171] Optionally, in the embodiments of the present application, the above preset distance condition may include, but is not limited to, a set of dynamic determination rules for screening the spatial correlation between the virtual speaker and the virtual source position. This condition can be set based on the absolute distance threshold, relative path loss or acoustic perception sensitivity. For example, select the top three speakers with the shortest straight-line distance from the virtual source position, or devices with no more than two reflections in the sound wave propagation path. Its calculation logic may include a geometric space fast retrieval algorithm (such as KD tree nearest neighbor search), weight screening based on the sound field energy distribution (such as the direct sound ratio exceeding the set ratio), or equivalent distance correction combined with environmental obstacle detection (such as sound wave diffraction path length compensation). The role of the preset distance condition is to balance the rendering accuracy and the computational complexity, and avoid resource waste caused by all-scene speakers participating in the calculation.

[0172] It should be noted that the distance calculation method can be extended according to the characteristics of the scenario and computing resources. For example, the distance metric can choose Euclidean distance to simplify the operation, an equivalent distance considering the sound wave propagation time (combining temperature and humidity to correct the sound speed), or a weighted distance introducing the diffraction path of obstacles; the computing optimization strategy can adopt spatial grid pre-computation (reducing the real-time computing amount), GPU parallel acceleration (processing large-scale speaker arrays), or incremental update (only calculating the speakers whose position change exceeds the threshold); the exception handling mechanism can include invalid coordinate filtering (automatically removing speakers outside the scene boundary), normalization of the distance calculation result (avoiding numerical overflow), or dynamic tolerance adjustment (adapting and relaxing conditions according to the environmental noise level). The present application does not make specific limitations in this regard.

[0173] In addition, the target virtual speaker screening strategy can be adapted to different rendering requirements. For example, the screening dimension can be extended to multi-condition combination (simultaneously meeting the distance threshold and frequency response range), priority grading (permanently retaining the lowest gain for the main speaker), or grouped processing (regarding adjacent speakers as a cluster and uniformly allocating gains); the dynamic adjustment mechanism can automatically enable standby devices when a speaker fails, re-screen devices according to the change of the listener's position, or switch the screening rules based on the content type (such as voice / music); the visualization assistance can include the display of a three-dimensional heat map of the screening results, a real-time detection panel of the device status, or a retrospective analysis function of historical screening data. The present application does not make specific limitations in this regard.

[0174] Moreover, the above gain coefficient can include but is not limited to a set of weight parameters for adjusting the output energy distribution of the target virtual speaker. This coefficient is calculated through a spatial audio algorithm and is used to precisely control the contribution degree of each speaker to the sound field reconstruction of the virtual source position. For example, in the beamforming algorithm, the gain coefficient can include amplitude weight and phase delay parameters, so that the collaborative output of multiple speakers forms a directional sound beam. Its calculation basis can include geometric acoustic models (such as inverse square law attenuation of distance), psychoacoustic perception models (such as binaural time difference enhanced localization), or machine learning-based data-driven optimization (such as a gain mapping table trained from historical listening data).

[0175] It should be noted that there are multiple possible implementations for the methods of vector operation and gain calculation. For example, the vector alignment algorithm can choose the least squares method to solve the overdetermined equation, an approximate solution based on geometric projection, or introduce a regularization term to prevent overfitting; the gain optimization goal can be set as uniform sound pressure level distribution, minimum positioning error, or power consumption balance; the computing framework can be adapted to real-time stream processing (low-latency requirements), offline batch processing (high-precision requirements), or edge-cloud collaborative computing (dynamic resource scheduling). The present application does not make specific limitations in this regard.

[0176] As an optional solution, such as Figure 10As shown, obtain the audio to be processed and at least one trajectory point, and generate initial audio data based on the audio to be processed and the trajectory point, including:

[0177] S202-1, obtain the audio to be processed and mark the audio timestamp;

[0178] S202-2, receive the trajectory points input by the user, where each trajectory point includes spatial position parameters and associated time information;

[0179] S202-3, bind the audio to be processed with the spatial position parameters based on the audio timestamp and the associated time information to generate initial audio data.

[0180] Optionally, in the embodiments of the present application, the above spatial position parameters may include, but are not limited to, a set of numerical values used to describe the positioning and motion state of the trajectory point in three-dimensional space. The parameter may include the x, y, and z axis coordinate values in a rectangular coordinate system, or the combination of radius, azimuth angle, and pitch angle in a polar coordinate system. For example, in an augmented reality scenario, the spatial position parameter may be embodied as the offset of a virtual object relative to the user's viewing point (such as x = 1.2m, y = 0.5m, z = -0.3m), or in the UAV flight path planning, it may be represented as a triple of GPS longitude, latitude, and altitude. Its data form may support floating-point arrays, binary encoded streams, or strings with unit identifiers (such as "2.3m@30°"), and can be collected by sensors, manually input, or generated by algorithms.

[0181] Optionally, in the embodiments of the present application, the above associated time information may include, but is not limited to, marker data used to bind the trajectory point and the audio timing relationship. The information may be embodied as an absolute timestamp (such as the UTC format "2024-05-01T08:30:00.000Z"), a relative time offset (such as 500ms after the start point of the audio), or a normalized time ratio (such as the 25% position of the total audio duration). For example, in an animation dubbing scenario, the associated time information may align a certain trajectory point with the pronunciation start moment of the character's line "start running", or in a sports event live broadcast, it may associate the commentary voice with a specific slow-motion replay segment. Its role is to ensure the precise synchronization of the spatial trajectory and the audio content in the time dimension.

[0182] Optionally, in the embodiments of the present application, the above audio timestamp may include, but is not limited to, a timing tag for identifying key events or segmentation positions in the audio to be processed. This tag can be generated based on audio waveform features (such as zero-crossing rate mutation points), semantic segmentation results (such as sentence intervals), or external input events (such as marker points triggered by users). For example, in the processing of conference recordings, the audio timestamp can mark the start and end times of the speeches of different speakers; in post-production of films and television, it can mark the synchronization points of explosion sound effects and corresponding pictures. Its storage form can include the sample point serial number (such as the 44,100th sample point), time code (such as 00:01:23.450), or a hash value index generated based on the audio fingerprint.

[0183] In an exemplary embodiment, Figure 12 is a schematic diagram of another optional method for processing audio data according to the embodiments of the present application. As Figure 12 shown, taking the creation and release of voice bullet screens in a video social platform as an example, it includes, but is not limited to, the following processes:

[0184] S1202: Voice bullet screen recording and timestamp marking:

[0185] The user records a 4-second audio to be processed through a mobile terminal device, with the sampling rate set to 48 kHz and compressed using the Opus encoding format. The system automatically inserts a timestamp sequence into the audio stream. For example, the starting point of the audio is marked as 00:00:00.000, and incremental marks (00:00:00.500, 00:00:01.000) are inserted every 500 ms. At the same time, the audio fingerprint generation algorithm is triggered to extract a 128-dimensional feature vector of this section of speech as a unique identifier.

[0186] Among them, the recording parameters can be set flexibly, including but not limited to supporting 16-bit / 24-bit bit depth selection, and the sampling rate can be set to 44.1 kHz or 96 kHz; the timestamp density can be set flexibly, including but not limited to at fixed intervals (such as every 100 ms), voice activity detection (insert only when there is human voice), or key frame synchronization (aligned with the video I-frame); the identifier generation can be set flexibly, including but not limited to generating a unique identifier using an MD5 hash value, Mel spectrogram features, or a voiceprint recognition model.

[0187] S1204: Trajectory point input and spatio-temporal parameter binding:

[0188] The user draws a dynamic trajectory in the three-dimensional coordinate system of the video preview interface and inputs three trajectory points:

[0189] Starting point: Spatial coordinates (1.2, 0.5, -0.8), associated time information 00:00:00.200;

[0190] Midpoint: Coordinates (0.6, 0.3, 0.2), associated time 00:00:01.800;

[0191] Endpoint: Coordinates (-0.4, 0.1, 1.2), associated time 00:00:03.400.

[0192] The trajectory point data is stored in a JSON structure, including Cartesian coordinates, associated timestamps, and motion interpolation types (such as Bezier curves).

[0193] Among them, the input method can be flexibly set, including but not limited to supporting gesture drawing of parabolas, voice commands (azimuth words parsed into coordinates), or importing preset trajectory templates; the coordinate system can be flexibly set, including but not limited to switching to a spherical coordinate system (radius 2.0m, azimuth angle 30°, elevation angle 45°) or a relative coordinate system (centered on the video subject); the data encapsulation can be flexibly set, including but not limited to using Protobuf binary encoding, adding a CRC32 checksum (such as 0x8A5D3E), or encrypting the signature to ensure data integrity.

[0194] S1206: Spatiotemporal data fusion and initial encapsulation:

[0195] Perform spatiotemporal alignment of the audio to be processed and the trajectory point data, including but not limited to:

[0196] Establish a timeline mapping table to perform offset compensation between the audio timestamp 00:00:01.000 and the associated time of the trajectory point 00:00:01.800;

[0197] For intermediate periods without explicit association (such as 00:00:00.500 - 00:00:01.800), use cubic spline interpolation to generate a continuous trajectory;

[0198] Encapsulate to generate an initial audio data packet, including the audio stream, trajectory parameter set, and metadata (such as coordinate system type WGS84, encoding format identifier OPUS_V1.3).

[0199] Among them, the alignment strategy can be flexibly set, including but not limited to supporting forward offset (the bullet screen lags behind the original sound), reverse offset (bullet screen preloading), or dynamic adaptive adjustment; the interpolation algorithm can be flexibly set, including but not limited to choosing linear interpolation, Catmull-Rom spline curves, or physical engine simulation of motion inertia; the metadata extension can be flexibly set, including but not limited to adding a copyright watermark, spatial audio rendering level (such as LOD2 level of detail), or device capability identifier (supporting Dolby Atmos).

[0200] Through the embodiments of the present application, by adopting the technologies of timing alignment and dynamic binding of spatial parameters, the precise mapping between audio content and spatial trajectories is achieved. The synchronization mechanism based on timestamps solves the problem of asynchronous audio-visual displacement in traditional solutions, for example, ensuring the real-time consistency of the azimuth between voice prompts and visual targets in virtual reality scenarios. The compatibility design of multi-dimensional trajectory point parameters supports diverse requirements from simple static positioning to complex motion paths, significantly enhancing the flexibility of spatial audio creation. The structured encapsulation scheme for initial audio data provides standardized input for subsequent sound field rendering, content retrieval, and cross-platform interaction, reducing the resource consumption and development complexity of multi-system collaboration.

[0201] As an alternative solution, as Figure 10 shown, obtain the audio to be processed and mark the audio timestamp, including at least one of the following:

[0202] S202-1-1, when the media information indicated by the source media information identifier includes a recorded video, mark the audio timestamp based on the playback progress of the recorded video, where the audio timestamp indicates that the audio to be processed is set to be played at the corresponding playback progress;

[0203] S202-1-2, when the media information indicated by the source media information identifier includes a live video, mark the audio timestamp based on the system time, where the audio timestamp indicates that the audio to be processed is set to be played at the corresponding system time.

[0204] Optionally, in the embodiments of the present application, the above source media information identifier may include, but is not limited to, index information for associating audio with target media content. This identifier can be embodied as a video file hash value, a unique live channel code, or a user-defined tag, such as a frame number sequence of a recorded video (e.g., the 12th minute and 35th second segment) or a combination of a live stream room number and timestamp (e.g., RoomID_20240501143000). Its function is to establish the playback timing mapping relationship between audio and media content, ensuring that the audio is triggered to play at a specified time point or media progress position.

[0205] Optionally, in the embodiments of the present application, the above recorded video may include, but is not limited to, media files that have been recorded and stored, and its playback progress advances linearly based on a fixed time axis. For example, movie and TV drama files, educational courseware, or pre-recorded product demonstration videos. The playback progress of such videos is usually marked based on time codes (such as HH:MM:SS.ms) or key frame numbers (such as the 1500th frame), and the audio timestamp can be associated with specific plot nodes (such as the starting moment of an explosion scene in a movie) or teaching knowledge point explanation segments.

[0206] Optionally, in the embodiments of the present application, the above playback progress may include, but is not limited to, timing parameters describing the current playback position of the recorded video. This parameter can be based on the internal media timeline (such as the number of milliseconds since the start of the video) or external event markers (such as bookmarks manually added by the user). For example, in the educational video scenario, the playback progress can be associated with the start time of the chapter in the course catalog; in the film and television scenario, it can be associated with the absolute time position where the end credits appear.

[0207] It should be noted that the association strategy of the playback progress can adapt to different media characteristics and interaction modes. For example, for multi-view recorded videos, the playback progress can be associated with the main view timeline or the local timeline of an independent view; for interactive educational videos, the playback progress can be bound to the nodes of the knowledge point tree structure to support non-linear jump associations; for film and television content with branching storylines, the playback progress can be extended to a multi-dimensional timeline (such as parallel time markers for different storyline branches). The present application does not make specific limitations on this.

[0208] Optionally, in the embodiments of the present application, the above live video may include, but is not limited to, real-time streaming media content, and its playback progress is dynamically synchronized with the system clock. For example, live sports events, online meetings, or real-time game commentaries. The playback progress of such videos is usually based on the NTP synchronized time or the streaming media shard sequence number, and the audio timestamp needs to be strictly aligned with the real-time data stream to ensure audio-visual synchronization at the viewer's end.

[0209] Optionally, in the embodiments of the present application, the above system time may include, but is not limited to, the real-time clock information maintained by the device or server. This time is usually based on Coordinated Universal Time (UTC) or local time zone time. For example, the UTC timestamp recorded when the host side collects audio in the live scenario (such as 2024-05-01T14:30:00.000Z). Its function is to provide a globally unified timing reference for the audio in the live stream and solve the clock drift problem during playback on multiple terminal devices.

[0210] It should be noted that there are various implementation methods for the system time synchronization mechanism in the live scenario. For example, the clock source can be selected as a GPS atomic clock, an NTP server, or a distributed blockchain timestamp; the time calibration strategy can include forward error correction (predicting clock deviation), backward compensation (dynamically adjusting the buffer), or master-slave synchronization based on heartbeat packets; the exception handling can be designed as a clock drift alarm, local silence filling, or degradation to relative time synchronization. The present application does not make specific limitations on this.

[0211] Through the embodiments of the present application, a differential timestamp marking strategy is adopted to achieve high-precision synchronization of audio and media content in the scenarios of video recording and live streaming. For the fixed time axis characteristic of the recorded video, an audio playback progress association mechanism is used to ensure that the audio is accurately triggered at the preset plot nodes; for the real-time requirement of the live video, the global synchronization of the system time is relied on to solve the clock difference problem of multi-terminal devices. This solution supports flexible timestamp marking granularity and association dimensions, can adapt to diverse scenarios such as film and television creation, online education, and real-time interaction, and significantly improves the reliability of audio-visual synchronization. At the same time, through the extensible design of the input source and the marking strategy, the technical threshold for processing multi-type media content is reduced, providing a unified timing benchmark for cross-platform audio-visual collaboration.

[0212] As an alternative solution, as Figure 10 shown, the audio to be processed is bound to the spatial position parameter based on the audio timestamp and the associated time information to generate the initial audio data, including at least one of the following:

[0213] When the media information indicated by the source media information identifier includes a recorded video, the associated time information is bound to the first playback progress of the recorded video, where the first playback progress indicates that when the recorded video is played to the first playback progress, the audio to be processed is played according to the speaker gain set at the corresponding track point;

[0214] When the media information indicated by the source media information identifier includes a live video, the associated time information is bound to the second playback progress of the audio to be processed, where the second playback progress indicates that when the audio to be processed is played to the second playback progress, the audio to be processed is played according to the speaker gain set at the corresponding track point.

[0215] Optionally, in the embodiments of the present application, the above first playback progress may include, but is not limited to, the timing marking parameters used to trigger audio playback in the recorded video. This parameter is defined based on the internal time axis of the recorded media, such as the absolute time code (e.g., 00:12:35.200) based on the starting point of the video, the key frame number (e.g., the 1500th frame), or the chapter node (e.g., the starting point of the third chapter). Its function is to establish an accurate association between audio playback and video images. For example, when the video is played to the moment of 00:05:23.500 of the protagonist's door-pushing action, the associated audio is synchronously triggered to play with the preset gain parameter. The first playback progress can be generated through video metadata parsing, user manual marking, or automatic detection based on content analysis (such as scene change recognition).

[0216] Optionally, in the embodiments of the present application, the above-mentioned second playback progress may include, but is not limited to, independent timing control parameters for audio playback in a live video scenario. This parameter is based on the starting point of the audio stream and defines the time offset of the audio itself (such as 500 ms after the start of playback), or a global trigger point based on the system time (such as UTC time 2024-05-01T14:30:00.000Z). For example, in a real-time commentary scenario, the second playback progress can be set to activate the surround sound enhancement effect 1.2 seconds after the start of the audio. Its generation method can include dynamic buffer monitoring (such as live stream delay compensation), user interaction triggering (such as the moment of sending a bullet screen), or external event synchronization (such as the sensor signal at the moment of a goal in a sports event).

[0217] It should be noted that the playback progress association method can adapt to different playback modes and interaction scenarios. For example, the dynamic adjustment mechanism can support real-time update of the playback progress (such as the progress jump caused by network fluctuations in a live stream), preloading based on a prediction algorithm (such as calculating the gain parameter for the next time period 100 ms in advance), or multi-dimensional progress synchronization (such as the linkage of video progress, audio progress, and user operation events); the cross-platform compatibility design can cover local file playback (such as Blu-ray discs), real-time streaming media (such as the low-latency HLS protocol), and mixed content (such as embedding live bullet screens in a recorded video); user interaction extensions can include bookmark marking (the user-defined gain parameter for a highlight moment), playback speed adaptation (compressing the gain change curve proportionally when fast-forwarding), or branch plot selection (different storylines corresponding to different sound field configurations). The present application does not make specific limitations in this regard.

[0218] Optionally, in the embodiments of the present application, the above-mentioned speaker gain may include, but is not limited to, an audio output adjustment coefficient dynamically calculated based on spatial position parameters. This coefficient is used to control the sound pressure level and frequency response characteristics of the target virtual speaker to simulate the perception of the azimuth and distance of the sound source in a virtual scene. For example, in a distance attenuation model, the speaker gain 2 meters away from the virtual sound source is set to 0.8, while the speaker gain 5 meters away is reduced to 0.3; in an azimuth perception model, the speaker located 30 degrees to the left front of the listener can be additionally increased by a gain of 0.15 to enhance the sense of localization. The calculation of the gain parameter can be combined with physical acoustic models, psychoacoustic weights, or machine learning prediction results.

[0219] It should be noted that the configuration dimension of the speaker gain can be extended according to the sound field characteristics and device capabilities. For example, the gain adjustment model can select a physically based spherical wave attenuation formula, a psychoacoustic azimuth weight table (HRTF data), or frequency band-specific coefficients generated by deep learning; the device adaptation strategy can optimize the binaural gain difference for the headphone virtualization scenario, optimize the multi-channel energy distribution for the home theater system, or optimize the directivity parameters for the outdoor sound reinforcement system; the dynamic control mechanism can integrate automatic level control (to prevent sudden volume changes), environmental noise compensation (dynamically increasing the gain according to the noise collected by the microphone), or emotional response adaptation (enhancing the low-frequency shock feeling according to the user's heart rate data). This application does not make specific limitations in this regard.

[0220] Through the embodiments of this application, by adopting a differential playback progress binding mechanism and a dynamic gain configuration strategy, the precise coordination of audio-visual content and the spatial sound field is achieved. For the fixed time axis characteristic of the recorded video, through the first playback progress association, the azimuth and volume expressiveness of the audio at the preset plot nodes are ensured; for the real-time requirement of the live video, relying on the second playback progress independent control mechanism, the sound field parameters are guaranteed to be updated dynamically with the live broadcast. The compatibility design of the multi-dimensional binding strategy supports full-scenario coverage from film and television production to real-time interaction, significantly improving the creative freedom and playback stability of spatial audio. The adaptive calculation of the gain parameters and the device collaborative optimization effectively enhance the sound image positioning accuracy and environmental immersion, providing users with a more layered and spatially realistic auditory experience.

[0221] As an alternative solution, as Figure 10 shown, obtaining the audio to be processed and at least one track point, and generating initial audio data based on the audio to be processed and the track point, includes:

[0222] S202-4, in response to the first interaction operation performed on the application interface, start recording the audio to be processed;

[0223] S202-5, in response to the end of the recording of the audio to be processed, display the track point generation interface, and display at least one track point in the virtual playback scene in the track point generation interface.

[0224] Optionally, in an embodiment of the present application, the above-mentioned first interactive operation may include but is not limited to a set of control instructions for the user to trigger audio recording. This operation can be implemented by clicking a button, long pressing a touch area, voice commands or gesture recognition, for example, clicking a circular recording icon on a mobile device screen to start recording, or activating voice capture through a pinch gesture in a VR device. Its technical feature is that it converts user intentions into signals that can be recognized by the system and triggers the initialization of subsequent audio capture modules. Application scenarios cover video barrage creation, real-time voice interaction and spatial audio content generation. For example, a user triggers barrage recording by double-clicking the screen while watching a video, and preloads audio encoding parameters (such as a sampling rate of 48kHz and mono mode).

[0225] It should be noted that there are many possibilities for the triggering method and feedback mechanism of the first interactive operation. For example, the triggering conditions can be expanded to device posture recognition (the phone starts recording when it is tilted more than 30 degrees), environmental sound triggering (automatic recording when specific keywords are detected) or cross-device collaboration (pressing the smart watch triggers the phone to record); operation feedback can include visual prompts (button color gradient), tactile vibration (short vibration prompts to start recording) or sound feedback (prompts the scale to rise); permission control can be designed to verify user identity before recording (face recognition), limit the length of a single recording (such as a maximum of 10 seconds) or dynamically adjust the sampling rate according to the network status (reduced to 16kHz when the network is weak). This application does not make specific restrictions on this.

[0226] Optionally, in an embodiment of the present application, the above-mentioned trajectory point generation interface may include but is not limited to a visual interactive platform for editing spatial trajectories. The interface usually includes a three-dimensional coordinate system of a virtual playback scene, trajectory point operation controls, and a preview window. For example, a virtual grid space superimposed on the video screen is projected in AR glasses, and the user adds trajectory points by dragging gestures. Its core functions include real-time rendering of trajectory paths, displaying coordinate parameters, and supporting dynamic adjustments. For example, in the bullet screen editing interface, the user can see the path curve of the trajectory point moving along the video timeline, and adjust the trajectory point spacing or movement speed through the slider. In terms of technical implementation, a physical engine can be integrated to simulate the inertia of trajectory motion, or provide an intelligent alignment function that is adsorbed to scene objects.

[0227] It should be noted that the presentation form and interaction logic of the trajectory point generation interface can be adapted to different hardware and scenario requirements. For example, the interface layout can be designed as a split-screen mode (video preview on the left, trajectory editing on the right), a floating translucent window (superimposed on the video) or a holographic projection (free manipulation in three-dimensional space); the interaction method can support precise drawing with a stylus, control with the direction keys of a game controller or eye-tracking focus selection; the trajectory generation mode can include manual frame-by-frame addition, AI-assisted automatic completion (generating matching trajectories based on voice content) or templated quick application (preset paths for explosion effects). This application does not make specific restrictions on this.

[0228] Through the embodiments of the present application, the technical threshold for users to create spatialized voice barrages is significantly reduced by adopting interactive operations with clear intentions and visual trajectory editing solutions. The diversified triggering methods of the first interactive operation adapt to the operating habits of multiple platforms such as mobile terminals and XR devices, and improve the immediacy and accuracy of voice input. The three-dimensional visualization design of the trajectory point generation interface enables users to intuitively perceive the temporal and spatial relationship between the sound source motion trajectory and the video screen, avoiding the abstractness of traditional parameter configuration. The two-stage operation process (recording first and then editing) effectively separates the core links in the creative process, reduces the probability of misoperation, and enhances the operation guidance through dynamic feedback of interface elements. This solution provides a standardized tool chain for the personalized expression of voice barrages and the creation of spatial sound fields, and promotes the immersive experience upgrade of user-generated content (UGC).

[0229] As an alternative, Figure 10 As shown, in response to the completion of the audio recording to be processed, a track point generation interface is displayed, and at least one track point is displayed in the virtual playback scene in the track point generation interface, including:

[0230] S202-5-1-1, in response to the completion of recording of the audio to be processed, determining the number of track points based on the duration of the audio to be processed;

[0231] S202-5-1-2, displaying a track point generation interface, and displaying track points in a virtual playback scene in the track point generation interface according to the number of track points.

[0232] Optionally, in an embodiment of the present application, the number of trajectory points mentioned above may include but is not limited to the total number of spatial marker points dynamically calculated or preset according to the audio duration. This value maps the audio time axis to a spatial trajectory density parameter through an algorithm. For example, 3 seconds of audio generates 3 trajectory points (one node per 1 second), or generates non-uniformly distributed points based on audio rhythm analysis (dense high-frequency paragraphs and sparse silent paragraphs). The calculation basis may include audio waveform energy distribution (such as zero-crossing rate peak points), speech semantic segmentation (such as sentence intervals) or user-defined rules (such as setting a point to be generated every 500ms). The role of the number of trajectory points is to balance the complexity of the spatial trajectory and the consumption of rendering resources, ensuring that the motion path can express the audio characteristics while avoiding the burden of redundant data.

[0233] It should be noted that the method for determining the number of trajectory points can be flexibly adjusted according to the audio content and business requirements. For example, the calculation strategy can be extended to be based on speech emotion recognition (increasing the trajectory point density in exciting paragraphs), environmental noise level (reducing the number of points in high-noise scenarios to reduce interference), or device performance constraints (limiting the maximum number of points on mobile devices to no more than 5); the dynamic adjustment mechanism can support users to manually override the algorithm results (such as deleting redundant points or appending key points), real-time preview feedback (prompting performance warnings when the number of points is too large), or automatic optimization (dynamically increasing or decreasing the number of points according to the rendering load). This application does not make specific limitations on this.

[0234] Optionally, in the embodiments of this application, the above-mentioned trajectory point generation interface may include, but is not limited to, an interactive operation platform integrating virtual scene rendering and trajectory editing functions. This interface usually includes a three-dimensional coordinate system visualization module, trajectory point operation controls (such as drag / delete / snap tools), and a parameter setting panel (such as movement speed, interpolation type). For example, in the voice bullet screen editing scenario, the interface can display a virtual space grid synchronized with the video screen, and users can add trajectory points at specific coordinate positions with a stylus while real-time previewing the moving path of the sound source. Its technical implementation can be combined with a GPU-accelerated rendering engine, a physical collision detection algorithm, or an AI-assisted path generation module.

[0235] It should be noted that there are multiple implementation solutions for the presentation form and interaction logic of the trajectory point generation interface. For example, the interface layout can be adapted to split-screen mode (left audio waveform diagram and right three-dimensional trajectory editing area), AR overlay mode (trajectory points floating in the real scene), or holographic projection mode (naked-eye three-dimensional manipulation); the interaction method can support gesture recognition (pinching with both hands to zoom in and out of the scene), eye movement following (gaze focus to select trajectory points), or cross-device collaboration (drawing a path on a mobile phone and synchronizing it to a VR headset); the auxiliary functions can integrate automatic path optimization (avoiding virtual obstacles according to the acoustic model), a historical trajectory template library (one-key application of common motion curves), or collaborative editing (multiple people modifying different trajectory segments simultaneously). This application does not make specific limitations on this.

[0236] Through the embodiments of this application, the dynamic generation mechanism of the number of trajectory points driven by the audio duration and the design of the visualization editing interface significantly reduce the operation threshold for users to create spatial voice bullet screens. The physical modeling and acoustic simulation capabilities of the virtual playback scene enable the trajectory points to not only have spatial coordinate attributes but also accurately reflect the characteristics of sound wave propagation. For example, when blocked by an obstacle, it can automatically calculate the gain attenuation of the diffraction path. The diverse adaptation solutions for interface interaction support full-scene coverage from lightweight editing on mobile devices to high-precision control on professional XR devices, improving the creation efficiency and flexibility. The intelligent optimization strategy for the number and density of trajectory points avoids data overload while ensuring spatial expressiveness, providing a reliable data basis for real-time rendering and cross-platform transmission.

[0237] As an alternative, as Figure 10 shown, determining the speaker gain corresponding to the target virtual speaker based on the initial audio data includes:

[0238] S202-5-2-1, when the number of track points meets the preset number condition, determining the target position vector corresponding to each track point based on the position information corresponding to each track point, and determining the position indicated by each target position vector as the virtual source position corresponding to each track point, where the target position vector is used to indicate the spatial position of the corresponding track point in the virtual playback scenario;

[0239] S202-5-2-2, obtaining the distances between each virtual speaker and each virtual source position in the virtual playback scenario;

[0240] S202-5-2-3, determining the speakers whose distances meet the preset distance condition as the target virtual speakers corresponding to each track point, and determining the speaker position vectors corresponding to each target virtual speaker corresponding to each track point;

[0241] S202-5-2-4, based on the target position vector and the speaker position vector, determining the gain coefficient of each target virtual speaker corresponding to each track point, where the speaker position vector is the same as the target position vector after weighted summation according to the gain coefficient, and the speaker gain includes the gain coefficient.

[0242] Optionally, in the embodiments of the present application, the above preset number condition may include, but is not limited to, a trigger threshold rule for determining whether to perform multi-track point sound field calculation. This condition can be dynamically adjusted based on the audio duration, device performance, or user-set values. For example, when the number of track points is greater than or equal to 3, a high-order sound field rendering algorithm is enabled, or the maximum number of track points is limited to 5 according to the mobile GPU load. Its function is to balance the computational complexity and spatial positioning accuracy. For example, in the voice bullet screen scenario, if the user only sets 1 track point, static sound source rendering is used; if the number of track points exceeds the preset value (such as 4), the dynamic path interpolation algorithm is activated to generate a smooth movement trajectory.

[0243] It should be noted that there are various possibilities for the setting dimension and determination logic of the preset number condition. For example, the condition type can be extended to be based on the audio content complexity (such as increasing the point limit when the proportion of high-frequency components exceeds the threshold), user privilege level (VIP users are allowed to set more track points), or scene type (the game scene forces a minimum of 3 points); the dynamic adjustment strategy can be designed as real-time performance monitoring (automatically reducing the number of points when the CPU occupancy exceeds 80%), network status adaptation (reducing the number of points in a weak network environment to ensure real-time transmission), or AI prediction (recommending the optimal number of points based on historical data). The present application does not make specific limitations in this regard.

[0244] Optionally, in the embodiments of the present application, the above gain coefficient may include, but is not limited to, an energy distribution weight dynamically calculated according to the spatial relationship between the sound source and the speaker. This coefficient is determined by a vector projection algorithm, so that the synthesized sound field of multiple speakers reaches a preset sound pressure level in the direction of the target position vector. For example, the gain of the left front speaker is 0.7 and the gain of the right rear speaker is 0.3 to simulate that the sound source is located in front of the listener's left. Its physical meaning is to reconstruct the azimuth perception of the virtual sound source at the listener's position through the collaborative output of multiple speakers.

[0245] It should be noted that the calculation and distribution method of the gain coefficient can be extended according to the device type and acoustic target. For example, the optimization target can be set to minimize the positioning error (infer the weight by reverse engineering the psychoacoustic test data), maximize the energy efficiency ratio (prioritize the use of high-sensitivity speakers), or balance the power consumption (avoid overloading a single speaker); the calculation framework can be adapted to offline pre-calculation (generate a gain table in advance), real-time iterative optimization (update the coefficient for each frame), or hybrid mode (pre-calculate the key frames + real-time interpolation); the exception handling can include smooth gain transition (enable fade-in and fade-out when there is a mutation), invalid speaker filtering (shield the faulty device), or degraded rendering (only retain the main channel coefficient). The present application does not make specific limitations on this.

[0246] Through the embodiments of the present application, by adopting a multi-threshold condition judgment and vectorized space mapping mechanism, the dynamic adaptation and efficient calculation of the spatial sound field of the voice bullet screen are realized. The introduction of the preset number condition effectively avoids resource waste in low-complexity scenarios, and at the same time automatically enables a high-order algorithm to ensure the positioning accuracy when the trajectory points are dense. The standardized expression of the target position vector provides a unified interface for multi-source data collaboration and cross-platform rendering, and solves the problem of aligning the spatial reference systems between heterogeneous devices. The vector projection calculation strategy of the gain coefficient accurately restores the sound field energy distribution at the physical level, significantly improving the accuracy of sound image positioning and the sense of environmental immersion. This solution ensures the full-link stability from mobile devices to professional audio workstations through a dynamic strategy adaptation and exception handling mechanism, providing a reliable technical foundation for real-time interaction and high-quality spatial audio creation.

[0247] As an optional solution, the above method further includes:

[0248] When the number of trajectory points meets the preset number condition, obtain the playback time corresponding to each trajectory point;

[0249] Based on the playback time interval between the first trajectory point and the second trajectory point with adjacent playback times, increase the trajectory points by interpolation.

[0250] Optionally, in the embodiments of the present application, the above playback time may include, but is not limited to, the time-axis marking parameters associated with the track points. This parameter can be embodied as an absolute timestamp (such as 500 ms after the start point of the audio), a relative time ratio (such as the 30% position of the total audio duration), or an event trigger marker (such as the moment when a specific voice keyword appears). For example, in the bullet screen editing scenario, the playback time can be associated with the exact moment when a character in the video starts speaking, or the start point of a highlight segment manually marked by the user. Its data source can include audio waveform analysis (such as energy peak detection), external event signals (such as video key frames), or user interaction inputs.

[0251] It should be noted that the acquisition and calibration methods of the playback time can be adapted to different application scenarios. For example, the time marking source can include audio fingerprint matching (automatically aligning repeated paragraphs), multi-modal sensor synchronization (such as associating accelerometer data with the movement trajectory), or cross-platform time synchronization (aligning with the streaming media server clock in the live broadcast scenario); the calibration strategy can involve dynamic delay compensation (automatically offsetting the timestamp during network fluctuations), error smoothing processing (eliminating device clock jitter), or outlier filtering (removing error markers that deviate significantly from the time axis). The present application does not make specific limitations on this.

[0252] Optionally, in the embodiments of the present application, the above interpolation method may include, but is not limited to, an algorithm set for generating transition track points between adjacent track points. This method can generate uniformly distributed points based on linear interpolation, generate a smooth path based on a Bezier curve, or simulate the motion inertia of a physical engine (such as a parabolic trajectory). For example, in the voice bullet screen scenario, if the interval between two track points is 1 second and the spatial distance is large, 3 intermediate points can be inserted through cubic spline interpolation to make the movement speed of the sound image match the rhythm of the voice. The interpolation parameters can be dynamically adjusted, such as controlling the path curvature or acceleration according to the audio spectrum characteristics.

[0253] It should be noted that the selection of the interpolation algorithm and the parameter configuration can be flexibly extended according to the rendering target. For example, the interpolation dimension can support spatial coordinate interpolation (generating position transition points), time-space joint interpolation (synchronously adjusting the position and timestamp), or attribute interpolation (such as the volume gradually changing along the path); the parameter optimization can be based on an acoustic model (the interpolation point density matches the reverberation decay characteristics), user preferences (manually setting the path curvature), or real-time performance monitoring (degrading to linear interpolation when computing resources are scarce). The present application does not make specific limitations on this.

[0254] Through the embodiments of the present application, by adopting a dynamic threshold judgment and an adaptive interpolation mechanism, the smoothness and rhythm adaptability of the spatial trajectory of the voice bullet screen are achieved. The intelligent determination of the preset quantity condition effectively avoids inefficient calculations, starts interpolation only when the trajectory complexity reaches the requirement, and saves system resources. The multi-source acquisition and calibration strategy of the playback time ensure the accurate time synchronization of the trajectory points and the audio content. For example, it accurately triggers the change of the surround sound field at the highlight moment of the video. The diverse adaptation ability of the interpolation algorithm supports the flexible generation from simple straight paths to complex motion curves, making the movement trajectory of the sound image highly consistent with the emotional ups and downs of the voice. This solution provides high-availability technical support for spatial audio creation in different scenarios by dynamically optimizing the interpolation density and computational load, taking into account the performance differences between mobile devices and professional devices.

[0255] As an alternative solution, as Figure 10 shown, based on the speaker gain, the audio to be processed and the media information indicated by the source media information identifier are mixed to generate target audio data, including:

[0256] S206-1, obtaining the information type corresponding to the media information indicated by the source media information identifier;

[0257] S206-2, determining the gain allocation ratio based on the information type;

[0258] S206-3, mixing the audio to be processed and the media information indicated by the source media information identifier according to the gain allocation ratio to generate target audio data.

[0259] Optionally, in the embodiments of the present application, the above information type may include, but is not limited to, the functional classification identifier of the media content. This type can be divided according to content attributes (such as background music, environmental sound effects, voice narration), scene uses (such as movie soundtracks, game sound effects, live interactions), or emotional labels (such as tense, cheerful, suspenseful). For example, in the voice bullet screen scenario, the information type can be identified as "user voice bullet screen" and "video original sound track". The former needs to improve the voice clarity, and the latter needs to control the background volume to avoid interference. Its role is to guide the differential processing of the mixing strategy. For example, set the gain ratio of the voice bullet screen to 0.8 and the background music to 0.3 to ensure the clear hierarchy of the primary and secondary audio.

[0260] It should be noted that there are multiple possibilities for the acquisition method of information types and the classification dimensions. For example, the classification basis can be extended to copyright attributes (original / authorized / public materials), technical formats (mono / stereo / Ambisonics), or interaction modes (editable / read-only / dynamic response); the acquisition methods can include automatic content analysis (spectrum feature clustering), user manual marking (checking type labels in the editing interface), or third-party metadata import (reading classification identifiers from video platform interfaces); the dynamic update mechanism can support real-time type switching (temporarily inserting an advertisement sound effect type during a live broadcast) or learning-based classification (optimizing type identifiers according to the user operation history). This application does not make specific limitations in this regard.

[0261] Optionally, in the embodiments of this application, the above gain allocation ratio may include, but is not limited to, a set of multi-channel volume weight parameters dynamically adjusted according to information types. This ratio is determined by audio content priority, scene requirements, or user preferences. For example, in a multi-person voice bullet chat scenario, the host's voice is allocated a gain of 0.7, the audience's bullet chat voice is allocated a gain of 0.4, and the background music is allocated a gain of 0.2. The calculation basis can include semantic analysis (such as keyword detection to increase the gain of emergency notifications), acoustic features (such as appropriately reducing noise for voices with a high proportion of high-frequency components), or device capabilities (such as enabling spatial gain compensation for headphone devices).

[0262] It should be noted that the determination strategy of the gain allocation ratio can be extended according to business scenarios and device characteristics. For example, the ratio calculation model can select fixed priority rules (voice bullet chat is always higher than background music), dynamic loudness equalization (dynamically adjusting according to real-time audio energy), or emotion-driven models (automatically increasing the low-frequency gain in tense scenarios); the optimization objectives can include maximizing speech intelligibility, enhancing the sense of sound field space, or power consumption equalization (limiting the number of high-gain channels); the exception handling mechanism can include ratio clipping (preventing gain from exceeding the limit and causing distortion), conflict resolution (enabling weighted voting when multiple types compete), or degradation strategies (merging channels when device resources are insufficient). This application does not make specific limitations in this regard.

[0263] Optionally, in the embodiments of this application, the above target audio data may include, but is not limited to, a playable signal for a terminal device formed after mixing multi-source audio. This data includes multi-channel audio streams (such as 5.1 surround sound), metadata tags (such as loudness normalization parameters), and device adaptation instructions (such as headphone virtualization identifiers). For example, in a voice bullet chat scenario, the target audio data can be generated by mixing the user's voice bullet chat (left front channel gain of 0.6), video original sound (center channel gain of 0.4), and environmental sound effects (surround channel gain of 0.3), and encapsulated as a multi-track audio file or a real-time streaming media data packet.

[0264] It should be noted that there are various implementation solutions for the technical path of mixing processing and output optimization. For example, the processing flow can be divided into preprocessing (noise suppression / equalization filtering), core mixing (multi-track gain superposition / phase alignment), and post-processing (dynamic compression / limiter protection); the mixing mode can support offline rendering (high-precision multi-track synthesis), real-time low-latency processing (game voice interaction), or layered mixing (outputting tracks separately for later adjustment); output optimization can include metadata embedding (such as dynamic range control tags), format adaptive transcoding (switching the encoding scheme according to the playback device), or security enhancement (digital watermark / encrypted stream encapsulation). This application does not make specific limitations on this.

[0265] Through the embodiments of this application, an information type-driven dynamic gain allocation strategy is adopted to achieve intelligent mixing and hierarchical presentation of multi-source audio. Based on the differential processing mechanism of type identifiers, it ensures that the gain ratio of voice bullet screens and background sound tracks adapts to the scene requirements. For example, in the emergency notification scenario, the voice clarity is preferentially improved, and in the music appreciation scenario, the original sound quality is enhanced. The structured output design of the target audio data supports full-scenario coverage from professional production to real-time interaction, improving cross-platform playback compatibility. The flexible expansion ability of the mixing processing flow takes into account the requirements of high-precision offline rendering and low-latency real-time synthesis, providing a technical basis for personalized audio experience and efficient content creation.

[0266] The following will further explain and illustrate this application with specific examples:

[0267] Exemplarily, taking the voice bullet screen as an example, it includes but is not limited to the following process:

[0268] As Figure 3 shown, record user interaction, add a voice bullet screen recording entry button on the playback page, and click the entry button to open the voice bullet screen recording page.

[0269] As Figure 4 shown, on the voice bullet screen recording page, the left half is for voice recording to allow users to record sounds. The maximum recording time is only allowed to be 5 seconds. (The maximum recording duration can be changed according to business requirements). After the recording is completed, the sound azimuth editor on the right can be operated.

[0270] It should be noted that the sound track can be added through the "Add Track" button, and the sound spatial playback position can be added. T is the start playback time of the current sound position, with the unit of milliseconds. X, Y, and Z are the coordinate values of the spatial coordinate system centered on the human head respectively. After successful addition, the current specified position will be displayed on the spatial position schematic diagram. In addition, multiple positions can be added to the currently recorded bullet screen. The interval between each position should be no less than 500 ms. Each voice bullet screen shall not exceed 5 position coordinates. (Both of the above parameters can be adjusted by the service according to the actual situation). Moreover, multiple spatial coordinates can form a bullet screen movement track, and the bullet screen voice is moved between each position based on a smoothing algorithm using the bullet screen movement track. If there is only one spatial coordinate, it can be fixed at the specified position for playback.

[0271] S3, as Figure 8 shown, after the bullet screen recording is completed, click to send the bullet screen, and the generated voice bullet screen will be sent to the server background for storage. The service can charge members according to the voice bullet screen. The voice bullet screen uploaded by users will be distributed to all users at the playback end. Users can click the button of the voice bullet screen style to perform rendering and playback.

[0272] In an exemplary embodiment, as Figure 11 shown, the vector pointing from the origin to the sound source becomes the direction vector of the sound source. θ and φ respectively represent the elevation angle and the azimuth angle. θ ∈ [0, π] is defined as the angle between the direction vector and the positive direction of the z-axis; φ ∈ [0, 2π] is the angle between the direction vector and the positive direction of the x-axis, and r represents the radial distance from the coordinate origin. The relationship between spherical coordinates (r, θ, φ) and rectangular coordinates (x, y, z) is:

[0273]

[0274] Among them, the bullet screen position information includes the elevation angle, azimuth angle and radius of the direction vector of the sound, the start time, and the bullet screen duration. Each voice bullet screen includes the content of the voice bullet screen (audio adts header + aac-encoded voice raw data), multiple bullet screen position information (for forming a track), the bullet screen duration, the time when the voice bullet screen appears (for the appearance of the voice bullet screen), and the video id matched by the bullet screen (for matching the corresponding video source).

[0275] It should be noted that the voice bullet screen generation method can include but is not limited to the following processes:

[0276] The voice content recorded on the voice danmu recording page is stored as adts header + aac rawdata. The relevant parameters include sample rate, channels, sample format, etc. The compression format is aac. This data is recorded as the voice danmu content, which is mono audio;

[0277] On the voice danmu recording page, each trajectory point is recorded as danmu position information, and multiple trajectory points can be recorded as a group of danmu position information;

[0278] Based on the danmu content and danmu position information, complete danmu data is formed, and the duration of the voice danmu, the time when the danmu appears in the video source, and the corresponding video source identifier are recorded;

[0279] In response to the user clicking to send the danmu, the complete danmu data is sent to the background for storage.

[0280] In an exemplary embodiment, Figure 13 is a schematic diagram of another optional audio data processing method according to an embodiment of the present application, as Figure 13 shown, wherein:

[0281] Source audio stream is the video source. For example, movies, TV series, etc. launched on a certain video sharing platform.

[0282] Danmu stream is the user danmu data stream, which can be downloaded locally according to the time when the user sends the danmu.

[0283] Danmu audio decoder is a separate danmu audio data decoder, which is responsible for decoding the danmu audio data. The decoded data is combined with the corresponding danmu position information as the input of the danmu renderer.

[0284] Danmu Spatial Render is a module for spatial rendering. According to the pcm data and spatial position information of the danmu, and through an algorithm (such as vbap), it is mapped to a 5.1.4 speaker bed (the 5.1.4 bed can select 7.1.2 or other channal layouts according to the actual scenario).

[0285] Danmu Spatial Render is internally configured with n mono channels, which means how many voice danmus can be rendered simultaneously. N is configurable.

[0286] The Source audio stream forms the layout of the speaker bed through upmixing. For example, stereo upmixing to a 5.1.4 speaker bed.

[0287] The Mixer module mixes the voice barrage and the video source, and will perform a volume gain adjustment. To maintain the priority of the voice barrage, the volume of the video source can be reduced and the proportion of the voice barrage can be increased. For example, the volume adjustment is: source volume: voice barrage volume = 3:7.

[0288] Among them, the calculation of the barrage position information is as Figure 6 shown. The barrage information contains the position information where the voice barrage should be played: P (pitch angle, horizontal angle, distance). By combining the pitch angle, horizontal angle, and distance of the barrage with the VBAP algorithm, the corresponding three speakers l m , l k , l n .

[0289] The VBAP calculation process is as follows:

[0290] As Figure 7 shown: Given the azimuth angle and pitch angle of the input object in the spherical coordinate system, convert the spherical coordinates to a three-dimensional unit vector p = [p1 p2 p3] T .

[0291]

[0292] Among them, the gain of the target virtual speaker is calculated in the following way: Find the three speakers closest to the virtual source: l m , l k , l n . Map the 5.1.4 channel layout to the spherical space, and respectively obtain three three-dimensional vectors pointing to lm, lk, ln, L1 = [l 11 , l 12 , l 13 pointing to l m , L2 = [l 21 , l 22 , l 23 pointing to l k , L3 = [l 31 , l 32 , l 33 pointing to l n , such that:

[0293] p = g1l1 + g2l2 + g3l3;

[0294] Here g = [g1, g2, g3], representing the gains of the three speakers. If the sound is close to l m , the gain of g1 will approach 1, and the same applies to other positions. And the energies of the final speakers p1, p2, p3 are obtained.

[0295] It should be noted that multiple voice barrage position information can be added, so that the movement trajectory of the sound can be formed. Figure 14 It is a schematic diagram of another optional method for processing audio data according to an embodiment of the present application. The voice barrage positions are as Figure 14 shown. For each position P1, P2, P3, P4, calculations will be performed. For P n where the time interval between them is greater than 500 ms, interpolation can be performed, and position information P n1 …P nn is inserted every 500 ms. According to the inserted P n1 , P n2 , P n3 …, the aforementioned gain calculation is performed.

[0296] Through the embodiments of the present application, the rendering and playing of the voice barrage can have a sense of spatial orientation. It allows users to have a sense of spatial movement when listening to the voice barrage, providing a more immersive experience compared to traditional voice barrages, and can also be used as a membership privilege of the service, etc.

[0297] It can be understood that in the specific implementation of the present application, data related to user information, etc. is involved. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.

[0298] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0299] According to another aspect of the embodiments of the present application, there is also provided an audio data processing device for implementing the above-mentioned audio data processing method. As Figure 15 shown, the device includes:

[0300] An acquisition module 1502, configured to acquire the audio to be processed and at least one trajectory point, and generate initial audio data based on the audio to be processed and the trajectory point, where the initial audio data includes the source media information identifier corresponding to the audio to be processed and the position information of the trajectory point;

[0301] A determination module 1504 is configured to determine a speaker gain corresponding to a target virtual speaker based on initial audio data. The position information of the trajectory point is used to determine the virtual source position in the virtual playback scene. The virtual playback scene is deployed with virtual speakers. The target virtual speaker is at least one of the virtual speakers, and the target virtual speaker represents the virtual speaker determined based on the virtual source position. The speaker gain is used to simulate the playback of the audio to be processed at the virtual source position.

[0302] A generation module 1506 is configured to perform a mixing process on the audio to be processed and the media information indicated by the source media information identifier based on the speaker gain to generate target audio data.

[0303] As an alternative solution, the above device determines the speaker gain corresponding to the target virtual speaker based on the initial audio data in the following manner: determining the virtual source position based on the position information; determining the target virtual speaker and the speaker gain according to the virtual source position and the speaker positions corresponding to the respective virtual speakers in the virtual playback scene, where the target virtual speaker is the speaker in the virtual playback scene whose distance from the virtual source position meets a preset distance condition.

[0304] As an alternative solution, the above device determines the virtual source position based on the position information in the following manner: converting the position information into spatial trajectory point coordinates in a spherical coordinate system, where the position information represents the trajectory point coordinates in a rectangular coordinate system; determining a target position vector based on the spatial trajectory point coordinates, and determining the position indicated by the target position vector as the virtual source position, where the target position vector is used to indicate the spatial position of the corresponding trajectory point in the virtual playback scene.

[0305] As an alternative solution, the above device determines the target virtual speaker and the speaker gain according to the virtual source position and the speaker positions corresponding to the respective virtual speakers in the virtual playback scene in the following manner: obtaining the distances between the respective virtual speakers in the virtual playback scene and the virtual source position; determining the speakers whose distances meet the preset distance condition as the target virtual speakers; determining the speaker position vectors corresponding to the target virtual speakers; determining a gain coefficient based on the target position vector and the speaker position vectors, where the sum of the speaker position vectors weighted by the gain coefficient is the same as the target position vector, and the speaker gain includes the gain coefficient.

[0306] As an alternative solution, the above-mentioned device is used to obtain the audio to be processed and at least one trajectory point, and generate initial audio data based on the audio to be processed and the trajectory point: obtain the audio to be processed and mark the audio timestamp; receive the trajectory points input by the user, where each trajectory point includes a spatial position parameter and associated time information; bind the audio to be processed with the spatial position parameter based on the audio timestamp and the associated time information to generate the initial audio data.

[0307] As an alternative solution, the above-mentioned device is used to obtain the audio to be processed and mark the audio timestamp in at least one of the following ways: when the media information indicated by the source media information identifier includes a recorded video, mark the audio timestamp based on the playback progress of the recorded video, where the audio timestamp indicates that the audio to be processed is set to be played at the corresponding playback progress; when the media information indicated by the source media information identifier includes a live video, mark the audio timestamp based on the system time, where the audio timestamp indicates that the audio to be processed is set to be played at the corresponding system time.

[0308] As an alternative solution, the above-mentioned device is used to bind the audio to be processed with the spatial position parameter based on the audio timestamp and the associated time information in at least one of the following ways to generate the initial audio data: when the media information indicated by the source media information identifier includes a recorded video, bind the associated time information with the first playback progress of the recorded video, where the first playback progress indicates that when the recorded video is played to the first playback progress, the audio to be processed is played according to the speaker gain set by the corresponding trajectory point; when the media information indicated by the source media information identifier includes a live video, bind the associated time information with the second playback progress of the audio to be processed, where the second playback progress indicates that when the audio to be processed is played to the second playback progress, the audio to be processed is played according to the speaker gain set by the corresponding trajectory point.

[0309] As an alternative solution, the above-mentioned device is used to obtain the audio to be processed and at least one trajectory point, and generate initial audio data based on the audio to be processed and the trajectory point: in response to a first interaction operation performed on the corresponding application interface, start recording the audio to be processed; in response to the end of the recording of the audio to be processed, display a trajectory point generation interface, and display at least one trajectory point in the virtual playback scene in the trajectory point generation interface.

[0310] As an alternative solution, the above-mentioned device is used to, in response to the end of the recording of the audio to be processed, display a trajectory point generation interface and display at least one trajectory point in the virtual playback scene in the trajectory point generation interface in the following way: in response to the end of the recording of the audio to be processed, determine the number of trajectory points based on the duration of the audio to be processed; display the trajectory point generation interface, and display the trajectory points in the virtual playback scene in the trajectory point generation interface according to the number of trajectory points.

[0311] As an alternative, the above device is used to determine the speaker gain corresponding to the target virtual speaker based on the initial audio data in the following manner: when the number of track points meets the preset number condition, determine the target position vector corresponding to each track point based on the position information corresponding to each track point, and determine the virtual source position corresponding to each track point as the position indicated by each target position vector, where the target position vector is used to indicate the spatial position of the corresponding track point in the virtual playback scene; obtain the distances between each virtual speaker and each virtual source position in the virtual playback scene; determine the speakers whose distances meet the preset distance condition as the target virtual speakers corresponding to each track point, and determine the speaker position vector corresponding to each target virtual speaker corresponding to each track point; based on the target position vector and the speaker position vector, determine the gain coefficient of each target virtual speaker corresponding to each track point, where the sum of the speaker position vectors weighted by the gain coefficient is the same as the target position vector, and the speaker gain includes the gain coefficient.

[0312] As an alternative, the above device is further used to: when the number of track points meets the preset number condition, obtain the playback time corresponding to each track point; increase the track points by interpolation based on the playback time interval between the first track point and the second track point with adjacent playback times.

[0313] As an alternative, the above device is used to mix the audio to be processed and the media information indicated by the source media information identifier based on the speaker gain in the following manner to generate the target audio data: obtain the information type corresponding to the media information indicated by the source media information identifier; determine the gain allocation ratio based on the information type; mix the audio to be processed and the media information indicated by the source media information identifier according to the gain allocation ratio to generate the target audio data.

[0314] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0315] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.

[0316] According to one aspect of the present application, there is provided a computer program product, which includes a computer program.

[0317] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.

[0318] Figure 16 Schematically shown is a block diagram of a computer system of an electronic device for implementing the embodiments of the present application.

[0319] It should be noted that Figure 16 The computer system 1600 of the shown electronic device is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0320] As Figure 16 shown, the computer system 1600 includes a central processing unit 1601 (Central Processing Unit, CPU), which can perform various appropriate actions and processes according to the program stored in the read-only memory 1602 (Read-Only Memory, ROM) or the program loaded from the storage section 1608 into the random access memory 1603 (Random Access Memory, RAM). In the random access memory 1603, various programs and data required for system operation are also stored. The central processing unit 1601, the read-only memory 1602, and the random access memory 1603 are connected to each other via a bus 1604. The input / output interface 1605 (Input / Output interface, i.e., I / O interface) is also connected to the bus 1604.

[0321] The following components are connected to the input / output interface 1605: an input section 1606 including a keyboard, a mouse, etc.; an output section 1607 including such as a cathode ray tube (Cathode Ray Tube, CRT), a liquid crystal display (Liquid Crystal Display, LCD), etc. and a speaker, etc.; a storage section 1608 including a hard disk, etc.; and a communication section 1609 including a network interface card such as a local area network card, a modem, etc. The communication section 1609 performs communication processing via a network such as the Internet. A drive 1610 is also connected to the input / output interface 1605 as needed. A removable medium 1611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1610 as needed so that the computer program read from it can be installed into the storage section 1608 as needed.

[0322] In particular, according to the embodiments of the present application, the processes described in each method flowchart can be implemented as computer software programs. For example, the embodiments of the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1609, and / or installed from the removable medium 1611. When the computer program is executed by the central processing unit 1601, various functions defined in the system of the present application are executed.

[0323] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1609, and / or installed from the removable medium 1611. When the computer program is executed by the central processing unit 1601, various functions provided by the embodiments of the present application are executed.

[0324] According to another aspect of the embodiments of the present application, an electronic device for implementing the above-mentioned audio data processing method is further provided. The electronic device may be Figure 1 the terminal device or server shown. This embodiment takes the electronic device as the terminal device as an example for illustration. As Figure 17 shown, the electronic device includes a memory 1702 and a processor 1704. A computer program is stored in the memory 1702, and the processor 1704 is configured to execute the steps in any of the above method embodiments through the computer program.

[0325] Optionally, in this embodiment, the above-mentioned electronic device may be at least one of multiple network devices in a computer network.

[0326] Optionally, in this embodiment, the above-mentioned processor may be configured to execute the methods in the embodiments of the present application through the computer program.

[0327] Optionally, those of ordinary skill in the art can understand that Figure 17 the structure shown is only schematic, Figure 17 and it does not limit the structure of the above-mentioned electronic device. For example, the electronic device may further include more or fewer components (such as a network interface, etc.) than those shown in Figure 17 , or have a different configuration from that shown in Figure 17 .

[0328] Among them, the memory 1702 can be used to store software programs and modules, such as the program instructions / modules corresponding to the audio data processing method and device in the embodiments of the present application. The processor 1704 executes various functional applications and data processing by running the software programs and modules stored in the memory 1702, that is, to implement the above-mentioned audio data processing method. The memory 1702 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 1702 may further include a memory remotely disposed relative to the processor 1704, and these remote memories can be connected to the terminal device through a network. Examples of the above network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 1702 can specifically but not limitedly be used to store information such as audio. As an example, as Figure 17 shown, the above-mentioned memory 1702 may include but is not limited to the acquisition module 1502, the determination module 1504, and the generation module 1506 in the above-mentioned audio data processing device. In addition, it may also include but is not limited to other module units in the above-mentioned audio data processing device, which will not be elaborated in this example.

[0329] Optionally, the above-mentioned transmission device 1706 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wired network and a wireless network. In one instance, the transmission device 1706 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers through a network cable, thereby enabling communication with the Internet or a local area network. In one instance, the transmission device 1706 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0330] In addition, the above-mentioned electronic device further includes: a display 1708 for displaying the above-mentioned bullet screen information; and a connection bus 1710 for connecting each module component in the above-mentioned electronic device.

[0331] In other embodiments, the above-mentioned terminal device or server may be a node in a distributed system. Among them, the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting the multiple nodes through network communication. Among them, the nodes can form a peer-to-peer network, and any form of computing device, such as an electronic device such as a server or a terminal device, can become a node in the blockchain system by joining the peer-to-peer network.

[0332] According to one aspect of the present application, a computer-readable storage medium is provided. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the audio data processing method provided in various alternative implementations of the above-mentioned audio data processing aspect.

[0333] Optionally, in this embodiment, the above-mentioned computer-readable storage medium may be configured to store instructions for executing the methods in the various embodiments of the present application.

[0334] Optionally, in this embodiment, those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program. The program can be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disc, or the like.

[0335] The serial numbers of the above embodiments of the present application are only for description and do not represent the superiority or inferiority of the embodiments.

[0336] If the integrated unit in the above embodiment is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in the above-mentioned computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing one or more electronic devices to execute all or part of the steps of the methods described in the various embodiments of the present application.

[0337] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0338] In the several embodiments provided by the present application, it should be understood that the disclosed application program can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of the units or modules can be in an electrical or other form.

[0339] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0340] In addition, each functional unit in various embodiments of the present application may be integrated into a processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0341] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for processing audio data, characterized in that Including: Obtain the audio to be processed and at least one track point, and generate initial audio data based on the audio to be processed and the track point, where the initial audio data includes the source media information identifier corresponding to the audio to be processed and the position information of the track point; Determine the speaker gain corresponding to the target virtual speaker based on the initial audio data, where the position information of the track point is used to determine the virtual source position in the virtual playback scenario, the virtual playback scenario is deployed with virtual speakers, the target virtual speaker is at least one of the virtual speakers, the target virtual speaker represents the virtual speaker determined based on the virtual source position, and the speaker gain is used to simulate the playback of the audio to be processed at the virtual source position; Perform a mixing process on the audio to be processed and the media information indicated by the source media information identifier based on the speaker gain to generate target audio data.

2. The method according to claim 1, characterized in that, The determining the speaker gain corresponding to the target virtual speaker based on the initial audio data includes: Determine the virtual source position based on the position information; Determine the target virtual speaker and the speaker gain according to the virtual source position and the speaker positions corresponding to the respective virtual speakers in the virtual playback scenario, where the target virtual speaker is the speaker in the virtual playback scenario whose distance from the virtual source position satisfies a preset distance condition.

3. The method according to claim 2, wherein The determining the virtual source position based on the position information includes: Convert the position information into spatial track point coordinates in the spherical coordinate system, where the position information represents the track point coordinates in the rectangular coordinate system; Determine a target position vector based on the spatial track point coordinates, and determine the position indicated by the target position vector as the virtual source position, where the target position vector is used to indicate the spatial position of the corresponding track point in the virtual playback scenario.

4. The method according to claim 3, characterized in that The determining the target virtual speaker and the speaker gain according to the virtual source position and the speaker positions corresponding to the respective virtual speakers in the virtual playback scenario includes: Obtain the distances between the respective virtual speakers in the virtual playback scenario and the virtual source position; Determine the virtual speakers whose distances satisfy the preset distance condition as the target virtual speakers; Determine the speaker position vector corresponding to the target virtual speaker; Determine a gain coefficient based on the target position vector and the speaker position vector, where the speaker position vector is the same as the target position vector after being weighted and summed according to the gain coefficient, and the speaker gain includes the gain coefficient.

5. The method according to claim 1, wherein The obtaining the audio to be processed and at least one track point, and generating initial audio data based on the audio to be processed and the track point includes: Obtain the audio to be processed and mark the audio timestamp; Receive the track points input by the user, where each track point includes a spatial position parameter and associated time information; Bind the audio to be processed and the spatial position parameter based on the audio timestamp and the associated time information to generate the initial audio data.

6. The method according to claim 5, characterized in that, Obtaining the audio to be processed and marking the audio timestamp includes at least one of the following: Obtaining the audio to be processed and the media information indicated by the source media information identifier, and when the media information includes a recorded video, marking the audio timestamp based on the playback progress of the recorded video, where the audio timestamp indicates that the audio to be processed is set to be played at the corresponding playback progress; Obtaining the audio to be processed and the media information indicated by the source media information identifier, and when the media information includes a live video, marking the audio timestamp based on the system time, where the audio timestamp indicates that the audio to be processed is set to be played at the corresponding system time.

7. The method according to claim 5, wherein Binding the audio to be processed and the spatial position parameter based on the audio timestamp and the associated time information to generate the initial audio data includes at least one of the following: When the media information indicated by the source media information identifier includes a recorded video, binding the associated time information to the first playback progress of the recorded video to generate the initial audio data, where the first playback progress indicates that when the recorded video is played to the first playback progress, the audio to be processed is played according to the speaker gain set at the corresponding trajectory point; When the media information indicated by the source media information identifier includes a live video, binding the associated time information to the second playback progress of the audio to be processed to generate the initial audio data, where the second playback progress indicates that when the audio to be processed is played to the second playback progress, the audio to be processed is played according to the speaker gain set at the corresponding trajectory point.

8. The method according to claim 1, wherein Obtaining the audio to be processed and at least one trajectory point, and generating the initial audio data based on the audio to be processed and the trajectory point includes: In response to a first interaction operation performed on the corresponding application interface, obtaining the audio to be processed, where the first interaction operation is used to trigger the start of recording the audio to be processed; In response to the end of recording the audio to be processed, displaying a trajectory point generation interface, and obtaining the at least one trajectory point in the virtual playback scene in the trajectory point generation interface; Generating the initial audio data based on the audio to be processed and the trajectory point.

9. The method according to claim 8, characterized in that, In response to the end of recording the audio to be processed, displaying a trajectory point generation interface, and displaying the at least one trajectory point in the virtual playback scene in the trajectory point generation interface includes: In response to the end of recording the audio to be processed, determining the number of trajectory points based on the duration of the audio to be processed; Displaying the trajectory point generation interface, and displaying the trajectory points in the virtual playback scene in the trajectory point generation interface according to the number of trajectory points.

10. The method according to claim 8, characterized in that, Determining the speaker gain corresponding to the target virtual speaker based on the initial audio data includes: When the number of the trajectory points meets a preset number condition, determine a target position vector corresponding to each trajectory point based on the position information corresponding to each trajectory point, and determine the position indicated by each target position vector as the virtual source position corresponding to each trajectory point, where the target position vector is used to indicate the spatial position of the corresponding trajectory point in the virtual playback scene; Obtain the distances between each virtual speaker in the virtual playback scene and each virtual source position; Determine the speakers whose distances meet a preset distance condition as the target virtual speakers corresponding to each trajectory point, and determine the speaker position vectors corresponding to each target virtual speaker corresponding to each trajectory point; Based on the target position vector and the speaker position vector, determine the gain coefficient of each target virtual speaker corresponding to each trajectory point, where the speaker position vector is the same as the target position vector after weighted summation according to the gain coefficient, and the speaker gain includes the gain coefficient.

11. The method according to claim 8, wherein The method further includes: When the number of the trajectory points meets a preset number condition, obtain the playback time corresponding to each trajectory point; Increase trajectory points in an interpolation manner based on the playback time interval between a first trajectory point and a second trajectory point with adjacent playback times.

12. The method according to claim 1, characterized in that, The mixing the to-be-processed audio and the media information indicated by the source media information identifier based on the speaker gain to generate target audio data includes: Obtain the information type corresponding to the media information indicated by the source media information identifier; Determine a gain allocation ratio based on the information type; Mix the to-be-processed audio and the media information indicated by the source media information identifier according to the gain allocation ratio to generate target audio data.

13. An audio data processing device, characterized in that, including: An obtaining module, configured to obtain to-be-processed audio and at least one trajectory point, and generate initial audio data based on the to-be-processed audio and the trajectory point, where the initial audio data includes a source media information identifier corresponding to the to-be-processed audio and position information of the trajectory point; A determining module, configured to determine a speaker gain corresponding to a target virtual speaker based on the initial audio data, where the position information of the trajectory point is used to determine a virtual source position in a virtual playback scene, the virtual playback scene is deployed with virtual speakers, the target virtual speaker is at least one of the virtual speakers, the target virtual speaker represents a virtual speaker determined based on the virtual source position, and the speaker gain is used to simulate the playback of the to-be-processed audio at the virtual source position; A generating module, configured to mix the to-be-processed audio and the media information indicated by the source media information identifier based on the speaker gain to generate target audio data.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, where the computer program, when run by a processor, executes the method according to any one of claims 1 to 12.

15. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 12.

16. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the method described in any one of claims 1 to 12 through the computer program.

Citation Information

Cited By

  • Sound effect configuration method and device and electronic equipment

    CN120929041A

  • Distributed rendering node display picture synchronization method

    CN121070298A

  • Online automatic thematic map making method and system based on artificial intelligence

    CN121616699A