Virtual panoramic sound generation method and device, electronic equipment and storage medium

By acquiring the user's real head and location information, and matching a suitable head model and transfer function, the accuracy problem caused by individual user and scene differences in virtual panoramic sound generation is solved, achieving higher accuracy in panoramic sound generation.

CN122093732APending Publication Date: 2026-05-26SHENZHEN LONGWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN LONGWEI TECH CO LTD
Filing Date
2026-01-30
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing virtual panoramic sound generation methods, due to their use of a general head transfer function model, cannot effectively adapt to differences in individual user characteristics and scene differences, resulting in distortion of panoramic sound orientation perception and reduced generation accuracy.

Method used

By acquiring the target user's real head size and position information, matching a suitable head model, determining the spacing and deviation angle parameters, filtering the target head transfer function, constructing panoramic sound parameters, and generating virtual panoramic sound.

Benefits of technology

It improves the accuracy of virtual panoramic sound generation, reduces interference caused by individual user differences and scene differences, and ensures the adaptability and accuracy of panoramic sound playback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122093732A_ABST
    Figure CN122093732A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a virtual panoramic sound generation method and device, electronic equipment and a storage medium, and belongs to the technical field of audio processing. The method comprises the following steps: acquiring real head size information and head position information of a target user, and acquiring a target audio clip of audio playing equipment; determining a target head model of a head transfer function set matched with the real head size information based on the real head size information; based on the head position information, determining a distance parameter and a deviation angle parameter between the target user and the audio playing device; based on the real head size information, the spacing parameter and the deviation angle parameter, performing function matching on the head transfer function set to obtain a target head transfer function; based on the target audio clip, performing audio parameter adjustment on the target head transfer function to obtain a panoramic sound playback target sound field parameter; and performing virtual panoramic sound generation based on the panoramic sound playback target sound field parameters. According to the embodiment of the invention, the generation accuracy of the virtual panoramic sound can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to a virtual panoramic sound generation method and apparatus, electronic device and storage medium. Background Technology

[0002] Virtual panoramic sound generation is used to simulate auditory perception from different spatial directions. For example, by generating virtual panoramic sound for the sound of rain in movies and TV shows, viewers can be provided with a virtual scene experience of being on a rainy day, thereby improving their viewing experience.

[0003] Currently, common virtual panoramic sound generation methods all require the use of a general head transfer function model to adjust the time difference, volume difference, and phase difference of the audio playback device's output signal, thereby simulating the auditory perception of sound coming from different spatial directions. However, due to individual differences in user characteristics and scene differences, using a general head transfer function model can easily cause distortion in panoramic sound location perception, thus reducing the accuracy of virtual panoramic sound. Therefore, how to improve the generation accuracy of virtual panoramic sound has become an urgent technical problem to be solved. Summary of the Invention

[0004] The main objective of this application is to provide a virtual panoramic sound generation method, apparatus, electronic device, and storage medium, aiming to improve the accuracy of virtual panoramic sound generation.

[0005] To achieve the above objectives, a first aspect of this application proposes a virtual panoramic sound generation method, applied to an audio playback device, the method comprising:

[0006] Obtain the target user's actual head size and head position information, and obtain the target audio's channel information; Based on the real head size information, model matching is performed on the pre-built head model library to obtain the target head model, wherein the target head model contains a set of head transfer functions; Based on the head position information, the distance parameter and deviation angle parameter between the target user and the audio playback device are determined; Based on the actual head size information, the spacing parameters, and the deviation angle parameters, function matching is performed on the head transfer function set to obtain the target head transfer function; Based on the target head transfer function and the audio channel information, panoramic sound parameters are constructed to obtain the panoramic sound reproduction target sound field parameters. Based on the target sound field parameters of the panoramic sound playback, virtual panoramic sound is generated for the target audio.

[0007] In some embodiments, before performing model matching on a pre-built head model library based on the actual head size information to obtain the target head model, the method further includes: A human head model library is constructed based on a preset head size dataset, wherein the human head model library includes multiple human head models of different sizes; Based on the aforementioned head model, audio transmission parameters are measured to obtain a set of head transfer functions; Data binding is performed on the human head model and the set of head transfer functions to obtain the head model; Based on the aforementioned head model library, the head models are assembled to obtain the head model library.

[0008] In some embodiments, the head model library includes multiple head models, each head model containing head size information. The step of matching a pre-built head model library with the actual head size information to obtain a target head model includes: Based on the actual head size information and the model head size information, the head size deviation is calculated to obtain head size deviation data; Based on a preset head size weight allocation strategy, the head size deviation data is weighted to obtain head size deviation weight data. Based on the multiple head models, the head size deviation weight data are compared to obtain the target deviation weight data. Based on the target deviation weight data, a model is selected from the head model library to obtain the target head model.

[0009] In some embodiments, determining the distance parameter and deviation angle parameter between the target user and the audio playback device based on the head position information includes: The dimensions of the audio playback device are measured to obtain the playback device size parameters; Based on the size parameters of the playback device, a reference three-dimensional spatial coordinate system is constructed, wherein the reference three-dimensional spatial coordinate system includes the coordinates of the device center origin; Based on the reference three-dimensional spatial coordinate system and the head position information, the head coordinates of the target user are measured to obtain the head center coordinates; Based on the coordinates of the device center origin and the coordinates of the head center, the relative positions of the coordinate points are estimated to obtain the spacing parameter and the deviation angle parameter.

[0010] In some embodiments, the step of estimating the relative positions of coordinate points based on the coordinates of the device center origin and the coordinates of the head center to obtain the spacing parameter and the deviation angle parameter includes: Based on the coordinates of the device center origin and the coordinates of the head center, the distance between coordinate points is calculated to obtain the spacing parameter; Based on the coordinates of the device center origin and the reference three-dimensional spatial coordinate system, determine the reference plane of the head center without offset; Based on the head center coordinates and the unoffset head center reference plane, the head offset is calculated to obtain the head offset distance; Based on the head offset distance, the spacing parameter, and the reference plane of the head center without offset, the offset angle is calculated to obtain the deviation angle parameter.

[0011] In some embodiments, the step of performing function matching on the set of head transfer functions based on the actual head size information, the spacing parameter, and the deviation angle parameter to obtain the target head transfer function includes: Based on the actual head size information, the spacing parameters, and the deviation angle parameters, the transmission angle of the playback device is calculated to obtain the transmission angle of the speaker. Based on the speaker transmission angle, the head transmission function set is angle-matched to obtain the target head transmission function.

[0012] In some embodiments, the step of constructing panoramic sound parameters based on the target head transfer function and the audio channel information to obtain panoramic sound playback target sound field parameters includes: Based on the target head transfer function and the vocal tract information, an audio tract mapping function is generated; The parameters of the target sound field for panoramic sound reproduction are obtained by integrating the parameters of the audio channel mapping function.

[0013] To achieve the above objectives, a second aspect of this application provides a virtual panoramic sound generation apparatus, the apparatus comprising: The information acquisition module is used to acquire the target user's actual head size and head position information, and to acquire the target audio's channel information; The model matching module is used to perform model matching on a pre-built head model library based on the real head size information to obtain a target head model, wherein the target head model contains a set of head transfer functions; The parameter calculation module is used to determine the distance parameter and deviation angle parameter between the target user and the audio playback device based on the head position information; The function selection module is used to perform function matching on the set of head transfer functions based on the actual head size information, the spacing parameter and the deviation angle parameter to obtain the target head transfer function; The parameter adjustment module is used to construct panoramic sound parameters based on the target head transfer function and the channel information to obtain the panoramic sound reproduction target sound field parameters. The panoramic sound generation module is used to generate virtual panoramic sound from the target audio based on the target sound field parameters of the panoramic sound playback target.

[0014] To achieve the above objectives, a third aspect of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method of the first aspect described above.

[0015] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method of the first aspect described above.

[0016] The virtual panoramic sound generation method, apparatus, electronic device, and storage medium proposed in this application obtain the target user's real head size and head position information, and then match a target head model adapted to the target user from a preset head model library based on the real head size information. This ensures that the set of head transfer functions in the target head model fits the target user, avoiding the adaptation deviation between general head transfer functions and the target user, thereby improving the generation accuracy of virtual panoramic sound. Subsequently, based on the head position information, the distance parameters and deviation angle parameters between the target user and the audio playback device are determined. Then, combined with the target user's real head size information, distance parameters, and deviation angle parameters, a target head transfer function is selected from the set of head transfer functions. This ensures the adaptability of the transfer function to the panoramic sound playback application scenario, thereby further improving the generation accuracy of virtual panoramic sound. Finally, based on the target head transfer function and the channel information of the target audio, panoramic sound parameters are constructed to obtain the target sound field parameters for panoramic sound playback. Based on the target sound field parameters for panoramic sound playback, virtual panoramic sound is generated for the target audio, reducing interference factors in the target head transfer function, thereby improving the generation accuracy of virtual panoramic sound. Attached Figure Description

[0017] Figure 1 This is a flowchart of the virtual panoramic sound generation method provided in the embodiments of this application; Figure 2 This is a flowchart of a virtual panoramic sound generation method provided in another embodiment of this application; Figure 3 yes Figure 1 The flowchart of step S102 in the document; Figure 4 yes Figure 1 The flowchart of step S103 in the process; Figure 5 yes Figure 4 The flowchart of step S404 in the document; Figure 6 yes Figure 1 The flowchart of step S104 in the process; Figure 7 yes Figure 1 The flowchart of step S105 in the process; Figure 8 This is a schematic diagram of the virtual panoramic sound generation device provided in the embodiments of this application; Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0019] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0021] First, let's analyze some of the terms used in this application: Virtual surround sound: Virtual surround sound is a three-dimensional sound field restoration technology based on audio signal processing. It does not rely on multi-channel physical speaker arrays. Virtual surround sound accurately analyzes the original characteristics of the audio, the user's head physiological parameters and spatial position information, and combines the head transfer function to optimize key acoustic parameters such as time delay, volume difference, and spectrum difference. It simulates the propagation law of sound in three-dimensional space, allowing users to perceive the location of virtual sound sources from all directions, such as front, back, left, right, up and down, on ordinary playback devices such as headphones and stereo speakers. It constructs an enveloping sound field environment. Virtual surround sound can be widely used in film, music, games and other fields. It can significantly improve the spatial immersion and location realism of audio, making users feel as if they are in the center of the audio scene and have an immersive listening experience.

[0022] Head-Related Transfer Function (HRTF): The HRTF is a core function characterizing the changes in acoustic properties of sound as it travels from a sound source in space to the listener's ears. Essentially, the HRTF maps the signal differences between the two ears after the sound source signal has been reflected, diffracted, and filtered by physiological structures such as the head, auricle, and torso. The HRTF varies significantly depending on individual physiological characteristics such as head size, ear distance, and auricle shape, as well as the spatial position of the sound source relative to the listener. The core function of the HRTF is to quantify the differences in key acoustic parameters such as time delay, volume difference, and spectral difference between the received signals from both ears. These differences are the core basis for the human brain to perceive the spatial location of sound and construct a three-dimensional sound field. In technologies such as virtual surround sound and spatial audio, the HRTF is fundamental to achieving personalized, high-fidelity sound field reproduction. By matching the HRTF to the user's individual physiological characteristics and spatial position parameters, listeners can accurately perceive the three-dimensional spatial location of sound on ordinary playback devices.

[0023] Virtual panoramic sound generation is used to simulate auditory perception from different spatial directions. For example, by generating virtual panoramic sound for the sound of rain in movies and TV shows, viewers can be provided with a virtual scene experience of being on a rainy day, thereby improving their viewing experience.

[0024] Currently, common virtual panoramic sound generation methods all require the use of a general head transfer function model to adjust the time difference, volume difference, and phase difference of the audio playback device's output signal, thereby simulating the auditory perception of sound coming from different spatial directions. However, due to individual differences in user characteristics and scene differences, using a general head transfer function model can easily cause distortion in panoramic sound location perception, thus reducing the accuracy of virtual panoramic sound. Therefore, how to improve the generation accuracy of virtual panoramic sound has become an urgent technical problem to be solved.

[0025] Based on this, embodiments of this application provide a virtual panoramic sound generation method and apparatus, electronic device and storage medium, aiming to improve the accuracy of virtual panoramic sound generation.

[0026] The virtual panoramic sound generation method, apparatus, electronic device, and storage medium provided in this application are specifically described through the following embodiments. First, the virtual panoramic sound generation method in this application embodiment is described.

[0027] The virtual panoramic sound generation method provided in this application relates to the field of audio processing technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the virtual panoramic sound generation method, but is not limited to the above forms.

[0028] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0029] Figure 1 This is an optional flowchart of the virtual panoramic sound generation method provided in this application embodiment. The virtual panoramic sound generation method can be applied to devices with audio playback functions. Figure 1 The method may include, but is not limited to, steps S101 to S106.

[0030] Step S101: Obtain the target user's actual head size and head position information, and obtain the target audio's channel information; Step S102: Based on the real head size information, perform model matching on the pre-built head model library to obtain the target head model, wherein the target head model contains a set of head transfer functions; Step S103: Based on the head position information, determine the distance parameters and deviation angle parameters between the target user and the audio playback device; Step S104: Based on the actual head size information, spacing parameters and deviation angle parameters, perform function matching on the head transfer function set to obtain the target head transfer function; Step S105: Based on the target head transfer function and the channel information, construct the panoramic sound parameters to obtain the panoramic sound playback target sound field parameters.

[0031] Step S106: Based on the target sound field parameters of the panoramic sound playback, generate virtual panoramic sound for the target audio.

[0032] Steps S101 to S106 of this embodiment involve obtaining the target user's actual head size and position information, and then matching a target head model adapted to the target user from a preset head model library based on the actual head size information. This ensures that the set of head transfer functions in the target head model matches the target user, avoiding the adaptation deviation between the general head transfer function and the target user, thereby improving the accuracy of virtual panoramic sound generation. Subsequently, based on the head position information, the distance parameters and deviation angle parameters between the target user and the audio playback device are determined. Then, combined with the target user's actual head size information, distance parameters, and deviation angle parameters, a target head transfer function is selected from the set of head transfer functions. This ensures the adaptability of the transfer function to the panoramic sound playback scenario, further improving the accuracy of virtual panoramic sound generation. Finally, based on the target head transfer function and the channel information of the target audio, panoramic sound parameters are constructed to obtain the target sound field parameters for panoramic sound playback. Based on the target sound field parameters for panoramic sound playback, virtual panoramic sound is generated for the target audio, reducing interference factors in the target head transfer function and thus improving the accuracy of virtual panoramic sound generation.

[0033] In step S101 of some embodiments, the audio playback device refers to a device that has audio playback function and can generate virtual panoramic sound, such as a mobile phone or tablet computer. It should be noted that the audio playback device usually only has 2 channels.

[0034] The target user refers to an individual who experiences virtual immersive sound. For example, in the scenario of watching an immersive sound movie on a mobile phone, the target user can be a consumer who is holding a mobile phone and watching the movie.

[0035] Actual head size information refers to the physical head size data of the target user. Actual head size information usually includes the horizontal width between the cheekbones on both sides of the head, the horizontal width between the tips of the ears on both sides of the head, the horizontal width between one ear and the eyebrows, the horizontal width between the back of the head and the eyebrows, the vertical length between the top of the head and one ear, and the vertical length between the top of the head and the tip of the chin.

[0036] Head position information refers to data related to the spatial position of the target user's head relative to the audio playback device.

[0037] Target audio refers to the audio content that needs to be processed into virtual surround sound and output through an audio playback device. For example, in the scenario of watching a surround sound movie, the target audio can be background sound effects, character dialogues, background music, and other audio materials from the movie.

[0038] Channel information refers to the basic attributes of the vocal tract, which typically includes the number of channels and the definition of channel orientation.

[0039] In this embodiment, the audio playback device can acquire multi-view image sequences of the target user's head from the front and sides using a front-facing camera. Further, an image recognition algorithm is used to locate key features such as the head contour vertices, ear positions, and facial markers in the multi-view image sequence. Combined with imaging parameters such as focal length and pixel ratio from the front-facing camera, the pixel distances in the multi-view image sequence are converted into actual physical dimensions. This yields real head size information such as the lateral width between the cheekbones on both sides of the head, the lateral width between the ear tips on both sides of the head, the lateral width between one ear and the brow, the lateral width between the back of the brain and the brow, the vertical length between the top of the head and one ear, and the vertical length between the top of the head and the chin tip. Simultaneously, during the acquisition of the multi-view image sequence, the front-facing camera can calculate the distance between the target user's head and the audio playback device using a distance sensor, thereby determining the target user's head position information. Furthermore, by extracting the audio content currently to be played by the audio playback device, the target audio can be obtained. Further, by analyzing the target audio signal, the number of channels and the channel orientation definitions corresponding to the target audio can be obtained.

[0040] In step S102 of some embodiments, the head model library refers to a database stored in an audio playback device that contains head models with multiple different head size information.

[0041] The target head model refers to the model selected from the head model library that best matches the actual head size information of the target user.

[0042] The head transfer function set refers to the set of acoustic parameters bound to the head model that describe the sound transmitted from the speaker of the audio playback device to the ears of the target user. The head transfer function set usually contains a set of parameters for the time delay, spectral difference, and phase difference of the sound received by the target head model from the audio playback device at different angles.

[0043] It is important to understand that virtual surround sound is often related to the user's head size. For example, if user A has a distance of 20cm between their ears and user B has a distance of 18cm between their ears, when listening to the same audio clip played by the same audio playback device, the difference in ear distance will cause differences in the time and decibels at which users A and B receive the sound. This can easily lead to user A's virtual location judgment of the virtual surround sound generated by the audio playback device differing from user B's virtual location judgment of the virtual surround sound generated by the audio playback device. Therefore, this application embodiment can ensure that users with different head sizes have corresponding head transfer functions at different angles by designing head models of different sizes and measuring the head transfer function at different angles for each head model of different sizes.

[0044] For details, please refer to Figure 2 In some embodiments, before step S102 is implemented, the virtual panoramic sound generation method may also include, but is not limited to, steps S201 to S204: Step S201: Based on the preset head size dataset, construct a human head model library, wherein the human head model library includes multiple human head models of different sizes; Step S202: Based on the human head model, measure the audio transfer parameters to obtain the head transfer function set; Step S203: Data binding is performed on the head model and the head transfer function set to obtain the head model; Step S204: Based on the head model library, the head models are assembled to obtain the head model library.

[0045] In step S201 of some embodiments, the head size dataset refers to a dataset containing multiple combinations of key head dimensions. For example, in the scenario of building a head model library, the head size dataset can be a standardized dataset containing different combinations of key parameters such as head width, ear distance, and head height.

[0046] A head model library refers to a basic database containing multiple head models of different sizes. A head model is a head simulation model built based on the size information in a head size dataset.

[0047] In this embodiment, a head size dataset containing multiple combinations of key head sizes can be compiled by surveying head size characteristic data of different populations. Furthermore, 3D modeling software such as Blender can be used to build corresponding 3D simulation models, i.e., human head models, based on each set of head size parameters in the head size dataset. Finally, all the completed human head models of different sizes are summarized and stored in the same database to form a human head model library.

[0048] In step S202 of some embodiments, the head transfer function set refers to the combination of audio transfer parameters measured for a single head model, covering the target user's head at different angles.

[0049] In this embodiment, a speaker with performance matching that of an audio playback device can be placed directly in front of the head model, and test audio can be played through the speaker to collect the original sound signal received by the head model. Then, the same speaker can be placed at different angles, such as the left and right sides of the same head model, and the same test audio can be played through the speaker to collect the test sound signal received by the head model at this time. Furthermore, by comparing the audio transmission parameters such as time delay, spectral attenuation, and phase shift between the original sound signal and the test sound signal, and finally, the audio transmission parameters of the head model under all measurement angles can be classified and integrated according to the angle to form a head transfer function set specific to the head model.

[0050] In step S203 of some embodiments, the head model refers to a model that includes head size information and a set of head transfer functions.

[0051] In this embodiment, a unique identifier such as an ID number can be assigned to each head model, and the head size data corresponding to the head model can be recorded. Furthermore, an association mapping can be established between the unique identifier of the head model and the corresponding head transfer function set, thereby realizing the binding and storage of the unique identifier, head size data and head transfer function set, and thus forming a head model.

[0052] In step S204 of some embodiments, the head models that have completed data binding are stored in a database in the manner of a human head model library to form a head model library. It should be noted that, in order to improve the indexing efficiency of the head models in the head model library, the head models that have completed data binding can be classified and sorted and a retrieval index can be established before the head model library is formed.

[0053] Steps S201 to S204 of this application embodiment involve building a multi-size head model library based on a pre-acquired head size dataset. This ensures that the head models in the library cover various individual differences among users, providing a model foundation for personalized matching of target users. Then, audio transfer parameters are measured for each head model to obtain its corresponding set of head transfer functions. This ensures that the set of head transfer functions covers head transfer functions at different angles, thereby improving the fit between the head transfer functions and the user's head. Finally, the head models and the set of head transfer functions are bound together to form a head model. Based on the head model library, a model set is created for the head models to obtain the head model library, which improves the matching efficiency of the head transfer functions.

[0054] In this embodiment of the application, after obtaining the head model library, head size deviation data can be calculated by comparing the real head size information with the model head size information of each head model. Then, weights are assigned to the deviation data according to a preset weight allocation strategy to obtain head size deviation weight data. Subsequently, the deviation weight data of multiple head models are compared and the target deviation weight data is selected. Finally, the target head model that best matches the user's real head size is selected from the head model library based on the target deviation weight data.

[0055] For details, please refer to Figure 3 In some embodiments, the head model library includes multiple head models, each containing head size information. Step S102 may include, but is not limited to, steps S301 to S304: Step S301: Based on the actual head size information and the model head size information, calculate the head size deviation to obtain head size deviation data; Step S302: Based on the preset head size weight allocation strategy, the head size deviation data is weighted to obtain head size deviation weight data. Step S303: Based on multiple head models, compare the head size deviation weight data to obtain the target deviation weight data; Step S304: Based on the target deviation weight data, select a model from the head model library to obtain the target head model.

[0056] In step S301 of some embodiments, the model head size information refers to the standard size data of each head model in the preset head model library. The model head size information usually includes the horizontal width between the cheekbones on both sides of the head, the horizontal width between the ear tips on both sides of the head, the horizontal width between one ear and the eyebrow, the horizontal width between the back of the brain and the eyebrow, the vertical length between the top of the head and one ear, and the vertical length between the top of the head and the chin tip.

[0057] Head size deviation data refers to the difference between the actual head size information of the target user and the model head size information of each head model. For example, the actual head size information of the target user includes a horizontal width between the cheekbones on both sides of the head of 18cm, and the model head size information of head model A includes a horizontal width between the cheekbones on both sides of the head of 17.5cm. Then the head size deviation data can include the horizontal width deviation between the cheekbones on both sides of the head of 18cm-17.5cm=0.5cm.

[0058] In this embodiment, the actual head size information is compared with the model head size information of each head model in the same dimension. This allows for the calculation of the difference between the actual head size information and the model head size information of each head model under the same dimension, i.e., head size deviation data. For example, the actual head size information a of target user A, the model head size information b of head model B, and the model head size information c of head model C all include dimensional information such as the horizontal width between the cheekbones on both sides of the head and the horizontal width between the ear tips on both sides of the head. Specifically, in actual head size information a, the horizontal width between the cheekbones on both sides of the head is 18cm and the horizontal width between the ear tips on both sides of the head is 18.4cm; in model head size information b, the horizontal width between the cheekbones on both sides of the head is 19cm and the horizontal width between the ear tips on both sides of the head is 19.6cm; and in model head size information c, the horizontal width between the cheekbones on both sides of the head is 17cm and the horizontal width between the ear tips on both sides of the head is 17cm. The lateral width between the ear tips on both sides of the head is 17.3cm. Therefore, the head size deviation data between the actual head size information a and the model head size information b includes the lateral width deviation between the cheekbones on both sides of the head, which is 18cm-19cm=-1cm, and the lateral width deviation between the ear tips on both sides of the head, which is 18.4cm-19.6cm=-1.2cm. The head size deviation data between the actual head size information a and the model head size information c includes the lateral width deviation between the cheekbones on both sides of the head, which is 18cm-17cm=1cm, and the lateral width deviation between the ear tips on both sides of the head, which is 18.4cm-17.3cm=1.1cm.

[0059] In step S302 of some embodiments, the head size weighting strategy refers to the weighting rules used to distinguish the importance of different head size parameters. For example, among head size parameters such as the horizontal width between the cheekbones on both sides of the head, the horizontal width between the ear tips on both sides of the head, the horizontal width between one ear and the eyebrow, the horizontal width between the back of the head and the eyebrow, the vertical length between the top of the head and one ear, and the vertical length between the top of the head and the tip of the chin, the head size weighting strategy can be 15% for the horizontal width between the cheekbones on both sides of the head, 25% for the horizontal width between the ear tips on both sides of the head, 20% for the horizontal width between one ear and the eyebrow, 15% for the horizontal width between the back of the head and the eyebrow, 15% for the vertical length between the top of the head and one ear, and 10% for the vertical length between the top of the head and the tip of the chin.

[0060] Head size deviation weighting data refers to the comprehensive deviation data obtained by weighting head size deviation data according to a weighting strategy. For example, when the head size weighting strategy is as follows: 15% for the horizontal width between the cheekbones on both sides of the head, 25% for the horizontal width between the ear tips on both sides of the head, 20% for the horizontal width between one ear and the glabella, 15% for the horizontal width between the back of the head and the glabella, 15% for the vertical length between the top of the head and one ear, and 10% for the vertical length between the top of the head and the tip of the chin, and the head size deviation weighting data is calculated by weighting the head size deviation data between the cheekbones on both sides of the head, the deviation ... When the lateral width deviation between the cheekbones is 0.5cm, the lateral width deviation between the tips of the ears on both sides of the head is 1cm, the lateral width deviation between one ear and the space between the eyebrows is 0.3cm, the lateral width deviation between the back of the head and the space between the eyebrows is 0.6cm, the vertical length deviation between the top of the head and one ear is 0.8cm, and the vertical length deviation between the top of the head and the tip of the chin is 1.2cm, the weighting data for head size deviation can be 0.5×15%+1×25%+0.3×20%+0.6×15%+0.8×15%+1.2×10%=0.715.

[0061] In this embodiment of the application, before matching the target head model adapted to the target user from the head model library based on the actual head size information, the size information weights can be set according to the degree of influence of different dimensions of size information on the head transfer function. For example, if the head size information only includes two dimensions: the horizontal width between the cheekbones on both sides of the head and the horizontal width between the ear tips on both sides of the head, and the influence of the horizontal width between the cheekbones on the head transfer function is less than the influence of the horizontal width between the ear tips on the head transfer function, then the weight of the horizontal width between the cheekbones on both sides of the head can be set to 45%, and the weight of the horizontal width between the ear tips on both sides of the head can be set to 55%.

[0062] Furthermore, in this embodiment of the application, after determining the head size weight allocation strategy, the head size deviation data of different dimensions can be multiplied with the weights specified in the head size weight allocation strategy to obtain the head size weight product data of different dimensions. Finally, the head size deviation weight data of the target user and different head models can be obtained by summing the head size weight product data of different dimensions.

[0063] In step S303 of some embodiments, the target deviation weight data refers to the smallest data selected from the head size deviation weight data of all head models.

[0064] In this embodiment of the application, by repeating the above steps S301 and S302, the head size deviation weight data corresponding to each head model can be calculated. Furthermore, by sorting the data or comparing the data one by one, the data with the smallest value can be located from the multiple head size deviation weight data. Finally, the head size deviation weight data with the smallest value is extracted from the multiple head size deviation weight data to obtain the target deviation weight data.

[0065] In step S304 of some embodiments, the target head model can be identified by searching the head model library for the head model corresponding to the target deviation weight data.

[0066] Steps S301 to S304 of this embodiment involve comparing the actual head size information of the target user with the model head size information of each head model in the head model library one by one to obtain head size deviation data. Then, according to a pre-set head size weight allocation strategy, the head size deviation data is weighted to obtain head size deviation weight data, which provides a quantitative basis for head model adaptation. Furthermore, based on multiple head models, the head size deviation weight data is compared to obtain target deviation weight data. Based on the target deviation weight data, models are selected from the head model library to ensure the compatibility between the selected target head model and the target user, thereby improving the accuracy of virtual panoramic sound generation.

[0067] In step S103 of some embodiments, the spacing parameter refers to the spatial straight-line distance data between the target user and the audio playback device.

[0068] The deviation angle parameter refers to the offset angle data of the target user's head center relative to the center of the audio playback device.

[0069] In this embodiment, the size parameters of the audio playback device can be obtained by measuring the size of the audio playback device. Then, a reference three-dimensional spatial coordinate system containing the coordinates of the device's center origin can be constructed based on the size parameters. Subsequently, by combining the coordinate system of the device's center origin with the head position information, the head coordinates of the target user can be measured to obtain the head center coordinates. Finally, by estimating the relative position of the device's center origin coordinates and the head center coordinates, the distance parameters and deviation angle parameters between the target user and the audio playback device can be obtained.

[0070] For details, please refer to Figure 4 In some embodiments, the head model library includes multiple head models, each containing head size information. Step S103 may include, but is not limited to, steps S401 to S404: Step S401: Measure the dimensions of the audio playback device to obtain the playback device dimension parameters; Step S402: Based on the size parameters of the playback device, construct a reference three-dimensional spatial coordinate system, wherein the reference three-dimensional spatial coordinate system includes the coordinates of the device center origin; Step S403: Based on the reference three-dimensional spatial coordinate system and head position information, measure the head coordinates of the target user to obtain the head center coordinates; Step S404: Based on the coordinates of the equipment center origin and the head center, estimate the relative position of the coordinate points to obtain the spacing parameters and deviation angle parameters.

[0071] In step S401 of some embodiments, the playback device size parameters refer to the physical size-related data of the audio playback device. The playback device size parameters include, but are not limited to, the length, width, and spacing between the two speakers of the audio playback device.

[0072] In this embodiment, the size parameters of the audio playback device can be determined by retrieving data from the module storing hardware parameters in the audio playback device. Alternatively, high-precision measuring tools such as electronic calipers can be used to physically measure the audio playback device to obtain its length, width, thickness, as well as the installation coordinates of the dual speakers, speaker spacing, and other dimensional data. The above dimensional data can then be summarized into the size parameters of the playback device.

[0073] In step S402 of some embodiments, the reference three-dimensional spatial coordinate system refers to a three-dimensional reference system used to locate the spatial position of the target user and the audio playback device. Generally speaking, the reference three-dimensional spatial coordinate system is a three-dimensional coordinate system constructed with the center of the audio playback device as the origin, the width direction of the device as the x-axis, the height direction of the device as the y-axis, and the thickness direction of the device as the z-axis. It should be noted that the origin coordinates (0, 0, 0) of the reference three-dimensional spatial coordinate system constructed in the above state are the origin coordinates of the device center.

[0074] In this embodiment, a three-dimensional rectangular coordinate system can be established with the physical center of the audio playback device as the origin. The x-axis of this three-dimensional rectangular coordinate system is along the width direction of the audio playback device, the y-axis is along the height direction of the audio playback device, and the z-axis is along the thickness direction of the audio playback device. The origin coordinates are (0, 0, 0). Further, the scale of the constructed three-dimensional rectangular coordinate system is set according to the size parameters of the playback device. Finally, the three-dimensional rectangular coordinate system with the determined scale is used as the reference three-dimensional spatial coordinate system, and the coordinates of the physical center of the audio playback device in the reference three-dimensional spatial coordinate system are used as the origin coordinates of the device center.

[0075] In step S403 of some embodiments, the head center coordinates refer to the coordinate values ​​corresponding to the head center of the target user in the reference three-dimensional space coordinate system.

[0076] In this embodiment of the application, the spatial position of the target user's head in the reference three-dimensional spatial coordinate system can be clearly determined through the above-mentioned head position information. Furthermore, based on the spatial position, the coordinates of the key feature points of the target user's head in the reference three-dimensional spatial coordinate system can be calculated. The key feature points can be eyes, nose, and mouth, etc. Finally, based on the spatial coordinates of the eyes, nose, and mouth, the coordinates of the geometric center of the target user's head in the reference three-dimensional spatial coordinate system, i.e., the head center coordinates, can be calculated.

[0077] In step S404 of some embodiments, the spatial straight-line distance between the audio playback device and the target user's head center can be calculated based on the device center origin coordinates and the head center coordinates to obtain the spacing parameter. Then, based on the device center origin coordinates and the reference three-dimensional spatial coordinate system, the non-offset head center reference plane can be determined. Subsequently, by combining the head center coordinates and the non-offset head center reference plane, the distance of the target user's head offset relative to the non-offset head center reference plane can be calculated, i.e., the head offset distance. Finally, by converting the head offset distance, the spacing parameter, and the non-offset head center reference plane into offset angles, the deviation angle parameter of the target user's head can be calculated.

[0078] For details, please refer to Figure 5 In some embodiments, the head model library includes multiple head models, each containing head size information. Step S404 may include, but is not limited to, steps S501 to S504: Step S501: Based on the coordinates of the equipment center origin and the head center, calculate the distance between coordinate points to obtain the spacing parameters; Step S502: Based on the coordinates of the equipment center origin and the reference three-dimensional space coordinate system, determine the reference plane of the head center without offset; Step S503: Based on the head center coordinates and the unoffset head center reference plane, calculate the head offset to obtain the head offset distance. Step S504: Based on the head offset distance, spacing parameters, and the reference plane of the head center without offset, the offset angle is calculated to obtain the deviation angle parameters.

[0079] In step S501 of some embodiments, the differences between the coordinates of the device center origin and the head center on each axis can be calculated to obtain the x-axis coordinate difference, y-axis coordinate difference, and z-axis coordinate difference. Further, the x-axis coordinate difference, y-axis coordinate difference, and z-axis coordinate difference are squared and summed to obtain the coordinate sum. Finally, the square root of the coordinate sum is taken to obtain the spatial straight-line distance between the target user and the audio playback device, i.e., the spacing parameter. For example, when the coordinates of the device center origin are (0, 0, 0) and the coordinates of the head center are (8, 4, 50), the spacing parameter can be √[(8-0)²+(4-0)²+(50-0)²]≈50.8cm.

[0080] In step S502 of some embodiments, the non-offset head center reference plane refers to the reference plane where the head center of the target user is located when the head has no horizontal offset, determined based on the coordinates of the device center origin and the reference three-dimensional spatial coordinate system.

[0081] In this embodiment, the coordinate axes of the reference three-dimensional spatial coordinate system can be used as a reference. Combined with the coordinates of the device center origin, a reference coordinate axis (usually the x-axis, i.e. the device width direction) related to the horizontal offset direction of the target user's head can be selected. A target plane perpendicular to the reference coordinate axis and passing through the device center origin can be constructed. At this time, the target plane is the head center reference plane without offset.

[0082] In step S503 of some embodiments, the head offset distance refers to the distance that the target user's head center is offset relative to the reference plane of the unoffset head center.

[0083] In this embodiment of the application, the head offset distance can be obtained by calculating the vertical distance between the head center point corresponding to the head center coordinates and the reference plane of the head center without offset.

[0084] In step S504 of some embodiments, the vertical projection parameter can be obtained by calculating the distance of the spacing parameter in the straight direction of the non-offset head center reference plane. Furthermore, the angle of the target user's head offset relative to the non-offset head center reference plane, i.e., the deviation angle parameter, can be calculated by using trigonometric functions.

[0085] Steps S501 to S504 of this embodiment calculate the spatial distance between the coordinates of the device center origin and the head center, ensuring the accuracy of the obtained spacing parameters. Then, a non-offset head center reference plane is established based on the device center origin coordinates and the reference three-dimensional spatial coordinate system, forming a unified and stable offset reference, thus avoiding the problem of reference ambiguity in offset judgment. Then, the head offset distance is calculated by combining the head center coordinates and the non-offset head center reference plane, which can accurately quantify the degree of head offset of the target user. Finally, the offset angle is converted by the head offset distance, spacing parameters and reference plane, improving the calculation accuracy of the deviation angle parameters, thereby improving the scene adaptation accuracy of virtual panoramic sound generation.

[0086] Steps S401 to S404, as illustrated in this embodiment, involve measuring the dimensions of the audio playback device to obtain its size parameters, providing a data foundation for establishing a coordinate system. Then, using the center of the audio playback device as the coordinate system's far point, the coordinate system scale is determined based on the device's size parameters, thereby constructing a reference three-dimensional spatial coordinate system. This standardizes the calculation of spatial parameters, improving the calculation efficiency of spacing and deviation angle parameters. Next, based on head position information, the coordinates of the target user's head center in the reference three-dimensional spatial coordinate system—the head center coordinates—can be accurately measured, achieving a quantitative representation of the user's head spatial position. Finally, based on the device's center origin coordinates and the head center coordinates, the relative positions of the coordinate points are estimated, improving the accuracy and calculation efficiency of the spacing and deviation angle parameters.

[0087] In step S104 of some embodiments, the target head transfer function refers to the acoustic transfer rule selected from the set of head transfer functions that adapts to the actual head size information, spacing parameters and deviation angle parameters.

[0088] In this embodiment, the angle at which the speaker of the audio playback device transmits sound signals to the target user can be calculated by combining real head size information, spacing parameters, and deviation angle parameters. Furthermore, a transfer function matching the angle, i.e., the target head transfer function, can be selected from the head transfer function set of the above-mentioned head model.

[0089] For details, please refer to Figure 6 In some embodiments, step S104 may include, but is not limited to, steps S601 to S602: Step S601: Based on the actual head size information, spacing parameters and deviation angle parameters, calculate the transmission angle of the playback device to obtain the speaker transmission angle; Step S602: Based on the speaker transmission angle, perform angle matching on the head transfer function set to obtain the target head transfer function.

[0090] In step S601 of some embodiments, the speaker transmission angle refers to the angle between the acoustic transmission path from the speaker of the audio playback device to the target user's ear and the direction of the vertical line of the device's width center.

[0091] In this embodiment, the ear coordinates of the target user's ears in the reference three-dimensional spatial coordinate system and the speaker coordinates of the audio playback device's two speakers in the reference three-dimensional spatial coordinate system can be determined based on the actual head size information, spacing parameters, and deviation angle parameters. The ear coordinates include the left ear coordinates and the right ear coordinates, and the speaker coordinates include the left speaker coordinates and the right speaker coordinates. Further, acoustic transmission paths are constructed between the corresponding points of the left speaker coordinates and the corresponding points of the left and right ear coordinates, and acoustic transmission paths are constructed between the corresponding points of the right speaker coordinates and the corresponding points of the left and right ear coordinates. Then, the angle between each transmission path and the vertical line of the device width center is calculated to obtain multiple sets of speaker transmission angles.

[0092] In step S602 of some embodiments, by traversing the set of head transfer functions corresponding to the target head model, the angle parameter represented by each head transfer function in the set of head transfer functions can be identified. Furthermore, by finding the head transfer function that has the same transfer angle as the speaker, the target head transfer function can be extracted from the set of head transfer functions.

[0093] Steps S601 to S602, as illustrated in this embodiment, combine real head size information, spacing parameters, and deviation angle parameters to calculate the transmission angle of the playback device. This can convert the key features and spatial position parameters of the target user's head into quantitative data of the acoustic transmission angle, thereby ensuring that the obtained speaker transmission angle fits the actual acoustic propagation scenario. Furthermore, based on the speaker transmission angle, angle matching is performed on the head transmission function set to obtain the target head transmission function. This can avoid parameter adaptation deviation problems caused by angle differences and improve the generation accuracy of virtual panoramic sound.

[0094] In step S105 of some embodiments, the target sound field parameter for panoramic sound reproduction refers to the sound field feature parameters used to generate virtual panoramic sound after the audio parameters have been adjusted. For example, in the scenario of generating a virtual overhead sound source, the target sound field parameter for panoramic sound reproduction can be an adaptation parameter that describes the time difference, volume difference, and spectral difference between the two ears.

[0095] In this embodiment, key acoustic parameters such as time delay, volume difference, and spectrum difference in the target head transfer function can be adjusted according to the original audio signal characteristics represented by the target audio segment and the virtual sound source orientation requirements. This can counteract the transmission path interference from the speaker of the audio playback device to the target user's ears and make the adjusted parameters meet the virtual sound source orientation requirements of the virtual panoramic sound, thus enabling panoramic sound playback of the target sound field parameters.

[0096] For details, please refer to Figure 7 In some embodiments, step S105 may include, but is not limited to, steps S701 to S702: Step S701: Generate an audio channel mapping function based on the target head transfer function and channel information; Step S702: Perform parameter integration on the audio channel mapping function to obtain the target sound field parameters for immersive sound playback.

[0097] In step S701 of some embodiments, the audio channel mapping function refers to a set of dedicated acoustic mapping parameters adapted to a single target audio channel. The audio channel mapping function can be used to describe the acoustic parameter rules of the audio signal of that channel propagating from the virtual location to the user's ears.

[0098] In this embodiment, based on the general acoustic propagation law in the target head transfer function, personalized parameters can be calculated for the specific orientation and signal characteristics of each channel in the target audio according to the channel information. For example, for the left front 20° channel, combined with the basic parameters of the general acoustic propagation law of "sound propagation at 20° orientation" in the target head transfer function, the time delay from this channel to the user's left ear is adjusted to 0.12ms and the volume difference is +2dB. Furthermore, by combining these calculated parameters, the audio channel mapping function corresponding to the left front 20° channel can be formed.

[0099] In step S702 of some embodiments, the audio channel mapping functions of all channels of the target audio are summarized, and the core acoustic parameters in each function are extracted to form an initial set containing all channel parameters. Then, the core acoustic parameter benchmark is unified. Specifically, the time delay parameters of all channels of the target audio can be corrected with the center of the user's head as the reference point. For example, the delay of each channel can be uniformly corrected to "the absolute propagation time from the speaker of the audio playback device to the user's ears". Alternatively, the average volume of the center channel can be used as the anchor point to align the volume difference parameters of all channels. Then, all the core acoustic parameters with the unified benchmark are integrated together, and the channel orientation labels corresponding to each core acoustic parameter are marked to obtain the target sound field parameters for panoramic sound reproduction.

[0100] Steps S701 to S702, as illustrated in this embodiment, generate audio channel mapping functions specific to each channel based on the target head transfer function and channel information. This enables deep integration of the user's individual acoustic propagation patterns with the channel spatial orientation and signal characteristics, avoiding positioning deviations caused by channel parameters deviating from the user's physiological characteristics and channel orientation. Then, the parameters of each audio channel mapping function are integrated, and the integrated parameters are standardized to obtain the target sound field parameters for panoramic sound reproduction. This eliminates the orientation conflicts and auditory fragmentation problems between multi-channel parameters, resulting in a better panoramic sound experience for the user.

[0101] In step S106 of some embodiments, the target sound field parameters for panoramic sound reproduction can be converted into the output signal parameters of the dual speakers of the audio playback device according to the mapping relationship between the sound field parameters and the speaker output signal. Further, the audio playback device drives the dual speakers to play the target audio segment according to the output signal parameters, so that when the sound reaches the target user's ears through the transmission path, a difference signal that conforms to the target sound field parameters for panoramic sound reproduction is formed. Finally, the target user's brain can perceive that the sound from the audio playback device comes from a virtual location based on the difference signal received by the ears, thereby realizing the generation of virtual panoramic sound.

[0102] This application obtains the target user's actual head size and position information, and then matches a target head model adapted to the target user from a pre-set head model library based on the actual head size information. This ensures that the set of head transfer functions in the target head model fits the target user, avoiding the adaptation deviation between general head transfer functions and the target user, thereby improving the accuracy of virtual panoramic sound generation. Subsequently, based on the head position information, the distance parameters and deviation angle parameters between the target user and the audio playback device are determined. Combining the target user's actual head size information, distance parameters, and deviation angle parameters, a target head transfer function is selected from the set of head transfer functions. This ensures the adaptability of the transfer function to the panoramic sound playback application scenario, thereby further improving the accuracy of virtual panoramic sound generation. Finally, based on the target head transfer function and the channel information of the target audio, panoramic sound parameters are constructed to obtain the target sound field parameters for panoramic sound playback. Based on the target sound field parameters for panoramic sound playback, virtual panoramic sound is generated for the target audio, reducing interference factors in the target head transfer function and thus improving the accuracy of virtual panoramic sound generation.

[0103] Please see Figure 8 This application also provides a virtual panoramic sound generation device that can implement the above-described virtual panoramic sound generation method. The device includes: The information acquisition module 801 is used to acquire the target user's real head size information and head position information, and to acquire the target audio channel information; The model matching module 802 is used to perform model matching on a pre-built head model library based on real head size information to obtain a target head model, wherein the target head model contains a set of head transfer functions; The parameter calculation module 803 is used to determine the distance parameters and deviation angle parameters between the target user and the audio playback device based on the head position information. The function selection module 804 is used to perform function matching on the set of head transfer functions based on the real head size information, spacing parameters and deviation angle parameters to obtain the target head transfer function. The parameter adjustment module 805 is used to construct panoramic sound parameters based on the target head transfer function and the channel information to obtain the panoramic sound playback target sound field parameters. The panoramic sound generation module 806 is used to generate virtual panoramic sound from the target audio based on the target sound field parameters of the panoramic sound playback target.

[0104] The specific implementation of the virtual panoramic sound generation device is basically the same as the specific embodiment of the virtual panoramic sound generation method described above, and will not be repeated here.

[0105] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described virtual panoramic sound generation method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0106] Please see Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 902 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 902 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called and executed by the processor 901 using the virtual panoramic sound generation method of the embodiments of this application. The input / output interface 903 is used to implement information input and output; The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904); The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.

[0107] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described virtual panoramic sound generation method.

[0108] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0109] The virtual panoramic sound generation method, virtual panoramic sound generation device, electronic device, and storage medium provided in this application embodiment obtain the target user's real head size information and head position information, obtain the target audio's channel information, and perform model matching on a pre-built head model library based on the real head size information to obtain a target head model. The target head model includes a set of head transfer functions. Based on the head position information, the distance parameters and deviation angle parameters between the target user and the audio playback device are determined. Based on the real head size information, distance parameters, and deviation angle parameters, the set of head transfer functions is matched to obtain a target head transfer function. Based on the target head transfer function and channel information, panoramic sound parameters are constructed to obtain panoramic sound playback target sound field parameters. Based on the panoramic sound playback target sound field parameters, virtual panoramic sound is generated for the target audio.

[0110] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0111] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0112] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0113] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0114] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0115] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0116] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. The coupling or direct coupling or communication connection between the shown or discussed units may be through some interfaces, or indirect coupling or communication connection between the apparatus or units, and may be electrical, mechanical, or other forms.

[0117] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0118] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0119] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0120] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A method for generating virtual panoramic sound, applied to an audio playback device, characterized in that, The method includes: Obtain the target user's actual head size and head position information, and obtain the target audio's channel information; Based on the real head size information, model matching is performed on the pre-built head model library to obtain the target head model, wherein the target head model contains a set of head transfer functions; Based on the head position information, the distance parameter and deviation angle parameter between the target user and the audio playback device are determined; Based on the actual head size information, the spacing parameters, and the deviation angle parameters, function matching is performed on the head transfer function set to obtain the target head transfer function; Based on the target head transfer function and the audio channel information, panoramic sound parameters are constructed to obtain the panoramic sound reproduction target sound field parameters. Based on the target sound field parameters of the panoramic sound playback, virtual panoramic sound is generated for the target audio.

2. The method according to claim 1, characterized in that, Before performing model matching on a pre-built head model library based on the actual head size information to obtain the target head model, the process further includes: A human head model library is constructed based on a preset head size dataset, wherein the human head model library includes multiple human head models of different sizes; Based on the aforementioned head model, audio transmission parameters are measured to obtain a set of head transfer functions; Data binding is performed on the human head model and the set of head transfer functions to obtain the head model; Based on the aforementioned head model library, the head models are assembled to obtain the head model library.

3. The method according to claim 2, characterized in that, The head model library includes multiple head models, and each head model contains head size information. The step of matching a pre-built head model library with the actual head size information to obtain a target head model includes: Based on the actual head size information and the model head size information, the head size deviation is calculated to obtain head size deviation data; Based on a preset head size weight allocation strategy, the head size deviation data is weighted to obtain head size deviation weight data. Based on the multiple head models, the head size deviation weight data are compared to obtain the target deviation weight data. Based on the target deviation weight data, a model is selected from the head model library to obtain the target head model.

4. The method according to claim 1, characterized in that, The step of determining the distance parameters and deviation angle parameters between the target user and the audio playback device based on the head position information includes: The dimensions of the audio playback device are measured to obtain the playback device size parameters; Based on the size parameters of the playback device, a reference three-dimensional spatial coordinate system is constructed, wherein the reference three-dimensional spatial coordinate system includes the coordinates of the device center origin; Based on the reference three-dimensional spatial coordinate system and the head position information, the head coordinates of the target user are measured to obtain the head center coordinates; Based on the coordinates of the device center origin and the coordinates of the head center, the relative positions of the coordinate points are estimated to obtain the spacing parameter and the deviation angle parameter.

5. The method according to claim 4, characterized in that, The step of estimating the relative positions of coordinate points based on the coordinates of the device's center origin and the head's center to obtain the spacing parameter and the deviation angle parameter includes: Based on the coordinates of the device center origin and the coordinates of the head center, the distance between coordinate points is calculated to obtain the spacing parameter; Based on the coordinates of the device center origin and the reference three-dimensional spatial coordinate system, determine the reference plane of the head center without offset; Based on the head center coordinates and the unoffset head center reference plane, the head offset is calculated to obtain the head offset distance; Based on the head offset distance, the spacing parameter, and the reference plane of the head center without offset, the offset angle is calculated to obtain the deviation angle parameter.

6. The method according to any one of claims 1-5, characterized in that, The step of performing function matching on the set of head transfer functions based on the actual head size information, the spacing parameters, and the deviation angle parameters to obtain the target head transfer function includes: Based on the actual head size information, the spacing parameters, and the deviation angle parameters, the transmission angle of the playback device is calculated to obtain the transmission angle of the speaker. Based on the speaker transmission angle, the head transmission function set is angle-matched to obtain the target head transmission function.

7. The method according to any one of claims 1-5, characterized in that, The step of constructing panoramic sound parameters based on the target head transfer function and the audio channel information to obtain the panoramic sound reproduction target sound field parameters includes: Based on the target head transfer function and the vocal tract information, an audio tract mapping function is generated; The parameters of the target sound field for panoramic sound reproduction are obtained by integrating the parameters of the audio channel mapping function.

8. A virtual panoramic sound generation device, characterized in that, The device includes: The information acquisition module is used to acquire the target user's actual head size and head position information, and to acquire the target audio's channel information; The model matching module is used to perform model matching on a pre-built head model library based on the real head size information to obtain a target head model, wherein the target head model contains a set of head transfer functions; The parameter calculation module is used to determine the distance parameter and deviation angle parameter between the target user and the audio playback device based on the head position information; The function selection module is used to perform function matching on the set of head transfer functions based on the actual head size information, the spacing parameter and the deviation angle parameter to obtain the target head transfer function; The parameter adjustment module is used to construct panoramic sound parameters based on the target head transfer function and the channel information to obtain the panoramic sound reproduction target sound field parameters. The panoramic sound generation module is used to generate virtual panoramic sound from the target audio based on the target sound field parameters of the panoramic sound playback target.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the virtual panoramic sound generation method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the virtual panoramic sound generation method according to any one of claims 1 to 7.