A virtual reality-based method for obtaining a spatial direction parameter of a sound production device
By selecting associated sound source signals from reference sound source data in virtual reality, generating a set of sound direction parameters and performing parameter analysis, the problem of inaccurate sound direction parameters in virtual reality is solved, achieving higher recognition accuracy and reducing data processing volume.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU MOVIE POWER TECH CO LTD
- Filing Date
- 2022-10-25
- Publication Date
- 2026-05-12
AI Technical Summary
In virtual reality, interference factors exist when sound is transmitted in different directions, resulting in inaccurate sound direction parameters.
By obtaining the audio corresponding to the sound source signal to be processed from the reference sound source data, the associated sound source signals are selected based on key descriptions, a set of pronunciation directions is generated, and parameter analysis is performed, thereby reducing the amount of sound source data processing and improving the accuracy of parameter recognition.
It achieves accurate acquisition of pronunciation direction parameters while reducing data processing workload, thus improving recognition accuracy.
Smart Images

Figure CN115691550B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data acquisition technology, and more specifically, to a method for acquiring spatial orientation parameters of a sound-emitting device based on virtual reality. Background Technology
[0002] Virtual reality (VR) technology is a computer simulation system that creates and allows users to experience virtual worlds. It uses computers to generate a simulated environment, immersing users in that environment. VR technology utilizes real-world data, generates electronic signals through computer technology, and combines these signals with various output devices to transform them into phenomena that people can perceive. These phenomena can be real objects or substances invisible to the naked eye, represented through three-dimensional models.
[0003] In real life, sound travels in different directions, giving the sound a sense of depth and layering. However, the transmission of sound through these different directions is subject to various interference factors, leading to inaccurate parameters regarding the direction of sound transmission. Therefore, a technical solution is urgently needed to address these problems. Summary of the Invention
[0004] To address the technical problems existing in related technologies, this application provides a method for obtaining spatial orientation parameters of a sound-producing device based on virtual reality.
[0005] In a first aspect, a method for obtaining spatial direction parameters of a sound-emitting device based on virtual reality is provided. The method includes at least: obtaining a voiceprint dataset corresponding to the audio of at least one sound source signal to be processed from reference sound source data; selecting reference associated sound source signals associated with the reference sound source signal from the at least one sound source signal to be processed based on a first key description of the audio of the reference mined sound source signal and a second key description of the audio of each voiceprint dataset to be processed; obtaining a reference pronunciation direction set corresponding to the pronunciation direction of the reference associated sound source signal from the reference sound source data based on the voiceprint dataset to be processed of the reference associated sound source signal; and performing parameter analysis on the reference pronunciation direction set to determine the pronunciation direction parameters of the reference associated sound source signal.
[0006] In one independently implemented embodiment, obtaining the voiceprint dataset corresponding to the audio of at least one sound source signal to be processed from the reference sound source data includes: identifying important information in the reference sound source data to generate reference important nodes corresponding to the audio of each sound source signal to be processed in the reference sound source data; and generating the voiceprint dataset corresponding to the sound source signal to be processed based on the reference important nodes corresponding to each sound source signal to be processed in the reference sound source data.
[0007] In one independently implemented embodiment, the step of selecting a reference associated sound source signal from the not less than one sound source signal to be processed, based on the first sound source data key description corresponding to the audio of the reference mined sound source signal and the second sound source data key description corresponding to each voiceprint dataset to be processed, includes: associating the first sound source data key description corresponding to the audio of the reference mined sound source signal with the second sound source data key description corresponding to each voiceprint dataset to be processed, and determining the sound source signal to be processed corresponding to the second sound source data key description associated with the first sound source data key description as the reference associated sound source signal.
[0008] In one standalone embodiment, before obtaining the audio-text dataset corresponding to at least one audio source signal to be processed from the reference audio source data, the method further includes: obtaining mining audio source data; and combining the mining audio source data to generate a reference mining audio source signal and a first audio source data key description corresponding to the audio of the reference mining audio source signal in the mining audio source data.
[0009] In one standalone embodiment, generating a reference excavation sound source signal by combining the excavation sound source data includes: identifying important information in the excavation sound source data to generate reference important nodes corresponding to the audio of each original sound source signal in at least one original sound source signal included in the excavation sound source data; and selecting and determining the reference excavation sound source signal from the at least one original sound source signal based on the important node offset information of the reference important nodes corresponding to the audio of each original sound source signal.
[0010] In one standalone embodiment, the process of obtaining a reference pronunciation direction set corresponding to the pronunciation direction of the reference associated sound source signal from the reference sound source data based on the voiceprint dataset to be processed includes: generating a second mapping relationship data between the voiceprint dataset to be processed and the reference pronunciation direction set by combining the first mapping relationship data between the audio and the pronunciation direction; and obtaining the reference pronunciation direction set corresponding to the pronunciation direction of the reference associated sound source signal from the reference sound source data by combining the second mapping relationship data and the voiceprint dataset to be processed.
[0011] In one standalone embodiment, the method further includes: generating a new reference mining sound source signal based on the important node offset information of the reference important nodes corresponding to the audio of each sound source signal to be processed, after determining that there is no second sound source data key description associated with the first sound source data key description.
[0012] In one independently implemented embodiment, the step of performing parameter analysis on the reference sound direction set to determine the sound direction parameters of the reference associated sound source signal includes: combining the reference sound direction set, reading the sound source feature data corresponding to the reference sound source set from the reference sound source data; obtaining the sound source vibration description and sound source azimuth description of the sound source feature data through a pre-configured artificial intelligence thread; obtaining the sound source vibration description corresponding one-to-one with the feature indicators that differ in the sound source feature data; wherein, among the several feature indicators, there are similar first feature indicators and second feature indicators, and the first feature indicator is smaller than the second feature indicator; the sound source vibration description corresponding to the first feature indicator is obtained by combining the sound source vibration description corresponding to the second feature indicator and the sound source azimuth description of the sound source vibration description corresponding to the second feature indicator; the sound source azimuth description of the sound source vibration description corresponding to the second feature indicator is obtained through the pre-configured artificial intelligence thread; and performing parameter analysis on the sound source feature data based on the sound source vibration description corresponding one-to-one with the feature indicators that differ to determine the sound direction parameters of the reference associated sound source signal.
[0013] In one independently implemented embodiment, the parameter analysis of the sound source feature data based on the sound source vibration description corresponding to the distinguishing feature indicators includes: for each distinguishing feature indicator, generating a first sound source analysis result of the sound source feature data under that feature indicator based on the sound source vibration description corresponding to that feature indicator; combining the first sound source analysis results of the sound source feature data under each feature indicator to generate the probability that each sound source in the sound source feature data is located as a sound source corresponding to the sounding direction; and combining the probability that each sound source in the sound source feature data is located as a sound source corresponding to the sounding direction with a specified mining probability vector to perform parameter analysis on the sound source feature data.
[0014] In one independently implemented embodiment, the step of combining the first sound source analysis results of the sound source feature data under each feature indication to generate the probability that each sound source in the sound source feature data is located as a sound source corresponding to the direction of sound production includes: after performing multiple rounds of splicing processing according to the feature indications with differences, determining the probability that each sound source in the sound source feature data is located as a sound source corresponding to the direction of sound production; wherein, the m-th round of splicing processing in the multiple rounds of splicing processing includes: generating first bias information of the first sound source analysis result under the first feature indication; splicing the first sound source analysis result under the first feature indication and the first sound source analysis result under the second feature indication using the first bias information of the first sound source analysis result under the first feature indication to determine the reference sound source analysis result under the second feature indication; and optimizing the reference sound source analysis result to the first sound source analysis result of the first feature indication in the (m+1)-th round of splicing.
[0015] In one standalone embodiment, after determining the pronunciation direction parameter of the reference associated sound source signal, the method further includes: configuring the analysis variable for sound source localization corresponding to the pronunciation direction in the reference sound source data as a first reference variable in conjunction with the pronunciation direction parameter; and configuring the analysis variable for sound source localization other than the pronunciation direction in the reference sound source data as a second reference variable.
[0016] This application provides a method for obtaining spatial direction parameters of a virtual reality-based sound source device. Based on a first key description of the audio data corresponding to a reference-mined sound source signal and second key descriptions of the voiceprint dataset corresponding to each sound source signal to be processed, the method can accurately select second key descriptions associated with the first key description from several second key descriptions. This means it can accurately select reference associated sound source signals related to the reference-mined sound source signal, achieving precise mining of the reference-mined sound source signal. Then, using the selected voiceprint dataset, a reference sound direction set relative to the sound direction is generated. Therefore, only parameter analysis needs to be performed on the reference sound direction set, without processing the remaining data in the reference sound source data. This reduces the workload of sound source data processing while ensuring the accuracy of the processed sound source data, thus reducing the amount of data processing work while minimizing the number of sound source data segments to be processed. Furthermore, since only the sound source data of the reference sound direction set needs processing, it helps improve the accuracy of parameter recognition for sound direction identification, resulting in more accurate sound direction parameters. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating a method for obtaining spatial orientation parameters of a virtual reality-based sound-producing device, as provided in an embodiment of this application. Detailed Implementation
[0019] To better understand the above technical solutions, the technical solutions of this application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of this application and the specific features in the embodiments are detailed descriptions of the technical solutions of this application, rather than limitations on the technical solutions of this application. In the absence of conflict, the embodiments of this application and the technical features in the embodiments can be combined with each other.
[0020] Please see Figure 1 This paper illustrates a method for obtaining spatial orientation parameters of a sound-producing device based on virtual reality. This method may include the technical solutions described in steps S101-S104.
[0021] S101: Obtain the audio corresponding to the audio of at least one audio source signal to be processed from the reference sound source data.
[0022] The reference sound source data can be one sound source data in the pronunciation segment, and can include no less than one sound source signal to be processed.
[0023] The sound source signal to be processed can be a sound source signal appearing in the reference sound source data, or it can appear in the remaining sound source data in the corresponding pronunciation segment of the reference sound source data. The sound source signal to be processed can be human sounds or sounds emitted by devices in the reference sound source data, etc.
[0024] To obtain reference sound source data, one approach is to determine the currently processed sound source data as the reference sound source data after obtaining the pronunciation segment. Alternatively, one can specify the reference sound source data to be obtained.
[0025] For example, after obtaining the reference sound source data, the reference sound source data can be processed to generate the sound source signals to be processed included in the reference sound source data. Then, for each sound source signal to be processed, the audio of the sound source signal to be processed and the location of the audio in the reference sound source data can be generated. Thus, based on the location of the audio of the sound source signal to be processed in the reference sound source data, the audio of the sound source signal to be processed can be generated as a voiceprint dataset to be processed.
[0026] S102: Based on the key description of the first sound source data corresponding to the audio of the reference mined sound source signal and the key description of the second sound source data corresponding to each voiceprint dataset to be processed, select a reference associated sound source signal that is related to the reference mined sound source signal from at least one sound source signal to be processed.
[0027] Furthermore, the reference source signal is the source signal that needs to be identified for the direction of sound production. The key description of the first source data corresponding to the audio of the reference source signal can be cached before the source data is processed, or it can be obtained based on the source data processed earlier during the source signal mining and source data identification process. No specific limitation is made here.
[0028] The reference associated sound source signal is the sound source signal to be processed that is the same as the reference excavation sound source signal. The key description of the sound source data may include important node information, sound wave information, sound wave vibration information, etc.
[0029] In this embodiment, for each of the at least one sound source signals to be processed, after determining the voiceprint dataset corresponding to the audio of the sound source signal to be processed, a key description of the second sound source data corresponding to the voiceprint dataset to be processed can be determined. Based on this, a key description of the second sound source data corresponding to each of the at least one sound source signals to be processed can be determined.
[0030] Then, based on the key description of the second source data corresponding to each source signal to be processed and the key description of the first source data corresponding to the audio of the reference source signal, the key description of the second source data that matches the audio of the key description of the first source data can be selected. Thus, based on the key description of the second source data, the source signal to be processed corresponding to it can be determined as the reference associated source signal.
[0031] S103: Based on the voiceprint dataset to be processed from the reference associated sound source signal, obtain the reference pronunciation direction set corresponding to the pronunciation direction of the reference associated sound source signal from the reference sound source data.
[0032] For example, based on the matching between the audio and the direction of sound production in the voiceprint dataset to be processed, the localization relationship and audio relationship between the voiceprint dataset to be processed and the reference direction of sound production can be generated. Then, based on the voiceprint dataset to be processed, the obtained localization relationship and audio relationship, the reference direction of sound production corresponding to the direction of sound production of the reference associated sound source signal can be determined from the reference sound source data.
[0033] In one possible implementation, after determining the reference pronunciation direction set, the sound source data corresponding to the reference pronunciation direction set can be read from the reference sound source data to obtain the sound source feature data corresponding to the reference pronunciation direction set.
[0034] S104: Perform parameter analysis on the reference sound direction set to obtain the sound direction parameters of the reference associated sound source signal.
[0035] In this embodiment, after determining the reference sound direction set, parameter analysis can be performed on the reference sound direction set to generate the first information for locating each sound source in the reference sound direction set.
[0036] Therefore, based on the first information of the sound source localization corresponding to the sound direction and the first information of each sound source localization, the sound source localization belonging to the sound direction in the reference sound direction set can be generated. Then, the sound source localization belonging to the sound direction can be used to identify the reference sound direction set to obtain the sound direction parameters of the reference associated sound source signal.
[0037] In addition, if the sound source feature data corresponding to the reference sound direction set is determined, the sound source feature data can be directly analyzed to obtain the sound direction parameters of the reference associated sound source signal.
[0038] Furthermore, based on the key descriptions of the first sound source data corresponding to the audio of the reference mined sound source signal and the key descriptions of the second sound source data corresponding to the voiceprint datasets to be processed for each sound source signal to be processed, it is possible to accurately select the second sound source data key descriptions associated with the first sound source data key description from several second sound source data key descriptions. That is, it is possible to accurately select the reference associated sound source signals associated with the reference mined sound source signal, thus achieving accurate mining of the reference mined sound source signal. Then, through the selected voiceprint dataset to be processed, a reference pronunciation direction set relative to the pronunciation direction is generated. Therefore, only parameter analysis needs to be performed on the reference pronunciation direction set, without processing the remaining data in the reference sound source data. This reduces the workload of sound source data processing and ensures the accuracy of the sound source data to be processed. It reduces the workload of data processing while reducing the number of sound source data segments to be processed. Moreover, since only the sound source data of the reference pronunciation direction set needs to be processed, it is beneficial to improve the accuracy of parameter recognition of pronunciation direction, thus enabling more accurate pronunciation direction parameters.
[0039] In an alternative embodiment, for S101, a method for determining a voiceprint dataset to be processed provided by the embodiments of this disclosure may specifically include the following.
[0040] S201: Identify important information in the reference sound source data and generate reference important nodes corresponding to the audio of each sound source signal to be processed in at least one sound source signal to be processed included in the reference sound source data.
[0041] Furthermore, the important information identification thread can be used to identify important information in the reference sound source data in order to determine the reference important nodes corresponding to the audio in the reference sound source data.
[0042] For example, reference sound source data can be loaded into a pre-configured important information recognition thread. This thread processes the reference sound source data to generate all important nodes within it. Then, based on the location relationships between these important reference nodes, reference important nodes for audio signals belonging to the same source signal can be generated. This allows us to determine the reference important nodes corresponding to the audio signals of at least one source signal included in the reference sound source data.
[0043] S202: For each of the sound source signals to be processed in at least one sound source signal to be processed, generate a voiceprint dataset corresponding to the sound source signal to be processed based on the reference important node corresponding to the sound source signal to be processed.
[0044] For example, for each sound source signal to be processed, based on the reference importance node corresponding to that sound source signal, the range of sound source data determined by the reference importance node can be determined. Therefore, the obtained range of sound source data can be defined as the voiceprint dataset to be processed corresponding to that sound source signal.
[0045] Preferably, the voiceprint dataset corresponding to each voice source signal to be processed in the reference voice source data can be determined. In this way, because the configured important information recognition thread carries high-performance computing power, it can output accurate important node information. Therefore, based on the obtained important node information, the voiceprint dataset corresponding to each voice source signal to be processed can be accurately determined.
[0046] In an alternative embodiment, the key descriptions of the audio source data of the reference mining sound source signal are necessarily the same or similar across different sound source data. Therefore, for S102, after determining the second key descriptions of the sound source data corresponding to each sound source signal to be processed, the first key description of the audio source data corresponding to the reference mining sound source signal can be obtained. Then, the first key descriptions of the sound source data corresponding to the reference mining sound source signal are associated with each of the second key descriptions of the sound source data to generate a query to determine whether there are associated key descriptions of the sound source data. That is, to determine whether there are second key descriptions of the sound source data that match the audio source corresponding to the first key description of the sound source data. If so, the sound source data to be processed corresponding to the obtained associated second key descriptions of the sound source data can be determined as reference associated sound source data associated with the reference mining sound source signal. In other words, the reference mining sound source signal is selected.
[0047] In an alternative embodiment, after determining the second key description of the sound source data corresponding to a sound source signal to be processed, it can be associated with the first key description of the sound source data to generate a result indicating whether it is associated with the first key description of the sound source data. If so, the sound source data to be processed corresponding to the second key description of the sound source data can be directly determined as the reference associated sound source data; if not, after determining the second key description of the sound source data corresponding to the next sound source signal to be processed, the step of associating with the first key description of the sound source data can be returned. In this way, if the first key description of the sound source data obtained first is an associated key description of the sound source data, the reference associated sound source signal can be directly determined without having to determine the second key descriptions of the remaining sound source signals to be processed, thus improving the association efficiency.
[0048] In an alternative embodiment, before determining the voiceprint dataset to be processed in the reference sound source data, it is also necessary to obtain the mining sound source data corresponding to the reference mining sound source signal. Then, based on the mining sound source data, a reference mining sound source signal and a corresponding sound source data key description of the reference mining sound source signal's audio in the mining sound source data can be generated, and this key description is then determined as the first sound source data key description for association. In this way, based on the obtained first sound source data key description, accurate mining of the reference mining sound source signal can be achieved.
[0049] For example, the reference excavation sound source signal can be determined by following these steps.
[0050] (1) Identify important information in the excavated sound source data and generate reference important nodes corresponding to the audio of each original sound source signal in at least one original sound source signal included in the excavated sound source data.
[0051] Furthermore, the original sound source signal can be the sound source signal included in the mined sound source data.
[0052] In this embodiment, after obtaining the excavation sound source data, the important information of the excavation sound source data can be identified through a pre-configured important information identification thread to generate reference important nodes corresponding to the audio of each original sound source signal in at least one original sound source signal included in the excavation sound source data.
[0053] Furthermore, in parallel with determining the reference important nodes corresponding to the audio, the important information recognition thread can also output the important node offset information of the reference important nodes corresponding to each original sound source signal.
[0054] (2) Based on the important node offset information of the reference important nodes corresponding to the audio of each original sound source signal, a reference mining sound source signal is selected from no less than one original sound source signal.
[0055] For example, based on the bias degree of the important node corresponding to each original sound source signal, the reference important node with the best bias degree can be selected. Then, based on the reference important node, its corresponding voiceprint dataset can be generated. Thus, the original sound source signal corresponding to the voiceprint dataset can be determined, and the original sound source signal can be identified as the reference mining sound source signal.
[0056] In an alternative embodiment, after obtaining the excavation sound source data, the excavation range specified by the excavation sound source data can be determined. Then, important information can be directly identified in the specified excavation range to generate the original sound source signal of that range, and the original sound source signal can be determined as the reference excavation sound source signal.
[0057] In addition, after determining the reference important node with the best bias or the reference important node corresponding to the specified mining range, it is also necessary to determine the first sound source data key description corresponding to the audio of the reference important node in the mining sound source data or reference sound source data, so as to facilitate the subsequent sound source signal mining based on the first sound source data key description.
[0058] In an alternative embodiment, regarding S102, if it is determined that there is no second sound source data key description associated with the first sound source data key description, it can be concluded that the reference sound source signal to be mined is not in the reference sound source data. Therefore, based on the important node bias information of the reference important nodes corresponding to the audio of each sound source signal to be processed, the sound source signal corresponding to the reference important node with the highest bias can be generated and identified as the new reference sound source signal to be mined. In parallel, the sound source data key description of the voiceprint dataset corresponding to the new reference sound source signal can be determined, identified as the new first sound source data key description, and cached. In this way, the mining of the new reference sound source signal can be achieved.
[0059] Alternatively, in an alternative embodiment, if it is determined that there is no second sound source data key description associated with the first sound source data key description, the identification of the current reference sound source data can be abandoned, the remaining sound source data in the pronunciation segment can be obtained, and the step of determining the reference associated sound source data can be returned.
[0060] In an alternative embodiment, for S103: the reference pronunciation direction set corresponding to the pronunciation direction can be determined according to the following steps.
[0061] (1) Based on the first mapping relationship data between audio and pronunciation direction, generate the second mapping relationship data between the voiceprint dataset to be processed and the reference pronunciation direction set.
[0062] Furthermore, the first mapping relationship data is used to represent the localization relationship and amplitude relationship between audio and pronunciation direction. The second mapping relationship data is used to represent the association between the voiceprint dataset to be processed and the reference pronunciation direction set.
[0063] For example, based on the first mapping relationship data between audio and pronunciation direction, the relative positioning relationship and amplitude relationship between audio and pronunciation direction in the reference sound source data can be generated. Then, based on the obtained relative positioning relationship, amplitude relationship, and the audio corresponding to the voiceprint dataset to be processed, the association relationship between the reference pronunciation direction set and the voiceprint dataset to be processed can be generated. That is, the second mapping relationship data between the voiceprint dataset to be processed and the reference pronunciation direction set can be generated.
[0064] (2) Based on the second mapping relationship data and the voiceprint dataset to be processed, obtain the reference pronunciation direction set corresponding to the pronunciation direction of the reference associated sound source signal from the reference sound source data.
[0065] In this embodiment, the voiceprint dataset to be processed can be used as a reference, and the association relationship corresponding to the obtained second mapping relationship data can be used to convert the voiceprint dataset to be processed into a reference pronunciation direction set in the reference sound source data. That is, the reference pronunciation direction set corresponding to the pronunciation direction is obtained.
[0066] In an alternative embodiment, for S104, parameter analysis is performed on the reference pronunciation direction set to obtain the pronunciation direction parameters of the reference associated sound source signal. This is a method for parameter analysis of the reference pronunciation direction set provided by the embodiments of this disclosure, which may specifically include the following contents.
[0067] S501: Based on the reference pronunciation direction set, read the sound source feature data corresponding to the reference pronunciation direction set from the reference sound source data.
[0068] Furthermore, after determining the reference pronunciation direction set, the local sound source data corresponding to the reference pronunciation direction set can be read from the reference sound source data based on the location of the reference pronunciation direction set in the reference sound source data, and generated as the sound source feature data corresponding to the reference pronunciation direction set.
[0069] S502: Obtain sound source vibration description and sound source location description from the pre-configured artificial intelligence thread.
[0070] For example, after determining the reference sound direction set, the sound source feature data corresponding to the reference sound direction set can be loaded into the pre-configured artificial intelligence thread. Then, the artificial intelligence thread can obtain the sound source vibration description of the sound source feature data, and at the same time, it can also obtain the sound source orientation description of the sound source feature data.
[0071] S503: Obtain the sound source vibration description corresponding one-to-one with the feature indicators that differ from the sound source feature data.
[0072] Among the several feature indicators, there are similar first feature indicators and second feature indicators, where the first feature indicator is smaller than the second feature indicator.
[0073] The sound source vibration description corresponding to the first feature indication is obtained based on the sound source vibration description corresponding to the second feature indication and the sound source orientation description of the sound source vibration description corresponding to the second feature indication.
[0074] The second feature indicates the sound source location description corresponding to the sound source vibration description, which is obtained by the pre-configured artificial intelligence thread.
[0075] In this embodiment, the step of obtaining the sound source vibration description corresponding one-to-one with the feature indicators that differ from the sound source feature data is executed by a pre-configured artificial intelligence thread. The pre-configured artificial intelligence thread includes feature acquisition units corresponding to several feature indicators; each feature acquisition unit can acquire the sound source vibration description under its corresponding feature indicator. Based on the several feature acquisition units, the sound source vibration description corresponding one-to-one with the feature indicators that differ from the sound source feature data can be obtained respectively.
[0076] The feature indicator can be the sound source data recognition degree, and the sound source feature data corresponding to the reference sound direction set has the original sound source data recognition degree. Specifically, the original sound source data recognition degree can also be the sound source data recognition degree possessed by the reference sound source data.
[0077] For example, firstly, a pre-configured artificial intelligence thread can be used to obtain the original sound source data, identifying the sound source vibration description and sound source location description corresponding to the sound source feature data. Then, the recognition rate of this original sound source data is determined as a second feature indicator. Based on the sound source vibration description and sound source location description corresponding to this second feature indicator, a sound source vibration description corresponding to the first feature indicator is generated. Furthermore, while determining the sound source vibration description corresponding to the first feature indicator, the sound source location description of this vibration description can also be determined.
[0078] Then, the first feature indication can be determined as the new second feature indication, and a sound source vibration description and a sound source orientation description corresponding to the next first feature indication that is smaller than the new second feature indication can be generated. Thus, it is possible to obtain, one by one, the sound source vibration description and the sound source orientation description from the sound source vibration description that correspond to the feature indications that differ from the sound source feature data. Among these, the feature indication identified from the original sound source data is the optimal feature indication.
[0079] S504: Based on the characteristic indications of the differences, the sound source vibration description is matched one by one to perform parameter analysis on the sound source feature data to obtain the sound direction parameters of the reference associated sound source signal.
[0080] Furthermore, based on the range corresponding to the obtained pronunciation direction, it is possible to perform parameter analysis on the sound source feature data and obtain the pronunciation direction parameters of the reference associated sound source signal.
[0081] In an alternative embodiment, the sound source location description of the sound source vibration description corresponding to the second feature indication includes the positioning relationship between the first feature indications in the sound source vibration description corresponding to the second feature indication. A method for determining the sound source vibration description corresponding to the first feature indication provided by an embodiment of this disclosure may specifically include the following:
[0082] S601: For each of the second feature indications in the first feature indication, based on the positioning information of the second feature indication, select the first reference feature indication corresponding to the second feature indication from the first feature indications corresponding to the second feature indication.
[0083] Here, the sound source vibration description corresponding to each feature indication includes a different number of feature indications. The number of second feature indications in the sound source vibration description corresponding to the first feature indication is less than the number of first feature indications in the sound source vibration description corresponding to the second feature indication. That is, the number of feature indications corresponding to low sound source data recognition is less than the number of feature indications corresponding to high sound source data recognition.
[0084] Each second feature indication corresponding to the first feature indication has a first feature indication that corresponds to the second feature indication.
[0085] In this embodiment, for each second feature indication in the first feature indication, the positioning information of that second feature indication can be determined, and for each first feature indication in the second feature indication, the positioning information of that first feature indication can also be determined. Then, based on the positioning information of each second feature indication in the first feature indication and each first feature indication in the second feature indication, first feature indications and second feature indications with the same positioning can be determined from the first feature indication and the second feature indication respectively, and the aforementioned second feature indications are determined as the selected first reference feature indications corresponding to the aforementioned first feature indications. That is, a first reference feature indication corresponding to each second feature indication in the first feature indication can be generated from the first feature indications corresponding to the second feature indications.
[0086] S602: Based on the first feature indication and the second feature indication, generate a reference number of the second feature indication in the first feature indication corresponding to the first feature indication in the second feature indication.
[0087] Furthermore, the sound source vibration description of one of the second feature indicators in the first feature indication can be determined based on the sound source vibration description of several first feature indicators in the second feature indication.
[0088] For example, based on the association between the first feature indication and the second feature indication, a reference number can be generated in which one second feature indication in the first feature indication corresponds to one first feature indication in the second feature indication.
[0089] S603: Based on the positioning relationship between the first feature indicators and the positioning information of the first reference feature indicators, select a reference number of second reference feature indicators from the first feature indicators corresponding to the second feature indicators.
[0090] In this embodiment, based on the sound source orientation description corresponding to the second feature indication, the positioning relationship between the first feature indications in the second feature indication can be determined. Then, for each obtained first reference feature indication, a reference number of first feature indications can be selected from the first feature indications corresponding to the second feature indication according to the positioning information of the first reference feature indication and the positioning relationship between the first feature indications, and generated as the second reference feature indication.
[0091] For example, based on the positioning information of the first reference feature indication and the positioning relationship between the first feature indication and the first feature indication, the first feature indication with a specified number of reference differences that are different from the first reference feature indication can be selected from the first feature indications corresponding to the second feature indication and determined as the second reference feature indication.
[0092] S604: Based on the sound source vibration description indicated by the second reference feature, generate the sound source vibration description indicated by the second feature, and based on the sound source vibration descriptions of each of the second feature indicators in the obtained first feature indication, generate the sound source vibration description corresponding to the first feature indication.
[0093] Furthermore, based on the sound source vibration description of each of the second reference feature indications in the obtained number of references, the sound source vibration description of the second feature indication in the first feature indication corresponding to the second reference feature indication can be determined.
[0094] Therefore, based on the above steps, the sound source vibration description of each of the second feature indications in the first feature indication can be determined, and based on the sound source vibration description of each of the second feature indications, the sound source vibration description corresponding to the first feature indication can be determined.
[0095] In this way, it is possible to obtain the sound source vibration descriptions corresponding one by one to the feature indicators that differ from the reference sound direction set in the sound source feature data.
[0096] After obtaining the sound source vibration description corresponding to the different feature indicators, the method for parametric analysis of a reference sound direction set provided in this embodiment of the disclosure may include the following:
[0097] S701: For each feature indication that has differences, based on the sound source vibration description corresponding to the feature indication, generate the first sound source analysis result of the sound source feature data under the feature indication.
[0098] S702: Based on the first sound source analysis results under each feature indication of the sound source feature data, generate the probability of locating each sound source in the sound source feature data as the sound source location corresponding to the direction of sound production.
[0099] Here, after obtaining the first sound source analysis results under each feature indication, multiple rounds of splicing processing can be performed according to the feature indications that have differences. After that, the probability of each sound source in the sound source feature data being located as the sound source corresponding to the direction of sound production can be obtained.
[0100] S703: Based on the probability of locating each sound source in the sound source feature data as the direction of sound production and the specified mining probability vector, perform parameter analysis on the sound source feature data.
[0101] For example, the probability of locating each sound source in the sound source feature data as a sound source location corresponding to the direction of sound production can be compared with a specified mining probability vector. If the probability is greater than the specified mining probability vector, the sound source location is determined to be a sound source location corresponding to the direction of sound production. If the probability is not greater than the specified mining probability vector, the sound source location is generated as a sound source location that is not a sound source location corresponding to the direction of sound production.
[0102] Therefore, it is possible to determine the location of sound sources belonging to the direction of sound production and the location of sound sources not belonging to the direction of sound production in the sound source feature data, and based on the obtained results, to complete the parameter analysis of the sound source feature data and obtain the sound production direction parameters.
[0103] For example, the pronunciation direction parameter can be the pronunciation direction segmentation sound source data corresponding to the pronunciation direction.
[0104] In an alternative embodiment, for the m-th round of splicing in a multi-round splicing process, a splicing process is provided according to the embodiments of this disclosure, which may specifically include the following steps.
[0105] S801: Determine the first bias information of the first sound source analysis result under the first feature indication.
[0106] S802: Using the first bias information of the first sound source analysis result under the first feature indication, the first sound source analysis result under the first feature indication and the first sound source analysis result under the second feature indication are spliced together to obtain the reference sound source analysis result under the second feature indication.
[0107] In this embodiment, after obtaining the first bias information of the first sound source analysis result under the first feature indication, the splicing structure can splice the first sound source analysis result under the first feature indication and the first sound source analysis result under the second feature indication based on the first bias information to obtain the reference sound source analysis result under the second feature indication.
[0108] For example, based on the first bias information of the first sound source analysis results for each sound source localization, a first bias of the first sound source analysis results for each sound source localization is generated. The first bias of the first sound source analysis results for each sound source localization is compared with a specified bias target value. If the first bias is not less than the specified bias target value, the first sound source analysis result for the sound source localization corresponding to that first bias is determined as the first sound source analysis result under the second feature indication. If the first bias is less than the specified bias target value, the first sound source analysis result for the sound source localization corresponding to that first bias under the second feature indication is determined as the reference sound source analysis result.
[0109] S803: Optimize the reference sound source analysis result into the first sound source analysis result indicated by the first feature during the (m+1)th round of splicing.
[0110] For example, the aforementioned second feature indication can be determined as a new first feature indication, and the reference sound source analysis result under the aforementioned second feature indication can be optimized into the new first feature indication first sound source analysis result in the (m+1)th round of splicing process.
[0111] Based on the above steps, the reference sound source analysis results under the best feature indication can be determined. The reference sound source analysis results are also used to indicate the probability that each sound source in the sound source feature data is located as the sound source location corresponding to the direction of sound production.
[0112] Therefore, based on the reference sound source analysis results under the optimal feature indication, it is possible to determine the probability that each sound source in the sound source feature data is located in the direction of sound production.
[0113] Furthermore, the important information identification thread mentioned in the embodiments of this disclosure can also generate reference important nodes corresponding to the audio of each sound source signal to be processed based on the key descriptions and sound source orientation descriptions under the feature indications of differences in the acquired reference sound source data. The specific implementation process can refer to the process of acquiring the sound source vibration description and sound source orientation description of the sound source feature data corresponding to the reference sound direction set, and is not limited here.
[0114] In an alternative embodiment, after determining the sound direction parameter, the analysis variable for sound source localization corresponding to the sound direction in the reference sound source data can be configured as a first reference variable based on the obtained sound direction parameter. Furthermore, the analysis variables for sound source localization other than the sound direction in the reference sound source data can be configured as a second reference variable.
[0115] Based on the above, a virtual reality-based spatial orientation parameter acquisition device 200 is provided, which is applied to a virtual reality-based spatial orientation parameter acquisition system for a sound-emitting device. The device includes:
[0116] The data acquisition module 210 is used to obtain the audio corresponding to the audio of at least one audio source signal to be processed from the reference sound source data.
[0117] The signal selection module 220 is used to select a reference associated sound source signal that is associated with the reference mined sound source signal from the not less than one sound source signal to be processed, based on the first sound source data key description corresponding to the audio of the reference mined sound source signal and the second sound source data key description corresponding to each voiceprint dataset to be processed.
[0118] The direction determination module 230 is used to obtain a reference pronunciation direction set corresponding to the pronunciation direction of the reference associated sound source signal from the reference sound source data based on the voiceprint dataset to be processed by the reference associated sound source signal;
[0119] The parameter determination module 240 is used to perform parameter analysis on the reference pronunciation direction set and determine the pronunciation direction parameters of the reference associated sound source signal.
[0120] Based on the above, a virtual reality-based sound device spatial direction parameter acquisition system 300 is shown, including a processor 310 and a memory 320 that communicate with each other. The processor 310 is used to read computer programs from the memory 320 and execute them to implement the above method.
[0121] Based on the above, a computer-readable storage medium is also provided, on which a computer program stored implements the above method during runtime.
[0122] In summary, based on the above scheme, using the key descriptions of the first sound source data corresponding to the audio of the reference mined sound source signal and the key descriptions of the second sound source data corresponding to the voiceprint datasets of each sound source signal to be processed, it is possible to accurately select the second sound source data key descriptions associated with the first sound source data key description from several second sound source data key descriptions. That is, it is possible to accurately select the reference associated sound source signals, achieving precise mining of the reference mined sound source signals. Then, using the selected voiceprint datasets to be processed, a reference pronunciation direction set relative to the pronunciation direction is generated. Therefore, only parameter analysis needs to be performed on the reference pronunciation direction set, without processing the remaining data in the reference sound source data. This reduces the workload of sound source data processing while ensuring the accuracy of the sound source data to be processed, reducing the amount of data processing workload while reducing the number of sound source data segments to be processed. Furthermore, since only the sound source data of the reference pronunciation direction set needs to be processed, it is beneficial to improve the accuracy of parameter recognition of pronunciation direction, thus enabling more accurate pronunciation direction parameters.
[0123] It should be understood that the systems and modules described above can be implemented in various ways. For example, in some embodiments, the systems and modules can be implemented by hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated-design hardware. Those skilled in the art will understand that the methods and systems described above can be implemented using computer-executable instructions and / or included in processor control code, for example, such code provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The systems and modules of this application can be implemented not only by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field-programmable gate arrays, programmable logic devices, etc., but also by software executed by various types of processors, or by a combination of the aforementioned hardware circuits and software (e.g., firmware).
[0124] It should be noted that different embodiments may produce different beneficial effects. In different embodiments, the beneficial effects may be any one or a combination of the above, or any other possible beneficial effects.
[0125] The basic concepts have been described above. Obviously, for those skilled in the art, the detailed disclosure above is merely illustrative and does not constitute a limitation of this application. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this application. Such modifications, improvements, and corrections are suggested in this application, and therefore remain within the spirit and scope of the exemplary embodiments of this application.
[0126] Furthermore, this application uses specific terms to describe embodiments of the application. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of the application. Therefore, it should be emphasized and noted that "an embodiment," "one embodiment," or "an alternative embodiment" mentioned twice or more in different locations in this specification do not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of the application can be appropriately combined.
[0127] Furthermore, those skilled in the art will understand that aspects of this application can be described and illustrated through several patentable types or situations, including any new and useful combination of processes, machines, products, or substances, or any new and useful improvements thereof. Accordingly, aspects of this application can be implemented entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. All of the above hardware or software may be referred to as a “data block,” “module,” “engine,” “unit,” “component,” or “system.” Furthermore, aspects of this application may manifest as a computer product located on one or more computer-readable media, the product including computer-readable program code.
[0128] Computer storage media may contain a propagated data signal containing computer program code, for example, on baseband or as part of a carrier wave. This propagated signal may take various forms, including electromagnetic, optical, and suitable combinations thereof. Computer storage media can be any computer-readable medium other than a computer-readable storage medium, which can be connected to an instruction execution system, apparatus, or device to enable communication, propagation, or transmission of a program for use. The program code located on the computer storage medium can be propagated through any suitable medium, including radio, cable, fiber optic cable, RF, or similar media, or any combination of the above media.
[0129] The computer program code required for the operation of each part of this application can be written in any one or more programming languages, including object-oriented programming languages such as Java, Scala, Smalltalk, Eiffel, JADE, Emerald, C++, C#, VB.NET, Python, etc., conventional procedural programming languages such as C, Visual Basic, Fortran 2003, Perl, COBOL 2002, PHP, ABAP, dynamic programming languages such as Python, Ruby, and Groovy, or other programming languages. This program code can run entirely on the user's computer, or as a standalone software package on the user's computer, or partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any network, such as a local area network (LAN) or wide area network (WAN), or connected to an external computer (e.g., via the Internet), or in a cloud computing environment, or used as a service such as Software as a Service (SaaS).
[0130] Furthermore, unless expressly stated in the claims, the order of processing elements and sequences, the use of numbers and letters, or other names described in this application are not intended to limit the order of the processes and methods of this application. Although the foregoing disclosure has discussed some currently considered useful embodiments of the invention through various examples, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments; rather, the claims are intended to cover all modifications and equivalent combinations that conform to the substance and scope of the embodiments of this application. For example, while the system components described above can be implemented using hardware devices, they can also be implemented solely through software solutions, such as installing the described system on existing servers or mobile devices.
[0131] Similarly, it should be noted that, in order to simplify the description of the present application and thus aid in the understanding of one or more embodiments of the invention, the foregoing description of the embodiments of the present application sometimes combines multiple features into a single embodiment, drawing, or description thereof. However, this disclosure method does not imply that the subject matter of the application requires more features than those mentioned in the claims. In fact, the embodiments contain fewer features than all the features of the single embodiments disclosed above.
[0132] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of embodiments are modified in some examples with the terms "approximately," "approximately," or "generally." Unless otherwise stated, "approximately," "approximately," or "generally" indicates that the numbers are open to adaptive variation. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may be changed depending on the characteristics required by individual embodiments. In some embodiments, numerical parameters are taken into account a specified number of significant digits and employ a general method of digit reservation. Although the numerical ranges and parameters used to confirm their breadth of application in some embodiments of this application are approximate values, in specific embodiments, such values are set as precisely as feasible.
[0133] For each patent, patent application, patent application publication, and other material such as articles, books, specifications, publications, and documents referenced in this application, the entire contents of that patent are incorporated herein by reference. This excludes historical application documents that are inconsistent with or conflict with the content of this application, as well as documents that limit the broadest scope of the claims in this application (currently or subsequently appended to this application). It should be noted that if there are any inconsistencies or conflicts between the descriptions, definitions, and / or terminology used in the supplementary materials of this application and the content of this application, the descriptions, definitions, and / or terminology used in this application shall prevail.
[0134] Finally, it should be understood that the embodiments described in this application are merely illustrative of the principles of the embodiments of this application. Other modifications may also fall within the scope of this application. Therefore, alternative configurations of the embodiments of this application are considered as examples and not limitations, and are regarded as consistent with the teachings of this application. Accordingly, the embodiments of this application are not limited to the embodiments explicitly described and illustrated in this application.
[0135] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for obtaining spatial orientation parameters of a sound-producing device based on virtual reality, characterized in that, The method includes at least: From the reference sound source data, obtain the audio corresponding to the sound source signal to be processed, and obtain the voiceprint dataset to be processed. Based on the key description of the first sound source data corresponding to the audio of the reference mined sound source signal and the key description of the second sound source data corresponding to each voiceprint dataset to be processed, a reference associated sound source signal associated with the reference mined sound source signal is selected from the not less than one sound source signal to be processed. Based on the voiceprint dataset to be processed from the reference associated sound source signal, a set of reference pronunciation directions corresponding to the pronunciation direction of the reference associated sound source signal is obtained from the reference sound source data; Perform parameter analysis on the reference sound direction set to determine the sound direction parameters of the reference associated sound source signal; The voiceprint dataset to be processed based on the reference associated sound source signal obtains a reference pronunciation direction set corresponding to the pronunciation direction of the reference associated sound source signal from the reference sound source data, including: By combining the first mapping relationship data between the audio and the pronunciation direction, a second mapping relationship data between the voiceprint dataset to be processed and the reference pronunciation direction set is generated; By combining the second mapping relationship data and the voiceprint dataset to be processed, a reference pronunciation direction set corresponding to the pronunciation direction of the reference associated sound source signal is obtained from the reference sound source data.
2. The method according to claim 1, characterized in that, The step of obtaining the audio corresponding to at least one audio source signal to be processed from the reference sound source data includes: Important information is identified in the reference sound source data to generate reference important nodes corresponding to the audio of each sound source signal to be processed in at least one sound source signal to be processed included in the reference sound source data. For each of the at least one source signal to be processed, a voiceprint dataset corresponding to the source signal to be processed is generated based on the reference important nodes corresponding to the source signal to be processed.
3. The method according to claim 1 or 2, characterized in that, The step of selecting a reference associated sound source signal from the at least one unprocessed sound source signal by means of: associating the first sound source data key description corresponding to the audio of the reference mined sound source signal with the second sound source data key description corresponding to each unprocessed sound source dataset; and determining the unprocessed sound source signal corresponding to the second sound source data key description associated with the first sound source data key description as the reference associated sound source signal.
4. The method according to claim 3, characterized in that, Before obtaining the audio-text dataset corresponding to the audio of at least one audio source signal to be processed from the reference audio source data, the method further includes: obtaining the mined audio source data; and combining the mined audio source data to generate a reference mined audio source signal and a first audio source data key description corresponding to the audio of the reference mined audio source signal in the mined audio source data.
5. The method according to claim 4, characterized in that, The step of generating a reference excavation sound source signal by combining the excavation sound source data includes: Important information is identified in the excavated sound source data, and reference important nodes are generated for the audio corresponding to each original sound source signal in at least one original sound source signal included in the excavated sound source data. Based on the important node offset information of the reference important nodes corresponding to the audio of each original sound source signal, the reference mining sound source signal is selected and determined from the not less than one original sound source signal.
6. The method according to claim 3, characterized in that, The method further includes: on the basis of determining that there is no second sound source data key description associated with the first sound source data key description, generating a new reference mining sound source signal based on the important node offset information of the reference important nodes corresponding to the audio of each sound source signal to be processed.
7. The method according to claim 6, characterized in that, The step of performing parameter analysis on the reference pronunciation direction set to determine the pronunciation direction parameters of the reference associated sound source signal includes: Based on the reference pronunciation direction set, the sound source feature data corresponding to the reference pronunciation direction set is read from the reference sound source data; By using a pre-configured artificial intelligence thread, the vibration description and location description of the sound source feature data are obtained. Obtain the sound source vibration description corresponding one-to-one with the feature indicators that differ from the sound source feature data; wherein, among the several feature indicators, there are similar first feature indicators and second feature indicators, and the first feature indicator is smaller than the second feature indicator; The sound source vibration description corresponding to the first feature indication is obtained by combining the sound source vibration description corresponding to the second feature indication and the sound source orientation description of the sound source vibration description corresponding to the second feature indication; The second feature indicates that the sound source location description corresponding to the sound source vibration description is obtained by the pre-configured artificial intelligence thread; Based on the sound source vibration description corresponding to the different feature indicators, parameter analysis is performed on the sound source feature data to determine the sound direction parameters of the reference associated sound source signal.
8. The method according to claim 7, characterized in that, The parameter analysis of the sound source feature data based on the sound source vibration description corresponding to the differences in feature indicators includes: For each feature indication among the differing feature indications, based on the sound source vibration description corresponding to the feature indication, a first sound source analysis result of the sound source feature data under that feature indication is generated; Based on the first sound source analysis results of the sound source feature data under each feature indication, the probability of locating each sound source in the sound source feature data as the sound source location corresponding to the sound direction is generated. By combining the probability of locating each sound source in the sound source feature data as the sound source location corresponding to the direction of sound production with the specified mining probability vector, parameter analysis is performed on the sound source feature data.
9. The method according to claim 8, characterized in that, The step of combining the sound source feature data with the first sound source analysis results under each feature indication to generate the probability that each sound source in the sound source feature data is located as a sound source corresponding to the direction of sound production includes: After performing multiple rounds of splicing processing according to the feature indications that show differences, the probability of each sound source in the sound source feature data being located as a sound source corresponding to the direction of sound production is determined; wherein, the m-th round of splicing processing in the multiple rounds of splicing processing includes: generating the first bias information of the first sound source analysis result under the first feature indication; Using the first bias information of the first sound source analysis result under the first feature indication, the first sound source analysis result under the first feature indication and the first sound source analysis result under the second feature indication are spliced together to determine the reference sound source analysis result under the second feature indication; the reference sound source analysis result is optimized to the first sound source analysis result under the first feature indication in the (m+1)th round of splicing. The method further includes, after determining the sound direction parameters of the reference associated sound source signal: Based on the aforementioned sound direction parameter, the analysis variable for sound source localization corresponding to the sound direction in the reference sound source data is configured as the first reference variable; The analysis variables for locating sound sources other than the direction of sound production in the reference sound source data are configured as the second reference variables.