A Virtual Reality-Based Audio Acquisition Method

By utilizing audio optimization and artificial intelligence threads to filter standard digital audio data descriptions in virtual reality technology, the problems of noise interference and excessive data volume in audio data acquisition in virtual reality are solved, achieving efficient and accurate audio data acquisition.

CN115562616BActive Publication Date: 2025-11-14GUANGZHOU MOVIE POWER TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211309248.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-25
Publication Date
2025-11-14
Estimated Expiration
2042-10-25

AI Technical Summary

Technical Problem

In virtual reality technology, the acquisition of audio data is hampered by large data volumes and severe noise interference, making it difficult to guarantee the accuracy of the audio data.

Method used

A virtual reality-based audio acquisition method is adopted, which uses audio optimization threads and artificial intelligence threads to filter standard digital audio data descriptions, and generates partial description notes, partial description data and overall description data through parallel processing, thereby optimizing the audio data acquisition process.

Benefits of technology

It improved the noise interference problem, increased the data filtering speed, avoided the problem of excessive data volume caused by multiple filtering, and ensured the accuracy and efficiency of audio data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115562616B_ABST
    Figure CN115562616B_ABST
Patent Text Reader

Abstract

This application provides an audio acquisition method based on virtual reality. It utilizes an artificial intelligence thread to first filter standard digital audio data descriptions, and then, based on these filtered descriptions, enables the parallel determination of partial descriptive notes, partial descriptive data, and overall descriptive data. This improves noise interference and effectively increases the data filtering speed. Furthermore, since the standard digital audio data descriptions are required for filtering the aforementioned descriptive data, obtaining the descriptive data through the selected standard digital audio data descriptions avoids the problem of repeatedly filtering the standard digital audio data descriptions to find existing audio data descriptions, thus mitigating the issue of excessive data volume during data filtering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio data acquisition technology, and more specifically, to an audio acquisition method based on virtual reality. Background Technology

[0002] Virtual reality technology has gained increasing recognition, allowing users to experience the most realistic sensations in the virtual reality world. The realism of its simulated environment is indistinguishable from the real world, giving people a sense of immersion. At the same time, virtual reality possesses all the sensory functions that humans have, such as hearing, vision, touch, taste, and smell. Finally, it has a powerful simulation system that truly realizes human-computer interaction, allowing people to operate freely and receive the most realistic feedback from the environment during operation.

[0003] Currently, improving audio quality in virtual reality technology is a difficult technical problem to overcome. Due to the large amount of audio data, the acquired audio data may be defective, making it difficult to guarantee accurate audio data acquisition. Summary of the Invention

[0004] To address the technical problems existing in related technologies, this application provides an audio acquisition method based on virtual reality.

[0005] In a first aspect, a virtual reality-based audio acquisition method is provided, the method comprising at least: obtaining target digital audio data; selecting a standard digital audio data description of the target digital audio data; and combining the standard digital audio data description to generate partial description notes of the target digital audio data, partial description data corresponding to the partial description notes, and overall description data of the target digital audio data.

[0006] In one standalone embodiment, generating partial descriptive notes, partial descriptive data corresponding to the partial descriptive notes, and overall descriptive data of the target digital audio data by combining the standard digital audio data description includes: using an audio optimization thread, combining the standard digital audio data description, to output partial descriptive notes, partial descriptive data corresponding to the partial descriptive notes, and overall descriptive data of the target digital audio data.

[0007] In one standalone embodiment, configuring the audio optimization thread includes: obtaining several example digital audio data; for each example digital audio data, generating at least one example-mined digital audio data that is associated with the example digital audio data, and fusing the example digital audio data with each example-mined digital audio data one by one to obtain at least one audio data tuple; configuring the audio optimization thread using the audio data tuple.

[0008] In one independently implemented embodiment, the exemplary digital audio data includes first digital audio data, and the audio data tuple includes a first digital audio data tuple; the step of generating at least one example-mined digital audio data with a correlation to each example digital audio data, and fusing the example digital audio data with each example-mined digital audio data one by one to obtain at least one audio data tuple, includes: obtaining a plurality of first digital audio data under a target topic; generating topic data corresponding to the target topic based on the plurality of first digital audio data; and selecting at least one group of first digital audio data with a correlation from the first digital audio data in combination with the topic data to obtain at least one first digital audio data tuple.

[0009] In one independently implemented embodiment, the exemplary digital audio data includes second digital audio data, and the audio data tuple includes a second digital audio data tuple; the step of generating at least one example-mined digital audio data with a correlation to each example digital audio data, and fusing the example digital audio data with each example-mined digital audio data one by one to obtain at least one audio data tuple, includes: obtaining several second digital audio data under non-target topics; performing a first optimization process on each second digital audio data to obtain at least one target-mined digital audio data with a correlation to the second digital audio data, and fusing the second digital audio data with each target-mined digital audio data one by one to obtain at least one second digital audio data tuple.

[0010] In one standalone embodiment, configuring the audio optimization thread using the audio data tuple includes: performing a second optimization process on a portion of the audio data tuple to obtain at least one optimized audio data tuple; the audio data tuple includes a first audio data tuple and / or a second audio data tuple; and configuring the audio optimization thread using the portion of the audio data tuple and the optimized audio data tuple.

[0011] In one standalone embodiment, configuring the audio optimization thread using the audio data tuple includes: generating partial sound quality recording points for each example digital audio data in the audio data tuple; selecting a first timbre and a second timbre corresponding to the partial sound quality recording points from the example-mined digital audio data in the audio data tuple; and configuring the audio optimization thread using the partial sound quality recording points and the first and second timbres corresponding to the partial sound quality recording points.

[0012] In one independently implemented embodiment, selecting the first timbre and the second timbre corresponding to the partial sound quality recording points from the example-mined digital audio data in the audio data tuple includes: determining the sound quality recording points in the example-mined digital audio data that correspond to the partial sound quality recording points as the first timbre; selecting the sound quality recording point with the best matching degree from the sound quality recording points other than the first timbre in the example-mined digital audio data, and determining the selected sound quality recording point with the best matching degree as the second timbre.

[0013] In one independent embodiment, configuring the audio optimization thread using the partial audio quality recording points and the first and second timbres corresponding to the partial audio quality recording points includes: loading the audio data tuple into the audio optimization thread to obtain partial evaluation audio quality recording points, first evaluation description data corresponding to the partial audio quality recording points, second evaluation description data corresponding to the first timbre, and third evaluation description data corresponding to the second timbre; generating a first quantization evaluation model using the partial evaluation audio quality recording points and the partial audio quality recording points; generating a second quantization evaluation model using the first evaluation description data corresponding to the partial audio quality recording points, the second evaluation description data corresponding to the first timbre, and the third evaluation description data corresponding to the second timbre; and configuring the audio optimization thread using the first quantization evaluation model and the second quantization evaluation model.

[0014] In one independently implemented embodiment, loading the audio data tuple into the audio optimization thread to obtain partial evaluation sound quality recording points, first evaluation description data corresponding to the partial sound quality recording points, second evaluation description data corresponding to the first timbre, and third evaluation description data corresponding to the second timbre includes: loading the audio data tuple into the audio optimization thread; obtaining the evaluation standard digital audio data description corresponding to the audio data tuple by a standard description filtering sub-thread in the audio optimization thread; and obtaining the partial evaluation sound quality recording points, the first evaluation description data corresponding to the partial sound quality recording points, the second evaluation description data corresponding to the first timbre, and the third evaluation description data corresponding to the second timbre by a description evaluation sub-thread in the audio optimization thread in combination with the evaluation standard digital audio data description.

[0015] In one independently implemented embodiment, configuring the audio optimization thread using the partial audio quality recording points and the first and second timbres corresponding to the partial audio quality recording points includes: inputting a plurality of first audio data tuples into the audio optimization thread, obtaining the overall fourth evaluation description data corresponding to each digital audio data tuple in each of the first audio data tuples one by one, and storing the overall fourth evaluation description data into a specified matrix; configuring the audio optimization thread based on a plurality of second audio data tuples and the fourth evaluation description data in the specified matrix.

[0016] In one standalone embodiment, configuring the audio optimization thread based on a plurality of second audio data tuples and fourth evaluation description data in the specified matrix includes: inputting the plurality of second audio data tuples into the audio optimization thread to obtain, one by one, the overall fifth evaluation description data of each digital audio data in each of the second audio data tuples; combining the fourth evaluation description data in the specified matrix and the fifth evaluation description data to generate a third quantization evaluation model, and configuring the audio optimization thread through the third quantization evaluation model.

[0017] In one standalone embodiment, the fifth evaluation description data includes a sixth evaluation description data and a seventh evaluation description data; the step of combining the fourth evaluation description data in the specified matrix and the fifth evaluation description data to generate a third quantitative evaluation model includes: for each second audio data tuple, generating sixth evaluation description data for example digital audio data in the second audio data tuple, determining the example mined digital audio data in the second audio data tuple as the first sample audio data corresponding to the audio data tuple, and generating seventh evaluation description data for the first sample audio data; selecting the fourth evaluation description data from the specified matrix that has the best matching degree with the sixth evaluation description data of the example digital audio data in the second audio data tuple, and determining the selected fourth evaluation description data as the eighth evaluation description data corresponding to the negative audio data tuple; generating the third quantitative evaluation model through the sixth evaluation description data, the seventh evaluation description data, and the eighth evaluation description data.

[0018] In one standalone embodiment, after obtaining the overall fifth evaluation description data of each digital audio data in each second audio data tuple, the method further includes: optimizing the specified matrix based on the overall fifth evaluation description data of each digital audio data.

[0019] The virtual reality-based audio acquisition method provided in this application utilizes an artificial intelligence thread to first filter standard digital audio data descriptions. Then, based on the filtered standard digital audio data descriptions, it can simultaneously determine partial descriptive notes, partial descriptive data, and overall descriptive data, improving noise interference and effectively increasing the data filtering speed. Furthermore, since the standard digital audio data descriptions are the descriptive data needed to filter the aforementioned descriptive data, obtaining the aforementioned descriptive data by selecting the obtained standard digital audio data descriptions avoids the problem of repeatedly filtering the standard digital audio data descriptions to find the existing audio data descriptions, thus mitigating the problem of excessive data volume in data filtering. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart of a virtual reality-based audio acquisition method provided in an embodiment of this application. Detailed Implementation

[0022] To better understand the above technical solutions, the technical solutions of this application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of this application and the specific features in the embodiments are detailed descriptions of the technical solutions of this application, rather than limitations on the technical solutions of this application. In the absence of conflict, the embodiments of this application and the technical features in the embodiments can be combined with each other.

[0023] Please see Figure 1 This paper illustrates an audio acquisition method based on virtual reality, which may include the technical solutions described in steps S101-S103.

[0024] S101: Obtain the target digital audio data.

[0025] Furthermore, the target digital audio data can be digital audio data under the target theme at the audio receiver, or digital audio data under a non-target theme. When there are partial descriptive notes in the target digital audio data, the partial descriptive notes can be the most important information that best describes the target digital audio data within a specified area at the corresponding binary node of the target digital audio data.

[0026] Partial descriptive notes can reflect the sample description of a portion of the target digital audio data and can be used to mark target notes in the target digital audio data. By associating the partial descriptive notes of the target digital audio data with the partial descriptive notes of the remaining digital audio data, the association between the target digital audio data and the remaining digital audio data can be realized.

[0027] S102: Standard digital audio data description for filtering target digital audio data.

[0028] Furthermore, based on the obtained target digital audio data, a standard digital audio data description of the target digital audio data can be selected. The standard digital audio data description can be the digital audio data description that is needed when determining the partial description notes of the target digital audio data, the partial description data corresponding to the partial description notes, and the overall description data of the target digital audio data.

[0029] S103: Based on standard digital audio data description, generate partial description notes of the target digital audio data, partial description data corresponding to the partial description notes, and overall description data of the target digital audio data.

[0030] Furthermore, partial description data refers to global description data used to describe the specified area at the node corresponding to the partial description note; each partial description note can have one corresponding partial description data. Global description data refers to global description data used to represent the target digital audio data; each target digital audio data has a unique global digital audio data description.

[0031] Furthermore, partial and overall descriptive data can be represented by multidimensional descriptive variables.

[0032] In this embodiment, after obtaining the standard digital audio data description, the standard digital audio data description and the target digital audio data can be used to determine, based on the premise that the target digital audio data contains partial descriptive notes, the partial descriptive notes of each part of the target digital audio data and the corresponding partial descriptive data can be determined in parallel. Specifically, based on the premise that the target digital audio data contains partial descriptive notes, a target digital audio data set can contain at least one partial descriptive note; that is, a target digital audio data set can correspond to at least one partial descriptive data set. Furthermore, based on the standard digital audio data description and the target digital audio data, the overall descriptive data of the target digital audio data can also be determined in parallel.

[0033] For example, determining the descriptive data for each part requires not only the standard digital audio data description and the target digital audio data, but also the corresponding descriptive notes. Furthermore, the descriptive data for each part can be called a partial sub-description, and the overall descriptive data can be called an overall sub-description.

[0034] Preferably, by using an artificial intelligence thread to first filter standard digital audio data descriptions, and then based on the filtered standard digital audio data descriptions, it is possible to determine some descriptive notes, some descriptive data, and the overall descriptive data in parallel, which improves the noise interference problem, effectively increases the data filtering speed, and since the standard digital audio data descriptions are the descriptive data needed to filter the above descriptive data, obtaining the above descriptive data by selecting the obtained standard digital audio data descriptions can avoid the problem of filtering the standard digital audio data descriptions multiple times to find the existing audio data descriptions when obtaining the descriptive data, thus reducing the problem of excessive data volume in the data filtering process.

[0035] In a standalone embodiment, for S103, an audio optimization thread can be used to output partial description notes of the target digital audio data, partial description data corresponding to the partial description notes, and overall description data of the target digital audio data based on standard digital audio data description.

[0036] Furthermore, the virtual reality-based audio acquisition method provided in this embodiment can be applied in an audio optimization thread, and steps S101 to S103 can all be executed using this audio optimization thread. Specifically, the audio optimization thread can be used to process the acquired target digital audio data, output a standard digital audio data description of the target digital audio data, and then, using the standard digital audio data description and the target digital audio data, output partial description notes of the target digital audio data, partial description data corresponding to the partial description notes, and overall description data of the target digital audio data.

[0037] For example, an audio optimization thread may include a standard description filtering sub-thread and a description evaluation sub-thread. The description evaluation sub-thread may include a partial description note recognition sub-thread, a partial description data recognition sub-thread, and a total description data recognition sub-thread. An audio optimization thread provided in this embodiment outputs partial description notes, corresponding partial description data, and total description data of the target digital audio data. Specifically, the standard description filtering sub-thread is used to filter standard digital audio data descriptions of the target digital audio data; the partial description note recognition sub-thread is used to output partial description notes of the target digital audio data; the partial description data recognition sub-thread is used to output the corresponding partial description data; and the total description data recognition sub-thread is used to output the total description data of the target digital audio data.

[0038] In an alternative embodiment, the audio optimization thread needs to be configured to output evaluation data with high reliability and accuracy. Therefore, this disclosure also provides a method for configuring the audio optimization thread, which may specifically include the following steps.

[0039] S301: Obtain several sample digital audio data.

[0040] S302: For each example digital audio data, generate at least one example-mined digital audio data that is related to the example digital audio data, and fuse the example digital audio data with each example-mined digital audio data one by one to obtain at least one audio data tuple.

[0041] Among them, audio data tuples carrying relationships include digital audio data tuples carrying locally repeating nodes.

[0042] S303: Configure the audio optimization thread using audio data tuples.

[0043] Furthermore, the example digital audio data can be digital audio data under random themes, and the correlation can reflect the changes in direction and the number and location of repeated nodes between the two digital audio data. The configured audio optimization thread can, based on the processing of the example digital audio data, output partially descriptive notes of the example digital audio data with high reliability and accuracy, partially descriptive data corresponding to the partially descriptive notes, and overall descriptive data of the target digital audio data.

[0044] After selecting the audio data tuples, each pair of example digital audio data in the audio data tuples can be loaded into the audio optimization thread. Then, based on the output of the audio optimization thread, a quantitative evaluation model for optimizing the configuration of the audio optimization thread is constructed. The audio optimization thread is optimized using the constructed quantitative evaluation model tuples to obtain the configured audio optimization thread. Furthermore, the audio data tuples include several pairs.

[0045] Furthermore, after obtaining several audio data tuples, a second optimization process can be performed on a portion of these audio data tuples to obtain at least one optimized audio data tuple.

[0046] In one standalone embodiment, the exemplary digital audio data may include first digital audio data, and the audio data tuple includes the first digital audio data tuple.

[0047] For step S302, the audio data tuple can be determined by following these steps.

[0048] (1) Obtain several first digital audio data under the target topic.

[0049] (2) Generate topic data corresponding to the target topic based on several first digital audio data.

[0050] (3) Based on the topic data, select at least one set of first digital audio data that carries a correlation from the first digital audio data to obtain at least one first digital audio data tuple.

[0051] Furthermore, in order to improve the evaluation accuracy of the configured audio optimization thread for digital audio data under the target topic, the audio optimization thread can be configured using the digital audio data tuple under the target topic.

[0052] The first digital audio data can be digital audio data obtained under the target topic. The first digital audio data tuple includes two first digital audio data selected from several first digital audio data with a correlation.

[0053] For example, after obtaining several different digital audio data, an artificial intelligence thread can be performed on the target topic corresponding to the digital audio data binary based on the several digital audio data to determine the topic data corresponding to the target topic. Then, for each of the several first digital audio data, using the topic data, the remaining first digital audio data that are related to the first digital audio data can be selected from the several first digital audio data. Then, the first digital audio data can be fused with each of the remaining first digital audio data that are related to the first digital audio data to obtain the first digital audio data binary corresponding to the first digital audio data binary. Then, the first digital audio data binary corresponding to each of the several first digital audio data binary can be determined. Furthermore, at least one first digital audio data binary can be determined.

[0054] In an alternative embodiment, the exemplary digital audio data may further include second digital audio data, and similarly, the audio data tuple may also include a second digital audio data tuple. For example, the second digital audio data is digital audio data outside the target topic. After obtaining several second digital audio data sets, a first optimization process can be performed on each second digital audio data set. This allows for the identification of at least one target-mined digital audio data set associated with the second digital audio data set. Then, the second digital audio data set is fused with each target-mined digital audio data set identified in the at least one target-mined digital audio data set to obtain at least one second digital audio data tuple corresponding to the second digital audio data tuple. Further, based on the first optimization process, the second digital audio data tuple corresponding to each second digital audio data tuple can be determined.

[0055] After determining the second digital audio data tuple, the audio optimization thread can be optimized using both the first and second digital audio data tuples. For example, a first number of first digital audio data tuples and a second number of second digital audio data tuples for configuration can be determined using a specified confidence level. Then, the first number of first digital audio data tuples and the second number of second digital audio data tuples can be selected together to optimize the audio optimization thread.

[0056] Furthermore, for the selected first number of first digital audio data tuples and the selected second number of second digital audio data tuples, a portion of the first digital audio data tuples and / or the second digital audio data tuples can be selected and subjected to a second optimization process to obtain optimized first digital audio data tuples and / or optimized second digital audio data tuples. Subsequently, the first digital audio data tuples and / or the second digital audio data tuples, the optimized first digital audio data tuples and / or the optimized second digital audio data tuples can be used together to optimize the configuration of the audio optimization thread.

[0057] Furthermore, regarding S303, after determining the audio data tuples, partial descriptive notes included in the digital audio data of each audio data tuple can be generated. Then, the partial descriptive notes included in the example digital audio data of the determined audio data tuples can be identified as the partial timbre recording points corresponding to the audio data tuples. Moreover, the first timbre and the second timbre corresponding to the determined partial timbre recording points can be selected from the example mined digital audio data in the audio data tuples. Afterward, the partial timbre recording points, the first timbre, and the second timbre can be used to optimize the configuration of the audio optimization thread to obtain the configured audio optimization thread.

[0058] Furthermore, the example digital audio data in each audio data tuple can include several partial description notes. That is, from the example-mined digital audio data corresponding to the audio data tuple, a first timbre and a second timbre corresponding to each partial description keypoint among the partial description notes can be generated. Then, the determined partial description keypoints and their corresponding first and second timbres can be used to optimize the configuration of the audio optimization thread.

[0059] In one independent embodiment, regarding the determination of the first and second timbres corresponding to a random partial timbre recording point included in the example digital audio data within the audio data tuple, a partial descriptive note corresponding to that partial descriptive keypoint can be generated from the example-mined digital audio data and identified as the first timbre corresponding to that partial descriptive keypoint. Then, the second timbre corresponding to that partial descriptive keypoint can be generated from the partial descriptive notes other than the first timbre included in the example-mined digital audio data. For example, for each partial descriptive note among the partial descriptive notes other than the first timbre included in the example-mined digital audio data, the matching degree between each partial descriptive note and that partial descriptive keypoint can be determined, and then the partial descriptive note with the best matching degree can be selected as the second timbre.

[0060] Furthermore, after determining the partial sound quality recording points, the first timbre, and the second timbre included in the audio data tuple, the audio data tuple can be loaded into the audio optimization thread. For each partial sound quality recording point included in the audio data tuple and its corresponding first and second timbres, the audio optimization thread can, based on the processing of the digital audio data in the audio data tuple, output the partial evaluation sound quality recording point corresponding to that partial sound quality recording point, the first evaluation description data corresponding to that partial sound quality recording point, the second evaluation description data for the first timbre corresponding to that partial sound quality recording point, and the third evaluation description data for the second timbre corresponding to that partial sound quality recording point. Here, the first, second, and third evaluation description data are all output partial description data.

[0061] Then, by using partial sound quality recording points and their corresponding partial evaluation sound quality recording points, and taking the partial sound quality recording points as the target values, a first quantitative evaluation model for partial description note recognition can be built. Then, the partial description note recognition sub-thread of the first quantitative evaluation model can be optimized and configured. In this way, the partial description note recognition sub-thread in the configured audio optimization thread can output partial evaluation sound quality recording points with high accuracy.

[0062] Furthermore, a second quantitative evaluation model can be generated using the first evaluation description data, the second evaluation description data, and the third evaluation description data corresponding to some audio quality recording points. The second quantitative evaluation model can be a quantitative evaluation model thread.

[0063] Then, the partial description data recognition sub-thread of the second quantitative evaluation model binary can be optimized and configured. In this way, the partial description data recognition sub-thread in the configured audio optimization thread can output partial evaluation description data with higher accuracy.

[0064] Furthermore, the first and second quantitative evaluation models can be generated simultaneously, and then the binary pairs of the first and second quantitative evaluation models can be utilized.

[0065] The audio optimization thread performs optimization configuration to obtain a fully configured audio optimization thread.

[0066] In addition, the virtual reality-based audio acquisition method provided in this embodiment can also simultaneously configure the overall description data recognition sub-thread.

[0067] For example, several first audio data tuples can be loaded into the audio optimization thread simultaneously. These first audio data tuples can be digital audio data tuples selected from a predetermined set of audio data tuples. Then, for each of the input first audio data tuples, the standard description filtering sub-thread within the audio optimization thread can filter the evaluation standard descriptions of the digital audio data included in that first audio data tuple. Finally, the partial description note recognition sub-thread within the audio optimization thread can generate the corresponding digital audio data tuples based on the determined evaluation standard descriptions and the digital audio data included in the first audio data tuples. The audio data binary corresponds to partial evaluation sound quality recording points. The partial description data recognition sub-thread in the audio optimization thread can, based on the partial evaluation sound quality recording points corresponding to the first audio data binary, the evaluation standard description, and the digital audio data included in the first audio data binary, output the first evaluation description data, second evaluation description data, and third evaluation description data corresponding to each digital audio data binary in the first audio data binary. The overall description data recognition sub-thread in the audio optimization thread can, based on the evaluation standard description and the digital audio data included in the first audio data binary, output the overall fourth evaluation description data corresponding to each digital audio data binary in the first audio data binary. The fourth evaluation description data is the output overall description data.

[0068] Furthermore, by utilizing partial evaluation sound quality recording points, first evaluation description data, second evaluation description data, and third evaluation description data, a second quantization evaluation model and a third quantization evaluation model can be built. The partial description note recognition sub-thread and the partial description data recognition sub-thread can be optimized and configured. Moreover, the third quantization evaluation model can be built using the fourth evaluation description data, and the overall description data recognition sub-thread can be optimized and configured in parallel.

[0069] It is understandable that during the optimization configuration of the overall description data recognition sub-thread using audio data tuples, for a random example digital audio data, its corresponding first sample audio data can be digital audio data that carries a correlation with the example digital audio data, and the second sample audio data is all digital audio data that has no correlation with the current digital audio data. Therefore, for each input audio data tuple, the example digital audio data and the example mined digital audio data included therein carry a correlation. Thus, for each input audio data tuple, only the first sample audio data exists, and the second sample audio data does not exist. However, in order to improve the recognition accuracy of the overall evaluation description data output by the configured overall description data recognition sub-thread, this disclosure proposes a matrix-based configuration method. The overall evaluation description data generated by each optimization configuration is stored in a specified matrix and used as the overall evaluation description data corresponding to the negative audio data tuple in the next round of optimization configuration, thereby improving the recognition accuracy of the overall evaluation description data output by the configured overall description data recognition sub-thread.

[0070] For example, a second audio data tuple, which is different from the first audio data tuple, can be selected from a certain number of audio data tuples and loaded into the feature recognition artificial intelligence thread. The overall fifth evaluation description data of each digital audio data in each second audio data tuple can be obtained. Then, the fifth evaluation description data and the fourth evaluation description data in the specified matrix can be used to generate a third quantitative evaluation model for configuring the audio optimization thread.

[0071] In one standalone embodiment, the fifth evaluation description data may include a sixth evaluation description data corresponding to the audio data binary in the second audio data binary, and a seventh evaluation description data corresponding to the example-mined digital audio data binary in the second audio data binary.

[0072] For example, for each of the several second audio data tuples loaded into the audio optimization thread, the sixth evaluation description data corresponding to the audio data tuple in each second audio data tuple, and the seventh evaluation description data corresponding to the example-mined digital audio data tuple in the second audio data tuple can be determined. Then, from the fourth evaluation description data stored in a specified matrix, which is determined to be the overall evaluation description data corresponding to the negative audio data tuple, the fourth evaluation description data with the best matching degree with the sixth evaluation description data can be generated. This fourth evaluation description data is determined as the overall evaluation description data corresponding to the negative audio data tuple used to generate the third quantization evaluation model, i.e., the eighth evaluation description data. Then, the sixth evaluation description data, the seventh evaluation description data, and the eighth evaluation description data can be used to generate the third quantization evaluation model. The overall description data of the third quantization evaluation model tuple is used to identify the sub-thread for the current round configuration. The third quantization evaluation model can be a comparative quantization evaluation model.

[0073] Furthermore, each round of optimization configuration for the audio optimization thread also involves configuring the standard description filtering sub-thread. Once the audio optimization thread is configured, the standard description filtering sub-thread is also configured, enabling it to output highly accurate evaluation standard descriptions.

[0074] Preferably, the configured audio optimization thread obtained by using synchronous configuration can synchronously output the partial description notes of the target digital audio data, the partial description data corresponding to the partial description notes, and the overall description data of the target digital audio data based on the filtered standard digital audio data description, thereby improving the data recognition speed of the target digital audio data.

[0075] In one independent implementation, the virtual reality-based audio acquisition method provided in this disclosure can further optimize the standard description filtering sub-thread, partial description note recognition sub-thread, and partial description data recognition sub-thread in the audio optimization thread by first using audio data tuples. After the configuration is completed, the overall description data recognition sub-thread in the audio optimization thread is optimized by using audio data tuples. Finally, the configured audio optimization thread is obtained. No specific limitations are made here.

[0076] Based on the above, a virtual reality-based audio acquisition device 200 is provided, which is applied to a virtual reality-based audio acquisition system. The device includes:

[0077] Data acquisition module 210 is used to acquire target digital audio data;

[0078] Description selection module 220 is used to select a standard digital audio data description for the target digital audio data;

[0079] The data generation module 230 is used to combine the standard digital audio data description to generate partial description notes of the target digital audio data, partial description data corresponding to the partial description notes, and overall description data of the target digital audio data.

[0080] Based on the above, a virtual reality-based audio acquisition system 300 is shown, including a processor 310 and a memory 320 that communicate with each other. The processor 310 is used to read computer programs from the memory 320 and execute them to implement the above-described method.

[0081] Based on the above, a computer-readable storage medium is also provided, on which a computer program stored implements the above method during runtime.

[0082] In summary, based on the above scheme, by using an artificial intelligence thread to first filter standard digital audio data descriptions, and then using these filtered descriptions, it is possible to determine some descriptive notes, some descriptive data, and the overall descriptive data in parallel. This improves the noise interference problem, effectively increases the data filtering speed, and since the standard digital audio data descriptions are needed to filter all the aforementioned descriptive data, obtaining the descriptive data through the selected standard digital audio data descriptions avoids the problem of repeatedly filtering the standard digital audio data descriptions to find existing audio data descriptions. This also mitigates the problem of excessive data volume in the data filtering process.

[0083] It should be understood that the systems and modules described above can be implemented in various ways. For example, in some embodiments, the systems and modules can be implemented by hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated-design hardware. Those skilled in the art will understand that the methods and systems described above can be implemented using computer-executable instructions and / or included in processor control code, for example, such code provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The systems and modules of this application can be implemented not only by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field-programmable gate arrays, programmable logic devices, etc., but also by software executed by various types of processors, or by a combination of the aforementioned hardware circuits and software (e.g., firmware).

[0084] It should be noted that different embodiments may produce different beneficial effects. In different embodiments, the beneficial effects may be any one or a combination of the above, or any other possible beneficial effects.

[0085] The basic concepts have been described above. Obviously, for those skilled in the art, the detailed disclosure above is merely illustrative and does not constitute a limitation of this application. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this application. Such modifications, improvements, and corrections are suggested in this application, and therefore remain within the spirit and scope of the exemplary embodiments of this application.

[0086] Furthermore, this application uses specific terms to describe embodiments of the application. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of the application. Therefore, it should be emphasized and noted that "an embodiment," "one embodiment," or "an alternative embodiment" mentioned twice or more in different locations in this specification do not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of the application can be appropriately combined.

[0087] Furthermore, those skilled in the art will understand that aspects of this application can be described and illustrated through several patentable types or situations, including any new and useful combination of processes, machines, products, or substances, or any new and useful improvements thereof. Accordingly, aspects of this application can be implemented entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. All of the above hardware or software may be referred to as a “data block,” “module,” “engine,” “unit,” “component,” or “system.” Furthermore, aspects of this application may manifest as a computer product located on one or more computer-readable media, the product including computer-readable program code.

[0088] Computer storage media may contain a propagated data signal containing computer program code, for example, on baseband or as part of a carrier wave. This propagated signal may take various forms, including electromagnetic, optical, and suitable combinations thereof. Computer storage media can be any computer-readable medium other than a computer-readable storage medium, which can be connected to an instruction execution system, apparatus, or device to enable communication, propagation, or transmission of a program for use. The program code located on the computer storage medium can be propagated through any suitable medium, including radio, cable, fiber optic cable, RF, or similar media, or any combination of the above media.

[0089] The computer program code required for the operation of each part of this application can be written in any one or more programming languages, including object-oriented programming languages ​​such as Java, Scala, Smalltalk, Eiffel, JADE, Emerald, C++, C#, VB.NET, Python, etc., conventional procedural programming languages ​​such as C, Visual Basic, Fortran 2003, Perl, COBOL 2002, PHP, ABAP, dynamic programming languages ​​such as Python, Ruby, and Groovy, or other programming languages. This program code can run entirely on the user's computer, or as a standalone software package on the user's computer, or partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any network, such as a local area network (LAN) or wide area network (WAN), or connected to an external computer (e.g., via the Internet), or in a cloud computing environment, or used as a service such as Software as a Service (SaaS).

[0090] Furthermore, unless expressly stated in the claims, the order of processing elements and sequences, the use of numbers and letters, or other names described in this application are not intended to limit the order of the processes and methods of this application. Although the foregoing disclosure has discussed some currently considered useful embodiments of the invention through various examples, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments; rather, the claims are intended to cover all modifications and equivalent combinations that conform to the substance and scope of the embodiments of this application. For example, while the system components described above can be implemented using hardware devices, they can also be implemented solely through software solutions, such as installing the described system on existing servers or mobile devices.

[0091] Similarly, it should be noted that, in order to simplify the description of the present application and thus aid in the understanding of one or more embodiments of the invention, the foregoing description of the embodiments of the present application sometimes combines multiple features into a single embodiment, drawing, or description thereof. However, this disclosure method does not imply that the subject matter of the application requires more features than those mentioned in the claims. In fact, the embodiments contain fewer features than all the features of the single embodiments disclosed above.

[0092] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of embodiments are modified in some examples with the terms "approximately," "approximately," or "generally." Unless otherwise stated, "approximately," "approximately," or "generally" indicates that the numbers are open to adaptive variation. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may be changed depending on the characteristics required by individual embodiments. In some embodiments, numerical parameters are taken into account a specified number of significant digits and employ a general method of digit reservation. Although the numerical ranges and parameters used to confirm their breadth of application in some embodiments of this application are approximate values, in specific embodiments, such values ​​are set as precisely as feasible.

[0093] For each patent, patent application, patent application publication, and other material such as articles, books, specifications, publications, and documents referenced in this application, the entire contents of that patent are incorporated herein by reference. This excludes historical application documents that are inconsistent with or conflict with the content of this application, as well as documents that limit the broadest scope of the claims in this application (currently or subsequently appended to this application). It should be noted that if there are any inconsistencies or conflicts between the descriptions, definitions, and / or terminology used in the supplementary materials of this application and the content of this application, the descriptions, definitions, and / or terminology used in this application shall prevail.

[0094] Finally, it should be understood that the embodiments described in this application are merely illustrative of the principles of the embodiments of this application. Other modifications may also fall within the scope of this application. Therefore, alternative configurations of the embodiments of this application are considered as examples and not limitations, and are regarded as consistent with the teachings of this application. Accordingly, the embodiments of this application are not limited to the embodiments explicitly described and illustrated in this application.

[0095] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for acquiring audio based on virtual reality, characterized in that, The method includes at least: Obtain the target digital audio data; Select a standard digital audio data description for the target digital audio data; Based on the standard digital audio data description, generate partial description notes of the target digital audio data, partial description data corresponding to the partial description notes, and overall description data of the target digital audio data; The step of generating partial description notes, corresponding partial description data, and overall description data of the target digital audio data by combining the standard digital audio data description includes: using an audio optimization thread, combining the standard digital audio data description, to output partial description notes, corresponding partial description data, and overall description data of the target digital audio data; configuring the audio optimization thread includes: Obtain several sample digital audio data; For each example digital audio data, generate at least one example-mined digital audio data that is related to the example digital audio data, and fuse the example digital audio data with each example-mined digital audio data one by one to obtain at least one audio data tuple. The audio optimization thread is configured using the audio data tuple; The example digital audio data includes first digital audio data, and the audio data tuple includes the first digital audio data tuple; the step of generating at least one example-mined digital audio data that is associated with each example digital audio data, and fusing the example digital audio data with each example-mined digital audio data one by one to obtain at least one audio data tuple, includes: Obtain several first-digit audio data points under the target topic; Based on several first digital audio data, generate topic data corresponding to the target topic; Combining the topic data, select at least one set of first digital audio data that carries a correlation from the first digital audio data to obtain at least one first digital audio data tuple; The example digital audio data includes second digital audio data, and the audio data tuple includes the second digital audio data tuple; the step of generating at least one example-mined digital audio data with a correlation to each example digital audio data, and fusing the example digital audio data with each example-mined digital audio data one by one to obtain at least one audio data tuple, includes: Obtain several second-digit audio data points from non-target topics; For each second digital audio data, a first optimization process is performed on the second digital audio data to obtain at least one target mining digital audio data that is related to the second digital audio data. The second digital audio data is then fused with each target mining digital audio data one by one to obtain at least one second digital audio data tuple.

2. The method as described in claim 1, characterized in that, The configuration of the audio optimization thread using the audio data tuple includes: Perform a second optimization process on a portion of the audio data tuples to obtain at least one optimized audio data tuple; The audio data tuple includes a first audio data tuple and / or a second audio data tuple; The audio optimization thread is configured using one portion of the audio data tuple and the optimized audio data tuple.

3. The method as described in claim 2, characterized in that, The configuration of the audio optimization thread using the audio data tuple includes: Generate partial audio quality recording points for each example digital audio data in the audio data tuple respectively; From the examples mined from the audio data tuples, the first timbre and the second timbre corresponding to the partial sound quality recording points are selected; The audio optimization thread is configured using the aforementioned partial sound quality recording points and the first and second timbres corresponding to the aforementioned partial sound quality recording points; The step of selecting the first and second timbres corresponding to the partial sound quality recording points from the digital audio data mined from the audio data tuples includes: The audio quality recording points in the digital audio data mined in the example that correspond to the aforementioned partial audio quality recording points are determined as the first timbre. From the digital audio data mined from the example, excluding the first timbre, select the timbre recording point with the best matching degree with the aforementioned timbre recording points, and determine the selected timbre recording point with the best matching degree as the second timbre. The step of configuring the audio optimization thread using the partial sound quality recording points and the first and second timbres corresponding to the partial sound quality recording points includes: The audio data tuple is loaded into the audio optimization thread to obtain partial evaluation sound quality recording points, first evaluation description data corresponding to the partial sound quality recording points, second evaluation description data corresponding to the first timbre, and third evaluation description data corresponding to the second timbre. A first quantitative evaluation model is generated by using the aforementioned partial sound quality recording points and the aforementioned partial sound quality recording points; A second quantitative evaluation model is generated by using the first evaluation description data corresponding to the partial sound quality recording points, the second evaluation description data corresponding to the first timbre, and the third evaluation description data corresponding to the second timbre. The audio optimization thread is configured using the first quantization evaluation model and the second quantization evaluation model.

4. The method as described in claim 3, characterized in that, The step of loading the audio data tuple into the audio optimization thread to obtain partial evaluation sound quality recording points, first evaluation description data corresponding to the partial sound quality recording points, second evaluation description data corresponding to the first timbre, and third evaluation description data corresponding to the second timbre includes: The audio data tuple is loaded into the audio optimization thread, and the standard description filtering sub-thread in the audio optimization thread obtains the evaluation standard digital audio data description corresponding to the audio data tuple. The description and evaluation sub-thread in the audio optimization thread, combined with the evaluation standard digital audio data description, obtains partial evaluation sound quality recording points, first evaluation description data corresponding to the partial sound quality recording points, second evaluation description data corresponding to the first timbre, and third evaluation description data corresponding to the second timbre.

5. The method as described in claim 4, characterized in that, The configuration of the audio optimization thread using the partial audio quality recording points and the first and second timbres corresponding to the partial audio quality recording points includes: Several first audio data tuples are input into the audio optimization thread to obtain the overall fourth evaluation description data corresponding to each digital audio data tuple in each first audio data tuple, and the overall fourth evaluation description data is stored in a specified matrix. The audio optimization thread is configured based on several second audio data tuples and the fourth evaluation description data in the specified matrix.

6. The method as described in claim 5, characterized in that, The configuration of the audio optimization thread based on several second audio data tuples and the fourth evaluation description data in the specified matrix includes: Several second audio data tuples are input into the audio optimization thread to obtain the overall fifth evaluation description data of each digital audio data in each second audio data tuple. By combining the fourth evaluation description data and the fifth evaluation description data in the specified matrix, a third quantitative evaluation model is generated, and the audio optimization thread is configured using the third quantitative evaluation model. The fifth evaluation description data includes the sixth and seventh evaluation description data; the step of combining the fourth and fifth evaluation description data in the specified matrix to generate the third quantitative evaluation model includes: For each second audio data tuple, a sixth evaluation description data for example digital audio data in the second audio data tuple is generated, and the example mined digital audio data in the second audio data tuple is determined as the first sample audio data corresponding to the audio data tuple, and a seventh evaluation description data for the first sample audio data is generated. Select the fourth evaluation description data that best matches the sixth evaluation description data of the example digital audio data in the second audio data binary from the specified matrix, and determine the selected fourth evaluation description data as the eighth evaluation description data corresponding to the negative audio data binary. A third quantitative evaluation model is generated using the sixth evaluation description data, the seventh evaluation description data, and the eighth evaluation description data. The process includes, after obtaining the overall fifth evaluation description data of each digital audio data in each second audio data tuple, optimizing the specified matrix based on the overall fifth evaluation description data of each digital audio data.

Citation Information

Patent Citations

  • Melody extraction method and melody recognition system for audio files

    CN102063904A

  • Audio processing method and device, electronic equipment and storage medium

    CN112489682A