A multi-channel audio stream splitting processing method based on virtual reality
By generating a reference data stream feature set to clean and identify audio anomalies, the problem of anomalies in audio stream splitting processing in virtual reality technology is solved, ensuring the confidence level of the audio splitting processing results.
Patent Information
- Application Number
- CN202211319044.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-26
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2042-10-26
AI Technical Summary
In virtual reality technology, there are abnormal issues in audio stream splitting processing, which affect the confidence level of audio stream splitting processing.
By generating a reference data stream feature set, the audio type information and audio anomaly information in the audio data text to be processed are cleaned and the reference intelligent optimization thread is used for stream splitting.
Ensure the confidence level of the audio splitting processing results and accurately separate abnormal information from the audio data text to be processed.
Smart Images

Figure CN115686428B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio data processing technology, and more specifically, to a method for multi-channel audio stream splitting processing based on virtual reality. Background Technology
[0002] In recent years, virtual reality (VR) technology has had a significant impact on the film and entertainment market due to its widespread application in the film and television industry. This experience center allows viewers to feel as if they are in a real-world scene, immersing themselves in the virtual environment created by the film. However, with the continuous innovation of VR technology, when it is combined with audio stream splitting processing technology, audio information anomalies may occur, potentially interfering with the audio stream. This makes it difficult to ensure the confidence level of the audio splitting processing results. Summary of the Invention
[0003] To address the technical problems existing in related technologies, this application provides a multi-channel audio stream splitting processing method based on virtual reality.
[0004] In a first aspect, a method for multi-channel audio stream splitting based on virtual reality is provided. The method includes at least: obtaining reference audio interaction data and audio data text to be processed corresponding to the reference audio interaction data; performing feature extraction processing on the audio data text to be processed and the reference audio interaction data through a configured reference intelligent optimization thread for text training to determine an audio feature file; wherein the audio feature file is used to represent audio type information and / or audio anomaly information corresponding to the audio data text to be processed; and performing splitting processing on the audio data text to be processed through the reference intelligent optimization thread and the audio feature file to obtain an audio splitting processing result.
[0005] In one standalone embodiment, the step of splitting the audio data text to be processed using the reference intelligent optimization thread and the audio feature file to obtain an audio splitting processing result includes: performing feature extraction processing on the audio feature file and the audio data text to be processed using the reference intelligent optimization thread to determine a reference data stream feature set; wherein, the reference data stream feature set is used to clean the audio type information in the audio data text to be processed, and / or to identify audio anomaly information in the audio data text to be processed; and performing splitting processing on the audio data text to be processed using the reference data stream feature set to generate the audio splitting processing result.
[0006] In one independently implemented embodiment, the step of performing feature extraction processing on the audio feature file and the audio data text to be processed through the reference intelligent optimization thread to determine a reference data stream feature set includes: performing feature extraction processing on the audio feature file and the audio data text to be processed through the reference intelligent optimization thread to determine a plurality of first data stream feature sets with different timbres; determining the first data stream feature set with the first timbres as the data stream feature set to be processed, performing recognition processing on the data stream feature set to be processed to determine a second data stream feature set; wherein, the recognition processing includes feature extraction processing and / or feature derivation processing; combining the second data stream feature set and the first data stream feature set with the same timbre as the second data stream feature set to determine a third data stream feature set; determining the third data stream feature set as the optimized data stream feature set to be processed, and feeding it back to the step of performing recognition processing on the data stream feature set to be processed to determine the second data stream feature set, until the timbre of the determined third data stream feature set is the same as the second timbre corresponding to the first data stream feature set, and determining the third data stream feature set corresponding to the second timbre as the reference data stream feature set.
[0007] In one standalone embodiment, the reference intelligent optimization thread is configured according to the following steps: obtaining a configuration example, wherein the configuration example includes example audio interaction data, a first audio data text and a second audio data text corresponding to the example audio interaction data, the accuracy of the first audio data text exceeding the accuracy of the second audio data text; configuring the intelligent optimization thread to be configured through the configuration example to obtain the reference intelligent optimization thread.
[0008] In one independent implementation, the configuration example further includes an example digital signal set and an example analog signal set corresponding to the example audio interaction data. Based on the reference intelligent optimization thread including multiple timbre recognition threads and noise optimization threads, configuring the intelligent optimization thread to be configured through the configuration example to obtain the reference intelligent optimization thread includes: loading the second audio data text and the example audio interaction data into the multiple timbre recognition threads to obtain an evaluation audio feature file, an evaluation analog signal set, and an evaluation digital signal set corresponding to the example audio interaction data; loading the evaluation audio feature file and the second audio data text into the noise optimization thread to generate the evaluation audio data text; and configuring the intelligent optimization thread to be configured by combining at least one of the following text information: the evaluation analog signal set and the example analog signal set, the evaluation digital signal set and the example digital signal set, the evaluation audio data text and the first audio data text, and the evaluation audio data text and the second audio data text, to obtain the reference intelligent optimization thread.
[0009] In one standalone embodiment, configuring the intelligent optimization thread to be configured, and obtaining the reference intelligent optimization thread, by combining at least one of the following textual information: the evaluation analog signal set and the example analog signal set, the evaluation digital signal set and the example digital signal set, the evaluation audio data text and the first audio data text, and the evaluation audio data text and the second audio data text, includes: determining a first quantization evaluation vector representing audio anomaly information error and a second quantization evaluation vector representing error tolerance by combining the evaluation audio data text and the first audio data text; combining the evaluation analog signal set and the example analog signal set... The system uses a signal set to determine a third quantization evaluation vector to represent semantic information error; combines the evaluation digital signal set and the example digital signal set to determine a fourth quantization evaluation vector to represent normal information error; and based on the second audio data text and the evaluation audio data text, determines a fifth quantization evaluation vector to represent invalidity; combines at least one of the first, second, third, fourth, and fifth quantization evaluation vectors to determine a reference quantization evaluation vector; and combines the reference quantization evaluation vector to configure the intelligent optimization thread to obtain the reference intelligent optimization thread.
[0010] In one standalone embodiment, based on the reference quantization evaluation vector including a fourth quantization evaluation vector, the step of determining the fourth quantization evaluation vector for representing the normal information error by combining the evaluation digital signal set and the example digital signal set includes: combining the evaluation digital signal set and the example digital signal set to generate an association relationship between first discrete information of each first audio waveform in the evaluation digital signal set and second discrete information of second audio waveforms associated with the first audio waveforms in the example digital signal set; and generating the fourth quantization evaluation vector by using the association relationship corresponding to each first audio waveform and the number of first audio waveforms.
[0011] In one independent embodiment, based on the reference quantization evaluation vector including the second quantization evaluation vector, the step of determining the second quantization evaluation vector for representing the error tolerance by combining the evaluation audio data text and the first audio data text includes: generating a first feature in a first dimension and a second feature in a second dimension for each audio waveform in the evaluation audio data text; generating a third feature in a first dimension and a fourth feature in a second dimension for each audio waveform in the first audio data text; and generating the second quantization evaluation vector by combining the first feature and the second feature corresponding to each third audio waveform in the evaluation audio data text, and the third feature and the fourth feature corresponding to a fourth audio waveform associated with the third audio waveform in the first audio data text.
[0012] In one standalone embodiment, based on the reference quantization evaluation vector including the fifth quantization evaluation vector, the step of determining the fifth quantization evaluation vector to represent invalidity based on the second audio data text and the evaluation audio data text includes: performing feature extraction processing on the evaluation audio data text and the second audio data text respectively through a configured feature extraction processing thread to generate a first reference data stream feature set corresponding to the evaluation audio data text and a second reference data stream feature set corresponding to the second audio data text; and combining the first reference data stream feature set and the second reference data stream feature set to generate the fifth quantization evaluation vector.
[0013] The present application provides a method for multi-channel audio stream splitting based on virtual reality. By generating a reference data stream feature set, the reference data stream feature set can clean the audio type information in the audio data text to be processed and / or identify the audio abnormal information in the audio data text to be processed. Then, the audio data text to be processed is split through the reference data stream feature set to accurately separate the abnormal information in the audio data text to be processed. In this way, the confidence of the audio splitting processing result can be ensured. Attached Figure Description
[0014] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This is a flowchart of a multi-channel audio stream splitting processing method based on virtual reality, provided as an embodiment of this application. Detailed Implementation
[0016] To better understand the above technical solutions, the technical solutions of this application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of this application and the specific features in the embodiments are detailed descriptions of the technical solutions of this application, rather than limitations on the technical solutions of this application. In the absence of conflict, the embodiments of this application and the technical features in the embodiments can be combined with each other.
[0017] Please see Figure 1 This paper illustrates a multi-channel audio stream splitting processing method based on virtual reality, which may include the technical solutions described in steps S101-S103.
[0018] S101, obtain the reference audio interaction data and the corresponding audio data text to be processed.
[0019] S102, through the configured reference intelligent optimization thread used for text training, feature extraction processing is performed on the audio data text to be processed and the reference audio interaction data to determine the audio feature file; wherein, the audio feature file is used to represent the audio type information and / or audio anomaly information corresponding to the audio data text to be processed.
[0020] S103, by referring to the intelligent optimization thread and audio feature file, performs split processing on the audio data text to be processed, and obtains the audio split processing result.
[0021] Understandably, by referencing the intelligent optimization thread, feature extraction processing is performed on the audio data text to be processed and the reference audio interaction data to determine the audio feature file. This audio feature file can be used to represent the audio type information and / or audio anomaly information contained in the audio data text to be processed. Then, by referring to the intelligent optimization thread and the audio feature file, the audio data text to be processed can be split. For example, the audio type information in the audio data text to be processed can be cleaned, and / or abnormal audio anomaly information in the audio data text to be processed can be identified, resulting in a highly accurate audio splitting processing result.
[0022] The following provides further limitations on S101-S103.
[0023] In S101, the reference audio interaction data can be a random type of audio interaction data in the real-time scene. The reference audio interaction data can be either the first audio interaction data or the second audio interaction data.
[0024] In this way, the audio data text to be processed and the reference audio interaction data are associated audio interaction data.
[0025] In S102 and S103, the reference depth thread can include multiple timbre recognition threads and noise optimization threads. The audio data text to be processed and the reference audio interaction data can be considered as multi-dimensional first audio interaction data. The multiple timbre recognition threads in the reference intelligent optimization thread perform feature extraction processing on the first audio interaction data to determine the audio feature file. The audio feature file represents the audio category information and / or audio anomaly information corresponding to the audio data text to be processed. Then, the audio feature file and the audio data text to be processed can be considered as multi-dimensional second audio interaction data. The noise optimization thread in the reference intelligent optimization thread performs feature extraction processing on the second audio interaction data to obtain the audio splitting processing result after splitting.
[0026] In one possible implementation, in S103, the audio data text to be processed is split by referring to the intelligent optimization thread and the audio feature file to obtain the audio splitting processing result, which may specifically include the following steps.
[0027] S1031, the audio feature file and the audio data text to be processed are processed by reference intelligent optimization thread to determine the reference data stream feature set; wherein, the reference data stream feature set is used to clean the audio type information in the audio data text to be processed, and / or to identify audio abnormal information in the audio data text to be processed.
[0028] S1032, the audio data text to be processed is split into streams using the reference data stream feature set, and the audio splitting processing result is determined.
[0029] It is understandable that by generating a reference data stream feature set, the audio type information in the audio data text to be processed can be cleaned and / or the audio abnormal information in the audio data text to be processed can be identified. Then, by using the reference data stream feature set to perform split processing on the audio data text to be processed, the abnormal information in the audio data text to be processed can be accurately separated. In this way, the confidence of the audio split processing result can be ensured.
[0030] In S1031, the audio feature file and the audio data text to be processed can be processed by referring to the noise optimization thread in the intelligent optimization thread to determine the reference data stream feature set.
[0031] In one possible implementation, in S1031, feature extraction processing is performed on the audio feature file and the audio data text to be processed by referring to the intelligent optimization thread to determine the reference data stream feature set, which may include the content described in steps A1-A4.
[0032] Step A1 involves using a reference intelligent optimization thread to perform feature extraction on the audio feature file and the audio data text to be processed, thereby determining several first data stream feature sets with different timbres.
[0033] Step A2: The first data stream feature set of the first timbre is determined as the data stream feature set to be processed, and the data stream feature set to be processed is identified to determine the second data stream feature set; wherein, the identification process includes feature extraction processing and / or feature derivation processing.
[0034] Step A3: Determine the third data stream feature set based on the second data stream feature set and the first data stream feature set with the same timbre as the second data stream feature set.
[0035] Step A4: The third data stream feature set is determined as the optimized data stream feature set to be processed, and fed back to the step of identifying and processing the data stream feature set to be processed and determining the second data stream feature set, until the timbre of the determined third data stream feature set is the same as the second timbre corresponding to the first data stream feature set, and the third data stream feature set corresponding to the second timbre is determined as the reference data stream feature set.
[0036] Furthermore, the audio feature file and the audio data text to be processed can be considered as multi-dimensional second audio interaction data. The noise optimization thread in the reference intelligent optimization thread is used to perform feature extraction processing on the second audio interaction data to obtain a first data stream feature set of the first timbre. Then, feature extraction processing can be performed on the first data stream feature set of the first timbre to obtain a first data stream feature set of the second timbre, resulting in several first data stream feature sets of different timbres. For example, the feature extraction process can be performed through a feature extraction unit.
[0037] For example, a first data stream feature set of several different timbres may include a first data stream feature set of a first timbre, a first data stream feature set of a second timbre, and a first data stream feature set of a third timbre, wherein the first timbre exceeds the second timbre, and the second timbre exceeds the third timbre.
[0038] Furthermore, the first data stream feature set of the third timbre (i.e., the first data stream feature set of the first timbre) can be determined as the data stream feature set to be processed. This data stream feature set (i.e., the first data stream feature set of the third timbre) is then subjected to recognition processing (for example, feature extraction processing can be performed through a feature extraction unit, or feature derivation processing can be performed through a derivation unit) to determine the second data stream feature set after the first processing. At this point, the timbre of the second data stream feature set after the first processing can be the third timbre. Then, the second data stream feature set after the first processing can be concatenated with the first data stream feature set that has the same timbre as the second data stream feature set (here, the first data stream feature set of the third timbre) to determine the third data stream feature set after the first processing. Here, the timbre of the third data stream feature set after the first processing can be the third timbre.
[0039] For example, the second data stream feature set after the first processing (i.e., the second data stream feature set of the third timbre) can be associated with the first data stream feature set of the third timbre to determine the third data stream feature set after the first processing; or, the second data stream feature set of the third timbre can be associated with the first data stream feature set of the third timbre, and feature extraction processing can be performed on the associated data stream feature set to determine the third data stream feature set after the first processing; or, the comparison result of the feature vectors at the same feature location in the second data stream feature set of the third timbre and the first data stream feature set of the third timbre can be determined to obtain the abnormal data stream feature set, and the abnormal data stream feature set can be determined as the third data stream feature set after the first processing.
[0040] The third data stream feature set after the first processing can be determined as the optimized data stream feature set to be processed. This optimized feature set is then subjected to recognition processing (e.g., feature extraction and / or feature derivation) to determine the second data stream feature set after the second processing. In this case, the timbre of the second data stream feature set after the second processing can be the second timbre. Furthermore, the second data stream feature set after the second processing and the first data stream feature set with the second timbre can be concatenated to determine the third data stream feature set after the second processing. Here, the timbre of the third data stream feature set after the second processing can also be the second timbre.
[0041] Therefore, the third data stream feature set after the second processing can be determined as the optimized data stream feature set to be processed. This optimized data stream feature set is then subjected to recognition processing to determine the second data stream feature set after the third processing. At this point, the timbre of the second data stream feature set after the third processing can be the first timbre. Furthermore, the third data stream feature set after the third processing can be concatenated with the first data stream feature set of the first timbre to determine the third data stream feature set after the third processing. Here, the timbre of the third data stream feature set after the third processing can be the first timbre. It is known that the first timbre is the second timbre corresponding to the first data stream feature set. Therefore, the third data stream feature set corresponding to the second timbre is determined as the reference data stream feature set, that is, the third data stream feature set of the first timbre is determined as the reference data stream feature set.
[0042] For example, after generating several first data stream feature sets with different timbres, the first data stream feature sets of the first timbre (third timbre) can be processed for identification to determine the second data stream feature set of the second timbre; then the second data stream feature set of the second timbre can be associated with the first data stream feature set of the second timbre, and the associated data stream feature set of the second timbre can be processed for feature extraction to determine the second data stream feature set of the first timbre; then the second data stream feature set of the first timbre can be associated with the first data stream feature set of the first timbre, and the associated data stream feature set of the first timbre can be processed for feature extraction to determine the second data stream feature set of the first timbre. This second data stream feature set of the first timbre (second timbre) can be a reference data stream feature set.
[0043] In step S1032, the reference data stream feature set can be concatenated with the audio data text to be processed to determine the audio splitting processing result. Alternatively, at least one feature extraction process can be performed on the reference data stream feature set, and the data stream feature set after at least one feature extraction process can be concatenated with the audio data text to be processed to determine the audio splitting processing result. For example, the feature concatenation process can involve weighting the feature vectors at the same feature locations in the reference data stream feature set and the audio data text to be processed; or, the reference data stream feature set can be associated with the audio data text to be processed, and the associated data stream feature set can be processed by the feature extraction unit.
[0044] The multi-channel audio stream splitting processing method based on virtual reality is further explained. The audio data text to be processed and the reference audio interaction data are loaded into multiple timbre recognition threads within a reference intelligent optimization thread. The feature extraction processing architecture in these threads performs feature extraction on the audio data text and the reference audio interaction data to determine a second intermediate data stream feature set. This second intermediate data stream feature set is then loaded into a first verification sub-thread corresponding to the audio feature file, resulting in the audio feature file. Furthermore, the configuration of the reference intelligent optimization thread may also include a second verification sub-thread corresponding to the analog signal set and a third verification sub-thread corresponding to the digital signal set.
[0045] In one possible implementation, the reference intelligent optimization thread can be configured according to the following steps.
[0046] Step B1: Obtain a configuration example, wherein the configuration example includes example audio interaction data, a first audio data text and a second audio data text corresponding to the example audio interaction data, and the accuracy of the first audio data text exceeds the accuracy of the second audio data text.
[0047] Step B2: Configure the smart optimization thread to be configured using the configuration example to obtain the reference smart optimization thread.
[0048] The example audio interaction data can be either the second audio interaction data or the first audio interaction data. The example audio interaction data corresponds to the first audio data text and the second audio data text. The accuracy of the first audio data text exceeds that of the second audio data text, meaning that the second audio data text can be considered as real-time data of the first audio data text.
[0049] The example audio interaction data here may also include the example digital signal set and the example analog signal set corresponding to the example audio interaction data.
[0050] For example, sample audio interaction data can be loaded into the configured intelligent differentiation thread to obtain the sample analog signal set corresponding to the sample audio interaction data.
[0051] In one possible implementation, the configuration example further includes a set of example digital signals and a set of example analog signals corresponding to the example audio interaction data. Based on the reference intelligent optimization thread including multiple timbre recognition threads and noise optimization threads, step B2 involves configuring the intelligent optimization thread to be configured through the configuration example to obtain the reference intelligent optimization thread, which may specifically include the following:
[0052] Step B21: Load the second audio data text and the example audio interaction data into multiple timbre recognition threads to obtain the evaluation audio feature file, evaluation analog signal set, and evaluation digital signal set corresponding to the example audio interaction data.
[0053] Step B22: Load the evaluation audio feature file and the second audio data text into the noise optimization thread to determine the evaluation audio data text.
[0054] Step B23: Combining at least one type of text information from the evaluation analog signal set and the example analog signal set, the evaluation digital signal set and the example digital signal set, the evaluation audio data text and the first audio data text, and the evaluation audio data text and the second audio data text, configure the intelligent optimization thread to be configured, and obtain the reference intelligent optimization thread.
[0055] The processes of steps B21 and B22 can be referred to the aforementioned explanation of the reference intelligent optimization thread, and will not be described in detail here.
[0056] In step B23, an audio set is constructed by evaluating the analog signal set and the example analog signal set, an audio set is constructed by evaluating the digital signal set and the example digital signal set, an audio set is constructed by evaluating the audio data text and the first audio data text, and an audio set is constructed by evaluating the audio data text and the second audio data text, resulting in four audio sets. Based on at least one of the four audio sets, a smart optimization thread to be configured can be configured to obtain a reference smart optimization thread.
[0057] In one possible implementation, step B23, combining at least one of the following text information—the evaluation analog signal set and the example analog signal set, the evaluation digital signal set and the example digital signal set, the evaluation audio data text and the first audio data text, and the evaluation audio data text and the second audio data text—configures the intelligent optimization thread to be configured, resulting in a reference intelligent optimization thread, which may include the following:
[0058] Step C1 involves determining a first quantization evaluation vector to represent audio anomaly information error and a second quantization evaluation vector to represent error tolerance based on the evaluation audio data text and the first audio data text; determining a third quantization evaluation vector to represent semantic information error based on the evaluation analog signal set and the example analog signal set; determining a fourth quantization evaluation vector to represent normal information error based on the evaluation digital signal set and the example digital signal set; and determining a fifth quantization evaluation vector to represent invalidity based on the second audio data text and the evaluation audio data text.
[0059] Step C2: Determine a reference quantization evaluation vector based on at least one of the first quantization evaluation vector, the second quantization evaluation vector, the third quantization evaluation vector, the fourth quantization evaluation vector, and the fifth quantization evaluation vector.
[0060] Step C3: Based on the reference quantization evaluation vector, configure the intelligent optimization thread to be configured to obtain the reference intelligent optimization thread.
[0061] In this way, by configuring several quantization evaluation vectors, a reference quantization evaluation vector can be determined through at least one quantization evaluation vector. When configuring the intelligent optimization thread through the reference quantization evaluation vector, the accuracy of the obtained reference intelligent optimization thread can be improved.
[0062] In one possible implementation, the determination of a second quantization evaluation vector to represent the error tolerance, based on the evaluation audio data text and the first audio data text, on the basis of a reference quantization evaluation vector including a second quantization evaluation vector, may include the following.
[0063] Step D1: Determine the first feature of each audio waveform in the evaluation audio data text in the first dimension and the second feature in the second dimension.
[0064] Step D2: Determine the third feature of each audio waveform in the first audio data text in the first dimension and the fourth feature in the second dimension.
[0065] Step D3: Based on the first and second features corresponding to each third audio waveform in the evaluation audio data text, and the third and fourth features corresponding to the fourth audio waveform associated with the third audio waveform in the first audio data text, determine the second quantization evaluation vector.
[0066] In one possible implementation, based on the reference quantization evaluation vector including the fourth quantization evaluation vector, a fourth quantization evaluation vector for representing the normal information error is determined based on the evaluation digital signal set and the example digital signal set, which may include the following.
[0067] Step E1: Based on the evaluation digital signal set and the example digital signal set, determine the correlation between the first discrete information of each first audio waveform in the evaluation digital signal set and the second discrete information of the second audio waveform associated with the first audio waveform in the example digital signal set.
[0068] Step E2: Determine the fourth quantization evaluation vector by considering the correlation between each first audio waveform and the number of first audio waveforms.
[0069] In this way, the first audio waveform can evaluate various audio waveforms in the digital signal set, and the second audio waveform can be an audio waveform in the example digital signal set that has the same location as the first audio waveform.
[0070] In one possible implementation, based on the reference quantization evaluation vector including the fifth quantization evaluation vector, determining a fifth quantization evaluation vector to represent invalidity based on the second audio data text and the evaluation audio data text may include the following.
[0071] Step F1 involves performing feature extraction processing on the evaluation audio data text and the second audio data text respectively through the configured feature extraction processing thread, to determine the first reference data stream feature set corresponding to the evaluation audio data text and the second reference data stream feature set corresponding to the second audio data text.
[0072] Step F2: Determine the fifth quantization evaluation vector based on the first reference data stream feature set and the second reference data stream feature set.
[0073] In step F2, the fifth quantization evaluation vector is determined based on the first reference data stream feature set and the second reference data stream feature set, which may include the following:
[0074] Step F21: Determine the comparison result between the first feature vector of each first feature point in the first reference data stream feature set and the second feature vector of the second feature point associated with the first feature point in the second reference data stream feature set.
[0075] Step F22: The sum of squares of the comparison results corresponding to each first feature point is determined as the fifth quantization evaluation vector.
[0076] In implementation, the configured feature extraction processing thread can be a CNN thread. The configured CNN thread performs feature extraction processing on the evaluation audio data text and the second audio data text respectively, determining the first reference data stream feature set corresponding to the evaluation audio data text and the second reference data stream feature set corresponding to the second audio data text.
[0077] In steps C2 and C3, during implementation, any one of the first, second, third, fourth, and fifth quantization evaluation vectors can be determined as the reference quantization evaluation vector; alternatively, multiple quantization evaluation vectors among the first, second, third, fourth, and fifth quantization evaluation vectors can be weighted to obtain the reference quantization evaluation vector. For example, the sum of the first and second quantization evaluation vectors can be used as the reference quantization evaluation vector, or the sum of the first, second, third, fourth, and fifth quantization evaluation vectors can be used as the reference quantization evaluation vector. Then, the reference quantization evaluation vector can be used to configure the intelligent optimization thread to be configured, thus obtaining the reference intelligent optimization thread.
[0078] Based on the above, a multi-channel audio stream splitting and processing device 200 based on virtual reality is provided, which is applied to a multi-channel audio stream splitting and processing system based on virtual reality. The device includes:
[0079] Data acquisition module 210 is used to acquire reference audio interaction data and the audio data text to be processed corresponding to the reference audio interaction data;
[0080] The file determination module 220 is used to perform feature extraction processing on the audio data text to be processed and the reference audio interaction data through a configured reference intelligent optimization thread for text training, and determine the audio feature file; wherein, the audio feature file is used to represent the audio type information and / or audio anomaly information corresponding to the audio data text to be processed;
[0081] The result determination module 230 is used to perform split processing on the audio data text to be processed through the reference intelligent optimization thread and the audio feature file to obtain the audio split processing result.
[0082] Based on the above, a virtual reality-based multi-channel audio stream splitting and processing system 300 is shown, including a processor 310 and a memory 320 that communicate with each other. The processor 310 is used to read computer programs from the memory 320 and execute them to implement the above-described method.
[0083] Based on the above, a computer-readable storage medium is also provided, on which a computer program stored implements the above method during runtime.
[0084] In summary, based on the above scheme, by generating a reference data stream feature set, the audio type information in the audio data text to be processed can be cleaned and / or the audio abnormal information in the audio data text to be processed can be identified. Then, by using the reference data stream feature set to perform split processing on the audio data text to be processed, the abnormal information in the audio data text to be processed can be accurately separated. In this way, the confidence of the audio split processing result can be ensured.
[0085] It should be understood that the systems and modules described above can be implemented in various ways. For example, in some embodiments, the systems and modules can be implemented by hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated-design hardware. Those skilled in the art will understand that the methods and systems described above can be implemented using computer-executable instructions and / or included in processor control code, for example, such code provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The systems and modules of this application can be implemented not only by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field-programmable gate arrays, programmable logic devices, etc., but also by software executed by various types of processors, or by a combination of the aforementioned hardware circuits and software (e.g., firmware).
[0086] It should be noted that different embodiments may produce different beneficial effects. In different embodiments, the beneficial effects may be any one or a combination of the above, or any other possible beneficial effects.
[0087] The basic concepts have been described above. Obviously, for those skilled in the art, the detailed disclosure above is merely illustrative and does not constitute a limitation of this application. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this application. Such modifications, improvements, and corrections are suggested in this application, and therefore remain within the spirit and scope of the exemplary embodiments of this application.
[0088] Furthermore, this application uses specific terms to describe embodiments of the application. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of the application. Therefore, it should be emphasized and noted that "an embodiment," "one embodiment," or "an alternative embodiment" mentioned twice or more in different locations in this specification do not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of the application can be appropriately combined.
[0089] Furthermore, those skilled in the art will understand that aspects of this application can be described and illustrated through several patentable types or situations, including any new and useful combination of processes, machines, products, or substances, or any new and useful improvements thereof. Accordingly, aspects of this application can be implemented entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. All of the above hardware or software may be referred to as a “data block,” “module,” “engine,” “unit,” “component,” or “system.” Furthermore, aspects of this application may manifest as a computer product located on one or more computer-readable media, the product including computer-readable program code.
[0090] Computer storage media may contain a propagated data signal containing computer program code, for example, on baseband or as part of a carrier wave. This propagated signal may take various forms, including electromagnetic, optical, and suitable combinations thereof. Computer storage media can be any computer-readable medium other than a computer-readable storage medium, which can be connected to an instruction execution system, apparatus, or device to enable communication, propagation, or transmission of a program for use. The program code located on the computer storage medium can be propagated through any suitable medium, including radio, cable, fiber optic cable, RF, or similar media, or any combination of the above media.
[0091] The computer program code required for the operation of each part of this application can be written in any one or more programming languages, including object-oriented programming languages such as Java, Scala, Smalltalk, Eiffel, JADE, Emerald, C++, C#, VB.NET, Python, etc., conventional procedural programming languages such as C, Visual Basic, Fortran 2003, Perl, COBOL 2002, PHP, ABAP, dynamic programming languages such as Python, Ruby, and Groovy, or other programming languages. This program code can run entirely on the user's computer, or as a standalone software package on the user's computer, or partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any network, such as a local area network (LAN) or wide area network (WAN), or connected to an external computer (e.g., via the Internet), or in a cloud computing environment, or used as a service such as Software as a Service (SaaS).
[0092] Furthermore, unless expressly stated in the claims, the order of processing elements and sequences, the use of numbers and letters, or other names described in this application are not intended to limit the order of the processes and methods of this application. Although the foregoing disclosure has discussed some currently considered useful embodiments of the invention through various examples, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments; rather, the claims are intended to cover all modifications and equivalent combinations that conform to the substance and scope of the embodiments of this application. For example, while the system components described above can be implemented using hardware devices, they can also be implemented solely through software solutions, such as installing the described system on existing servers or mobile devices.
[0093] Similarly, it should be noted that, in order to simplify the description of the present application and thus aid in the understanding of one or more embodiments of the invention, the foregoing description of the embodiments of the present application sometimes combines multiple features into a single embodiment, drawing, or description thereof. However, this disclosure method does not imply that the subject matter of the application requires more features than those mentioned in the claims. In fact, the embodiments contain fewer features than all the features of the single embodiments disclosed above.
[0094] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of embodiments are modified in some examples with the terms "approximately," "approximately," or "generally." Unless otherwise stated, "approximately," "approximately," or "generally" indicates that the numbers are open to adaptive variation. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may be changed depending on the characteristics required by individual embodiments. In some embodiments, numerical parameters are taken into account a specified number of significant digits and employ a general method of digit reservation. Although the numerical ranges and parameters used to confirm their breadth of application in some embodiments of this application are approximate values, in specific embodiments, such values are set as precisely as feasible.
[0095] For each patent, patent application, patent application publication, and other material such as articles, books, specifications, publications, and documents referenced in this application, the entire contents of that patent are incorporated herein by reference. This excludes historical application documents that are inconsistent with or conflict with the content of this application, as well as documents that limit the broadest scope of the claims in this application (currently or subsequently appended to this application). It should be noted that if there are any inconsistencies or conflicts between the descriptions, definitions, and / or terminology used in the supplementary materials of this application and the content of this application, the descriptions, definitions, and / or terminology used in this application shall prevail.
[0096] Finally, it should be understood that the embodiments described in this application are merely illustrative of the principles of the embodiments of this application. Other modifications may also fall within the scope of this application. Therefore, alternative configurations of the embodiments of this application are considered as examples and not limitations, and are regarded as consistent with the teachings of this application. Accordingly, the embodiments of this application are not limited to the embodiments explicitly described and illustrated in this application.
[0097] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for virtual reality based multi-channel audio stream offloading, the method comprising: The method comprises at least: obtaining reference audio interaction data and the corresponding to-be-processed audio data text of the reference audio interaction data; performing feature extraction processing on the to-be-processed audio data text and the reference audio interaction data through the configured reference intelligent optimization thread for text training, to determine an audio feature file; wherein the audio feature file is used to represent audio category information and / or audio abnormal information corresponding to the to-be-processed audio data text; performing shunt processing on the to-be-processed audio data text through the reference intelligent optimization thread and the audio feature file, to obtain an audio shunt processing result; The shunt processing on the to-be-processed audio data text through the reference intelligent optimization thread and the audio feature file to obtain an audio shunt processing result comprises: performing feature extraction processing on the audio feature file and the to-be-processed audio data text through the reference intelligent optimization thread, to determine a reference data stream feature set; wherein the reference data stream feature set is used to clean audio category information in the to-be-processed audio data text, and / or is used to identify audio abnormal information in the to-be-processed audio data text; performing shunt processing on the to-be-processed audio data text through the reference data stream feature set, to generate the audio shunt processing result; The feature extraction processing on the audio feature file and the to-be-processed audio data text through the reference intelligent optimization thread to determine a reference data stream feature set comprises: performing feature extraction processing on the audio feature file and the to-be-processed audio data text through the reference intelligent optimization thread, to determine a plurality of first data stream feature sets of different tones; determining a first data stream feature set of a first tone as a to-be-processed data stream feature set, performing identification processing on the to-be-processed data stream feature set, to determine a second data stream feature set; wherein the identification processing comprises feature extraction processing and / or feature derivation processing; combining the second data stream feature set and a first data stream feature set of the same tone as the second data stream feature set, to determine a third data stream feature set; determining the third data stream feature set as an optimized to-be-processed data stream feature set, and feeding back to the step of performing identification processing on the to-be-processed data stream feature set to determine a second data stream feature set, until the tone of the determined third data stream feature set is the same as a second tone corresponding to the first data stream feature set, and the third data stream feature set corresponding to the second tone is determined as the reference data stream feature set.
2. The method of claim 1, wherein, The reference intelligent optimization thread is configured according to the following steps: obtaining a configuration example, wherein the configuration example comprises example audio interaction data, a first audio data text and a second audio data text corresponding to the example audio interaction data, and the accuracy of the first audio data text is higher than that of the second audio data text; configuring a to-be-configured intelligent optimization thread through the configuration example, to obtain the reference intelligent optimization thread.
3. The method of claim 2, wherein, In the configuration example further includes an example audio interaction data corresponding to the example digital signal set and the example analog signal set, the reference intelligent optimization thread includes a plurality of tone recognition threads and noise optimization threads, based on the configuration example, the configuration of the to-be-configured intelligent optimization thread is obtained by the reference intelligent optimization thread, including: Load the second audio data text and the example audio interaction data into the plurality of tone recognition threads to obtain the evaluation audio feature file, the evaluation analog signal set, and the evaluation digital signal set corresponding to the example audio interaction data; Load the evaluation audio feature file and the second audio data text into the noise optimization thread to generate evaluation audio data text; Combine at least one of the evaluation analog signal set and the example analog signal set, the evaluation digital signal set and the example digital signal set, the evaluation audio data text and the first audio data text, and the evaluation audio data text and the second audio data text to configure the to-be-configured intelligent optimization thread to obtain the reference intelligent optimization thread.
4. The method according to claim 2 or 3, characterized in that, Combine at least one of the evaluation analog signal set and the example analog signal set, the evaluation digital signal set and the example digital signal set, the evaluation audio data text and the first audio data text, and the evaluation audio data text and the second audio data text to configure the to-be-configured intelligent optimization thread to obtain the reference intelligent optimization thread, including: Determine a first quantization evaluation vector for representing audio abnormal information error and a second quantization evaluation vector for representing error tolerance based on the evaluation audio data text and the first audio data text; and determine a third quantization evaluation vector for representing semantic information error based on the evaluation analog signal set and the example analog signal set; Determine a fourth quantization evaluation vector for representing normal information error based on the evaluation digital signal set and the example digital signal set; and determine a fifth quantization evaluation vector for representing invalid based on the second audio data text and the evaluation audio data text; Determine a reference quantization evaluation vector based on at least one of the first quantization evaluation vector, the second quantization evaluation vector, the third quantization evaluation vector, the fourth quantization evaluation vector, and the fifth quantization evaluation vector; Configure the to-be-configured intelligent optimization thread based on the reference quantization evaluation vector to obtain the reference intelligent optimization thread.
5. The method of claim 4, wherein, Based on the reference quantization evaluation vector including the fourth quantization evaluation vector, the fourth quantization evaluation vector for representing normal information error based on the evaluation digital signal set and the example digital signal set includes: Generate the association relationship between the first discrete information of each first audio waveform in the evaluation digital signal set and the second discrete information of the second audio waveform associated with the first audio waveform in the example digital signal set based on the evaluation digital signal set and the example digital signal set. The fourth quantization evaluation vector is generated according to the association relationship corresponding to each first audio waveform and the number of the first audio waveforms.
6. The method of claim 4, wherein, On the basis that the reference quantization evaluation vector comprises a second quantization evaluation vector, the second quantization evaluation vector used for representing error permission is determined by combining the evaluation audio data text and the first audio data text, and the second quantization evaluation vector comprises: first features of each audio waveform in the evaluation audio data text in a first dimension and second features of each audio waveform in the evaluation audio data text in a second dimension are generated; third features of each audio waveform in the first audio data text in the first dimension and fourth features of each audio waveform in the first audio data text in the second dimension are generated; the second quantization evaluation vector is generated by combining the first features and the second features of each third audio waveform in the evaluation audio data text and the third features and the fourth features of a fourth audio waveform associated with the third audio waveform in the first audio data text.
7. The method of claim 4, wherein, On the basis that the reference quantization evaluation vector comprises a fifth quantization evaluation vector, the fifth quantization evaluation vector used for representing invalidity is determined based on the second audio data text and the evaluation audio data text, and the fifth quantization evaluation vector comprises: first reference data stream feature sets corresponding to the evaluation audio data text and second reference data stream feature sets corresponding to the second audio data text are generated by performing feature extraction processing on the evaluation audio data text and the second audio data text respectively through the configured feature extraction processing thread; and the fifth quantization evaluation vector is generated by combining the first reference data stream feature sets and the second reference data stream feature sets.
Citation Information
Patent Citations
Audio processing method and device, readable storage medium and electronic equipment
CN112562649A
Method and system for detecting forged audio and storage medium
CN113409771A