Distributed computing processing method applied to voiceprint waveform visualization
Through distributed computing processing methods and advanced visualization tools, the problem of low efficiency of large-scale voice data processing in the existing technology is solved, and the rapid and accurate visualization of voice ripple waveforms is achieved, and the efficiency and visualization effect of data processing are improved.
Patent Information
- Application Number
- CN202510317385.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-17
AI Technical Summary
The existing voice ripple waveform visualization methods are computationally expensive, inefficient, and lack real-time performance when processing large-scale voice data, making it difficult to detect abnormalities in a timely manner, resulting in potentially dangerous accidents.
The distributed computing processing method is used to convert and process the sound ripple waveform data format into an array, and visual data is generated through CPU parallel computing or multi-threading methods. Advanced visualization tools such as ECharts, Matplotlib, and D3.js are used for multi-angle display.
It significantly improves the calculation efficiency of voice ripple waveforms and the accuracy of data expression, realizes the fast and accurate visualization of large-scale voice data, supports a variety of visualization forms, has efficient user interaction functions, and is suitable for the analysis and processing of various voice data.
Smart Images

Figure CN120164484A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of distributed computing processing, and particularly relates to a distributed computing processing method applied to the visualization of voiceprint waveforms. Background Art
[0002] With the development of information technology and the advent of the data era, voice data, as an important information carrier, has been widely used in the fields of speech recognition, speech analysis, speech synthesis, etc. The voiceprint waveform is an important manifestation of the voice signal. By visualizing the voiceprint waveform, the characteristics of the voice signal can be intuitively displayed, which is convenient for analysis and processing.
[0003] However, the existing voiceprint waveform visualization methods often face the problems of large computational complexity and low efficiency when processing large-scale voice data. The real-time importance of the voiceprint waveform is that if abnormalities are not detected in time, it is very likely to lead to dangerous accidents.
[0004] The processing cycle for visualizing the traditional large amount of voiceprint waveform data after processing is: the processing time for 200,000 pieces of data is 1000 - 3000 ms, and due to the limitation of hardware configuration, the time may be longer; there is no corresponding prompt and alarm instruction issued, which is not convenient for managers to understand in real time and conduct emergency processing, and the applicability has certain limitations.
[0005] Therefore, proposing an efficient distributed computing processing method that can achieve fast and accurate visualization of voiceprint waveforms has important application value. Summary of the Invention
[0006] Compared with the prior art, the present invention significantly improves data accuracy and efficiency, accurately represents the characteristics of the voiceprint waveform with fewer parameters, greatly improves the computing efficiency, overcomes the problems of large computational complexity and low efficiency when the prior art processes large-scale voice data, realizes fast and accurate visualization of the voiceprint waveform, has high computing efficiency, and improves the accuracy of data expression.
[0007] The present invention adopts the following technical solutions to solve the above problems:
[0008] A distributed computing processing method applied to the visualization of voiceprint waveforms, the method comprising the following steps:
[0009] S1: Distributed computing processing: Convert the voiceprint waveform data format, process and store it in an array, and process the array to generate visualization data of the voiceprint waveform;
[0010] S2: Visualization display: Use a visualization tool to visualize the processed data to generate a voiceprint waveform diagram.
[0011] Further, in S1, it specifically includes the following steps:
[0012] S101: Voiceprint waveform feature extraction: Use a function to extract the spectral features of the voiceprint waveform;
[0013] S102: Data acquisition and preprocessing: Acquire the voiceprint waveform data extracted in S101 and perform format conversion into an array;
[0014] S103: Array splitting and parallel processing: Divide the preprocessed array in S102 into multiple sub-arrays, calculate and process each sub-array, and generate visualization data of the voiceprint waveform.
[0015] Furthermore, in S2, the visualization display further includes a user interaction function, provides a user interface, and supports zooming in, zooming out, and panning operations.
[0016] Furthermore, in S2, the visualization tool is ECharts, Matplotlib, D3.js, or other visualization tools with similar functions. The generated visualization graph of the full data volume supports multiple visualization forms, including but not limited to waveform graphs, spectrograms, and three-dimensional voiceprint graphs, realizing multi-angle display of voice signals.
[0017] Furthermore, in S101, the function adopts the short-time Fourier transform (STFT).
[0018] Furthermore, in S102, the voiceprint waveform data is obtained from an HTTP interface or Socket.io. The obtained json or xml data is stored in an array through format conversion processing, and the data integrity is verified.
[0019] Furthermore, in S102, when the length of the voiceprint waveform data exceeds the maximum limit, the array is intercepted and then new data is stored in the array.
[0020] Furthermore, in S103, the way to calculate and process each sub-array is CPU parallel computing or multi-threaded method.
[0021] Furthermore, in S2, for the visualization graphs generated from multiple groups of data, they are integrated through a layer style merging algorithm, and the visualization graphs generated from multiple groups of data are merged to form a visualization graph of the full data volume.
[0022] Furthermore, the voiceprint waveform data includes different languages and dialects, and this method can process and visualize the voice data of different languages and dialects.
[0023] The beneficial effects of the present invention are as follows:
[0024] 1. The present invention proposes a distributed computing and processing method for voiceprint waveform visualization. By adopting distributed computing technology, the burden on a single processing node is reduced, and the overall data processing efficiency is improved. It can efficiently and accurately process and display the voiceprint waveforms of large-scale voice data, and uses a multi-core processor for parallel computing to accelerate the data processing process.
[0025] 2. Through the application of a distributed file system and a database, the efficient storage and management of voice data are realized;
[0026] 3. Through the application of a distributed computing processing framework, the parallel processing and fast calculation of acoustic feature data are realized.
[0027] 4. Through advanced visualization tools and graphics processing algorithms, the multi-angle display of voiceprint waveforms is realized. This method has the advantages of high computing efficiency, fast processing speed, and good visualization effect, and is applicable to the analysis and processing of various types of voice data.
[0028] 5. This method not only improves the visualization effect of voiceprint waveforms, can quickly and accurately realize the visualization display of voiceprint waveforms, significantly improves the data loading speed and display effect, but also has important practical significance in improving application efficiency, reliability, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the specific embodiments of the present invention, the accompanying drawings required for use in the description of the specific embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0030] Figure 1 is the flowchart of the calculation and processing method of the present invention;
[0031] Figure 2 is the schematic diagram of the method for data processing and visualization of the present invention;
[0032] Figure 3 is the flowchart of data processing and visualization of the present invention;
[0033] Figure 4 is the schematic diagram of the sampling data submission sample diagram of Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] To enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0035] The present invention discloses a distributed computing processing method applied to voiceprint waveform visualization, focusing on the field of voiceprint technology. With the help of distributed computing technology, it can effectively process large-scale voiceprint data, and then display voiceprint features, achieving fast visualization of voiceprint waveforms, improving data loading speed and accuracy. Its innovation lies in a unique distributed computing architecture. After obtaining data, through format conversion and array segmentation, parallel computing or multi-threaded processing is used to generate visualization data, which is finally presented in a variety of visualization tools, supporting various forms such as waveform diagrams, spectrograms, and three-dimensional voiceprint diagrams, realizing multi-angle display.
[0036] Compared with traditional technologies, the present invention significantly improves data accuracy and efficiency, accurately represents the characteristics of voiceprint waveforms with fewer parameters, greatly improves computing efficiency, overcomes the problems of large computing volume and low efficiency in processing large-scale voice data in the prior art, realizes fast and accurate visualization of voiceprint waveforms, has high computing efficiency, and improves the accuracy of data expression. It supports rich user interaction functions such as zooming in, zooming out, panning, and annotation. In practical applications, it helps to quickly detect equipment abnormalities. This method not only improves the visualization effect of voiceprint waveforms, but also has important practical significance in improving application efficiency, reliability, etc.
[0037] It should be noted that the distributed computing processing method used in the embodiments of the present invention is the Distributed Computing method, but it is not limited to this method. Other distributed computing methods with similar functions can also achieve the same effect.
[0038] Embodiment 1
[0039] As Figures 1-3 shown, the present invention discloses a distributed computing processing method applied to voiceprint waveform visualization, and this method includes the following steps:
[0040] S1: Distributed computing processing: Convert the format of voiceprint waveform data, process and store it in an array, and process the array to generate visualization data of the voiceprint waveform, which specifically includes the following steps:
[0041] S101: Voiceprint waveform feature extraction: Use a function to extract the spectral features of the voiceprint waveform. Here, the function can adopt the Short-Time Fourier Transform (STFT).
[0042] S102: Data acquisition and preprocessing: Acquire the voiceprint waveform data extracted in S101, and perform format conversion into an array;
[0043] The voiceprint waveform data is obtained from the HTTP interface or Socket.io. The obtained json or xml data is stored in an array through format conversion processing, and the data integrity is verified. When the length of the voiceprint waveform data exceeds the maximum limit, the array is intercepted and then the new data is stored in the array.
[0044] S103: Array splitting and parallel processing: The preprocessed array in S102 is divided into multiple sub-arrays, and each sub-array is calculated and processed to generate the visualization data of the voiceprint waveform. The way to calculate and process each sub-array is CPU parallel computing or multi-threading.
[0045] S2: Visualization display: Use a visualization tool to visualize the processed data and generate a voiceprint waveform diagram.
[0046] The visualization display also includes user interaction functions, provides a user interface, and supports operations such as zooming in, zooming out, and panning; the visualization tool is ECharts, Matplotlib, D3.js or other visualization tools with similar functions. The generated visualization diagram of the full data volume supports multiple visualization forms, including but not limited to waveform diagrams, spectrograms, and three-dimensional voiceprint diagrams, realizing multi-angle display of voice signals; a user-defined visualization interface and analysis tool are also provided to meet the personalized needs of different users.
[0047] For the visualization diagrams generated from multiple groups of data, they are integrated through a layer style merging algorithm. The visualization diagrams generated from multiple groups of data are merged to form a visualization diagram of the full data volume. For waveform diagrams, multiple waveform diagrams can be aligned along the time axis and superimposed for display; for spectrograms, multiple spectrograms can be aligned along the frequency axis and the transparency can be adjusted for display.
[0048] This method further includes data compression and storage optimization to improve the storage and retrieval efficiency of large-scale voice data.
[0049] It should be noted that the voiceprint waveform data in S1 includes different languages and dialects, that is, this method can process and visualize the voice data of different languages and dialects.
[0050] Embodiment 2
[0051] In this embodiment, a distributed computing method is adopted to realize the voiceprint waveform visualization processing of 240,000 pieces of voice data per second, which specifically includes the following steps:
[0052] S1: Distributed computing processing:
[0053] S101: Voiceprint waveform feature extraction:
[0054] Use the short-time Fourier transform (STFT) to extract the spectral features of the voiceprint waveform.
[0055]
[0056] Among them, x(n) is the voice signal, w(n) is the window function, and ω is the frequency.
[0057] S102: Data acquisition and preprocessing:
[0058] Push data to the client through the HTTP interface or WebSocket Sokect.io, obtain the voiceprint waveform data, and provide the data in the form of JSON XML data format, etc. Verify the data integrity and discard abnormal frames (such as signal-to-noise ratio < 20dB). Use the Shannon sampling theorem to ensure that the sampling frequency f s ≥2f max , to avoid spectral aliasing; use the corresponding scientific computing parsing library to parse the data into a dictionary or list form, and then convert it into an array. If the length exceeds the maximum limit, the array will be intercepted and the new data will be stored in the data. Assume that the data array is data_array and the maximum length limit is max_length. If len(data_array) > max_length, the interception operation will be performed: data_array = data_array[:max_length].
[0059] S103: Array segmentation and parallel processing:
[0060] Divide the preprocessed array data_array into n sub-arrays sub_arrays. The division method can be to divide it according to a fixed length. Assume that the length of each sub-array is sub_length, then sub_arrays = [data_array[i:i+sub_length] for i in range(0, len(data_array), sub_length)]; then divide this array into multiple arrays, and use CPU parallel computing or multi-threading to process each sub-array to generate the visualization data of the voiceprint waveform.
[0061] The parallelism calculation formula is P = N / n, where P is the parallelism, N is the total number of tasks, and n is the number of processor cores. For time complexity analysis, assume that the time complexity of each sub-task is O(f(n)), then the total time complexity is O(P·f(n)) = O(n / n·f(n)) = O(f(n)), indicating that parallel computing can significantly reduce the processing time.
[0062] S2: Visualization display:
[0063] (1) Data visualization:
[0064] Use tools such as ECharts, Matplotlib or D3.js to draw the generated voiceprint waveform data into waveform graphs, spectrum graphs, etc.; for visualization graphs generated by multiple sets of data, integrate them through layer style merging algorithms. For example, for waveform graphs, multiple sets of waveform graphs can be aligned by time axis and displayed in superposition; for spectrum graphs, multiple sets of spectrum graphs can be aligned by frequency axis and displayed with transparency adjusted; layer style merging, merge visualization graphs generated by multiple sets of data to form a visualization graph with full data volume; dynamic rendering: accelerate spectrum graph generation based on WebGL, frame rate ≥ 30fps.
[0065] (2) User interaction function: Provides a user interface that supports operations such as zooming in, zooming out, and panning. For example, zooming in and out can be achieved by scrolling the mouse wheel, and panning can be achieved by dragging the mouse.
[0066] The present invention significantly improves the efficiency and accuracy of voiceprint waveform visualization and reduces resource consumption through distributed computing and advanced visualization technology, and has high practicality and superiority.
[0067] In order to prove the advancement of this method, an experiment was designed to verify it. The specific method is as follows:
[0068] (1) Selecting speech data of the same scale, 240,000 sampling data are used in this embodiment, and the traditional method and the method of the present invention are used to perform voiceprint waveform visualization processing respectively;
[0069] (2) Record indicators such as processing time, resource usage (such as CPU usage, memory usage), and visualization effects (such as image clarity, response speed).
[0070] The experimental results are as follows:
[0071] (1) Processing time comparison:
[0072] Traditional method: It takes about 3000ms to process 240,000 pieces of data; The method of the present invention: The computing efficiency is improved. Through actual tests, the time for serial processing and parallel processing of voiceprint data of the same scale is compared.
[0073] Assuming that serial processing of 240,000 data requires T_serial = 2000ms, after parallel processing using a 4-core CPU, T_parallel = 500ms, the acceleration ratio S = 2000 / 500 = 4, and the processing speed is increased by about 4 times, proving that parallel computing effectively improves data processing efficiency. Specific data are shown in the table below:
[0074]
[0075] (2) Resource usage comparison:
[0076] Traditional method: CPU usage rate is as high as 90%, and memory occupancy is about 2GB; The method of the present invention: CPU usage rate is on average 60%, and memory occupancy is about 1.2GB, with more efficient resource utilization.
[0077] (3) Comparison of visualization effects:
[0078] Traditional method: Obvious image loading delay and slow response to interactive operations; The method of the present invention: Visualization effect evaluation, by observing the generated voiceprint waveform diagram, spectrogram and three-dimensional voiceprint diagram, to evaluate the visualization effect. The image loads quickly, the interactive operation is smooth, and the visualization effect is clearer.
[0079] The present invention has been described in detail through the embodiments, but the content is only the preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the application of the present invention should still fall within the scope covered by the patent of the present invention.
Claims
1. A distributed computing processing method for voiceprint waveform visualization, characterized in that: The method comprises the following steps: S1: Distributed computing processing: converting the voiceprint waveform data format, processing and storing it in an array, and processing the array to generate visual data of the voiceprint waveform; S2: Visual display: Use visualization tools to visualize the processed data and generate a voiceprint waveform.
2. The distributed computing processing method for voiceprint waveform visualization according to claim 1, characterized in that: S1 specifically includes the following steps: S101: voiceprint waveform feature extraction: using a function to extract the frequency spectrum features of the voiceprint waveform; S102: Data acquisition and preprocessing: Acquire the voiceprint waveform data extracted in S101 and convert the format into an array; S103: Array segmentation and parallel processing: the array preprocessed in S102 is divided into a plurality of sub-arrays, and each sub-array is processed by calculation to generate visual data of the voiceprint waveform.
3. The distributed computing processing method for voiceprint waveform visualization according to claim 1, characterized in that: In S2, the visual display also includes a user interaction function, providing a user interface and supporting zooming in, zooming out, and panning operations.
4. The distributed computing processing method for voiceprint waveform visualization according to claim 1, characterized in that: In S2, the visualization tool is ECharts, Matplotlib, D3.js or other visualization tools with similar functions. The generated visualization graph of the full data volume supports multiple visualization forms, including but not limited to waveform graph, spectrum graph and three-dimensional voiceprint graph, to achieve multi-angle display of voice signals.
5. The distributed computing processing method for voiceprint waveform visualization according to claim 2, characterized in that: In S101, the function uses short-time Fourier transform (STFT).
6. The distributed computing processing method for voiceprint waveform visualization according to claim 2, characterized in that: In S102, the voiceprint waveform data is obtained from the HTTP interface or Socket.io, the obtained json or xml data is stored in an array through format conversion, and the data integrity is verified.
7. The distributed computing processing method for voiceprint waveform visualization according to claim 2, characterized in that: In S102, when the length of the voiceprint waveform data exceeds the maximum limit, the array is intercepted and then the new data is stored in the array.
8. The distributed computing processing method for voiceprint waveform visualization according to claim 2, characterized in that: In S103, the method of calculating and processing each of the sub-arrays is CPU parallel calculation or multi-threading.
9. The distributed computing processing method for voiceprint waveform visualization according to claim 7, characterized in that: In S2, the visualization graphs generated by multiple sets of data are integrated through a layer style merging algorithm to merge the visualization graphs generated by multiple sets of data to form a visualization graph with the entire data volume.
10. A distributed computing processing method for voiceprint waveform visualization according to any one of claims 1 to 9, characterized in that: The voiceprint waveform data includes different languages and dialects, and the method can process and visualize the voice data of the different languages and dialects.