A speech training recognition method, device, equipment and medium
Patent Information
- Application Number
- CN202610985704.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-10-09
AI Technical Summary
[0002]随着门店规模扩大及培训频次提升,依赖人工抽检或简单规则校验来对培训进行监控,已难以满足高覆盖率、高一致性和实时反馈的需求
[0015]本发明的有益效果在于,本发明提供的上述语音培训识别方法,通过先获取培训人员佩戴录音设备采集的原始音频,再对音频进行人声识别提取与拼接形成目标音频流,能够有效过滤无效音频片段,大幅降低后续数据处理成本;将目标音频流转换为结构化文本数据,并结合文本数据与目标音频流开展多维度检测,可显著提升识别准确率与检测稳定性,同时实现音频质量、培训作业规范、违规话术的全面质检,保证检测覆盖广度与执行效率;综合各维度检测结果对异常音频标记并生成质检工单推送至业务端,能够及时发现并干预培训不规范行为,实现对培训执行效果的自动化识别与全程监控,整体方案场景适配性强、扩展灵活,可在各类培训场景中稳定落地应用。
Smart Images

Figure CN122889009A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a speech training and recognition method, apparatus, device, and medium. Background Technology
[0002] As stores expand and training frequency increases, relying on manual spot checks or simple rule verification to monitor training is no longer sufficient to meet the demands for high coverage, high consistency, and real-time feedback.
[0003] Among the relevant technical solutions, training and recognition schemes based on multimodal large models require unified modeling and inference of various data such as audio, video, and text, resulting in high overall computing power consumption and recognition costs. Furthermore, multimodal models require a large amount of labeled data for training or fine-tuning before deployment, leading to long training cycles and high costs. In complex environments like actual stores, recognition performance is easily affected by noise, scene differences, and other factors, making it difficult to guarantee long-term stable and controllable recognition accuracy. Summary of the Invention
[0004] The purpose of this invention is to provide a speech training recognition method, device, equipment, and medium that can effectively reduce processing costs, improve recognition accuracy and detection efficiency, realize automated monitoring of the training process and timely intervention for non-standard behaviors, and has strong scenario scalability.
[0005] To address the aforementioned technical problems, this invention provides a speech training and recognition method, comprising: Obtain the raw audio collected by the recording equipment worn by the trainees; The original audio is subjected to human voice segment recognition and extraction, and spliced together to obtain the target audio stream; Convert the target audio stream into structured text data; By combining the structured text data with the target audio stream, multi-dimensional detection is performed to obtain detection results for each dimension; Based on the comprehensive test results from various dimensions, unexpected audio is marked, a quality inspection work order is generated, and it is pushed to the business side.
[0006] In a first aspect, in the speech training and recognition method provided by the present invention, multi-dimensional detection is performed by combining the structured text data and the target audio stream to obtain detection results for each dimension, including: Using the structured text data as the basis for semantic analysis and the target audio stream as the basis for audio feature analysis, a multi-dimensional detection model is used to simultaneously detect audio quality, training assignment standards, and inappropriate language, and the detection results for each dimension are obtained.
[0007] On the other hand, in the above-mentioned speech training recognition method provided by the present invention, a multi-dimensional detection model is used to simultaneously detect from three dimensions: audio quality, training assignment specifications, and inappropriate speech, to obtain the detection results for each dimension, including: The structured text data and the target audio stream are input into the multi-dimensional detection model; Based on the signal characteristics of the target audio stream, the audio reception status and device wearing compliance are detected to obtain audio quality detection results; Based on the structured text data, the training duration, number of training sessions, and completeness of the course content are verified to obtain the training assignment standardization test results. Based on the structured text data, semantic analysis and risk content matching are performed to identify whether there are any illegal statements, and the results of illegal statement detection are obtained.
[0008] On the other hand, in the speech training and recognition method provided by the present invention, the process of recognizing and extracting human voice segments from the original audio and splicing them together to obtain the target audio stream includes: The original audio is input into the speech activity detection model; The speech activity detection model performs frame-by-frame analysis on the entire signal of the original audio to identify and locate speech segments containing valid human voices; Remove invalid segments from the original audio and extract the speech segments containing valid human voices; The extracted speech segments are spliced together sequentially in chronological order to form a continuous target audio stream containing only valid speech.
[0009] On the other hand, in the speech training and recognition method provided by the present invention, converting the target audio stream into structured text data includes: The target audio stream is input into a speech recognition model for segment-by-segment parsing, and the speech signal in the target audio stream is converted into text content; The text content is synchronously associated with the time information of the corresponding audio segments to form structured text data with time stamps and standardized format.
[0010] On the other hand, in the speech training and recognition method provided by this invention, by comprehensively analyzing the detection results from various dimensions, unexpected audio is marked, a quality inspection work order is generated, and pushed to the business end, including: The audio quality test results, the training operation specification test results, and the non-compliant script test results are summarized; Based on the veto or tiered warning rules, identify and mark unexpected audio that exceeds the set range; Based on the marked unexpected audio, a quality inspection work order is generated, which includes the unexpected type, evidence fragments, and relevant detection basis, and the quality inspection work order is pushed to the business end.
[0011] On the other hand, in the above-mentioned speech training recognition method provided by the present invention, obtaining the original audio collected by the recording device worn by the trainee includes: The recording equipment is pre-assigned to each trainee on a one-to-one basis; During the training process, the raw audio is obtained by near-field sound pickup after the trainees wear the recording equipment as required.
[0012] To address the aforementioned technical problems, the present invention also provides a voice training and recognition device, comprising: The raw audio acquisition module is used to acquire the raw audio collected by the recording equipment worn by the trainees; The audio stream splicing module is used to identify and extract human voice segments from the original audio and splice them to obtain the target audio stream; A text conversion module is used to convert the target audio stream into structured text data; A multi-dimensional detection module is used to combine the structured text data with the target audio stream to perform multi-dimensional detection and obtain detection results for each dimension. The summary and tagging module is used to integrate the detection results from various dimensions, tag unexpected audio, generate quality inspection work orders, and push them to the business side.
[0013] To address the aforementioned technical problems, the present invention also provides an electronic device, comprising: Memory, used to store computer programs; A processor is used to implement the steps of the above-described speech training and recognition method when executing the computer program.
[0014] To address the aforementioned technical problems, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the aforementioned speech training and recognition method.
[0015] The beneficial effects of this invention are as follows: The above-mentioned voice training recognition method provided by this invention, by first acquiring the original audio collected by the recording device worn by the trainee, and then performing voice recognition extraction and splicing on the audio to form a target audio stream, can effectively filter invalid audio segments and significantly reduce subsequent data processing costs; converting the target audio stream into structured text data and combining the text data with the target audio stream to carry out multi-dimensional detection can significantly improve the recognition accuracy and detection stability, while realizing comprehensive quality inspection of audio quality, training operation specifications, and non-compliant language, ensuring the breadth of detection coverage and execution efficiency; by comprehensively analyzing the detection results of various dimensions, abnormal audio is marked and a quality inspection work order is generated and pushed to the business end, which can promptly detect and intervene in non-standard training behaviors, realize automated identification and full-process monitoring of training execution effects, and the overall solution has strong scenario adaptability and flexible expansion, and can be stably applied in various training scenarios.
[0016] In addition, the present invention also provides a corresponding speech training and recognition device, electronic device and computer-readable storage medium for the speech training and recognition method, which have the same or corresponding technical features as the speech training and recognition method mentioned above, and have the same effect. Attached Figure Description
[0017] To more clearly illustrate the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A flowchart of the speech training and recognition method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the framework corresponding to the speech training and recognition method provided in the embodiments of the present invention; Figure 3 This is a schematic diagram of the multi-dimensional detection process provided in the embodiments of the present invention; Figure 4 This is a schematic diagram of the structure of the speech training and recognition device provided in an embodiment of the present invention. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.
[0020] It should be noted that, in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0021] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0022] The specific application environment architecture or specific hardware architecture on which the speech training and recognition method depends is described here.
[0023] The embodiments of the present invention provide a speech training and recognition method, and the method is described in detail in conjunction with the execution flow of the speech training and recognition method. Figure 1 A flowchart of the speech training and recognition method provided in the embodiments of the present invention is shown below. Figure 1 As shown, the method includes: S101. Obtain the raw audio collected by the recording equipment worn by the trainees.
[0024] It should be noted that "training personnel" refers to staff members who participate in training or perform training tasks in stores or other business scenarios, and are the subjects of this audio collection. The recording equipment is a specialized recording device used for close-range speech capture with high-fidelity sound pickup capabilities, and must be worn and attached to the training personnel. The raw audio refers to the complete audio signal directly captured by the recording equipment, without any filtering, trimming, or algorithmic processing.
[0025] Step S101 involves training personnel to wear recording devices to acquire raw audio, achieving near-field high-fidelity acquisition and providing high-quality, reliable basic data for subsequent speech processing, text conversion, and multi-dimensional quality inspection.
[0026] S102. Identify and extract human voice segments from the original audio and splice them together to obtain the target audio stream.
[0027] In practice, step S102 involves recognizing, extracting, and splicing human voice segments from the original audio. This effectively removes invalid data such as silence and background noise, significantly reducing the amount of data to be processed while retaining valid voice information. This reduces the computational power consumption and operating costs of subsequent voice conversion and intelligent detection, and improves overall processing efficiency.
[0028] S103. Convert the target audio stream into structured text data.
[0029] In implementation, step S103 transforms the preprocessed target audio stream into structured text data, realizing the standardized conversion of unstructured speech information into computable and searchable text data. This provides a unified and stable analytical foundation for subsequent multi-dimensional intelligent detection, effectively eliminating the subjectivity and variability caused by manual judgment, and improving the consistency, traceability, and quantifiability of the detection results.
[0030] S104. Combine structured text data with target audio stream to perform multi-dimensional detection and obtain detection results for each dimension.
[0031] In implementation, step S104 combines structured text data and target audio stream to conduct multi-dimensional detection. It uses text data to accurately determine semantic aspects such as training standards and inappropriate language, and relies on target audio stream to objectively verify signal characteristics such as audio quality and device wearing status. Through dual-dimensional data collaborative analysis, the comprehensiveness and accuracy of the detection results are greatly improved.
[0032] S105. Based on the comprehensive detection results from various dimensions, mark the unexpected audio, generate a quality inspection work order, and push it to the business end.
[0033] In implementation, step S105 summarizes the detection results from various dimensions to achieve anomaly judgment and processing. It can accurately mark unexpected audio and automatically generate quality inspection work orders to push to the business end. It transforms intelligent detection results into executable and traceable management actions, promptly discovers and intervenes in non-standard training behaviors, which not only improves the response speed and handling efficiency of training monitoring, but also realizes the full-process automated control from audio collection and analysis to anomaly intervention, providing reliable support for continuously optimizing the training execution effect.
[0034] In the above-mentioned voice training recognition method provided by the embodiments of the present invention, by first acquiring the original audio collected by the recording device worn by the trainee, and then performing voice recognition extraction and splicing on the audio to form a target audio stream, invalid audio segments can be effectively filtered out, significantly reducing the subsequent data processing cost. Converting the target audio stream into structured text data and combining the text data with the target audio stream to carry out multi-dimensional detection can significantly improve the recognition accuracy and detection stability. At the same time, it can realize comprehensive quality inspection of audio quality, training operation specifications, and non-compliant language, ensuring the breadth of detection coverage and execution efficiency. By comprehensively analyzing the detection results of various dimensions, abnormal audio is marked and a quality inspection work order is generated and pushed to the business end, which can promptly detect and intervene in non-standard training behaviors, realize automated identification and full-process monitoring of training execution effects. The overall solution has strong scenario adaptability and flexible expansion, and can be stably applied in various training scenarios.
[0035] Furthermore, in a specific implementation, in the above-mentioned voice training recognition method provided in the embodiments of the present invention, step S101 of acquiring the original audio collected by the recording device worn by the trainee may specifically include: pre-binding the recording device to the trainee one-to-one; during the training process, acquiring the original audio collected by the trainee through near-field sound pickup after the trainee wears the recording device in accordance with the specifications.
[0036] It should be noted that, taking freight store operations as an example, the training process typically involves several steps: training conducted by store trainers, quality inspectors manually checking in-store surveillance videos, relying on manual experience to check for problems in the training, and handling and penalizing violations. The "manual checking of in-store surveillance videos" step relies heavily on the audio-visual quality of far-field monitoring and the subjective experience of quality inspectors, which has drawbacks such as loud audio noise making it difficult to hear clearly, low efficiency of manual checks, highly subjective scoring standards, and difficulty in achieving large-scale, comprehensive coverage.
[0037] This invention upgrades the "manual sampling inspection of surveillance videos" to "wearable device data collection" and introduces artificial intelligence (AI) quality inspection capabilities. This ensures data quality and inspection efficiency from the source. Building upon the existing model that relies on fixed monitoring positions for far-field recording, it upgrades to having employees wear recording devices for near-field recording, significantly improving audio signal-to-noise ratio and clarity, thereby enhancing subsequent recognition accuracy. Simultaneously, it introduces automated AI quality inspection capabilities to replace traditional manual listening, using algorithms to perform 7... 24-hour full-coverage testing completely solves the problems of narrow coverage, low efficiency, and inconsistent subjective judgment standards in manual sampling.
[0038] During implementation and the initial business launch phase, this invention can bind the recording device one-to-one with specific positions in the store (such as front desk staff and trainers), requiring employees to wear it properly during work hours (such as when giving lectures in the training room). By using near-field sound pickup, it significantly reduces environmental background noise interference, obtaining clear audio data with a high signal-to-noise ratio from the source, providing a data foundation for subsequently improving the accuracy of AI recognition.
[0039] Furthermore, in a specific implementation, in the speech training and recognition method provided in the embodiments of the present invention, step S102 involves recognizing and extracting human voice segments from the original audio and splicing them together to obtain a target audio stream. Specifically, this may include: inputting the original audio into a speech activity detection model; having the speech activity detection model analyze the entire signal of the original audio frame by frame to identify and locate speech segments containing valid human voices; removing invalid segments from the original audio and extracting speech segments containing valid human voices; and splicing the extracted speech segments sequentially in chronological order to form a continuous target audio stream containing only valid speech.
[0040] In implementation, this invention can preprocess all collected store audio data and uniformly integrate it with the Voice Activity Detection (VAD) model. Specifically, it applies VAD's intelligent audio extraction method to achieve pre-filtering of invalid information and cost control. Store audio streams are uniformly integrated into the VAD model for intelligent preprocessing. The VAD model can detect the voice activity status in the audio stream in real time, automatically extract valid human voice segments, and splice them into a continuous, valid target audio stream, filtering out invalid segments such as long periods of silence and background noise. This process not only significantly reduces the time spent manually or automatically processing "blank audio" and achieves pre-filtering of invalid audio, but also significantly compresses the amount of data to be processed, allowing subsequent speech recognition and analysis to be performed only on audio segments containing valid voice content. Furthermore, it effectively reduces the computational consumption and calling costs of audio stream conversion, significantly improving the overall system operating efficiency.
[0041] Furthermore, in a specific implementation, in the speech training and recognition method provided in the embodiments of the present invention, step S103 converts the target audio stream into structured text data, which may specifically include: inputting the target audio stream into a speech recognition model for segment-by-segment parsing, converting the speech signals in the target audio stream into text content; and synchronously associating the text content with the time information of the corresponding speech segments to form structured text data with time stamps and format specifications.
[0042] In implementation, this invention can call the Automatic Speech Recognition (ASR) interface to convert unstructured speech data into structured text data, using valid audio segments extracted and spliced by VAD. The text output by ASR will contain timestamped dialogue content, serving as the core input data for the subsequent AI intelligent recognition module. In other words, this invention uses ASR technology to convert unstructured data into structured text, eliminating subjective human interference. High-precision ASR technology converts high-quality audio segments cleaned by VAD into text data, which is then used by an AI model for semantic analysis. This step transforms the subjective, non-standardized process of "human hearing and human judgment" into an objective, standardized computation process based on text data, ensuring high consistency and traceability in the analysis of training content, phrasing, and logic, and solving the problem of the difficulty in scaling up manual quality inspection.
[0043] Furthermore, in specific implementation, in the above-mentioned speech training recognition method provided in the embodiments of the present invention, step S104 combines structured text data and target audio stream to perform multi-dimensional detection and obtain detection results for each dimension. Specifically, it may include: using structured text data as the basis for semantic analysis and target audio stream as the basis for audio feature analysis, using a multi-dimensional detection model to simultaneously detect from three dimensions: audio quality, training operation specifications, and illegal speech, and obtaining detection results for each dimension.
[0044] Figure 2 This is a schematic diagram of the framework corresponding to the speech training and recognition method provided in an embodiment of the present invention. In implementation, as... Figure 2 As shown, this invention, based on the features of the escaped text and the original audio, initiates three-dimensional deep AI quality inspection in parallel. This constructs a three-dimensional AI quality inspection system integrating "audio quality + operational standards + prohibited language," achieving end-to-end risk control. A multi-dimensional detection model is established to perform a three-level deep scan of the audio: first, detecting "audio quality" (e.g., whether the equipment is worn, whether there is malicious microphone obstruction); second, identifying "operational standards" (e.g., whether the training duration is sufficient, whether key processes are omitted); and finally, monitoring "prohibited language" (e.g., whether there is insult or false promise). The judgment process mechanism employs a "one-vote veto" or "tiered early warning" mechanism. If any abnormality is identified in any of the above stages (e.g., substandard audio quality or detected prohibited words), the system will automatically mark it and push it to the business personnel for review and penalty. This mechanism, through audio quality detection, operational standard identification, and language compliance analysis, significantly improves business processing efficiency while achieving automated identification and continuous monitoring of training execution effectiveness and proactive interception of store service risks, effectively reducing user complaints and public opinion risks caused by inadequate training.
[0045] Furthermore, in specific implementation, in the above steps, a multi-dimensional detection model is used to simultaneously detect three dimensions: audio quality, training assignment specifications, and inappropriate language, obtaining detection results for each dimension. Specifically, this may include: inputting structured text data and the target audio stream into the multi-dimensional detection model; detecting the audio reception status and device wearing specifications based on the signal characteristics of the target audio stream to obtain audio quality detection results; verifying the training duration, number of training sessions, and completeness of the teaching content based on structured text data to obtain training assignment specifications detection results; and performing semantic analysis and risk content matching based on structured text data to identify the existence of inappropriate language, obtaining inappropriate language detection results.
[0046] Figure 3 This is a schematic diagram illustrating the multi-dimensional detection process provided in an embodiment of the present invention. In implementation, as... Figure 3As shown, structured text data and target audio stream are input into a multi-dimensional detection model. The model performs detection from three aspects: audio quality, training assignment standards, and inappropriate language. The specific content is as follows: In terms of audio quality detection, it mainly includes sound reception anomaly detection and wearing compliance detection. Sound reception anomaly detection uses signal processing technology to extract the spectral characteristics of the audio and sets empirical thresholds such as volume decibels and signal strength. When the audio signal indicators are consistently lower than the preset thresholds, it is determined that there is an abnormality such as equipment failure or obstruction of the sound reception. Wearing compliance detection comprehensively analyzes five dimensions of audio: signal-to-noise ratio, spectral distribution characteristics, dynamic range, low-frequency noise components, and spectral stability. The system calculates feature values for each dimension and compares them with empirical values. Then, it uses a weighted algorithm to fuse the scores and outputs the judgment result of whether the recording is worn correctly. If any of the above detections is abnormal, the audio quality is directly determined to be unqualified, and the corresponding abnormality type is marked for the business side to handle.
[0047] Regarding the inspection of training assignments, compliance checks are conducted sequentially for training session numbers, training duration, and training content completeness. Historical training data for a time period (e.g., the last X days) set by the trainer is automatically aggregated. The actual number of training sessions is statistically analyzed and compared with the standard number of sessions stipulated in the business regulations to determine if the training frequency is sufficient. Simultaneously, the training duration is checked against the metadata of the effective duration of the audio files to ensure it meets the standards. Then, based on the text content converted from automatic speech recognition, a Large Language Model (LLM) and prompt word engineering are used to semantically compare the actual teaching content with the standard training process to identify any omissions of key presentation elements and determine content completeness. If a trainer fails to meet the standards in any of the three aspects—number of sessions, duration, or content completeness—the assignment is deemed non-compliant, and a corresponding violation report is generated.
[0048] In terms of identifying prohibited language, a targeted prompt word project is built by combining store management regulations with the existing risk language database. Through a large language model, the converted text content is subjected to deep semantic analysis to accurately identify whether there are violations such as insulting customers, making false promises, or leaking privacy. If prohibited semantics are detected or risk keywords are hit, it is determined that there is a violation of the operation, and relevant text fragments are extracted as evidence.
[0049] Furthermore, in specific implementation, in the above-mentioned voice training recognition method provided in the embodiments of the present invention, step S105 integrates the detection results of various dimensions, marks the unexpected audio, generates a quality inspection work order and pushes it to the business end, which may specifically include: summarizing the audio quality detection results, training operation specification detection results, and violation script detection results; identifying and marking the unexpected audio that exceeds the set range according to the judgment rules of veto or graded warning; generating a quality inspection work order containing the unexpected type, evidence fragments and relevant detection basis based on the marked unexpected audio, and pushing the quality inspection work order to the business end.
[0050] In implementation, this invention can aggregate audio quality testing results, training assignment specification testing results, and non-compliant script testing results. Based on a veto or tiered warning system, it identifies and marks any abnormal behavior by trainers, such as substandard audio quality, non-standard assignments, or non-compliant assignments. Based on the marked unexpected audio, it automatically generates a quality inspection work order containing the unexpected type, evidence fragments, and relevant testing data, and pushes the work order to the business management system. Business personnel can quickly verify the evidence provided by the system, such as audio fragments, text records, and testing reports, and execute corresponding penalties according to relevant management regulations, thus achieving closed-loop management of the entire process from detection, identification, work order generation to execution.
[0051] It should be noted that this invention can construct a complete intelligent processing chain for store training management, achieving unified management and intelligent identification of the entire training process. It covers multiple key aspects such as intelligent audio extraction, audio quality detection, work standard identification, and judgment of inappropriate language, avoiding the problem of traditional solutions focusing only on inappropriate language identification and having incomplete coverage. This method can comprehensively identify and manage various types of non-standard work behaviors, such as insufficient training time, missing training sessions, and incomplete training content, improving the comprehensiveness and consistency of store training management. This invention can also introduce LLM as the core identification and understanding module to uniformly analyze and judge the language content and work execution during the training process, avoiding the high cost and low efficiency problems caused by repeatedly developing identification logic and rule models in different business scenarios. Based on the universal understanding capability of LLM, the need for manual rule maintenance and repeated model training can be reduced, improving the overall reusability and implementation efficiency of the system. In addition, this invention uses an intelligent audio extraction method based on voice activity detection to effectively segment long training audio, performing subsequent speech recognition and analysis only on segments containing valid speech content, significantly reducing the speech recognition computation cost while ensuring recognition effectiveness. This method avoids full transcription and analysis of invalid silent or noisy segments, improving overall processing efficiency and reducing system resource consumption.
[0052] This invention boasts excellent versatility and scalability, making it suitable not only for store training management scenarios but also for various other business scenarios such as front desk reception, after-sales service, and customer follow-up. Through a unified audio processing and violation recognition framework, only minor scene configurations or adjustments to prompts are required to achieve rapid reuse and capability expansion across different business scenarios.
[0053] From the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by software plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0054] Embodiments of the present invention also provide a voice training and recognition device. Figure 4 This is a schematic diagram of the speech training and recognition device provided in an embodiment of the present invention. This embodiment is based on functional modules, such as… Figure 4 As shown, the device includes: The raw audio acquisition module 10 is used to acquire the raw audio collected by the recording device worn by the trainees; The audio stream splicing module 11 is used to identify and extract human voice segments from the original audio and splice them to obtain the target audio stream; Text conversion module 12 is used to convert the target audio stream into structured text data; The multi-dimensional detection module 13 is used to combine structured text data and target audio stream to perform multi-dimensional detection and obtain detection results for each dimension; The summary and marking module 14 is used to integrate the detection results from various dimensions, mark the unexpected audio, generate quality inspection work orders, and push them to the business end.
[0055] In the speech training recognition device provided in this embodiment of the invention, the interaction of the above five modules can first acquire the original audio collected by the recording device worn by the trainee, and then perform human voice recognition extraction and splicing to form a target audio stream. This can effectively filter invalid audio segments and significantly reduce subsequent data processing costs. The target audio stream is converted into structured text data, and multi-dimensional detection is carried out in combination with the text data and the target audio stream, which can significantly improve the recognition accuracy and detection stability. At the same time, it can achieve comprehensive quality inspection of audio quality, training operation specifications, and non-compliant speech, ensuring the breadth of detection coverage and execution efficiency. By combining the detection results of various dimensions, abnormal audio is marked and a quality inspection work order is generated and pushed to the business end, which can promptly detect and intervene in non-standard training behaviors, realize automated identification and full-process monitoring of training execution effects. The overall solution has strong scenario adaptability and flexible expansion, and can be stably applied in various training scenarios.
[0056] Since the embodiments of the speech training and recognition device correspond to the embodiments of the speech training and recognition method, the descriptions of the features in the embodiments corresponding to the speech training and recognition device can be found in the relevant descriptions of the embodiments corresponding to the speech training and recognition method, and will not be repeated here. Furthermore, it has the same beneficial effects as the speech training and recognition method mentioned above.
[0057] Furthermore, in a specific implementation, in the above-mentioned voice training and recognition device provided in the embodiments of the present invention, the original audio acquisition module 10 can be used to pre-bind the recording device to the trainee one-to-one; during the training process, the original audio is acquired by the trainee wearing the recording device in accordance with the specifications and collected by near-field sound pickup.
[0058] Furthermore, in a specific implementation, in the speech training and recognition device provided in the embodiments of the present invention, the audio stream splicing module 11 can be used to input the original audio into the speech activity detection model; the speech activity detection model analyzes the entire signal of the original audio frame by frame, identifies and locates the speech segments containing valid human voices; invalid segments in the original audio are removed, and speech segments containing valid human voices are extracted; the extracted multiple speech segments are spliced sequentially in chronological order to form a continuous target audio stream containing only valid speech.
[0059] Furthermore, in a specific implementation, in the speech training and recognition device provided in the embodiments of the present invention, the text conversion module 12 can be used to input the target audio stream into the speech recognition model for segment-by-segment parsing, convert the speech signal in the target audio stream into text content; and synchronously associate the text content with the time information of the corresponding speech segment to form structured text data with time stamps and format specifications.
[0060] Furthermore, in specific implementation, in the above-mentioned voice training recognition device provided in this embodiment of the invention, the multi-dimensional detection module 13 can be used to simultaneously detect from three dimensions—audio quality, training assignment specifications, and inappropriate language—using structured text data as the basis for semantic analysis and the target audio stream as the basis for audio feature analysis, to obtain detection results for each dimension. Specifically, the structured text data and the target audio stream are input into the multi-dimensional detection model; based on the signal characteristics of the target audio stream, the audio reception status and device wearing specifications are detected to obtain audio quality detection results; based on the structured text data, the training duration, number of training sessions, and completeness of the teaching content are verified to obtain training assignment specifications detection results; based on the structured text data, semantic analysis and risk content matching are performed to identify whether inappropriate language exists, to obtain inappropriate language detection results.
[0061] Furthermore, in specific implementation, in the above-mentioned voice training recognition device provided in the embodiments of the present invention, the summary marking module 14 can be used to summarize the audio quality detection results, training operation specification detection results, and violation speech detection results; identify and mark unexpected audio that exceeds the set range according to the judgment rules of veto or graded warning; generate a quality inspection work order containing unexpected type, evidence fragments and relevant detection basis based on the marked unexpected audio, and push the quality inspection work order to the business end.
[0062] Embodiments of the present invention also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above-described embodiments of the speech training and recognition method.
[0063] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program configured to execute the steps in any of the above-described embodiments of the speech training and recognition method.
[0064] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0065] Embodiments of the present invention also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described embodiments of the speech training and recognition method.
[0066] Embodiments of the present invention also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described embodiments of the speech training and recognition method.
[0067] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0068] The above provides a detailed description of the speech training and recognition method, apparatus, device, and medium provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the embodiments above are only intended to help understand the method and core ideas of the present invention. It should be noted that those skilled in the art can make various improvements and modifications to the present invention without departing from its principles, and these improvements and modifications also fall within the protection scope of the present invention.
Claims
1. A speech training and recognition method, characterized in that, include: Obtain the raw audio collected by the recording equipment worn by the trainees; The original audio is subjected to human voice segment recognition and extraction, and spliced together to obtain the target audio stream; Convert the target audio stream into structured text data; By combining the structured text data with the target audio stream, multi-dimensional detection is performed to obtain detection results for each dimension; Based on the comprehensive test results from various dimensions, unexpected audio is marked, a quality inspection work order is generated, and it is pushed to the business side.
2. The speech training and recognition method according to claim 1, characterized in that, Multi-dimensional detection is performed by combining the structured text data with the target audio stream to obtain detection results for each dimension, including: Using the structured text data as the basis for semantic analysis and the target audio stream as the basis for audio feature analysis, a multi-dimensional detection model is used to simultaneously detect audio quality, training assignment standards, and inappropriate language, and the detection results for each dimension are obtained.
3. The speech training and recognition method according to claim 2, characterized in that, A multi-dimensional detection model was used to simultaneously detect violations from three dimensions: audio quality, training assignment standards, and inappropriate language. The results for each dimension were obtained, including: The structured text data and the target audio stream are input into the multi-dimensional detection model; Based on the signal characteristics of the target audio stream, the audio reception status and device wearing compliance are detected to obtain audio quality detection results; Based on the structured text data, the training duration, number of training sessions, and completeness of the course content are verified to obtain the training assignment standardization test results. Based on the structured text data, semantic analysis and risk content matching are performed to identify whether there are any illegal statements, and the results of illegal statement detection are obtained.
4. The speech training and recognition method according to claim 1, characterized in that, The process of identifying and extracting human voice segments from the original audio and then splicing them together to obtain the target audio stream includes: The original audio is input into the speech activity detection model; The speech activity detection model performs frame-by-frame analysis on the entire signal of the original audio to identify and locate speech segments containing valid human voices; Remove invalid segments from the original audio and extract the speech segments containing valid human voices; The extracted speech segments are spliced together sequentially in chronological order to form a continuous target audio stream containing only valid speech.
5. The speech training and recognition method according to claim 1, characterized in that, Converting the target audio stream into structured text data includes: The target audio stream is input into a speech recognition model for segment-by-segment parsing, and the speech signal in the target audio stream is converted into text content; The text content is synchronously associated with the time information of the corresponding audio segments to form structured text data with time stamps and standardized format.
6. The speech training and recognition method according to claim 3, characterized in that, Based on the comprehensive detection results from various dimensions, unexpected audio is marked, a quality inspection work order is generated, and pushed to the business side, including: The audio quality test results, the training operation specification test results, and the non-compliant script test results are summarized; Based on the veto or tiered warning rules, identify and mark unexpected audio that exceeds the set range; Based on the marked unexpected audio, a quality inspection work order is generated, which includes the unexpected type, evidence fragments, and relevant detection basis, and the quality inspection work order is pushed to the business end.
7. The speech training and recognition method according to claim 1, characterized in that, Obtain the raw audio captured by the recording equipment worn by the trainees, including: The recording equipment is pre-assigned to each trainee on a one-to-one basis; During the training process, the raw audio is obtained by near-field sound pickup after the trainees wear the recording equipment as required.
8. A voice training recognition device, characterized in that, include: The raw audio acquisition module is used to acquire the raw audio collected by the recording equipment worn by the trainees; The audio stream splicing module is used to identify and extract human voice segments from the original audio and splice them to obtain the target audio stream; A text conversion module is used to convert the target audio stream into structured text data; A multi-dimensional detection module is used to combine the structured text data with the target audio stream to perform multi-dimensional detection and obtain detection results for each dimension. The summary and tagging module is used to integrate the detection results from various dimensions, tag unexpected audio, generate quality inspection work orders, and push them to the business side.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the speech training and recognition method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the speech training and recognition method as described in any one of claims 1 to 7.