Abnormal speaking monitoring method and device, equipment, storage medium and program product

Through the recurrent neural network and anomaly speech detection model, automatic identification of the abnormal speech of the reviewer in the review activities is achieved, and the problems of high cost and high error detection rate under the human resources supervision model are solved, and the effectiveness and consistency of supervision are improved.

CN119943093APending Publication Date: 2025-05-06CHINA SOUTHERN POWER GRID MATERIALS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510019601.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, when the human resources supervision model is used in the evaluation activities to monitor the abnormal speech of the reviewer, there are problems such as high cost, high mis-detection and missed detection rates, which affects the effectiveness of supervision.

Method used

An abnormal speech monitoring method is adopted, by obtaining the audio data to be tested, converting it into a language text sequence using a recurrent neural network, and identifying it through an abnormal speech detection model to obtain the recognition results of abnormal speech.

Benefits of technology

This method can effectively reduce labor costs, reduce false detection and missed detection rates, improve the regulatory effectiveness of abnormal speeches, and ensure the consistency of supervision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943093A_ABST
    Figure CN119943093A_ABST
Patent Text Reader

Abstract

The invention relates to an abnormal speaking monitoring method and device, computer equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: acquiring to-be-tested audio data; converting the to-be-tested audio data into a language text sequence through a recurrent neural network; and performing abnormal speech recognition based on the language text sequence through an abnormal speech detection model to obtain an abnormal speech recognition result of the to-be-detected audio data, the abnormal speech recognition result comprising at least one of an abnormal discrimination result and a target abnormal type. By adopting the method, the abnormal speaking supervision effectiveness can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of review monitoring, and in particular to a method, apparatus, computer equipment, computer-readable storage medium and computer program product for monitoring abnormal speech. Background Art

[0002] In review activities such as bid evaluation, test marking, academic paper review, and competition review, reviewers are required to complete their work tasks autonomously and independently without interacting with each other, so as to minimize deviations caused by human factors. At the same time, it can also prevent any form of conflict of interest or improper interference, and ensure the professionalism and authority of the entire review process.

[0003] In traditional technology, a human supervision model is usually adopted, hiring and training professional supervisors to supervise the reviewers' abnormal remarks, such as illegal remarks and misleading remarks during the review process.

[0004] However, the manual supervision model is not only costly, but also subject to human supervision, which is bound to result in false detections and missed detections, thus affecting the effectiveness of supervision of abnormal speech. Summary of the invention

[0005] Based on this, it is necessary to provide an abnormal speech monitoring method, device, computer equipment, computer-readable storage medium and computer program product that can improve the effectiveness of abnormal speech supervision in response to the above-mentioned technical problems.

[0006] In a first aspect, the present application provides a method for monitoring abnormal speech, comprising:

[0007] Get the audio data to be tested;

[0008] The audio data to be tested is converted into a language text sequence through a recurrent neural network;

[0009] Through the abnormal speech detection model, abnormal speech recognition is performed based on the language text sequence to obtain the abnormal speech recognition result of the audio data to be tested, and the abnormal speech recognition result includes at least one of the abnormal discrimination result and the target abnormal type.

[0010] In one embodiment, obtaining the audio data to be tested includes:

[0011] Acquire an audio signal to be tested, where the audio signal to be tested is collected from a target scene by a sound collector;

[0012] The audio signal to be tested is processed by dividing into frames to obtain multiple frames of audio data to be tested.

[0013] In one embodiment, the audio signal to be tested is subjected to frame processing to obtain multiple frames of audio data to be tested, including:

[0014] Obtain the frame length value and frame shift value corresponding to the target scene, wherein the frame shift value is smaller than the frame length value;

[0015] According to the frame length value and the frame shift value, the audio signal to be tested is processed by frame division to obtain multiple frames of audio data to be tested.

[0016] In one embodiment, performing frame processing on the audio signal to be tested to obtain multiple frames of audio data to be tested includes:

[0017] Performing frame processing on the audio signal to be tested to obtain multiple frames;

[0018] Each sub-frame is processed by adding a Hanning window to obtain a plurality of audio data to be tested.

[0019] In one embodiment, the recurrent neural network includes an acoustic model and a language model, and the acoustic model and the language model are modeled by the recurrent neural network; the audio data to be tested is converted into a language text sequence by the recurrent neural network, including:

[0020] Extract features of the audio data to be tested and obtain a Mel frequency cepstral coefficient sequence;

[0021] The Mel frequency cepstral coefficient sequence is encoded through an acoustic model to obtain acoustic coding features;

[0022] The acoustic coding features are decoded through the language model to obtain a language text sequence.

[0023] In one embodiment, the abnormal speech detection model includes a natural language large model, which is obtained by supervised training using annotated text data in the business field; the abnormal speech detection model is used to perform abnormal speech recognition based on the language text sequence to obtain abnormal speech recognition results of the audio data to be tested, including:

[0024] When semantic retrieval of the language text sequence is performed through the natural language big model and at least one target anomaly type is obtained, determining the anomaly discrimination result as the presence of an anomaly, wherein the matching degree between the target anomaly type and the language text sequence is higher than a preset matching degree threshold;

[0025] When a semantic search is performed on a language text sequence through a large natural language model and the target anomaly type is not retrieved, the anomaly discrimination result is determined to be that no anomaly exists.

[0026] In a second aspect, the present application also provides an abnormal speech monitoring device, comprising:

[0027] An acquisition module, used to acquire audio data to be tested;

[0028] A conversion module, used for converting the audio data to be tested into a language text sequence through a recurrent neural network;

[0029] The recognition module is used to perform abnormal speech recognition based on the language text sequence through the abnormal speech detection model to obtain the abnormal speech recognition result of the audio data to be tested, and the abnormal speech recognition result includes at least one of the abnormal judgment result and the abnormal type.

[0030] In a third aspect, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0031] Get the audio data to be tested;

[0032] The audio data to be tested is converted into a language text sequence through a recurrent neural network;

[0033] Through the abnormal speech detection model, abnormal speech recognition is performed based on the language text sequence to obtain the abnormal speech recognition result of the audio data to be tested, and the abnormal speech recognition result includes at least one of the abnormal discrimination result and the target abnormal type.

[0034] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the following steps are implemented:

[0035] Get the audio data to be tested;

[0036] The audio data to be tested is converted into a language text sequence through a recurrent neural network;

[0037] Through the abnormal speech detection model, abnormal speech recognition is performed based on the language text sequence to obtain the abnormal speech recognition result of the audio data to be tested, and the abnormal speech recognition result includes at least one of the abnormal discrimination result and the target abnormal type.

[0038] In a fifth aspect, the present application further provides a computer program product, including a computer program, which implements the following steps when executed by a processor:

[0039] Get the audio data to be tested;

[0040] The audio data to be tested is converted into a language text sequence through a recurrent neural network;

[0041] Through the abnormal speech detection model, abnormal speech recognition is performed based on the language text sequence to obtain the abnormal speech recognition result of the audio data to be tested, and the abnormal speech recognition result includes at least one of the abnormal discrimination result and the target abnormal type.

[0042] The above abnormal speech monitoring method, device, computer equipment, computer readable storage medium and computer program product, by acquiring the audio data to be tested, converting the audio data to be tested into a language text sequence through a recurrent neural network, realizes the speech recognition of the audio data to be tested, converts the audio data to be tested into a language text sequence that can be recognized by the abnormal speech detection model, and then through the abnormal speech detection model, performs abnormal speech recognition based on the language text sequence to obtain the abnormal speech recognition result of the audio data to be tested, and the abnormal speech recognition result includes at least one of the abnormal discrimination result and the target abnormal type, thereby realizing the abnormal speech recognition of the audio data to be tested. Compared with the human supervision mode, on the one hand, after the model is trained, it can be repeatedly used in each review activity in the same field, thereby effectively reducing the labor cost. On the other hand, using the model to accurately identify the abnormal speech in the audio data to be tested can effectively reduce the probability of false detection, and the model can efficiently identify each sentence in the review process, so it can effectively reduce the probability of missed detection, and the model has a unified evaluation standard for each sentence, so it can ensure the consistency of supervision. In summary, the effectiveness of supervision of abnormal speech can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0044] Figure 1 A flowchart of an abnormal speech monitoring method in one embodiment;

[0045] Figure 2 It is a flowchart of the steps of obtaining the audio data to be tested in one embodiment;

[0046] Figure 3 A flowchart of an abnormal speech monitoring method in another embodiment;

[0047] Figure 4 is a structural block diagram of an abnormal speech monitoring device in one embodiment;

[0048] Figure 5 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0050] In an exemplary embodiment, Figure 1 As shown, a method for monitoring abnormal speech is provided. This embodiment takes the method applied to a terminal as an example, wherein the terminal may be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The Internet of Things devices may be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, projection devices, etc. Portable wearable devices may be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted devices may be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. It can be understood that the method may also be applied to a server, and may also be applied to a system including a terminal and a server, and may be implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps S10-S30. Among them:

[0051] Step S10, obtaining audio data to be tested.

[0052] The audio data to be tested may refer to data containing voice information of the assessors in the assessment activity.

[0053] Exemplarily, during the review activity, the sound stream of the reviewer can be captured in real time through a sound collector; then, the real-time captured sound stream can be used as the audio data to be tested, or the real-time captured sound stream can be pre-processed in real time and the pre-processed sound stream can be used as the audio data to be tested, or the voice information of the reviewer can be recorded through a sound collector and stored as a sound file, and the sound file can be used as the audio data to be tested, or the recorded sound file can be pre-processed and the pre-processed sound file can be used as the audio data to be tested. The specific method can be determined according to the actual situation, and this embodiment does not impose any limitation on this.

[0054] Among them, the preprocessing may include at least one of noise reduction, speech enhancement, speech separation, adaptive filtering, etc.

[0055] In some feasible implementations, the sound stream captured in real time can be filtered by a filter to filter out the sound streams in other frequency bands except the preset human voice frequency band, so as to obtain the target sound stream in the preset human voice frequency band; then the energy value of the target sound stream is monitored, and when it is detected that the energy value at a certain moment in the target sound stream is higher than the preset energy value threshold, the target sound stream is sent to the recurrent neural network for conversion, and the energy value of the target sound stream is continuously monitored; when it is detected that the energy value of the target sound stream is less than or equal to the preset energy value threshold, the timing is started, and if the energy value of the target sound stream is still not detected to be higher than the preset energy value threshold when the timing reaches the preset timing threshold, the sending of the target sound stream to the recurrent neural network for conversion is suspended; if the energy value of the target sound stream is detected to be higher than the preset energy value threshold before the timing reaches the preset timing threshold, the step of continuously monitoring the energy value of the target sound stream is returned. The model algorithm has high requirements for computing power, so the recurrent neural network and abnormal speech detection model are usually deployed on the server. In the real-time monitoring scenario, the audio data to be tested collected by the terminal needs to be sent to the server for the recurrent neural network and abnormal speech detection model deployed on the server to monitor abnormal speech. If all the sounds collected during the review process are uploaded and monitored, a large amount of communication resources and computing power resources will be required. First, non-human voices are filtered out through a filter, and only human voices are retained. Then, the energy value is monitored to determine whether the sound stream contains valid information. Only data containing valid information is transmitted and monitored. This can effectively save communication resources and computing power resources while ensuring the accuracy of abnormal speech monitoring.

[0056] Step S20: converting the audio data to be tested into a language text sequence through a recurrent neural network.

[0057] Among them, the recurrent neural network is a special type of neural network, which has unique advantages in processing sequence data. The recurrent neural network has a memory characteristic and can retain the information of the previous time step during the calculation process and use it in subsequent calculations. The recurrent neural network can be deployed on a terminal or a server, which can be determined according to actual conditions, and this embodiment does not limit this.

[0058] As an example, when a recurrent neural network is deployed on a terminal, the audio data to be tested can be input into a pre-trained recurrent neural network deployed on the terminal to perform speech recognition and convert the audio data to be tested into a corresponding language text sequence.

[0059] As another example, when the recurrent neural network is deployed on a server, the terminal can send the audio data to be tested to the server, so that the server can input the received audio data to be tested into a pre-trained recurrent neural network for speech recognition, convert the audio data to be tested into a corresponding language text sequence, and return the converted language text sequence to the terminal. The terminal receives the language text sequence returned by the server based on the audio data to be tested.

[0060] Step S30, using an abnormal speech detection model, performing abnormal speech recognition based on a language text sequence to obtain an abnormal speech recognition result of the audio data to be tested, wherein the abnormal speech recognition result includes at least one of an abnormal discrimination result and a target abnormal type.

[0061] Among them, the abnormal speech detection model may refer to an artificial intelligence model for abnormal speech detection, which may include at least one of a neural network model, a large-scale language model, a clustering model, etc. The abnormal speech detection model may be pre-trained using training samples, and the training samples may be annotated data after annotating the historical audio data in this field. For example, the historical audio data in this field may be annotated to determine whether it is abnormal and the type of abnormality, and the annotated data may be determined as the training sample. The abnormal speech detection model may be deployed on a terminal or a server, and may be specifically determined according to actual conditions, and this embodiment does not limit this.

[0062] As an example, when an abnormal speech detection model is deployed on a terminal, a language text sequence can be input into a pre-trained abnormal speech detection model deployed on the terminal, and the abnormal speech detection model is used to detect whether there are abnormal speeches in the language text sequence to obtain an abnormal discrimination result. In some feasible implementations, the abnormal discrimination result may include the presence of abnormal speeches or the absence of abnormal speeches; when an abnormal discrimination result of the presence of abnormal speeches is obtained, the abnormal speech detection model can be used to further detect the target abnormal type corresponding to the abnormal speech in the language text sequence. Therefore, when abnormal speeches are detected in the language text sequence, the abnormal speech recognition result output by the abnormal speech detection model may include the abnormal discrimination result of the presence of abnormal speeches and the target abnormal type; when abnormal speeches are detected in the language text sequence without abnormal speeches, the abnormal speech recognition result output by the abnormal speech detection model may include the abnormal discrimination result of the absence of abnormal speeches.

[0063] As another example, when the abnormal speech detection model is deployed on the server, the terminal can send the language text sequence to the server so that the server can input the received language text sequence into the pre-trained abnormal speech detection model, and detect whether there is an abnormal speech in the language text sequence through the abnormal speech detection model to obtain an abnormal discrimination result. In some feasible implementations, the abnormal discrimination result may include the presence of abnormal speech or the absence of abnormal speech; in the case of obtaining an abnormal discrimination result of the presence of abnormal speech, the target abnormal type corresponding to the abnormal speech in the language text sequence can be further detected by the abnormal speech detection model. Therefore, in the case of detecting that there is an abnormal speech in the language text sequence, the abnormal speech recognition result output by the abnormal speech detection model may include the abnormal discrimination result of the presence of abnormal speech and the target abnormal type; in the case of detecting that there is no abnormal speech in the language text sequence, the abnormal speech recognition result output by the abnormal speech detection model may include the abnormal discrimination result of the absence of abnormal speech. After obtaining the abnormal speech recognition result, the server returns the abnormal speech recognition result to the terminal, and the terminal receives the abnormal speech recognition result returned by the server based on the language text sequence.

[0064] In the above abnormal speech monitoring method, by obtaining the audio data to be tested, the audio data to be tested is converted into a language text sequence through a recurrent neural network, so as to realize the speech recognition of the audio data to be tested, and the audio data to be tested is converted into a language text sequence that can be recognized by the abnormal speech detection model, and then the abnormal speech detection model is used to perform abnormal speech recognition based on the language text sequence to obtain the abnormal speech recognition result of the audio data to be tested, and the abnormal speech recognition result includes at least one of the abnormal discrimination result and the target abnormal type, so as to realize the abnormal speech recognition of the audio data to be tested. Compared with the human supervision mode, on the one hand, after the model is trained, it can be repeatedly used in each review activity in the same field, so as to effectively reduce the labor cost, and on the other hand, the model can be used to accurately identify the abnormal speech in the audio data to be tested, which can effectively reduce the probability of false detection, and the model can efficiently identify each sentence in the review process, so as to effectively reduce the probability of missed detection, and the model has a unified evaluation standard for each sentence, so as to ensure the consistency of supervision, and in summary, the effectiveness of supervision of abnormal speech can be improved.

[0065] In an exemplary embodiment, Figure 2 As shown, obtaining the audio data to be tested includes steps S11 to S12. Among them:

[0066] Step S11, obtaining an audio signal to be tested, where the audio signal to be tested is collected from a target scene by a sound collector.

[0067] The audio signal to be tested may refer to an audio signal collected from the target scene by a sound collector, which contains the voice information of the assessor in the assessment activity. The target scene can be judged and distinguished based on the actual situation. For example, it can be divided into bid evaluation scene, test paper correction scene, academic paper review scene, competition review scene, etc. It can also be divided into academic scene, engineering scene, art scene, etc. according to the business scene. In some feasible implementations, it can be divided according to whether the abnormal speech monitoring standards are the same or similar. The abnormal speech monitoring standards corresponding to the same scene can be the same or similar.

[0068] Exemplarily, a sound collector may be deployed in the target scene, and voice information generated in the target scene may be collected in real time by the sound collector.

[0069] Step S12, performing frame processing on the audio signal to be tested to obtain multiple frames of audio data to be tested.

[0070] Exemplarily, after the audio signal to be tested is acquired, the audio signal to be tested may be further subjected to frame processing, and the audio signal to be tested may be divided into a plurality of sub-signals with a preset frame length, and the sub-signals are used as the audio data to be tested.

[0071] In this embodiment, in the scenario of real-time transmission and real-time monitoring, the amount of data transmitted and monitored each time can be reduced by framing, thereby improving the system response speed, allowing users to obtain feedback more promptly, and reducing the burden of a single transmission and monitoring, thereby maintaining system stability.

[0072] In an exemplary embodiment, the audio signal to be tested is subjected to frame processing to obtain multiple frames of audio data to be tested, including steps S1211 to S1212. In which:

[0073] Step S1211, obtaining a frame length value and a frame shift value corresponding to the target scene, wherein the frame shift value is smaller than the frame length value.

[0074] The frame length value may refer to the length of time covered by each frame. The frame shift value may refer to the offset of the starting position of the next frame relative to the starting position of the previous frame after processing a frame. If the frame shift value is smaller than the frame length value, partial overlap between frames can be achieved to improve the integrity of the semantic information contained in each frame, avoid the situation of taking the information out of context in subsequent speech recognition and abnormal speech recognition, reduce the false detection rate of abnormal speech, and avoid the situation of missing abnormal speech due to segmentation into different frames.

[0075] In some feasible implementations, the frame length value and the frame shift value can be determined in advance based on the duration of abnormal speeches in historical audio data collected in different scenarios. As an example, the longer the average duration of the abnormal speeches, the longer the frame length value can be, and the shorter the frame shift value can be; the frame length value can be greater than the maximum value of the duration of the abnormal speeches to ensure that no missed detection occurs.

[0076] In some feasible implementations, the overlap length between frames may be 30%-50% of the frame length. For example, the frame length may be 30 seconds, the frame shift value may be 20 seconds, and the overlap length between frames may be 10 seconds.

[0077] Exemplarily, after the target scene is determined, the frame length value and frame shift value corresponding to the target scene can be queried from the correspondence between the preset scene and the frame length value and the frame shift value.

[0078] Step S1212: performing frame processing on the audio signal to be tested according to the frame length value and the frame shift value to obtain multiple frames of audio data to be tested.

[0079] Exemplarily, the initial moment of the audio signal to be tested can be used as the initial moment of the first frame, and starting from the initial moment of the first frame, the first frame sub-signal with a frame length equal to the frame length value is segmented out, and then the initial moment of the first frame is added with the frame shift value to determine the initial moment of the next frame, and starting from the initial moment of the next frame, the next frame sub-signal with a frame length equal to the frame length value is segmented out, and then the initial moment of the current frame is added with the frame shift value to re-determine the initial moment of the next frame, and this cycle is repeated until all the audio signals to be tested are segmented, and these sub-signals are used as the audio data to be tested.

[0080] In this embodiment, by partially overlapping frames, the integrity of the semantic information contained in each frame can be improved, and the subsequent speech recognition and abnormal speech recognition can be avoided from being taken out of context, the false detection rate of abnormal speeches can be reduced, and the abnormal speeches can be avoided from being missed due to being divided into different frames.

[0081] In an exemplary embodiment, the audio signal to be tested is subjected to frame processing to obtain multiple frames of audio data to be tested, including steps S1221 to S1222.

[0082] Step S1221, performing frame processing on the audio signal to be tested to obtain multiple frames.

[0083] Exemplarily, after the audio signal to be tested is acquired, the audio signal to be tested may be further subjected to frame processing to obtain a plurality of frames.

[0084] In some feasible implementations, performing frame processing on the audio signal to be tested to obtain multiple frames may include: obtaining a frame length value and a frame shift value corresponding to the target scene, wherein the frame shift value is less than the frame length value; and performing frame processing on the audio signal to be tested according to the frame length value and the frame shift value to obtain multiple frames.

[0085] Step S1222: Perform Hanning window processing on each sub-frame to obtain a plurality of audio data to be tested.

[0086] The Hanning window is a weighted function that reduces spectrum leakage by smoothing the edges of the signal and reducing discontinuities at the frame boundaries.

[0087] Exemplarily, a window coefficient array with the same length as the frame length can be generated for each subframe according to the Hanning window formula, and then the window coefficient array is multiplied element by element by the audio sample value of each subframe to obtain the audio data to be tested corresponding to each subframe.

[0088] In this embodiment, the Hanning window processing can smooth the edge of the signal, reduce the discontinuity at the frame boundary, and thus reduce spectrum leakage.

[0089] In an exemplary embodiment, the ring neural network includes an acoustic model and a language model, and the acoustic model and the language model are modeled by a recurrent neural network; the audio data to be tested is converted into a language text sequence by the recurrent neural network, including steps S21 to S23. Among them:

[0090] Step S21, extracting features from the audio data to be tested to obtain a Mel-frequency cepstral coefficient sequence.

[0091] Exemplarily, after obtaining the audio data to be tested, the audio data to be tested can be first subjected to a fast Fourier transform, and the audio data to be tested can be converted from the time domain to the frequency domain to obtain a spectrum, and then a Mel filter bank is used to weight and integrate the energy in different frequency ranges to simulate the nonlinear perception of the human auditory system to frequency; and then a discrete cosine transform is performed on the output of the Mel filter bank to obtain a Mel-frequency cepstral coefficient sequence. In some feasible implementations, only the first preset number of Mel-frequency cepstral coefficient sequences can be retained, and the preset number can be 12, 13, etc., because these Mel-frequency cepstral coefficient sequences contain most of the information.

[0092] Step S22: Encode the Mel-frequency cepstral coefficient sequence through an acoustic model to obtain acoustic coding features.

[0093] Among them, the acoustic model can refer to a model built using a recurrent neural network or its variants that can process sequence data. The acoustic model can learn the patterns of audio features changing over time and map these patterns to the probability distribution of corresponding phonemes, characters, or other language units.

[0094] Exemplarily, after obtaining the Mel-frequency cepstral coefficient sequence, the acoustic model can be used as an encoder. The acoustic model encodes the Mel-frequency cepstral coefficient sequence by learning the time dependency and contextual relationship of the Mel-frequency cepstral coefficient sequence, and outputs the probability that the audio data to be tested corresponds to each possible phoneme or character, thereby obtaining acoustic coding features. The acoustic coding features can be represented by high-level audio features.

[0095] In some feasible implementations, before encoding, each Mel-frequency cepstral coefficient in the Mel-frequency cepstral coefficient sequence may be de-meaned and normalized, and then the processed Mel-frequency cepstral coefficient sequence may be encoded through an acoustic model to obtain acoustic coding features.

[0096] Step S23, decoding the acoustic coding features through a language model to obtain a language text sequence.

[0097] Among them, the language model can refer to a deep learning model based on a recurrent neural network architecture, which is used to predict the probability distribution of the next word in a text sequence by capturing the temporal dependency and contextual information in the text.

[0098] Exemplarily, the acoustic coding features output by the acoustic model can be input into the language model for word-by-word decoding. The language model is combined with context information to select the most likely word as the output, and all the output words are combined in order into a final language text sequence.

[0099] In an exemplary embodiment, the abnormal speech detection model includes a natural language large model, which is obtained by supervised training using annotated text data in the business field; the abnormal speech detection model is used to perform abnormal speech recognition based on the language text sequence to obtain the abnormal speech recognition result of the audio data to be tested, including steps S31 to S32. Among them:

[0100] Step S31, when semantic retrieval is performed on the language text sequence through the natural language big model and at least one target abnormality type is obtained, the abnormality judgment result is determined to be the existence of an abnormality, wherein the matching degree between the target abnormality type and the language text sequence is higher than a preset matching degree threshold.

[0101] Among them, large natural language models can refer to deep learning models with large parameters, rich training data and complex architecture. They are designed to understand and generate natural language text. Large natural language models are obtained by supervised training using annotated text data from the business domain. They can also be pre-trained using non-business domain data and then fine-tuned using annotated text data from the business domain.

[0102] Exemplarily, a language text sequence can be input into a natural language large model, and through the natural language large model, a semantic comparison is performed between the language text sequence and the abnormal speech vocabulary corresponding to each abnormal type, and the semantic similarity is determined as the matching degree between the language text sequence and each abnormal type, and the abnormal type with a matching degree higher than a preset matching degree threshold is determined as the target abnormal type. Furthermore, if at least one target abnormal type is retrieved, the abnormality discrimination result can be determined as the presence of an abnormality, and the abnormality discrimination result of the presence of an abnormality and the target abnormal type are output together as the abnormal speech recognition result.

[0103] In some feasible implementations, when there is no abnormality in the abnormal identification result, the abnormal speech recognition result can be displayed through the terminal's display or sound output device. For example, the abnormal identification result and the target abnormal type can be displayed in the form of a message on the display interface of the terminal, to assist staff in discovering illegal speech and stopping it in time, maintain a fair and just working environment, and be responsible for work results.

[0104] Step S32, when semantic retrieval is performed on the language text sequence through the natural language large model and the target abnormality type is not retrieved, the abnormality discrimination result is determined to be that an abnormality exists.

[0105] Exemplarily, if the target abnormality type with a matching degree higher than a preset matching degree threshold is not retrieved through the natural language large model, the abnormality judgment result can be determined as no abnormality exists, and the abnormality judgment result of no abnormality exists can be output as an abnormal speech recognition result.

[0106] In this embodiment, the accuracy of abnormal speech recognition can be improved by using a large natural language model.

[0107] In some exemplary embodiments, Figure 3As shown, first, the voice information at the review site is collected through a microphone to obtain the original audio data, and then a lossless audio encoder can be used to convert the original audio data into a digital format. Then, FFmpeg (an open source multimedia framework) can be used to encode the PCM (Pulse Code Modulation) audio into the AAC (Advanced Audio Coding) format and encapsulate it into the M4A (an audio file format) format for network transmission. Then, the continuous audio signal is divided into short-time frames with overlapping frames, where each frame is 30 seconds long, containing 30 seconds * 16000 samples / second = 480,000 samples are collected, with an overlap of 10 seconds between frames, so as to ensure that the 30-second short time frame contains complete semantics to the greatest extent possible, and avoid false alarms of illegal speeches caused by the subsequent large language model due to out-of-context content; then, a window function is added to each frame to smooth and weight the signal in the time domain, reduce discontinuities and spectrum leakage at the frame boundaries; then, the time domain signal is converted into a frequency domain signal, the frequency information is extracted, and then the spectral features are extracted through the Mel filter bank and discrete cosine transform to obtain the Mel frequency cepstral coefficients, and then the extracted Mel frequency cepstral coefficients are de-meaned and normalized; then, the processed The processed Mel-frequency cepstral coefficients are input into the recurrent neural network architecture, and the input Mel-frequency cepstral coefficients are encoded through the acoustic model, and the probability of the audio data to be tested corresponding to each possible phoneme or character is output. The output of the acoustic model is then input into the language model to generate a language text sequence; further, the language text sequence is input into the business field vertical big model for semantic retrieval, and illegal speech is retrieved through semantics. The recognition results and reasons for the violation are displayed in the form of a message on the terminal display interface, which assists staff in discovering illegal speech and stopping the behavior in a timely manner, maintaining a fair and just working environment, and being responsible for work results.

[0108] Among them, the business domain vertical big model can be obtained by pre-supervised training with business domain annotated data. Specifically, the business domain related annotated data can be collected first, for example, the content of illegal speech in the bid evaluation scenario and the categories involved; the illegal speech actually appearing in the invigilation scenario and the categories involved; the illegal speech appearing in the paper marking scenario and the corresponding categories, etc. The pre-trained natural language big model is supervised by annotated data to perform generalization capabilities corresponding to the business domain and scenario, and the business domain vertical big model is obtained.

[0109] It should be understood that, although the steps in the flowcharts involved in the above embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0110] Based on the same inventive concept, the embodiment of the present application also provides an abnormal speech monitoring device for implementing the abnormal speech monitoring method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more abnormal speech monitoring device embodiments provided below can refer to the limitations of the abnormal speech monitoring method above, and will not be repeated here.

[0111] In an exemplary embodiment, Figure 4 As shown, an abnormal speech monitoring device is provided, including: an acquisition module 402, a conversion module 404 and an identification module 406, wherein:

[0112] An acquisition module 402 is used to acquire audio data to be tested;

[0113] The conversion module 404 is used to convert the audio data to be tested into a language text sequence through a recurrent neural network;

[0114] The recognition module 406 is used to perform abnormal speech recognition based on the language text sequence through the abnormal speech detection model to obtain the abnormal speech recognition result of the audio data to be tested, and the abnormal speech recognition result includes at least one of the abnormal judgment result and the abnormal type.

[0115] In an exemplary embodiment, the acquisition module 402 is further configured to:

[0116] Acquire an audio signal to be tested, where the audio signal to be tested is collected from a target scene by a sound collector;

[0117] The audio signal to be tested is processed by dividing into frames to obtain multiple frames of audio data to be tested.

[0118] In an exemplary embodiment, the acquisition module 402 is further configured to:

[0119] Obtain the frame length value and frame shift value corresponding to the target scene, wherein the frame shift value is smaller than the frame length value;

[0120] According to the frame length value and the frame shift value, the audio signal to be tested is processed by frame division to obtain multiple frames of audio data to be tested.

[0121] In an exemplary embodiment, the acquisition module 402 is further configured to:

[0122] Performing frame processing on the audio signal to be tested to obtain multiple frames;

[0123] Each sub-frame is processed by adding a Hanning window to obtain a plurality of audio data to be tested.

[0124] In an exemplary embodiment, the recurrent neural network includes an acoustic model and a language model, and the acoustic model and the language model are obtained by adopting recurrent neural network modeling; the conversion module 404 is also used for:

[0125] Extract features of the audio data to be tested and obtain a Mel frequency cepstral coefficient sequence;

[0126] The Mel frequency cepstral coefficient sequence is encoded through an acoustic model to obtain acoustic coding features;

[0127] The acoustic coding features are decoded through the language model to obtain a language text sequence.

[0128] In an exemplary embodiment, the abnormal speech detection model includes a natural language large model, which is obtained by supervised training using annotated text data in the business field; the recognition module 406 is also used to:

[0129] When semantic retrieval of the language text sequence is performed through the natural language big model and at least one target anomaly type is obtained, determining the anomaly discrimination result as the presence of an anomaly, wherein the matching degree between the target anomaly type and the language text sequence is higher than a preset matching degree threshold;

[0130] When a semantic search is performed on a language text sequence through a large natural language model and the target anomaly type is not retrieved, the anomaly discrimination result is determined to be that no anomaly exists.

[0131] Each module in the abnormal speech monitoring device can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.

[0132] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 5As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC) or other technologies. When the computer program is executed by the processor, an abnormal speech monitoring method is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse.

[0133] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0134] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.

[0135] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0136] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0137] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0138] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.

[0139] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0140] The above embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A method for monitoring abnormal speech, characterized in that: The method comprises: Get the audio data to be tested; The audio data to be tested is converted into a language text sequence through a recurrent neural network; By using an abnormal speech detection model, abnormal speech recognition is performed based on the language text sequence to obtain an abnormal speech recognition result of the audio data to be tested, wherein the abnormal speech recognition result includes at least one of an abnormal discrimination result and a target abnormal type.

2. The method according to claim 1, characterized in that The obtaining of the audio data to be tested comprises: Acquire an audio signal to be tested, where the audio signal to be tested is collected from a target scene by a sound collector; The audio signal to be tested is processed by frame division to obtain multiple frames of audio data to be tested.

3. The method according to claim 2, characterized in that The step of performing frame processing on the audio signal to be tested to obtain multiple frames of audio data to be tested includes: Acquire a frame length value and a frame shift value corresponding to the target scene, wherein the frame shift value is smaller than the frame length value; According to the frame length value and the frame shift value, the audio signal to be tested is processed by frame division to obtain multiple frames of audio data to be tested.

4. The method according to claim 2, characterized in that: The performing frame processing on the audio signal to be tested to obtain multiple frames of audio data to be tested comprises: Performing frame processing on the audio signal to be tested to obtain multiple frames; Each sub-frame is processed by adding a Hanning window to obtain a plurality of audio data to be tested.

5. The method according to any one of claims 1 to 4, characterized in that: The recurrent neural network includes an acoustic model and a language model, and the acoustic model and the language model are obtained by adopting recurrent neural network modeling; The step of converting the audio data to be tested into a language text sequence by using a recurrent neural network includes: Extracting features of the audio data to be tested to obtain a Mel-frequency cepstral coefficient sequence; Encoding the Mel-frequency cepstral coefficient sequence through an acoustic model to obtain an acoustic coding feature; The acoustic coding features are decoded through a language model to obtain a language text sequence.

6. The method according to any one of claims 1 to 4, characterized in that: The abnormal speech detection model includes a large natural language model, which is obtained by supervised training using annotated text data in the business field; The abnormal speech detection model is used to perform abnormal speech recognition based on the language text sequence to obtain the abnormal speech recognition result of the audio data to be tested, including: When semantic retrieval is performed on the language text sequence through a natural language big model and at least one target abnormality type is obtained, determining that the abnormality discrimination result is that an abnormality exists, wherein the matching degree between the target abnormality type and the language text sequence is higher than a preset matching degree threshold; When semantic retrieval is performed on the language text sequence through a large natural language model and no target abnormality type is retrieved, the abnormality determination result is determined to be that no abnormality exists.

7. An abnormal speech monitoring device, characterized in that: The device comprises: An acquisition module, used to acquire audio data to be tested; A conversion module, used for converting the audio data to be tested into a language text sequence through a recurrent neural network; The recognition module is used to perform abnormal speech recognition based on the language text sequence through an abnormal speech detection model to obtain an abnormal speech recognition result of the audio data to be tested, wherein the abnormal speech recognition result includes at least one of an abnormal discrimination result and an abnormal type.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.