Audio detection method and apparatus, electronic device, and storage medium
By monitoring and processing emotional abnormalities during voice communication in real time, using the first and second detection modes, the lack of emotion monitoring in voice communication is solved, disputes are avoided, and detection efficiency is improved.
Patent Information
- Application Number
- CN202210725935.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-23
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2042-06-23
AI Technical Summary
The lack of effective methods in existing technologies to monitor the emotions of both parties during voice communication can lead to overly aggressive language and disputes when emotions are abnormal.
The first detection mode monitors the emotions of the communication partner in real time. If an abnormality is detected, it is dealt with. The second detection mode is then used to monitor the emotions of the next communication partner. Combined with audio processing technology, the volume is reduced or interference audio is added to avoid disputes.
This effectively avoids disputes between communication partners caused by emotional abnormalities and improves the efficiency of audio detection.
Smart Images

Figure CN115273901B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of voice communication, and more particularly, to an audio detection method and device, an electronic device, and a storage medium. BACKGROUND
[0002] In daily life, people are prone to say dirty words and cause quarrels and hurt each other's feelings when communicating with others through telephone, voice, and other voice communication methods due to bad mood.
[0003] In the process of research and practice of related technologies, the present application found that there is currently no effective method for monitoring the emotions of both parties in the voice communication process. Therefore, how to monitor the emotions of the communication object in the voice communication process to avoid aggressive language due to emotional abnormalities is a problem that needs to be solved at present. SUMMARY
[0004] In view of the above problems, the present application provides an audio detection method, device, electronic device, and storage medium, which can monitor the emotions of the communication object through real-time audio detection, and improve the efficiency of audio detection.
[0005] In a first aspect, the embodiments of the present application provide an audio detection method, which comprises: obtaining a first emotion detection result of first to-be-detected audio of a first object according to a first detection mode; performing first audio processing on the first to-be-detected audio when the emotion of the first to-be-detected audio of the first user is abnormal; obtaining a second emotion detection result of second to-be-detected audio of a second object according to a second detection mode; wherein the collection time of the second to-be-detected audio is later than that of the first to-be-detected audio; and performing second audio processing on the second to-be-detected audio when the emotion of the second to-be-detected audio of the second user is abnormal.
[0006] In a second aspect, the embodiments of the present application provide an audio detection device, which comprises: a first detection module configured to obtain a first emotion detection result of first to-be-detected audio of a first object according to a first detection mode; a first processing module configured to perform first audio processing on the first to-be-detected audio when the emotion of the first to-be-detected audio of the first user is abnormal; a second detection module configured to obtain a second emotion detection result of second to-be-detected audio of a second object according to a second detection mode; wherein the collection time of the second to-be-detected audio is later than that of the first to-be-detected audio; and a second processing module configured to perform second audio processing on the second to-be-detected audio when the emotion of the second to-be-detected audio of the second user is abnormal.
[0007] In a third aspect, an electronic device is provided, which includes one or more processors, a memory, and one or more application programs. The one or more application programs are stored in the memory and configured to be executed by the one or more processors. The one or more application programs are configured to perform the audio detection method according to the first aspect.
[0008] In a fourth aspect, a computer-readable storage medium is provided, which stores program codes. The program codes can be invoked by a processor to perform the audio detection method according to the first aspect.
[0009] The audio detection method, device, electronic device, and storage medium provided in the present application relate to the technical field of voice communication. Specifically, the method comprises: obtaining a first emotion detection result of first to-be-detected audio of a first object according to a first detection mode; performing first audio processing on the first to-be-detected audio when the emotion of the first to-be-detected audio of the first user is abnormal; obtaining a second emotion detection result of second to-be-detected audio of a second object according to a second detection mode; wherein the collection time of the second to-be-detected audio is later than that of the first to-be-detected audio; and performing second audio processing on the second to-be-detected audio when the emotion of the second to-be-detected audio of the second user is abnormal. Thus, the emotion of each object can be monitored in real time through the first detection mode, and the next communication object can be monitored and processed through the corresponding second detection mode after an abnormal emotion of a communication object is detected, which can effectively avoid disputes between communication objects caused by abnormal emotions and greatly improve the efficiency of audio detection. BRIEF DESCRIPTION OF DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort.
[0011] Figure 1 A flowchart of an audio detection method according to an embodiment of the present application is shown;
[0012] Figure 2 A detection processing flowchart of an audio detection method according to an embodiment of the present application is shown;
[0013] Figure 3 A structural block diagram of an audio detection device according to an embodiment of the present application is shown;
[0014] Figure 4 A structural block diagram of an electronic device according to an embodiment of the present application is shown.
[0015] Figure 5 A structural block diagram of a computer readable storage medium according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0016] In order to enable those skilled in the art to better understand the scheme of the present application, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0017] In today's society, people often communicate with others through voice communication, which may be, for example, mobile phone, voice phone and the like. During voice communication, if one party has a bad mood and swears due to some reasons, the other party is also likely to have a bad mood. If no corresponding measures are taken, the two parties are likely to quarrel, which may cause damage to each other's feelings.
[0018] In the process of research and practice of related technologies, the inventors of the present application found that there is no effective method for monitoring the emotions of both parties during voice communication. Therefore, how to monitor the emotions of the communication object during voice communication to avoid aggressive language due to emotional abnormalities is a problem that needs to be solved at present.
[0019] Therefore, in order to overcome the above-mentioned defects, the inventors of the present application propose an audio detection method, device, electronic equipment and storage medium provided by the embodiments of the present application, which specifically includes: obtaining a first emotion detection result of a first to-be-detected audio of a first object according to a first detection mode; when the emotion of the first to-be-detected audio of the first user is abnormal, performing first audio processing on the first to-be-detected audio; obtaining a second emotion detection result of a second to-be-detected audio of a second object according to a second detection mode; wherein the collection time of the second to-be-detected audio is later than that of the first to-be-detected audio; when the emotion of the second to-be-detected audio of the second user is abnormal, performing second audio processing on the second to-be-detected audio. Thus, the emotions of each object can be monitored in real time through the first detection mode, and the next communication object is monitored and processed in the corresponding second detection mode after an abnormal emotion of a communication object is detected, which greatly improves the efficiency of audio detection and effectively avoids the situation that the communication objects quarrel due to emotional abnormalities.
[0020] The embodiments will be described below.
[0021] Please refer to Figure 1 , Figure 1 An audio detection method provided by the embodiments of the present application is shown. Specifically, the method can include steps 110 to 140.
[0022] In step 110, a first emotion detection result of a first to-be-detected audio of the first object is obtained according to the first detection mode.
[0023] In the embodiments of the present application, the audio detection method can be applied to a server. The server can receive audio data sent by a communication object who is speaking, and forward the audio data to other communication objects to realize voice communication. The server can be a separate server or a server cluster, and can be a local server or a cloud server.
[0024] In some embodiments, the communication objects can communicate through terminal devices (mobile phones, wearable watches, telephones, etc.) to realize voice communication. Alternatively, the terminal devices can be used for mobile communication, that is, communication through dialing a mobile phone number. Alternatively, the terminal devices can be used for voice communication through a client installed on the terminal devices, such as WeChat voice / video call, QQ voice / video call, etc.
[0025] Further, the client can be a computer application (APP) installed on the terminal device, or a Web client, which refers to an application developed based on a Web architecture.
[0026] In some embodiments, the terminal device and the server can communicate through a network, which is usually the Internet, but can also be any network, including but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or any combination of virtual private networks. In addition, the terminal device and the server can also communicate through a specific communication protocol, including but not limited to a BLE (Bluetooth low energy) protocol, a WLAN (Wireless Local Area Network) protocol, a Bluetooth protocol, a ZigBee protocol or a Wi-Fi (Wireless Fidelity) protocol, etc.
[0027] In the embodiments of the present application, if there is no speaking object with emotional abnormality before the current speaking communication object (the subsequent speaking communication object is referred to as a speaking object) in the voice communication process, the current speaking communication object is taken as the first object, and the first to-be-detected audio is the corresponding communication audio in the current speaking process of the first object.
[0028] For example, there are communication objects A and B in the communication process, and the communication object A speaks first. There is no emotionally abnormal speaking object before the communication object A speaks, so the communication object A is determined as the first object A during the communication object A speaking, and the communication audio corresponding to the current speaking process of the first object A is regarded as the first to-be-detected audio A. If it is determined that the first object A has no emotional abnormality, the communication object B can be determined as the first object B when the communication object B speaks, so that the communication audio corresponding to the current speaking process of the first object B is the first to-be-detected audio B.
[0029] In the embodiment of the application, the first emotion detection result is an emotion detection result of the first user in the first to-be-detected audio, and the emotion detection result includes emotional normality and emotional abnormality. In the voice communication process, there are at least two communication objects, and one to-be-detected audio can include multiple pieces of audio data, i.e., one to-be-detected audio can include multiple pieces of voice audio spoken by one communication object.
[0030] In the embodiment of the application, the first detection mode can be a detection mode used when there is no emotionally abnormal speaking object before the current speaking object, i.e., the first detection mode can be used to detect the to-be-detected audio corresponding to the first object. It can be understood that if the current speaking object is the first speaking object, there is no emotional abnormality of the previous speaking object, so the first speaking object is the first object, and the first detection mode is used to perform emotion detection on the communication audio corresponding to the first object, i.e., the first to-be-detected audio.
[0031] Specifically, the server performs emotion detection on the first to-be-detected audio of the first object using the first detection mode to obtain the emotion detection result of the first to-be-detected audio, i.e., the first emotion detection result. It can be understood that the emotional abnormality here can refer to that the communication object is in an emotional agitation state, i.e., is about to be angry, or is already in an angry state, etc.
[0032] In some embodiments, the preset conditions corresponding to each detection mode can be different, i.e., the standards for judging emotional abnormality in different detection modes can be different.
[0033] In some embodiments, the first detection mode can determine whether the communication object has emotional abnormality in the to-be-detected audio through the first preset condition, and the first preset condition can be at least one of the existence of sensitive content in the to-be-detected audio and the increase of the audio parameter exceeding a first threshold. If the to-be-detected audio satisfies the first preset condition, it can be considered that the speaking object corresponding to the to-be-detected audio has emotional abnormality in the to-be-detected audio. Therefore, the first detection mode can be to detect the sensitive information and the audio information of the to-be-detected audio.
[0034] Specifically, the server performs sensitive information and audio information detection on the first to-be-detected audio of the first object after obtaining the first to-be-detected audio of the first object. If at least one of the following conditions is met, i.e., if the first to-be-detected audio contains sensitive content or the increment of the audio parameter exceeds the first threshold, or both, it is determined that the emotion of the first to-be-detected audio of the first user is abnormal. If neither of the above conditions is met, it is determined that the emotion of the first to-be-detected audio of the first user is normal.
[0035] Further, the server can include a sensitive information detection module and an audio information detection module, which are respectively used to detect sensitive content and the increment of the audio parameter in the audio segment constituting the to-be-detected audio.
[0036] In some embodiments, the sensitive content can be a pre-set sensitive word. Specifically, the to-be-detected audio can be input into a pre-determined sensitive word library, and then an algorithm such as AC (Aho-Corasick matching algorithm), DFA (Deterministic Finite Automaton), or the like can be used for sensitive content detection.
[0037] In some embodiments, the audio parameter can be the volume, and the increment of the audio parameter can be the increment of the volume, which can be represented by the difference between the highest point and the lowest point of the energy of the audio. Specifically, the difference between the lowest point and the highest point of the energy of the to-be-detected audio within a preset time can be recorded using the RMS (Rate Monotonic Scheduling) value, so as to determine the difference in the volume increase or decrease of the to-be-detected audio.
[0038] In the embodiments of the present application, in order to ensure that other communication objects can receive the communication audio content of the current speaking object in real time, the to-be-detected audio is not received in its entirety before detection, and then sent to other communication objects after detection is completed. Instead, the sensitive information and audio information detection are performed on each audio segment of a preset length as it is received. If it is detected that the audio segment meets the preset condition, the audio segment is processed, and after processing, the audio segment is sent to the terminal device where the other communication object is located, and the next audio segment is detected simultaneously. The to-be-detected audio can be a communication audio composed of one or more audio segments.
[0039] Therefore, the detection and processing of the to-be-detected audio in the embodiments of the present application are actually the detection and processing of the audio segments constituting the to-be-detected audio.
[0040] In step 120, when the emotion of the first to-be-detected audio of the first user is abnormal, the first audio processing is performed on the first to-be-detected audio.
[0041] In the embodiments of the present application, after determining that the emotion of the first to-be-detected audio of the first user is abnormal, that is, the first user has emotional abnormalities in the first to-be-detected audio, the server will perform first audio processing on the first to-be-detected audio to reduce the probability that other communication objects also have emotional abnormalities after receiving the first to-be-detected audio.
[0042] In some embodiments, the method of the first audio processing may, for example, be reducing the volume of the to-be-detected audio, adding a first interference audio to the to-be-detected audio, etc. Wherein the specific way of reducing the volume of the to-be-detected audio can be to reduce the volume of the to-be-detected audio to a first preset volume; the first interference audio can be a preset noise, music, etc. Audio content.
[0043] Further, if there is sensitive content in the audio segments constituting the to-be-detected audio, the method of the first audio processing can further include de-audio processing of the sensitive content. Alternatively, the de-audio processing can be to de-audio the entire audio segment where the sensitive content exists. Alternatively, the de-audio processing can also be to de-audio the part of the audio segment where the sensitive content exists, and the rest is not processed.
[0044] It can be understood that the specific method of the first audio processing can also be a combination of the above methods. Further, the specific method of the first audio processing can also have other processing methods, which are not limited.
[0045] In some embodiments, since other communication objects may interrupt or interrupt during the speech of one communication object, in order to determine the detection mode of the corresponding to-be-detected audio of the other communication object in time when the other communication object speaks, an emotional abnormality identifier can be recorded in the server after detecting that the audio segment constituting the to-be-detected audio of the speaker object has emotional abnormalities. Wherein the emotional abnormality identifier can be an identifier ID, or a temporary file, indicating that the previous speaker object of the current speaker object has emotional abnormalities.
[0046] Specifically, the server can determine that the speaking object has emotion abnormality in the current audio to be detected after detecting for the first time that the audio segment of the speaking object meets the preset condition, and then record an emotion abnormality identifier in the server, so that no matter whether other audio segments in the audio to be detected of the speaking object have been detected, as long as it is determined that the emotion abnormality identifier exists in the server before detecting the audio to be detected of the next speaking object, the detection mode of the audio to be detected of the next speaking object can be determined.
[0047] It can be understood that after determining for the first time that the audio segment meets the preset condition and then recording the emotion abnormality identifier in the server, the audio detection and processing of the audio segment after the speaking object are not affected, that is, as long as the audio to be detected of the speaking object is not detected, the detection and processing will be continued normally.
[0048] In step 130, a second emotion detection result of a second audio to be detected of a second object is obtained according to a second detection mode.
[0049] In the embodiment of the present application, if there is a speaking object with emotion abnormality before the current speaking object, the current speaking object is taken as the second object; and the second audio to be detected is the corresponding communication audio in the current speaking process of the second object.
[0050] For example, there are communication objects A, B and C in the communication process, the communication object A is the first object A, and it is determined through detection that the first object A has emotion abnormality, then when the communication object B speaks, the communication object B can be determined as the second object B, so that the corresponding communication audio of the second object B in the current speaking process is the second audio to be detected B, and if it is detected that the second object B also has emotion abnormality, the communication object C speaking after the communication object B is taken as the second object C, and the corresponding communication audio of the second object C in the current speaking process is the second audio to be detected C.
[0051] In the embodiment of the present application, the second emotion detection result is the emotion detection result of the second user in the second audio to be detected; and the second detection mode can be the detection mode used when there is a speaking object with emotion abnormality before the current speaking object. That is, if the server detects that there is an emotion abnormality identifier when the current speaking object speaks, the second detection mode is used to detect whether there is emotion abnormality in the audio to be detected of the current speaking object.
[0052] Specifically, after detecting that there is an emotion abnormality identifier, the server can determine that the first user has emotion abnormality in the first audio to be detected, so that after obtaining the second audio to be detected of the second object, the emotion detection result of the second audio to be detected of the second object, that is, the second emotion detection result, is obtained by using the second detection mode.
[0053] Further, in the embodiments of the present application, since the detection mode of the to-be-detected audio of the next speaker object is determined according to the emotion detection result of the previous speaker object, in order to make the emotion abnormality identifier stored in the server as the emotion detection result of the previous speaker object when the next speaker object speaks, the server needs to delete the stored emotion abnormality identifier after the to-be-detected audio corresponding to the speaker object with emotion abnormality is detected by using the second detection mode, and then determine whether to record the emotion abnormality identifier again according to the emotion detection result of the next speaker object.
[0054] In some embodiments, the presence of several communication objects in the communication process and which communication object the communication audio corresponding to each time point belongs to can be determined by using a human voice segmentation technology. Specifically, the server determines which communication object the to-be-detected audio corresponding to each time point belongs to by using the human voice segmentation technology after obtaining the to-be-detected audio sent by the terminal device each time. The time point at which the server receives the to-be-detected audio can be regarded as the collection time of the to-be-detected audio.
[0055] Further, since multiple communication objects may speak at the same time during the voice communication process, in order to more accurately determine the order in which the to-be-detected audio is received by the server, that is, more accurately determine the order in which each communication object speaks, the collection time can be at least accurate to the millisecond level. It can be understood that the collection time can also be accurate to the microsecond level, nanosecond level, etc. according to actual needs. In the embodiments of the present application, the collection time of the second to-be-detected audio is later than the collection time of the first to-be-detected audio.
[0056] In some embodiments, the server can determine the first to-be-detected audio of the first object and the collection time of the first to-be-detected audio from the multiple communication objects in the communication process according to the human voice segmentation technology. Specifically, the server determines the first object from the multiple communication objects according to the human voice segmentation technology, takes the first to-be-detected audio corresponding to the first object as the first to-be-detected audio, and takes the time point at which the first to-be-detected audio is received as the collection time of the first to-be-detected audio. Further, the second speaker object can be determined as the second object from the multiple communication objects by using a similar method, the first to-be-detected audio corresponding to the second object is taken as the second to-be-detected audio, and the time point at which the second to-be-detected audio is received is taken as the collection time of the second to-be-detected audio.
[0057] In some embodiments, in the second detection mode, the second preset condition can be used to determine whether the communication object has emotional abnormality in the to-be-detected audio. Since the speaker object after the previous speaker object has emotional abnormality is more likely to have emotional abnormality, a stricter detection condition can be used to detect the to-be-detected audio of the speaker object after the previous speaker object.
[0058] As an example, the second preset condition can be that at least one of the to-be-detected audio has sensitive content and the increase value of the audio parameter exceeds the second threshold. If the to-be-detected audio meets the second preset condition, it can be considered that the speaker object corresponding to the to-be-detected audio has emotional abnormality in the to-be-detected audio, and the second threshold is lower than the first threshold. Specifically, after the server determines that the first to-be-detected audio of the first object has emotional abnormality, the second detection mode is used to detect the second to-be-detected audio of the second object. If at least one of the second to-be-detected audio has sensitive content and the increase value of the audio parameter exceeds the second threshold, it is determined that the second to-be-detected audio of the second user has emotional abnormality, otherwise, it is determined that the second to-be-detected audio of the second user has emotional normality. It can be understood that the detection process of the second detection mode is similar to the detection process when the preset condition in the first detection mode is the first preset condition, and the detailed process will not be described here.
[0059] As an example, the second preset condition can also be that the preset audio part of the to-be-detected audio has sensitive information, that is, the second detection mode can be sensitive information detection on the preset audio part of the to-be-detected audio. Specifically, after the server determines that the first to-be-detected audio of the first object has emotional abnormality, the second detection mode is used to detect the preset audio part of the second to-be-detected audio of the second object. If the preset audio part of the second to-be-detected audio has sensitive content, it is determined that the second to-be-detected audio of the second user has emotional abnormality, otherwise, it is determined that the second to-be-detected audio of the second user has emotional normality.
[0060] Since in the case that the previous speaker object has emotional abnormality, the next speaker object is most likely to have emotional abnormality at the beginning of speaking, the preset audio part can be the first sentence audio part in the to-be-detected audio, or the audio part of the preset time length in front of the to-be-detected audio. The specific setting can be made according to actual needs, and the embodiments of the present application do not limit this.
[0061] In step 140, when the second to-be-detected audio of the second user has emotional abnormality, the second audio processing is performed on the second to-be-detected audio.
[0062] In the embodiments of the present application, after determining that the second emotion of the second user in the second to-be-detected audio is abnormal, i.e., the second user has an emotional abnormality in the second to-be-detected audio, the server performs second audio processing on the second to-be-detected audio to reduce the probability of emotional abnormality of other communication objects after receiving the second to-be-detected audio.
[0063] The specific method of the second audio processing is similar to the first audio processing method described above, and will not be described here.
[0064] In some embodiments, after detecting that the to-be-detected audio of the communication object has an emotional abnormality, in addition to processing the to-be-detected audio, a prompt information can also be displayed on the display interface of the terminal device corresponding to the communication object with the emotional abnormality to remind the communication object to keep the emotion stable. The prompt information can be displayed in the form of text, pictures, etc., and the specific content of the prompt information can be set according to actual needs, which is not limited.
[0065] In some embodiments, if it is determined that the second to-be-detected audio of the second user has an emotional abnormality, a third to-be-detected audio of a communication object after the second object is detected using a second detection mode to obtain a third emotional detection result until the third emotional detection result is normal. Wherein, in the case of determining that the second object has an emotional abnormality, the communication object after the second object is still the second object, but is a different communication object from the current communication object determined to have an emotional abnormality, so that the third to-be-detected audio is also the communication audio corresponding to the current speaking process of the second object.
[0066] For example, there are communication objects A and B, the first speaker communication object A is the second object A, and there is an emotional abnormality. Then the subsequent speaker communication object B is taken as the second object B, and the communication audio corresponding to the current speaking process of the second object B is taken as the third to-be-detected audio B.
[0067] In some embodiments, if the emotion of the second to-be-detected audio of the second user is normal, a fourth to-be-detected audio of a communication object after the second object is detected according to a first detection mode. Wherein, in the case of determining that the second object does not have an emotional abnormality, i.e., the emotion is normal, the communication object after the second object is the first object, so that the fourth to-be-detected audio is also the communication audio corresponding to the current speaking process of the first object.
[0068] That is, in the embodiments of the present application, as long as the previous speaking object has an emotional abnormality, the to-be-detected audio corresponding to the subsequent speaking object is detected using the second detection mode; as long as the previous speaking object has a normal emotion, the to-be-detected audio of the subsequent speaking object is detected using the first detection mode. Wherein, the to-be-detected audio corresponding to the first speaking object is detected using the first detection mode.
[0069] Please refer to Figure 2 , Figure 2 A detection process flow chart of the audio detection method is shown, specifically:
[0070] In step S1, a first to-be-detected audio of a first object is acquired;
[0071] In step S2, emotion detection is performed on the first to-be-detected audio of the first object using a first detection mode; if the first object is normal in emotion in the first to-be-detected audio, return to step S1 to acquire the first to-be-detected audio of the next first object; if the first object is abnormal in emotion in the first to-be-detected audio, proceed to step S3;
[0072] In step S3, first audio processing is performed on the first to-be-detected audio of the first object;
[0073] In step S4, a second to-be-detected audio of a second object is acquired;
[0074] In step S5, emotion detection is performed on the second to-be-detected audio of the second object using a second detection mode; if the second object is abnormal in emotion in the second to-be-detected audio, proceed to step S6; if the second object is normal in emotion in the second to-be-detected audio, return to step S1;
[0075] In step S6, second audio processing is performed on the second to-be-detected audio of the second object; after the processing, return to step S4.
[0076] Among them, after returning from step S6 to step S4, the second to-be-detected audio of the next second object is acquired, but in order to better describe, in the embodiment of the application, the specific operation of the returned step is described as acquiring the third to-be-detected audio of the next second object; similarly, the operation after returning from step S5 to step S1 is described as acquiring the fourth to-be-detected audio of the next first object.
[0077] It can be understood that although there are differences in the description, in essence, it is to Figure 2 The detection process flow chart shown continuously detects and processes the communication audio corresponding to the communication object in the voice communication process in real time, the audio detection efficiency is high, and the occurrence of disputes between communication objects caused by emotional abnormalities can be effectively avoided.
[0078] The embodiment of the application obtains a first emotion detection result of first to-be-detected audio of a first object according to a first detection mode; when the emotion of the first to-be-detected audio of the first user is abnormal, performing first audio processing on the first to-be-detected audio; obtaining a second emotion detection result of second to-be-detected audio of a second object according to a second detection mode; wherein the collection time of the second to-be-detected audio is later than that of the first to-be-detected audio; when the emotion of the second to-be-detected audio of the second user is abnormal, performing second audio processing on the second to-be-detected audio. Thus, the emotion of each object can be monitored in real time through the first detection mode, and after the emotion of a communication object is monitored to be abnormal, processing is performed, and the next communication object is monitored and processed through the corresponding second detection mode, which greatly improves the efficiency of communication audio detection, and effectively avoids the situation that communication objects quarrel because of emotion abnormality.
[0079] Please refer to Figure 3 , Figure 3 A structure block diagram of an audio detection device 200 provided by the embodiment of the application is shown. The audio detection device 200 includes a first detection module 210, a first processing module 220, a second detection module 230, and a second processing module 240, specifically:
[0080] The first detection module 210 is configured to obtain a first emotion detection result of first to-be-detected audio of a first object according to a first detection mode.
[0081] The first processing module 220 is configured to perform first audio processing on the first to-be-detected audio when the emotion of the first to-be-detected audio of the first user is abnormal.
[0082] The second detection module 230 is configured to obtain a second emotion detection result of second to-be-detected audio of a second object according to a second detection mode; wherein the collection time of the second to-be-detected audio is later than that of the first to-be-detected audio.
[0083] The second processing module 240 is configured to perform second audio processing on the second to-be-detected audio when the emotion of the second to-be-detected audio of the second user is abnormal.
[0084] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described device and module can refer to the corresponding process in the foregoing method embodiments, which will not be described herein.
[0085] In several embodiments provided in the application, the coupling between the modules can be electrical, mechanical or other forms of coupling.
[0086] In addition, each of the functional modules in the embodiments of the present application can be integrated in one processing module, or each of the modules can be physically present individually, or two or more modules can be integrated in one module. The integrated module can be implemented in the form of hardware or in the form of a software functional module.
[0087] Please refer to Figure 4 , Figure 4 A structural block diagram of an electronic device 300 provided by an embodiment of the present application is shown. The electronic device 300 can be a notebook computer, a desktop computer, or the like, which can run an application program. The electronic device 300 in the present application can include one or more of the following components: a processor 310, a memory 320, and one or more application programs, wherein the one or more application programs can be stored in the memory 320 and configured to be executed by the one or more processors 310, and the one or more programs are configured to perform the method as described in the foregoing method embodiments.
[0088] The processor 310 can include one or more processing cores. The processor 310 connects various parts in the entire electronic device 300 by various interfaces and lines, performs various functions of the electronic device 300 and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 320, and calling data stored in the memory 320. Optionally, the processor 310 can be implemented in at least one of a hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 310 can integrate a combination of one or several of a central processing unit (CPU), a graphics processor (GPU), and a modem, and the like. Among them, the CPU mainly processes an operating system, a user interface, and an application program, and the like; the GPU is responsible for rendering and drawing display content; and the modem is used for processing wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor 310, but can be implemented by a separate communication chip.
[0089] The memory 320 can include a random access memory (RAM) and can also include a read-only memory (ROM). The memory 320 can be used to store instructions, programs, codes, code sets, or instruction sets. The memory 320 can include a program storage area and a data storage area, where the program storage area can store instructions for implementing an operating system, instructions for implementing at least one function (such as a detection function, an audio processing function, etc.), instructions for implementing each of the method embodiments described below, and the like. The data storage area can also store data created by the electronic device 300 in use (such as first audio to be detected, first emotion detection result, second audio to be detected, second emotion detection result, etc.).
[0090] Referring to Figure 5 , Figure 5 A structural block diagram of a computer-readable storage medium provided by an embodiment of the present application is shown. The computer-readable storage medium 400 stores program codes therein, which can be invoked by a processor to execute the audio detection method described in the above method embodiments.
[0091] The computer-readable storage medium 400 can be an electronic memory such as a flash memory, an EEPROM (electrically erasable programmable read-only memory), an EPROM, a hard disk, or a ROM. Alternatively, the computer-readable storage medium 400 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 400 has a storage space for program codes 410 for executing any of the method steps described above. These program codes can be read from or written to one or more computer program products. The program codes 410 can be compressed in an appropriate form, for example.
[0092] The embodiments of the present application also provide a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to cause the computer device to perform the audio detection method described in the various optional embodiments described above.
[0093] The audio detection method, device, electronic equipment and storage medium of the present application specifically disclose the following: obtaining a first emotion detection result of first to-be-detected audio of a first object according to a first detection mode; when the emotion of the first to-be-detected audio of the first user is abnormal, performing first audio processing on the first to-be-detected audio; obtaining a second emotion detection result of second to-be-detected audio of a second object according to a second detection mode; wherein the collection time of the second to-be-detected audio is later than that of the first to-be-detected audio; when the emotion of the second to-be-detected audio of the second user is abnormal, performing second audio processing on the second to-be-detected audio. Thus, the emotion of each object can be monitored in real time through the first detection mode, and after it is monitored that the emotion of a communication object is abnormal, processing is performed, and the next communication object is monitored and processed in the corresponding second detection mode, greatly improving the efficiency of audio detection, and effectively avoiding the situation that communication objects quarrel because of emotional abnormalities.
[0094] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not drive the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An audio detection method, characterized in that, The method includes: According to the first detection mode, the first emotion detection result of the first audio to be detected of the first object is obtained. The first detection mode determines that there is an emotional abnormality by satisfying the first preset condition. The first preset condition includes that the increase of the audio parameters of the audio to be detected exceeds the first threshold. When the first object's first audio to be detected has an abnormal emotion, the first audio to be detected is subjected to first audio processing; The second emotion detection result of the second audio to be detected of the second object is obtained according to the second detection mode; wherein, the acquisition time of the second audio to be detected is later than that of the first audio to be detected, and the second detection mode determines that there is an emotional abnormality by satisfying the second preset condition, the second preset condition including the increase of the audio parameters of the audio to be detected exceeding the second threshold, and the second threshold being lower than the first threshold; When the emotion of the second audio to be detected in the second object is abnormal, the second audio to be detected is subjected to second audio processing.
2. The method according to claim 1, characterized in that, After performing the second audio processing on the second audio to be detected, the method further includes: According to the second detection mode, the third emotion detection result of the third audio to be detected of the communication object after the second object is obtained until the third emotion detection result is normal.
3. The method according to claim 1, characterized in that, The method further includes: If the emotion of the second audio to be detected of the second object is normal, the fourth audio to be detected of the communication object after the second object is detected according to the first detection mode.
4. The method according to claim 1, characterized in that, The step of obtaining the first emotion detection result of the first audio to be detected of the first object according to the first detection mode includes: Sensitive information and audio information are detected in the first audio file of the first object to be detected; If the first audio to be detected contains sensitive content and the increase in audio parameters exceeds the first threshold, then it is determined that the first audio to be detected of the first object has an abnormal emotion. Otherwise, it is determined that the emotion of the first audio to be detected of the first object is normal.
5. The method according to claim 4, characterized in that, The step of obtaining the second emotion detection result of the second audio to be detected of the second object according to the second detection mode includes: Sensitive information and audio information are detected in the second audio of the second object; If the second audio to be detected contains sensitive content and the increase in audio parameters exceeds the second threshold, then it is determined that the second audio to be detected of the second object has an abnormal emotion; wherein, the second threshold is lower than the first threshold; Otherwise, it is determined that the emotion of the second audio to be detected by the second object is normal.
6. The method according to claim 1, characterized in that, Before obtaining the first emotion detection result of the first audio to be detected of the first object according to the first detection mode, the method further includes: The first audio to be detected of the first object is determined from multiple communication objects based on human voice segmentation technology, and the acquisition time of the first audio to be detected is also determined. Before obtaining the second emotion detection result of the second audio to be detected of the second object according to the second detection mode, the method further includes: The second audio to be detected of the second object is determined from multiple communication objects based on human voice segmentation technology, and the acquisition time of the second audio to be detected is also determined.
7. An audio detection device, characterized in that, The device includes: The first detection module is used to obtain the first emotion detection result of the first audio to be detected of the first object according to the first detection mode. The first detection mode determines that there is an emotional abnormality by satisfying the first preset condition. The first preset condition includes the increase of the audio parameters of the audio to be detected exceeding the first threshold. The first processing module is used to perform first audio processing on the first audio to be detected when the first object has an abnormal emotion in the first audio to be detected. The second detection module is used to obtain the second emotion detection result of the second audio to be detected of the second object according to the second detection mode; wherein, the acquisition time of the second audio to be detected is later than that of the first audio to be detected, and the second detection mode determines that there is an emotional abnormality by satisfying the second preset condition, the second preset condition including that the increase of the audio parameters of the audio to be detected exceeds the second threshold, and the second threshold is lower than the first threshold; The second processing module is used to perform second audio processing on the second audio to be detected if the emotion of the second audio to be detected of the second object is abnormal.
8. An electronic device, characterized in that, include: One or more processors; Memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, the one or more applications being configured to perform the audio detection method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains program code that can be called by a processor to execute the audio detection method as described in any one of claims 1-6.
Citation Information
Patent Citations
Communication control method and apparatus in voice communication
CN106790957A