A voice processing method, device and equipment of a mobile device and a storage medium

By integrating an AI voice chip into mobile devices and using preset emotion tags and keywords for voice processing, the complexity caused by the reliance on servers in existing voice detection devices is solved. This enables timely judgment of the target person's thoughts, behaviors, and emotional changes, improving processing efficiency and accuracy.

CN114420162BActive Publication Date: 2025-11-28SEAWAY TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111668937.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-11-28
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

Existing voice detection devices rely on servers, resulting in complex voice processing methods and an inability to promptly assess the thoughts, behaviors, and psychological changes of target individuals.

Method used

By integrating an AI voice chip into mobile devices, user voice information is collected in real time. Voice processing is performed using preset emotion tags and keywords to generate voice processing results, enabling timely analysis of the target person's thoughts, behaviors, psychological changes, and emotional fluctuations.

Benefits of technology

It enables rapid local recognition and processing of user voice information on mobile devices, timely analysis of target personnel's thoughts, behaviors, and emotional changes, reduces reliance on servers, and improves processing efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114420162B_ABST
    Figure CN114420162B_ABST
Patent Text Reader

Abstract

The application relates to a voice processing method and device of a mobile device, equipment and a storage medium, and relates to the technical field of voice processing. The voice processing method of the mobile device comprises the following steps: acquiring user voice information collected by the mobile device; if the user voice information contains key voice information corresponding to a preset emotion label, the user voice information is determined as target detection voice information; voice processing is performed according to the target voice detection information, and a voice processing result corresponding to the user voice information is obtained. The application solves the problem that the existing voice detection device is dependent on a server, the voice processing mode is complex, and the thought behavior of target personnel cannot be timely analyzed and judged through conversation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of voice processing, and particularly relates to a voice processing method and device of a mobile device, a mobile device and a storage medium. BACKGROUND

[0002] At present, an existing voice detection transmission device mainly detects user voice information through a microphone, stores the voice information locally, transmits the user voice information to a server through a wireless network when accessing the wireless network, and obtains a thought behavior of a target person through the server, so that the user logs in the server through another terminal and obtains the corresponding research and judgment result from the server.

[0003] It can be seen that the existing voice detection device depends on the server, so that the voice processing method is complex and cannot timely research and judge the thought behavior of the target person through conversation. SUMMARY

[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present application provides a voice processing method and device of a mobile device, a mobile device and a storage medium.

[0005] In a first aspect, the present application provides a voice processing method of a mobile device, characterized in that comprising:

[0006] acquiring user voice information collected by the mobile device;

[0007] if the user voice information contains key voice information corresponding to a preset emotion label, determining the user voice information as target detection voice information;

[0008] performing voice processing according to the target voice detection information to obtain a voice processing result corresponding to the user voice information.

[0009] Optionally, after acquiring the user voice information collected by the mobile device, the method further comprises:

[0010] performing text conversion according to the user voice information to obtain text conversion information;

[0011] determining whether the text conversion information contains a target keyword corresponding to the preset emotion label;

[0012] if the text conversion information contains the target keyword, determining that the user voice information contains key voice information corresponding to a preset emotion label;

[0013] if the text conversion information does not contain the target keyword, deleting the user voice information.

[0014] Optionally, the voice processing according to the target voice detection information to obtain a voice processing result corresponding to the user voice information comprises:

[0015] The voice processing system of the mobile device is woken up for the target voice detection information;

[0016] The voice processing result corresponding to the target voice detection information is determined through the voice processing system.

[0017] Optionally, the voice processing result corresponding to the target voice detection information is determined through the voice processing system, and the method comprises:

[0018] The server address information is obtained through the voice processing system;

[0019] The voice processing request is generated according to the server address information and the target voice detection information, and the voice processing request is used to trigger the server to obtain the monitoring image information corresponding to the target voice detection information;

[0020] The target voice detection information is subjected to emotional analysis processing in combination with the monitoring image information to obtain the voice processing result.

[0021] Optionally, the method further comprises:

[0022] The alarm prompt information corresponding to the target keyword is generated based on the voice processing result;

[0023] The alarm prompt information is outputted according to the alarm prompt information.

[0024] In a second aspect, the application provides a voice processing device of a mobile device, comprising:

[0025] A user voice acquisition module is configured to acquire user voice information collected by the mobile device;

[0026] A target detection voice information determination module is configured to determine the user voice information as target detection voice information when the user voice information contains keyword voice information corresponding to a preset emotional label;

[0027] A voice processing result generation module is configured to perform voice processing according to the target voice detection information to obtain a voice processing result corresponding to the user voice information.

[0028] Optionally, after the user voice information collected by the mobile device is acquired, the method further comprises:

[0029] A voice text conversion module is configured to perform text conversion according to the user voice information to obtain text conversion information;

[0030] The target keyword judgment module is configured to judge whether the text conversion information contains a target keyword corresponding to the preset emotional label.

[0031] The key voice information determination module is configured to determine that the user voice information contains key voice information corresponding to a preset emotional label when the text conversion information contains the target keyword.

[0032] The voice information deletion module is configured to delete the user voice information when the text conversion information does not contain the target keyword.

[0033] Optionally, the voice processing result generation module comprises:

[0034] The voice processing system wake-up sub-module is configured to wake up a voice processing system of the mobile device for the target voice detection information.

[0035] The voice processing result determination sub-module is configured to determine a voice processing result corresponding to the target voice detection information by using the voice processing system.

[0036] In a third aspect, the present application provides a voice processing device of a mobile device, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory complete mutual communication through the communication bus.

[0037] The memory is configured to store a computer program.

[0038] The processor is configured to execute the program stored on the memory, and implement the steps of the voice processing method of the mobile device according to any one of the embodiments of the first aspect.

[0039] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the voice processing method of the mobile device according to any one of the embodiments of the first aspect.

[0040] In summary, the present application acquires user voice information collected by a mobile device, determines the user voice information as target detection voice information if the user voice information contains key voice information corresponding to a preset emotional label, performs voice processing according to the target voice detection information, and obtains a voice processing result corresponding to the user voice information, thereby solving the problem that the existing voice transmission device needs to transmit acquired voice information to a server for voice detection, and thus cannot timely determine the thought behavior, psychological change and emotional fluctuation of a caller. BRIEF DESCRIPTION OF DRAWINGS

[0041] The accompanying drawings, which are incorporated into and constitute a part of this specification, illustrate embodiments consistent with the present application and serve to explain the principles of the present application together with the specification.

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings required by the embodiments or prior art description will be briefly introduced as follows. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without any creative effort.

[0043] Figure 1 A flowchart of a voice processing method of a mobile device provided by an embodiment of the present application;

[0044] Figure 2 A step flowchart of a voice processing method of a mobile device provided by an optional embodiment of the present application;

[0045] Figure 3 A structural diagram of a mobile device provided by an embodiment of the present application;

[0046] Figure 4 A structural block diagram of a voice processing device of a mobile device provided by an embodiment of the present application;

[0047] Figure 5 A structural diagram of a voice processing device of a mobile device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0048] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without any creative effort fall within the scope of protection of the present application.

[0049] In a specific implementation, when the existing voice detection device performs voice detection, it needs to be powered by 220 volts separately, which makes the voice detection device inconvenient to carry and work, has a large power consumption, and cannot timely analyze the thought behavior and psychological changes of the target personnel due to its dependence on the server.

[0050] One of the core ideas of the embodiments of the present application is to provide a voice processing method of a mobile device. The voice processing method detects key voice in collected user voice information, processes the user voice to obtain a voice processing result when the key voice is detected, timely analyzes the thought behavior, psychological change and emotional fluctuation of a target person through the collection of user voice information, avoids the existing voice processing device that sends the collected voice information to a server, and then performs voice detection recognition and voice processing through the server, which leads to a complex voice recognition process and the problem of not timely analyzing the thought behavior and psychological change of the target person.

[0051] To facilitate the understanding of the embodiments of the present application, further explanation and description will be made in combination with the drawings and specific embodiments. The embodiments do not constitute a limitation on the embodiments of the present application.

[0052] Figure 1 A flowchart of a voice processing method of a mobile device provided by the embodiments of the present application is shown in FIG. 1. As shown in FIG. 1, the voice processing method of the mobile device provided by the present application can specifically include the following steps: Figure 1

[0053] Step 110: Obtain user voice information collected by the mobile device.

[0054] Specifically, the mobile device can be a portable voice detection storage processing transmission device, and the present application does not limit this. The mobile device can collect the conversation content of the target person in real time when the user and the target person have a voice conversation, as user voice information. Specifically, a double-channel or multi-channel microphone (Microphone, Mic) can be built into the mobile device to realize high-quality and real-time collection of user voice information through the double-channel or multi-channel Mic.

[0055] In a specific implementation, the mobile device can be built-in with a battery. The voltage value of the battery can be 5 volts, or a battery with a different voltage value can be set according to actual use requirements. The present application does not limit this. The battery supplies power to each chip and module in the mobile device, so that the mobile device can work normally, improves the portability of the mobile device, and avoids the problem that the traditional voice detection storage device is difficult to carry due to the need for 220-volt voltage for separate power supply when working.

[0056] Step 120: If the user voice information contains key voice information corresponding to a preset emotional label, the user voice information is determined as target detection voice information.

[0057] ​Specifically, one or more emotion labels can be preset according to actual use requirements and actual use scenarios, and each emotion label can be preset with one or more corresponding keywords as key voice information. When the user voice information is detected, the user voice information can be compared with the preset key voice information, so as to determine the user voice information as target detection voice information when the key voice information is detected in the user voice information.

[0058] For example, in the case of judging whether the target person has a violent tendency, the preset keyword can be a word such as "destroy" or "attack", and the present example is not limited thereto.

[0059] In actual processing, after the mobile device collects the user voice information, the user voice information can be converted into text information, and the text information can be compared with the preset key voice information corresponding to the preset emotion label. If words such as "destroy" or "attack" are detected in the text information, the user voice information can be determined as target detection voice information. In subsequent processing, the thought behavior, psychological change, and emotional fluctuation of the target person can be judged in time based on the target detection voice information. For example, if it is judged that the target person has a violent tendency, the user can be warned in time.

[0060] In a specific implementation, an artificial intelligence (AI) voice chip that is pre-trained can be integrated in the mobile device. Based on the key voice information corresponding to the preset emotion label, the AI voice chip can quickly identify the user voice information, which can be completed locally, making the voice processing method more convenient and fast.

[0061] Step 130, performing voice processing according to the target voice detection information to obtain a voice processing result corresponding to the user voice information.

[0062] Specifically, the voice processing result can include the judgment result of the thought behavior, psychological change, and emotional fluctuation of the target person, such as whether the target person has a violent tendency. If it is determined that the target person has a violent tendency, the mobile device can generate corresponding alarm information and output the alarm information to the user of the mobile device for early warning. In addition, in some application scenarios, the psychological change of the target person can also be judged to realize a lie detection function.

[0063] In a specific implementation, after determining the target voice detection information, the mobile device can store the user voice information corresponding to the target voice detection information, which can be stored in the built-in memory of the mobile device, or in a micro SD card (SD) mounted in the mobile device, or in a server. The present application does not limit this. Specifically, the mobile device can be configured with network information in advance, such as wired network, wireless network, or mobile network, etc. The present application does not limit this. After determining the target voice detection information, the target voice detection information and the user voice information corresponding to the target voice detection information can be stored in the server through the network, so that the user can check the stored user voice information at any time through the server, saving local memory space.

[0064] Further, the server can also perform voice processing on the target voice detection information uploaded by the mobile device to obtain a server voice processing result, and can send the server voice processing result to the mobile device through the network. The mobile device can compare the local voice processing result and the server voice processing result, thereby improving the accuracy of the research and judgment result of the target person's thinking behavior, psychological change, and emotional fluctuation, etc.

[0065] To sum up, the embodiments of the present application obtain user voice information collected by a mobile device, determine the user voice information as target detection voice information if the user voice information contains key voice information corresponding to a preset emotion label, perform voice processing according to the target voice detection information to obtain a voice processing result corresponding to the user voice information, thereby solving the problem that the existing voice detection device depends on a server, making the voice processing mode complex, and unable to research and judge the thinking behavior of the target person in time through conversation.

[0066] Reference Figure 2 , a step flow diagram of a voice processing method of a mobile device provided by an optional embodiment of the present application is shown. The voice processing method of the mobile device can specifically include the following steps:

[0067] Step 210, obtaining user voice information collected by the mobile device.

[0068] Step 220, performing text conversion according to the user voice information to obtain text conversion information.

[0069] Step 230, determining whether the text conversion information contains a target keyword corresponding to the preset emotion label.

[0070] Step 240, if the text conversion information contains the target keyword, determining that the user voice information contains key voice information corresponding to a preset emotion label.

[0071] In step 250, if the text conversion information does not contain the target keyword, the user voice information is deleted.

[0072] In step 260, if the user voice information contains the key voice information corresponding to the preset emotional label, the user voice information is determined as the target detection voice information.

[0073] In actual processing, the mobile device can be internally provided with a flash memory module. After the user voice information is collected, the user voice information can be temporarily stored in the flash memory module internally provided in the mobile device, and the user voice information can be subjected to text conversion to obtain text conversion information corresponding to the user voice information. If the text conversion information contains a target keyword corresponding to a preset emotional label, the user voice information can be determined as target detection voice information, and the user voice information can be stored in the internal memory or the mounted SD card of the mobile device. Specifically, the user voice information can be subjected to encryption processing by a preset encryption decoding algorithm to obtain encrypted voice information, and the encrypted voice information can be stored in the internal memory or the mounted SD card of the mobile device. When the user needs to view the relevant voice information, the encrypted voice information can be decrypted and played. If the text conversion information does not contain the target keyword corresponding to the preset emotional label, the user voice information can be deleted from the flash memory module to reduce memory occupation.

[0074] In step 270, the voice processing system of the mobile device is woken up in response to the target voice detection information.

[0075] In step 280, the voice processing system is used to determine a voice processing result corresponding to the target voice detection information.

[0076] Specifically, when the mobile device is powered on, the voice processing system of the mobile device can be in a sleep state to save power. When the AI voice chip determines that the user voice information collected by the two microphones contains key voice information corresponding to a preset emotional label, the target language detection information is determined, and the voice processing system is triggered from the sleep state to the wake-up state based on the target language detection information, the target voice detection information corresponding to the user voice information is subjected to voice processing to obtain a voice processing result, and the thought behavior, psychological change, and emotional fluctuation of the target person are timely analyzed and judged. In subsequent processing, the user of the mobile device can also be timely warned based on the voice processing result, such as timely reminding the user when it is judged that the target person has a violent tendency.

[0077] In a specific implementation, after the mobile device is powered on, the user can configure the network through the voice processing system, such as configuring a wireless network, a wired network, or a mobile network, etc. After completing the network configuration, the user can import the key voice information corresponding to the preset emotion label from the server online through the corresponding network, or the user can preset the key voice information corresponding to the emotion label according to the actual use demand, which is not limited in the present application. After completing the corresponding network configuration, the voice processing system can enter a sleep state to save the power of the portable voice detection storage processing and transmission device.

[0078] Further, the mobile device can also prompt the user of the current working state of the mobile device through a Light-Emitting Diode (LED) lamp. For example, the mobile device can emit different colored light through the LED lamp to prompt the user whether the network configuration is successful when configuring the network, and can also warn the user through different colored light based on the voice processing result, etc., which is not limited in the present example.

[0079] As an example of the present application, referring to Figure 3 , a structural schematic diagram of a mobile device provided by an embodiment of the present application is shown. Specifically, the mobile device can specifically include a wireless transmission module, an AI voice chip, a voice processing module, a dual Pulse Density Modulation (PDM) microphone, a storage module, and a power management module, wherein the power management module can include a charging management, a battery, and a plurality of power conversion chips. Specifically, the voltage value of the battery can be 5 volts, the charging management submodule can charge the battery through an external power supply when the battery power is insufficient, such as the charging management submodule can preset a Micro Universal Serial Bus (USB) or a Type-C USB, and charge the battery through the preset USB interface, and can realize the upgrade of the wireless transmission module and other modules through the USB interface cooperating with the upgrade program. The power conversion chip can be used to convert the 5-volt voltage output by the battery to realize power supply for the wireless transmission module, the voice processing module, the AI voice chip, and the dual PDM microphone, such as it can be converted according to the actual power supply demand of each module, which is not limited in the present example.

[0080] In the process of language processing, the PDM microphone can collect the voice conversation between the user and the target person in real time as user voice information, store the user voice information in the flash memory module, and output the user voice information to the AI voice chip. The AI voice chip can convert the user voice information into text to obtain text conversion information, and can detect the text conversion information based on the preset target keyword information. If the target keyword information is not detected, the user voice information in the flash memory module can be deleted. If the target keyword is detected, the user voice information can be encrypted by a preset encryption and decryption algorithm to store the encrypted user voice information in the storage module, such as an SD card. The present application does not limit this. The unencrypted user voice information can be used as target voice detection information and output to the voice processing module. If the current voice processing system is not connected to the network, the voice processing module can generate a voice processing result based on the target language detection information, timely analyze the target person's behavior, psychological changes, and emotional fluctuations, and output an alarm information to notify the user. If the current voice processing system is connected to the network and connected to the server through the network, the voice processing module can generate a voice processing result based on the target language detection information, and can output the user voice information and the target language detection information corresponding to the user voice information to the server through the wireless transmission module, so as to store the user voice information and the target language detection information in the server, process the voice processing result, and then the voice processing module can generate and output an alarm information based on the local voice processing result and the server voice processing result, thereby improving the accuracy of the analysis result.

[0081] In an optional embodiment, the above-mentioned determination of the voice processing result corresponding to the target voice detection information by the voice processing system can specifically include the following sub-steps:

[0082] Sub-step 2801: obtaining, by the voice processing system, server address information.

[0083] Sub-step 2802: generating a voice processing request according to the server address information and the target voice detection information, the voice processing request being used to trigger the server to obtain monitoring image information corresponding to the target voice detection information.

[0084] Specifically, the server address information can be an Internet Protocol (IP) address of the server, and the application does not make any limitation on this. The corresponding server address information can be preset before the mobile device is shipped, or the user can configure the server address information through the corresponding network after the mobile device is started, and the application does not make any limitation on this. When the voice processing system obtains the target voice detection information, the corresponding voice processing request can be generated according to the target voice detection information, and sent to the server corresponding to the server address information through the preset network. After receiving the voice processing request, the server can obtain the monitoring image information corresponding to the target voice detection information. Specifically, after the user configures the server address information through the corresponding network, the user can set the corresponding camera IP address in the server. When the server receives the voice processing request, the camera can be controlled to collect the monitoring pictures of the current scene based on the camera IP address, such as collecting the monitoring pictures within 10 seconds, and uploading the collected monitoring pictures to the server as the monitoring image information. Of course, the server can also control the camera to upload the pictures collected by the camera within 10 seconds before the receiving time of the voice processing request as the monitoring image information based on the receiving time of the voice processing request, and the application does not make any limitation on this.

[0085] In substep 2803, the target voice detection information is analyzed in combination with the monitoring image information to obtain the voice processing result.

[0086] Specifically, the server can analyze the emotion in combination with the monitoring image information and the target voice detection information, such as analyzing the thought behavior, psychological change, and emotional fluctuation of the target person based on the behavior action of the target person in the monitoring image information in combination with the target voice detection information to obtain the corresponding analysis result as the voice processing result, thereby improving the accuracy of the voice processing result.

[0087] In a specific implementation, the mobile device can generate the corresponding voice processing result based on the target voice detection information, and the server can send the voice processing result to the mobile device after obtaining the voice processing result. In subsequent processing, the mobile device can combine the locally generated voice processing result and the server generated voice processing result, so that the final voice processing result is more accurate, and the analysis accuracy is improved.

[0088] In substep 2804, the alarm prompt information corresponding to the target keyword is generated based on the voice processing result.

[0089] In substep 2805, the alarm prompt information is output.

[0090] Specifically, the alarm prompt information can be generated based on the voice processing result, and the alarm prompt information can be output to remind the user to pay attention to the thought behavior, psychological change and emotional fluctuation of the target person.

[0091] In a specific implementation, if the user does not configure a corresponding network for the mobile device, the mobile device can generate the alarm prompt information based on the local voice processing result. If the mobile device has configured a corresponding network and is connected to a corresponding server, the mobile device can receive the voice processing result sent by the server, and can generate the alarm prompt information in combination with the local voice processing result and the server voice processing result, and output the alarm prompt information.

[0092] It can be seen that, in the embodiments of the present application, the user voice information collected by the mobile device is obtained, text conversion is performed according to the user voice information to obtain text conversion information, it is judged whether the text conversion information contains a target keyword corresponding to a preset emotion label, if the text conversion information contains the target keyword, it is determined that the user voice information contains key voice information corresponding to the preset emotion label, if the text conversion information does not contain the target keyword, the user voice information is deleted, if the user voice information contains key voice information corresponding to the preset emotion label, the user voice information is determined as target detection voice information, the voice processing system of the mobile device is woken up for the target voice detection information, the voice processing result corresponding to the target voice detection information is determined through the voice processing system, and the existing voice detection device is solved, which depends on the server, so that the voice processing mode is complex, and the thought behavior of the target person cannot be timely analyzed and judged through conversation.

[0093] It should be noted that, for the method embodiments, in order to simply describe, they are all described as a series of action combinations, but those skilled in the art should know that the embodiments of the present application are not limited to the action order described, because according to the embodiments of the present application, certain steps can be performed in other order or simultaneously.

[0094] As shown in Figure 4 The embodiments of the present application also provide a voice processing device 400 of a mobile device, which comprises:

[0095] A user voice acquisition module 410 is configured to acquire user voice information collected by the mobile device.

[0096] A target detection voice information determination module 420 is configured to determine the user voice information as target detection voice information when the user voice information contains key voice information corresponding to a preset emotion label.

[0097] The voice processing result generation module 430 is configured to perform voice processing according to the target voice detection information, to obtain a voice processing result corresponding to the user voice information.

[0098] Optionally, after the user voice information collected by the mobile device is obtained, the method further includes:

[0099] The voice text conversion module is configured to perform text conversion according to the user voice information, to obtain text conversion information.

[0100] The target keyword judgment module is configured to judge whether the text conversion information contains a target keyword corresponding to the preset emotional label.

[0101] The key voice information determination module is configured to, when the text conversion information contains the target keyword, determine that the user voice information contains key voice information corresponding to a preset emotional label.

[0102] The voice information deletion module is configured to, when the text conversion information does not contain the target keyword, delete the user voice information.

[0103] Optionally, the voice processing result generation module includes:

[0104] The voice processing system wake-up sub-module is configured to wake up a voice processing system of the mobile device for the target voice detection information.

[0105] The voice processing result determination sub-module is configured to determine, by the voice processing system, a voice processing result corresponding to the target voice detection information.

[0106] Optionally, the voice processing result determination sub-module includes:

[0107] The server address information acquisition unit is configured to acquire, by the voice processing system, server address information.

[0108] The voice processing request generation unit is configured to generate a voice processing request according to the server address information and the target voice detection information, the voice processing request being used to trigger a server to acquire monitoring image information corresponding to the target voice detection information.

[0109] The voice processing result determination unit is configured to perform emotional analysis processing on the target voice detection information in combination with the monitoring image information, to obtain the voice processing result.

[0110] Optionally, the method further includes:

[0111] The alarm prompt information module is configured to generate alarm prompt information corresponding to the target keyword based on the voice processing result.

[0112] The alarm prompt information output module is configured to output the alarm prompt information.

[0113] It should be noted that the speech processing apparatus of the mobile device provided in the embodiments of the present application can execute the speech processing method of the mobile device provided in any of the embodiments of the present application, and has the corresponding functions and advantages of the execution method.

[0114] In a specific implementation, the speech processing apparatus of the mobile device can be integrated in a device, so that the device can perform key speech information detection and speech processing according to the collected user speech information, and obtain a speech processing result corresponding to the user speech information, to serve as a speech processing device of the mobile device, and realize user speech keyword detection. The speech processing device of the mobile device can be composed of two or more physical entities, or can be composed of one physical entity, for example, the device can be a personal computer (PC), a computer, a server, and the like, and the embodiments of the present application do not make a specific limitation in this regard.

[0115] As shown in Figure 5 The embodiments of the present application provide a speech processing device of a mobile device, which includes a processor 111, a communication interface 112, a memory 113, and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 complete mutual communication through the communication bus 114; the memory 113 is configured to store a computer program; and the processor 111 is configured to execute the program stored in the memory 113, to implement the steps of the speech processing method of the mobile device provided in any of the preceding method embodiments. For example, the steps of the speech processing method of the mobile device can include the following steps: obtaining user speech information collected by the mobile device; if the user speech information contains key speech information corresponding to a preset emotion label, determining the user speech information as target detection speech information; and performing speech processing according to the target speech detection information, to obtain a speech processing result corresponding to the user speech information.

[0116] The embodiments of the present application further provide a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the speech processing method of the mobile device provided in any of the preceding method embodiments.

[0117] It has to be noted that, in the present document, relational terms are intended only to convey a possible relationship between elements or

[0118] The above description is merely that of the specific embodiments of the application and therefore is not intended to limit the application. Many modifications are possible in the specific implementation of the application. The general principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Thus, the present application is not to be limited to the embodiments described herein but is to be accorded the full scope of the claims, the full scope being defined in the language of the claims.

Claims

1. A voice processing method of a mobile device, the method comprising: The method comprises the following steps: acquiring user voice information collected by a mobile device; if the user voice information contains key voice information corresponding to a preset emotion label, determining the user voice information as target voice detection information, wherein the preset emotion label contains one or more corresponding keywords; performing voice processing according to the target voice detection information to obtain a voice processing result corresponding to the user voice information; wherein the voice processing result generation module comprises: a voice processing system wake-up sub-module configured to wake up a voice processing system of the mobile device for the target voice detection information; determining a voice processing result corresponding to the target voice detection information through the voice processing system; the voice processing system comprises: acquiring server address information through the voice processing system; generating a voice processing request according to the server address information and the target voice detection information, sending the voice processing request to a server to trigger the server to obtain monitoring image information corresponding to the target voice detection information based on a camera IP address set in the server, and perform emotion analysis processing based on a behavior action of a target person in the monitoring image information and the target voice detection information to obtain a server voice processing result; 2. The method of claim 1, wherein, the mobile device generates a mobile device voice processing result based on the target voice detection information; comparing the mobile device voice processing result and the server voice processing result to generate the voice processing result. after acquiring the user voice information collected by the mobile device, the method further comprises the following steps: performing text conversion on the user voice information to obtain text conversion information; determining whether the text conversion information contains a target keyword corresponding to the preset emotion label; 3. The method of claim 2, wherein, if the text conversion information contains the target keyword, determining that the user voice information contains key voice information corresponding to a preset emotion label; if the text conversion information does not contain the target keyword, deleting the user voice information. the method further comprises the following steps:

4. A speech processing apparatus of a mobile device, characterized by generating an alarm prompt information corresponding to the target keyword based on the voice processing result; outputting the alarm prompt information. The method comprises the following steps: a user voice acquisition module configured to acquire user voice information collected by a mobile device; a target detection voice information determination module configured to determine the user voice information as target voice detection information if the user voice information contains key voice information corresponding to a preset emotion label, wherein the preset emotion label contains one or more corresponding keywords; a voice processing result generation module configured to perform voice processing according to the target voice detection information to obtain a voice processing result corresponding to the user voice information; the voice processing result generation module comprises: a voice processing system wake-up sub-module configured to wake up a voice processing system of the mobile device for the target voice detection information; The voice processing result determination submodule is configured to determine a voice processing result corresponding to the target voice detection information by the voice processing system, including: acquiring server address information by the voice processing system; generating a voice processing request according to the server address information and the target voice detection information, and sending the voice processing request to a server to trigger the server to acquire monitoring image information corresponding to the target voice detection information based on a camera IP address set in the server, and perform emotion analysis processing based on a behavior action of a target person in the monitoring image information and the target voice detection information to obtain a server voice processing result; The mobile device generates a mobile device voice processing result based on the target voice detection information; The mobile device voice processing result and the server voice processing result are compared to generate the voice processing result.

5. The apparatus of claim 4, wherein, After the user voice information collected by the mobile device is acquired, the method further includes: A voice text conversion module is configured to perform text conversion based on the user voice information to obtain text conversion information; A target keyword judgment module is configured to determine whether the text conversion information contains a target keyword corresponding to the preset emotion label; A key voice information determination module is configured to determine that the user voice information contains key voice information corresponding to a preset emotion label when the text conversion information contains the target keyword; A voice information deletion module is configured to delete the user voice information when the text conversion information does not contain the target keyword.

6. A speech processing device of a mobile device, characterized by The mobile device includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; The memory is configured to store a computer program; The processor is configured to execute the program stored in the memory to implement the steps of the voice processing method of the mobile device according to any one of claims 1-3.

7. A computer-readable storage medium having stored thereon a computer program, characterized in that The computer program is executed by the processor to implement the steps of the voice processing method of the mobile device according to any one of claims 1-3.

Citation Information

Patent Citations

  • Method for obtaining pictures via remote control and server

    CN105120159A

  • Emotional prompting method, device, mobile terminal, and storage medium

    CN109040471A

  • Emotion recognition method and device, terminal device, server and storage medium

    CN112116925A