Information Processing Method, Apparatus, Device, and Computer-Readable Storage Medium

By identifying and caching high-frequency inquiries, the method reduces server processing load and response times, improving user experience in voice interactions.

CN114927134BActive Publication Date: 2025-07-15CHINA MOBILEHANGZHOUINFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210639832.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-06
Publication Date
2025-07-15
Estimated Expiration
2042-06-06

AI Technical Summary

Technical Problem

During the existing voice interaction process, the platform server performs voice synthesis processing on all inquiry requests, resulting in an increase in burden and a longer time to feedback and reply to voice, affecting the user experience.

Method used

By obtaining user voice information within the preset time, preprocessing is performed to determine high-frequency information, natural language generation process is performed to determine reply text information, and then buffering reply voice information after voice synthesis processing is performed, reducing the voice synthesis process of subsequent repeated requests.

Benefits of technology

It reduces the burden on the platform server, increases the speed of user terminals to obtain reply voice information, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114927134B_ABST
    Figure CN114927134B_ABST
Patent Text Reader

Abstract

The present invention discloses an information processing method, apparatus, device and computer-readable storage medium. Among them, the method includes: obtaining all first user voice information within a preset time, preprocessing the first user voice information to determine high-frequency information, performing natural language generation processing on the high-frequency information to determine first reply text information, performing speech synthesis processing on the first reply text information to determine first reply voice information and caching the first reply voice information. The present invention can, based on the high-frequency information in the first user voice information, cache the first reply voice information corresponding to the high-frequency information. When subsequently receiving user voice information matching the high-frequency information, the corresponding reply voice information can be determined through the cached first reply voice information, reducing the speech synthesis process, thereby reducing the burden on the platform server, enabling the user terminal to quickly obtain the reply voice information, and improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information processing, and in particular, to an information processing method, apparatus, device, and computer-readable storage medium. Background Art

[0002] In the existing voice interaction process between a user terminal and a platform server, the platform server obtains an inquiry request sent by the user terminal. After being processed by the platform server, a reply voice message is fed back to the user terminal, and finally the user terminal receives the reply voice message and broadcasts it.

[0003] In actual application scenarios, among all the inquiry requests initiated by user terminals, a considerable number of them are the same. No matter how many times the same inquiry request is initiated by the user terminal, the platform server performs voice synthesis processing on this inquiry request. The process of repeated voice synthesis processing undoubtedly greatly increases the burden on the platform server, makes the processing time of the inquiry request by the platform server longer, thereby extending the time for the user terminal to broadcast the reply voice message, resulting in a poor user experience. Summary of the Invention

[0004] The main objective of the present invention is to provide an information processing method, device, and computer-readable storage medium, aiming to solve the technical problems that in the existing voice interaction process, the platform server performs voice synthesis processing on all inquiry requests, resulting in an increased burden on the platform server and a longer time for feedback of the reply voice.

[0005] To achieve the above objective, the present invention provides an information processing method applied to a platform server. The information processing method includes the following steps:

[0006] Obtain all first user voice messages within a preset time, where the first user voice messages are sent from a user terminal to the platform server;

[0007] Preprocess the first user voice messages to determine high-frequency information;

[0008] Perform natural language generation processing on the high-frequency information to determine first reply text information;

[0009] Perform voice synthesis processing on the first reply text information to determine first reply voice messages and cache the first reply voice messages.

[0010] Further, the step of preprocessing the first user voice messages to determine high-frequency information includes:

[0011] Perform speech recognition processing on the first user voice messages to determine first user text information;

[0012] Perform natural language understanding processing on the first user text information to determine the first domain information, first intent information, and first key slot information corresponding to each piece of first user text information;

[0013] Based on the first domain information, first intent information, and first key slot information corresponding to each piece of first user text information, determine the high-frequency information.

[0014] Further, the step of determining the high-frequency information based on the first domain information, first intent information, and first key slot information corresponding to each piece of first user text information includes:

[0015] Determine the first user text information with the same first domain information, first intent information, and first key slot information as the same type of text information, and record the quantity of the same type of text information;

[0016] If the quantity reaches the first preset quantity, determine the same type of text information as the high-frequency information; or, sort the same type of text information in descending order based on the quantity to obtain a sorting result, and use the first second preset quantity of the same type of text information in the sorting result as the high-frequency information.

[0017] Further, the information processing method further includes:

[0018] Obtain the second user voice information sent by the user terminal;

[0019] Perform speech recognition processing on the second user voice information to determine the second user text information;

[0020] Perform natural language understanding processing on the second user text information to determine the second domain information, second intent information, and second key slot information corresponding to the second user text information;

[0021] Based on the second domain information, second intent information, and second key slot information, determine whether there is a target high-frequency information corresponding to the second user text information in the high-frequency information;

[0022] If there is the target high-frequency information, obtain the second reply voice information corresponding to the target high-frequency information from the cached first reply voice information, and send the second reply voice information to the user terminal.

[0023] Further, the step of determining whether there is a target high-frequency information corresponding to the second user text information in the high-frequency information based on the second domain information, second intent information, and second key slot information includes:

[0024] If there is target domain information in the first domain information of the high-frequency information that matches the second domain information, determine whether there is target intention information in the first intention information of the high-frequency information corresponding to the target domain information that matches the second intention information;

[0025] If there is target intention information in the first intention information of the high-frequency information corresponding to the target domain information that matches the second intention information, determine whether there is target key slot information in the first key slot information of the high-frequency information corresponding to the target intention information that matches the second key slot information. Among them, if there is target key slot information in the first key slot information of the high-frequency information corresponding to the target intention information that matches the second key slot information, it is determined that there is the target high-frequency information.

[0026] Further, after the step of determining whether there is target high-frequency information corresponding to the second user text information based on the second domain information, second intention information, and second key slot information, the information processing method further includes:

[0027] If there is no such target high-frequency information, perform natural language generation processing on the second user text information to obtain a user reply message, and send the user reply message to the user terminal, where the user terminal receives the user reply message;

[0028] If a voice synthesis request corresponding to the user reply message sent by the user terminal is received, perform voice synthesis processing on the user reply message to obtain a third reply voice message.

[0029] Further, the information processing method further includes:

[0030] Obtain user operation information sent by the user terminal;

[0031] If the user operation information is activation operation information, send a first fixed reply voice to the user terminal, where the user terminal stores the first fixed reply voice;

[0032] If the user operation information is tone modification information, send a second fixed reply voice to the user terminal, where the user terminal stores the second fixed reply voice.

[0033] In addition, to achieve the above object, the present invention further provides an information processing device, and the information processing device includes:

[0034] An acquisition module, configured to acquire all first user voice messages within a preset time, where the first user voice messages are sent from a user terminal to a platform server;

[0035] The first processing module is used to preprocess the first user voice information to determine high-frequency information;

[0036] The second processing module is used to perform natural language generation processing on the high-frequency information to determine the first reply text information;

[0037] The cache module is used to perform speech synthesis processing on the first reply text information to determine the first reply voice information and cache the first reply voice information.

[0038] In addition, to achieve the above object, the present invention also provides an information processing device, which includes: a memory, a processor, and an information processing program stored on the memory and executable on the processor. When the information processing program is executed by the processor, the steps of the foregoing information processing method are implemented.

[0039] In addition, to achieve the above object, the present invention also provides a computer-readable storage medium, on which an information processing program is stored. When the information processing program is executed by a processor, the steps of the foregoing information processing method are implemented.

[0040] The present invention obtains all the first user voice information within a preset time. The first user voice information is sent from a user terminal to a platform server. Then, the first user voice information is preprocessed to determine high-frequency information. Then, natural language generation processing is performed on the high-frequency information to determine the first reply text information. Finally, speech synthesis processing is performed on the first reply text information to determine the first reply voice information and cache the first reply voice information. It can, according to the high-frequency information in the first user voice information, cache the first reply voice information corresponding to the high-frequency information. When a user voice information matching the high-frequency information is received subsequently, the corresponding reply voice information can be determined through the cached first reply voice information, reducing the speech synthesis process, thereby reducing the burden on the platform server, enabling the user terminal to quickly obtain the reply voice information, and improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is a schematic structural diagram of an information processing device in a hardware operating environment related to an embodiment of the present invention;

[0042] Figure 2 is a schematic flowchart of a first embodiment of the information processing method of the present invention;

[0043] Figure 3 is a schematic diagram of functional modules of an embodiment of the information processing device of the present invention.

[0044] The implementation, functional features, and advantages of the present invention will be further described in conjunction with embodiments with reference to the accompanying drawings. Detailed implementation manners

[0045] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0046] As Figure 1 shown, Figure 1 is a schematic structural diagram of an information processing device in the hardware operating environment involved in the embodiment solution of the present invention.

[0047] The information processing device in the embodiment of the present invention may be a PC. As Figure 1 shown, the information processing device may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. Among them, the communication bus 1002 is used to implement connection communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. The memory 1005 may optionally be a storage device independent of the aforementioned processor 1001.

[0048] Optionally, the information processing device may further include a camera, an RF (Radio Frequency) circuit, sensors, an audio circuit, a WiFi module, and so on. Among them, the sensors such as a light sensor, a motion sensor, and other sensors. Of course, the information processing device may also be configured with other sensors such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, etc., which will not be elaborated here.

[0049] Those skilled in the art can understand that Figure 1 the terminal structure shown in

[0050] As Figure 1 shown, the memory 1005, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an information processing program.

[0051] In Figure 1In the information processing device shown, the network interface 1004 is mainly used to connect to the background server and communicate data with the background server; the user interface 1003 is mainly used to connect to the client (user side) and communicate data with the client; and the processor 1001 can be used to call the information processing program stored in the memory 1005.

[0052] In this embodiment, the information processing device includes: a memory 1005, a processor 1001, and an information processing program stored on the memory 1005 and executable on the processor 1001. When the processor 1001 calls the information processing program stored in the memory 1005, it executes the steps of the information processing method in each of the following embodiments.

[0053] The present invention also provides a method. Refer to Figure 2 , Figure 2 which is a schematic flowchart of the first embodiment of the method of the present invention.

[0054] In this embodiment, the information processing method is applied to a platform server and includes the following steps:

[0055] Step S101: Obtain all first user voice messages within a preset time, where the first user voice messages are sent from user terminals to the platform server;

[0056] Step S102: Preprocess the first user voice messages to determine high-frequency information;

[0057] Step S103: Perform natural language generation processing on the high-frequency information to determine first reply text messages;

[0058] Step S104: Perform speech synthesis processing on the first reply text messages to determine first reply voice messages and cache the first reply voice messages.

[0059] In this embodiment, each user terminal monitors the user's voice in real time and collects first user voice messages. Specifically, the user terminal can obtain first user voice messages in real time and send them to the platform server when it detects a collection start instruction. Then, the platform server obtains all the first user voice messages sent from user terminals within a preset time in real time. The preset time can be set manually. For example, the preset time can be set to one day or one week.

[0060] Next, preprocess the first user voice information to determine high-frequency information. In one implementation, the platform server performs natural language understanding processing on the first user text information, which can be carried out by the natural language processing service NLP (Natural Language Processing) in the platform server. A core function of NLP is the natural language understanding service NLU (Natural Language Understanding). When NLU processes the first user text information, it returns text parameters corresponding to each user text information in the first user text information, such as: the first domain information, the first intent information, the first key slot information, etc. The first user text information with the same first domain information, first intent information, and first key slot information is the same type of text information, and the number of the same type of text information is recorded. Or in another implementation, the platform server calculates the similarity between each user text information in the first user text information. If the similarity between a user text information in the first user text information and other user text information in the first user text information reaches the artificially preset similarity value, it is determined as the same type of text information, and the number of the same type of text information is recorded. Specifically, the cosine similarity can be used, and the cosine value of the angle between the vector of user A's text information and the vector of user B's text information in the vector space is used as a measure of the difference between the vector of user A's text information and the vector of user B's text information. The closer the cosine value is to 1, the closer the angle is to 0 degrees, and the more similar the two vectors are. Finally, if the quantity reaches the first preset quantity, the same type of text information is determined as high-frequency information. Or, the same type of text information is sorted in descending order of quantity to obtain a sorting result, and the first second preset quantity of the same type of text information in the sorting result is used as high-frequency information.

[0061] Perform natural language generation processing on the high-frequency information to determine the first reply text information. Specifically, through a core function of NLP, the natural language generation service NLG (Natural Language generation), perform NLG processing on the high-frequency information, and the first reply text information that is specific and true for the high-frequency information can be obtained. For example, if the high-frequency information is "What's the weather like today", then after natural language generation processing, the first reply text information for the high-frequency information will be generated: "It will be cloudy during the day today with scattered showers, cloudy tonight, the highest temperature is 28°C, the lowest temperature is 20°C, the humidity is 55 - 95%, and the northerly wind is 3 to 4 levels turning to 2 to 3 levels."

[0062] Finally, perform speech synthesis processing on the first reply text message, determine the first reply speech message, and cache the first reply speech message. Specifically, the first user text message can be processed by the Text-to-Speech (TTS) service in the platform server to obtain the first reply speech message. There are mainly two types of TTS: "concatenation method" and "parametric method". The concatenation method selects the required basic units from a large number of pre-recorded voices and concatenates them into the first user text message. The parametric method refers to generating the speech parameters at each moment according to the statistical model, and then obtaining the first user text message according to the parameters. Then, the obtained first reply speech message is cached in the TTS speech synthesis service, or the obtained first reply speech message is cached in the platform server.

[0063] It should be noted that after performing speech synthesis processing on the first reply text message, determining the first reply speech message, and caching the first reply speech message, in one implementation, the high-frequency information is associated with the first reply speech message, or in another implementation, the tag high-frequency information and the first reply speech message corresponding to the high-frequency information are used, so that when a user speech message matching the high-frequency information is received subsequently, the corresponding reply speech message can be determined through the cached first reply speech message.

[0064] The information processing method proposed in this embodiment can, by obtaining all the first user speech messages within a preset time (where the first user speech messages are sent from the user terminal to the platform server), then performing preprocessing on the first user speech messages to determine the high-frequency information, then performing natural language generation processing on the high-frequency information to determine the first reply text message, and finally performing speech synthesis processing on the first reply text message to determine the first reply speech message and cache the first reply speech message, be able to, based on the high-frequency information in the first user speech message, cache the first reply speech message corresponding to the high-frequency information, and when a user speech message matching the high-frequency information is received subsequently, determine the corresponding reply speech message through the cached first reply speech message, reducing the speech synthesis process, thereby reducing the burden on the platform server, enabling the user terminal to quickly obtain the reply speech message, and improving the user experience.

[0065] Based on the first embodiment, a second embodiment of the information processing method of the present invention is proposed. In this embodiment, step S102 includes:

[0066] Step S201, perform speech recognition processing on the first user speech message to determine the first user text message;

[0067] Step S202: Perform natural language understanding processing on the first user text information to determine the first domain information, first intent information, and first key slot information corresponding to each piece of first user text information.

[0068] Step S203: Determine the high-frequency information based on the first domain information, first intent information, and first key slot information corresponding to each piece of first user text information.

[0069] Specifically, the platform server performs speech recognition processing on the first user voice information to determine the first user text information. The speech recognition service ASR (Automatic Speech Recognition) in the platform server can be used to perform speech recognition processing on the first user voice information to obtain the first user text information. The first user text information includes the text information corresponding to each piece of voice information in the first user voice information.

[0070] Next, perform natural language understanding processing on the first user text information, that is, perform semantic understanding, intent recognition, key slot determination, etc. on the first user text information. It should be noted that the natural language processing service NLP (Natural Language Processing) in the platform server can be used to perform natural language understanding processing. A core function of NLP is NLU (Natural Language Understanding). When NLU processes the first user text information, it returns parameter information such as the first domain information, first intent information, and first key slot information. For example: The user terminal detects that the first user voice information of the user is "What's the weather like today", and sends the first user voice information to the platform server in real time. Then the platform server performs speech recognition processing on the first user voice information of "What's the weather like today" to obtain the first user text information. Then the first user text information is processed by NLU to obtain the first domain information of the first user text information as "weather", the first intent information of the first user text information as "inquiry about weather conditions", and the first key slot information of the first user text information as "today". It should be noted that the key slot can be understood as the content that has been clearly defined, such as the key slot being time, place, or person.

[0071] The platform server regards first user text information with the same first domain information, first intent information, and first key slot information as the same type of text information, and records the quantity of the same type of text information. If the quantity reaches the first preset quantity, it determines that the same type of text information is high-frequency information. Or, it sorts the same type of text information according to the order from large to small based on the quantity, obtains the sorting result, and takes the first second preset quantity of the same type of text information in the sorting result as the high-frequency information.

[0072] The information processing method proposed in this embodiment determines the first user text information by performing speech recognition processing on the first user voice information. Then, it performs natural language understanding processing on the first user text information to determine the first domain information, first intent information, and first key slot information corresponding to each first user text information. Then, based on the first domain information, first intent information, and first key slot information corresponding to each first user text information, it determines the high-frequency information, can accurately obtain the high-frequency information according to the first domain information, first intent information, and first key slot information, caches the first reply voice information corresponding to the high-frequency information subsequently. When receiving user voice information matching the high-frequency information, it can determine the corresponding reply voice information through the cached first reply voice information, reducing the process of speech synthesis, thereby reducing the burden on the platform server, enabling the user terminal to quickly obtain the reply voice information, and improving the user experience.

[0073] Based on the second embodiment, the third embodiment of the information processing method of the present invention is proposed. In this embodiment, step S202 includes:

[0074] Step S301, determine first user text information with the same first domain information, first intent information, and first key slot information as the same type of text information, and record the quantity of the same type of text information;

[0075] Step S302, if the quantity reaches the first preset quantity, determine that the same type of text information is high-frequency information; or, sort the same type of text information according to the order from large to small based on the quantity, obtain the sorting result, and take the first second preset quantity of the same type of text information in the sorting result as the high-frequency information.

[0076] The platform server detects the first domain information, first intent information, and first key slot information corresponding to each first user text information. If there is user text information with the same first domain information, first intent information, and first key slot information among the first user text information, it determines that these user text information with the same first domain information, first intent information, and first key slot information are the same type of text information, and calculates the quantity of the same type of text information.

[0077] Then, when the number of text messages of the same type reaches the first preset number, the text messages of the same type are determined to be high-frequency information. It should be noted that if the same user terminal repeatedly sends the first user text messages with the same first field information, first intent information, and first key slot information within the same preset time period, the number of the same type of text messages of the first user text messages with the same first field information, first intent information, and first key slot information will not increase, or the increase in the number will be reduced according to other artificially set rules. For example, the same user terminal sends multiple first user text messages within a two-hour artificially preset time period. After the platform server voice recognition processing and natural speech understanding, the first field information, first intent information, and first key slot information of the multiple first user text messages are the same, then the number of recorded text messages of the same type is still 1. Among them, the first preset number can be 5, 10, etc.

[0078] Alternatively, the text messages of the same type are sorted in descending order of quantity to obtain a sorting result, and the text messages of the same type with the second preset number of the top ranked in the sorting result are used as the high-frequency information. Then, the high-frequency information is processed by natural language generation to determine the first reply text message, and finally, the first reply text message is processed by speech synthesis to determine the first reply voice message and cache the first reply voice message. For example, if there are 10 text messages of the same type and the second preset number is 3, the text messages of the same type are sorted from large to small according to the number of the text messages of the same type, and the top three text messages of the same type are used as the high-frequency information.

[0079] The information processing method proposed in this embodiment determines that the first user text information with the same first field information, first intent information and first key slot information is the same type of text information, and records the number of the same type of text information. Then, if the number reaches a first preset number, the same type of text information is determined to be high-frequency information; or, the same type of text information is sorted in descending order based on the number to obtain a sorting result, and the second preset number of the same type of text information ranked at the top in the sorting result is used as the high-frequency information. The high-frequency information can be accurately obtained according to the first field information, the first intent information and the first key slot information, and the first reply voice information corresponding to the high-frequency information is subsequently cached. When the user voice information matching the high-frequency information is received, the corresponding reply voice information can be determined through the cached first reply voice information, thereby reducing the voice synthesis process, thereby reducing the burden on the platform server, allowing the user terminal to quickly obtain the reply voice information, and improving the user experience.

[0080] Based on the first embodiment, a fourth embodiment of the information processing method of the present invention is proposed. In this embodiment, the information processing method includes:

[0081] Step 401, obtaining second user voice information sent by the user terminal;

[0082] Step 402, performing speech recognition processing on the second user voice information to determine second user text information;

[0083] Step 403, performing natural language understanding processing on the second user text information to determine second domain information, second intention information, and second key slot information corresponding to the second user text information;

[0084] Step 404, based on the second domain information, second intention information, and second key slot information, determining whether there is target high-frequency information corresponding to the second user text information in the high-frequency information;

[0085] Step 405, if there is the target high-frequency information, obtaining a second reply voice message corresponding to the target high-frequency information from the cached first reply voice messages, and sending the second reply voice message to the user terminal.

[0086] In this embodiment, the user terminal monitors the user's voice in real time and collects second user voice information. Specifically, the user terminal can obtain the second user voice information in real time and send it to the platform server when detecting a collection start instruction. Then the platform server obtains the second user voice information in real time. Perform speech recognition processing on the second user voice information to obtain second user text information, and then perform natural language understanding processing on the second user text information to determine second domain information, second intention information, and second key slot information corresponding to the second user text information. According to the second domain information, second intention information, and second key slot information, determine whether there is target high-frequency information corresponding to the second user text information in the high-frequency information.

[0087] Further, in one embodiment, step S404 includes:

[0088] Step a, if there is target domain information in the first domain information of the high-frequency information that matches the second domain information, determining whether there is target intention information in the first intention information of the high-frequency information corresponding to the target domain information that matches the second intention information;

[0089] Step b, if there is a target intention information that matches the second intention information among the first intention information of the high-frequency information corresponding to the target domain information, determine whether there is a target key slot information that matches the second key slot information among the first key slot information of the high-frequency information corresponding to the target intention information. Among them, if there is a target key slot information that matches the second key slot information among the first key slot information of the high-frequency information corresponding to the target intention information, it is determined that there is the target high-frequency information.

[0090] In this embodiment, the platform server checks whether there is a target domain information that matches the second domain information among the first domain information of the high-frequency information in the cache. If there is a target domain information, continue to check whether there is a target intention information that matches the second intention information among the first intention information of the high-frequency information corresponding to the target domain information in the cache. If there is a target intention information, continue to check whether there is a target key slot information that matches the second key slot information among the first key slot information of the high-frequency information corresponding to the target intention information in the cache. If there is a target key slot information, it is determined that the high-frequency information corresponding to the target key slot information matches the second user text information, that is, the high-frequency information corresponding to the target key slot information is the target high-frequency information. Furthermore, it can be accurately determined whether the second user text information is high-frequency information. If so, the corresponding reply voice information can be determined through the cached first reply voice information, reducing the voice synthesis process, thereby reducing the burden on the platform server, enabling the user terminal to quickly obtain the reply voice information, and improving the user experience.

[0091] For example, if the second user voice information is "What's the weather like today", the platform server performs speech recognition processing on the second user voice information to obtain the second user text information. Then, the second user text information is subjected to NLU processing to obtain the second domain information of the second user text information as "weather", the second intention information of the second user text information as "inquiry about weather conditions", and the second key slot information of the second user text information as "today". Then, the platform server checks whether there is a target domain information that matches "weather" among the first domain information of the high-frequency information in the cache. If there is a target domain information, continue to check whether there is a target intention information that matches "inquiry about weather conditions" among the first intention information of the high-frequency information corresponding to the multiple target domain information in the cache. If there is a target intention information, continue to check whether there is a target key slot information that matches "today" among the first key slot information of the high-frequency information corresponding to the multiple target intention information in the cache. If there is a target key slot information, it is determined that the high-frequency information corresponding to the target key slot information matches the second user text information, that is, the first domain information of the target high-frequency information is "weather", the first intention information is "inquiry about weather conditions", and the first key slot information is "today".

[0092] Finally, the platform server obtains the second reply voice message corresponding to the target high-frequency information from the cache, and sends the second reply voice message to the user terminal, where the user terminal obtains the second reply voice message and plays the second reply voice message.

[0093] The information processing method proposed in this embodiment obtains the second user voice message sent by the user terminal, then performs speech recognition processing on the second user voice message to determine the second user text message, and then performs natural language understanding processing on the second user text message to determine the second domain information, the second intent information, and the second key slot information corresponding to the second user text message. Then, based on the second domain information, the second intent information, and the second key slot information, it determines whether there is a target high-frequency information corresponding to the second user text message in the high-frequency information. Finally, if there is the target high-frequency information, it obtains the second reply voice message corresponding to the target high-frequency information from the cached first reply voice message, and sends the second reply voice message to the user terminal. It can accurately obtain the target high-frequency information from the cache according to the second domain information, the second intent information, and the second key slot information, and then determine the reply voice message corresponding to the target high-frequency information, reducing the speech synthesis process, thereby reducing the burden on the platform server, enabling the user terminal to quickly obtain the reply voice message, and improving the user experience.

[0094] Based on the fourth embodiment, a fifth embodiment of the information processing method of the present invention is proposed. In this embodiment, after step 404, it includes:

[0095] Step 501, if there is no such target high-frequency information, perform natural language generation processing on the second user text message to obtain a user reply message, and send the user reply message to the user terminal, where the user terminal receives the user reply message;

[0096] Step 502, if a voice synthesis request corresponding to the user reply message sent by the user terminal is received, perform voice synthesis processing on the user reply message to obtain a third reply voice message.

[0097] In this embodiment, the platform server checks whether there is a target domain information in the first domain information of the high-frequency information in the cache that matches the second domain information. If there is no target domain information in the first domain information of the high-frequency information that matches the second domain information, perform natural language generation processing on the second user text message to obtain a user reply message.

[0098] Alternatively, the platform server checks whether there is a target domain information in the first domain information of the high-frequency information in the cache that matches the second domain information. If there is target domain information, it continues to check whether there is a target intent information in the first intent information of the high-frequency information corresponding to the target domain information in the cache that matches the second intent information. If there is no target intent information, natural language generation processing is performed on the second user text information to obtain a user reply information, and natural language generation processing is performed on the second user text information to obtain a user reply information.

[0099] Alternatively, the platform server checks whether there is a target domain information in the first domain information of the high-frequency information in the cache that matches the second domain information. If there is target domain information, it continues to check whether there is a target intent information in the first intent information of the high-frequency information corresponding to the target domain information in the cache that matches the second intent information. If there is target intent information, it continues to check whether there is a target key slot information in the first key slot information of the high-frequency information corresponding to the target intent information in the cache that matches the second key slot information. If there is no target key slot information, natural language generation processing is performed on the second user text information to obtain a user reply information.

[0100] Next, the platform server sends the user reply information to the user terminal. The user terminal checks whether there is a fixed voice reply corresponding to the user reply information in its local cache. If there is, the user terminal plays the fixed voice reply corresponding to the user reply information. If there is no fixed voice reply corresponding to the user reply information in the user terminal, the user terminal sends a voice synthesis request corresponding to the user reply information to the platform server. The TTS server in the platform server performs voice synthesis processing on the user reply information to obtain a third reply voice message.

[0101] Finally, the platform server sends the third reply voice message to the user terminal. Then the user terminal obtains the third reply voice message and plays the third reply voice message.

[0102] For the information processing method proposed in this embodiment, if there is no such target high-frequency information, natural language generation processing is performed on the second user text information to obtain a user reply information, and the user reply information is sent to the user terminal. Among them, the user terminal receives the user reply information. Then, if a voice synthesis request corresponding to the user reply information sent by the user terminal is received, voice synthesis processing is performed on the user reply information to obtain a third reply voice message. This can enable the user terminal to obtain a reply voice message through voice synthesis when there is no target high-frequency information and fixed voice reply, improving the user experience.

[0103] Based on the above various embodiments, a sixth embodiment of the information processing method of the present invention is proposed. In this embodiment, the information processing method further includes:

[0104] Step 601, obtaining user operation information sent by the user terminal;

[0105] Step 602, if the user operation information is activation operation information, sending a first fixed reply voice to the user terminal, where the user terminal stores the first fixed reply voice;

[0106] Step 603, if the user operation information is timbre modification information, sending a second fixed reply voice to the user terminal, where the user terminal stores the second fixed reply voice.

[0107] In this embodiment, the user operation information sent by the user terminal is obtained. The user operation information can be activation operation information or timbre modification information. When the user operation information is activation operation information, a first fixed reply voice is sent to the user terminal, where the user terminal stores the first fixed reply voice. For example, the first fixed reply voice can be "Okay", "Hello, master", etc. When the user operation information is timbre modification information, a second fixed reply voice is sent to the user terminal, where the user terminal stores the second fixed reply voice. The second fixed reply voice can be a change in the timbre of the first fixed reply voice. For example, if the first fixed reply voice is in a male timbre, changing the male timbre of the first fixed reply voice to a female timbre can be used as the second fixed reply voice.

[0108] The information processing method proposed in this embodiment obtains the user operation information sent by the user terminal. Then, if the user operation information is activation operation information, a first fixed reply voice is sent to the user terminal, where the user terminal stores the first fixed reply voice. Then, if the user operation information is timbre modification information, a second fixed reply voice is sent to the user terminal, where the user terminal stores the second fixed reply voice. Through the activation operation information or timbre modification information of the user operation information, fixed reply voices are stored in the user terminal and the timbres of the fixed reply voices are switched, improving the user experience.

[0109] The present invention also provides an information processing device, which is applied to a platform server. Referring to Figure 3 , the information processing device includes:

[0110] An obtaining module 10, configured to obtain all first user voice information within a preset time, where the first user voice information is sent from the user terminal to the platform server;

[0111] A first processing module 20, configured to pre-process the first user voice information to determine high-frequency information;

[0112] The second processing module 30 is used to perform natural language generation processing on the high-frequency information to determine the first reply text information;

[0113] The cache module 40 is used to perform speech synthesis processing on the first reply text information, determine the first reply voice information and cache the first reply voice information.

[0114] Furthermore, the first processing module 20 is also used for:

[0115] Performing voice recognition processing on the first user voice information to determine the first user text information;

[0116] Performing natural language understanding processing on the first user text information to determine first field information, first intent information, and first key slot information corresponding to each first user text information;

[0117] Based on the first leading corresponding to each first user text information

[0118] Domain information, first intent information, and first key slot information determine high-frequency information.

[0119] Furthermore, the first processing module 20 is also used for:

[0120] Determine that the first user text messages having the same first field information, first intent information, and first key slot information are the same type of text messages, and record the number of the same type of text messages;

[0121] If the number reaches a first preset number, the same type of text information is determined to be high-frequency information; or, the same type of text information is sorted in descending order based on the number to obtain a sorting result, and the second preset number of the same type of text information ranked at the top in the sorting result is used as the high-frequency information.

[0122] Furthermore, the information processing device is also used for:

[0123] Acquire the second user voice information sent by the user terminal;

[0124] Performing voice recognition processing on the second user's voice information to determine the second user's text information;

[0125] Performing natural language understanding processing on the second user text information to determine second domain information, second intent information, and second key slot information corresponding to the second user text information;

[0126] Based on the second domain information, second intent information, and second key slot information, determine whether there is target high-frequency information corresponding to the second user text information in the high-frequency information;

[0127] If there is the target high-frequency information, obtain the second reply voice message corresponding to the target high-frequency information from the cached first reply voice messages, and send the second reply voice message to the user terminal.

[0128] Further, the information processing device is further configured to:

[0129] If there is target domain information in the first domain information of the high-frequency information that matches the second domain information, determine whether there is target intent information in the first intent information of the high-frequency information corresponding to the target domain information that matches the second intent information;

[0130] If there is target intent information in the first intent information of the high-frequency information corresponding to the target domain information that matches the second intent information, determine whether there is target key slot information in the first key slot information of the high-frequency information corresponding to the target intent information that matches the second key slot information, where if there is target key slot information in the first key slot information of the high-frequency information corresponding to the target intent information that matches the second key slot information, determine that there is the target high-frequency information.

[0131] Further, the information processing device is further configured to:

[0132] If there is no such target high-frequency information, perform natural language generation processing on the second user text information to obtain a user reply message, and send the user reply message to the user terminal, where the user terminal receives the user reply message;

[0133] If a voice synthesis request corresponding to the user reply message sent by the user terminal is received, perform voice synthesis processing on the user reply message to obtain a third reply voice message.

[0134] Further, the information processing device is further configured to:

[0135] Obtain user operation information sent by the user terminal;

[0136] If the user operation information is activation operation information, send a first fixed reply voice to the user terminal, where the user terminal stores the first fixed reply voice;

[0137] If the user operation information is tone color modification information, send a second fixed reply voice to the user terminal, where the user terminal stores the second fixed reply voice.

[0138] The methods executed by the above program units can refer to the various embodiments of the information processing method of the present invention, which will not be elaborated here.

[0139] In addition, an embodiment of the present invention further provides an information processing device, which includes: a memory, a processor, and an information processing program stored on the memory and operable on the processor. When the information processing program is executed by the processor, it implements the steps of the information processing method described above.

[0140] In addition, an embodiment of the present invention further provides a computer-readable storage medium, on which an information processing program is stored. When the information processing program is executed by a processor, it implements the steps of the information processing method described above.

[0141] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or system including the element.

[0142] The serial numbers of the above embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments.

[0143] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing an information processing device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0144] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, are equally included in the patent protection scope of the present invention.

Claims

1. An information processing method, characterized in that, Applied to a platform server, the information processing method further includes: Obtain all first user voice information within a preset time, where the first user voice information is sent from a user terminal to the platform server; Preprocess the first user voice information to determine high-frequency information; Perform natural language generation processing on the high-frequency information to determine first reply text information; Perform speech synthesis processing on the first reply text information to determine first reply voice information and cache the first reply voice information; The step of preprocessing the first user voice information to determine high-frequency information includes: Based on the first domain information, first intent information, and first key slot information corresponding to each first user text information, determine that the first user text information with the same first domain information, first intent information, and first key slot information is the same type of text information, and record the quantity of the same type of text information, where the quantity of the same type of text information repeatedly sent by the same user terminal within the same preset time period does not increase; Determine high-frequency information according to the quantity.

2. The information processing method according to claim 1, wherein Before the step of determining that the first user text information with the same first domain information, first intent information, and first key slot information is the same type of text information and recording the quantity of the same type of text information, it further includes: Perform speech recognition processing on the first user voice information to determine first user text information; Perform natural language understanding processing on the first user text information to determine the first domain information, first intent information, and first key slot information corresponding to each first user text information.

3. The information processing method according to claim 1, wherein The step of determining high-frequency information according to the quantity includes: If the quantity reaches a first preset quantity, determine that the same type of text information is high-frequency information; or, sort the same type of text information in descending order of the quantity to obtain a sorting result, and use the first second preset quantity of the same type of text information in the sorting result as the high-frequency information.

4. The information processing method according to claim 1, wherein The information processing method further includes: Obtain second user voice information sent by the user terminal; Perform speech recognition processing on the second user voice information to determine second user text information; Perform natural language understanding processing on the second user text information to determine the second domain information, second intent information, and second key slot information corresponding to the second user text information; Based on the second domain information, second intent information, and second key slot information, determine whether there is target high-frequency information corresponding to the second user text information in the high-frequency information; If there is the target high-frequency information, obtain the second reply voice information corresponding to the target high-frequency information from the cached first reply voice information and send the second reply voice information to the user terminal.

5. The information processing method according to claim 4, wherein The step of determining whether there is target high-frequency information corresponding to the second user text information in the high-frequency information based on the second domain information, second intent information, and second key slot information includes: If there is target domain information in the first domain information of the high-frequency information that matches the second domain information, determine whether there is target intent information in the first intent information of the high-frequency information corresponding to the target domain information that matches the second intent information; If there is target intent information in the first intent information of the high-frequency information corresponding to the target domain information that matches the second intent information, determine whether there is target key slot information in the first key slot information of the high-frequency information corresponding to the target intent information that matches the second key slot information. Among them, if there is target key slot information in the first key slot information of the high-frequency information corresponding to the target intent information that matches the second key slot information, determine that there is the target high-frequency information.

6. The information processing method according to claim 4, wherein After the step of determining whether there is target high-frequency information corresponding to the second user text information based on the second domain information, the second intent information, and the second key slot information, the information processing method further includes: If there is no such target high-frequency information, perform natural language generation processing on the second user text information to obtain a user reply message, and send the user reply message to the user terminal, where the user terminal receives the user reply message; If a voice synthesis request corresponding to the user reply message sent by the user terminal is received, perform voice synthesis processing on the user reply message to obtain a third reply voice message.

7. The information processing method according to any one of claims 1 to 6, characterized in that, The information processing method further includes: Obtain user operation information sent by the user terminal; If the user operation information is activation operation information, send a first fixed reply voice to the user terminal, where the user terminal stores the first fixed reply voice; If the user operation information is tone modification information, send a second fixed reply voice to the user terminal, where the user terminal stores the second fixed reply voice.

8. An information processing apparatus, characterized in that, The information processing device includes: An acquisition module, configured to acquire all first user voice messages within a preset time, where the first user voice messages are sent from the user terminal to the platform server; A first processing module, configured to perform preprocessing on the first user voice messages to determine high-frequency information; A second processing module, configured to perform natural language generation processing on the high-frequency information to determine a first reply text message; A cache module, configured to perform voice synthesis processing on the first reply text message to determine a first reply voice message and cache the first reply voice message; The first processing module is further configured to: Based on the first domain information, the first intent information, and the first key slot information corresponding to each first user text message, determine that the first user text messages with the same first domain information, first intent information, and first key slot information are the same type of text messages, and record the quantity of the same type of text messages, where the quantity of the same type of text messages repeatedly sent by the same user terminal within the same preset time period does not increase; Determine the high-frequency information according to the quantity.

9. An information processing apparatus, characterized in that, The information processing device includes: a memory, a processor, and an information processing program stored on the memory and executable on the processor. When the information processing program is executed by the processor, the steps of the information processing method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that, An information processing program is stored on the computer-readable storage medium. When the information processing program is executed by a processor, the steps of the information processing method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Data processing method and device, electronic equipment and computer storage medium

    CN113779204A

  • Replay statement determination method and device based on knowledge graph and electronic equipment

    CN114090755A