Method and apparatus for determining echo audio, storage medium, and electronic device
By analyzing the user's voice commands, combining language expression and preference information, and determining and sending personalized reply audio, the problem of the device's inability to respond according to the user's language expression habits and preferences is solved, thereby improving the user experience.
Patent Information
- Application Number
- CN202210284536.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-22
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2042-03-22
AI Technical Summary
After receiving a user's voice command, existing devices are unable to determine the corresponding reply audio based on the user's language expression habits and preferences.
By obtaining the target object's voice commands, analyzing their language expressions and preference information, and using pre-collected sample audio and historical operation data, the corresponding reply audio is determined and sent.
It implements intelligent replies based on the user's language expression habits and preferences, improves the user experience, and solves the problem of the device's inability to respond in a personalized manner.
Smart Images

Figure CN114817514B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of smart home technology, and more specifically, to a method and device for determining reply audio, a storage medium, and an electronic device. Background Art
[0002] With the advent of an intelligent society, more and more smart devices have emerged. A key feature of smart devices is that they can provide different services for different users. For example, existing devices can recognize the user's voiceprint, determine the user's identity, and then provide predetermined services based on the user's identity.
[0003] But people are different. Different people have different speaking accents (Mandarin or dialects), different speaking habits, different ways of expressing emotions (such as speaking calmly or angrily), and other different ways of expressing language, and their corresponding preferences are also different.
[0004] Regarding related technologies, when the device receives a user's voice command, it is unable to determine the corresponding reply audio based on the user's language expression habits and preferences. No effective solution has been proposed yet.
[0005] Therefore, it is necessary to improve the related technology to overcome the above-mentioned defects in the related technology. Summary of the Invention
[0006] Embodiments of the present invention provide a method and device for determining a reply audio, a storage medium, and an electronic device, so as to at least solve the problem that when the device obtains a user voice command, it is unable to determine the corresponding reply audio according to the user's language expression habits and preferences.
[0007] According to one aspect of an embodiment of the present invention, a method for determining a reply audio is provided, comprising: obtaining a voice instruction of the target object; obtaining a predetermined target language expression and target preference information of the target object based on the voice instruction, wherein the target language expression is a language expression determined based on a plurality of sample audios of the target object collected in advance, and the target preference information is preference information determined based on a set of historical operations of the target object on a plurality of devices; determining a first reply audio of the voice instruction based on the target language expression and the target preference information, and sending the first reply audio to an audio playback device.
[0008] According to another aspect of the embodiments of the present application, a device for determining a reply audio is also provided, comprising: a first obtaining module, configured to obtain a voice instruction of a target object; a second obtaining module, configured to obtain a target language expression manner and target preference information of the target object according to the voice instruction, wherein the target language expression manner is determined according to a plurality of sample audios of the target object collected in advance, and the target preference information is determined according to a group of historical operations of the target object on a plurality of devices; and a determining module, configured to determine a first reply audio of the voice instruction according to the target language expression manner and the target preference information, and send the first reply audio to an audio playing device.
[0009] According to still another aspect of the embodiments of the present application, a computer readable storage medium is also provided, wherein the computer readable storage medium stores a computer program, and the computer program is configured to execute the above-mentioned method for determining a reply audio when running.
[0010] According to still another aspect of the embodiments of the present application, an electronic device is also provided, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein the processor executes the above-mentioned method for determining a reply audio through the computer program.
[0011] According to the present application, in the case of obtaining a voice instruction of a target object, a target language expression manner and target preference information of the target object are obtained in advance, and then a first reply audio of the voice instruction is determined according to the target language expression manner and the target preference information, and the first reply audio is sent to an audio playing device. By using the above technical solution, the user can be intelligently replied according to the language expression manner and preference information, and the user can be better served, and the problem that the device cannot determine the corresponding reply audio according to the language expression habit and preference of the user in the case of obtaining the voice instruction of the user is solved. BRIEF DESCRIPTION OF DRAWINGS
[0012] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate embodiments consistent with the present application and, together with the description, further serve to explain the principles behind the application.
[0013] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced as follows, and obviously, other drawings can also be obtained by those skilled in the art without any creative effort on the basis of these drawings.
[0014] Figure 1 is a hardware environment schematic diagram of a method for determining a reply audio according to an embodiment of the present application;
[0015] Figure 2 is a flowchart of a method for determining a reply audio according to an embodiment of the present invention (I);
[0016] Figure 3 is a flowchart (II) of a method for determining a reply audio according to an embodiment of the present invention;
[0017] Figure 4 1 is a structural block diagram of a device for determining a reply audio according to an embodiment of the present invention (I);
[0018] Figure 5 2 is a structural block diagram of a device for determining a reply audio according to an embodiment of the present invention. DETAILED DESCRIPTION
[0019] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0020] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0021] According to one aspect of the embodiment of the present application, a method for determining a reply audio is provided. The method for determining a reply audio is widely used in smart home (Smart Home), smart home, smart home device ecology, smart residential (Intelligence House) ecology and other whole-house intelligent digital control application scenarios. Optionally, in this embodiment, the above-mentioned method for determining a reply audio can be applied to Figure 1 In the hardware environment shown in FIG. 1 , which is composed of a terminal device 102 and a server 104. Figure 1As shown, the server 104 is connected to the terminal device 102 via a network, and can be used to provide services (such as application services, etc.) for the terminal or the client installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for the server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data computing services for the server 104.
[0022] The aforementioned network may include, but is not limited to, at least one of the following: a wired network and a wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: a wide area network, a metropolitan area network, and a local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity) and Bluetooth. The terminal device 102 may be, but is not limited to, a PC, a mobile phone, a tablet computer, a smart air conditioner, a smart range hood, a smart refrigerator, a smart oven, a smart stove, a smart washing machine, a smart water heater, a smart washing machine, a smart dishwasher, a smart projection device, a smart TV, a smart clothes drying rack, smart curtains, smart audio and video, a smart socket, a smart speaker, a smart fresh air device, smart kitchen and bathroom equipment, smart bathroom equipment, a smart sweeping robot, a smart window cleaning robot, a smart mopping robot, a smart air purifier, a smart steamer, a smart microwave oven, a smart kitchen treasure, a smart purifier, a smart water dispenser, a smart door lock, etc.
[0023] In order to solve the technical problem of the application, a method for determining a reply audio is provided in this embodiment. Figure 2 Flowchart (1) of a method for determining a reply audio according to an embodiment of the present invention, the process includes the following steps:
[0024] Step S202, obtaining the target object's voice command;
[0025] It should be noted that the target objects include but are not limited to users of terminal devices, which include but are not limited to audio playback devices.
[0026] In an exemplary embodiment, the voice instructions include but are not limited to interactive voice instructions, query voice instructions, operational semantic instructions, etc., for example: Please play a few stories for me at random.
[0027] Step S204: acquiring a predetermined target language expression and target preference information of the target subject based on the voice command, wherein the target language expression is a language expression determined based on a plurality of pre-collected sample audios of the target subject, and the target preference information is preference information determined based on a set of historical operations of the target subject on a plurality of devices;
[0028] In one exemplary embodiment, voice commands can be subjected to voiceprint recognition to determine the identity of the target user, thereby obtaining the target user's pre-determined target language expression and preferences. Target language expression includes, but is not limited to, speech rate, emotion, word order, dialect, etc. Target preference information includes, but is not limited to, frequently used devices, frequently used device functions, frequently listened to music, and frequently viewed content.
[0029] In an exemplary embodiment, the target object's historical operation records on multiple devices include but are not limited to: various operations performed on a single device (such as what kind of music is played, what kind of content is queried, what kind of software is browsed, what kind of function is used, etc.), binding the device to a cloud server, etc. (such as binding fetal heart monitors, blood pressure monitors and other devices to a cloud server), setting up linkage operations between devices (such as setting the air conditioner to turn on when the smart door lock is opened, etc.), etc.
[0030] Step S206: determining a first reply audio of the voice instruction according to the target language expression and the target preference information, and sending the first reply audio to an audio playback device.
[0031] It should be noted that the technical solution of the embodiment of the present application can be applied on a cloud server, and this embodiment does not make any specific limitations here.
[0032] Through the above steps, when the target object's voice command is obtained, the target language expression and target preference information of the predetermined target object are obtained, and then the first reply audio of the voice command is determined based on the target language expression and target preference information, and the first reply audio is sent to the audio playback device. The above technical solution can intelligently reply to the user based on the user's language expression and preference information, better serve the user, and solve the problem that the device cannot determine the corresponding reply audio based on the user's language expression habits and preferences when obtaining the user's voice command.
[0033] In an exemplary embodiment, determining the first reply audio of the voice instruction according to the target language expression and the target preference information includes the following steps:
[0034] Step S1: determining reply information according to the target semantic information and the target preference information carried by the language instruction;
[0035] For example, assuming that the target semantic information carried by the language instruction is: please help introduce a historical figure of a dynasty at random. According to the target preference information of the user, it is determined that the user likes the Song Dynasty, and then the related information of a person in the Song Dynasty is searched on the Internet at random, and the searched related information is determined as the reply information; assuming that the target semantic information is: play some knowledge at random, and according to the target preference information, it is determined that the target object is bound with a sphygmomanometer at home, and then the knowledge related to blood pressure is searched, and it is determined as the reply information.
[0036] Step S2: voice synthesis of the reply information into the first reply audio according to the target language expression mode.
[0037] In an exemplary embodiment, the above step S2 can be implemented in the following manner:
[0038] Manner one: in the case where the target language expression mode indicates that the sentence of the first syntax is adjusted to the sentence of the second syntax, it is determined whether the reply information includes the first syntax sentence; in the case where the reply information includes the first syntax sentence, the syntax of the first sentence is adjusted from the first syntax to the second syntax; and the second sentence obtained after adjustment is voice synthesized into the first reply audio.
[0039] It should be noted that the first syntax includes but is not limited to: inverted sentence, word sentence, and character sentence. For example, if the target language expression mode of the user indicates that the user likes to use inverted sentences and the like, then the sentence in the reply information that can be expressed in inverted form is inverted, and then the reply information obtained after inversion is converted into the first reply audio. For example, "do you play games tonight" is converted into "do you play games, tonight".
[0040] Manner two: in the case where the target language expression mode indicates a target speech rate feature, wherein the target speech rate feature is used to indicate a target speech rate corresponding to the target object; the reply information is voice synthesized into the first reply audio according to the target speech rate indicated by the target speech rate feature, wherein the playing speed of the first reply audio is the target speech rate.
[0041] For example, if the target language expression mode of the user indicates that the speech rate of the user is 500 words per minute, then the reply information is voice synthesized into the first reply audio with a speech rate of 500 words per minute.
[0042] Manner three: in the case where the target language expression mode indicates a target emotion feature, wherein the target speech rate feature is used to indicate a target emotion corresponding to the target object; the reply information is voice synthesized into the first reply audio according to the target emotion feature, wherein the playing volume, playing pitch, and / or playing timbre of the first reply audio match the target emotion.
[0043] In one exemplary embodiment, after determining the target emotional characteristics of the user, a voice matching the target emotion can be selected from multiple recorded voices and used to synthesize the reply message into the first reply audio. For example, if the target language expression indicates that the user is easily irritable, a soothing voice can be selected and used to synthesize the reply message.
[0044] Method 4: When the target language expression indicates a target dialect, converting the reply message into a dialect reply message in the target dialect; and synthesizing the dialect reply message into the first reply audio. Alternatively, obtaining target dialect information, wherein the target dialect information is used to indicate the dialect used in the voice instruction; and synthesizing the reply message into the first reply audio based on the target dialect information and the target language expression.
[0045] For example, if the target language expression indicates that the user frequently speaks Cantonese, the reply message will be converted into a Cantonese dialect reply message for playback. Alternatively, if it is recognized that the user is currently asking in Cantonese, the reply message will be synthesized into the first reply audio using Cantonese and the user's target language expression (such as emotion, speaking speed, word order, etc.).
[0046] It should be noted that, when synthesizing the reply message speech into the first reply audio according to the target language expression, it can be achieved through method 1, and / or method 2, and / or method 3, and / or method 4. That is, it can be achieved by using at least one of the following methods: method 1, method 2, method 3, and method 4.
[0047] In an exemplary embodiment, determining the first reply audio of the voice instruction according to the target language expression and the target preference information further includes the following steps:
[0048] Step 1: Determine a reply audio according to the target semantic information and the target preference information carried by the voice command;
[0049] In an exemplary embodiment, assuming that the target semantic information is: play some opera at random, and according to the target preference information it is determined that the target object likes Peking opera, then a Peking opera can be randomly determined and determined as the reply audio.
[0050] Step 2: Adjust the reply audio to the first reply audio according to the target language expression.
[0051] In an example embodiment, the step two can be implemented in the following manner: in the case that the target language expression mode represents a target speed feature, adjusting the speed of the reply audio according to the target speed feature represented by the target speed, and determining the adjusted reply audio as the first reply audio. It should be noted that the target speed feature is used to represent the target speed corresponding to the target object, for example, 200 words per minute.
[0052] It should be noted that in the case that the target language expression mode represents a target volume feature, adjusting the volume of the reply audio according to the target volume feature represented by the target volume, and determining the adjusted reply audio as the first reply audio.
[0053] In an example embodiment, after obtaining the voice instruction of the target object, in the case that the voice instruction represents target operation information to be executed, obtaining the target language expression mode of the target object determined in advance; selecting a target operation reply audio matching the target language expression mode from a pre-set group of operation reply audios, wherein each operation reply information in the group of operation reply audios corresponds to a language expression mode; determining the target operation reply audio as the second reply audio, and sending the second reply audio to the audio playback device; or in the case that the voice instruction represents target operation information to be executed, obtaining the target language expression mode of the target object determined in advance; according to the target language expression mode, voice synthesizing a pre-set target operation reply information into a second reply audio, and sending the second reply audio to the audio playback device.
[0054] For better understanding, the following is a specific description, assuming that the voice instruction is to turn on the air conditioner, and after turning on the air conditioner, the operation reply audio needs to be broadcast to the user, for example, "the air conditioner is turned on", if the cloud server saves a group of operation reply audios, and different operation reply audios correspond to different language expression modes, then only the target operation reply audio matching the target language expression mode of the user needs to be selected from the pre-set group of operation reply audios to reply. If the device or the cloud does not save a group of operation reply audios, then the pre-set operation reply information "the air conditioner is turned on" needs to be voice synthesized into a second reply audio corresponding to the target language expression mode of the user.
[0055] It should be noted that the target language expression mode determined from the plurality of sample audios of the target object collected in advance can be implemented in the following manner: determining the sample language expression mode of the target object in each sample audio of the plurality of sample audios, obtaining a plurality of sample language expression modes; determining the target language expression mode of the target object through the plurality of sample language expression modes.
[0056] Specifically, determining the sample language expression of the target object in each sample audio of the multiple sample audios can be achieved in the following manner: determining the sample language type of each sample audio, and performing speech recognition on each sample audio through a speech recognition model corresponding to the sample language type to obtain a sample recognition result, wherein the sample speech recognition result includes: sample volume, sample text information, sample speaking speed, and sample emotion; parsing the sample text information to obtain a sample word order corresponding to the sample text information, wherein the sample language expression includes at least one of the following: sample language type, the sample volume, the sample speaking speed, the sample word order, and the sample emotion.
[0057] In an exemplary embodiment, determining the sample language type of each sample audio can be achieved by: performing feature extraction on each sample audio to obtain sample speech features, wherein the sample speech features include: Mel-frequency cepstral coefficient features and shifted differential cepstral features; sending the sample speech features to a language recognition neural network model to obtain the sample language type.
[0058] Obviously, the above-described embodiments are only part of the embodiments of the present invention, not all of them. To better understand the above-mentioned method for determining the reply audio, the above process is described below in conjunction with the embodiments, but is not intended to limit the technical solutions of the embodiments of the present invention. Specifically:
[0059] In an optional embodiment, Figure 3 Flowchart (II) of the method for determining reply audio according to an embodiment of the present invention, specifically:
[0060] Step S302: Analyze basic user information, for example:
[0061] (1) Determine the user's geographic location;
[0062] Based on the user's geographic location, if the user has not manually set the dialect type, the dialect type will be dynamically activated for recognition. In areas with a large number of immigrants, the Mandarin recognition model will be prioritized, and the dialect recognition type will be dynamically switched after accumulating user speech. In areas with a large number of local residents, the recognition preference of the local dialect type will be prioritized to improve recognition quality.
[0063] (2) Determine the bound feature device
[0064] For family accounts that are bound to the following devices, recommended content prioritizes content for those with the following characteristics, as shown in Table 1:
[0065] Table 1
[0066]
[0067] (3) Determine the user's voiceprint through the user's voice, and obtain multiple sets of user's voice data based on the voiceprint.
[0068] Step S304: The voice learning system determines the user's voice information (equivalent to the target language expression in the above embodiment). The voice learning system analyzes the user's identity and preferences from multiple voice dimensions by accumulating voice data. Specifically, the user's voice data is recorded and then clustered in the cloud. Cluster analysis can be based on, but not limited to, the following voice information dimensions: gender, age, language type (e.g., Chinese, English), accent (dialect type), speaking habits (e.g., favorite internet slang, speech inversion, etc.), volume, speaking rate, emotion, and other acoustic dimensions.
[0069] By extracting the phonemes of user voice information and reading environmental information as described above, data analysis is formed, and user voice information is recorded according to multiple and cross-analysis dimensions.
[0070] Step S306: The content learning system determines the user's preferences (equivalent to the target preference information in the above-mentioned embodiment). Specifically, the content learning system learns the preferred content of the family or family members, and forms independent family member voice identity information records based on the voiceprint information recorded by the user or the user identities learned through clustering in step S304. Based on different user identities, each user's preferred usage content is recorded, and cluster analysis can be formed based on, but not limited to, the following user usage habit information dimensions: preferred voice operation function categories; preferred device control modes / functions; preferred resources for on-demand resources; preferred chat content; preferred records of other voice skills; and preferred operation time periods.
[0071] Step S308: Dynamically switch the speaker and speaker-related parameters based on the identified different members, and provide preferred content for each user; specifically, after the information is recorded, the voice personality preferences of each member's information are recorded, and the speaker is dynamically matched and switched based on different preferences, achieving real-time dynamic switching of the speaker's timbre type, accent, reply content, reply tone, reply volume, reply content, etc. The dynamic changes of the speaker parameters are shown in Table 2:
[0072] Table 2
[0073]
[0074] Example of dynamically switching speakers and content: A family consists of the following members:
[0075] (1) User 1: elderly, speaking in dialect, noisy background, prefers Henan opera, and uses the service at 3 p.m.; query: "play radio"; reply: reply in dialect, increase volume appropriately, speak slower, and play Henan opera content;
[0076] (2) User 2: adult female voice, speaking Mandarin, quiet background noise, preferred content is news, used during dinner time; query content: "play radio"; reply: replied in Mandarin, with normal volume and speed, and the radio content played is China National Radio News.
[0077] In addition, the above-mentioned technical solution of the embodiment of the present invention accumulates user data and the system automatically learns and clusters user voice habits through the user's basic information and daily interaction with the voice device through voice. After the user enters the basic voiceprint information in advance, or through the voice information accumulated by the user using voice, voice information data statistics are formed in different dimensions such as family, account, environment, and bound device.
[0078] In this embodiment of the present invention, users do not need to manually select a dialect type. Instead, their identity can be confirmed based on the user's spoken information, allowing for dynamic switching of dialect types. Furthermore, a scheme is proposed for dynamically adjusting speaker parameters based on the user's personality and language habits, enabling dynamic switching of speakers and reply content across multiple dimensions, including volume, speaking speed, and content. This allows for smoother voice communication, tailored to the user's habits and preferences, with minimal or no user intervention. This creates a dedicated voice assistant for each family member, providing personalized, dynamic services across multiple dimensions, including voice experience and content.
[0079] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of various embodiments of the present invention.
[0080] This embodiment also provides a device for determining a reply audio signal, which is used to implement the above-mentioned embodiments and preferred embodiments. Details already described will not be repeated here. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
[0081] Figure 4 1 is a structural block diagram of a device for determining a reply audio according to an embodiment of the present invention (I), the device comprising:
[0082] A first acquisition module 42 is used to acquire a voice instruction of a target object;
[0083] a second acquisition module 44 configured to acquire, based on the voice command, a predetermined target language expression and target preference information of the target subject, wherein the target language expression is a language expression determined based on a plurality of pre-collected sample audios of the target subject, and the target preference information is preference information determined based on a set of historical operations of the target subject on a plurality of devices;
[0084] The determination module 46 is configured to determine a first reply audio of the voice instruction according to the target language expression and the target preference information, and send the first reply audio to an audio playback device.
[0085] By means of the above-mentioned device, when a voice instruction of a target object is obtained, the target language expression and target preference information of the predetermined target object are obtained, and then the first reply audio of the voice instruction is determined based on the target language expression and target preference information, and the first reply audio is sent to the audio playback device. By adopting the above-mentioned technical solution, the user can be intelligently replied according to the user's language expression and preference information, thus better serving the user, and solving the problem that the device cannot determine the corresponding reply audio according to the user's language expression habits and preferences when obtaining the user's voice instruction.
[0086] In an exemplary embodiment, the determination module 46 is further used to determine the reply information based on the target semantic information and the target preference information carried by the voice instruction; and synthesize the reply information into the first reply audio according to the target language expression.
[0087] In an exemplary embodiment, the determination module 46 is further used to determine the reply audio based on the target semantic information and the target preference information carried by the voice instruction; and adjust the reply audio to the first reply audio according to the target language expression.
[0088] In an exemplary embodiment, the determining module 46 is further configured to, in a case where the target language expression manner indicates that the target language expression manner represents adjusting a sentence in a first word order into a sentence in a second word order, and the reply information comprises a first sentence in the first word order, adjusting the word order of the first sentence from the first word order into the second word order, and performing speech synthesis on the second sentence obtained after the adjustment to obtain the first reply audio; in a case where the target language expression manner indicates that the target language expression manner represents a target speech rate feature, wherein the target speech rate feature is used to represent a target speech rate corresponding to the target object, performing speech synthesis on the reply information to obtain the first reply audio according to the target speech rate indicated by the target speech rate feature; in a case where the target language expression manner indicates that the target language expression manner represents a target emotion feature, wherein the target emotion feature is used to represent a target emotion corresponding to the target object, performing speech synthesis on the reply information to obtain the first reply audio according to the target emotion feature; and in a case where the target language expression manner indicates that the target language expression manner represents a target dialect, converting the reply information into dialect reply information in the target dialect, and performing speech synthesis on the dialect reply information to obtain the first reply audio.
[0089] In an exemplary embodiment, the determining module 46 is further configured to obtain target dialect information, wherein the target dialect information is used to represent a dialect adopted by the speech instruction; and perform speech synthesis on the reply information to obtain the first reply audio according to the target dialect information and the target language expression manner.
[0090] In an exemplary embodiment, the determining module 46 is further configured to, in a case where the target language expression manner indicates that the target language expression manner represents a target speech rate feature, wherein the target speech rate feature is used to represent a target speech rate corresponding to the target object, adjust the speech rate of the reply audio according to the target speech rate indicated by the target speech rate feature, and determine the reply audio obtained after the adjustment as the first reply audio.
[0091] Figure 5 FIG. 2 is a structural block diagram of a reply audio determination apparatus according to an embodiment of the present application, which comprises a determining module 46.
[0092] In an exemplary embodiment, the processing module 48 is used to obtain the predetermined target language expression of the target object when the voice instruction indicates the target operation information to be executed; select the target operation response audio that matches the target language expression from a preset set of operation response audios, wherein each operation response information in the set of operation response audios corresponds to a language expression; determine the target operation response audio as the second response audio, and send the second response audio to the audio playback device; or when the voice instruction indicates the target operation information to be executed, obtain the predetermined target language expression of the target object; synthesize the preset target operation response information into the second response audio according to the target language expression, and send the second response audio to the audio playback device.
[0093] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any one of the above method embodiments when running.
[0094] Optionally, in this embodiment, the storage medium may be configured to store a computer program for performing the following steps:
[0095] S1, obtain the target object’s voice instructions;
[0096] S2, acquiring a predetermined target language expression and target preference information of the target subject according to the voice command, wherein the target language expression is a language expression determined based on a plurality of sample audios of the target subject collected in advance, and the target preference information is preference information determined based on a set of historical operations of the target subject on a plurality of devices;
[0097] S3: Determine a first reply audio of the voice instruction according to the target language expression and the target preference information, and send the first reply audio to an audio playback device.
[0098] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0099] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail here.
[0100] An embodiment of the present invention further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0101] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:
[0102] S1, obtain the target object’s voice instructions;
[0103] S2, acquiring a predetermined target language expression and target preference information of the target subject according to the voice command, wherein the target language expression is a language expression determined based on a plurality of sample audios of the target subject collected in advance, and the target preference information is preference information determined based on a set of historical operations of the target subject on a plurality of devices;
[0104] S3: Determine a first reply audio of the voice instruction according to the target language expression and the target preference information, and send the first reply audio to an audio playback device.
[0105] In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0106] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail here.
[0107] Obviously, those skilled in the art will appreciate that the various modules or steps of the present invention described above can be implemented using a general-purpose computing device, can be centralized on a single computing device, or can be distributed across a network of multiple computing devices. They can be implemented using program code executable by the computing device, and thus, can be stored in a storage device and executed by the computing device. In some cases, the steps shown or described herein can be performed in a different order than that shown, or can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0108] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for determining a reply audio, characterized in that: include: Obtaining voice commands from the target object; Acquiring a predetermined target language expression and target preference information of the target object according to the voice command, wherein the target language expression is a language expression determined based on a plurality of sample audios of the target object collected in advance, and the target preference information is preference information determined based on a set of historical operations of the target object on a plurality of devices; Determining a first reply audio of the voice instruction according to the target language expression and the target preference information, and sending the first reply audio to an audio playback device; The method further includes: pre-collecting a plurality of sample audios of the target object; determining a sample language expression of the target object in each of the plurality of sample audios to obtain a plurality of sample language expressions; and determining a target language expression of the target object based on the plurality of sample language expressions; Determining the sample language expression of the target object in each sample audio of the multiple sample audios includes: determining the sample language type of each sample audio, and performing speech recognition on each sample audio through a speech recognition model corresponding to the sample language type to obtain a sample recognition result, wherein the sample recognition result includes: sample volume, sample text information, sample speaking speed, and sample emotion; parsing the sample text information to obtain a sample word order corresponding to the sample text information, wherein the sample language expression includes: sample language type, the sample volume, the sample speaking speed, the sample word order, and the sample emotion; Determining the sample language type of each sample audio includes: extracting features from each sample audio to obtain sample speech features, wherein the sample speech features include: Mel-frequency cepstral coefficient features and shifted differential cepstral features; and sending the sample speech features to a language recognition neural network model to obtain the sample language type.
2. The method according to claim 1, characterized in that Determining a first reply audio of the voice instruction according to the target language expression and the target preference information includes: Determining reply information according to the target semantic information and the target preference information carried by the voice command; The reply information is voice-synthesized into the first reply audio according to the target language expression.
3. The method according to claim 1, characterized in that Determining a first reply audio of the voice instruction according to the target language expression and the target preference information includes: Determining a reply audio according to the target semantic information and the target preference information carried by the voice command; According to the target language expression, the reply audio is adjusted to the first reply audio.
4. The method according to claim 2, characterized in that The step of synthesizing the reply message into the first reply audio according to the target language expression includes: When the target language expression indicates adjusting a sentence in a first word order to a sentence in a second word order, and the reply information includes a first sentence in the first word order, adjusting the word order of the first sentence from the first word order to the second word order, and synthesizing the second sentence obtained after the adjustment into the first reply audio; and / or In a case where the target language expression represents a target speech rate feature, wherein the target speech rate feature is used to represent a target speech rate corresponding to the target object; synthesizing the reply information speech into the first reply audio according to the target speech rate represented by the target speech rate feature; and / or In the case where the target language expression represents a target emotion feature, wherein the target speech rate feature is used to represent the target emotion corresponding to the target object; synthesizing the reply information speech into the first reply audio according to the target emotion feature; and / or In a case where the target language expression represents a target dialect, the reply information is converted into dialect reply information of the target dialect, and the dialect reply information is speech-synthesized into the first reply audio.
5. The method according to claim 2, characterized in that The step of synthesizing the reply message into the first reply audio according to the target language expression includes: Acquiring target dialect information, wherein the target dialect information is used to indicate the dialect used in the voice instruction; The reply information is speech-synthesized into the first reply audio according to the target dialect information and the target language expression.
6. The method according to claim 3, characterized in that The step of adjusting the reply audio to the first reply audio according to the target language expression includes: In the case where the target language expression mode represents a target speech rate feature, wherein the target speech rate feature is used to represent a target speech rate corresponding to the target object; The speaking speed of the reply audio is adjusted according to the target speaking speed represented by the target speaking speed feature, and the adjusted reply audio is determined as the first reply audio.
7. The method according to claim 1, characterized in that After acquiring the voice instruction of the target object, the method further includes: In the case where the voice instruction represents target operation information to be executed, obtaining the predetermined target language expression of the target object; selecting a target operation reply audio that matches the target language expression from a preset set of operation reply audios, wherein each operation reply audio in the set of operation reply audios corresponds to a language expression; determining the target operation reply audio as a second reply audio, and sending the second reply audio to the audio playback device; or In the case where the voice instruction represents the target operation information to be executed, the predetermined target language expression of the target object is obtained; according to the target language expression, the preset target operation reply information is voice synthesized into a second reply audio, and the second reply audio is sent to the audio playback device.
8. A device for determining a reply audio, characterized in that: include: A first acquisition module is used to acquire a voice instruction of a target object; a second acquisition module, configured to acquire, based on the voice command, a predetermined target language expression and target preference information of the target object, wherein the target language expression is a language expression determined based on a plurality of pre-collected sample audios of the target object, and the target preference information is preference information determined based on a set of historical operations of the target object on a plurality of devices; a determination module, configured to determine a first reply audio of the voice instruction according to the target language expression and the target preference information, and send the first reply audio to an audio playback device; The second acquisition module is further configured to pre-collect a plurality of sample audios of the target object; determine a sample language expression of the target object in each of the plurality of sample audios to obtain a plurality of sample language expressions; and determine a target language expression of the target object based on the plurality of sample language expressions; The second acquisition module is further configured to determine the sample language type of each sample audio, and perform speech recognition on each sample audio using a speech recognition model corresponding to the sample language type to obtain a sample recognition result, wherein the sample recognition result includes: sample volume, sample text information, sample speaking speed, and sample emotion; and parse the sample text information to obtain a sample word order corresponding to the sample text information, wherein the sample language expression includes: sample language type, the sample volume, the sample speaking speed, the sample word order, and the sample emotion; Among them, the second acquisition module is also used to extract features from each sample audio to obtain sample speech features, wherein the sample speech features include: Mel-frequency cepstral coefficient features and shifted difference cepstral features; and send the sample speech features to the language recognition neural network model to obtain the sample language type.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the method according to any one of claims 1 to 7 when executed.
10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 7 through the computer program.
Citation Information
Patent Citations
Speech synthesis method and related equipment
CN108962217A
Information interaction method and device, electronic equipment and storage medium
CN111639162A
Man-machine cooperation non-inductive control method and device for multiple robots
CN113067952A