Voice interaction method, device and computer readable storage medium
By acquiring the voice attributes and object tags of the target object and adjusting the configuration parameters of the virtual voice object, the problem of limited digital human image and acoustic configuration is solved, and a rich human-computer interaction experience and fun are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- VOICEAI TECH CO LTD
- Filing Date
- 2022-12-27
- Publication Date
- 2026-04-21
AI Technical Summary
In existing human-computer interaction scenarios, the digital human image and acoustic configuration are relatively simple, resulting in boring and uninteresting human-computer interaction scenarios and reducing user experience.
By acquiring the voice attributes and object attribute tags of the target object, the voice interaction scenario category is identified, and the configuration parameters of the virtual voice object, including the voice interaction model and body model, are adjusted to enrich the image and acoustic configuration of the digital human.
It enhances the fun and user experience of human-computer interaction scenarios, making interactions more diverse and in line with user needs through personalized voice and avatar configurations.
Smart Images

Figure CN116092489B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, specifically to a voice interaction method, device, and computer-readable storage medium. Background Technology
[0002] With the development of information technology, human-computer interaction (HCI) technology has been widely applied in various industries, such as intelligent robots and welcoming robots, all of which require the support of HCI technology. However, while current HCI scenarios can recognize user voice and provide corresponding responses, such as through voice or human-like feedback, this makes the HCI scenarios relatively simple and mechanical. Digital humans, on the other hand, are digitized human figures created using digital technology that closely resemble human figures. Applying digital human creations to HCI scenarios allows for human-computer dialogue in a digital form, giving voice interaction devices a more human-like character.
[0003] However, the current digital humans in human-computer interaction scenarios are relatively simple in terms of appearance and acoustic configuration, which makes the human-computer interaction scenarios too boring and lacks fun, thus reducing the user's experience in human-computer interaction scenarios. Summary of the Invention
[0004] This application provides a voice interaction method, device, and computer-readable storage medium. These features enrich the visual and acoustic configuration of human-computer interaction scenarios, enhance their appeal, and improve the user experience.
[0005] This application provides a voice interaction method, including:
[0006] Retrieve the voice attributes and object attribute tags of the target object;
[0007] Identify the scene category in the current voice interaction scenario and select the virtual voice object corresponding to the scene category;
[0008] Adjust the configuration parameters of the virtual voice object according to the voice attributes and object attribute tags to obtain the target virtual voice object;
[0009] The voice information of the target object is obtained, and voice interaction is performed on the voice information through the target virtual voice object.
[0010] Accordingly, embodiments of this application provide a voice interaction device, including:
[0011] The acquisition unit is used to acquire the voice attributes and object attribute tags of the target object;
[0012] The selection unit is used to identify the scene category in the current voice interaction scenario and select the virtual voice object corresponding to the scene category;
[0013] An adjustment unit is used to adjust the configuration parameters of the virtual voice object according to the voice attribute and object attribute label to obtain the target virtual voice object;
[0014] An interaction unit is used to acquire the voice information of the target object and perform voice interaction on the voice information through the target virtual voice object.
[0015] In some embodiments, the adjustment unit is further configured to:
[0016] Determine the voice interaction model and body posture model associated with the virtual voice object;
[0017] Adjust the sound configuration parameters of the voice interaction model associated with the virtual voice object according to the voice attributes to obtain the target voice interaction model;
[0018] Adjust the image configuration parameters of the body model according to the object attribute tags to obtain the target body model;
[0019] Based on the target voice interaction model and the target body posture model, construct the target virtual voice object.
[0020] In some embodiments, the adjustment unit is further configured to:
[0021] Extract the acoustic feature values corresponding to the speech attributes;
[0022] Adjust the sound configuration parameters of the voice interaction model associated with the virtual voice object according to the voice attributes to obtain the target voice interaction model.
[0023] In some embodiments, the adjustment unit is further configured to:
[0024] Retrieve body posture attributes from the database that match the object's attribute labels;
[0025] The body posture configuration parameters of the body posture model are adjusted according to the body posture attribute characteristics to obtain the target body posture model.
[0026] In some embodiments, the adjustment unit is further configured to:
[0027] The clothing characteristics of the target object are determined based on the clothing label;
[0028] Based on feature similarity, target clothing features that are similar to the clothing features of the target object are determined from the database;
[0029] Adjust the color feature value of the target clothing feature according to the color difference label to obtain the body shape attribute feature.
[0030] In some embodiments, the selection unit is further configured to:
[0031] Obtain the selected voice interaction text of the target object;
[0032] Identify the semantic information corresponding to the voice interaction text;
[0033] The scene category in the current voice interaction scenario is determined based on the semantic information.
[0034] In some embodiments, the interaction unit is further configured to:
[0035] Based on semantic similarity, query the voice response text that matches the voice information;
[0036] The voice response text is converted into speech to obtain acoustic speech features;
[0037] Based on the target virtual voice object, voice interaction is performed based on the acoustic voice features.
[0038] In some embodiments, the acquiring unit is further configured to:
[0039] Determine the object identifier of the target object and query the interactive feedback information associated with the object identifier;
[0040] Semantic recognition is performed on the interactive feedback information to obtain the target feedback semantics;
[0041] Based on the target feedback semantics, determine the voice attribute and object attribute labels.
[0042] Furthermore, this application also provides a computer device, including a processor and a memory, wherein the memory stores a computer program, and the processor is used to run the computer program in the memory to implement the steps in the voice interaction method provided in this application.
[0043] Furthermore, embodiments of this application also provide a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute steps in any of the voice interaction methods provided in embodiments of this application.
[0044] Furthermore, embodiments of this application also provide a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps in any of the voice interaction methods provided in embodiments of this application.
[0045] This application embodiment can obtain the voice attributes and object attribute tags of a target object; identify the scene category in the current voice interaction scenario and select the virtual voice object corresponding to the scene category; adjust the configuration parameters of the virtual voice object according to the voice attributes and object attribute tags to obtain the target virtual voice object; obtain the voice information of the target object, and perform voice interaction on the voice information through the target virtual voice object. Therefore, this solution can first determine the voice attributes and object attribute tags of the target object, identify the current voice interaction scene category of the target object, select a virtual voice object that matches the scene category, then adjust the acoustic and visual configuration parameters of the virtual voice object according to the voice attributes and object attribute tags to obtain the adjusted target virtual voice object, and finally, perform voice interaction on each sentence of voice information of the target object through the target virtual voice object. In this way, the visual and acoustic configuration of the human-computer interaction scenario can be enriched according to the user's actual needs, improving the fun of the human-computer interaction scenario and enhancing the user's experience in the human-computer interaction scenario. Attached Figure Description
[0046] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0047] Figure 1 This is a schematic diagram of a scenario for the voice interaction system provided in an embodiment of this application;
[0048] Figure 2 This is a flowchart illustrating the steps of the voice interaction method provided in the embodiments of this application;
[0049] Figure 3 This is a schematic flowchart of another step of the voice interaction method provided in the embodiments of this application;
[0050] Figure 4 This is a schematic diagram of the structure of the voice interaction device provided in the embodiments of this application;
[0051] Figure 5 This is a schematic diagram of the structure of the computer device provided in the embodiments of this application. Detailed Implementation
[0052] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0053] This application provides a voice interaction method, apparatus, and computer-readable storage medium. This application will describe the voice interaction apparatus from the perspective of the voice interaction device, which can be integrated into a computer device. The computer device can be a server, which can be a standalone physical server, a server cluster composed of multiple physical servers, or a distributed system. It can also provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, etc. Furthermore, the computer device can be a terminal device. The terminal can be a television, smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, wearable device, in-vehicle terminal, etc., but is not limited to these.
[0054] For example, see Figure 1 This is a schematic diagram of a voice interaction system provided in an embodiment of this application. The scenario includes a terminal or a server.
[0055] The terminal or server can obtain the voice attributes and object attribute tags of the target object; identify the scene category in the current voice interaction scenario and select the virtual voice object corresponding to the scene category; adjust the configuration parameters of the virtual voice object according to the voice attributes and object attribute tags to obtain the target virtual voice object; obtain the voice information of the target object and perform voice interaction on the voice information through the target virtual voice object.
[0056] Voice interaction can include processing methods such as acquiring sound signals, determining the weight coefficients and volume values of the sound signals, determining the acoustic events to be processed first, determining the target event information corresponding to the acoustic events to be processed first, and voice interaction.
[0057] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the preferred order of the embodiments.
[0058] In this application embodiment, the description will be from the perspective of a voice interaction device, specifically one that can be integrated into a computer device such as a terminal device or a server. See also Figure 2 , Figure 2This application provides a flowchart illustrating the steps of a voice interaction method. For example, when a processor on a server or terminal device executes the program corresponding to the voice interaction method, the specific flow of the voice interaction method is as follows:
[0059] 101. Obtain the voice attribute and object attribute tag of the target object.
[0060] In this embodiment, the voice interaction method can be applied to voice interaction scenarios between users and "digital humans," which are digitally created using digital technology and closely resemble human figures. However, to enrich the image and acoustic configuration of digital humans in human-computer interaction scenarios, the voice attributes and object attribute tags of the target users participating in the voice interaction scenario can be obtained. This allows for the subsequent shaping of the digital human's image and sound effects based on the user's voice attributes and object attribute tags, enriching the image and sound effects of the digital human. Furthermore, the language attributes and object attribute tags of the target object can be used to adjust the configuration of the language system, enabling the target virtual voice object to perform voice interaction, thereby enhancing the fun and richness of the human-computer interaction scenario.
[0061] The target audience can be the current users participating in the voice interaction scenario or the users who have participated in the voice interaction scenario before, such as students who have received training or customer service representatives.
[0062] Among these, speech attributes can be various properties contained in the sound emitted by the target object, including but not limited to pitch, intensity, duration, and timbre. It is worth noting that different target objects possess different speech attributes.
[0063] Among them, object attribute tags can be object information of the target object, including but not limited to a series of tag information that can reflect the user's characteristics, such as the user's age, occupation, clothing, clothing style, color, etc.
[0064] Specifically, the target object's current voice can be collected and its voice attributes can be obtained through speech spectrum analysis. At the same time, images containing the target object can be collected and its object attribute labels can be analyzed through image information. Specifically, the object attribute labels of the target object can be evaluated through a trained neural network model. For example, the model can predict and evaluate the target object's age range, clothing style, clothing color, etc. In this way, the configuration parameters of the digital human in the current voice interaction scenario can be configured based on the voice attributes and object attribute labels obtained from the analysis.
[0065] It should be noted that when analyzing speech, the pitch, intensity, duration, and timbre of the target speech can also be analyzed. It can also be used to determine the speech speed, whether the speech characteristics are regional, and whether there are other speech characteristics compared to standard speech, thereby determining the speech attributes of the target speech.
[0066] For example, let user A be the target object. User A is from Northeast China and has a distinct Northeast regional accent. User A also speaks quickly. Therefore, user A's voice attributes include the two attributes of Northeast regional accent and fast speaking speed. Through a user survey of user A, user A is 24 years old and works as a salesperson. Therefore, the corresponding object attribute tags for Xiaoming include the three object tags of name, age, and occupation.
[0067] In some implementations, the desired image and voice characteristics of the digital human can be understood from information provided by the target object. For example, the user's (demanded) voice characteristics and image can be obtained from user complaint information. Step 101, "obtaining the target object's voice attributes and object attribute tags," may include:
[0068] (101.1) Determine the object identifier of the target object and query the interactive feedback information associated with the object identifier;
[0069] (101.2) Perform semantic recognition on the interactive feedback information to obtain the target feedback semantics;
[0070] (101.3) Determine the speech attribute and object attribute labels based on the target feedback semantics.
[0071] The object identifier can be the account identifier of the target object when logging into the voice interaction application (platform), representing the target object's online identity on the voice interaction platform. The object identifier can be an account, nickname, number, custom identifier, etc.
[0072] The interactive feedback information can be information provided by the target object in response to the interactive scenario at a historical moment. This interactive feedback information can include the voice characteristics and image characteristics provided by the target object, reflecting the target object's experience with human-computer voice interaction. For example, the interactive feedback information includes the target object's experience with the "digital human," such as the "digital human's" pitch being too high, its timbre being too poor, its voice quality being too harsh, and its image being too old; or the "digital human's" timbre value not being in the range of (12, 26), its image not resembling that of a 7-12 year old, its clothing color not being red, and its clothing style not resembling a school uniform, etc.
[0073] In this context, target feedback semantics can be semantics composed of keywords in the interactive feedback information, reflecting the voice effect and image of the "digital human" desired by the target object. For example, if the interactive feedback information is "The image of the digital human is not suitable for 7-12 year olds, the pitch value is not AB, and the clothing style is not a school uniform," then after semantic recognition, the target feedback semantics can be obtained as "The image design of the digital human is for 7-12 year olds, the clothing style is a school uniform, and the pitch value is AB." Then, the subsequent voice attribute is "the pitch value is AB," and the object attribute label is "the image design is for 7-12 year olds, and the clothing style is a school uniform."
[0074] Specifically, after identifying the target object, the next step is to determine its corresponding object identifier. This object identifier is used to query the information provided by the target object in the database and then return it as interactive feedback information. For example, if a salesperson is identified as the target object, their data in the database is associated with the identifier 001. Therefore, using the object identifier 001, the system can find related data in the user database, such as suggestions or complaints made by the salesperson during historical voice interactions. Further, the interactive system performs semantic recognition on the returned interactive feedback information to obtain the processed target semantic information, i.e., target feedback semantics. Based on the information in the target feedback semantics, the system identifies the voice attributes and object attribute tags corresponding to the target object. For example, the salesperson's object identifier is 001. The voice interaction system uses the identifier 001 to query the database for related complaint information, then performs semantic recognition on the complaint information to obtain the content of the information, such as the digital human's voice characteristics, speaking habits, content keyword extraction, etc., to obtain the target feedback semantics. Finally, the target feedback semantics are used to extract and determine the corresponding voice attributes and object attribute tags in the complaint information, such as speaking speed, clarity of pronunciation, clothing, etc.
[0075] By using the above methods, the voice attributes and object attribute tags of the target object can be obtained, so that the language system configuration can be adjusted using the language attributes and object attribute tags of the target object. In this way, the target virtual voice object can perform voice interaction, thereby improving the fun and richness of the human-computer interaction scenario.
[0076] 102. Identify the scene category in the current voice interaction scenario and select the virtual voice object corresponding to the scene category.
[0077] In this embodiment of the application, in order to enable diverse voice interaction through the target virtual voice object and thus enrich the human-computer interaction scenario, the current voice interaction scenario can be used to obtain the scenario category, and then a suitable virtual voice object can be selected to make the human-computer interaction scenario more targeted.
[0078] The scenario category can be the specific scenario required for users to use digital humans for human-computer interaction, such as oral practice training for primary school students, customer service skills training, sales skills training, and interview skills training for news reporters.
[0079] Among them, virtual voice objects can be digital human figures created using digital technology that are close to human figures, or they can be virtual humans with human forms simulated by computers, i.e., digital humans.
[0080] Specifically, the voice interaction system identifies the category of the voice interaction scenario based on the current voice context, i.e., the scenario category. For example, in the current customer service training scenario, the identified scenario type is a training scenario; in a buyer-seller dialogue scenario, the identified scenario type is a dialogue scenario established from the perspectives of both the buyer and seller. After identifying the scenario category in the current voice interaction scenario, a computer-simulated virtual human-like character corresponding to that scenario category is selected, i.e., a digital human (virtual voice object). For example, if the current voice interaction scenario category is a training scenario, then the digital human for that scenario is the corresponding training instructor character.
[0081] In some implementations, the target object information associated with the target object can be queried based on the target object's target identifier, and semantic recognition can be performed to obtain the target object's voice attributes and object attribute tags. For example, step 102, "identifying the scene category in the current voice interaction scenario," may include:
[0082] (102.1) Obtain the selected voice interaction text of the target object;
[0083] (102.2) Recognize the semantic information corresponding to the voice interaction text;
[0084] (102.3) Determine the scene category in the current voice interaction scene based on semantic information.
[0085] Among them, voice interaction text can be the interactive content text required in the human-computer interaction scenario between the target object and the digital human. For example, in the scenario of sales skills training, voice interaction text is the text of daily sales conversations.
[0086] Semantic information can be information in the relevant format of the voice interaction text, used to concisely express the meaning of related dialogue questions through keywords. For example, if the dialogue question is "Sir, do you need insurance, medical insurance or pension insurance?", the semantic information could be "Insurance = Pension insurance or Medical insurance". The above is just an example and is not limited here.
[0087] Specifically, after acquiring the target object, the voice interaction system determines the target object information. Then, based on this information, it determines the voice interaction text. For example, if the target object is a salesperson, the corresponding voice interaction text is a sales script. Next, based on the determined voice interaction text, semantic recognition is performed. By using the contextual information within the text, the specific semantics of the text are determined, i.e., semantic information. Finally, based on this semantic information, the scenario category corresponding to the current target object is determined. For example, if the voice interaction text is a sales script, and after semantic recognition, the semantic information indicates that the sales script is an insurance sales presentation practice text, and it is within a sales presentation practice scenario, then the scenario category in the current voice interaction scenario is a sales presentation practice scenario.
[0088] By using the above methods, the scene category in the current voice interaction scenario can be identified, and the virtual voice object corresponding to that scene category can be selected. This allows the scene category to be obtained using the current voice interaction scenario, and then a suitable virtual voice object can be selected, making subsequent human-computer interaction scenarios more targeted.
[0089] 103. Adjust the configuration parameters of the virtual voice object according to the voice attribute and object attribute tags to obtain the target virtual voice object.
[0090] In this embodiment of the application, in order to make the digital human more suitable for the interaction scenario, the configuration parameters of the virtual voice object can be adjusted by the voice attribute and object attribute tag, thereby improving the digital human, making the human-computer interaction scenario more targeted, and improving the user's experience in the human-computer interaction scenario.
[0091] The configuration parameters can be virtual parameters corresponding to a complete virtual voice object, such as the height and weight parameters, facial parameters, voice parameters, clothing parameters, and style parameters of the virtual voice object.
[0092] The target virtual voice object can be a virtual voice object (digital human) obtained by adjusting the configuration parameters of a virtual voice object. For example, a male sales digital human on the screen who wears glasses and a black suit and speaks Cantonese.
[0093] Specifically, after determining the target object's language attributes and object attribute tags, the configuration parameters of the virtual voice object are adjusted based on the target object's voice attributes and object attribute tags to obtain the adjusted digital human, i.e., the target virtual voice object. For example, the current target object has three voice attributes: a Northeastern accent, a relatively high-pitched voice, and a speech rate of 200 words per minute. The voice configuration parameters of the virtual voice object are adjusted based on these three voice attributes. Then, based on the current target object's three object attribute tags: salesperson, black suit, the digital human's tag configuration parameters are adjusted to obtain a sales-type digital human with a Northeastern accent, a relatively high-pitched voice, a speech rate of 200 words per minute, and wearing a black suit.
[0094] In some implementations, the configuration parameters of the virtual voice object can be adjusted according to the voice attributes and object attribute tags of the target object to obtain the target virtual voice object. For example, step 103, "adjusting the configuration parameters of the virtual voice object according to the voice attributes and object attribute tags to obtain the target virtual voice object," may include:
[0095] (103.1) Determine the voice interaction model and body posture model associated with the virtual voice object;
[0096] (103.2) Adjust the sound configuration parameters of the voice interaction model associated with the virtual voice object according to the voice attributes to obtain the target voice interaction model;
[0097] (103.3) Adjust the image configuration parameters of the body model according to the object attribute tags to obtain the target body model;
[0098] (103.4) Construct the target virtual voice object based on the target voice interaction model and the target body posture model.
[0099] The voice interaction model can be a vocalization model for a digital human (virtual voice object) to interact with the user. Virtual voice objects can use this model to interact with users, making human-computer interaction simpler and more natural.
[0100] Among them, the body model can be a digital human image model in which the virtual voice object appears in a visualized three-dimensional form. By adjusting the body model, the image features such as the virtual voice object's body, clothing, and expression can be adjusted.
[0101] The target voice interaction model can be a new voice interaction model obtained by adjusting the configuration parameters of the voice interaction model of the virtual voice object, and the target voice interaction model conforms to the use of the target virtual voice object in the current scene category.
[0102] The target body model can be a new body model obtained by adjusting the configuration parameters of the body model of the virtual voice object, and the target body model conforms to the use of the target virtual voice object in the current scene category.
[0103] Specifically, after determining the current virtual voice object based on the current scene category, the voice interaction system then reads the corresponding voice interaction model and body model to obtain the configuration parameters of the current digital human. Next, based on the voice attributes of the target object and the configuration parameters in the object attribute tags, the system adjusts the configuration parameters of the voice interaction model and body model corresponding to the current virtual voice object. After adjustment, the target voice interaction model and target body model are obtained respectively. Finally, a new digital human, namely the target virtual voice object, is constructed using the target voice interaction model and target body model.
[0104] For example, the current scenario is an interview training scenario. The corresponding digital human (virtual voice object) has a voice interaction model with a pitch of 50 Hz, a body model with a height of 170cm, a gender of male, a body type of thin, and a clothing configuration of a black suit. Based on the voice attributes of the target object, the current configuration parameters of the target object are determined to be a pitch of 60 Hz. The object attribute tags include a gender of male, a body type of fat, and a clothing configuration of a white suit. After adjustment, the original digital human configuration parameters are reconstructed. The voice configuration parameters of the newly constructed digital human are changed to a pitch of 60 Hz, and the image configuration parameters are configured as a gender of male, a body type of fat, and a clothing configuration of a white suit.
[0105] In some implementations, step (103.2), "adjusting the sound configuration parameters of the voice interaction model associated with the virtual voice object according to the voice attributes to obtain the target voice interaction model," may include:
[0106] (103.2.1) Extract the acoustic feature values corresponding to the speech attributes;
[0107] (103.2.2) Adjust the sound configuration parameters of the voice interaction model associated with the virtual voice object according to the voice attributes to obtain the target voice interaction model.
[0108] Among them, acoustic feature values can be physical quantities that represent the acoustic characteristics of speech, including values for pitch, timbre, duration, and intensity. Examples include energy concentration areas, formant frequencies, formant intensities, and bandwidth that represent timbre, as well as duration, fundamental frequency, and average speech power that represent the prosodic characteristics of speech.
[0109] Specifically, after obtaining the voice attributes of the target object, the voice attributes are analyzed to obtain the values of pitch, timbre, duration, and intensity, i.e., acoustic feature values. Then, based on the extracted acoustic feature values, the configuration parameters of pitch, timbre, duration, and intensity of the voice interaction model associated with the digital human (virtual voice object) are adjusted to construct a new voice interaction model, i.e., the target voice interaction model. For example, after extracting the acoustic feature values of the voice attributes in the complaint information of target object A, it is determined that the formant frequency corresponding to the timbre of the digital human should be 50 Hz, while the current timbre configuration parameter of the voice interaction model of the digital human (virtual voice object) is 30 Hz. Therefore, the original formant frequency of 30 Hz is adjusted to 30 Hz to obtain a corresponding new voice interaction model, i.e., the target voice interaction model.
[0110] In some implementations, step (103.3), "adjusting the image configuration parameters of the body model according to the object attribute tags to obtain the target body model," may include:
[0111] (103.3.1) Query the database for body posture attributes that match the object's attribute labels;
[0112] (103.3.2) Adjust the posture configuration parameters of the posture model according to the posture attribute characteristics to obtain the target posture model.
[0113] Among them, physical attributes can be the characteristics of the target object's face, appearance, body shape, etc., such as a person's height, weight, or thinness, the color of clothing (yellow, white, black, red), and the style of clothing.
[0114] Specifically, after obtaining the object attribute tags of the target object, the next step is to analyze these object attribute tags to obtain the characteristics of the object attribute tags in terms of body shape, clothing, etc., in the database, i.e., body posture attribute features. Then, based on the matched body posture attribute features, the configuration parameters of the features of the body posture model associated with the digital human (virtual voice object) in terms of face, appearance, body shape, etc., are adjusted to construct a new body posture model, i.e., the target body posture model. For example, after extracting the body posture attribute features of the object attribute tags of the target object A, it is determined that A's image is a child wearing blue sportswear, while the image of the digital human (virtual voice object) body posture model is an adult wearing a black suit. Then, the body posture model is adjusted according to the extracted body posture attribute features, so that the adjusted body posture model presents the image of "blue sportswear" and "child", resulting in a corresponding new body posture model, i.e., the target body posture model. It should be noted that the digital human is not a real object, but a virtual "animated character" displayed on the screen. Its corresponding target body posture model is to bring users a more realistic visual effect and improve the experience of people participating in voice interaction scenarios.
[0115] In some implementations, the object attribute tags include at least color tags and clothing tags, then step (103.3.1) may include:
[0116] (103.3.1.1) Determine the clothing characteristics of the target object based on the clothing label;
[0117] (103.3.1.2) Based on feature similarity, determine the target clothing features that are similar to the clothing features of the target object from the database;
[0118] (103.3.1.3) Adjust the color feature value of the target clothing feature according to the color label to obtain the body shape attribute feature.
[0119] Among them, clothing tags can be the clothing categories and specific elements displayed or associated with the target object, which can include the types of clothes and clothing types, such as: suits, dresses, T-shirts, trousers, etc.
[0120] Among them, clothing features can be the feature points of the clothing tag corresponding to the target object. For example, the clothing features corresponding to a tailcoat are: two feature points: tailcoat feature and suit feature.
[0121] Feature similarity can be the degree of similarity between the current clothing tag of the target object and the clothing tags stored in the database in advance. For example, the feature similarity between a white crew neck T-shirt of the target object and a white crew neck T-shirt stored in the database is 100%.
[0122] In this context, "target clothing features" can refer to clothing features in the database that share similar characteristics with the clothing tags of the target object. Specifically, these characteristics can be identical or similar. For example, if the target object's clothing features include "suit" and "trousers," then the target clothing features can be clothing features that contain similar characteristics to "suit" and "trousers."
[0123] The color tag can be the color of the clothing in the attribute tag of the target object, including the color category, color code, etc. For example, the color tag of a white suit is the white tag.
[0124] The color feature value can be the color code parameter of the selected target clothing feature, such as the color code parameter of the adjustable clothing configuration of the virtual voice object.
[0125] Specifically, after determining the object attribute tags of the target object, the clothing tags are extracted from the object attribute tags. Then, feature extraction is performed on the clothing tags. For example, if the complaint information of the target object states that the digital human should wear a jacket and trousers, the corresponding clothing features are the two feature points of jacket and trousers. Next, the database storing the clothing features of virtual voice objects is accessed, and the clothing features of the target object are compared to find the target clothing feature with the highest similarity in the database. This similar clothing feature is then used to configure the body posture model. In addition, the color tags in the object attribute tags are extracted to determine the color or color code of the target object's clothing. This color or color code is then adjusted to the color feature value corresponding to the target clothing feature, thus obtaining the body posture attribute features. For example, if the color tag of the shirt in the object attribute tags is white, then the color of the shirt in the target clothing feature is adjusted to white.
[0126] By using the above methods, the configuration parameters of the virtual voice object can be adjusted according to the voice attributes and object attribute tags to obtain the target virtual voice object, i.e., the target digital human. This makes the digital human image that provides voice interaction more specific and targeted, improves the diversity of machines in human-computer interaction, and enhances the user's experience in human-computer interaction scenarios.
[0127] 104. Obtain the voice information of the target object and perform voice interaction on the voice information through the target virtual voice object.
[0128] In this embodiment of the application, in order to make the dialogue of human-computer interaction more in line with the interaction scenario, the interactive content and direction of the machine can be determined by the voice information corresponding to the target object, thereby making the dialogue of human-computer interaction scenario more targeted and improving the user's experience in human-computer interaction scenario.
[0129] Among them, voice information can be text or audio information corresponding to the target object. For example, if the target object is a salesperson, then the corresponding text or audio information is sales-related speech.
[0130] Specifically, after obtaining the voice information of the target object, the voice information of the target object is analyzed to obtain the corresponding text or audio information of the target object. Then, the text or audio information is transmitted to the voice system corresponding to the target digital human, that is, the voice system of the target virtual voice object. After the target digital human reads the text or audio information, it interacts with the user through the human-computer interaction function.
[0131] In some implementations, the target object information associated with the target object can be queried based on the target object's target identifier, and semantic recognition can be performed to obtain the target object's voice attributes and object attribute tags. For example, step 104, "voice interaction with voice information through the target virtual voice object," may include:
[0132] (104.1) Based on semantic similarity, query the voice response text that matches the voice information;
[0133] (104.2) Perform speech conversion on the speech response text to obtain acoustic speech features;
[0134] (104.3) Based on the target virtual speech object, perform voice interaction based on acoustic speech features.
[0135] Semantic similarity can be defined as the degree of similarity between the semantic information in the target object's speech information and the semantic information in the database.
[0136] Among them, the voice response text can be the text with the highest semantic similarity that the server queries from the database (or from the scenario) based on the semantics of the text. Taking the training scenario as an example, the voice interaction text is the training content. By matching similarity, the randomness of the training content can be increased, achieving a flexible, vivid and interesting effect, and avoiding the training content from being too templated.
[0137] Specifically, after identifying the target's voice information, the semantic information of this voice information is analyzed to derive its keywords. Then, the similarity of this semantic information is compared with that of semantic information in the database. Voice information with higher semantic similarity is selected and saved in text form, i.e., voice response text. The matched voice interaction text is then converted into speech to obtain voice with acoustic features. For example, if the acoustic feature is Cantonese, the resulting voice is converted into Cantonese audio. Finally, the voice with acoustic features is transmitted to the target digital human (target virtual voice object) for human-computer interaction with the user.
[0138] Using the above methods, the corresponding voice interaction text can be obtained based on the voice information of the target object. After voice conversion, the target virtual voice object can perform voice interaction based on the voice information, making the dialogue in the human-computer interaction scenario more targeted and improving the user's experience in the human-computer interaction scenario.
[0139] By implementing any one or a combination of implementation methods in the embodiments of this application, voice interaction application scenarios can be realized.
[0140] As can be seen from the above, the embodiments of this application obtain the voice attributes and object attribute tags of the target object; identify the scene category in the current voice interaction scenario and select the virtual voice object corresponding to the scene category; adjust the configuration parameters of the virtual voice object according to the voice attributes and object attribute tags to obtain the target virtual voice object; obtain the voice information of the target object, and perform voice interaction on the voice information through the target virtual voice object. Therefore, this solution can first determine the voice attributes and object attribute tags of the target object, identify the current voice interaction scene category of the target object, select a virtual voice object that matches the scene category, then adjust the acoustic and visual configuration parameters of the virtual voice object through the voice attributes and object attribute tags to obtain the adjusted target virtual voice object, and finally, perform voice interaction on each sentence of voice information of the target object through the target virtual voice object; thus, the visual and acoustic configuration of the human-computer interaction scenario can be enriched according to the actual needs of the user, improving the fun of the human-computer interaction scenario and enhancing the user's experience in the human-computer interaction scenario.
[0141] Based on the method described in the above embodiments, the following examples will provide further detailed explanations.
[0142] This application uses a voice interaction device as an example to further describe the voice interaction method provided in this application. Wherein, Figure 3 This is a schematic flowchart of another step of the voice interaction method provided in the embodiments of this application. For ease of understanding, the embodiments of this application are combined with... Figure 3 Describe it.
[0143] In this embodiment, the description will focus on a voice interaction device, which can be integrated into a computer device such as a server. When the processor on the in-vehicle terminal executes the program instructions corresponding to the data transmission method, the specific flow of the voice interaction method is as follows:
[0144] 201. The server determines the object identifier of the target object and queries the interaction feedback information corresponding to the object identifier and the target interaction feedback information.
[0145] The target audience can be the current users participating in the voice interaction scenario or the users who have participated in the voice interaction scenario before, such as students who have received training or customer service representatives.
[0146] The object identifier can be the account identifier of the target object when logging into the voice interaction application (platform), representing the target object's online identity on the voice interaction platform. The object identifier can be an account, nickname, number, custom identifier, etc.
[0147] The interactive feedback information can be information provided by the target object in response to the interactive scenario at a historical moment. This interactive feedback information can include the voice characteristics and image characteristics provided by the target object, reflecting the target object's experience with human-computer voice interaction. For example, the interactive feedback information includes the target object's experience with the "digital human," such as the "digital human's" pitch being too high, its timbre being too poor, its voice quality being too harsh, and its image being too old; or the "digital human's" timbre value not being in the range of (12, 26), its image not resembling that of a 7-12 year old, its clothing color not being red, and its clothing style not resembling a school uniform, etc.
[0148] In this context, target feedback semantics can be semantics composed of keywords in the interactive feedback information, reflecting the voice effect and image of the "digital human" desired by the target object. For example, if the interactive feedback information is "The image of the digital human is not suitable for 7-12 year olds, the pitch value is not AB, and the clothing style is not a school uniform," then after semantic recognition, the target feedback semantics can be obtained as "The image design of the digital human is for 7-12 year olds, the clothing style is a school uniform, and the pitch value is AB." Then, the subsequent voice attribute is "the pitch value is AB," and the object attribute label is "the image design is for 7-12 year olds, and the clothing style is a school uniform."
[0149] Specifically, after identifying the target object, the next step is to determine its corresponding object identifier. This object identifier is used to query the database for information provided by the associated target object and then return it as interactive feedback information. For example, if a salesperson is identified as the target object, their data in the database is associated with the identifier 001. Therefore, using the object identifier 001, the system can find related data in the user database, such as suggestions or complaints made by the salesperson during historical voice interactions. Further, the interactive system performs semantic recognition on the returned interactive feedback information to obtain the processed target semantic information, i.e., the target feedback semantics. Based on the information in the target feedback semantics, the system identifies the target object's corresponding voice attributes and object attribute tags. For instance, if the salesperson's object identifier is 001, the voice interaction system uses this identifier to query the database for associated complaint information. Then, it performs semantic recognition on this complaint information to obtain its content, such as the digital human's voice characteristics, speaking habits, and keyword extraction, thus deriving the target feedback semantics.
[0150] 202. The server determines the voice attributes and object attribute tags of the target object based on the interactive feedback information.
[0151] Among these, speech attributes can be various properties contained in the sound emitted by the target object, including but not limited to pitch, intensity, duration, and timbre. It is worth noting that different target objects possess different speech attributes.
[0152] Among them, object attribute tags can be object information of the target object, including but not limited to a series of tag information that can reflect the user's characteristics, such as the user's age, occupation, clothing, clothing style, color, etc.
[0153] Specifically, the target object's current voice can be collected and its voice attributes can be obtained through speech spectrum analysis. At the same time, images containing the target object can be collected and its object attribute labels can be analyzed through image information. Specifically, the object attribute labels of the target object can be evaluated through a trained neural network model. For example, the model can predict and evaluate the target object's age range, clothing style, clothing color, etc. In this way, the configuration parameters of the digital human in the current voice interaction scenario can be configured based on the voice attributes and object attribute labels obtained from the analysis.
[0154] It should be noted that when analyzing speech, the pitch, intensity, duration, and timbre of the target speech can also be analyzed. It can also be used to determine the speech speed, whether the speech characteristics are regional, and whether there are other speech characteristics compared to standard speech, thereby determining the speech attributes of the target speech.
[0155] For example, let user A be the target object. User A is from Northeast China and has a distinct Northeast regional accent. User A also speaks quickly. Therefore, user A's voice attributes include the two attributes of Northeast regional accent and fast speaking speed. Through a user survey of user A, user A is 24 years old and works as a salesperson. Therefore, the corresponding object attribute tags for Xiaoming include the three object tags of name, age, and occupation.
[0156] 203. The server obtains the selected voice interaction text of the target object and recognizes the corresponding semantic information.
[0157] Among them, voice interaction text can be the interactive content text needed in the human-computer interaction scenario between the target object and the digital human. For example, in the scenario of sales skills training, voice interaction text is the text of daily sales conversations.
[0158] Semantic information can be information in the relevant format of the voice interaction text, used to concisely express the meaning of related dialogue questions through keywords. For example, if the dialogue question is "Sir, do you need insurance, medical insurance or pension insurance?", the semantic information could be "Insurance = Pension insurance or medical insurance". The above is just an example and is not limited here.
[0159] Specifically, after acquiring the target object, the voice interaction system determines the target object information. Then, based on this information, it determines the voice interaction text. For example, if the target object is a salesperson, the corresponding voice interaction text is a sales script. Next, based on the determined voice interaction text, semantic recognition is performed. By using the contextual information within the text, the specific semantics of the text are determined, i.e., semantic information. For example, if the voice interaction text is a sales script, semantic recognition will reveal that the sales script is an insurance sales presentation practice text.
[0160] 204. The server determines the scene category in the current voice interaction scenario based on semantic information and selects the virtual voice object corresponding to the scene category.
[0161] The scenario category can be the specific scenario required for users to use digital humans for human-computer interaction, such as oral practice training for primary school students, customer service skills training, sales skills training, and interview skills training for news reporters.
[0162] Among them, virtual voice objects can be digital human figures created using digital technology that are close to human figures, or they can be virtual humans with human forms simulated by computers, i.e., digital humans.
[0163] Specifically, the voice interaction system identifies the category of the voice interaction scenario based on the context corresponding to the current semantic information. For example, in the current customer service training scenario, the identified scenario type is a training scenario; in a buyer-seller dialogue scenario, the identified scenario type is a dialogue scenario established from the perspectives of both the buyer and seller. After identifying the scenario category in the current voice interaction scenario, a computer-simulated virtual human-like figure, i.e., a digital human (virtual voice object), is selected corresponding to that scenario category. For instance, if the current voice interaction scenario category is a training scenario, then the digital human for that scenario is the corresponding training instructor figure.
[0164] 205. The server adjusts the sound configuration parameters of the voice interaction model associated with the virtual voice object according to the voice attributes of the target object, and obtains the target voice interaction model.
[0165] The configuration parameters can be virtual parameters corresponding to a complete virtual voice object, such as the height and weight parameters, facial parameters, voice parameters, clothing parameters, and style parameters of the virtual voice object.
[0166] The voice interaction model can be a vocalization model for a digital human (virtual voice object) to interact with the user. Virtual voice objects can use this model to interact with users, making human-computer interaction simpler and more natural.
[0167] The target voice interaction model can be a new voice interaction model obtained by adjusting the configuration parameters of the voice interaction model of the virtual voice object, and the target voice interaction model conforms to the use of the target virtual voice object in the current scene category.
[0168] Specifically, after obtaining the speech attributes of the target object, the speech attributes are analyzed to obtain the values of pitch, timbre, duration, and intensity, i.e., acoustic feature values. These acoustic feature values can be physical quantities representing the acoustic characteristics of speech, including values of pitch, timbre, duration, and intensity. For example, timbre represents the energy concentration area, formant frequency, formant intensity, and bandwidth, while prosodic characteristics include duration, fundamental frequency, and average speech power. Then, based on the extracted acoustic feature values, the configuration parameters of pitch, timbre, duration, and intensity of the speech interaction model associated with the digital human (virtual speech object) are adjusted to construct a new speech interaction model, i.e., the target speech interaction model. For example, after extracting acoustic feature values from the voice attributes in the complaint information of target object A, it is determined that the formant frequency corresponding to the digital human's timbre should be 50 Hz. However, the current timbre configuration parameter of the digital human (virtual voice object) voice interaction model is 30 Hz. Therefore, the original formant frequency of 30 Hz is adjusted to 30 Hz to obtain a corresponding new voice interaction model, namely the target voice interaction model.
[0169] 206. The server adjusts the configuration parameters of the body model associated with the virtual voice object based on the object attribute tags of the target object to obtain the target body model.
[0170] Among them, the body model can be a digital human image model in which the virtual voice object appears in a visualized three-dimensional form. By adjusting the body model, the image features such as the virtual voice object's body, clothing, and expression can be adjusted.
[0171] The target body model can be a new body model obtained by adjusting the configuration parameters of the body model of the virtual voice object, representing the new image features of the digital human after configuration.
[0172] Specifically, after obtaining the object attribute tags of the target object, the system then queries the database for body posture attributes that match the object attribute tags. These body posture attributes can be features exhibited by the target object's face, appearance, body shape, etc., such as a person's height, weight, and build, the color of clothing (yellow, white, black, red), and clothing style. Then, based on the matched body posture attributes, the configuration parameters of the facial, appearance, body shape, and other features of the body posture model associated with the digital human (virtual voice object) are adjusted to construct a new body posture model, i.e., the target body posture model. For example, after extracting the body posture features from the object attribute tags of target object A, we determine that A's image is a child wearing blue sportswear, while the body posture model of the digital human (virtual voice object) is an adult wearing a black suit. The body posture model is then adjusted according to the extracted body posture features, so that the adjusted body posture model presents the image of "blue sportswear" and "child," resulting in a corresponding new body posture model, i.e., the target body posture model. It should be noted that the digital human is not a real object, but a virtual "animated character" displayed on a screen. Its corresponding target body posture model is designed to provide users with a more realistic visual effect and improve the experience of those participating in voice interaction scenarios.
[0173] In addition, object attribute tags also include color tags and clothing tags. Among them, the color tag can be the color of the clothing in the corresponding object attribute tag of the target object, including the color category, color code, etc. For example, the color tag of a white suit is the white tag; the clothing tag can be the clothing category and specific elements displayed or associated with the target object, which can include the type of clothing and the type of clothing, such as: suit, dress, T-shirt, trousers, etc.
[0174] Specifically, after determining the object attribute tags of the target object, the clothing tags are extracted from the object attribute tags. Then, feature extraction is performed on the clothing tags. For example, if the complaint information of the target object states that the digital human should wear a jacket and trousers, the corresponding clothing features are the two feature points of jacket and trousers. Next, the database storing the clothing features of virtual voice objects is accessed, and the clothing features of the target object are compared to find the target clothing feature with the highest similarity in the database. This similar clothing feature is then used to configure the body posture model. In addition, the color tags in the object attribute tags are extracted to determine the color or color code of the target object's clothing. This color or color code is then adjusted to the color feature value corresponding to the target clothing feature, thus obtaining the body posture attribute features. For example, if the color tag of the shirt in the object attribute tags is white, then the color of the shirt in the target clothing feature is adjusted to white.
[0175] 207. The server determines the target virtual voice object based on the target voice interaction model and the target body posture model.
[0176] The target virtual voice object can be a virtual voice object (digital human) obtained by adjusting the configuration parameters of a virtual voice object. For example, a digital human on the screen who is male, wearing glasses and a black suit, and speaking Cantonese.
[0177] Specifically, after determining the target object's language attributes and object attribute tags, the configuration parameters of the virtual voice object are adjusted based on these attributes to obtain the target voice interaction model and the target body model, which together form a new digital human, i.e., the target virtual voice object. For example, the current target object has three voice attributes: a Northeastern accent, a relatively high-pitched voice, and a speech rate of 200 words per minute. The virtual voice object's voice configuration parameters are adjusted based on these three attributes. Then, based on the target object's three object attribute tags—salesperson, black suit—the digital human's tag configuration parameters are adjusted to obtain a sales-type digital human with a Northeastern accent, a relatively high-pitched voice, a speech rate of 200 words per minute, and wearing a black suit.
[0178] 208. The server queries the corresponding voice reply text based on the target object's voice information.
[0179] Among them, voice information can be text or audio information corresponding to the target object. For example, if the target object is a salesperson, then the corresponding text or audio information is sales-related speech.
[0180] Among them, the voice response text can be the text with the highest semantic similarity that the server queries from the database (or from the scenario) based on the semantics of the text. Taking the training scenario as an example, the voice interaction text is the training content. By matching similarity, the randomness of the training content can be increased, achieving a flexible, vivid and interesting effect, and avoiding the training content from being too templated.
[0181] Specifically, after identifying the target's voice information, the semantic information of the voice information is analyzed to obtain the keywords of the voice information. Then, the similarity of the semantic information is compared with the semantic information in the database. Voice information with higher semantic similarity is selected and saved in text form, i.e., voice reply text.
[0182] 209. The server performs speech conversion on the voice reply text to obtain acoustic speech features, and uses the target virtual voice object to perform voice interaction based on the acoustic speech features.
[0183] Specifically, after acquiring the voice information of the target object, the voice information is analyzed to obtain the corresponding text, i.e., the voice response text. The matched voice interaction text is then converted into speech to obtain speech with acoustic features; for example, if the acoustic feature is Cantonese, the resulting speech is also converted into Cantonese audio. Finally, the speech with acoustic features is transmitted to the target digital human (target virtual voice object) for human-computer interaction with the user.
[0184] Through the above application scenario examples, the following effects can be achieved: the visual and acoustic configurations of human-computer interaction scenarios can be enriched according to the actual needs of users, thereby increasing the fun of human-computer interaction scenarios and improving the user's experience in human-computer interaction scenarios.
[0185] To better implement the above methods, this application also provides a voice interaction device that can be integrated into computer equipment, such as in-vehicle terminals.
[0186] For example, such as Figure 4 As shown, the voice interaction device may include an acquisition unit 301, a selection unit 302, an adjustment unit 303, and an interaction unit 304.
[0187] Acquisition unit 301 is used to acquire the voice attributes and object attribute tags of the target object;
[0188] The selection unit 302 is used to identify the scene category in the current voice interaction scenario and select the virtual voice object corresponding to the scene category;
[0189] The adjustment unit 303 is used to adjust the configuration parameters of the virtual voice object according to the voice attributes and object attribute tags to obtain the target virtual voice object;
[0190] The interaction unit 304 is used to acquire the voice information of the target object and to perform voice interaction on the voice information through the target virtual voice object.
[0191] In some embodiments, the adjustment unit 303 is further configured to: determine the voice interaction model and body model associated with the virtual voice object; adjust the sound configuration parameters of the voice interaction model associated with the virtual voice object according to the voice attributes to obtain the target voice interaction model; adjust the image configuration parameters of the body model according to the object attribute tags to obtain the target body model; and construct the target virtual voice object according to the target voice interaction model and the target body model.
[0192] In some embodiments, the adjustment unit 303 is further configured to: extract acoustic feature values corresponding to the speech attributes; and adjust the sound configuration parameters of the speech interaction model associated with the virtual speech object according to the speech attributes to obtain the target speech interaction model.
[0193] In some embodiments, the adjustment unit 303 is further configured to: query body posture attribute features that match the object attribute tags from the database; adjust the body posture configuration parameters of the body posture model according to the object attribute features to obtain the target body posture model.
[0194] In some embodiments, the adjustment unit 303 is further configured to: determine the clothing features of the target object based on the clothing label; determine the target clothing features similar to the clothing features of the target object from the database based on feature similarity; and adjust the color feature value of the target clothing features according to the color difference label to obtain the body shape attribute features.
[0195] In some embodiments, the selection unit 302 is further configured to: acquire the voice interaction text selected by the target object; identify the semantic information corresponding to the voice interaction text; and determine the scene category in the current voice interaction scene based on the semantic information.
[0196] In some embodiments, the interaction unit 304 is further configured to: query voice interaction text that matches the voice information based on semantic similarity; perform voice conversion on the voice interaction text to obtain acoustic voice features; and perform voice interaction based on the target virtual voice object and targeting the acoustic voice features.
[0197] In some embodiments, the acquisition unit 301 is further configured to: determine the object identifier of the target object and query the interactive feedback information associated with the object identifier; perform semantic recognition on the interactive feedback information to obtain the target feedback semantics; and determine the voice attribute and object attribute label based on the target feedback semantics.
[0198] As can be seen from the above, the embodiments of this application can first determine the voice attributes and object attribute tags of the target object, and identify the current voice interaction scenario category of the target object, so as to select a virtual voice object that matches the scenario category. Then, the acoustic and visual configuration parameters of the virtual voice object are adjusted through the voice attributes and object attribute tags to obtain the adjusted target virtual voice object. Finally, for each voice information of the target object, voice interaction can be performed through the target virtual voice object. In this way, the visual and acoustic configuration of the human-computer interaction scenario can be enriched according to the actual needs of the user, thereby improving the fun of the human-computer interaction scenario and enhancing the user's experience in the human-computer interaction scenario.
[0199] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0200] This application also provides a computer device, such as... Figure 5 As shown, it illustrates a schematic diagram of the computer device involved in the embodiments of this application, specifically:
[0201] The computer device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, and an input unit 404. Those skilled in the art will understand that... Figure 5 The computer device structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:
[0202] The processor 401 is the control center of the computer device. It connects various parts of the computer device via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in the memory 402, and by calling data stored in the memory 402, thereby providing overall monitoring of the computer device. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 401.
[0203] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and voice interactions by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.
[0204] The computer device also includes a power supply 403 that supplies power to the various components. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 403 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0205] The computer device may also include an input unit 404, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0206] Although not shown, the computer device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the computer device loads the executable files corresponding to the processes of one or more applications into the memory 402 according to the following instructions, and the processor 401 runs the applications stored in the memory 402 to realize various functions, as follows:
[0207] Obtain the voice attributes and object attribute tags of the target object; identify the scene category in the current voice interaction scenario and select the virtual voice object corresponding to the scene category; adjust the configuration parameters of the virtual voice object according to the voice attributes and object attribute tags to obtain the target virtual voice object; obtain the voice information of the target object and perform voice interaction based on the voice information through the target virtual voice object.
[0208] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0209] As can be seen from the above, the embodiments of this application can first determine the voice attributes and object attribute tags of the target object, and identify the current voice interaction scenario category of the target object, so as to select a virtual voice object that matches the scenario category. Then, the acoustic and visual configuration parameters of the virtual voice object are adjusted through the voice attributes and object attribute tags to obtain the adjusted target virtual voice object. Finally, for each voice information of the target object, voice interaction can be performed through the target virtual voice object. In this way, the visual and acoustic configuration of the human-computer interaction scenario can be enriched according to the actual needs of the user, thereby improving the fun of the human-computer interaction scenario and enhancing the user's experience in the human-computer interaction scenario.
[0210] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0211] Therefore, embodiments of this application provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the voice interaction methods provided in embodiments of this application. For example, the instructions can execute the following steps:
[0212] Obtain the voice attributes and object attribute tags of the target object; identify the scene category in the current voice interaction scenario and select the virtual voice object corresponding to the scene category; adjust the configuration parameters of the virtual voice object according to the voice attributes and object attribute tags to obtain the target virtual voice object; obtain the voice information of the target object and perform voice interaction based on the voice information through the target virtual voice object.
[0213] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0214] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0215] This application also provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the voice interaction methods provided in the various optional implementations of the above embodiments.
[0216] Since the instructions stored in the computer-readable storage medium can execute the steps of any of the voice interaction methods provided in the embodiments of this application, the beneficial effects that any of the voice interaction methods provided in the embodiments of this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.
[0217] The above provides a detailed description of a voice interaction method, apparatus, device, and computer-readable storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A voice interaction method, characterized in that, include: Retrieve the voice attributes and object attribute tags of the target object; Identify the scene category in the current voice interaction scenario and select the virtual voice object corresponding to the scene category, wherein the scene category is a dialogue scenario; The configuration parameters of the virtual voice object are adjusted according to the voice attributes and object attribute tags to obtain the target virtual voice object; the voice interaction model and body model associated with the virtual voice object are determined; the voice configuration parameters of the voice interaction model associated with the virtual voice object are adjusted according to the voice attributes to obtain the target voice interaction model; the object attribute tags include at least color tags and clothing tags, and the clothing features of the target object are determined according to the clothing tags; Based on feature similarity, target clothing features that are similar to the clothing features of the target object are determined from the database; Adjust the color feature values of the target clothing feature according to the color label to obtain the body posture attribute features; adjust the body posture configuration parameters of the body posture model according to the body posture attribute features to obtain the target body posture model; construct the target virtual voice object according to the target voice interaction model and the target body posture model; The voice information of the target object is obtained, and voice interaction is performed on the voice information through the target virtual voice object.
2. The method according to claim 1, characterized in that, The step of adjusting the sound configuration parameters of the voice interaction model associated with the virtual voice object according to the voice attributes to obtain the target voice interaction model includes: Extract the acoustic feature values corresponding to the speech attributes; Adjust the sound configuration parameters of the voice interaction model associated with the virtual voice object according to the voice attributes to obtain the target voice interaction model.
3. The method according to claim 1, characterized in that, The identification of scene categories in the current voice interaction scenario includes: Obtain the selected voice interaction text of the target object; Identify the semantic information corresponding to the voice interaction text; The scene category in the current voice interaction scenario is determined based on the semantic information.
4. The method according to claim 1, characterized in that, The step of interacting with the voice information through the target virtual voice object includes: Based on semantic similarity, query the voice response text that matches the voice information; The voice response text is converted into speech to obtain acoustic speech features; Based on the target virtual voice object, voice interaction is performed based on the acoustic voice features.
5. The method according to claim 1, characterized in that, The process of obtaining the voice attributes and object attribute tags of the target object includes: Determine the object identifier of the target object and query the interactive feedback information associated with the object identifier; Semantic recognition is performed on the interactive feedback information to obtain the target feedback semantics; Based on the target feedback semantics, determine the voice attribute and object attribute labels.
6. A voice interaction device, characterized in that, include: The acquisition unit is used to acquire the voice attributes and object attribute tags of the target object; The selection unit is used to identify the scene category in the current voice interaction scenario and select the virtual voice object corresponding to the scene category, wherein the scene category is a dialogue scenario; An adjustment unit is used to adjust the configuration parameters of the virtual voice object according to the voice attributes and object attribute tags to obtain a target virtual voice object: determine the voice interaction model and body model associated with the virtual voice object; adjust the sound configuration parameters of the voice interaction model associated with the virtual voice object according to the voice attributes to obtain a target voice interaction model; the object attribute tags include at least color tags and clothing tags, and the clothing features of the target object are determined according to the clothing tags; Based on feature similarity, target clothing features that are similar to the clothing features of the target object are determined from the database; Adjust the color feature values of the target clothing feature according to the color label to obtain the body posture attribute features; adjust the body posture configuration parameters of the body posture model according to the body posture attribute features to obtain the target body posture model; construct the target virtual voice object according to the target voice interaction model and the target body posture model; An interaction unit is used to acquire the voice information of the target object and perform voice interaction on the voice information through the target virtual voice object.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium is computer-readable and stores a plurality of instructions adapted for loading by a processor to perform the steps of the voice interaction method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Virtual image voice interaction method and device, projection equipment and computer medium
CN113436602A
Information display method and device, equipment, medium and product
CN114092669A