Equipment interaction method and device, electronic equipment and storage medium
By configuring the target role shape for intelligent interactive devices and using the target language model to generate and play voice data, the problem of low interactiveness of existing AI interaction methods is solved, the interaction efficiency and interactivity are improved, and the value of artificial intelligence technology is fully utilized.
Patent Information
- Application Number
- CN202510125269.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-30
AI Technical Summary
Among the existing interaction methods based on AI technology, the interactivity between AI products and users is low, resulting in low device interaction efficiency and inability to fully utilize the value of artificial intelligence technology.
By configuring the appearance of the intelligent interactive device to a specific target role, using the target language model corresponding to the target role, the voice data of the target role for the user interaction trigger event is generated based on the event information of the user interaction trigger event, and the voice data is played through the intelligent interactive device.
It effectively improves the interaction effect between users and intelligent interactive devices, improves the interactivity between artificial intelligence products and users, and further exerts the value of artificial intelligence technology.
Smart Images

Figure CN120066259A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and particularly relates to a device interaction method, apparatus, electronic device, and storage medium. Background Art
[0002] With the rapid development of Internet technology, artificial intelligence (AI) technology has also been widely applied. Currently, most of the scenarios applying AI technology in the market interact with users through computers and mobile phone software, focusing on the office field or as a replacement for mobile phone functions.
[0003] In the process of researching and practicing the existing technology, it is found that in the existing interaction methods based on AI technology, the interactivity between AI products and users is still relatively low, and the value that AI technology can bring cannot be fully exerted, resulting in low device interaction efficiency. Summary of the Invention
[0004] Embodiments of the present application provide a device interaction method, apparatus, electronic device, and storage medium, which can effectively improve the interaction effect between users and intelligent interaction devices, thereby enhancing the interactivity between artificial intelligence products and users, and further exerting the value that artificial intelligence technology can bring.
[0005] Embodiments of the present application provide a device interaction method, including:
[0006] In response to a user interaction trigger event for an intelligent interaction device, obtain event information of the user interaction trigger event, and the shape of the intelligent interaction device is configured as a specific target role;
[0007] Based on the event information, generate voice data of the target role for the user interaction trigger event through the target language model corresponding to the target role, and the target language model is trained based on the role information and conversation information of the target role;
[0008] Play the voice data through the intelligent interaction device based on the role attribute information of the target role.
[0009] Correspondingly, embodiments of the present application further provide a device interaction apparatus, including:
[0010] An obtaining unit, configured to obtain event information of a user interaction trigger event in response to the user interaction trigger event for the intelligent interaction device, and the shape of the intelligent interaction device is configured as a specific target role;
[0011] A generation unit, configured to generate, by using a target language model corresponding to the target role, speech data of the target role for the user interaction trigger event based on the event information, where the target language model is trained based on role information and conversation information of the target role;
[0012] A playback unit, configured to play the speech data through the intelligent interaction device based on role attribute information of the target role.
[0013] In addition, an embodiment of the present application further provides an electronic device, including a processor and a memory. The memory stores a computer program. When the computer program is executed by the processor, the processor is caused to execute the steps of any one of the device interaction methods provided by the embodiments of the present application.
[0014] In addition, an embodiment of the present application further provides a computer-readable storage medium, including a computer program. When the computer program runs on an electronic device, the computer program is configured to cause the electronic device to execute the steps of any one of the device interaction methods provided by the embodiments of the present application.
[0015] In addition, an embodiment of the present application further provides a computer program product, including a computer program. The computer program is stored in a computer-readable storage medium. When a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device executes the steps of any one of the device interaction methods provided by the embodiments of the present application.
[0016] In the embodiments of the present application, in response to a user interaction trigger event for an intelligent interaction device, event information of the user interaction trigger event is obtained, and the appearance of the intelligent interaction device is configured as a specific target role; through the target language model corresponding to the target role, speech data of the target role for the user interaction trigger event is generated based on the event information, where the target language model is trained based on role information and conversation information of the target role; through the intelligent interaction device, the speech data is played based on role attribute information of the target role. In this way, through the intelligent interaction device whose appearance is configured as a specific target role, event information of the user interaction trigger event is obtained, so that through the target language model corresponding to the target role, speech data of the target role for the user interaction trigger event is generated, and in this way, through the intelligent interaction device, the speech data is played based on role attribute information of the target role, which can realize replying to the event triggered by the user through the target role, effectively improve the interaction effect between the user and the intelligent interaction device, thereby improving the interactivity between the artificial intelligence product and the user, and further effectively bringing into play the value that the artificial intelligence technology can bring. Description of the Drawings
[0017] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0018] Figure 1 It is a schematic diagram of an implementation scenario of a device interaction method provided in an embodiment of the present application;
[0019] Figure 2 It is a schematic flowchart of a device interaction method provided in an embodiment of the present application;
[0020] Figure 3a It is a schematic diagram of an intelligent interaction device of a device interaction method provided in an embodiment of the present application;
[0021] Figure 3b It is a schematic diagram of a specific framework of a device interaction method provided in an embodiment of the present application;
[0022] Figure 3c It is another schematic diagram of a specific framework of a device interaction method provided in an embodiment of the present application;
[0023] Figure 3d It is a schematic diagram of the overall process of a device interaction method provided in an embodiment of the present application;
[0024] Figure 3e It is a schematic diagram of the base structure of a device interaction method provided in an embodiment of the present application;
[0025] Figure 3f It is another schematic diagram of the base structure of a device interaction method provided in an embodiment of the present application;
[0026] Figure 3g It is a schematic diagram of the light effect setting of a device interaction method provided in an embodiment of the present application;
[0027] Figure 3h It is another schematic diagram of the base structure of a device interaction method provided in an embodiment of the present application;
[0028] Figure 3i It is a schematic diagram of the base design of a device interaction method provided in an embodiment of the present application;
[0029] Figure 3j It is another schematic diagram of the base design of a device interaction method provided in an embodiment of the present application;
[0030] Figure 3k It is another schematic diagram of the base design of a device interaction method provided in an embodiment of the present application;
[0031] Figure 4 It is a schematic structural diagram of the device interaction device provided in the embodiment of the present application;
[0032] Figure 5 It is a schematic structural diagram of the electronic device provided in the embodiment of the present application. Specific Embodiments
[0033] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0034] At the same time, in the description of the embodiments of the present application, terms such as "first" and "second" are only used for distinguishing descriptions, and cannot be understood as indicating or implying relative importance. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of the present application, "a plurality" means two or more, unless otherwise specifically defined.
[0035] The embodiment of the present application provides a device interaction method, device, electronic device, and storage medium. Among them, the device interaction device can be integrated in the electronic device, and the electronic device can be a server or a terminal device, etc.
[0036] Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, network acceleration services (Content Delivery Network, CDN), and big data and artificial intelligence platforms. The terminal can include, but is not limited to, action figures, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle terminals, aircraft, etc. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any restrictions here.
[0037] Please refer to Figure 1 , taking the device interaction device integrated in the electronic device as an example, Figure 1Schematic diagram of an implementation scenario of the device interaction method provided by an embodiment of this application. Among them, the electronic device can be a terminal. The electronic device can respond to a user interaction trigger event for an intelligent interaction device, obtain event information of the user interaction trigger event, and the appearance of the intelligent interaction device is configured as a specific target role; through a target language model corresponding to the target role, voice data of the target role for the user interaction trigger event is generated based on the event information, and the target language model is trained based on the role information and conversation information of the target role; through the intelligent interaction device, the voice data is played based on the role attribute information of the target role.
[0038] It should be noted that Figure 1 The schematic diagram of the implementation environment scenario of the device interaction method shown is only an example. The implementation environment scenario of the device interaction method described in the embodiments of this application is to more clearly illustrate the technical solutions of the embodiments of this application, and does not constitute a limitation on the technical solutions provided by the embodiments of this application. Those of ordinary skill in the art know that with the evolution of data processing and the emergence of new business scenarios, the technical solutions provided by this application are equally applicable to similar technical problems.
[0039] The solution provided by the embodiments of this application is specifically illustrated by the following embodiments. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.
[0040] This embodiment will be described from the perspective of a device interaction device. The device interaction device can be specifically integrated in an electronic device, and the electronic device can be a terminal, which is not limited in this application.
[0041] Please refer to Figure 2 , Figure 2 which is a flowchart of the device interaction method provided by an embodiment of this application. The device interaction method can be applied to a target device, and the method includes:
[0042] In step 101, in response to a user interaction trigger event for an intelligent interaction device, event information of the user interaction trigger event is obtained.
[0043] Among them, the appearance of the intelligent interaction device is configured as a specific target role.
[0044] Among them, the intelligent interaction device can be a device that interacts with users based on artificial intelligence technology. For example, the intelligent interaction device can be a figurine, that is, an intelligent voice interaction figurine based on artificial intelligence (abbreviated as AI). The appearance of the intelligent interaction device can be configured as a specific target character, which can be a virtual character or an authorized real person. For example, the target character can include characters such as heroes, anime characters, robots, and virtual characters. For example, please refer to Figure 3a , Figure 3a which is a schematic diagram of an intelligent interaction device for a device interaction method provided in an embodiment of this application. The intelligent interaction device can be configured as a virtual character with a robot appearance.
[0045] The user interaction trigger event can be an event for triggering the interaction between the intelligent interaction device and the user. For example, the user interaction trigger event can include the user's voice for initiating interaction with the intelligent interaction device, or can include the user's interaction actions with the intelligent device, and can also include events such as the triggering of a preset reminder event or a wake-up event.
[0046] The event information can be information generated based on the user interaction trigger event. For example, when the user interaction trigger event is the user's voice for initiating interaction with the intelligent interaction device, the event information can be the voice data emitted by the user. When the user interaction trigger event is the user's interaction actions with the intelligent device, the event information can be the action information of the user's interaction actions, and the action information can indicate the actions performed by the user. When the user interaction trigger event is the triggering of a preset reminder event, the event information can be the relevant information of the reminder event. For example, it can include the object related to the reminder event and other relevant information. The object can be a person, animal, plant, etc. related to the reminder event, and this application embodiment does not make any limitations here.
[0047] Optionally, the product layout for the intelligent interaction device can include two dimensions: horizontal and vertical. Vertically, it can include product themes such as cyberpunk, anime characters, and superheroes. Horizontally, products with different characters can be laid out under each theme, and product designs for different characters are carried out around the same theme. For example, taking the cyberpunk theme as an example, the design concept of the cyberpunk theme can be: in the current upsurge of the explosion of artificial intelligence AI, to think about the possible highest form carriers [robots] and [humans] after the highly evolved future AI, the relationship between [intelligence] and [humanity], the relationship between [confrontation] and [symbiosis], and what will the future technology + people + social structure look like? In this regard, the concept expression of this theme can be that the fusion expression of real human skin texture and mechanical mecha style on a carrier is a beautiful ideal wish: harmonious coexistence. Therefore, for products under the cyberpunk theme, their designs have the characteristics of a sense of brokenness in the overall structure + exposed mechanical style + technological lighting + robot voice.
[0048] In step 102, through the target language model corresponding to the target role, voice data for the target role in response to the user interaction trigger event is generated based on the event information.
[0049] Among them, the target language model is trained based on the role information and dialogue information of the target role. Specifically, it can be obtained by fine-tuning the basic model based on the artificial intelligence toolkit (OpenAI). For example, relevant information corresponding to the target role can be obtained as the training samples of the target language model. For example, it can include information such as the interview videos, articles, self-introductions, personal experiences, and family backgrounds of the target role as the training samples of the model to fine-tune the model and customize the tone, characteristics, vocabulary, personal background, life history, etc. of the large language model, so as to realize the dialogue of the target role through the target language model.
[0050] Among them, the target language model can be a large language model (Large Language Model, abbreviated as LLM), the voice data can be the data for the target role to reply to the user interaction trigger event, the role information can be the relevant information of the target role, for example, it can include information such as the life story, personality, background, worldviews, thinking patterns, and conversation tones of the target role. The dialogue information can be the conversation content between the target role and other people. In this way, training the target language model through the role information and dialogue information of the target role can enable the trained target language model to learn the personality and style of the target role, so that the reply of the target role to the user interaction trigger event can be simulated based on the target language model, thereby improving the interaction effect between the user and the intelligent interaction device.
[0051] Optionally, the intelligent interaction device can upload the event information to the cloud to call the target language model corresponding to the target role, and generate the voice data of the target role for the event triggered by the user interaction based on the input event information.
[0052] Among them, there are various ways to generate the voice data of the target role for the event triggered by the user interaction based on the event information through the target language model corresponding to the target role. For example, the event triggered by the user interaction may include the user uttering a voice. The step of obtaining the event information of the event triggered by the user interaction may include: obtaining the first voice data uttered by the user. Accordingly, the step of generating the voice data of the target role for the event triggered by the user interaction based on the event information through the target language model corresponding to the target role may include: based on the first voice data, generating, through the target language model corresponding to the target role, the voice data of the target role for replying to the first voice data.
[0053] Among them, the first voice data may be the data of the voice uttered by the user for the intelligent interaction device.
[0054] Among them, there are various ways to generate the voice data of the target role for replying to the first voice data based on the first voice data through the target language model corresponding to the target role. For example, the dialogue mode matching the first voice data may be identified; if the dialogue mode is the chat mode, semantic feature extraction is performed on the first voice data through the target language model corresponding to the target role to obtain semantic features; and based on the semantic features, the target language model generates the voice data of the target role for replying to the first voice data.
[0055] Among them, the dialogue mode may be the interaction mode between the intelligent interaction device and the user. For example, it may include the chat mode and the instruction mode. In the chat mode, the intelligent interaction device can have a free conversation with the user. In the instruction mode, the intelligent interaction device can respond based on the instruction triggered by the user. The semantic feature may be the information characterizing the semantics of the first voice data.
[0056] For example, in the chat mode, the user can initiate a free conversation, so that through the intelligent interaction device, based on the target language model corresponding to the target role, combined with the attributes such as the life story, background, three views and thinking mode of the corresponding target role and the tone of the role, a free conversation can be carried out with the user. For example, when the target role is the role of an entrepreneur, the user can initiate a conversation: "Hi XXX, why did you found XXX company?", that is, the first voice data. Accordingly, the intelligent interaction device can generate voice data based on the first voice data and play it: "Because I hope to contribute to the future living space of mankind and improve the risk resistance ability of all mankind. What do you think about the future living space of mankind?" and so on.
[0057] In one embodiment, the user can actively request the intelligent interaction device to summarize important events within a specific time period. For example, the first voice data can be "Summarize the major events that occurred in 2024", thereby triggering the intelligent interaction device to summarize and reply based on the conversation content with the user during 2024 and the important moments the user asked the intelligent interaction product to remember. Specifically, information extraction and analysis can be performed by including but not limited to using technologies such as vector databases, traditional databases, and large language models, so as to generate the voice data corresponding to the major events that occurred to the user in 2024 and play it.
[0058] Optionally, based on the first voice data, there can be various ways to generate the voice data for the target role to reply to the first voice data through the target language model corresponding to the target role. For example, the conversation mode matched by the first voice data can be identified; if the conversation mode is the instruction mode, the instruction content in the first voice data is extracted based on the target language model corresponding to the target role, an instruction reply voice is generated based on the instruction content, and the multimedia content indicated by the instruction content is obtained; based on the instruction reply voice and the multimedia content, the voice data for the target role to reply to the first voice data is generated.
[0059] Among them, the instruction mode can be a mode of instruction interaction. In this mode, the intelligent interaction product has AI instruction type functions (Agents). The instruction content can be the instruction obtained from the first voice data. The instruction reply voice can be the voice for reply generated based on the instruction content. The multimedia content can be content such as music, news, weather forecast, etc. The instruction can include playing music, news broadcast, querying weather, etc.
[0060] Among them, there can be various ways to generate the voice data for the target role to reply to the first voice data based on the instruction reply voice and the multimedia content. For example, the multimedia content can be converted into voice modal data, so that the voice data for the target role to reply to the first voice data can be obtained by combining the instruction reply voice.
[0061] For example, assume the user issues the voice: "Hi XXX, play a piece of light music", which is the first voice data. Thus, the intelligent interaction device can obtain the instruction content "play music" based on the first voice data, and thus, in response to the instruction content, generate an instruction reply voice "Playing for you, XXX" based on the instruction content, and obtain the light music type of music indicated by the instruction content. Thus, the intelligent interaction device can play the voice "Playing for you, XXX" and the corresponding music.
[0062] For another example, assume the user issues the voice command: "Hi XXX, what are the important news in the technology field today?", so that the intelligent interaction device can obtain the command content "Play news in the technology field" based on the voice issued by the user. Then, in response to the command content, the device can generate a command reply voice "10 major events in the technology field today have been collected for you", and obtain the audio of the news in the technology field indicated by the command content. Thus, the intelligent interaction device can play the voice "10 major events in the technology field today have been collected for you" and broadcast the corresponding news.
[0063] Optionally, there can be multiple ways to generate the voice data for the target role to reply to the first voice data based on the first voice data. For example, based on the first voice data, it can be determined whether the user is the target user bound to the intelligent interaction device; if the user is the target user, based on the historical interaction data between the target user and the intelligent interaction device and the first voice data, the voice data for the target role to reply to the first voice data is generated through the target language model corresponding to the target role; if not, based on the first voice data, the voice data for the target role to reply to the first voice data is generated through the target language model corresponding to the target role.
[0064] Among them, the target user can be a user bound to the intelligent interaction device. For example, it can be the owner of the intelligent interaction device or the main user, that is, the primary user. The historical interaction data can be the interaction data between the target user and the intelligent interaction device during the historical interaction process. For example, it can include conversation data, data of command triggering and response, etc.
[0065] In this way, when the user is the target user, based on the historical conversations between the target user and the intelligent interaction device in the past, the reply to the user can be generated more accurately through the target language model, thereby enhancing the interaction experience between the user and the intelligent interaction device and the intelligent interaction value of the intelligent interaction device.
[0066] Among them, there can be multiple ways to determine whether the user is the target user bound to the intelligent interaction device based on the first voice data. For example, voiceprint recognition can be performed on the first voice data to obtain the first voiceprint information; the second voiceprint information corresponding to the target user bound to the intelligent interaction device is obtained; based on the first voiceprint information and the second voiceprint information, it is determined whether the user is the target user bound to the intelligent interaction device.
[0067] Among them, the first voiceprint information can be the voiceprint corresponding to the first voice data, and the second voiceprint information can be the voiceprint corresponding to the target user.
[0068] Among them, there are various ways to determine whether the user is the target user bound to the intelligent interaction device based on the first voiceprint information and the second voiceprint information. For example, the first voiceprint information and the second voiceprint information can be compared. If the first voiceprint information and the second voiceprint information belong to the same person, it can be determined that the user is the target user bound to the intelligent interaction device. If the first voiceprint information and the second voiceprint information are too different and do not belong to the same person, it can be determined that the user is not the target user bound to the intelligent interaction device.
[0069] In one embodiment, sound source localization can be performed through the intelligent interaction device, so that the position of the user to be interacted with can be determined, and the intelligent interaction device can be controlled to face the device to be interacted with, improving the sound pickup effect and the interaction experience between the user and the intelligent interaction device. Specifically, the sound source direction of the user can be obtained by identifying the sound source azimuth of the first voice data; based on the sound source direction, the intelligent interaction device can be controlled to face the user.
[0070] Among them, the sound source direction can be the location of the user who emits the first voice data.
[0071] Among them, there are various ways to obtain the sound source direction of the user by identifying the sound source azimuth of the first voice data. For example, the sound source azimuth of the first voice data can be identified through a multi-microphone array to obtain the sound source direction of the user, so that the intelligent interaction device can turn to the sound source direction, that is, face the interacting user after being awakened.
[0072] Optionally, in order to obtain the best speech-to-text (STT) effect, the intelligent interaction device can perform noise reduction processing on the collected first voice data. At the same time, when the intelligent interaction product plays voice or music, it is necessary to ensure normal sound collection.
[0073] In one embodiment, the user interaction trigger event can include a preset reminder event. The step of obtaining the event information of the user interaction trigger event can include: obtaining at least one object involved in the reminder event and the historical memory information of the intelligent interaction device for the object; accordingly, the step of generating the voice data of the target role for the user interaction trigger event based on the event information through the target language model corresponding to the target role can include: generating the voice data of the target role for the user interaction trigger event based on the object and the historical memory information through the target language model corresponding to the target role.
[0074] Among them, the historical memory information can be obtained based on data related to the object in the historical interaction data between the user and the intelligent interaction device. The reminder event can be an interaction event actively initiated by the intelligent interaction device, and the reminder event can be an event permitted by the user. When the reminder event is triggered, the intelligent interaction device can be in a state of being awakened by the user, or it can be not limited to the awakened state. The reminder event can include anniversary reminders, birthday reminders, important event reminders, memo event reminders, weather change reminders, festival reminders, date reminders, and other events.
[0075] For example, the reminder event can include, at a specific time point, important events of the user recorded according to past chat content records. The intelligent interaction device can actively initiate a conversation with the user based on the reminder event. For example, the intelligent interaction device can obtain the event information corresponding to the reminder event of the third anniversary of object A, and thus can control the intelligent interaction device to perform a raising hand action, light up the corresponding light effect, and play the corresponding voice data: "According to the records of past chats, tomorrow is the third anniversary of your acquaintance with object A. Don't forget to give him a surprise."
[0076] In one embodiment, the user interaction trigger event can include the user generating an interaction action for the intelligent device. The step of obtaining the event information of the user interaction trigger event can include: obtaining the action information of the interaction action generated by the user for the intelligent device; the step of generating, by the target language model corresponding to the target role, the voice data of the target role for the user interaction trigger event based on the event information can include: generating, by the target language model corresponding to the target role, the voice data of the target role for the user interaction trigger event based on the action information.
[0077] Among them, the interaction action can be an action of interacting with the intelligent interaction device. For example, it can include actions such as touching the head, pressing, patting, waving, and making a heart gesture on the intelligent interaction device. The action information can be information indicating the interaction action. For example, it can be the action name of the interaction action or the action identifier (ID), etc.
[0078] Among them, there are various ways to generate the speech data of the target role for the user interaction trigger event based on the action information through the target language model corresponding to the target role. For example, assuming the interaction action is waving, which can indicate that the user is greeting the intelligent interaction device to start a chat, then based on the action information of waving, the speech data corresponding to the target role can be generated through the target language model, such as speech data like "What do you want from me, human", "Hello, I'm here", etc. At the same time, the intelligent interaction device can also be controlled to perform the waving action. Another example, assuming the interaction action is making a heart, which can indicate that the user is showing friendliness to the intelligent interaction device, then based on the action information of making a heart, speech data corresponding to the target role in different styles can be generated through the target language model, such as speech data like "Humph, go away", "Thank you, I like you too", etc. At the same time, the intelligent interaction device can also be controlled to perform the making-a-heart action, etc.
[0079] In one embodiment, there are also various ways to generate the speech data of the target role for replying to the first speech data based on the first speech data through the target language model corresponding to the target role. For example, the device status information of the intelligent interaction device can be detected; the target language model matching the device status information can be determined from multiple large language models of the target role; and the speech data of the target role for replying to the first speech data can be generated based on the first speech data through the target language model.
[0080] Among them, the device status information can be the information indicating the status of the intelligent interaction device, and the device status information includes at least one of the posture information and the device component assembly information of the intelligent interaction device. The posture information can be the role posture of the intelligent interaction device. For example, it can include postures such as the target role raising a hand, making a fist, putting hands on the hips, lowering the head, bending the waist, etc. The device component assembly information can be the component assembly of the intelligent interaction device. For example, it can include assemblies such as carried props, weapons, equipment, accessories, clothing, etc., and can also include states such as component integrity, component mutilation, and component aging. Component mutilation can mean that the target role corresponding to the intelligent interaction device has missing parts such as missing arms or missing legs. Component aging can include states such as the hair color changing to white and the equipment being old equipment.
[0081] Different device posture information can indicate different emotional states of the target character corresponding to the intelligent interactive device. For example, when the intelligent interactive device makes a fist, it can indicate that the target character corresponding to the intelligent interactive device is in a confident and excited emotional state. When the intelligent interactive device lowers its head, it can indicate that the target character corresponding to the intelligent interactive device is in a depressed and sad emotional state. When the intelligent interactive device carries props, it can indicate that the target character corresponding to the intelligent interactive device is in a tenacious and fearless emotional state. When the equipment carried by the intelligent interactive device is old equipment, it can indicate that the target character corresponding to the intelligent interactive device is tired and pessimistic after the battle. In addition, different device posture information can indicate that the target character corresponding to the intelligent interactive device is in different stages of life. For example, when the hair of the intelligent interactive device is black, it can indicate that the target character corresponding to the intelligent interactive device is in the stage of young and middle-aged. When the hair of the intelligent interactive device is white, it can indicate that the target character corresponding to the intelligent interactive device is in the stage of old age. For another example, when the components of the smart interactive device are complete, it may indicate that the target character corresponding to the smart interactive device is in an optimistic attitude towards life; when the components of the smart interactive device are incomplete, it may indicate that the target character corresponding to the smart interactive device is in a pessimistic attitude towards life, etc.
[0082] Correspondingly, multiple large language models can be trained for the target character based on different life attitudes, emotional states or life stages. Different large language models can simulate the conversation style of the target character under different life attitudes, emotional states or life stages. In this way, based on the device status information of the intelligent interactive device, the target language model that matches the emotional state, life stage or life attitude of the device status information can be screened out from the multiple large language models corresponding to the target character. Therefore, the voice data of the target character's reply to the first voice data in the corresponding state can be generated based on the first voice data through the matched target language model, so that the language data replied by the intelligent interactive device is more interesting and playable, thereby enhancing the interactive experience between the user and the intelligent interactive device.
[0083] In step 103, voice data is played based on the role attribute information of the target role through the intelligent interactive device.
[0084] The character attribute information may include the target character's voice, tone, intonation, personality, age, gender and other attributes. For example, the voice data may be played using the target character's voice through an intelligent interactive device. In this way, the interaction between the target character and the user may be simulated through the intelligent interactive device, thereby improving the user's interaction experience with the intelligent interactive device.
[0085] Optionally, before playing the voice data through the intelligent interaction device, the intelligent interaction device can be controlled to perform a voice output prompt operation corresponding to the raising hand mode.
[0086] Among them, the raising hand mode can be a mode for controlling the intelligent interaction device to raise its hand, and the voice output prompt operation can be an operation to prompt the user that the intelligent interaction device will output voice. The voice output prompt operation can include operations such as controlling the intelligent interaction device to raise its hand, lighting up the light effect, and emitting a sound effect. In this way, before the intelligent interaction device plays the voice data, the intelligent interaction device can be controlled to perform a voice output prompt operation corresponding to the raising hand mode, thereby prompting the user, improving the interactivity between the user and the intelligent interaction device, and further enhancing the interaction experience between the user and the artificial intelligence product.
[0087] For example, taking the intelligent interaction device as an action figure, please refer to Figure 3b , Figure 3b which is a specific framework schematic diagram of a device interaction method provided in an embodiment of the present application. The user can ask questions about the intelligent interaction device. The action figure can receive the sound, and perform local processing through the hardware of the intelligent interaction device. At the same time, it can make sound, light, and dynamic effect adjustments according to events to obtain the first voice data. Then, the first voice data can be uploaded to the cloud server. The cloud server performs speech-to-text (STT) processing on the first voice data. Then, through the model (LLM) fine-tuned based on the large language model corresponding to the target role, that is, the target language model, and the corresponding vector database or retrieval-augmented generation (RAG), a reply text corresponding to the text of the first voice data can be generated. Thus, the reply file can be subjected to text-to-speech (TTS) processing to obtain voice data and send it back to the action figure, and the voice data is played through the action figure hardware. At the same time, sound, light, and dynamic effect adjustments can be made according to the playing event.
[0088] Optionally, a large language model based on multi-modal can be used to directly generate the voice data corresponding to the first voice data. For example, please refer to Figure 3c , Figure 3c which is another specific framework schematic diagram of a device interaction method provided in an embodiment of the present application. The multi-modal AI large model corresponding to the target role can be used to output the voice data replied by the target role for the first voice data based on the input first voice data.
[0089] Optionally, when playing voice data through the intelligent interaction device, the intelligent interaction device can also be controlled to perform corresponding actions simultaneously to enhance the interactivity between the user and the intelligent interaction device. For example, the role gesture information matching the voice data can be determined based on the role attribute information of the target role; the motion control information of at least one part of the intelligent interaction device and / or the device control information of the sub-devices on the intelligent interaction device can be determined based on the role gesture information; when playing voice data through the intelligent interaction device, the intelligent interaction device is also controlled to move through the motion control information and / or the device control information.
[0090] Among them, the role gesture information can indicate the information of the role gesture of the intelligent interaction device. The role gesture can include gestures such as making a fist, putting hands on the hips, bending over, waving, and spinning around. The part can include parts such as the hands, legs, waist, and head of the intelligent interaction device. The sub-devices can include devices such as the lighting device, heating device, and ventilation device assembled on the intelligent interaction device. The motion control information can be the information for controlling the intelligent interaction device to move, and the device control information can be the information for controlling the sub-devices of the intelligent interaction device to operate.
[0091] For example, when the voice data is "I'm very sorry. I didn't notice your request", the role gesture information can be bending over, and the motion control information can be the information for controlling the intelligent interaction device to bend over, so that when the intelligent interaction device plays the voice data, the intelligent interaction device can be controlled to bend over simultaneously. Another example, when the voice data is "Okay, let's work together", the role gesture information can be making a fist, and the motion control information can be the information for controlling the intelligent interaction device to make a fist, so that when the intelligent interaction device plays the voice data, the intelligent interaction device can be controlled to make a fist, etc. The specific control rules can be set according to the actual situation, and the embodiments of the present application do not make any limitations here.
[0092] In one embodiment, the intelligent interaction device provided by the embodiments of the present application can be used as the total control terminal of the smart home device and can control the connected smart home devices based on the user's instructions. Specifically, the smart home devices to be controlled and the corresponding control parameters can be determined based on the event information, and then the smart home devices to be controlled can be controlled based on the control parameters.
[0093] For example, when the user generates the first voice data "Turn on the air conditioner and set the temperature to 26 degrees" for the intelligent interaction device, based on the first voice data, the intelligent interaction device can determine the smart home device "air conditioner" to be controlled, as well as the corresponding control parameters "turn on" and "26 degrees". Thus, it can play the voice data "Okay, right away", and at the same time, based on the control parameters, control the smart home device "air conditioner" to turn on and set the cooling temperature of the air conditioner to 26 degrees.
[0094] For another example, when the user generates the first voice data "turn on the lights in the living room" for the intelligent interaction device, based on the first voice data, the intelligent interaction device can determine the smart home device to be controlled, "the lights in the living room", and the corresponding control parameter, "turn on". Thus, it can play the voice data "Okay, turning on the lights in the living room", and at the same time, based on the control parameter, control the smart home device corresponding to the lights in the living room to turn on.
[0095] In one embodiment, the user can interrupt when the intelligent interaction device is playing voice data to more accurately and realistically simulate a real - person conversation scenario, further enhancing the interaction experience between the user and the intelligent interaction device. For example, when detecting that the user has an interruption behavior for the voice data during the playback of voice data by the intelligent interaction device, the second voice data generated by the user based on the interruption behavior can be obtained; based on the data already played in the voice data and the second voice data, through the target language model, the target voice data for the target role to reply to the second voice data can be generated; the target voice data is played through the intelligent interaction device.
[0096] Among them, the interruption behavior can be the behavior of interrupting the intelligent interaction device from playing voice data, and the second voice data can be the voice for interaction emitted by the user after interrupting the intelligent interaction device from playing voice data. The target voice data can be the data generated by the target language model based on the already played data and the second voice data for replying to the second language data.
[0097] For example, when the intelligent interaction device plays the streamed - back voice data locally through the speaker, if it is interrupted by the user during the voice playback, the voice playback can be immediately stopped and the next listening mode can be entered, so that the second voice data emitted by the user can be collected. Thus, the already played voice data and the second voice data of the user's continued inquiry can be transmitted back to the cloud together for the next round of conversation through the target language model. In this way, the interaction experience between the user and the intelligent interaction device can be improved.
[0098] In one embodiment, please refer to Figure 3d , Figure 3dIt is a schematic diagram of the overall process of a device interaction method provided in an embodiment of the present application. The intelligent interaction device can have the capabilities of free conversation and active questioning. For free conversation, when the intelligent interaction device is awakened by the user, it can determine whether the current user is the primary user, and thus execute different background paths. When the user asks a question, it can answer based on the voice of the user's question and enable corresponding motion effects, lighting effects, etc. When the user actively exits, the intelligent interaction device can wait for 10 seconds. If the user does not continue to ask questions, it can control the listening lighting effect to gradually fade out and thus enter the standby state. For active questioning, during non-disturbance time, it can retrieve the RAG user-specific database or databases such as AGent, or trigger reminder events through a background preset scheduler, so as to have an active conversation with the user.
[0099] In one embodiment, the intelligent interaction device can be rotated. For example, the rotation of the intelligent interaction device can be achieved through the design of the base of the intelligent interaction device. Specifically, please refer to Figure 3e , Figure 3e It is a schematic diagram of the base structure of a device interaction method provided in an embodiment of the present application. A speaker, a functional lighting effect LED group, a power supply lighting effect LED group, a main board, a microphone array, and various switches and buttons can be configured in the base of the intelligent interaction device, so as to realize functions such as sound emission, light emission, and rotation of the intelligent interaction device. At the same time, a role model of the target role can be placed on the base, so that the appearance of the intelligent interaction device can be configured as the target role.
[0100] In one embodiment, please refer to Figure 3f , Figure 3f It is another schematic diagram of the base structure of a device interaction method provided in an embodiment of the present application. The main chip of the intelligent interaction device is as shown in the figure. The wake-up word, volume, voiceprint, rotation, opening and closing state, etc. of the intelligent interaction device can be controlled through a mobile application package (Android application package, abbreviated as APK). Specifically, voice data generated by the user can be captured based on the microphone array, and thus based on the voice data, control over the speaker, stepper motor, power supply light group (i.e., the power supply lighting effect LED group) and functional effect LED group (i.e., the functional lighting effect LED group) of the intelligent interaction device can be triggered to realize interactive control of the intelligent interaction device. It can also interact with the cloud platform through a wireless network (WiFi) to obtain reply voice data, etc.
[0101] In one embodiment, please refer to Figure 3g , Figure 3gIt is a schematic diagram of the light effect setting of a device interaction method provided in an embodiment of the present application. The intelligent interaction device can be set with corresponding light effects in different states. Among them, the types of light effects can include dynamic effect lights and decorative lights, etc. For example, at the moment of power-on and voice activation of the intelligent interaction device, the dynamic effect lights of the intelligent interaction device can be controlled to produce the effect of gradually changing color, lighting up, changing color, and then remaining on constantly, and at the same time, the decorative lights can be controlled to remain on constantly. In this way, the interaction effect of the intelligent interaction device can be improved.
[0102] Optionally, please refer to Figure 3h , Figure 3h It is another schematic diagram of the base structure of a device interaction method provided in an embodiment of the present application. Correspondingly, the embodiment of the present application also provides a corresponding circuit board design diagram. For example, please refer to Figure 3i , Figure 3i It is a schematic diagram of the base design of a device interaction method provided in an embodiment of the present application. Correspondingly, please refer to Figure 3j , Figure 3j It is another schematic diagram of the base design of a device interaction method provided in an embodiment of the present application, showing the specific circuit board of the intelligent interaction device provided in the embodiment of the present application. Correspondingly, please refer to Figure 3k , Figure 3k It is another schematic diagram of the base design of a device interaction method provided in an embodiment of the present application, showing the circuit design diagram of the base circuit board of the intelligent interaction device provided in the embodiment of the present application.
[0103] With the explosive development of AI technology, artificial intelligence has entered the vision of ordinary people. However, most of the scenarios in which AI technology is applied in the market are only micro-innovations based on computers and mobile phone software, focusing on the office field or as a substitute for mobile phone functions, and there is no effective landing scenario combined with physical hardware products. That is, there is no interactive medium that brings AI and humans closer, and AI cannot bring emotional value to users beyond basic functions. On the other hand, traditional hand-made products focus on restoring the appearance of the character image, and cannot restore its voice, personality, emotions, etc., lacking deeper interactivity and emotional connection. To this end, the embodiment of the present application combines artificial intelligence technology with hand-made products to create intelligent hand-made equipment, giving hand-made products functions such as memorizing dialogues, simulating the character's timbre, personality, and thinking mode to a great extent, greatly increasing the playability and emotional interaction attributes of hand-made products, and also making a greater degree of innovation in the product entity form of AI technology, so that it can provide users with practical functions and emotional value. Specifically, the embodiment of the present application can hardware AI technology, rather than allowing users to interact with AI only in the form of software on computers and mobile phones. Furthermore, the embodiments of the present application can give a "soul" to traditional figures, and enhance the interactivity and emotional connection of the figures. In addition, the AI-based intelligent interactive device in the embodiments of the present application not only satisfies practical functions, but also brings more emotional value and companionship value to users, thereby enhancing the interactivity between artificial intelligence products and users, and further effectively exerting the value that artificial intelligence technology can bring.
[0104] As can be seen from the above, the embodiment of the present application obtains the event information of the user interaction triggering event by responding to the user interaction triggering event for the intelligent interactive device, and the appearance of the intelligent interactive device is configured as a specific target role; through the target language model corresponding to the target role, the voice data of the target role for the user interaction triggering event is generated based on the event information, and the target language model is trained based on the role information and dialogue information of the target role; through the intelligent interactive device, the voice data is played based on the role attribute information of the target role. In this way, through the intelligent interactive device whose appearance is configured as a specific target role, the event information of the user interaction triggering event is obtained, so that through the target language model corresponding to the target role, the voice data of the target role for the user interaction triggering event is generated based on the event information, and through the intelligent interactive device, the voice data is played based on the role attribute information of the target role, so that the target role can be used to reply to the event triggered by the user, effectively improving the interaction effect between the user and the intelligent interactive device, thereby improving the interactivity between the artificial intelligence product and the user, and further effectively exerting the value that artificial intelligence technology can bring.
[0105] To better implement the above method, an embodiment of the present invention further provides a device interaction apparatus, which can be integrated in an electronic device, and the electronic device can be a terminal.
[0106] For example, as Figure 4 shown, it is a schematic structural diagram of the device interaction apparatus provided by an embodiment of the present application. The device interaction apparatus may include an acquisition unit 201, a generation unit 202, and a playback unit 203, as follows:
[0107] The acquisition unit 201 is configured to obtain event information of a user interaction trigger event in response to a user interaction trigger event for an intelligent interaction device, and the shape of the intelligent interaction device is configured as a specific target role;
[0108] The generation unit 202 is configured to generate voice data of the target role for the user interaction trigger event based on the event information through a target language model corresponding to the target role, and the target language model is trained based on the role information and conversation information of the target role;
[0109] The playback unit 203 is configured to play the voice data based on the role attribute information of the target role through the intelligent interaction device.
[0110] In one embodiment, the user interaction trigger event includes the user uttering a voice. The acquisition unit 201 is configured to: obtain the first voice data uttered by the user;
[0111] The generation unit 202 includes:
[0112] The first generation subunit is configured to generate voice data of the target role for replying to the first voice data through the target language model corresponding to the target role based on the first voice data.
[0113] In one embodiment, the first generation subunit is configured to:
[0114] Identify the conversation mode matched by the first voice data;
[0115] If the conversation mode is a chat mode, perform semantic feature extraction on the first voice data through the target language model corresponding to the target role to obtain semantic features;
[0116] Generate voice data of the target role for replying to the first voice data through the target language model based on the semantic features.
[0117] In one embodiment, the first generation subunit is configured to:
[0118] Identify the conversation mode matched by the first voice data;
[0119] If the dialogue mode is the instruction mode, extract the instruction content in the first voice data based on the target language model corresponding to the target role, generate an instruction reply voice based on the instruction content, and obtain the multimedia content indicated by the instruction content;
[0120] Based on the instruction reply voice and the multimedia content, generate the voice data for the target role to reply to the first voice data.
[0121] In one embodiment, the first generation subunit includes:
[0122] A user identification module, configured to determine whether the user is the target user bound to the intelligent interaction device based on the first voice data;
[0123] A first generation module, configured to, if the user is the target user, generate the voice data for the target role to reply to the first voice data through the target language model corresponding to the target role based on the historical interaction data between the target user and the intelligent interaction device and the first voice data;
[0124] A second generation module, configured to, if not, generate the voice data for the target role to reply to the first voice data through the target language model corresponding to the target role based on the first voice data.
[0125] In one embodiment, the user identification module is configured to:
[0126] Perform voiceprint recognition on the first voice data to obtain the first voiceprint information;
[0127] Obtain the second voiceprint information corresponding to the target user bound to the intelligent interaction device;
[0128] Based on the first voiceprint information and the second voiceprint information, determine whether the user is the target user bound to the intelligent interaction device.
[0129] In one embodiment, the sound source localization unit is configured to:
[0130] Perform sound source azimuth recognition on the first voice data to obtain the sound source direction of the user;
[0131] Based on the sound source direction, control the intelligent interaction device to face the user.
[0132] In one embodiment, the first generation subunit is configured to:
[0133] Detect the device state information of the intelligent interaction device, where the device state information includes at least one of the attitude information of the intelligent interaction device and the device component assembly information;
[0134] Determine the target language model that matches the device state information from multiple large language models of the target role;
[0135] Generate, based on the first voice data, voice data for the target character to reply to the first voice data through the target language model.
[0136] In one embodiment, the user interaction trigger event includes a preset reminder event. The obtaining unit 201 is configured to:
[0137] Obtain at least one object involved in the reminder event and the historical memory information of the intelligent interaction device for the object;
[0138] The generating unit 202 is configured to:
[0139] Generate, through the target language model corresponding to the target character, voice data for the target character to reply to the user interaction trigger event based on the object and the historical memory information.
[0140] In one embodiment, the user interaction trigger event includes an interaction action generated by the user for the intelligent device. The obtaining unit 201 is configured to: obtain the action information of the interaction action generated by the user for the intelligent device
[0141] The generating unit 202 is configured to:
[0142] Generate, through the target language model corresponding to the target character, voice data for the target character to reply to the user interaction trigger event based on the action information.
[0143] In one embodiment, the device interaction device further includes a device control unit, configured to:
[0144] Determine the role posture information matching the voice data based on the role attribute information of the target character;
[0145] Determine the motion control information of at least one part of the intelligent interaction device and / or the device control information of the sub-devices on the intelligent interaction device based on the role posture information;
[0146] When playing the voice data through the intelligent interaction device, also control the intelligent interaction device to move through the motion control information and / or the device control information.
[0147] In one embodiment, the device interaction device further includes a voice output prompt unit, configured to:
[0148] Before playing the voice data through the intelligent interaction device, control the intelligent interaction device to execute the voice output prompt operation corresponding to the raising hand mode.
[0149] In one embodiment, the device interaction device further includes a playback interruption unit, configured to:
[0150] When playing voice data through an intelligent interaction device, if it is detected that the user interrupts the voice data, obtain the second voice data generated by the user based on the interruption behavior;
[0151] Based on the played data in the voice data and the second voice data, through the target language model, generate the target voice data for the target character to reply to the second voice data;
[0152] Play the target voice data through the intelligent interaction device.
[0153] As can be seen from the above, in the embodiment of the present application, the acquisition unit 201 obtains the event information of the user interaction trigger event in response to the user interaction trigger event for the intelligent interaction device, and the appearance of the intelligent interaction device is configured as a specific target character; the generation unit 202 generates the voice data for the target character to reply to the user interaction trigger event based on the event information through the target language model corresponding to the target character, and the target language model is trained based on the role information and conversation information of the target character; the playback unit 203 plays the voice data based on the role attribute information of the target character through the intelligent interaction device. In this way, through the intelligent interaction device whose appearance is configured as a specific target character, obtain the event information of the user interaction trigger event, so as to generate the voice data for the target character to reply to the user interaction trigger event based on the event information through the target language model corresponding to the target character, and thus play the voice data based on the role attribute information of the target character through the intelligent interaction device, which can realize replying to the event triggered by the user through the target character, effectively improve the interaction effect between the user and the intelligent interaction device, thereby enhancing the interactivity between the artificial intelligence product and the user, and further effectively exerting the value that the artificial intelligence technology can bring.
[0154] Correspondingly, the embodiment of the present application further provides an electronic device, which can be a terminal, and the terminal can be a terminal device such as a figurine, a smart phone, a tablet computer, a laptop computer, a touch screen, a game console, a personal computer (PC, Personal Computer), a personal digital assistant (Personal Digital Assistant, PDA), etc. Or, the electronic device can be a server.
[0155] As Figure 5 shown, Figure 5Schematic diagram of the structure of the electronic device provided by the embodiment of the present application. The electronic device 300 includes a processor 301 with one or more processing cores, a memory 302 with one or more computer-readable storage media, and a computer program stored on the memory 302 and executable on the processor. Among them, the processor 301 is electrically connected to the memory 302. Those skilled in the art can understand that the structure of the electronic device shown in the figure does not constitute a limitation on the electronic device, and it may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0156] The processor 301 is the control center of the electronic device 300, connecting various parts of the entire electronic device 300 through various interfaces and lines. By running or loading software programs and / or units stored in the memory 302, and by calling data stored in the memory 302, it executes various functions of the electronic device 300 and processes data. The processor 301 can be a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), etc., and can implement or execute the various methods, steps, and logical block diagrams disclosed in the embodiments of the present application.
[0157] In the embodiment of the present application, the processor 301 in the electronic device 300 will load the instructions corresponding to the processes of one or more application programs into the memory 302 according to the following steps, and the processor 301 will run the application programs stored in the memory 302 to implement various functions, such as:
[0158] In response to a user interaction trigger event for the intelligent interaction device, obtain the event information of the user interaction trigger event. The shape of the intelligent interaction device is configured as a specific target role; through the target language model corresponding to the target role, generate voice data of the target role for the user interaction trigger event based on the event information. The target language model is trained based on the role information and conversation information of the target role; through the intelligent interaction device, play the voice data based on the role attribute information of the target role.
[0159] This solution can obtain the event information of the user interaction trigger event by responding to the user interaction trigger event for the intelligent interaction device, and the appearance of the intelligent interaction device is configured as a specific target role; generate the voice data of the target role for the user interaction trigger event based on the event information through the target language model corresponding to the target role, and the target language model is trained based on the role information and conversation information of the target role; play the voice data based on the role attribute information of the target role through the intelligent interaction device. In this way, through the intelligent interaction device whose appearance is configured as a specific target role, obtain the event information of the user interaction trigger event, so as to generate the voice data of the target role for the user interaction trigger event through the target language model corresponding to the target role, and thus play the voice data based on the role attribute information of the target role through the intelligent interaction device, which can realize the reply to the event triggered by the user through the target role, effectively improve the interaction effect between the user and the intelligent interaction device, thereby enhancing the interactivity between the artificial intelligence product and the user, and further effectively exerting the value that the artificial intelligence technology can bring.
[0160] Furthermore, for the various functions realized by running the application program stored in the memory 302, reference can also be made to the descriptions in the foregoing embodiments, which will not be elaborated herein.
[0161] For the specific implementation of each of the above operations, reference can be made to the previous embodiments, which will not be elaborated herein.
[0162] Optionally, as Figure 5 shown, the electronic device 300 further includes: a touch display screen 303, a radio frequency circuit 304, an audio circuit 305, an input unit 306, and a power supply 307. Among them, the processor 301 is electrically connected to the touch display screen 303, the radio frequency circuit 304, the audio circuit 305, the input unit 306, and the power supply 307 respectively. Those skilled in the art can understand that Figure 5 the structure of the electronic device shown in
[0163] The touch display screen 303 can be used to display a graphical user interface and receive operation instructions generated by a user's interaction with the graphical user interface. The touch display screen 303 may include a display panel and a touch panel. Among them, the display panel can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the electronic device. These graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. The touch panel can be used to collect touch operations of the user on or near it (such as operations of the user using any suitable object or accessory such as a finger or a stylus on or near the touch panel), and generate corresponding operation instructions, and the operation instructions execute the corresponding program. Optionally, the touch panel can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user and detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 301, and can receive and execute the commands sent by the processor 301. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it is transmitted to the processor 301 to determine the type of touch event. Subsequently, the processor 301 provides a corresponding visual output on the display panel according to the type of touch event. In the embodiments of the present application, the touch panel and the display panel can be integrated into the touch display screen 303 to implement input and output functions. However, in some embodiments, the touch panel and the touch panel can be implemented as two independent components to implement input and output functions. That is, the touch display screen 303 can also be used as part of the input unit 306 to implement the input function.
[0164] The radio frequency circuit 304 can be used to transmit and receive radio frequency signals to establish wireless communication with a network device or other electronic devices through wireless communication, and transmit and receive signals with the network device or other electronic devices.
[0165] The audio circuit 305 can be used to provide an audio interface between the user and the electronic device through a speaker and a microphone. The audio circuit 305 can transmit the electrical signal converted from the received audio data to the speaker, and the speaker converts it into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 305 and then converted into audio data. After the audio data is output to the processor 301 for processing, it is transmitted through the radio frequency circuit 304 to, for example, another electronic device, or the audio data is output to the memory 302 for further processing. The audio circuit 305 may also include an earphone jack to provide communication between the peripheral earphone and the electronic device.
[0166] The input unit 306 can be used to receive an input target video and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0167] The power supply 307 is used to supply power to each component of the electronic device 300. Optionally, the power supply 307 can be logically connected to the processor 301 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 307 can also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0168] Although Figure 5 not shown in the figure, the electronic device 300 may also include a camera, a sensor, a Wi-Fi module, a Bluetooth module, etc., which will not be elaborated here.
[0169] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not elaborated in a certain embodiment, reference may be made to the relevant descriptions of other embodiments. It should be noted that the electronic device provided in the embodiments of the present application and the device interaction method in the above embodiments belong to the same concept. The specific implementation process is detailed in the above method embodiments and will not be elaborated here.
[0170] As can be seen from the above, the electronic device provided in the embodiments of the present application can obtain the event information of the user interaction trigger event by responding to the user interaction trigger event for the intelligent interaction device, and the appearance of the intelligent interaction device is configured as a specific target role; through the target language model corresponding to the target role, based on the event information, generate voice data of the target role for the user interaction trigger event, and the target language model is trained based on the role information and conversation information of the target role; through the intelligent interaction device, play the voice data based on the role attribute information of the target role. In this way, through the intelligent interaction device whose appearance is configured as a specific target role, obtain the event information of the user interaction trigger event, so as to generate, through the target language model corresponding to the target role, voice data of the target role for the user interaction trigger event based on the event information, and thus, through the intelligent interaction device, play the voice data based on the role attribute information of the target role, it is possible to realize replying to the event triggered by the user through the target role, effectively improving the interaction effect between the user and the intelligent interaction device, thereby enhancing the interactivity between the artificial intelligence product and the user, and further effectively bringing out the value that the artificial intelligence technology can bring.
[0171] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0172] To this end, an embodiment of the present application provides a computer-readable storage medium, which includes a computer program. When the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute any one of the device interaction methods provided by the embodiments of the present application. For example, the computer program can execute the steps of the following device interaction method:
[0173] In response to a user interaction trigger event for an intelligent interaction device, obtain the event information of the user interaction trigger event. The appearance of the intelligent interaction device is configured as a specific target role; through the target language model corresponding to the target role, generate voice data of the target role for the user interaction trigger event based on the event information. The target language model is trained based on the role information and dialogue information of the target role; through the intelligent interaction device, play the voice data based on the role attribute information of the target role.
[0174] This solution can obtain the event information of the user interaction trigger event by responding to the user interaction trigger event for the intelligent interaction device. The appearance of the intelligent interaction device is configured as a specific target role; through the target language model corresponding to the target role, generate voice data of the target role for the user interaction trigger event based on the event information. The target language model is trained based on the role information and dialogue information of the target role; through the intelligent interaction device, play the voice data based on the role attribute information of the target role. In this way, through the intelligent interaction device whose appearance is configured as a specific target role, obtain the event information of the user interaction trigger event, so as to generate voice data of the target role for the user interaction trigger event based on the event information through the target language model corresponding to the target role. In this way, through the intelligent interaction device, play the voice data based on the role attribute information of the target role, and can realize answering the event triggered by the user through the target role, effectively improving the interaction effect between the user and the intelligent interaction device, thereby improving the interactivity between the artificial intelligence product and the user, and further effectively bringing out the value that the artificial intelligence technology can bring.
[0175] Furthermore, for the refinement steps of the above method steps, reference can also be made to the descriptions in the foregoing embodiments, which will not be elaborated herein.
[0176] For the specific implementation of each of the above operations, reference can be made to the previous embodiments, which will not be elaborated herein.
[0177] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.
[0178] Since the computer program stored in the computer-readable storage medium can execute any one of the device interaction methods provided in the embodiments of the present application, the beneficial effects achievable by any one of the device interaction methods provided in the embodiments of the present application can be realized. For details, refer to the previous embodiments and will not be elaborated here.
[0179] According to one aspect of the present application, there is also provided a computer program product, including a computer program, the computer program being stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device executes the methods provided in the various optional implementation manners in the above embodiments.
[0180] In the above embodiments of the device interaction device, computer-readable storage medium, electronic device, and computer program product, the descriptions of each embodiment have their own focuses. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments. Those skilled in the art can clearly understand that for the sake of convenience and brevity of description, the specific working processes and beneficial effects of the above-described device interaction device, computer-readable storage medium, computer program product, electronic device, and their corresponding units can refer to the description of the device interaction method in the above embodiments and will not be elaborated here specifically.
[0181] The above has introduced in detail a device interaction method, device, electronic device, computer-readable storage medium, and computer program product provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, based on the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A device interaction method, characterized in that: include: In response to a user interaction triggering event for an intelligent interactive device, acquiring event information of the user interaction triggering event, wherein the appearance of the intelligent interactive device is configured as a specific target character; Generate speech data of the target character for the user interaction triggering event based on the event information by using a target language model corresponding to the target character, wherein the target language model is trained based on the character information and dialogue information of the target character; The voice data is played through the intelligent interactive device based on the role attribute information of the target role.
2. The device interaction method according to claim 1, characterized in that: The user interaction triggering event includes a user speaking a voice, and obtaining event information of the user interaction triggering event includes: obtaining first voice data spoken by the user; The generating, using the target language model corresponding to the target role and based on the event information, the speech data of the target role for the user interaction triggering event includes: Based on the first voice data, voice data in which the target character responds to the first voice data is generated through a target language model corresponding to the target character.
3. The device interaction method according to claim 2, characterized in that: The generating, based on the first voice data and using a target language model corresponding to the target role, voice data in which the target role responds to the first voice data includes: identifying a conversation pattern matched by the first voice data; If the dialogue mode is a chat mode, extracting semantic features from the first voice data using a target language model corresponding to the target role to obtain semantic features; The target language model generates voice data for the target character to reply to the first voice data based on the semantic features.
4. The device interaction method according to claim 2, characterized in that: The generating, based on the first voice data and using a target language model corresponding to the target role, voice data in which the target role responds to the first voice data includes: identifying a conversation pattern matched by the first voice data; If the dialogue mode is a command mode, extracting command content in the first voice data based on a target language model corresponding to the target role, generating a command reply voice based on the command content, and acquiring multimedia content indicated by the command content; Based on the instruction reply voice and the multimedia content, voice data in which the target character responds to the first voice data is generated.
5. The device interaction method according to claim 2, characterized in that: The generating, based on the first voice data and using a target language model corresponding to the target role, voice data in which the target role responds to the first voice data includes: Based on the first voice data, determining whether the user is a target user bound to the intelligent interactive device; If the user is the target user, based on the historical interaction data between the target user and the intelligent interactive device and the first voice data, and using the target language model corresponding to the target role, generate voice data for the target role to reply to the first voice data; If not, based on the first voice data, voice data in which the target character responds to the first voice data is generated through a target language model corresponding to the target character.
6. The device interaction method according to claim 5, characterized in that: The determining, based on the first voice data, whether the user is a target user bound to the intelligent interactive device includes: Performing voiceprint recognition on the first voice data to obtain first voiceprint information; Acquire second voiceprint information corresponding to a target user bound to the intelligent interactive device; Based on the first voiceprint information and the second voiceprint information, it is determined whether the user is a target user bound to the intelligent interactive device.
7. The device interaction method according to claim 2, characterized in that: The method further comprises: Performing sound source direction recognition on the first voice data to obtain the sound source direction of the user; Based on the direction of the sound source, the intelligent interactive device is controlled to face the user.
8. The device interaction method according to claim 2, characterized in that: The generating, based on the first voice data and using a target language model corresponding to the target role, voice data in which the target role responds to the first voice data includes: Detecting device status information of the intelligent interactive device, where the device status information includes at least one of posture information of the intelligent interactive device and device component assembly information; Determining a target language model that matches the device state information from a plurality of large language models of the target role; The target language model generates voice data for the target character to reply to the first voice data based on the first voice data.
9. The device interaction method according to claim 1, characterized in that: The user interaction triggering event includes a preset reminder event, and the acquiring event information of the user interaction triggering event includes: Acquire at least one object involved in the reminder event and historical memory information of the intelligent interactive device for the object; The generating, using the target language model corresponding to the target role and based on the event information, the speech data of the target role for the user interaction triggering event includes: The target language model corresponding to the target role is used to generate voice data of the target role for the user interaction triggering event based on the object and the historical memory information.
10. The device interaction method according to claim 1, characterized in that: The user interaction triggering event includes an interaction action generated by the user on the smart device, and the acquiring event information of the user interaction triggering event includes: acquiring action information of the interaction action generated by the user on the smart device; The generating, using the target language model corresponding to the target role and based on the event information, the speech data of the target role for the user interaction triggering event includes: The target language model corresponding to the target role is used to generate voice data of the target role for the user interaction triggering event based on the action information.
11. The device interaction method according to any one of claims 1 to 10, characterized in that: The method further comprises: Determining character posture information matching the voice data based on the character attribute information of the target character; Determine motion control information of at least one part of the intelligent interactive device and / or device control information of a sub-device on the intelligent interactive device based on the character posture information; When the voice data is played through the intelligent interactive device, the intelligent interactive device is also controlled to move through the motion control information and / or the device control information.
12. The device interaction method according to any one of claims 1 to 10, characterized in that: The method further comprises: Before playing the voice data through the intelligent interactive device, the intelligent interactive device is controlled to perform a voice output prompt operation corresponding to the hand-raising mode.
13. The device interaction method according to any one of claims 1 to 10, characterized in that: The method further comprises: If, when the voice data is played through the intelligent interactive device, it is detected that the user has interrupted the voice data, obtaining second voice data generated by the user based on the interruption behavior; Based on the played data in the voice data and the second voice data, generating target voice data for the target character to reply to the second voice data through the target language model; The target voice data is played through the intelligent interactive device.
14. A device interaction apparatus, characterized in that: Applied to target devices, including: an acquisition unit, configured to acquire event information of a user interaction triggering event in response to a user interaction triggering event for an intelligent interactive device, wherein the appearance of the intelligent interactive device is configured as a specific target character; A generating unit, configured to generate speech data of the target character for the user interaction triggering event based on the event information by using a target language model corresponding to the target character, wherein the target language model is trained based on the character information and dialogue information of the target character; A playback unit is used to play the voice data based on the role attribute information of the target role through the intelligent interactive device.
15. An electronic device, characterized in that: The invention comprises a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the device interaction method according to any one of claims 1 to 13.
16. A storage medium, characterized in that: The invention comprises a computer program, and when the computer program is run on an electronic device, the computer program is used to make the electronic device execute the steps of the device interaction method according to any one of claims 1 to 13.