Interaction method, interaction system and electronic equipment
By combining external sensor identification information with a large language model, the AI chatbot achieves personalized and rich interactions with users, solving the problem of the single role of existing AI chatbots and improving user experience and device performance.
Patent Information
- Application Number
- CN202511426693.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2026-01-06
AI Technical Summary
Existing AI chatbots have a limited role, cannot provide a rich user experience, and have a limited scope of application.
By outputting response voice signals based on the identification information of external sensors, the system enables interactive communication of descriptive information from different external sensors. Combined with a large language model, it generates response voice signals to enrich the user experience, including automatic matching of role and scene description information.
It improves the usability and interactive fun of electronic devices, enables personalized and authentic interaction with users, and enhances the immersiveness and real-time nature of the interaction.
Smart Images

Figure CN121281518A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electronic product technology, and in particular to an interaction method, an interaction system, and an electronic device. Background Technology
[0002] Current AI (artificial intelligence) chatbots are fixed characters with a single personality, resulting in a rather monotonous interaction that fails to provide users with a richer experience, thus limiting their application scope. Summary of the Invention
[0003] To address the above issues, embodiments of the present invention provide an interaction method, an interaction system, and an electronic device. By identifying corresponding descriptive information through external sensors and outputting response voice signals, different roles can interact based on the descriptive information of different external sensors, enriching the user experience and improving the performance of the electronic device.
[0004] In some embodiments, the interaction method includes: when a non-contact sensor on an electronic device senses a signal from an external sensor, acquiring identification information of the external sensor; receiving a voice signal input by a user via a voice input device on the electronic device; and outputting a response voice signal via a voice playback device of the electronic device based on descriptive information associated with the identification information of the external sensor and the voice signal input by the user, wherein the response voice signal is generated by a large language model outputting response information based on the descriptive information and the voice signal input by the user, and the response voice signal is used to respond to the voice signal input by the user so that the electronic device can interact with the user via voice.
[0005] In some embodiments, the external sensor includes a card with an electronic sensor; and / or, the electronic sensor is an object with an NFC chip; and / or, the identification information includes the ID information of the external sensor.
[0006] In some embodiments, the identification information of the external sensor is transmitted to the server so that the server can obtain the description information corresponding to the external sensor based on the identification information of the external sensor; wherein, the description information includes role description information or scene description information.
[0007] In some embodiments, the user-inputted voice signal is transmitted to a server so that the server can recognize the user-inputted voice signal and generate speech-recognized text.
[0008] The response speech signal is generated by the large language model after outputting response information based on the prompt information and the speech recognition text corresponding to the speech signal;
[0009] The description information indicates the role that interacts with the user via voice when the large language model outputs response information.
[0010] In some embodiments, the interaction method includes:
[0011] Based on the audio signal input by the current user, identify the voiceprint information of the current user;
[0012] Based on the current user's voiceprint information, match it with pre-stored user registration information;
[0013] When a match is successful, the output audio information includes the user name from the user registration information;
[0014] The user registration information includes the user's name and voice characteristics.
[0015] In some embodiments, user registration information is stored through the following steps.
[0016] Receive instructions from the user to register a user, and obtain the user's voice signal and username;
[0017] The user's voice features are generated based on the user's input voice signal, and the user's voice features are associated with and stored with the user's name.
[0018] In some embodiments, the interactive content and / or status information of the electronic device are displayed through a display device, wherein the interactive content includes interactive role information or voice content, and the status information includes any one or more of the following: the battery level information of the electronic device, whether it is in voice input state, whether it is in voice playback state, card insertion state, and the name of the connected WiFi signal.
[0019] In some embodiments, the sensing status of the external sensor is displayed by a lighting effect display module disposed on the electronic device.
[0020] In some embodiments, the method further includes playing an initial prompt voice through a voice playback device of the electronic device before the non-contact sensor on the electronic device senses the sensing signal of the external sensor and before receiving a voice signal input by the user, wherein the initial prompt voice is generated based on the description information corresponding to the external sensor.
[0021] In some embodiments, the initial prompt voice is generated by the large language model based on the description information corresponding to the external sensor and the preset initial response description information.
[0022] In some embodiments, the initial prompt voice includes the name of the user who last interacted with the electronic device.
[0023] In some embodiments, an interactive system is provided, including one or more processors and a memory, wherein the memory stores a computer program, and when the processor executes the computer program, it can implement the steps of the interactive method of any one of the preceding claims.
[0024] In some embodiments, an electronic device is provided, including a housing and a processor disposed within the housing, the processor being configured to perform the method as described in any one of the preceding claims. Attached Figure Description
[0025] The technical features of this invention are detailed in the claims. To better understand the features and beneficial effects of this invention, various embodiments are described in detail below with reference to the accompanying drawings. This description is not restrictive, and the drawings include:
[0026] Figure 1 This is a schematic diagram of an interactive system described in an embodiment of the present invention.
[0027] Figure 2 This is a schematic diagram of an electronic device according to an embodiment of the present invention.
[0028] Figure 3 This is a schematic diagram of an interactive system described in another embodiment of the present invention.
[0029] Figure 4 This is a flowchart illustrating the interactive method of an embodiment of the present invention.
[0030] Figure 5 This is a schematic diagram of an electronic device according to an embodiment of the present invention.
[0031] Figure 6 This is a flowchart illustrating the interactive method of an embodiment of the present invention.
[0032] Figure 7 This is a flowchart illustrating the interactive method of an embodiment of the present invention. Specific Implementation
[0033] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0034] This application provides embodiments that Figure 1 The interactive system shown, such as Figure 1 As shown, the interactive system 1 may include an electronic device 11 and a server 12 that communicate wirelessly with each other. The electronic device receives voice signals input by a user and outputs a response voice signal based on the user's input. The electronic device includes a voice input device 111, a voice playback device 112, and a contactless sensor 113. The contactless sensor 113 is used to sense signals from external sensors, the voice input device 111 is used to receive voice signals input by the user, and the voice playback device 112 is used to output a response voice signal. In some other embodiments, the electronic device further includes a display screen or a touch screen. The display screen is used to display the interactive content and / or status information of the electronic device, and the touch screen is used to receive touch operations from the user. For example, the contactless sensor includes an NFC sensor, the voice input device includes a microphone, and the voice playback device includes a speaker. In other embodiments, such as... Figure 3 As shown, the interactive system 1 further includes a display device 13, which is used to display the interactive content and / or status information of the electronic device 11. The display device 14 is communicatively connected to the server 12.
[0035] For example, the display device 13 may include at least one of a smartphone, tablet computer, laptop computer, smart wearable device, etc., and an application (APP) for setting or controlling the electronic device may be installed on the display device 13. The display device communicates with the server or the electronic device. For example, the display device may also be an augmented reality (AR) device, a virtual reality (VR) device, etc.
[0036] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0037] The following detailed description, in conjunction with the accompanying drawings, outlines some embodiments of this application. Unless otherwise specified, the following embodiments and features can be combined with each other. Current AI chatbots are fixed characters with singular personalities, resulting in a limited and monotonous interaction experience that fails to provide users with a richer experience, thus restricting the robot's application scope. To address this problem, embodiments of the present invention provide an interaction method, such as... Figure 4 As shown, the steps include:
[0038] S1, When the non-contact sensor on the electronic device senses the signal of the external sensor, the identification information of the external sensor is acquired;
[0039] For example, the electronic device includes a housing / electronic device disposed within a cavity of the housing or on the housing / mechanical interaction device disposed on the housing; the electronic device includes a processor and the non-contact sensor. In this embodiment, by sensing the signal of an external sensing element through a non-contact sensor, manual alignment and insertion for contact sensing are eliminated, thereby improving the user experience and enhancing the ease of use of the electronic device.
[0040] In one embodiment, the external sensing element is an object with an NFC (near-field communication) chip. The non-contact sensor is a sensor capable of reading information from the object with the NFC chip, such as an NFC sensor.
[0041] In some embodiments, the external sensor is a card with an electronic sensor. Correspondingly, such as Figure 5 As shown, the housing 114 of the electronic device 11 is provided with a card slot 115 for accommodating the card. For example, the card slot 115 is an open card slot structure, so that the pattern on the card can be exposed to the outside, and the user can see the pattern on the card when interacting with the electronic device, and know the role of the current interaction during the interaction, making the interaction more user-friendly.
[0042] In one embodiment, acquiring the identification information of the external sensor includes acquiring the identification information of the NFC chip, such as stored ID information.
[0043] S2, Receive the voice signal input by the user through the voice input device on the electronic device;
[0044] The voice input device is a microphone mounted on the housing of the electronic device. The microphone can be an external device or a device mounted on the main body of the electronic device.
[0045] S3, based on the description information associated with the identification information of the external sensor and the voice signal input by the user, output a response voice signal through the voice playback device of the electronic device.
[0046] In step S3, the description information indicates the role that interacts with the user via voice when the large language model outputs response information. Specifically, the description information associated with the identification information of the external sensor includes role description information associated with the ID information of the NFC chip. For example, the role description information describes famous novel characters, such as Sun Wukong, Nezha, and Tang Sanzang, or characters of specific professions, such as teachers, fitness instructors, and police officers; or famous IP characters, such as Elsa and Anna from Frozen. By automatically acquiring role description information through contactless sensing and interacting with the user using AI based on this information, the interaction becomes more fun and personalized, improving the usability of the electronic device. In some embodiments, the description information associated with the identification information of the external sensor includes scene description information. For example, the scene description information describes the current dialogue scene, such as a classroom scene, a zoo scene, a fire scene, and a sleeping scene. When the description information includes scene description information, a role matching the scene is automatically assigned; for example, in a classroom scene, the dialogue role is assigned as a teacher, and in a zoo scene, the dialogue role is automatically assigned as the zookeeper.
[0047] For example, the sound attributes of the voice playback device are automatically set based on the role description information corresponding to the external sensor. The sound attributes include one or more of the speaker's timbre, tone, accent, etc.
[0048] For example, the server stores the identification information and corresponding description information of the external sensor. The electronic device then needs to transmit the identification information of the external sensor to the server so that the server can obtain the corresponding description information of the external sensor based on the identification information.
[0049] For example, the interaction method includes:
[0050] The user-inputted voice signal is transmitted to the server so that the server can recognize the user-inputted voice signal and generate speech-recognized text.
[0051] The response speech signal is generated by the large language model after outputting response information based on the prompt information and the speech recognition text corresponding to the speech signal. Specifically, the large language model does not directly process speech and output the speech itself. Instead, it needs to convert the speech signal input into text input into the large language model, or convert the text content output by the large language model into a speech signal for output.
[0052] The description information indicates the role that interacts with the user via voice when the large language model outputs response information.
[0053] Furthermore, the response voice signal is generated by the large language model after outputting response information based on the description information and the user-input voice signal. This response voice signal is used to respond to the user-input voice signal, enabling the electronic device to interact with the user via voice. Specifically, the large language model can engage in AI dialogue and interaction with the user in a specific role based on the input specific role description information. Compared to the one-way output of directly indexed content in existing technologies, this achieves a more realistic interactive effect. Additionally, the large language model can provide AI responses and interactions based on the user-input voice signal, providing real-time interactivity. Specifically, the electronic device or server system recognizes the user-input voice signal, generates corresponding text information, and then inputs it into the large language model. The large language model generates the response voice signal after outputting response information based on the prompt information and the speech-recognized text corresponding to the voice signal.
[0054] In some embodiments, the identification information of the external sensor is transmitted to the server so that the server can obtain the description information corresponding to the external sensor based on the identification information of the external sensor.
[0055] In some embodiments, such as Figure 6 As shown, the interaction method further includes the following steps:
[0056] S41, Based on the audio signal input by the current user, identify the voiceprint information of the current user;
[0057] S42, Match the pre-stored user registration information based on the current user's voiceprint information;
[0058] S43, When a match is successful, the output response voice signal includes the user name from the user registration information. The user registration information includes the user name and the user's voice characteristics.
[0059] Through steps S41 to S43, the user's name can be called during the conversation, improving the real-time performance and authenticity of the electronic device interaction. For example, if the pre-stored user registration information includes voiceprint feature A and the user name Jessica, then if the current user is Jessica herself and inputs an audio signal into the electronic device, the voiceprint information of the current user Jessica will be identified. When her voiceprint information successfully matches the pre-stored voiceprint feature A, the user name Jessica will be obtained, and the output response voice signal will call out the user name "Jessica".
[0060] For example, user registration information is stored through the following steps.
[0061] Receive instructions from the user to register a user, and obtain the user's voice signal and username;
[0062] The user's voice features are generated based on the user's input voice signal, and the user's voice features are associated with and stored with the user's name.
[0063] Specifically, the association between user voice characteristics and user name can be achieved in the following two ways:
[0064] When a user interacts with the electronic device for the first time, if the electronic device does not find matching user registration information, it will proactively prompt the user to enter a username. The prompt may be a voice prompt or a prompt displayed on a screen. The user can enter the username by voice or manually type the spelled username on the display. During a voice conversation with the user, upon receiving the username, the user's voiceprint characteristics are associated with and stored.
[0065] After a user interacts with the electronic device, the user's voice data is stored. In response to an instruction to register user information, historical conversation content, including the user's voice data, is displayed. Upon receiving a confirmation instruction to select a specific user's voice data, the user is prompted to enter a username. Once the username is received, the current user's voiceprint characteristics are associated with and stored.
[0066] In some embodiments, the interactive content and / or status information of the electronic device are displayed via a display device. The interactive content includes interactive character information or voice content, and the status information includes any one or more of the following: battery level information of the electronic device, whether it is in voice input mode, whether it is in voice playback mode, card insertion status, whether it is connected to Wi-Fi, and the name of the connected Wi-Fi signal. The display device may be located on the electronic device or may be a display device independent of the electronic device.
[0067] In some embodiments, the electronic device includes a lighting effect display module. The lighting effect display module, located on the electronic device, displays the sensing status of the external sensor. Different lighting effects are displayed before a sensing signal is detected, when a sensing signal from the external sensor is first detected, and when a stable and continuous sensing of the external sensor is achieved. Specifically, when the external sensor is a card with an NFC chip, different lighting effects are displayed before, during, and after the card is inserted.
[0068] In some embodiments, the method further includes:
[0069] When a stable and continuous sensing signal from an external sensor is detected, the operating state of the electronic device is determined.
[0070] Different lighting effects are displayed according to the different working states of the electronic device.
[0071] The operating states include voice playback state, voice input state, and sleep state. Specifically, when the electronic device is in voice playback state, it does not receive user voice input; when the electronic device is in sleep state, it does not output audio and does not receive user voice input.
[0072] When no voice signal is received within a preset time threshold during voice input, the electronic device is controlled to enter sleep mode.
[0073] Specifically, when the electronic device is in sleep mode, upon receiving a trigger signal, it exits sleep mode and enters either voice playback mode or voice input mode. In one embodiment, when the electronic device is in sleep mode, upon receiving a trigger signal, it exits sleep mode and plays an initial prompt voice through its voice playback device. The trigger signal can be manually activated via a physical button on the electronic device.
[0074] In some embodiments, during the dialogue, images or dynamic videos are automatically generated based on the current dialogue content to achieve a more realistic and engaging dialogue scenario.
[0075] In some embodiments, before the contactless sensor on the electronic device detects the sensing signal of the external sensor and before receiving a voice signal input from the user, an initial prompt voice is played through the voice playback device of the electronic device. The initial prompt voice is generated based on the descriptive information corresponding to the external sensor. For example, after detecting the sensing signal of the NFC card and obtaining the corresponding descriptive information, and before receiving a voice signal input from the user, a dialogue can be initiated with the character corresponding to the NFC card. The dialogue conforms to the character's language habits / greeting style and audio selection.
[0076] The initial prompt voice is generated by the large language model based on the description information corresponding to the external sensor and preset initial response description information. For example, the initial response description information is description information for initiating or guiding a dialogue, such as "You are XXX, greet the user," or "You are XXX, start an interesting conversation," or "You are XXX, guide the user into a specific scenario."
[0077] The initial prompt voice includes the username of the user who last interacted with the electronic device. For example, if the user who last interacted with the electronic device was Jessica, then the audio "Jessica" will be played in the initial prompt voice the next time a newly inserted NFC card is detected or the device is restarted. In this way, the immersive experience and deep connection with the electronic device can be increased, and the effectiveness of the interaction can be improved.
[0078] In some embodiments, the electronic device includes different voice input modes. When the electronic device is in different voice input modes, the voice playback response speed of the electronic device is different to adapt to the personalized needs of different users and improve the performance of the electronic device. In some embodiments, the method includes the steps of:
[0079] When the first input command is received, the electronic device is set to the first voice input mode;
[0080] When a second input command is received, the electronic device is set to the second voice input mode;
[0081] The output response speed of the electronic device differs depending on whether it is in a first voice input mode or a second voice input mode. The first voice mode includes an intermittent voice mode, and the second voice mode includes a continuous voice mode. The output response speed of the electronic device in the intermittent voice input mode is greater than that in the continuous voice input mode.
[0082] In the intermittent voice input mode, the voice input device can only receive voice input when a trigger signal is received; in the continuous voice input mode, the voice input device continuously receives voice input.
[0083] For example, the output response speed of the electronic device in interval voice input mode is greater than that in continuous voice input mode. For instance, during a conversation, when the user stops inputting audio signals, the output response speed of the electronic device in interval voice input mode is greater than that in continuous voice input mode. Specifically, the output response speed in interval voice input mode is at least 2 seconds or 0.5 seconds faster than that in continuous voice input mode. By setting different voice input modes and corresponding response speeds, different user needs and experiences can be personalized. For example, when a user needs simple conversations and quick responses, interval voice input mode can be used; when a user needs advanced conversations and slower responses that won't confuse the user, continuous voice input mode can be used. These two signal input modes can meet the needs of users with different cognitive abilities. By setting different signal input modes, the usability and playability of the electronic device can be further improved.
[0084] In some embodiments, when a first input instruction is received, the electronic device is set to an intermittent voice input mode, and when a second input instruction is received, the electronic device is set to a continuous voice input mode. In some embodiments, the first and second input instructions are implemented by manually adjusting a first physical device to different positions. In other embodiments, the first and second input instructions are implemented by operating a specific control on a touchscreen.
[0085] When the electronic device is in intermittent voice input mode, if no trigger signal is received, it stops receiving voice input and begins processing voice input and outputting a response. Exemplarily, the trigger signal is achieved by manually pressing a second physical device. In other embodiments, the trigger signal is achieved by manually operating a specific control on the touchscreen.
[0086] When the electronic device is in continuous voice input mode, if no voice input is detected within a preset time threshold, it begins processing the voice input and outputs a response. The preset time threshold is no greater than 1 second. For example, the preset time threshold is 0.5 seconds.
[0087] In some embodiments, by manually adjusting the first physical device to different positions, different signals sensed by the Hall sensor determine whether the electronic device is currently in an intermittent voice input mode or a continuous voice input mode. The first physical device includes a roller structure, a toggle structure, and a pressing structure. By rolling the roller structure, toggling the toggle structure, or pressing the pressing structure, the magnet on the first physical device is positioned at different locations, allowing the Hall sensor on the motherboard fixed inside the electronic device housing to sense electrical signals of varying intensities. Furthermore, a processor mounted on the motherboard can determine whether the electronic device is currently in an intermittent voice input mode or a continuous voice input mode based on these different electrical signals.
[0088] In some embodiments, the electronic device enters either a specific dialogue mode or a regular dialogue mode depending on whether the non-contact sensor detects a signal from an external sensor. When the non-contact sensor detects a signal from an external sensor, the electronic device enters the specific dialogue mode, and the electronic device outputs response audio with a corresponding role, which is related to the identification information of the detected external sensor. When the non-contact sensor does not detect a signal from an external sensor, the electronic device enters the regular dialogue mode, and the electronic device outputs response audio without a specific role. For example, in the regular dialogue mode, response audio with preset sound attributes is output. The preset attributes include preset timbre, tone, etc.
[0089] For example, this application provides an interaction method applied to the electronic device, wherein the electronic device is used to receive a voice signal input by a user and output a response voice signal based on the user-input voice signal, such as... Figure 7 As shown, the method includes:
[0090] S51, when it is detected that the identification information of the external sensor is not associated with any description information, the external sensor is confirmed to be a blank sensor;
[0091] S52, Receive the description information corresponding to the blank sensor, and associate the identification information of the blank sensor with the description information.
[0092] In some embodiments, when it is detected that the identification information of an external sensor is not associated with any description information, a reminder message is issued to prompt the user to set the description information of the external sensor.
[0093] When it is detected that the identification information of an external sensor is not associated with any descriptive information, a reminder message is issued to prompt the user to set the descriptive information of the external sensor, including:
[0094] The electronic device outputs a voice prompt via its voice playback device;
[0095] A prompt to enter descriptive information is sent through the computer interface.
[0096] The input description information reminders include reminders for users to bind roles.
[0097] In response to the input of descriptive information, the identification information of the blank sensor is associated with the descriptive information.
[0098] The operation of inputting description information includes the selection of existing roles and the input of descriptions for new roles.
[0099] After associating the identification information of the blank sensor with the description information, an initial prompt voice is played through the voice playback device of the electronic device according to the description information. The initial prompt voice is generated by the large language model based on the description information corresponding to the external sensor and preset initial response description information. For example, the initial response description information is description information that initiates or guides a dialogue, such as "You are XXX, greet the user," or "You are XXX, start an interesting conversation," or "You are XXX, guide the user into a specific scenario."
[0100] In some embodiments, the type of the external sensor is identified, and based on the type of the external sensor, it is determined whether it is capable of changing the associated descriptive information.
[0101] The identification of the type of the external sensor includes obtaining the type of the external sensor based on the identification information of the external sensor; wherein the type of the external sensor includes a fixed role type and a variable type.
[0102] When the type of the external sensor is variable, upon receiving an instruction to change the associated description information, the identification information of the blank sensor is associated with the new description information based on the input of the new description information.
[0103] In response to a viewing operation of the description information of the external sensor, the description information of the external sensor is displayed.
[0104] In some embodiments, in response to an instruction to create description information, user-inputted description information is stored; or in response to an instruction to create description information, new description information is automatically generated and stored based on user-inputted name information. For example, the name information includes a character name, and the description information includes a general introduction of the character, personality traits, speaking style, etc.
[0105] In some embodiments, in response to a command to create description information, the corresponding sound attribute is automatically selected based on the name information entered by the user.
[0106] In some embodiments, the corresponding sound attribute is automatically selected based on the input description information or the automatically generated description information.
[0107] In some embodiments, based on the input user voice dialogue content, the user's voice features and personality features are automatically extracted, and character description information is automatically created based on the personality features. The personality features include personality attributes, vocabulary habits, and speaking style. For example, based on the input user image, personality features, voice features, and user image are automatically associated. For instance, if the user inputs a voice dialogue between themselves and Jessica, in response to a selection operation, Jessica's voice features and personality features are automatically extracted. For example, personality features include that Jessica is a lively elementary school student, her catchphrase is "Oh my god!", and she likes dogs. Simultaneously, in response to the input user name and user image, the user name Jessica, the user image Jessica, Jessica's voice features, and personality features are associated to create the character of Jessica.
[0108] In some embodiments, a specific user's role is automatically created based on the historical dialogue content between a specific user and the electronic device. This specific user's role can be bound to a blank sensor to generate a virtual user, thereby enabling other users to interact with this virtual user. For example, after Jessica interacts with the electronic device, according to the instruction to create the Jessica role, Jessica's description information and voice characteristics are generated. When the blank sensor and the binding instruction are detected, the identification information of the blank sensor, the description information, and the voice characteristics are associated. Then, when a sensor bound to Jessica's description information and voice characteristics is detected, a new user can interact with the virtual Jessica.
[0109] Furthermore, when a specific user engages in a new conversation with the electronic device, the voiceprint information of the specific user is identified, and the description information of the specific user is updated based on the new conversation.
[0110] In some embodiments, if no user-inputted voice signal is received within a preset time threshold after the output response audio signal, a message reminding the user to input an audio signal is output. The output reminder audio signal is generated by the large language model based on preset timeout reminder description information. The timeout reminder description information includes calling out to the user, responding to the previous voice input with a different expression, or restarting a new topic.
[0111] In some embodiments, if no user-inputted voice signal is received within a preset time threshold after the output of the response audio signal, a dialogue termination response audio is output. This dialogue termination response audio is generated by the large language model based on preset timeout dialogue termination description information. The timeout dialogue termination description information includes description information based on the current dialogue scenario and description information based on the current time.
[0112] In some embodiments, if no user-inputted voice signal is received within a preset time threshold after outputting a response audio signal, the device automatically shuts down after outputting a response audio signal to end the dialogue; or it enters a sleep mode. When the electronic device is in sleep mode, it does not output audio and does not receive user voice input. Specifically, when the electronic device is in sleep mode, upon receiving a trigger signal, it exits the sleep mode and controls the electronic device to enter a voice playback mode or a voice input mode. Specifically, in one embodiment, when the electronic device is in sleep mode, upon receiving a trigger signal, it exits the sleep mode and plays an initial prompt voice through the electronic device's voice playback device. The trigger signal can be manually triggered by a physical button on the electronic device.
[0113] An interactive system according to this application may include a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it can implement the steps of any of the above-mentioned interactive methods.
[0114] For example, when a computer program is executed, the following steps can be achieved:
[0115] When a non-contact sensor on an electronic device senses a signal from an external sensor, it acquires the identification information of the external sensor.
[0116] Receives voice signals input by the user through a voice input device on an electronic device;
[0117] Based on the description information associated with the identification information of the external sensor and the voice signal input by the user, the electronic device outputs a response voice signal through its voice playback device. The response voice signal is generated by the large language model after outputting response information based on the description information and the voice signal input by the user. The response voice signal is used to respond to the voice signal input by the user so that the electronic device can interact with the user via voice.
[0118] For example, the external sensor includes a card with an electronic sensor; and / or, the electronic sensor is an object with an NFC chip; and / or, the identification information includes the ID information of the external sensor.
[0119] For example, when a computer program is executed, the following steps can be achieved:
[0120] The identification information of the external sensor is transmitted to the server, so that the server can obtain the description information corresponding to the external sensor based on the identification information; wherein, the description information includes character description information or scene description information.
[0121] For example, when a computer program is executed, the following steps can be achieved:
[0122] The user-inputted voice signal is transmitted to the server so that the server can recognize the user-inputted voice signal and generate speech-recognized text.
[0123] The response speech signal is generated by the large language model after outputting response information based on the prompt information and the speech recognition text corresponding to the speech signal.
[0124] The description information indicates the role that interacts with the user via voice when the large language model outputs response information.
[0125] For example, when a computer program is executed, the following steps can be achieved:
[0126] Based on the audio signal input by the current user, identify the voiceprint information of the current user;
[0127] Based on the current user's voiceprint information, match it with pre-stored user registration information;
[0128] When a match is successful, the output audio information includes the user name from the user registration information;
[0129] The user registration information includes the user's name and voice characteristics.
[0130] For example, when a computer program is executed, the following steps can be achieved:
[0131] Receive instructions from the user to register a user, and obtain the user's voice signal and username;
[0132] The user's voice features are generated based on the user's input voice signal, and the user's voice features are associated with and stored with the user's name.
[0133] For example, when a computer program is executed, the following steps can be achieved:
[0134] The interactive content and / or status information of the electronic device are displayed through a display device, wherein the interactive content includes interactive role information or voice content, and the status information includes any one or more of the following: the battery level information of the electronic device, whether it is in voice input mode, whether it is in voice playback mode, card insertion status, and the name of the connected WiFi signal.
[0135] For example, the sensing status of the external sensor is displayed by a lighting effect display module installed on the electronic device.
[0136] For example, before the non-contact sensor on the electronic device senses the sensing signal of the external sensor and before receiving the user's input voice signal, an initial prompt voice is played through the voice playback device of the electronic device. The initial prompt voice is, for example, generated by the large language model based on the description information corresponding to the external sensor and preset initial response description information.
[0137] For example, the initial prompt voice includes the name of the user who last interacted with the electronic device.
[0138] The interactive system is also used to execute the steps of any of the above interactive methods, which will not be elaborated here.
[0139] An electronic device according to this application includes a housing and a processor disposed within the housing, the processor being configured to perform the steps of the interaction method as described in any of the preceding claims.
[0140] One embodiment of this application provides a computer-readable storage medium storing a computer program that, when executed, implements the steps of any of the above-described interactive methods. The computer storage medium may be located in an electronic device or on a server; this application does not limit its location.
[0141] It is understood that computer-readable storage media can include: any entity or device capable of carrying computer programs, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), and software distribution media, etc. Computer programs include computer program code. Computer program code can be in the form of source code, object code, executable files, or certain intermediate forms, etc. Computer-readable storage media can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), and software distribution media.
[0142] In some cases of this application, the processor may be a central processing unit (CPU), a graphics processing unit (GPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0143] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the preferred scope of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the function involved, as will be understood by those skilled in the art to which embodiments of this application pertain.
[0144] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a system including a processing module or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).
[0145] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified.
Claims
1. An interaction method, characterized in that, The interaction method comprises When a non-contact sensor on the electronic device senses a signal of an external sensing member, obtaining identification information of the external sensing member; Receiving a voice signal input by a user through a voice input device on the electronic device; According to the description information associated with the identification information of the external sensing member and the voice signal input by the user, outputting a response voice signal through a voice playing device of the electronic device, wherein the response voice signal is generated after a large language model outputs response information according to the description information and the voice signal input by the user, and the response voice signal is used to respond to the voice signal input by the user so that the electronic device can perform voice interaction with the user.
2. The interaction method of claim 1, wherein, The external sensing member comprises a card with an electronic sensing member; and / or, the electronic sensing member is an object with an NFC chip; and / or, the identification information comprises ID information of the external sensing member.
3. The interaction method of claim 1, wherein, The interaction method comprises Transmitting the identification information of the external sensing member to a server, so that the server obtains description information corresponding to the external sensing member according to the identification information of the external sensing member; wherein the description information comprises role description information or scene description information.
4. The interaction method of claim 1, wherein, The interaction method comprises Transmitting the voice signal input by the user to the server, so that the server generates voice recognition text by recognizing the voice signal input by the user; Wherein, the response voice signal is generated after the large language model outputs response information according to the prompt information and the voice recognition text corresponding to the voice signal; Wherein, the description information indicates the role of the large language model when outputting response information for voice interaction with the user.
5. The interaction method of claim 1, wherein, The interaction method comprises According to the current user input audio signal, identifying the voiceprint information of the current user; According to the voiceprint information of the current user, matching the pre-stored user registration information; When the matching is successful, the output audio information comprises the user name in the user registration information; Wherein, the user registration information comprises the user name and user voice feature.
6. The interaction method of any of claim 5, wherein, The user registration information is stored by the following steps: Receiving an instruction of a registered user, obtaining a voice signal input by the user and a user name; According to the voice signal input by the user, generating a user voice feature, associating and storing the user voice feature and the user name.
7. The interaction method of claim 1, wherein, Displaying the interaction content and / or state information of the electronic device through a display device, wherein the interaction content comprises role information or voice content of interaction, and the state information comprises any one or more of the following: battery level information of the electronic device, whether in voice input state, whether in voice playing state, card insertion state, connected WiFi signal name.
8. The interaction method of claim 1, wherein, Displaying the sensing state of the external sensing member through a light effect display module arranged on the electronic device.
9. The interaction method of claim 1, wherein, The method further includes playing, by a voice playing device of the electronic device, an initial prompt voice before a non-contact sensor on the electronic device senses an induction signal of an external induction piece and does not receive a voice signal of a user input, wherein the initial prompt voice is generated according to description information corresponding to the external induction piece.
10. The interaction method of claim 9, wherein, The initial prompt voice is generated by the large language model according to the description information corresponding to the external induction piece and preset initial response description information.
11. The interaction method of claim 9, wherein, The initial prompt voice includes a user name that last interacted with the electronic device.
12. An interactive system, characterized by The electronic device includes one or more processors and a memory, the memory stores a computer program, and the processor executes the computer program to implement the steps of the interaction method of any one of claims 1-11.
13. An electronic device, comprising: The electronic device includes a housing and a processor disposed in the housing, and the processor is configured to execute the method of any one of claims 1-11.