Intelligent voice conversation interaction method, device and system

By reading role identification and user identity information, combining the intelligent agent model and functional interaction model, and adjusting the voice feedback data, the problem that existing intelligent interactive devices cannot adapt to different scenarios is solved, and personalized voice interaction experience and user-friendliness are achieved.

CN120708613APending Publication Date: 2025-09-26SHANGHAI XIANMO INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510886822.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing intelligent interactive devices cannot adapt to the exclusive interaction needs of different functional scenarios, user portraits cannot drive personalized interactions, and audio output devices cannot dynamically adjust the tone and speed of speech, resulting in a lack of immersive user experience.

Method used

The role identification is read through the voice interaction device, the corresponding role attribute set is retrieved, and personalized interaction feedback data is generated by combining the user identity information and voice interaction request. The voice feedback is adjusted using the user intelligent body model and the preset function interaction model to achieve coordinated response of software and hardware.

Benefits of technology

It achieves a personalized voice interaction experience, improves user participation and device adaptability, and enhances user friendliness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708613A_ABST
    Figure CN120708613A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent voice dialogue interaction method, device and system, and the method comprises the steps: reading a tag through a voice interaction device, carrying out the recognition, obtaining a role identifier, and calling a corresponding role attribute set according to the role identifier; generating interaction feedback data through a preset function interaction model according to user identity information and a user voice interaction request, and generating voice communication attributes through a user agent model according to the role attribute set and the user identity information; adjusting the interaction feedback data through the voice communication attribute to generate voice feedback data, and responding to the voice interaction request of the user; wherein the role attribute set comprises expression elements and an interaction normal form.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of smart home appliances, and in particular to an intelligent voice dialogue interaction method, device and system. Background Art

[0002] In the existing field of intelligent interaction, most people use a unified response strategy for voice interaction, which cannot adapt to the exclusive interaction needs of different functional scenarios such as psychological counseling and workplace guidance; the construction of user portraits ignores the differences in role attributes, resulting in the inability of portraits to drive personalized interactions; audio output devices cannot dynamically adjust to the tone / speed of speech that matches the role attributes (such as psychological counseling requires a gentle speed, and workplace training requires a formal tone), and existing hardware only supports fixed audio parameter output; as a result, it is difficult for users to feel involved in the process of intelligent voice interaction and they lack interest in communicating with it, which increases the difficulty of promoting intelligent interactive devices. Summary of the Invention

[0003] The purpose of this application is to provide an intelligent voice dialogue interaction method, device and system to solve the problem of fixed single interaction mode of intelligent interactive voice in the process, provide users with more personalized experience solutions, and at the same time, provide more accurate voice feedback based on the differences of different users to improve user participation.

[0004] To achieve the above-mentioned purpose, the intelligent voice dialogue interaction method provided by the present application specifically includes: obtaining a role identification by reading a tag through a voice interaction device, and retrieving a corresponding role attribute set according to the role identification; generating interaction feedback data through a preset functional interaction model according to user identity information and user voice interaction requests, and generating voice communication attributes through a user intelligent agent model according to the role attribute set and the user identity information; adjusting the interaction feedback data through the voice communication attributes to generate voice feedback data and responding to the user voice interaction request; wherein, the role attribute set includes performance elements and interaction paradigms.

[0005] In the above method, optionally, the construction of the user agent model includes: obtaining the voice dialogue flow between the user and different model objects, extracting the feedback features of the dialogue flow to the performance elements and the adaptation features to the interaction paradigm based on the voice dialogue flow; generating a user perception portrait of each role based on the feedback features and the adaptation features, and fusing the user perception portraits of multiple roles to train the user agent model.

[0006] In the above method, optionally, extracting feedback features for performance elements in the dialogue flow according to the voice dialogue flow includes: detecting occupational data involving preset occupational tags according to the voice dialogue flow, and obtaining an occupational feedback deviation value according to the occupational data and analysis of changes in user voice emotions in the performance features; obtaining a keyword repetition inquiry frequency corresponding to the occupation according to the occupational data analysis; and obtaining the feedback features according to the occupational feedback deviation value and the keyword repetition inquiry frequency.

[0007] In the above method, optionally, extracting adaptation features of the dialogue flow to the interaction paradigm based on the voice dialogue flow includes: detecting the pitch change rate of the user imitating the intelligent agent's voice style and the user's response delay time to different preset topics based on the voice dialogue flow; and obtaining the adaptation features based on the pitch change rate and the response delay time.

[0008] In the above method, optionally, generating interaction feedback data through a preset functional interaction model based on user identity information and user voice interaction request includes: retrieving the voice interaction data of the corresponding user within a preset period based on the user identity information, and performing correlation analysis on the voice interaction request and the voice interaction data to obtain a correlation value; when the correlation value is higher than a preset threshold, generating interaction feedback data through a preset functional interaction model based on the voice interaction data and the user voice interaction request.

[0009] The present application also provides a voice interaction device suitable for the intelligent voice dialogue interaction method, the device including a reading module, a main control module, a communication module and an audio interaction module; the reading module is used to read the label placed on the voice interaction device through near-field communication to obtain a role identification; the communication module is used to connect to an external device through Bluetooth to obtain user identity information and network configuration, and establish a real-time data channel with a cloud server through the network configuration; the main control module is used to transmit the role identification to the cloud server through the real-time data channel, and collect user voice interaction requests through the audio interaction module according to the verification results fed back by the cloud server; the user identity information and the user voice interaction request are transmitted to the cloud server through the real-time data channel, and the user voice interaction request is responded to through the audio interaction module according to the voice feedback data fed back by the cloud server.

[0010] The present application also provides a voice interaction system including a voice interaction device, the system including a cloud server, the cloud server being used to generate interaction feedback data through a preset functional interaction model based on received user identity information and user voice interaction requests; retrieve a corresponding role attribute set based on a received role identifier, and generate voice communication attributes through a user agent model based on the role attribute set and the user identity information; adjust the interaction feedback data through the voice communication attributes to generate voice feedback data, and provide the voice feedback data to the voice interaction device.

[0011] In the above-mentioned voice interaction system, optionally, the cloud server includes a verification device, which is used to verify the legitimacy of the role identification and feed back the verification result to the voice interaction device when the verification is passed.

[0012] The present application also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned method when executing the computer program.

[0013] The present application also provides a computer-readable storage medium, which stores a computer program for executing the above method.

[0014] The present application also provides a computer program product, comprising a computer program / instruction, which implements the steps of the above method when executed by a processor.

[0015] The beneficial technical effects of this application are: achieving differentiated responses to role attributes through a software-hardware collaborative architecture driven by role attributes; utilizing the portrait-to-attribute collaborative evolution method to construct user portraits based on role differentiation, and utilizing the portrait-feedback role attribute optimization method to make the conversation content more in line with user preferences and improve user friendliness. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings described herein are used to provide a further understanding of the present application, constitute a part of the present application, and do not constitute a limitation of the present application. In the drawings:

[0017] Figure 1 A flowchart of an intelligent voice dialogue interaction method provided in one embodiment of the present application;

[0018] Figure 2 A schematic diagram of the process of constructing a user agent model provided in one embodiment of the present application;

[0019] Figure 3 A schematic diagram of a feedback feature acquisition process provided in an embodiment of the present application;

[0020] Figure 4A schematic diagram of the adaptive feature acquisition process provided in one embodiment of the present application;

[0021] Figure 5 A schematic diagram of a flow chart for generating interactive feedback data provided in an embodiment of the present application;

[0022] Figure 6 A schematic diagram of the structure of a voice interaction device provided in one embodiment of the present application;

[0023] Figure 7 A schematic diagram of the structure of a voice interaction system provided in one embodiment of the present application;

[0024] Figure 8 A schematic diagram of the structure of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0025] The following will describe in detail the implementation methods of this application in conjunction with the accompanying drawings and examples, so that the application can fully understand how technical means are used to solve technical problems and achieve technical effects, and implement them accordingly. It should be noted that as long as there is no conflict, the various embodiments and the various features in each embodiment of this application can be combined with each other, and the resulting technical solutions are all within the scope of protection of this application.

[0026] Additionally, the steps shown in the flowcharts of the accompanying drawings may be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowcharts, in some cases the steps shown or described may be performed in an order different from that shown.

[0027] Please refer to Figure 1 As shown, the intelligent voice dialogue interaction method provided by this application specifically includes:

[0028] S101 reads the tag identification by the voice interaction device to obtain a role identification, and retrieves a corresponding role attribute set according to the role identification;

[0029] S102 generates interaction feedback data through a preset functional interaction model according to the user identity information and the user voice interaction request, and generates voice communication attributes through a user agent model according to the role attribute set and the user identity information;

[0030] S103 adjusts the interaction feedback data by using the voice communication attributes to generate voice feedback data and responds to the user voice interaction request; wherein the role attribute set includes performance elements and interaction paradigms.

[0031] In actual work, the application process of the above embodiment is as follows: first, the role identification is read and the attribute set is retrieved. This process can be achieved through a hardware trigger mechanism, that is, the user places a physical item (action figure, doll, etc.) with an NFC or RFID tag on the preset sensing area of ​​the interactive device; the interactive device reads the role identification data stored in the above tag through a built-in card reader component (a PN532 card reader can be used); thereafter, the role identification and the collected user identity information are provided to the cloud through the main control chip (an ESP32-S3 main control chip can be used) to obtain the corresponding role attribute set, wherein the user identity information can be obtained by interacting with the user's mobile terminal during the network pairing process using the voice interaction device, or by other means, such as passwords and fingerprints in the hardware, or by voice, facial recognition and other data verification, which is not further limited in this application.

[0032] After determining the above-mentioned role attribute set, the user's conversation history data can be retrieved according to the user's identity information, and the user's intention can be judged by performing association analysis in combination with the current voice request, thereby providing the corresponding interaction feedback data. In this process, a functional interaction model can be introduced. The functional interaction model is constructed by training the learning algorithm and the historical sample parameters of the corresponding functional field. This process can refer to the existing interaction logic, and this application does not impose any restrictions on this. Afterwards, it is necessary to generate voice communication attributes based on the user intelligent body model. In this process, parameters need to be dynamically generated based on the cross-role user portrait library, and the above-mentioned interaction feedback data needs to be adjusted based on the generated parameters.

[0033] Specifically, the method of adjusting the interactive feedback data may include text adjustment and voice adjustment. After the voice feedback data is generated after adjustment, the audio interaction module can be used for feedback interaction. Among them, the audio interaction module can adopt the NS4168 audio chip. The adjusted digital audio flows through the LM321DTR op amp to enhance the signal stability and is played through the speaker. At the same time, the LED indicator shows a green breathing light (indicating that the conversation is in progress). The user short presses the SD8233B capacitor button to immediately interrupt the playback (response delay <50ms).

[0034] Please refer to Figure 2 As shown, in one embodiment of the present application, the construction of the user agent model includes:

[0035] S201: obtaining a voice dialogue flow between a user and different model objects, and extracting feedback features of the dialogue flow on performance elements and adaptation features to the interaction paradigm based on the voice dialogue flow;

[0036] S202 generates a user perception portrait of each role based on the feedback features and the adaptation features, and integrates the user perception portraits of multiple roles to train the user agent model.

[0037] Please refer to Figure 3 As shown, in the above embodiment, extracting feedback features of performance elements in the dialogue flow according to the voice dialogue flow includes:

[0038] S301 detects occupational data related to a preset occupational tag according to the voice dialogue flow, and obtains an occupational feedback deviation value according to the occupational data and the change of the user's voice emotion in the performance characteristics;

[0039] S302 obtains the frequency of repeated keyword inquiries of the corresponding occupation according to the occupation data analysis;

[0040] S303 obtains the feedback feature according to the occupational feedback deviation value and the keyword repeated inquiry frequency.

[0041] For further information, please refer to Figure 4 As shown, extracting adaptation features of the dialogue flow to the interaction paradigm according to the voice dialogue flow includes:

[0042] S401 detects the pitch change rate of the user imitating the voice style of the intelligent agent and the user's response delay time to different preset topics according to the voice dialogue flow;

[0043] S402 obtains the adaptation feature according to the pitch change rate and the response delay time.

[0044] In actual work, the above embodiment can include three steps as a whole, which are as follows: first, multi-role voice dialogue flow collection is performed, during which the dialogue flows of users corresponding to roles with different role tags are collected separately, and these dialogue flows are stored separately by role.

[0045] Feedback features and adaptive features are then extracted, including the calculation of professional feedback deviation values, analysis of the slope of emotional change, and generation of feedback features. The analysis of the slope of emotional change involves using a pre-trained emotional model (such as BERT-Emotion) to output the emotional value of each sentence, then calculating the emotional change rate 30 seconds before and after the appearance of the professional topic, and then generating feedback features. Regarding the adaptive feature extraction of the interactive paradigm, it mainly includes the tone imitation pitch change rate, the topic response delay time, and the structured output of adaptive features. Among them, the tone imitation pitch change rate is obtained by extracting the fundamental frequency of the user's voice and the fundamental frequency of the agent's voice, and then calculating the dynamic follow-up rate. For the topic response delay time, the user's feedback time for different topics can be determined based on the predefined delay time. The above content is then used to form the adaptive feature structured output data, namely the tone imitation degree and the delay time of different topics.

[0046] Finally, user perception portrait generation and model training can be carried out. The user perception portrait generation process can include portrait data fusion and multi-role portrait fusion training. The portrait data fusion generates corresponding portrait output data according to the weight values ​​of the feedback feature weights and adaptive feature weight configurations of different role types. Subsequently, during the model training process, training can be performed through a dual-channel Transformer, where channel 1 inputs the multi-role portrait vector and channel 2 inputs the real-time voice interaction request, and then the user agent model is designed and trained according to the preset loss function.

[0047] Please refer to Figure 5 As shown, in one embodiment of the present application, generating interaction feedback data through a preset functional interaction model according to user identity information and user voice interaction request includes:

[0048] S501 retrieves voice interaction data of the corresponding user within a preset period according to the user identity information, and performs correlation analysis on the voice interaction request and the voice interaction data to obtain a correlation value;

[0049] S502: When the correlation value is higher than a preset threshold, interaction feedback data is generated according to the voice interaction data and the user voice interaction request through a preset functional interaction model.

[0050] In this embodiment, the main purpose is to associate the user's historical data with the current interaction request to avoid the user from repeatedly inputting a large amount of data to access the previous discussion topic. To this end, this application needs to first perform a relevance judgment during the communication process with the user to determine whether the user's current request is related to the previous content, and then reply. In this way, the user will not have an obvious sense of disconnection during the voice interaction process, and the linkage interaction of multiple roles also greatly improves the user experience.

[0051] Please refer to Figure 6 As shown, the present application also provides a voice interaction device suitable for the intelligent voice dialogue interaction method, the device including a reading module, a main control module, a communication module and an audio interaction module; the reading module is used to read the label placed on the voice interaction device through near-field communication to obtain a role identification; the communication module is used to connect to the external device through Bluetooth to obtain user identity information and network configuration, and establish a real-time data channel with the cloud server through the network configuration; the main control module is used to transmit the role identification to the cloud server through the real-time data channel, and collect user voice interaction requests through the audio interaction module according to the verification results fed back by the cloud server; the user identity information and the user voice interaction request are transmitted to the cloud server through the real-time data channel, and the user voice interaction request is responded to through the audio interaction module according to the voice feedback data fed back by the cloud server.

[0052] Since the principle of solving the problem by this device is similar to that of the intelligent voice dialogue interaction method, the implementation of this device can refer to the implementation of the intelligent voice dialogue interaction method, and the repeated parts will not be repeated.

[0053] Please refer to Figure 7 As shown, the present application also provides a voice interaction system including a voice interaction device, the system including a cloud server, the cloud server is used to generate interaction feedback data through a preset functional interaction model based on the received user identity information and user voice interaction request; retrieve the corresponding role attribute set based on the received role identification, generate voice communication attributes through the user intelligent body model based on the role attribute set and the user identity information; adjust the interaction feedback data through the voice communication attributes to generate voice feedback data, and provide the voice feedback data to the voice interaction device. Wherein, the cloud server includes a verification device, the verification device is used to verify the legitimacy of the role identification, and feedback the verification result to the voice interaction device when the verification is passed.

[0054] In this system, all data analysis items are set up on the cloud server in order to effectively utilize the computing power of the cloud while improving the security of user privacy data and preventing data leakage due to intrusion of local devices; and the process of role legitimacy verification can also effectively prevent criminals from using disguised intrusion methods to obtain user data.

[0055] The beneficial technical effects of this application are: achieving differentiated responses to role attributes through a software-hardware collaborative architecture driven by role attributes; utilizing the portrait-to-attribute collaborative evolution method to construct user portraits based on role differentiation, and utilizing the portrait-feedback role attribute optimization method to make the conversation content more in line with user preferences and improve user friendliness.

[0056] The present application also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned method when executing the computer program.

[0057] The present application also provides a computer-readable storage medium, which stores a computer program for executing the above method.

[0058] The present application also provides a computer program product, comprising a computer program / instruction, which implements the steps of the above method when executed by a processor.

[0059] like Figure 8As shown, the electronic device 600 may further include: a communication module 110, an input unit 120, an audio processor 130, a display 160, and a power supply 170. It is worth noting that the electronic device 600 does not necessarily have to include Figure 8 In addition, the electronic device 600 may also include all components shown in Figure 8 For components not shown, reference may be made to the prior art.

[0060] like Figure 8 As shown, the central processing unit 100 is sometimes also referred to as a controller or an operation control unit, and may include a microprocessor or other processor device and / or logic device. The central processing unit 100 receives inputs and controls the operations of various components of the electronic device 600 .

[0061] Memory 140 may be, for example, one or more of a cache, flash memory, hard drive, removable media, volatile memory, non-volatile memory, or other suitable devices. It may store the aforementioned failure-related information and may also store programs that execute the relevant information. The CPU 100 may execute the programs stored in memory 140 to implement information storage or processing.

[0062] The input unit 120 provides input to the CPU 100. The input unit 120 may be, for example, a keypad or touch input device. The power supply 170 is used to provide power to the electronic device 600. The display 160 is used to display objects such as images and text. The display may be, for example, an LCD display, but is not limited thereto.

[0063] The memory 140 may be a solid-state memory, such as a read-only memory (ROM), a random access memory (RAM), or a SIM card. Alternatively, it may be a memory that retains information even when power is off, can be selectively erased, and is provided with more data. Examples of such memory are sometimes referred to as EPROMs. The memory 140 may also be some other type of device. The memory 140 includes a buffer memory 141 (sometimes referred to as a buffer). The memory 140 may include an application / function storage unit 142 for storing application programs and function programs or processes for executing the operations of the electronic device 600 via the central processing unit 100.

[0064] The memory 140 may also include a data storage unit (data 143) for storing data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit (driver 144) of the memory 140 may include various driver programs for communication functions of the electronic device and / or for executing other functions of the electronic device (such as messaging applications, address book applications, etc.).

[0065] The communication module 110 is a transmitter / receiver 110 that transmits and receives signals via an antenna 111. The communication module (transmitter / receiver) 110 is coupled to the central processor 100 to provide input signals and receive output signals, which may be the same as in a conventional mobile communication terminal.

[0066] Based on different communication technologies, multiple communication modules 110 may be provided in the same electronic device, such as a cellular network module, a Bluetooth module, and / or a wireless local area network module. The communication module (transmitter / receiver) 110 is also coupled to a speaker 131 and a microphone 132 via an audio processor 130 to provide audio output via the speaker 131 and receive audio input from the microphone 132, thereby implementing common telecommunication functions. The audio processor 130 may include any suitable buffer, decoder, amplifier, etc. Furthermore, the audio processor 130 is also coupled to the central processing unit 100, enabling local recording via the microphone 132 and playback of stored audio via the speaker 131.

[0067] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0068] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0069] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0070] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0071] The specific embodiments described above further illustrate the purpose, technical solutions and beneficial effects of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. An intelligent voice dialogue interaction method, characterized in that: The method comprises: The voice interaction device reads the tag to identify the role and obtains the role identification, and retrieves the corresponding role attribute set according to the role identification; Generate interaction feedback data through a preset functional interaction model according to the user identity information and the user voice interaction request, and generate voice communication attributes through a user agent model according to the role attribute set and the user identity information; Adjusting the interaction feedback data by using the voice communication attribute to generate voice feedback data and responding to the user voice interaction request; The role attribute set includes performance elements and interaction paradigms.

2. The method according to claim 1, characterized in that The construction of the user agent model includes: Acquire the voice dialogue flow between the user and different model objects, and extract the feedback features of the dialogue flow on the performance elements and the adaptation features to the interaction paradigm based on the voice dialogue flow; Based on the feedback features and adaptation features, the user perception portrait of each role is generated, and the user agent model is trained by integrating the user perception portraits of multiple roles.

3. The method according to claim 2, characterized in that Extracting feedback features of performance elements in the speech dialogue flow according to the speech dialogue flow includes: Detecting occupational data involving preset occupational tags according to the voice dialogue flow, and obtaining an occupational feedback deviation value according to the occupational data and the emotional changes of the user's voice in the performance characteristics; Obtaining a frequency of repeated keyword inquiries for a corresponding occupation based on the occupation data analysis; The feedback feature is obtained according to the occupational feedback deviation value and the keyword repeated inquiry frequency.

4. The method according to claim 2, characterized in that Extracting adaptation features of the dialogue flow to the interaction paradigm according to the voice dialogue flow includes: Detecting the pitch change rate of the user's imitation of the agent's voice style and the user's response delay time to different preset topics based on the voice dialogue flow; The adaptation feature is obtained according to the pitch change rate and the response delay time.

5. The method according to claim 1, wherein The interaction feedback data generated by the preset functional interaction model based on user identity information and user voice interaction requests includes: Retrieving voice interaction data of the corresponding user within a preset period according to the user identity information, and performing correlation analysis on the voice interaction request and the voice interaction data to obtain a correlation value; When the correlation value is higher than a preset threshold, interaction feedback data is generated according to the voice interaction data and the user voice interaction request through a preset functional interaction model.

6. A voice interaction device suitable for the intelligent voice dialogue interaction method according to any one of claims 1 to 5, characterized in that: The device includes a reading module, a main control module, a communication module and an audio interaction module; The reading module is used to read the label placed on the voice interaction device through near field communication to obtain the role identification; The communication module is used to connect to the external device via Bluetooth, obtain user identity information and network configuration, and establish a real-time data channel with the cloud server through the network configuration; The main control module is used to transmit the role identification to the cloud server through the real-time data channel, and collect the user voice interaction request through the audio interaction module according to the verification result feedback from the cloud server; transmit the user identity information and the user voice interaction request to the cloud server through the real-time data channel, and respond to the user voice interaction request through the audio interaction module according to the voice feedback data feedback from the cloud server.

7. A voice interaction system comprising the voice interaction device according to claim 6, characterized in that: The system includes a cloud server, which is used to generate interaction feedback data through a preset functional interaction model based on received user identity information and user voice interaction requests; retrieve a corresponding role attribute set based on the received role identifier, and generate voice communication attributes through a user agent model based on the role attribute set and the user identity information; The interactive feedback data is adjusted according to the voice communication attribute to generate voice feedback data, and the voice feedback data is provided to the voice interaction device.

8. The voice interaction system according to claim 7, characterized in that: The cloud server includes a verification device, which is used to verify the legitimacy of the role identification and feed back the verification result to the voice interaction device when the verification passes.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.