Virtual Assistant Interaction Method and System for Hotel Rooms
By identifying and classifying interactive information of different modalities, and using task queues and processing modules to respond, the problem that virtual assistants are difficult to deal with multiple interactive information of different modalities is solved, and efficient and accurate response and personalized services are achieved.
Patent Information
- Application Number
- CN202411480450.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-23
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-10-23
AI Technical Summary
The virtual assistant receives multiple interactive information of different modes in a single hotel room, and these interactive information are not related, making it difficult for the virtual assistant to understand and respond to these interactive information at the same time.
By identifying the interactive content entered by the user within the preset period, determining its interaction type, and listing the interaction information of different modes into the task queue, responding to these interactive contents through the preset processing module. Identify the coordinated association information to generate modal conflict content, and dynamically adjust it according to the conflict content, represent the dependencies between different modalities through a graph model, and generate a response strategy.
Effectively manage interactive requests from multiple users, avoid instruction confusion and conflict handling, ensure that each user's needs are independently processed, improve the accuracy and consistency of responses, and provide a smooth and personalized user experience.
Smart Images

Figure CN119005994B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of interactive data processing, and particularly to a virtual assistant interaction method and system for hotel rooms. Background Art
[0002] With the continuous improvement of information technology and people's living standards, guests have put forward higher requirements for the check-in experience and management level of hotels. In this industry background, traditional hotel information service systems can no longer meet the requirements of new hotel management, and at the same time, it is difficult to solve problems such as many business links, low communication efficiency, inability to meet personalized service needs, and poor guest room service experience. Therefore, virtual assistants are added to existing hotel services to interact with residents in hotel rooms to meet the effective needs of residents when they are in the rooms.
[0003] If a virtual assistant receives multiple different-modal interaction messages in a single hotel room and these interaction messages are not related, it is difficult for the virtual assistant to understand and respond to these interaction messages simultaneously, and it cannot meet the needs of the room residents in a timely manner. Summary of the Invention
[0004] The present invention aims to solve the problem that a virtual assistant receives multiple different-modal interaction messages in a single hotel room and these interaction messages are not related, resulting in difficulty for the virtual assistant to understand and respond to these interaction messages simultaneously, and provides a virtual assistant interaction method and system for hotel rooms.
[0005] The present invention adopts the following technical means to solve the technical problems:
[0006] The present invention provides a virtual assistant interaction method for hotel rooms, including:
[0007] Based on the interaction permissions pre-unlocked for the hotel room, identify the interaction content input by the user into the hotel room within a preset time period;
[0008] Judge whether the interaction content belongs to the same interaction type, where the interaction type specifically includes image interaction, voice interaction, and text interaction;
[0009] If not, list the interaction content in the task queue preset by the virtual assistant, activate the processing module preset for the hotel room, allocate the interaction content from the task queue to the processing module, respond to the interaction content in the hotel room through the processing module, identify the collaborative association information pre-classified for the interaction content, and generate modal conflict content for the interaction content according to the collaborative association information, where the collaborative association information specifically includes keyword matching, spatial location matching, and task target matching;
[0010] Judge whether the modal conflict content can be dynamically adjusted;
[0011] If possible, identify the input data of the modal conflict content, represent the interaction information of different modalities as nodes in a preset graph model based on the input data, represent the edges between the nodes as the dependency relationships between different modalities, and generate a response strategy for the modal conflict content based on the nodes and the dependency relationships, where the input data specifically includes visual images, voice instructions, and tactile feedback.
[0012] Further, in the step of activating the preset processing module of the hotel room, allocating the interaction content from the task queue to the processing module, and responding to the interaction content in the hotel room through the processing module, it further includes:
[0013] Present the control content of the preset linked rooms in the hotel room based on the cross-room linkage setting preset by the user through the hotel terminal;
[0014] Determine whether the control content can be synchronized to the hotel room;
[0015] If possible, make the virtual assistant send a corresponding task request according to the interaction content, identify the pre-recorded content of the linked room, generate the task request from the linked room based on the pre-recorded content, and coordinate the hotel room and the linked room to execute the task request synchronously through the virtual assistant.
[0016] Further, before the step of generating the modal conflict content of the interaction content according to the collaborative association information, it further includes:
[0017] Obtain the task request obtained by converting the modal information based on the modal information pre-generated by the hotel room for the interaction content;
[0018] Determine whether the task request conforms to the preset execution content of the hotel room;
[0019] If not, identify the missing modality of the modal information, generate the content to be supplemented for the modal information according to the missing modality, obtain the response feedback of the user to the content to be supplemented, and reconstruct the response feedback into the modal information according to the interaction type, and then convert it to obtain the task request again.
[0020] Further, in the step of identifying the input data of the modal conflict content, it further includes:
[0021] Obtain the type of features input by the user into the hotel room through the virtual assistant based on the collection features preset by the hotel room;
[0022] Determine whether the type of features exceeds the preset upper limit;
[0023] If so, identify the modal conflict points of the feature type through the virtual assistant, generate the interceptable paragraph information of the modal conflict content according to the modal conflict points, transfer the interceptable paragraph information to a preset blank paragraph, construct the to-be-uploaded interactive content corresponding to the interceptable paragraph information, and synchronize the to-be-uploaded interactive content to the virtual assistant.
[0024] Further, in the step of determining whether the interactive content belongs to the same interaction type, it further includes:
[0025] Identify the corresponding interactive device in the hotel room based on the interactive content;
[0026] Determine whether the interactive device receives at least two pieces of the interactive content within a preset time period;
[0027] If so, obtain the interaction type of the interactive content, and classify the interactive content into the preset interaction options corresponding to the virtual assistant, where the interaction options specifically include query-type interaction, request-type interaction, and setting-type interaction.
[0028] Further, in the step of determining whether the modal conflict content can be dynamically adjusted, it further includes:
[0029] Obtain the conflict type of the modal conflict content, where the conflict type specifically includes space conflict, priority conflict, and time conflict;
[0030] Determine whether the conflict type is diverted to the same interaction path;
[0031] If not, generate the influencing device corresponding to the modal conflict content through the virtual assistant, identify the conflict range of the modal conflict content for the influencing device, and divide the dynamic adjustment interval of the influencing device based on the conflict range.
[0032] Further, in the step of identifying the interactive content input by the user into the hotel room within a preset time period based on the interactive permission for pre-unlocking the hotel room, it further includes:
[0033] Identify the total number of users staying in the hotel room;
[0034] Determine whether the total number of users reaches a preset threshold;
[0035] If so, based on the feature information pre-entered by the user, obtain the priority configuration of the user for other users in the hotel room, collect the interactive content received in the hotel room, and divide the execution sequence of the interactive content according to the priority configuration.
[0036] The present invention also provides a virtual assistant interaction system for hotel guest rooms, including:
[0037] An identification unit for identifying the interaction content input by a user into the hotel guest room within a preset time period based on the interaction permission for pre-unlocking the hotel guest room;
[0038] A judgment unit for judging whether the interaction content belongs to the same interaction type, where the interaction type specifically includes image interaction, voice interaction, and text interaction;
[0039] An execution unit for, if not, adding the interaction content to a task queue preset by the virtual assistant, activating a processing module preset for the hotel guest room, allocating the interaction content from the task queue to the processing module, responding to the interaction content in the hotel guest room through the processing module, identifying the collaborative association information pre-classified for the interaction content, and generating modal conflict content for the interaction content according to the collaborative association information, where the collaborative association information specifically includes keyword matching, spatial location matching, and task objective matching;
[0040] A second judgment unit for judging whether the modal conflict content can be dynamically adjusted;
[0041] A second execution unit for, if it can, identifying the input data of the modal conflict content, representing interaction information of different modalities as nodes in a preset graph model based on the input data, representing the edges between the nodes as the dependency relationships between different modalities, and generating a response strategy for the modal conflict content based on the nodes and the dependency relationships, where the input data specifically includes visual images, voice commands, and tactile feedback.
[0042] Further, the execution unit further includes:
[0043] A presentation subunit for presenting the control content of a preset linked guest room in the hotel guest room based on the cross-room linkage setting preset by the user through the hotel terminal;
[0044] A judgment subunit for judging whether the control content can be synchronized to the hotel guest room;
[0045] An execution subunit for, if it can, making the virtual assistant send a corresponding task request according to the interaction content, identifying the pre-recorded content of the linked guest room, generating the task request from the linked guest room based on the pre-recorded content, and coordinating the hotel guest room and the linked guest room to synchronously execute the task request through the virtual assistant.
[0046] Further, it further includes:
[0047] An acquisition unit, configured to acquire a task request obtained by converting the modal information based on the modal information pre-generated for the interactive content of the hotel room;
[0048] A third judgment unit, configured to judge whether the task request conforms to the preset execution content of the hotel room;
[0049] A third execution unit, configured to, if not, identify the missing modality of the modal information, generate the content to be complemented for the modal information according to the missing modality, acquire the response feedback of the user to the content to be complemented, reconstruct the response feedback into the modal information according to the interaction type, and convert it again to obtain the task request.
[0050] The present invention provides a virtual assistant interaction method and system for a hotel room, which has the following beneficial effects:
[0051] The present invention classifies different types of interactive content by identification, more accurately understands the user's intention, reduces the possibility of misunderstanding or misoperation, and at the same time lists different interactive content in a task queue and responds through a preset processing module, enabling the system to effectively allocate resources, avoid resource overload or response delay caused by processing multiple tasks simultaneously, and through identifying and processing modal conflict content, the system can dynamically adjust the conflict between different modalities to ensure that the user obtains a smooth and consistent experience when interacting with the virtual assistant. When the system identifies the modal conflict of the interactive content, by analyzing the dependency relationship between modalities to generate a response strategy, the system can maintain a high degree of flexibility in complex multi-task scenarios, automatically adjust the priority and execution order of tasks, and avoid task conflicts. Description of the Drawings
[0052] Figure 1 It is a schematic flowchart of an embodiment of the virtual assistant interaction method for a hotel room according to the present invention;
[0053] Figure 2 It is a structural block diagram of an embodiment of the virtual assistant interaction system for a hotel room according to the present invention. Detailed Embodiments
[0054] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. The implementation, functional features and advantages of the present invention will be further described in conjunction with the embodiments with reference to the drawings.
[0055] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0056] Referring to the attached Figure 1 , a virtual assistant interaction method for a hotel guest room in an embodiment of the present invention includes:
[0057] S1: Based on the interaction permission for pre-unlocking the hotel guest room, identify the interaction content input by the user into the hotel guest room within a preset time period;
[0058] S2: Determine whether the interaction content belongs to the same interaction type, where the interaction type specifically includes image interaction, voice interaction, and text interaction;
[0059] S3: If not, include the interaction content in the task queue preset by the virtual assistant, activate the processing module preset in the hotel guest room, allocate the interaction content from the task queue to the processing module, respond to the interaction content in the hotel guest room through the processing module, identify the collaborative association information pre-classified for the interaction content, and generate the modal conflict content of the interaction content according to the collaborative association information, where the collaborative association information specifically includes keyword matching, spatial position matching, and task objective matching;
[0060] S4: Determine whether the modal conflict content can be dynamically adjusted;
[0061] S5: If it can, identify the input data of the modal conflict content, represent the interaction information of different modalities as nodes in a preset graph model according to the input data, represent the edges between the nodes as the dependency relationships between different modalities, and generate a response strategy for the modal conflict content based on the nodes and the dependency relationships, where the input data specifically includes visual images, voice commands, and tactile feedback.
[0062] In this embodiment, after the system confirms that the user has checked into the hotel room, the hotel room will unlock the pre-set interaction permissions, identify the interaction content input by the user into the hotel room within the pre-set time period, and then the system determines whether these interaction contents belong to the same interaction type. The interaction types specifically include image interaction, sound interaction, and text interaction, so as to execute corresponding steps. For example, when the system determines that the interaction content input by the user into the hotel room belongs to the same interaction type, the system will consider that all inputs are concentrated in the same type, and the system does not need to perform cross-modal parsing or processing. The system will merge the contents of the same interaction type into an overall task and process them centrally. If the interaction contents have a time sequence, the system can execute these contents in the order of user input. At the same time, according to the unified interaction type, the feedback method can be optimized. For example, when the user controls the TV in different ways (such as voice and text), the system will give all feedbacks in a unified way through voice or screen display to simplify the user experience. And if the contents of the same interaction type contain repeated instructions, the system can identify and reduce redundancy to avoid performing the same operation multiple times. For example, if the user requests to play music twice continuously through voice, then the system will only execute it once. For example, when the system determines that the interaction content input by the user into the hotel room does not belong to the same interaction type, at this time the system will consider that there are multiple users in the hotel room simultaneously initiating interactions with the virtual assistant. The system will list the different interaction contents in the task queue pre-set by the virtual assistant, activate the processing module pre-set in the hotel room, allocate these interaction contents from the task queue to the processing module, and respond to these interaction contents in the hotel room through the processing module. Identify the pre-classified collaborative association information of these interaction contents. The collaborative association information specifically includes keyword matching, spatial location matching, and task target matching, and generate modal conflict contents of the interaction contents according to different collaborative association information. By listing different interaction contents in the task queue and allocating them to the processing module, the system can effectively manage the interaction requests of multiple users. It allows the system to avoid instruction confusion and processing conflicts when multiple users initiate interactions simultaneously, ensuring that the needs of each user in the hotel room can be independently and appropriately processed. At the same time, by identifying the collaborative association information of the interaction contents (such as keyword matching, spatial location matching, and task target matching), the intention and context of different interaction contents can be accurately distinguished. This precise identification and classification ability helps the system better understand user needs, so as to provide more relevant and personalized services for different users. And when a modal conflict is detected, conflict contents can be generated according to the collaborative association information and the processing strategy can be dynamically adjusted. This ability allows the system to flexibly handle conflicts when facing multiple interaction types, ensuring the interoperability and coordination between different modal inputs, thereby improving the accuracy and consistency of responses. Then the system determines whether these modal conflict contents can be dynamically adjusted to execute corresponding steps;For example, when the system determines that the modal conflict content of the interactive content cannot be dynamically adjusted, the system will consider that two or more conflicting interactive contents can only select one to be implemented. The system will evaluate the priority of the conflicting interactive contents, actively feedback the conflict situation to the user, and request the user to make a choice. The choices include the user's preference settings, the urgency of the interactive content, the time sequence of the user's input, and the importance of the target task of the content. At the same time, the system will prompt the user with the selected interactive content and explain the other contents that have not been executed, through voice prompts, text messages or image displays, etc., to ensure that the user understands the system's decision and processing results, and to avoid confusion or dissatisfaction caused by the system not executing a certain content. And if the unselected interactive content is equally important to the user, the system will prompt the user after the main task is executed and suggest whether compensation execution is needed; For example, when the system determines that the modal conflict content of the interactive content can be dynamically adjusted, at this time the system will consider that the conflicting interactive contents can be split and carried out simultaneously. The system will identify the input data of the modal conflict content. The input data specifically includes visual images, voice commands and tactile feedback. According to different input data, the interactive information of different modalities is represented as nodes in a preset graph model, and the edges between the nodes are represented as the dependency relationships between different modalities. Based on the nodes and the dependency relationships, a response strategy for the modal conflict content is generated; By splitting and processing the conflicting interactive contents in parallel, the system can flexibly handle the conflict situation of multi-modal information, avoid operation interruption or unresponsiveness caused by modal conflicts, dynamically adjust the processing strategy according to the real-time situation, better adapt to complex user needs, and at the same time represent the interactive information of different modalities as nodes in a graph model and process them through the dependency relationships between the nodes, so that the system can more efficiently allocate and utilize resources, reduce the waste of system resources, improve the overall operation efficiency, and be able to process more tasks in the same time period. And the use of the graph model allows the system to be more meticulous when dealing with complex scenarios. For example, when there are complex dependency relationships between visual images and voice commands, the system can accurately analyze and process these relationships, ensure that each interactive content can get appropriate responses and processing, generate more reasonable response strategies, and still make decisions that meet the user's expectations in case of conflicts, reducing misunderstandings and misoperations.;
[0063] It should be noted that the input data for identifying the modal conflict content, representing the interactive information of different modalities as nodes in a preset graph model according to the input data, representing the edges between the nodes as the dependency relationships between different modalities, and generating a response strategy for the modal conflict content based on the nodes and the dependency relationships are specifically exemplified as follows:
[0064] Suppose Mr. Zhang, a user staying in a smart guest room of a high-end hotel, enters the room in the evening and is ready to relax. The room is equipped with a virtual assistant that can interact with the user in multiple modalities such as voice, vision, and touch. At this time, Mr. Zhang hopes to make the room more comfortable, so he issues two different interaction requests simultaneously: while saying "Open the curtains so that I can see the view outside", he points at the TV with his finger, wanting to watch a movie.
[0065] At this time, the virtual assistant captures Mr. Zhang's voice command "Open the curtains so that I can see the view outside" through the microphone array in the room. At the same time, the camera in the room captures Mr. Zhang's gesture, and the system recognizes that he is pointing in the direction of the TV. Meanwhile, the system creates a "voice node" for Mr. Zhang's voice command, which records Mr. Zhang's voice content, voice emotion (because Mr. Zhang's tone is very relaxed, indicating that he is in a relaxed state), and relevant timestamp information. And the system creates a "visual node" for Mr. Zhang's gesture, which records the direction of the gesture, the pointing of the finger, and the object he is pointing at - the TV. The system first identifies the "curtains" and "view" mentioned in Mr. Zhang's voice command through keyword analysis and context understanding, which indicates that Mr. Zhang hopes to see the view outside and may hope that the light changes to create a comfortable atmosphere. However, the visual input shows that Mr. Zhang is pointing at the TV, which is not directly related to the content of the voice command (curtains and view). Therefore, the system creates a dependency edge between the "voice node" and the "visual node", but the weight of this dependency is relatively low because the two nodes point to different targets. The system will analyze that there is a certain conflict due to the different targets pointed to by the voice command and the gesture. For this reason, the system must decide how to coordinate these two requests by generating a dynamic adjustment strategy to decide whether to respond to these two requests separately. For example, the system gives priority to responding to Mr. Zhang's voice command because the voice expresses a clear need. Then the system sends a command to open the curtains and adjusts the lights so that the room light better presents the external view. At the same time, the system detects that Mr. Zhang is pointing at the TV but there is no clear voice command to match it. Therefore, to ensure that the user's intention is met, the system will ask Mr. Zhang through the virtual assistant: "Do you want to turn on the TV and play your favorite movie?" After Mr. Zhang hears the system's feedback and confirms: "Yes, please play a relaxing movie.", the system then executes the instruction to play the movie, ensuring that all of Mr. Zhang's needs are met.
[0066] By decomposing conflicting modal inputs into independent tasks and processing them separately, the system demonstrates its adaptability to complex user needs, accurately responding to both of Mr. Zhang's requirements, ensuring the consistency of the user experience. At the same time, through confirmation by asking, the system avoids possible misoperations (such as directly playing TV content instead of playing a movie first), which enhances the user's trust and satisfaction in the system. And the system optimizes resource allocation by modeling the dependency relationships between multi-modal inputs, ensuring that multiple user requests are responded to in the shortest possible time;
[0067] In summary, this processing flow can simultaneously achieve the multi-tasking ability of the virtual assistant and the interaction intelligence of the user, provide accurate and personalized services in complex scenarios, effectively solve the problem of multi-modal information conflict, and significantly improve the quality of the user experience when staying in a hotel guest room.
[0068] In this embodiment, when activating the preset processing module in the hotel guest room and allocating the interactive content from the task queue to the processing module, in step S3 of responding to the interactive content in the hotel guest room by the processing module, it further includes:
[0069] S31: Based on the cross-room linkage setting preset by the user through the hotel terminal, present the control content of the preset linked guest room in the hotel guest room;
[0070] S32: Determine whether the control content can be synchronized to the hotel guest room;
[0071] S33: If it can, then according to the interactive content, make the virtual assistant send a corresponding task request, identify the pre-recorded content of the linked guest room, generate the task request from the linked guest room based on the pre-recorded content, and coordinate the hotel guest room and the linked guest room to execute the task request synchronously through the virtual assistant.
[0072] In this embodiment, based on the cross-room linkage settings preset by the user through the hotel terminal, the system presents the control content of the preset linked guest rooms in the hotel guest room, and then the system determines whether these control contents can be synchronized to the hotel guest room to execute corresponding steps. For example, when the system determines that the control content of the linked guest room cannot be synchronized to the hotel guest room, the system will consider that the functional content of the hotel guest room cannot be applied to the linked guest room, and the system will display the specific reason, such as "Due to device incompatibility, the content of the linked guest room cannot be synchronized and controlled to the current room.", so that the user can understand the problem. At the same time, after notifying the user, the system will try to automatically reconnect or troubleshoot to see if the problem can be solved, to avoid the problem being caused by temporary network problems or device connection problems, and recommend that the user change the cross-room linkage settings, exclude the control content that cannot be synchronized in the current guest room, and provide feasible synchronization options according to the actual situation of the current guest room, such as replacing other hotel guest rooms that can achieve cross-room linkage with the linked guest room. For example, when the system determines that the control content of the linked guest room can be synchronized to the hotel guest room, at this time the system will consider that the functional content of the hotel guest room can be applied to the linked guest room for synchronous execution, and the system will make the virtual assistant issue a corresponding task request according to the interaction content, identify the executable content pre-recorded in the linked guest room, generate the same task request from the linked guest room according to different executable contents, and coordinate the hotel guest room and the linked guest room to synchronously execute the task request through the virtual assistant. The system can identify the executable content of the linked guest room and automatically generate the same task request, enabling the virtual assistant to intelligently coordinate the devices in multiple rooms and automatically execute the user's linkage settings, which not only reduces the user's operation burden, but also demonstrates the efficiency and intelligence of the system, meeting the user's needs for scenarios such as giving a speech or having a party in multiple rooms at the same time. At the same time, in the multi-room linkage scenario, the system makes the operations of each room consistent through a unified task request, avoiding the user experience differences caused by inconsistent settings in different rooms. For example, when the user synchronously plays background music or adjusts the lighting in multiple rooms, the system can ensure that the effects of each room are the same, improving the user's overall feeling, and through synchronous execution of the task request, it can quickly respond to the user's needs in multiple rooms without adjusting the settings room by room, which is especially suitable for scenarios such as business meetings and family gatherings, and users can quickly set the environment of multiple rooms to meet different needs.
[0073] In this embodiment, before step S3 of generating the modal conflict content of the interaction content according to the collaborative association information, it further includes:
[0074] S301: Based on the modal information pre-generated by the hotel guest room for the interaction content, obtain the task request obtained by converting the modal information;
[0075] S302: Determine whether the task request conforms to the execution content preset by the hotel guest room;
[0076] S303: If not, identify the missing modality of the modality information, generate the content to be completed for the modality information according to the missing modality, obtain the response feedback of the user to the content to be completed, reconstruct the response feedback into the modality information according to the interaction type, and then convert it again to obtain the task request.
[0077] In this embodiment, the system pre-generates modality information based on the interactive content of the hotel room, obtains the task requests converted from these modality information, and then the system determines whether these task requests conform to the execution content preset by the hotel room to execute the corresponding steps; for example, when the system determines that the task request converted from the modality information can conform to the execution content preset by the hotel room, the system will consider the modality information processing and task request generation processes to be effective, and the virtual assistant can correctly understand the user's intention. The system will directly execute these qualified task requests. For example, when the user requests to adjust the room temperature through a voice command, the system will immediately adjust the air conditioner settings to reach the temperature required by the user, and at the same time provide feedback to confirm that the task request has been successfully executed, including through voice prompts, display messages or mobile application notifications, such as "The room temperature has been adjusted to 22°C.", and record the result of this task execution for subsequent analysis and optimization. These data can also be used for personalized user experience, such as automatically adjusting the room settings according to the user's preferences; for example, when the system determines that the task request converted from the modality information cannot conform to the execution content preset by the hotel room, at this time the system will consider that the virtual assistant cannot understand the user's intention. The system will identify the missing modality of the modality information, generate the content to be completed for the modality information according to the missing modality, obtain the response feedback of the user to the content to be completed, reconstruct the response feedback into the modality information according to the interaction type, and then convert it again to obtain the task request; when the system cannot directly understand the user's intention, by identifying and completing the missing modality information, the robustness of the system can be effectively improved. Even if the initially converted task request does not conform to the preset content, the system can adaptively adjust by completing the information, enhancing the processing ability for complex interaction scenarios. At the same time, by generating the content to be completed and obtaining the user's feedback, the user can more clearly understand the system's requirements and give a response. This two-way interaction helps to eliminate misunderstandings, ensure that the final task request is more in line with the user's true intention, thereby improving the user's interaction satisfaction, and automatically identifying the missing modality information and proposing completion suggestions, reducing the need for the user to actively correct or repeat the input, making the user operation more convenient, and enabling the user to interact with the virtual assistant more easily. Through multiple conversions and feedback iterations, the system can more accurately generate task requests that meet the user's needs. This repeated correction process ensures the accuracy of the final task request, improving the success rate and efficiency of task execution.
[0078] It should be noted that the missing modality of the modality information is identified, the content to be completed of the modality information is generated according to the missing modality, the response feedback of the user to the content to be completed is obtained, the response feedback is reconstructed into the modality information according to the interaction type, and the task request is obtained again through conversion. The specific example is as follows:
[0079] Suppose user Li checks into a high-end hotel. He uses a virtual assistant in the room to control various facilities. Since the hotel room is fully functional, Li hopes to adjust the room lighting, temperature and other settings through voice commands. The hotel's virtual assistant can accept multi-modal commands such as voice, touch screen, gestures, and mobile APPs to execute tasks;
[0080] Since Li has just entered the hotel room and finds that the room lighting is a bit dazzling and hopes to dim the lights, he directly says to the virtual assistant: "Dim the lights a bit." However, since there are multiple lighting areas in the room (bedroom, living room, desk, bathroom, etc.), the system cannot determine which area of the lights Li wants to dim based on the voice command alone;
[0081] First, the system receives Xiao Li's voice command: "Dim the lights a bit." Since the specific lighting area is not clearly indicated in the voice command, after the system recognizes this, it determines that the key information of "lighting area" is missing in the voice command. Although the system understands the requirement of "dimming the lights" through voice recognition technology, because there are multiple lighting areas in the room, the system realizes that it is necessary to supplement the information of "which area to dim the lights". Therefore, to ensure the accuracy of the operation, the system actively feeds back information to Xiao Li. Through the voice of the virtual assistant, it prompts him and says: "Which area of the lights do you want to dim? Is it the bedroom, the living room, or the desk?"; By analyzing the room layout and the user's previous preference records, the system lists several possibilities and asks the user through voice prompts. This feedback can also be synchronously displayed through the prompt information on the display screen or the push message of the mobile APP to facilitate the user's selection; After hearing the system's question, Xiao Li replies through voice: "Dim the lights in the bedroom."; At this time, the user supplements the missing "area" information in the original command through voice. The system receives this supplementary information and begins to integrate it into the previous command; At the same time, if Xiao Li has multiple input methods, such as clicking on the "bedroom" icon through the mobile APP, this information can also be supplemented; The system integrates the "bedroom" information supplemented by Xiao Li into the initial command to form a complete modal information: "Dim the lights in the bedroom.", then the system updates the task request according to the completed modal information and gives it priority processing in the task queue; Therefore, the system internally associates the task request of "dimming the bedroom lights" with the original voice command of "dimming the lights" and maps this association to a graph model. In this graph model, "dim" is the action node and "bedroom lights" is the object node. The system generates specific control signals based on the relationship between the two; Finally, if the system converts the completed modal information into an accurate task request: "Dim the bedroom lights to 50% brightness.", then the system issues an instruction to dim the lights in the bedroom and completes the task; And when executing the task, the system also takes into account the user's historical preferences. For example, if the user has previously often dimmed the lights in the bedroom to 30%, the system will prompt the user or automatically apply this preference to further improve the user experience;
[0082] In summary, the system successfully understands Xiao Li's true intention by asking to supplement the missing information, thus avoiding incorrect operations caused by incomplete information. Even in the case of missing information or unclear instructions in multimodal interactions, the system can still respond flexibly and finally execute the task accurately. Through this active interaction method, users do not need to repeat instructions or perform complex operations, greatly improving the user experience and satisfaction. Especially in an intelligent hotel environment, this experience is particularly important. And the system can prioritize and allocate resources for tasks according to the completed information, improving the system's operation efficiency and avoiding unnecessary resource waste.
[0083] In this embodiment, in step S5 of identifying the input data of the modal conflict content, the following steps are further included:
[0084] S51: Based on the preset acquisition features of the hotel room, obtain the types of features input by the user into the hotel room through the virtual assistant;
[0085] S52: Determine whether the type of feature exceeds the preset upper limit;
[0086] S53: If so, identify the modal conflict points of the type of feature through the virtual assistant, generate the interceptable paragraph information of the modal conflict content according to the modal conflict points, transfer the interceptable paragraph information to a preset blank paragraph, construct the interaction content to be uploaded corresponding to the interceptable paragraph information, and synchronize the interaction content to be uploaded to the virtual assistant.
[0087] In this embodiment, the system obtains the types of features input by different users in the hotel room through the virtual assistant based on the pre-set acquisition features of the hotel room. Then, the system determines whether these types of features exceed the pre-set upper limit to execute corresponding steps. For example, when the system determines that the types of features input by different users through the virtual assistant do not exceed the pre-set upper limit, the system will consider that the input behavior of the user is within the range allowed by the system. The quantity and types of feature types are at normal levels, and the system can process these inputs normally. The system classifies the user input features collected according to the pre-set classification rules and distributes them to the corresponding processing modules for processing. At the same time, through multi-modal fusion technology, different types of input features are integrated to generate a comprehensive user request description. For example, combining voice commands with operations on the touch screen to form a complete task instruction. And after completing the task, the execution result is fed back to the user to confirm whether the task is completed according to the user's expectations. The virtual assistant is used to reply to the user and provide relevant feedback options, and the user can confirm or make further adjustments. For example, when the system determines that the types of features input by different users through the virtual assistant exceed the pre-set upper limit, at this time, the system will consider that the virtual assistant cannot complete the interactions corresponding to these interaction contents. The system will identify the modal conflict points of these types of features through the virtual assistant, generate the interceptable paragraph information of the modal conflict content, transfer these interceptable paragraph information to the pre-set blank paragraph, construct the to-be-uploaded interaction content corresponding to the interceptable paragraph information, and synchronize the to-be-uploaded interaction content to the virtual assistant. By identifying the modal conflict points and generating the interceptable paragraph information of the modal conflict content, the system can clarify and decompose complex user requests, help the system clearly understand the user's interaction content, avoid the understanding difficulties caused by the over-limit of feature types, make each paragraph focus on specific interaction points or request parts, thereby reducing confusion and misunderstanding. At the same time, synchronizing the interceptable paragraph information of the modal conflict content to the virtual assistant enables the system to process complex interaction requests step by step instead of processing all information at once, thereby improving the processing efficiency and accuracy, and clearly presenting the segmented request information to the user, allowing the user to complete the interaction step by step, avoiding the chaos caused by excessive information input at one time. Even if the virtual assistant cannot complete the interaction content at one time, the to-be-uploaded interaction content can be re-uploaded to the virtual assistant through the interceptable paragraph information to achieve the secondary interaction between the user and the virtual assistant.
[0088] In this embodiment, in step S2 of determining whether the interaction content belongs to the same interaction type, it further includes:
[0089] S21: Identify the corresponding interaction device in the hotel room based on the interaction content;
[0090] S22: Determine whether the interaction device has received at least two pieces of the interaction content within a preset time period;
[0091] S23: If so, obtain the interaction type of the interaction content, and classify the interaction content into preset interaction options corresponding to the virtual assistant, where the interaction options specifically include query-type interactions, request-type interactions, and setting-type interactions.
[0092] In this embodiment, the system identifies the settable interaction devices in the hotel room based on the interaction content, and then determines whether these interaction devices have received at least two different pieces of interaction content within a preset time period to perform corresponding steps. For example, when the system determines that the interaction device has not received two or more pieces of interaction content simultaneously, the system will consider that the current interaction demand of the user is relatively low, and the device does not need to perform complex multitasking. The system will directly execute the relevant tasks of the interaction content without performing complex task allocation or conflict management. For example, when the user requests to adjust the air conditioner temperature through a voice command, the system will directly adjust the temperature setting without considering the interference of other commands. At the same time, while performing the task, the system continues to monitor whether new interaction content is received by the device within the set time period. If new interaction content arrives subsequently, the system will re-judge and perform corresponding processing. And because the interaction demand is less, the system can save resources, reduce unnecessary computing and processing power consumption, thereby improving the overall operation efficiency of the system. For example, when the system determines that the interaction device has received two or more pieces of interaction content simultaneously, the system will consider that the interaction device may need to perform multitasking. The system will obtain the interaction types of these interaction contents and classify different interaction contents into preset interaction options corresponding to the virtual assistant. The interaction options specifically include query-type interactions, request-type interactions, and setting-type interactions. The system can quickly identify and classify multiple pieces of interaction content, avoiding processing delays caused by content confusion. By classifying the interaction content into different interaction options, the system can allocate resources more effectively, process various tasks, and improve the overall processing efficiency. At the same time, through reasonable interaction type classification, the system can provide more accurate and personalized responses to users. For example, query-type interactions are processed preferentially, allowing users to obtain the required information faster, while request-type interactions can be processed in parallel in the background to improve the execution speed. Such an experience makes users feel the intelligence and responsiveness of the system. And when classifying the interaction content, it is possible to identify possible conflicts between different tasks. For example, when a temperature setting request and a lighting adjustment request are received simultaneously, the system can prioritize according to the nature of the tasks, or perform non-conflicting tasks simultaneously to ensure that each task can be completed in an orderly manner and avoid task failures caused by conflicts. Because different types of interactions usually have different requirements for system resources, and through classification processing, the system can better allocate computing resources and network bandwidth.
[0093] In this embodiment, in step S4 of determining whether the modal conflict content can be dynamically adjusted, it further includes:
[0094] S41: Obtain the conflict type of the modal conflict content, where the conflict type specifically includes spatial conflict, priority conflict, and time conflict;
[0095] S42: Determine whether the conflict types are diverted to the same interaction path;
[0096] S43: If not, generate an affected device corresponding to the modal conflict content through a virtual assistant, identify the conflict range of the modal conflict content for the affected device, and divide a dynamic adjustment interval of the affected device based on the conflict range.
[0097] In this embodiment, the system obtains the conflict types of the modal conflict content. The conflict types specifically include spatial conflict, priority conflict, and time conflict. Then, it determines whether these conflict types are diverted to the same interaction path to execute corresponding steps. For example, when the system determines that the conflict types of the modal conflict content are diverted to the same interaction path, the system will consider that the interaction content of multiple conflict types ultimately points to the same system resource or device, which may lead to resource competition or chaos in the task execution order. The system will immediately lock the resource and perform conflict detection. For tasks that need to access the same resource simultaneously, the system will determine which task should be executed first through a task scheduling mechanism, and other tasks will queue up waiting for the resource to be released. At the same time, the priorities of the tasks will be adjusted according to preset rules. For example, some tasks may be given higher priorities due to special user requirements. At this time, the system should execute this task first, postpone or suspend tasks with lower priorities. If the priorities of the tasks are the same, the system can refer to historical records or user habits to determine the execution order. And if the conflict cannot be resolved automatically, a notification can be sent to the user to explain the current conflict situation and solicit the user's opinion. For example, the system can ask the user whether they are willing to execute one task first and then process another task. This way can not only resolve the conflict but also enhance the user's sense of control. For example, when the system determines that the conflict types of the modal conflict content are not diverted to the same interaction path, the system will consider that the interaction content of the conflict types can be executed separately. The system will generate an impact device corresponding to the modal conflict content through a virtual assistant, identify the conflict range of the modal conflict content on the impact device, and divide the dynamic adjustment range of the impact device based on the conflict range. Since the interaction content of the conflict types is executed separately, the system can process multiple tasks simultaneously without task delay or interruption due to resource conflicts. This strategy enables the virtual assistant to respond more quickly to multiple user requirements, enhancing the coherence and real-time nature of the user experience. At the same time, by identifying the conflict range of the modal conflict content on the impact device, the system can accurately divide the dynamic adjustment range of the device, avoiding unnecessary device adjustments or interferences, ensuring that the device is adjusted only within the required range, thereby maintaining the stability and efficient operation of the device. And the conflict content is reasonably diverted and processed, avoiding the competition of multiple tasks for the same device or resource, reducing the interference or chaos that the user may encounter when performing multiple tasks, and enabling the user to operate more smoothly without being confused or inconvenienced due to improper system processing.
[0098] It should be noted that generating an impact device corresponding to the modal conflict content through a virtual assistant, identifying the conflict range of the modal conflict content on the impact device, and dividing the dynamic adjustment range of the impact device based on the conflict range are specifically illustrated as follows:
[0099] Suppose there is a smart air conditioner and smart lights in a hotel room. User A requests the room temperature to be adjusted to 22°C through a voice command. At the same time, User B requests the lights to be dimmed to 30% brightness through a manual operation. The affected devices are identified by the system as the air conditioner and lights. Since the air conditioner and lights are affected by different commands, but these commands do not directly conflict; therefore, the system finds that the conflict scope of these commands is not on the same device, and a dynamic adjustment range can be divided. For the air conditioning system, it will be slightly adjusted around 22°C to ensure that the temperature change is within the acceptable range of the user and takes into account the energy-saving requirements; for the lights, the system will allow a slight adjustment based on 30% brightness to ensure that the light is neither dazzling nor meets the reading requirements;
[0100] In summary, through this method, the system ensures the balanced execution of multi-user commands, avoids direct conflicts, and improves the user experience at the same time. The setting of the dynamic adjustment range enables the system to find the best balance among different requirements, responds to user needs, and maintains the efficient operation of the device. And this detailed processing method ensures that in a complex environment, various interaction requirements can be properly processed, reducing conflicts and optimizing the user experience.
[0101] In this embodiment, in step S1 of identifying the interaction content input by the user into the hotel room within the preset time period based on the pre-unlock interaction permission of the hotel room, it further includes:
[0102] S11: Identify the total number of users staying in the hotel room;
[0103] S12: Determine whether the total number of users reaches a preset threshold;
[0104] S13: If so, based on the feature information pre-entered by the user, obtain the priority configuration corresponding to other users in the hotel room for the user, collect the interaction content received in the hotel room, and divide the execution sequence of the interaction content according to the priority configuration.
[0105] In this embodiment, the system identifies the total number of users checking into a hotel room, and then determines whether the total number of users reaches a preset threshold to execute corresponding steps. For example, when the system determines that the total number of users checking into the hotel room does not reach the preset threshold, the system will consider that the number of users in the room is still within the range allowed by the system. The virtual assistant does not need to perform additional multi-user management or complex interaction feature collection. The system will continue to manage according to the normal single-user or few-user mode. The virtual assistant will provide personalized services for the currently checked-in user, such as adjusting the room temperature, playing music, or providing other personalized intelligent services. At the same time, the system will default that all interactions and requests come from the existing few users. The virtual assistant can more precisely meet the needs of each user without performing complex task allocation or conflict management, and simplifies the user interaction path, reducing complex operations that may be caused in the multi-user situation. Therefore, the virtual assistant can quickly respond to and process each user's request without having to perform complex modal conflict handling or task queue management. For example, when the system determines that the total number of users checking into the hotel room reaches the preset threshold, at this time, the system will consider that the virtual assistant in the hotel room needs to collect different interaction features of multiple users during the check-in period. The system will identify and confirm the feature information of all checked-in users through face recognition, voice recognition, or other authentication means to ensure that the system can distinguish different users, and at the same time start collecting the interaction features of each user. These features include the user's voice command preferences, temperature setting habits, and lighting brightness preferences. According to the interaction features of each user, adjust its response method and optimize its response strategy to adapt to the presence of multiple users. For example, when multiple users issue instructions simultaneously, the system will determine which user to respond to first or how to process these instructions in parallel according to the preset priority or usage frequency.
[0106] Reference appendix Figure 2 , a virtual assistant interaction system for a hotel room in an embodiment of the present invention, includes:
[0107] An identification unit 10, configured to identify the interaction content input by a user into the hotel room within a preset time period based on the interaction permission pre-unlocked for the hotel room;
[0108] A judgment unit 20, configured to judge whether the interaction content belongs to the same interaction type, where the interaction type specifically includes image interaction, sound interaction, and text interaction;
[0109] An execution unit 30, configured to, if not, add the interactive content to a task queue preset by the virtual assistant, activate a processing module preset for the hotel room, allocate the interactive content from the task queue to the processing module, respond to the interactive content in the hotel room through the processing module, identify collaborative association information pre-classified for the interactive content, and generate modal conflict content for the interactive content according to the collaborative association information, where the collaborative association information specifically includes keyword matching, spatial location matching, and task objective matching;
[0110] A second determination unit 40, configured to determine whether the modal conflict content can be dynamically adjusted;
[0111] A second execution unit 50, configured to, if so, identify input data of the modal conflict content, represent interactive information of different modalities as nodes in a preset graph model based on the input data, represent edges between the nodes as dependencies between different modalities, and generate a response strategy for the modal conflict content based on the nodes and the dependencies, where the input data specifically includes visual images, voice commands, and tactile feedback.
[0112] In this embodiment, after the system confirms that the user has checked into the hotel room, the hotel room will unlock the pre-set interaction permissions, identify the interaction content input by the user into the hotel room within the pre-set time period, and then the system determines whether these interaction contents belong to the same interaction type. The interaction types specifically include image interaction, sound interaction, and text interaction, so as to execute corresponding steps. For example, when the system determines that the interaction contents input by the user into the hotel room belong to the same interaction type, the system will consider that all inputs are concentrated in the same type, and the system does not need to perform cross-modal parsing or processing. The system will merge the contents of the same interaction type into an overall task and process them centrally. If the interaction contents have a time sequence, the system can execute these contents in the order of user input. At the same time, according to the unified interaction type, the feedback method can be optimized. For example, when the user controls the TV in different ways (such as voice and text), the system will give all feedbacks in a unified way through voice or screen display to simplify the user experience. And if the contents of the same interaction type contain repeated instructions, the system can identify and reduce redundancy to avoid performing the same operation multiple times. For example, if the user requests to play music twice continuously through voice, then the system only executes it once. For example, when the system determines that the interaction contents input by the user into the hotel room do not belong to the same interaction type, at this time the system will consider that there are multiple users in the hotel room interacting with the virtual assistant at the same time. The system will list different interaction contents in the task queue pre-set by the virtual assistant, activate the processing module pre-set in the hotel room, allocate these interaction contents from the task queue to the processing module, and respond to these interaction contents in the hotel room through the processing module. Identify the pre-classified collaborative association information of these interaction contents. The collaborative association information specifically includes keyword matching, spatial location matching, and task target matching, and generate modal conflict contents of the interaction contents according to different collaborative association information. By listing different interaction contents in the task queue and allocating them to the processing module, the system can effectively manage the interaction requests of multiple users. It allows the system to avoid instruction confusion and processing conflicts when multiple users initiate interactions at the same time, ensuring that the needs of each user in the hotel room can be independently and appropriately processed. At the same time, by identifying the collaborative association information of the interaction contents (such as keyword matching, spatial location matching, and task target matching), the intention and context of different interaction contents can be accurately distinguished. This precise recognition and classification ability helps the system better understand user needs, so as to provide more relevant and personalized services for different users. And when a modal conflict is detected, it can generate conflict contents according to the collaborative association information and dynamically adjust the processing strategy. This ability allows the system to flexibly handle conflicts when facing multiple interaction types, ensuring the interoperability and coordination between different modal inputs, thereby improving the accuracy and consistency of responses. Then the system determines whether these modal conflict contents can be dynamically adjusted to execute corresponding steps;For example, when the system determines that the modal conflict content of the interactive content cannot be dynamically adjusted, the system will consider that two or more conflicting interactive contents can only select one to be implemented. The system will evaluate the priority of the conflicting interactive contents, actively feedback the conflict situation to the user, and request the user to make a choice. The choices include the user's preference settings, the urgency of the interactive content, the time sequence of the user input, and the importance of the target task of the content. At the same time, the system will prompt the user with the selected interactive content and explain the other unexecuted contents, through voice prompts, text messages or image displays, etc., to ensure that the user understands the system's decision and processing results, and avoid confusion or dissatisfaction caused by the system not executing a certain content. And if the unselected interactive content is equally important to the user, the system will prompt the user after the main task is executed and suggest whether compensation execution is needed; For example, when the system determines that the modal conflict content of the interactive content can be dynamically adjusted, at this time the system will consider that the conflicting interactive contents can be split and carried out simultaneously. The system will identify the input data of the modal conflict content. The input data specifically includes visual images, voice commands and tactile feedback. According to different input data, different modal interactive information is represented as nodes in a pre-set graph model, and the edges between the nodes are represented as the dependency relationships between different modalities. Based on the nodes and dependency relationships, a response strategy for the modal conflict content is generated; By splitting and processing the conflicting interactive contents in parallel, the system can flexibly handle the conflict situation of multi-modal information, avoid operation interruption or non-response caused by modal conflict, dynamically adjust the processing strategy according to the real-time situation, better adapt to complex user needs. At the same time, representing different modal interactive information as nodes in a graph model and processing through the dependency relationships between the nodes enables the system to more efficiently allocate and utilize resources, reduce the waste of system resources, improve the overall operation efficiency, be able to process more tasks in the same time period, and the use of the graph model allows the system to be more meticulous when dealing with complex scenarios. When there are complex dependency relationships between visual images and voice commands, the system can accurately analyze and process these relationships, ensure that each interactive content can get appropriate response and processing, generate a more reasonable response strategy, and still make a decision that meets the user's expectations in case of conflict, reducing misunderstandings and misoperations.;
[0113] In this embodiment, the execution unit further includes:
[0114] A presentation subunit, configured to present the control content of the preset linked guest rooms in the hotel guest room based on the cross-room linkage setting preset by the user through the hotel terminal;
[0115] A judgment subunit, configured to judge whether the control content can be synchronized to the hotel guest room;
[0116] An execution subunit is used to, if possible, enable the virtual assistant to issue a corresponding task request based on the interactive content, identify the pre-recorded content of the linked guest room, generate the task request from the linked guest room based on the pre-recorded content, and coordinate the hotel room and the linked guest room to synchronously execute the task request through the virtual assistant.
[0117] In this embodiment, the system presents the pre-set control contents of the linked guest rooms in the hotel rooms based on the cross-room linkage settings pre-set by the user through the hotel terminal, and then the system determines whether these control contents can be synchronized to the hotel rooms to execute the corresponding steps; for example, when the system determines that the control contents of the linked guest rooms cannot be synchronized to the hotel rooms, the system will consider that the functional contents of the hotel rooms cannot be applied to the linked guest rooms, and the system will display the specific reasons, such as "due to incompatible devices, the contents of the linked guest rooms cannot be synchronized to the current room." , let the user know where the problem lies. After notifying the user, the system will try to automatically reconnect or troubleshoot to see if the problem can be solved to avoid problems caused by temporary network problems or device connection problems. It also recommends that the user change the cross-room linkage settings to exclude the control content that cannot be synchronized in the current guest room, and provide feasible synchronization options based on the actual situation of the current guest room, such as replacing other hotel rooms and linkage rooms that can achieve cross-room linkage; for example, when the system determines that the control content of the linkage guest room can be synchronized to the hotel room, the system will then believe that the functional content of the hotel room can be applied to the linkage guest room for synchronous execution. The system will enable the virtual assistant to issue a corresponding task request based on the interactive content, identify the executable content pre-included in the linkage guest room, generate the same task request from the linkage guest room based on different executable content, and coordinate the hotel room and the linkage guest room to synchronously execute the task request through the virtual assistant; the system Identifying the executable content of the linked guest rooms and automatically generating the same task requests enables the virtual assistant to intelligently coordinate the devices in multiple rooms and automatically execute the user's linkage settings. This not only reduces the user's operating burden, but also demonstrates the efficiency and intelligence of the system, meeting the user's need to use multiple rooms for simultaneous speeches or parties. At the same time, in the multi-room linkage scenario, the system uses a unified task request to keep the operations of each room consistent, avoiding differences in user experience caused by inconsistent settings in different rooms. For example, when users play background music or adjust lighting in multiple rooms simultaneously, the system can ensure that the effects in each room are consistent, improving the user's overall experience, and by synchronously executing task requests, it can quickly respond to user needs in multiple rooms without having to adjust settings room by room. This is especially suitable for scenarios such as business meetings and family gatherings, where users can quickly set up the environment of multiple rooms to meet different needs.
[0118] In this embodiment, it also includes:
[0119] An acquisition unit, configured to acquire a task request obtained by converting the modal information based on the modal information pre-generated for the interactive content of the hotel room;
[0120] A third judgment unit, configured to judge whether the task request conforms to the preset execution content of the hotel room;
[0121] A third execution unit, configured to, if not, identify the missing modality of the modal information, generate the content to be complemented for the modal information according to the missing modality, acquire the response feedback of the user to the content to be complemented, and reconstruct the response feedback into the modal information according to the interaction type, and then convert it again to obtain the task request.
[0122] In this embodiment, the system obtains task requests obtained by converting such modal information based on the modal information pre-generated by the hotel room for interactive content, and then the system determines whether these task requests conform to the execution content preset by the hotel room to execute corresponding steps; for example, when the system determines that the task requests obtained by converting the modal information can conform to the execution content preset by the hotel room, the system will consider the modal information processing and task request generation processes to be effective, the virtual assistant can correctly understand the user's intention, and the system will directly execute these qualified task requests. For example, when the user requests to adjust the room temperature through a voice command, the system will immediately adjust the air conditioner settings to reach the temperature required by the user, and at the same time provide feedback to confirm that the task request has been successfully executed, including through voice prompts, display messages or mobile application notifications, such as "The room temperature has been adjusted to 22°C.", and record the result of this task execution for subsequent analysis and optimization. These data can also be used for personalized user experience, such as automatically adjusting the room settings according to the user's preferences; for example, when the system determines that the task requests obtained by converting the modal information cannot conform to the execution content preset by the hotel room, at this time the system will consider that the virtual assistant cannot understand the user's intention, the system will identify the missing modality of the modal information, generate the content to be complemented of the modal information according to the missing modality, obtain the response feedback of the user to the content to be complemented, and reconstruct the response feedback into the modal information according to the interaction type, and then convert it into a task request again; when the system cannot directly understand the user's intention, by identifying and complementing the missing modal information, the robustness of the system can be effectively improved. Even if the initially converted task request does not conform to the preset content, the system can adaptively adjust by complementing the information, enhance the processing ability for complex interaction scenarios, and at the same time, by generating the content to be complemented and obtaining the user's feedback, the user can more clearly understand the system's requirements and give a response. This two-way interaction helps to eliminate misunderstandings, ensure that the final task request is more in line with the user's true intention, thereby improving the user's interaction satisfaction, and automatically identifying the missing modal information and proposing complementation suggestions, reducing the need for the user to actively correct or repeat the input, making the user operation more convenient, and enabling the user to interact with the virtual assistant more easily. Through multiple conversions and feedback iterations, the system can more accurately generate task requests that meet the user's needs. This repeated correction process ensures the accuracy of the final task request, improves the success rate and efficiency of task execution.
[0123] In this embodiment, the second execution unit further includes:
[0124] An acquisition subunit, configured to obtain the type of features input by the user to the hotel room through the virtual assistant based on the acquisition features preset by the hotel room;
[0125] A second judgment subunit, configured to judge whether the type of features exceeds a preset upper limit;
[0126] A second execution subunit, configured to, if so, identify a modal conflict point of the feature type through the virtual assistant, generate interceptable paragraph information of the modal conflict content according to the modal conflict point, transfer the interceptable paragraph information to a preset blank paragraph, construct interactive content to be uploaded corresponding to the interceptable paragraph information, and synchronize the interactive content to be uploaded to the virtual assistant.
[0127] In this embodiment, the system obtains the types of features input by different users in the hotel room through the virtual assistant based on the pre-set collection features of the hotel room. Then, the system determines whether these types of features exceed the pre-set upper limit to execute corresponding steps. For example, when the system determines that the types of features input by different users through the virtual assistant do not exceed the pre-set upper limit, the system will consider that the user's input behavior is within the range allowed by the system, the quantity and types of feature types are at normal levels, the system can process these inputs normally, the system will classify the user input features collected according to the pre-set classification rules, allocate them to the corresponding processing modules for processing, and at the same time, through multi-modal fusion technology, integrate different types of input features to generate a comprehensive user request description. For example, combine voice commands with operations on the touch screen to form a complete task instruction. And after completing the task, the system will feedback the execution result to the user to confirm whether the task is completed according to the user's expectation, reply to the user through the virtual assistant, and provide relevant feedback options for the user to confirm or make further adjustments. For example, when the system determines that the types of features input by different users through the virtual assistant exceed the pre-set upper limit, at this time, the system will consider that the virtual assistant cannot complete the interactions corresponding to these interaction contents. The system will identify the modal conflict points of these types of features through the virtual assistant, generate the interceptable paragraph information of the modal conflict content according to the modal conflict points, transfer these interceptable paragraph information to the pre-set blank paragraph, construct the interaction content to be uploaded corresponding to the interceptable paragraph information, and synchronize the interaction content to be uploaded to the virtual assistant. By identifying the modal conflict points and generating the interceptable paragraph information of the modal conflict content, the system can clarify and decompose complex user requests, help the system clearly understand the user's interaction content, avoid the understanding difficulties caused by the over-limit of feature types, make each paragraph focus on specific interaction points or request parts, thereby reducing confusion and misunderstanding. At the same time, synchronize the interceptable paragraph information of the modal conflict content to the virtual assistant, so that the system can process complex interaction requests step by step instead of processing all information at once, thereby improving the processing efficiency and accuracy, and clearly presenting the segmented request information to the user, allowing the user to complete the interaction step by step, avoiding the chaos caused by too much information input at one time. Even if the virtual assistant cannot complete the interaction content at one time, the interaction content to be uploaded can be re-uploaded to the virtual assistant through the interceptable paragraph information to achieve the secondary interaction between the user and the virtual assistant.
[0128] In this embodiment, the judgment unit further includes:
[0129] An identification subunit, configured to identify the corresponding interaction device in the hotel room based on the interaction content;
[0130] A third judgment subunit, configured to judge whether the interaction device receives at least two pieces of the interaction content within a preset time period;
[0131] A third execution subunit, configured to, if so, obtain the interaction type of the interaction content and classify the interaction content into a preset interaction option corresponding to the virtual assistant, where the interaction option specifically includes a query-type interaction, a request-type interaction, and a setting-type interaction.
[0132] In this embodiment, the system identifies the interactable devices set in the hotel room based on the interaction content, and then determines whether these interactable devices receive at least two different interaction contents within a preset time period to perform corresponding steps. For example, when the system determines that the interaction device does not receive two or more interaction contents at the same time, the system will consider that the current interaction requirement of the user is low, and the device does not need to perform complex multitasking. The system will directly execute the relevant tasks of the interaction content without complex task allocation or conflict management. For example, when the user requests to adjust the air conditioner temperature through a voice command, the system will directly adjust the temperature setting without considering the interference of other commands. At the same time, while executing the task, the system continues to monitor whether new interaction contents are received by the device within the set time period. If new interaction contents arrive later, the system will re-judge and perform corresponding processing. And because the interaction requirement is less, the system can save resources, reduce unnecessary computing and processing power consumption, thereby improving the overall operation efficiency of the system. For example, when the system determines that the interaction device receives two or more interaction contents at the same time, the system will consider that the interaction device may need to perform multitasking. The system will obtain the interaction types of these interaction contents and classify different interaction contents into preset interaction options corresponding to the virtual assistant. The interaction options specifically include query-type interactions, request-type interactions, and setting-type interactions. The system can quickly identify and classify multiple interaction contents, avoiding processing delays caused by content confusion. By classifying the interaction contents into different interaction options, the system can more effectively allocate resources and process various tasks, improving the overall processing efficiency. At the same time, through reasonable interaction type classification, the system can provide more accurate and personalized responses to users. For example, query-type interactions are processed first to allow users to obtain the required information faster, while request-type interactions can be processed in parallel in the background to improve the execution speed. Such an experience makes users feel the intelligence and sensitivity of the system. And when classifying the interaction contents, it is possible to identify possible conflicts between different tasks. For example, when receiving a temperature setting request and a lighting adjustment request at the same time, the system can prioritize according to the nature of the tasks, or execute non-conflicting tasks at the same time to ensure that each task can be completed in an orderly manner and avoid task failures caused by conflicts. Because different types of interactions usually have different requirements for system resources, and through classification processing, the system can better allocate computing resources and network bandwidth.
[0133] In this embodiment, the second determination unit further includes:
[0134] A second acquisition subunit, configured to acquire a conflict type of the modal conflict content, where the conflict type specifically includes a spatial conflict, a priority conflict, and a time conflict;
[0135] A fourth determination subunit, configured to determine whether the conflict type is diverted to the same interaction path;
[0136] A fourth execution subunit, configured to, if not, generate an affected device corresponding to the modal conflict content through a virtual assistant, identify a conflict range of the modal conflict content with respect to the affected device, and divide a dynamic adjustment range of the affected device based on the conflict range.
[0137] In this embodiment, the system obtains the conflict types of modal conflict content. The conflict types specifically include spatial conflict, priority conflict, and time conflict. Then, it determines whether these conflict types are diverted to the same interaction path to execute corresponding steps. For example, when the system determines that the conflict types of modal conflict content are diverted to the same interaction path, the system will consider that the interactive content of multiple conflict types ultimately points to the same system resource or device, which may lead to resource competition or confusion in the task execution order. The system will immediately lock the resource and perform conflict detection. For tasks that need to access the same resource simultaneously, the system will decide which task to execute first through a task scheduling mechanism, and other tasks will queue up waiting for the resource to be released. At the same time, the priorities of the tasks will be adjusted according to preset rules. For example, some tasks may be given higher priorities due to special user requirements. At this time, the system should execute this task first, postpone or suspend tasks with lower priorities. If the priorities of the tasks are the same, the system can refer to historical records or user habits to determine the execution order. And if the conflict cannot be resolved automatically, a notice can be sent to the user to explain the current conflict situation and solicit the user's opinion. For example, the system can ask the user whether they are willing to execute one task first and then process another task. This way can not only resolve the conflict but also enhance the user's sense of control. For example, when the system determines that the conflict types of modal conflict content are not diverted to the same interaction path, the system will consider that the interactive content of the conflict types can be executed separately. The system will generate an impact device corresponding to the modal conflict content through a virtual assistant, identify the conflict range of the modal conflict content on the impact device, and divide the dynamic adjustment range of the impact device based on the conflict range. Since the interactive content of the conflict types is executed separately, the system can process multiple tasks simultaneously without task delay or interruption due to resource conflicts. This strategy enables the virtual assistant to respond more quickly to multiple user requirements, enhancing the coherence and real-time nature of the user experience. At the same time, by identifying the conflict range of the modal conflict content on the impact device, the system can accurately divide the dynamic adjustment range of the device, avoiding unnecessary device adjustments or interferences, ensuring that the device is adjusted only within the required range, thereby maintaining the stability and efficient operation of the device. And the conflict content is reasonably diverted and processed, avoiding the competition of multiple tasks for the same device or resource, reducing the interference or confusion that users may encounter when performing multiple tasks, and enabling users to operate more smoothly without being confused or inconvenienced due to improper system processing.
[0138] In this embodiment, the recognition unit further includes:
[0139] A second recognition subunit, configured to recognize the total number of users checking into the hotel room;
[0140] A fifth judgment subunit, configured to judge whether the total number of users reaches a preset threshold;
[0141] The fifth execution subunit is configured to, if so, based on the feature information pre-entered by the user, obtain the priority configuration corresponding to other users in the hotel room, collect the interaction content received in the hotel room, and divide the execution sequence of the interaction content according to the priority configuration.
[0142] In this embodiment, the system identifies the total number of users staying in the hotel room, and then determines whether the total number of users reaches a preset threshold to execute corresponding steps. For example, when the system determines that the total number of users staying in the hotel room does not reach the preset threshold, the system will consider that the number of users in the room is still within the range allowed by the system. The virtual assistant does not need to perform additional multi-user management or complex interaction feature collection. The system will continue to manage according to the normal single-user or small number of users mode. The virtual assistant will provide personalized services for the currently staying users, such as adjusting the room temperature, playing music, or providing other personalized intelligent services. At the same time, the system will default that all interactions and requests come from the existing small number of users. The virtual assistant can more accurately meet the needs of each user without performing complex task allocation or conflict management, and simplify the user's interaction path, reducing the complex operations that may be caused in the case of multiple users. Therefore, the virtual assistant can quickly respond to and process each user's request without performing complex modal conflict handling or task queue management. For example, when the system determines that the total number of users staying in the hotel room reaches the preset threshold, the system will consider that the virtual assistant in the hotel room needs to collect different interaction features of multiple users during the stay. The system will identify and confirm the feature information of all staying users through face recognition, voice recognition, or other identity verification means to ensure that the system can distinguish different users. At the same time, it will start to collect the interaction features of each user, including the user's voice command preferences, temperature setting habits, and lighting brightness preferences. According to the interaction features of each user, it will adjust its response method and optimize its response strategy to adapt to the presence of multiple users. For example, when multiple users issue instructions at the same time, the system will decide which user to respond to first or how to process these instructions in parallel according to the preset priority or usage frequency.
[0143] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A virtual assistant interaction method for a hotel room, characterized in that: The following steps are involved: Based on the pre-unlocked interactive rights of the hotel room, identifying the interactive content input by the user into the hotel room within a preset time period; Determining whether the interactive contents belong to the same interaction type, wherein the interaction type specifically includes image interaction, sound interaction, and text interaction; If not, the interactive content is included in the task queue preset by the virtual assistant, the processing module preset by the hotel room is activated, the interactive content is allocated from the task queue to the processing module, the interactive content is responded to in the hotel room by the processing module, the collaborative association information pre-classified by the interactive content is identified, and the modal conflict content of the interactive content is generated according to the collaborative association information, wherein the collaborative association information specifically includes keyword matching, spatial position matching and task target matching; Determining whether the modal conflict content can be dynamically adjusted; If yes, then identifying the input data of the modal conflict content, representing the interactive information of different modalities as nodes in a preset graph model according to the input data, representing the edges between the nodes as dependency relationships between different modalities, and generating a response strategy for the modal conflict content based on the nodes and the dependency relationships, wherein the input data specifically includes visual images, voice commands, and tactile feedback; The step of activating the preset processing module of the hotel room, allocating the interactive content from the task queue to the processing module, and responding to the interactive content in the hotel room through the processing module further includes: Based on the cross-room linkage setting preset by the user through the hotel terminal, presenting the control content of the preset linkage guest room in the hotel guest room; Determining whether the control content can be synchronized to the hotel room; If possible, the virtual assistant will issue a corresponding task request based on the interactive content, identify the pre-recorded content of the linked guest room, generate the task request from the linked guest room based on the pre-recorded content, and coordinate the hotel room and the linked guest room to synchronously execute the task request through the virtual assistant.
2. The virtual assistant interaction method for hotel rooms according to claim 1, characterized in that: Before the step of generating the modal conflict content of the interactive content according to the collaborative association information, the step further includes: Based on the modality information pre-generated by the hotel room for the interactive content, obtaining a task request converted from the modality information; Determining whether the task request meets the preset execution content of the hotel room; If not, identify the missing modality of the modal information, generate the content to be completed of the modal information according to the missing modality, obtain the user's response feedback to the content to be completed, reconstruct the response feedback into the modal information according to the interaction type, and convert it again to obtain the task request.
3. The virtual assistant interaction method for hotel rooms according to claim 1, characterized in that: The step of identifying the input data of the modal conflict content further includes: Based on the preset collected features of the hotel room, obtaining the feature type input by the user into the hotel room through the virtual assistant; Determining whether the feature type exceeds a preset upper limit; If so, the virtual assistant identifies the modal conflict point of the feature type, generates the interceptable paragraph information of the modal conflict content according to the modal conflict point, transfers the interceptable paragraph information to a preset blank paragraph, constructs the interactive content to be uploaded corresponding to the interceptable paragraph information, and synchronizes the interactive content to be uploaded to the virtual assistant.
4. The virtual assistant interaction method for hotel rooms according to claim 1, characterized in that: The step of determining whether the interactive contents belong to the same interaction type further includes: Identify a corresponding interactive device in the hotel room based on the interactive content; Determining whether the interactive device receives at least two pieces of interactive content within a preset time period; If so, the interaction type of the interactive content is obtained, and the interactive content is classified into preset interaction options corresponding to the virtual assistant, wherein the interaction options specifically include query type interaction, request type interaction and setting type interaction.
5. The virtual assistant interaction method for hotel rooms according to claim 1, characterized in that: The step of determining whether the modal conflict content can be dynamically adjusted further includes: Acquire the conflict type of the modality conflict content, wherein the conflict type specifically includes space conflict, priority conflict and time conflict; Determining whether the conflict types are diverted to the same interaction path; If not, the influencing device corresponding to the modal conflict content is generated through a virtual assistant, the conflict range of the modal conflict content on the influencing device is identified, and the dynamic adjustment interval of the influencing device is divided based on the conflict range.
6. The virtual assistant interaction method for hotel rooms according to claim 1, characterized in that: The step of identifying the interactive content input by the user into the hotel room within a preset time period based on the interactive authority of the hotel room pre-unlocking further includes: Identify the total number of users occupying said hotel rooms; Determining whether the total number of users reaches a preset threshold; If so, based on the feature information pre-entered by the user, the priority configuration corresponding to the user in the hotel room is obtained, the interactive content received in the hotel room is collected, and the execution sequence of the interactive content is divided according to the priority configuration.
7. A virtual assistant interactive system for hotel rooms, characterized in that: include: An identification unit, configured to identify the interactive content input into the hotel room by the user within a preset time period based on the interactive authority of the hotel room pre-unlocking; A judging unit, used to judge whether the interactive contents belong to the same interaction type, wherein the interaction types specifically include image interaction, sound interaction and text interaction; an execution unit, for, if not, listing the interactive content into a task queue preset by the virtual assistant, activating a processing module preset by the hotel room, allocating the interactive content from the task queue to the processing module, responding to the interactive content in the hotel room through the processing module, identifying collaborative association information pre-classified by the interactive content, and generating modal conflict content of the interactive content according to the collaborative association information, wherein the collaborative association information specifically includes keyword matching, spatial position matching, and task target matching; A second judgment unit, used to judge whether the modal conflict content can be dynamically adjusted; A second execution unit is configured to, if possible, identify input data of the modal conflict content, represent the interactive information of different modalities as nodes in a preset graph model according to the input data, represent the edges between the nodes as dependency relationships between different modalities, and generate a response strategy for the modal conflict content based on the nodes and the dependency relationships, wherein the input data specifically includes visual images, voice commands, and tactile feedback; Wherein, the execution unit further includes: A presentation subunit, configured to present control contents of preset linked guest rooms in hotel guest rooms based on the cross-room linkage settings preset by the user through the hotel terminal; A judging subunit, used for judging whether the control content can be synchronized to the hotel guest room; An execution subunit is used to, if possible, enable the virtual assistant to issue a corresponding task request based on the interactive content, identify the pre-recorded content of the linked guest room, generate the task request from the linked guest room based on the pre-recorded content, and coordinate the hotel room and the linked guest room to synchronously execute the task request through the virtual assistant.
8. The virtual assistant interactive system for hotel rooms according to claim 7, characterized in that: Also includes: An acquisition unit, configured to acquire, based on the modality information pre-generated by the hotel room for the interactive content, a task request obtained by converting the modality information; A third judgment unit is used to judge whether the task request meets the preset execution content of the hotel room; The third execution unit is used to, if not, identify the missing modality of the modal information, generate the content to be completed of the modal information according to the missing modality, obtain the user's response feedback to the content to be completed, reconstruct the response feedback into the modal information according to the interaction type, and convert it again to obtain the task request.
Citation Information
Patent Citations
Solution method of controlling multiple terminals by intelligent household core server
CN103888407A
Information fusion method and related equipment
CN117951637A
Smart home equipment control method and device, computer equipment and storage medium
CN118605195A