Physical therapy robot service method and system based on multi-modal identity recognition and streaming large model dialogue, and physical therapy robot

CN122802572APending Publication Date: 2026-09-22GUANGDONG EMBOSSED STORM ROBOT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610960785.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

其中,传统单体架构将理疗设备控制、业务管理、数据处理等全部功能集成于单一程序进程中各功能模块无法独立运行与迭代维护,系统扩展性差

Benefits of technology

[0020]有益效果:本申请提供了一种基于多模态身份识别与流式大模型对话的理疗机器人服务方法、系统及理疗机器人,所述基于多模态身份识别与流式大模型对话的理疗机器人服务系统包括外部设备层、服务层、引擎层以及流程编排器,通过服务层包括若干独立的服务进程,每个服务进程对应引擎层中的一个独立服务引擎,流程编排器与所述外部设备层和所述服务层通过事件驱动通信,将外部设备层接收的数据信息分发给服务进程,并调用服务进程执行理疗服务。本申请将理疗机器人的理疗服务系统拆分为若干独立的服务进程,各服务进程通过流程编排器进行调度以协同工作,实现高度松耦合与模块化,这样可以使得各服务进程能够独立开发、迭代与维护,提升了系统的可扩展性;同时通过统一的流程编排器进行协同调度,适配了多设备、多业务场景的联动运行需求,提高了系统的整体协同性与模块化程度,能够为用户提供更加流畅稳定的智能化理疗交互服务。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802572A_ABST
    Figure CN122802572A_ABST
Patent Text Reader

Abstract

The application discloses a physiotherapy robot service method and system based on multi-modal identity recognition and streaming large model conversation and a physiotherapy robot. The physiotherapy robot service system based on multi-modal identity recognition and streaming large model conversation comprises an external device layer, a service layer, an engine layer and a process arranger. The service layer comprises a plurality of independent service processes. Each service process corresponds to an independent service engine in the engine layer. The process arranger communicates with the external device layer and the service layer through event-driven communication. The process arranger distributes data information received by the external device layer to the service processes and calls the service processes to execute physiotherapy services. The physiotherapy service system of the physiotherapy robot is split into a plurality of independent service processes. The service processes are dispatched by the process arranger to work cooperatively, realizing high loose coupling and modularization. In this way, the service processes can be independently developed, iterated and maintained, improving the scalability of the system. Meanwhile, the unified process arranger is used for cooperative scheduling, adapting to the linkage operation requirements of multiple devices and multiple business scenarios, improving the overall cooperativeness and modularization degree of the system and providing smoother and more stable intelligent physiotherapy interactive services for users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of artificial intelligence and robotic service technology, and in particular to a physiotherapy robot service method, system and physiotherapy robot based on multimodal identity recognition and streaming large model dialogue. Background Technology

[0002] Existing physiotherapy robot control technology plays an increasingly important role in the field of rehabilitation medicine, and is widely used in post-stroke motor function recovery, neuromuscular training, and joint range of motion maintenance. Currently, mainstream physiotherapy robot service systems generally adopt either a traditional monolithic architecture or a point-to-point directly connected distributed architecture. The traditional monolithic architecture integrates all functions such as physiotherapy equipment control, business management, and data processing into a single program process, making it impossible for each functional module to operate independently and be iteratively maintained, resulting in poor system scalability. The point-to-point distributed architecture relies on hard-coded methods to achieve point-to-point communication between services, lacking an efficient unified collaboration mechanism. It cannot adapt to the collaborative operation requirements of multiple physiotherapy devices and multiple business scenarios, resulting in low overall system modularity and operational coordination.

[0003] Therefore, the existing technology still needs to be improved and enhanced. Summary of the Invention

[0004] The technical problem to be solved by this application is to provide a physiotherapy robot service method, system and physiotherapy robot based on multimodal identity recognition and streaming large model dialogue, which addresses the shortcomings of the existing technology.

[0005] To address the aforementioned technical problems, the first aspect of this application provides a physiotherapy robot service system based on multimodal identity recognition and streaming large-scale model dialogue, wherein the physiotherapy robot service system based on multimodal identity recognition and streaming large-scale model dialogue specifically includes: The external device layer is used to receive data information collected by external devices. The service layer includes several service processes, which are used to implement the physiotherapy services of the physiotherapy robot. The several service processes include at least speech recognition service, large model dialogue service and speech synthesis service. The engine layer consists of several independent service engines, each corresponding to a separate service process. Each service engine establishes a signal connection with its corresponding service process through a permanent online connection strategy to provide operational support for the service process. The process orchestrator communicates with the external device layer and the service layer via event-driven communication. It is used to distribute data information received by the external device layer to the service process and call the service process to execute physiotherapy services.

[0006] The aforementioned physiotherapy robot service system based on multimodal identity recognition and streaming large model dialogue, wherein the process orchestrator is configured with a first streaming interaction scheduling stage and a second streaming interaction scheduling stage; The first streaming interaction scheduling stage is used to use the large model dialogue service to infer the user's audio to obtain text data, and to drive the speech synthesis service to infer the text data for real-time broadcasting. The second streaming interactive scheduling stage is used to call the physiotherapy robot to perform the execution operation corresponding to the text data to obtain execution feedback data, and then use the execution feedback data to drive the speech synthesis service to perform inference for feedback broadcasting.

[0007] The aforementioned physiotherapy robot service system based on multimodal identity recognition and streaming large-scale model dialogue includes a voiceprint gating system configured in the voice recognition service. The voiceprint gating system is used to dynamically select the audio data to be released to the service engine corresponding to the voice recognition service according to the business scenario throughout the entire process of the physiotherapy service. The audio data is either user audio or silent audio.

[0008] The aforementioned physiotherapy robot service system based on multimodal identity recognition and streaming large model dialogue, wherein the voiceprint gating system is configured with multiple dynamically switchable working modes, including at least a silent mode, a direct access mode, a data acquisition mode, and a voiceprint verification mode. The mute mode is used to allow mute audio to be transmitted to the service engine corresponding to the speech recognition service. The direct access mode is used to allow user audio to be transmitted to the service engine corresponding to the speech recognition service; The acquisition mode is used to capture user audio segments for writing into the voiceprint template bank; The voiceprint verification mode is used to allow the user's audio to be sent to the service engine corresponding to the speech recognition service when the user's audio passes the voiceprint verification, and to allow silent audio to be sent to the service engine corresponding to the speech recognition service when the user's audio fails the voiceprint verification. The voiceprint verification mode includes a single-person voiceprint verification mode and a multi-person voiceprint verification mode.

[0009] The aforementioned physiotherapy robot service system based on multimodal identity recognition and streaming large model dialogue includes a voiceprint gating system equipped with a voiceprint scoring mechanism. This voiceprint scoring mechanism is used to extract voiceprint embedding features from user audio received by the external device layer distributed by the process orchestrator, and to match the voiceprint embedding features with the corresponding voiceprint template in the voiceprint template bank to verify the user audio.

[0010] The aforementioned physiotherapy robot service system based on multimodal identity recognition and streaming large model dialogue, wherein the voiceprint gating system is configured with a pre-buffering retransmission mechanism and / or an audio final review mechanism; The pre-buffering and resending mechanism is used to cache user audio when the voiceprint gating system is not enabled, and resend the cached user audio to the service engine corresponding to the speech recognition service when the voiceprint gating system is enabled. The voiceprint gating system is equipped with an audio final review mechanism. After a single round of voice interaction, the audio final review mechanism is used to extract the voiceprint embedding features of the cached user audio and match them with the corresponding voiceprint template in the voiceprint template bank to verify the user audio.

[0011] The aforementioned physiotherapy robot service system based on multimodal identity recognition and streaming large model dialogue, wherein the physiotherapy service is divided into multiple business stages according to the execution process, and the multiple business stages include at least the identity verification stage, the needs confirmation stage, the bed guidance stage, the physiotherapy service stage, the experience summary stage, and the recommendation stage.

[0012] The aforementioned physiotherapy robot service system based on multimodal identity recognition and streaming large-scale model dialogue includes a watershed mechanism between the demand confirmation stage and the bed-on guidance stage. The watershed mechanism uses the user's confirmation of starting the physiotherapy service as the watershed marker. Before the watershed marker, the absence of a detected user indicating that the physiotherapy service has been discontinued is used as a constraint condition. After the watershed marker, the absence of a detected user indicating that the physiotherapy service has been discontinued is not used as a constraint condition.

[0013] The aforementioned physiotherapy robot service system based on multimodal identity recognition and streaming large-scale dialogue, wherein the process orchestrator is further used for: When entering the next business phase, obtain the phase identifier and structured information of the next business phase; The structured information is updated into the user profile, and the interactive audio of the current business stage is cleared, so that the service engine corresponding to the large model dialogue service and / or speech synthesis service can form context information based on the interactive audio of the next business stage, the stage identifier of the next business stage, and the updated user profile.

[0014] The physiotherapy robot service system based on multimodal identity recognition and streaming large model dialogue, wherein, in the physiotherapy service stage, the process orchestrator is also used to call the service layer for active interaction according to the service progress of the physiotherapy service. The active interaction includes multiple active interaction stages, which include at least a sensory landing interaction stage, a deep companionship interaction stage, and a warm closing interaction stage. The sensory landing interaction stage is used to guide the interaction before the progress of the physiotherapy service reaches the first progress threshold in order to help the user experience the physiotherapy service. The deep companionship and interaction phase is used to conduct emotional interaction with the user when the progress of the physical therapy service is between the first progress threshold and the second progress threshold, so as to establish an emotional connection with the user. The warm closing interaction phase is used to conduct experiential interaction after the physiotherapy service progress reaches the second progress threshold in order to understand the user's physiotherapy service experience.

[0015] The aforementioned physiotherapy robot service system based on multimodal identity recognition and streaming large-scale model dialogue, wherein the proactive interaction by calling the service layer according to the service progress of the physiotherapy service is specifically as follows: Based on the monitoring phase trigger points of the physiotherapy service progress; When a stage trigger point is detected, obtain the large model instruction template corresponding to the detected stage trigger point; Based on the large model instruction template, obtain the information required for active interaction, and call the speech synthesis service to form active interaction speech based on the information required for active interaction, so as to trigger the active interaction stage corresponding to the trigger point based on the active interaction speech.

[0016] The aforementioned physiotherapy robot service system based on multimodal identity recognition and streaming large-scale dialogue, wherein the process orchestrator is further used for: Acquire facial images captured by an external device layer, and perform identity recognition based on the facial images; When the user's identity is recognized, the voice recognition service is invoked to verify the voiceprint. When the user's identity is not recognized, the speech recognition service is invoked to guide the user to register and build a voiceprint template. During the process of guiding the user to register, the user's name is obtained through interaction with the user.

[0017] The aforementioned physiotherapy robot service system based on multimodal identity recognition and streaming large-scale model dialogue includes a user data management service in the aforementioned service processes, and the engine layer includes a user data management engine; the user data management service is used to manage user profiles, voiceprint template bank, and facial images.

[0018] The second aspect of this application provides a physiotherapy robot service method based on multimodal identity recognition and streaming large-scale dialogue, applying the physiotherapy robot service system based on multimodal identity recognition and streaming large-scale dialogue as described above. The physiotherapy robot service method based on multimodal identity recognition and streaming large-scale dialogue specifically includes: Receive data information collected by external devices through the external device layer; The process orchestrator invokes service processes in the service layer based on the data information, and the service processes invoke their corresponding service engines to perform interactive services.

[0019] A third aspect of this application provides a physiotherapy robot equipped with the physiotherapy robot service system based on multimodal identity recognition and streaming large model dialogue as described above.

[0020] Beneficial Effects: This application provides a method, system, and physiotherapy robot service based on multimodal identity recognition and streaming large-scale model dialogue. The physiotherapy robot service system based on multimodal identity recognition and streaming large-scale model dialogue includes an external device layer, a service layer, an engine layer, and a process orchestrator. The service layer includes several independent service processes, each corresponding to an independent service engine in the engine layer. The process orchestrator communicates with the external device layer and the service layer via event-driven communication, distributing data information received by the external device layer to the service processes and invoking the service processes to execute physiotherapy services. This application decomposes the physiotherapy robot service system into several independent service processes. Each service process is scheduled to work collaboratively through the process orchestrator, achieving high loose coupling and modularity. This allows each service process to be independently developed, iterated, and maintained, improving the system's scalability. Simultaneously, collaborative scheduling through a unified process orchestrator adapts to the collaborative operation requirements of multiple devices and multiple business scenarios, improving the overall system's synergy and modularity, and providing users with a smoother and more stable intelligent physiotherapy interaction service. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 A schematic diagram of the principle of a physiotherapy robot service system based on multimodal identity recognition and streaming large model dialogue provided in the embodiments of this application.

[0023] Figure 2 This is a specific example diagram of a physiotherapy robot service system based on multimodal identity recognition and streaming large model dialogue provided in the embodiments of this application.

[0024] Figure 3 This is a schematic diagram of the workflow of a voiceprint gate control system.

[0025] Figure 4 A diagram illustrating the six business stages of interactive services.

[0026] Figure 5 This is a flowchart illustrating the collaborative process between the first and second streaming interactive scheduling phases.

[0027] Figure 6 A schematic block diagram of the terminal device provided in the embodiments of this application. Detailed Implementation

[0028] This application provides a method, system, and physiotherapy robot service based on multimodal identity recognition and streaming large-scale model dialogue. To make the objectives, technical solutions, and effects of this application clearer and more explicit, the following detailed description, with reference to the accompanying drawings and embodiments, further illustrates this application. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0029] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this application means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.

[0030] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0031] It should be understood that the sequence number and size of each step in this embodiment do not imply the order of execution. The execution order of each process is determined by its function and internal logic, and should not constitute any limitation on the implementation process of this application embodiment.

[0032] The application content will be further explained below with reference to the accompanying drawings and the description of the embodiments.

[0033] This embodiment provides a physiotherapy robot service system based on multimodal identity recognition and streaming large-scale model dialogue, such as... Figure 1 and Figure 2 As shown, the physiotherapy robot service system based on multimodal identity recognition and streaming large-scale dialogue specifically includes an external device layer, a service layer, an engine layer, and a process orchestrator. Both the external device layer and the service layer are connected to the process orchestrator and communicate via event-driven mechanisms. The service layer includes several independent service processes, and the engine layer includes several independent service engines. Each service engine corresponds one-to-one with a specific service process, and each service engine establishes a signal connection with its corresponding service process through a permanent online connection strategy to provide operational support for the service process. The process orchestrator acquires data information received from the external device layer and distributes it to the corresponding service processes in the service layer. Each service process then calls its corresponding service engine, which provides operational support to execute the physiotherapy service.

[0034] This application embodiment uses a process orchestrator as a global business scheduling bus (i.e., the system's core scheduling hub), connecting the external device layer, service layer, and engine layer. The process orchestrator unifies the collaborative scheduling of various service processes, changing the traditional point-to-point hard-coded communication scheduling method. It eliminates the need for complex cross-connections between functional modules; each service process only needs to communicate with the process orchestrator based on an agreed-upon event interface. This simplifies the implementation logic of multi-service collaboration and reduces the coupling between modules. Furthermore, in this embodiment, each service process and its corresponding service engine run independently without interference. Updates and iterations of a single service process do not affect the normal operation of other service processes, effectively reducing the difficulty of system development and maintenance, improving system development and iteration efficiency, and facilitating the flexible addition or removal of service processes according to the needs of different physiotherapy scenarios, resulting in stronger adaptability.

[0035] like Figure 2As shown, the external device layer can adapt to various external hardware and aggregate the signals from these devices into the process orchestrator. These compatible external hardware devices can include front-end cameras, bed-side cameras, massage bed devices, tablets / buttons, signal simulators, and microphones. The front-end camera is used to capture the user's facial image to trigger the identity verification stage in the physiotherapy service. The microphone is used to capture the user's audio, for example, combining it with facial images for identity verification during the physiotherapy service's identity verification stage, determining user needs during the needs confirmation stage, and obtaining user interaction information during the physiotherapy service stage, experience summary stage, and recommendation stage. The bed-side camera is used to capture the user's image for bed positioning recognition, for example, verifying bed placement and / or massage position during the physiotherapy service stage. The massage bed is used to perform therapeutic operations on users. For example, it adjusts the massage intensity and mode for corresponding areas according to the user's selected therapy plan, completing a preset therapy process. It also feeds back the real-time operating status information of the massage bed to the process orchestrator, facilitating timely responses to status changes during interactive services. The tablet / button is used for auxiliary interaction with the user, such as providing options for selecting therapy programs, adjusting therapy parameters, and confirming services, as well as receiving user interaction commands. The signal simulator is used during the system development and debugging phase to simulate the data acquisition and control signals output by external devices, assisting in the completion of system function testing.

[0036] like Figure 2 As shown, the service layer consists of several independent service processes. Each service process only calls its corresponding service engine, and there is no direct coupling or dependency between them. Interaction and collaboration are only achieved through event messages forwarded by the process orchestrator. When adding or adjusting service functions, only the corresponding service process needs to be developed or modified, without affecting the normal operation of other service processes. This reduces the development cost of function iteration and improves the flexibility of system function expansion. The service processes are only responsible for controlling the flow of business logic; the specific function execution is completed by the corresponding service engine. Each service process only maintains a stable connection with its corresponding service engine, obtains the execution result by sending service requests to the service engine, and then sends the processed result to the process orchestrator via event messages. The process orchestrator then uniformly schedules the service processes for the next stage.

[0037] The service process may include speech recognition service, large model dialogue service, and speech synthesis service. Correspondingly, the service engine includes speech recognition engine, large model inference engine, and speech synthesis engine. The speech recognition service and the speech recognition engine establish a communication connection through a permanent online connection strategy. The large model dialogue service and the large model inference engine establish a communication connection through a permanent online connection strategy. The speech synthesis service and the speech synthesis engine establish a communication connection through a permanent online connection strategy.

[0038] Specifically, the speech recognition service is used to invoke the speech recognition engine to recognize user audio sent by the process orchestrator. In other words, the speech recognition service receives user audio from the process orchestrator and uses it as input data for the speech recognition engine. The engine then recognizes the user audio and returns the generated text recognition result to the process orchestrator, which forwards the result to the large-scale dialogue service for further processing. The speech recognition engine supports segmented output of recognition results, which reduces the waiting time after the user finishes speaking, improves the smoothness of the interaction response, and optimizes the user experience.

[0039] The large-scale model dialogue service is used to invoke the large-scale model inference engine, which infers response data based on the text data recognized by the speech recognition service. Specifically, after the speech recognition service sends the recognized text results to the process orchestrator, the process orchestrator organizes the contextual dialogue information and forwards the integrated complete request to the large-scale model dialogue service. The large-scale model dialogue service then invokes the large-scale model inference engine to generate response text that conforms to the current physiotherapy scenario. The generated response text is then returned to the process orchestrator, which forwards it to the subsequent speech synthesis service for processing, ensuring the logical coherence of the entire voice interaction and that the response content is adapted to the service requirements of the physiotherapy robot.

[0040] The speech synthesis service invokes the speech synthesis engine to convert the response text generated by the large model dialogue service into speech output, thereby enabling natural speech interaction with the user. The speech synthesis engine supports streaming text input and streaming audio output. As the large model inference engine generates the response text segment by segment, the speech synthesis service can send these segments into the engine, starting speech synthesis processing in advance. This eliminates the need to wait for the entire response text to be generated, significantly reducing the waiting time for speech output and further improving the smoothness of the interaction.

[0041] This application embodiment achieves natural and fluent voice interaction through the collaborative work of speech recognition service, large-model dialogue service, and speech synthesis service. It provides users with a more user-friendly and intuitive physiotherapy interaction interface, lowering the barrier to entry. Simultaneously, the streaming processing mechanism effectively shortens the response time at each stage, improving the real-time performance and user experience. Furthermore, to manage user profiles and voiceprint templates, the service layer can include a user data management service. Correspondingly, the engine layer is configured with a user data management engine. The user data management service manages user profiles, voiceprint template banks, and facial images, while the user data management engine stores user profiles, facial images, and voiceprint template banks.

[0042] In one embodiment, the speech recognition service is configured with a voiceprint gating system. This system dynamically selects which audio data to allow to the speech recognition service engine based on the business scenario throughout the entire physiotherapy service process. This audio data can be user audio or silent audio. In other words, when the speech recognition service allows audio data to the speech recognition engine, the voiceprint gating system performs voiceprint verification and business scenario recognition on the user audio. Based on the voiceprint verification results and the business scenario, it determines whether to allow the user audio or silent audio. This filters out erroneously recorded environmental noise or audio data from other irrelevant personnel, preventing other users from interfering with the current interaction process and improving the accuracy of voice interaction. Furthermore, it prevents user audio from being allowed to the speech recognition engine in business scenarios where voice input is not required, reducing unnecessary computing power consumption and lowering system power consumption. The voiceprint gating system is only used to determine whether the audio data allowed to the service engine corresponding to the speech recognition service is user audio or silent audio. It does not control the disconnection between the speech recognition service and its corresponding service engine. Furthermore, the WebSocket connection between the speech recognition service and its corresponding service engine is always maintained. When the speech recognition service does not allow audio data to its corresponding service engine, the voiceprint gating system will send a frame of silent audio at preset intervals (such as the duration of one frame of audio data) to maintain the connection activity between the speech recognition service and its corresponding service engine. Simultaneously, if there is no real user voice for a certain period (such as 5 minutes), it will proactively disconnect and reconnect to release server resources. After disconnection, an exponential backoff strategy is used for automatic retry.

[0043] In one embodiment, the voiceprint gating system is configured with multiple dynamically switchable operating modes to control whether user audio is allowed to pass through or silenced audio to the service engine corresponding to the speech recognition service. These multiple dynamically switchable operating modes include at least a silence mode, a pass-through mode, a capture mode, and a voiceprint verification mode. Specifically, the silence mode allows silenced audio to be passed through the service engine corresponding to the speech recognition service; the pass-through mode allows user audio to be passed through the service engine corresponding to the speech recognition service; the capture mode captures user audio segments for writing to the voiceprint template bank; and the voiceprint verification mode allows user audio to be passed through the service engine corresponding to the speech recognition service when the user audio passes voiceprint verification, and allows silenced audio to be passed through the service engine corresponding to the speech recognition service when the user audio fails voiceprint verification.

[0044] It is understandable that physiotherapy robots operate across multiple business scenarios throughout the entire interactive physiotherapy service process (i.e., including all business scenarios). The process orchestrator sends mode switching commands to the voiceprint gating system based on the current business scenario and the corresponding voiceprint verification result. The voiceprint gating system then switches to the corresponding working mode based on the switching command to adapt to the differentiated needs of voice acquisition and recognition in different business scenarios. The correspondence between business scenarios, voiceprint verification results, and working modes can be summarized as follows: 1. Silent Mode The business scenarios corresponding to the silent mode are voice broadcast scenarios and spatial standby scenarios. The voiceprint verification result corresponding to the silent mode is no voiceprint verification, and the audio data released in the silent mode is silent mode; that is, when the physiotherapy robot is in the voice broadcast scenario or spatial standby scenario, the audio data released by the voice recognition service to its corresponding service engine (i.e., the voice recognition engine) is silent audio.

[0045] 2. Straight-through mode The business scenario corresponding to the direct access mode is the new user registration stage (i.e., there is no user voiceprint template). The voiceprint verification result corresponding to the direct access mode is no voiceprint verification. The audio data released by the direct access mode is the user's audio. That is, when the physiotherapy robot is in the new user registration stage, the audio data released by the speech recognition service to its corresponding service engine (i.e., speech recognition engine) is the user's audio.

[0046] 3. Acquisition Mode The business scenario corresponding to the collection mode is the voiceprint registration and recording scenario. The voiceprint verification result corresponding to the collection mode is no voiceprint verification. The audio data released by the collection mode is silent audio, and the user audio received in the voiceprint registration and recording scenario is cached so that voiceprint features can be extracted based on the cached user audio to form a voiceprint template. That is, when the physiotherapy robot is in the voiceprint registration and recording scenario, the audio data released by the speech recognition service to its corresponding service engine (i.e., speech recognition engine) is silent audio, and the user audio received in the voiceprint registration and recording scenario is cached.

[0047] 4. Voiceprint Verification Mode The business scenarios corresponding to the voiceprint verification mode are other business scenarios besides those corresponding to the silent mode, direct access mode, and collection mode. The voiceprint verification result corresponding to the voiceprint verification mode is verification passed, and the audio data released by the voiceprint verification mode is the user's audio. That is, when the physiotherapy robot is in a business scenario other than those corresponding to the silent mode, direct access mode, and collection mode, it will perform voiceprint verification on the user's audio, and release the user's audio to the service engine corresponding to the speech recognition service when the voiceprint verification is passed.

[0048] Furthermore, it should be noted that in practical applications, users can receive physiotherapy services alone or accompanied by others. Therefore, the voiceprint verification mode can be configured with single-user and multi-user voiceprint verification modes. Single-user voiceprint verification verifies the voiceprint information of a single user. Verification is only successful if the voiceprint features match a voiceprint template stored in the user's voiceprint template bank, allowing the user's audio to be released to the corresponding speech recognition service engine. Otherwise, silent audio is released to prevent accompanying persons from interfering with the current user's interaction. Multi-user voiceprint verification supports scenarios where multiple associated users interact with the user during physiotherapy. During verification, the collected user audio is compared sequentially with the voiceprint templates stored in the voiceprint template bank of all pre-set associated users participating in the current physiotherapy session. Verification is successful as long as the voiceprint matching degree of any associated user reaches a preset threshold, allowing the user's audio to be released. This meets the needs of multiple users jointly adjusting physiotherapy plans and engaging in interactive consultations, ensuring that voice interaction is not interfered with by irrelevant personnel while adapting to the actual needs of multi-user physiotherapy scenarios.

[0049] Furthermore, voiceprint verification can be achieved by scoring the user's audio. That is, the voiceprint gating system is configured with a voiceprint scoring mechanism. This mechanism extracts voiceprint embedding features from frames of user audio received by the external device layer, and matches these features with corresponding voiceprint templates in a voiceprint template bank to perform voiceprint verification on the user's audio. Specifically, as... Figure 3 As shown, the voiceprint verification process can be as follows: Each time the speech recognition service receives a frame of audio data (e.g., a 16kHz, 16-bit mono audio data frame output by the microphone every 200ms), it performs a sound detection (e.g., based on the root mean square energy threshold, peak amplitude threshold, and upper limit of silence ratio). Based on the detection result, it configures a sound identifier for each frame of audio data (this sound identifier can be either a sound identifier or a silent identifier) ​​and caches the audio data. Then, whenever the cached audio data meets preset conditions (e.g., the duration of audio data carrying a sound identifier reaches a first preset duration (e.g., 600ms), and the interval between the first and second preset durations of voiceprint verification reaches a second preset duration (e.g., 400ms), the first preset duration... If the length of the audio data is greater than the second preset duration, the system selects the window audio data with the third preset duration (e.g., 1600ms) closest to the current time from the cache, extracts the user's voiceprint embedding vector from the window audio data, and finally determines the similarity (e.g., cosine similarity) between the user's voiceprint embedding vector and the voiceprint template to determine the voiceprint verification score (e.g., weighting the similarity centroid with the mean of the largest number of similarities). The voiceprint verification score is then compared with a preset score threshold (e.g., 0.4). If the voiceprint verification score reaches the preset score threshold, the voiceprint verification is considered successful; otherwise, if the voiceprint verification score does not reach the preset score threshold, the voiceprint verification is considered unsuccessful.

[0050] In this embodiment, audio data is cached in frames and voiceprint verification is triggered periodically according to preset conditions. This not only ensures the timeliness of voiceprint verification, but also improves the verification accuracy by extracting effective voiceprint features through window segment selection. At the same time, it avoids performing feature extraction and comparison on every frame of audio, effectively reducing the computational power consumption of voiceprint verification and balancing verification efficiency and verification accuracy.

[0051] Furthermore, in practical applications, the initial voiceprint verification requires a pre-set audio duration, during which the user's audio is already waiting in the buffer. Therefore, while the voiceprint gating system is in voiceprint verification mode, upon successful initial verification, all cached user audio is sent to the speech recognition engine at once, ensuring that the beginning of the user's speech is not lost due to verification delay, and the user does not need to deliberately pause and wait. Based on this, the voiceprint gating system is configured with a pre-buffering retransmission mechanism. This mechanism caches user audio when the voiceprint gating system is not enabled and retransmits the cached user audio to the corresponding service engine for the speech recognition service when the voiceprint gating system is enabled.

[0052] Furthermore, to prevent user audio loss due to the voiceprint gating system not being activated before the user finishes speaking (when silence exceeds 500ms and a shutdown judgment is triggered), the voiceprint gating system is configured with an audio final review mechanism. This mechanism extracts the voiceprint embedding features of the cached user audio after a single round of voice interaction and matches them with the corresponding voiceprint template in the voiceprint template bank to verify the user audio. In other words, if the audio generated in a single round of interaction does not reach the first preset duration, voiceprint verification is performed on all audio generated in that round. If the verification result is successful, all unreleased user audio in the cache of that round of interaction is sent to the speech recognition engine all at once. This prevents the loss of valid voice from completed interactions due to verification delays, further ensuring the integrity of speech recognition and improving the accuracy of interaction in voiceprint verification mode.

[0053] like Figure 2 As shown, the workflow orchestrator, as the core scheduling hub of the system, is used to manage the entire process of physiotherapy services. Among them, such as... Figure 4 As shown, the entire process of physiotherapy services can be divided into the identity verification stage, the needs confirmation stage, the bed-sitting guidance stage, the physiotherapy service stage, the experience summary stage, and the recommendation stage. Among them, the identity verification stage is used to verify identity based on facial images and voiceprint features and create user profiles for new customers; the needs confirmation stage is used to determine the physiotherapy package; the bed-sitting guidance stage is used to guide users to their beds; the physiotherapy service stage is used to provide massage services; the experience summary stage is used to summarize the experience; and the recommendation stage is used to recommend products to users.

[0054] To adapt to different business stages, the workflow orchestrator has six business state machines, each corresponding one-to-one with one of the six business stages in the entire physiotherapy service process. The orchestrator can switch business state machines based on hardware triggers / voice interaction results to automatically navigate between business stages in the entire physiotherapy service process. Specifically, the business state machines for the identity verification stage, needs confirmation stage, bed-on guidance stage, physiotherapy service stage, experience summary stage, and recommendation stage are each configured with trigger conditions. When a business state machine's configured trigger condition is met, the workflow orchestrator triggers that business state machine to enter its corresponding business stage. The correspondence between business state machines and trigger conditions can be as follows: The trigger condition for the business state machine in the authentication phase is: the hardware device layer receives the user's facial image; The trigger condition for the business state machine in the requirement confirmation phase is: authentication completed; The trigger condition for the business state machine in the bed-on guidance phase is: receiving a confirmation instruction for the physiotherapy package; The trigger condition for the business state machine in the physiotherapy service phase is: completion of human acupoint recognition; The trigger condition for the business state machine in the experience summary phase is: the massage session is completed. The trigger condition for the business state machine in the recommendation phase is: completion of the experience summary.

[0055] Therefore, such as Figure 4 As shown, the entire process of interactive physiotherapy services can be summarized as follows: When the hardware device layer receives the user's face image, the process orchestrator triggers the business state machine of the authentication phase to enter the authentication phase and listens for the authentication completion signal. When the authentication completion signal is detected, the process orchestrator triggers the business state machine of the requirement confirmation phase to enter the requirement confirmation phase and listens for confirmation instructions for the physiotherapy package. When a confirmation instruction for a physiotherapy package is received, the process orchestrator triggers the business state machine of the bed-on guidance phase to enter the bed-on guidance phase and listens for the completion instruction of human acupoint recognition. When the instruction to complete the human acupoint recognition is received, the process orchestrator triggers the business state machine of the physiotherapy service stage to enter the physiotherapy service stage, and listens for the signal that the entire massage is completed. When the signal indicating that the entire massage session is complete is detected, the process orchestrator triggers the business state machine for the experience summary phase to enter the experience summary phase and listens for the experience summary completion signal. When the experience summary completion signal is detected, the process orchestrator triggers the business state machine of the recommendation phase to enter the recommendation phase, so as to realize the automated scheduling of the entire process of interactive physiotherapy service, and complete the entire process of physiotherapy service interaction without manual intervention.

[0056] Specifically, such as Figure 4 As shown, during the identity verification phase, the process orchestrator is also used to acquire facial images collected by the external device layer and perform identity recognition based on the facial images; when the user's identity is recognized, the voice recognition service is invoked to perform voiceprint verification; when the user's identity is not recognized, the voice recognition service is invoked to guide the user to register and construct a voiceprint template. In the process of guiding the user to register, the user's name is obtained through interaction with the user.

[0057] It is understandable that when a user's gaze is detected as directed towards the camera, a welcoming message is emitted and the user is guided to approach. When the distance between the user's face and the camera meets the preset requirements, facial image acquisition begins for facial comparison. If a facial information is matched in the preset facial database, the user's audio is acquired for voiceprint verification. Upon successful voiceprint verification, an identity verification completion signal is triggered, and the voiceprint gating system is activated. If no facial information is matched in the preset facial database, the user is guided to complete new user registration and enter facial and voiceprint information. Upon completion, an identity verification completion signal is triggered, and the voiceprint gating system is activated. Specifically, when guiding the user to complete new user registration and enter facial and voiceprint information, the user can be guided to control segmented audio recording via physical buttons. Each audio segment has its own voiceprint embedding vector extracted by the voiceprint acquisition module. The user's voiceprint embedding vectors from multiple audio segments are aggregated to construct a voiceprint template, which is then persistently stored in the user data management engine as a vector file.

[0058] Furthermore, during the process of guiding users to complete new user registration and enter facial and voiceprint information, the user's name can be obtained through natural language dialogue. Specifically, a prompt can be played to guide the user's response, activating the voiceprint gating system to receive only the user's audio. The voice recognition engine then identifies the user's audio to determine their name, and the name is written into the user's profile through the user data management service. This completes both voiceprint collection and simultaneous acquisition of the user's name, eliminating the need for additional information input through physical input devices, simplifying the new user registration process, and improving the smoothness of the registration interaction. Moreover, when a user changes their name, the voiceprint template and name information in the user profile can be automatically updated to adapt to the needs of user information changes, eliminating the need for administrators to manually adjust user profile information, further improving information management efficiency.

[0059] like Figure 4As shown, during the demand confirmation phase, the process orchestrator coordinates the work by calling the service processes and their corresponding service engines in the service layer to conduct multiple rounds of interaction with the user. During these interactions, the user is introduced to the physiotherapy package contents and price, asked whether they wish to begin the corresponding physiotherapy service, and a confirmation instruction is generated when the physiotherapy service begins. Furthermore, since the user is guided to their bed after the demand confirmation phase, the front-end camera used for authentication may fail to capture the user's image during this process, potentially misinterpreting the user's arrival at the bed as leaving. Therefore, a watershed mechanism is established between the demand confirmation phase and the bed-sitting guidance phase. This watershed mechanism uses the user's confirmation of starting the physiotherapy service as the dividing point. Before the watershed point, the absence of a detected user not discontinuing the physiotherapy service is a constraint; after the watershed point, this constraint is not applied. In other words, after a confirmation instruction for a physiotherapy package is generated, a watershed mechanism is activated to disable the "user departure detection" constraint logic. This prevents the physiotherapy service from being mistakenly terminated due to the camera failing to capture the user's face while the user is moving to the bed, ensuring the entire physiotherapy process can proceed continuously without the user needing to re-trigger verification and process restart. Simultaneously, the watershed mechanism automatically resets the marker at the end of the recommendation phase and after the entire physiotherapy process is completed, preparing the user for the next user's physiotherapy service without requiring manual reset of process nodes.

[0060] like Figure 4 As shown, during the bed-boarding guidance phase, the process orchestrator calls the service processes and their corresponding service engines in the service layer to coordinate work for multiple rounds of interaction with the user. During these interactions, the user is guided to the bed and guided to adjust their body posture so that the physiotherapy robot can recognize the acupoints corresponding to the physiotherapy package; and after recognizing the acupoints, a completion command for acupoint recognition is generated.

[0061] like Figure 4As shown, during the physiotherapy service phase, the process orchestrator coordinates the work by calling service processes and their corresponding service engines in the service layer to proactively interact with the user. This proactive interaction during the physiotherapy service phase includes multiple stages, at least a sensory engagement stage, a deep companionship stage, and a warm closing interaction stage. The sensory engagement stage guides the user's experience before the physiotherapy service progress reaches a first progress threshold, assisting them in feeling the service. The deep companionship stage facilitates emotional interaction between the first and second progress thresholds, establishing an emotional connection with the user. The warm closing interaction stage provides experiential interaction after the physiotherapy service progress reaches the second progress threshold, allowing the system to gauge the user's experience with the service. The first progress threshold can be 30%, the second progress threshold can be 70%, etc.

[0062] Furthermore, trigger points can be set within each active interaction phase. These trigger points allow for the creation of multiple refined interaction phases within each phase to accommodate different treatment durations, ensuring a moderate frequency of active interaction. This avoids excessive interaction that disrupts the user's treatment experience, while also preventing prolonged periods of inactivity that might make the user feel neglected. Simultaneously, the interactive content at each phase can be dynamically adjusted based on the type of treatment package and the user's historical preferences, further adapting to the needs of different users and enhancing the interactive experience during treatment.

[0063] Each stage's trigger point corresponds to a dedicated large-scale model instruction template, which is dynamically populated at runtime by extracting information from real-time status and customer profiles. Specifically, the service layer's proactive interaction based on the physiotherapy service progress involves listening for stage trigger points based on the physiotherapy service progress; when a stage trigger point is detected, the large-scale model instruction template corresponding to that trigger point is obtained; based on the large-scale model instruction template, the information required for proactive interaction is obtained, and a speech synthesis service is invoked to generate proactive interactive speech based on this information, so as to trigger the proactive interaction stage corresponding to the stage trigger point based on the proactive interactive speech. The large-scale model instruction template may include the current massage area name, massage device type, current acupoint name, customer history of problems (from structured customer profiles), current progress percentage, summary of the most recent user feedback, and previously recommended directions (to avoid repetition), etc. By dynamically populating the large-scale model instruction template, each proactive interaction has highly personalized characteristics, rather than being a fixed script.

[0064] For example, the physical therapy service includes a total of 7 trigger points in the sensory interaction stage, the deep companionship interaction stage, and the warm closing interaction stage. The multiple detailed interaction stages corresponding to the 7 trigger points are shown in Table 1.

[0065] Table 1

[0066] Among them, the windows are silent at 25%-38% and 60%-72%, where the system does not speak actively, leaving room for the user's physical experience.

[0067] Furthermore, each trigger point within a phase undergoes an adaptive judgment before being triggered: if the current conversation is ongoing, the corresponding interaction phase of the trigger point within the phase is skipped; if the massage is paused, the corresponding interaction phase of the trigger point within the phase is skipped; for the "Life Resonance" phase, if the user has spoken proactively within a certain timeframe, the "Life Resonance" phase is skipped to avoid repeated chatting; for the "Warm Recommendation" phase, if a recommendation has already been successfully made in this service, the "Warm Recommendation" phase is skipped to avoid repeated sales pitches. Each round of interaction polling triggers at most one checkpoint to prevent continuous interruptions.

[0068] like Figure 4 As shown, in the experience summary phase, users are proactively asked about their feelings, and these feelings are recorded in their user profiles. In the recommendation phase, personalized recommendations for physical therapy products or membership cards are made based on the user profiles, and the process enters a completed state after receiving a service end signal or the user says goodbye.

[0069] In one embodiment, in addition to routing and distributing data received from the external device layer, the process orchestrator also performs context management for the entire physiotherapy service process. This context management provides context information to the service engine in the engine layer, and the context information provided for each business stage may only include the dialogue context corresponding to that business stage (i.e., the dialogue history information of that business stage). To make the historical dialogue corresponding to the business stage serve as the dialogue context, the process orchestrator's context management may include managing the stage instructions of the business stages, managing user profiles, and maintaining the dialogue context of each business node. This allows for determining the business stage based on the stage instructions and obtaining the corresponding dialogue context and user profile to determine the context information for that business stage.

[0070] Therefore, in one embodiment, the process orchestrator is further configured to: When entering the next business phase, obtain the phase identifier and structured information of the next business phase; The structured information is updated into the user profile, and the interactive audio of the current business stage is cleared, so that the service engine corresponding to the large model dialogue service and / or speech synthesis service can form context information based on the interactive audio of the next business stage, the stage identifier of the next business stage, and the updated user profile.

[0071] Specifically, entering the next business stage refers to jumping from the current business stage to the next. Jump triggering conditions include at least one of the following: the business stage's preset conditions are met, the user actively initiates a stage jump request, or the external device output status meets the jump requirements. Upon entering the next business stage, the structured information processed in the current business stage is updated and stored in the corresponding user's profile. This ensures that the user profile continuously records user information, demand information, and interaction information throughout the entire physiotherapy service process, providing complete user data support for the service execution of subsequent business stages. Simultaneously, the interaction audio of the current business stage is cleared. This is because, in this embodiment, each business stage only retains the interaction audio of the current business stage as the dialogue context. This satisfies the dialogue logic understanding requirements of the current business stage while avoiding the redundancy of context information and reduced processing efficiency of large models caused by including all dialogue audio throughout the entire process, thus balancing the integrity of dialogue information and service operation efficiency.

[0072] In this embodiment, the context information provided to the service engine corresponding to the large-model dialogue service and / or speech synthesis service each time includes a stage identifier, user profile information, and the dialogue context corresponding to the business stage. This preserves key historical information while preventing the context from growing indefinitely. The user profile information can be content related to the business stage within the user profile, a summary of the user profile, or even include the user profile itself.

[0073] Furthermore, user profiles are gradually accumulated at each stage to support the acquisition of user personas at each business stage. Specifically, the information that can be written into the user profile at each business stage can be pre-defined. Then, when each business stage is executed, the information that can be written into the user profile at that stage is retrieved and converted into structured information to update the user profile. For example, the information that can be written in the identity verification stage may include user identity information and historical package records; the information that can be written in the demand confirmation stage may include user needs, user preferences, and user confirmed packages actively recorded by the large model through tools; the information that can be written in the physiotherapy service stage may include massage parameter adjustment history; and the information that can be written in the experience summary stage may include overall experience feedback, etc.

[0074] This application avoids redundancy in user profiles by defining the information that can be written to each business stage. By writing the structured information acquired at each business stage into the user profile, memory transfer across business stages is achieved. Therefore, this application's embodiments maintain the continuity of service memory throughout the entire process while controlling the input length of the large model, avoiding problems such as decreased inference accuracy and slower processing speed due to excessively long inputs, thus balancing service continuity and operational efficiency. Furthermore, throughout the entire process of providing physiotherapy services, this application's embodiments utilize a service engine to automatically record user information (such as user physical condition and needs, user preferences, and user experience) during interactive services, and input this recorded user information into the user profile for proactive knowledge accumulation. This eliminates the need for manual extraction or extraction according to manually set rules, automatically adapting to dynamic changes in user needs, ensuring the accuracy and completeness of user information records, and providing precise data support for services at each business stage.

[0075] In one embodiment, such as Figure 5 As shown, the process orchestrator is configured with a first streaming interaction scheduling stage and a second streaming interaction scheduling stage. These two stages enable streaming, step-by-step scheduling of speech recognition results and large model output results to adapt to business scenarios involving streaming output from the large model and streaming output from speech recognition, thereby improving the response efficiency of voice interaction. Specifically, the first streaming interaction scheduling stage utilizes the large model dialogue service to infer user audio to obtain text data, and then uses this text data to drive the speech synthesis service to perform inference for real-time playback. The second streaming interaction scheduling stage invokes the physiotherapy robot to execute the corresponding operation of the text data to obtain execution feedback data, and then uses this feedback data to again drive the speech synthesis service to perform inference for feedback playback. This embodiment of the application improves the smoothness of interaction through the collaborative work of the first and second streaming interaction scheduling stages.

[0076] For example, in the scenario of adjusting massage intensity: In the first streaming interactive scheduling phase: the large model inference engine outputs a live verbal response (such as "Okay, I'll turn the intensity down for you"), while declaring the device control tools and parameters to be invoked; the verbal response is immediately sent to the speech synthesis engine for speech synthesis and playback, so that the user can hear the reply immediately.

[0077] The second streaming interactive scheduling stage: The process orchestrator executes the device control commands and collects the execution results (such as the current intensity has been adjusted to level 3). The large model inference engine outputs detailed feedback in a streaming manner based on the tool execution results (such as "It has been adjusted to level 3, how do you feel?"). The detailed feedback is sent to the speech synthesis engine for speech synthesis and broadcasting.

[0078] Among them, the system tools that can perform broadcasting through the coordinated operation of the first streaming interactive scheduling phase and the second streaming interactive scheduling phase may include: Identity setting tool: used to advance the identity verification sub-state; Package Search Tool: Used to search for a list of available physiotherapy packages and their details; Package Confirmation Tool: Used to confirm the package selected by the user and write it into the user profile; Massage intensity adjustment tool: used to send intensity adjustment commands to the massage device; Massage speed adjustment tool: used to send speed increase / decrease commands to the massage device; Massage start / stop control tool: Used to control the start, pause, resume, and stop of the massage; Massage status query tool: used to obtain parameters such as current massage intensity, speed, and progress; Music control tools: used to control the playback, pause, toggle, and volume of background music; User information recording tools: used to write user needs, preferences, or feedback into structured customer profiles; Dialogue flow control tool: Used to control whether the current dialogue continues to wait for the user's response.

[0079] Furthermore, in practical applications, during the collaboration between the first and second streaming interactive scheduling phases, upon receiving text data reported by the speech recognition service, the orchestrator immediately constructs contextual information based on this text data and sends this contextual information to the large model dialogue service. The large model dialogue service then calls the large model inference engine to perform inference and output text fragments. These text fragments are immediately reported back to the orchestrator via the large model dialogue service, which then sends them to the speech synthesis service. The speech synthesis service calls the speech synthesis engine, which begins synthesizing and pushing audio after accumulating sufficient text fragments. This approach, by immediately forwarding each text fragment to the speech synthesis engine after its generation, without waiting for the entire sentence to be generated, achieves end-to-end streaming processing from speech input to speech output. This significantly reduces the delay from the end of the user's speech to the system's first word of response, thus reducing interactive response latency.

[0080] In summary, this embodiment provides a physiotherapy robot service system based on multimodal identity recognition and streaming large model dialogue. The physiotherapy robot service system based on multimodal identity recognition and streaming large model dialogue comprises an external device layer, a service layer, an engine layer and a process orchestrator. The service layer comprises a plurality of independent service processes, each service process corresponds to an independent service engine in the engine layer, and the process orchestrator performs event-driven communication with the external device layer and the service layer, distributes data information received by the external device layer to the service processes, and invokes the service processes to execute physiotherapy services. The present application splits the physiotherapy service system of a physiotherapy robot into a plurality of independent service processes, and each service process is scheduled by the process orchestrator to work cooperatively, achieving high loose coupling and modularity. This enables each service process to be independently developed, iterated and maintained, improving the scalability of the system; meanwhile, cooperative scheduling is performed through a unified process orchestrator to adapt to the linked operation requirements of multiple devices and multiple business scenarios, improving the overall synergy and modularity of the system, and providing users with a more smooth and stable intelligent physiotherapy interaction service.

[0081] Based on the above physiotherapy robot service system based on multimodal identity recognition and streaming large model dialogue, this embodiment provides a physiotherapy robot service method based on multimodal identity recognition and streaming large model dialogue, the method comprises: receiving, via an external device layer, data information collected by an external device; invoking, via a process orchestrator, a service process in a service layer based on the data information, so as to invoke a corresponding service engine thereof through the service process to perform interaction service.

[0082] Based on the above physiotherapy robot service method based on multimodal identity recognition and streaming large model dialogue, this embodiment provides a computer-readable storage medium, wherein the computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the physiotherapy robot service method based on multimodal identity recognition and streaming large model dialogue described in the above embodiment.

[0083] Based on the above physiotherapy robot service method based on multimodal identity recognition and streaming large model dialogue, the present application further provides a terminal device, such as Figure 6As shown, it includes at least one processor 20; a display screen 21; and a memory 22, and may also include a communications interface 23 and a bus 24. The processor 20, display screen 21, memory 22, and communications interface 23 can communicate with each other via the bus 24. The display screen 21 is configured to display a preset user guide interface in the initial setup mode. The communications interface 23 can transmit information. The processor 20 can invoke logical instructions in the memory 22 to execute the methods described in the above embodiments.

[0084] Furthermore, the logical instructions in the aforementioned memory 22 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.

[0085] The memory 22, as a computer-readable storage medium, can be configured to store software programs, computer-executable programs, such as program instructions or modules corresponding to the methods in the embodiments of this disclosure. The processor 20 executes functional applications and data processing by running the software programs, instructions, or modules stored in the memory 22, thereby implementing the methods in the above embodiments.

[0086] The memory 22 may include a program storage area and a data storage area. The program storage area may store the operating system and application programs required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 22 may include high-speed random access memory (RAM) and non-volatile memory. Examples include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, as well as transient storage media.

[0087] Furthermore, the specific process of loading and executing multiple instruction processors in the aforementioned storage medium and terminal device has been described in detail in the above method, and will not be repeated here.

[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A physiotherapy robot service system based on multimodal identity recognition and streaming large-scale model dialogue, characterized in that, The physiotherapy robot service system based on multimodal identity recognition and streaming large-scale model dialogue specifically includes: The external device layer is used to receive data information collected by external devices. The service layer includes several service processes, which are used to implement the physiotherapy services of the physiotherapy robot. The several service processes include at least speech recognition service, large model dialogue service and speech synthesis service. The engine layer consists of several independent service engines, each corresponding to a separate service process. Each service engine establishes a signal connection with its corresponding service process through a permanent online connection strategy to provide operational support for the service process. The process orchestrator communicates with the external device layer and the service layer via event-driven communication. It is used to distribute data information received by the external device layer to the service process and call the service process to execute physiotherapy services.

2. The physiotherapy robot service system based on multimodal identity recognition and streaming large-scale model dialogue according to claim 1, characterized in that, The process orchestrator is configured with a first streaming interactive scheduling phase and a second streaming interactive scheduling phase. The first streaming interaction scheduling stage is used to use the large model dialogue service to infer the user's audio to obtain text data, and to drive the speech synthesis service to infer the text data for real-time broadcasting. The second streaming interactive scheduling stage is used to call the physiotherapy robot to perform the execution operation corresponding to the text data to obtain execution feedback data, and then use the execution feedback data to drive the speech synthesis service to perform inference for feedback broadcasting.

3. The physiotherapy robot service system based on multimodal identity recognition and streaming large-scale model dialogue according to claim 1, characterized in that, The voice recognition service is equipped with a voiceprint gating system, which is used to dynamically select which audio data to release to the service engine corresponding to the voice recognition service according to the business scenario throughout the entire process of the physiotherapy service. The audio data is either user audio or silent audio.

4. The physiotherapy robot service system based on multimodal identity recognition and streaming large-scale model dialogue according to claim 3, characterized in that, The voiceprint gating system is configured with multiple dynamically switchable operating modes, including at least a silent mode, a direct-access mode, a data acquisition mode, and a voiceprint verification mode. The mute mode is used to allow mute audio to be transmitted to the service engine corresponding to the speech recognition service. The direct access mode is used to allow user audio to be transmitted to the service engine corresponding to the speech recognition service; The acquisition mode is used to capture user audio segments for writing into the voiceprint template bank; The voiceprint verification mode is used to allow the user's audio to be sent to the service engine corresponding to the speech recognition service when the user's audio passes the voiceprint verification, and to allow silent audio to be sent to the service engine corresponding to the speech recognition service when the user's audio fails the voiceprint verification. The voiceprint verification mode includes a single-person voiceprint verification mode and a multi-person voiceprint verification mode.

5. The physiotherapy robot service system based on multimodal identity recognition and streaming large-scale model dialogue according to claim 3, characterized in that, The voiceprint gating system is configured with a voiceprint scoring mechanism, which is used to extract voiceprint embedding features from user audio frames received by the external device layer distributed by the process orchestrator, and to match the voiceprint embedding features with the corresponding voiceprint templates in the voiceprint template bank to verify the user audio.

6. The physiotherapy robot service system based on multimodal identity recognition and streaming large-scale dialogue according to claim 3 or 5, characterized in that, The voiceprint gating system is equipped with a pre-buffering retransmission mechanism and / or an audio final review mechanism, wherein... The pre-buffering and resending mechanism is used to cache user audio when the voiceprint gating system is not enabled, and resend the cached user audio to the service engine corresponding to the speech recognition service when the voiceprint gating system is enabled. The voiceprint gating system is equipped with an audio final review mechanism. After a single round of voice interaction, the audio final review mechanism is used to extract the voiceprint embedding features of the cached user audio and match them with the corresponding voiceprint template in the voiceprint template bank to verify the user audio.

7. The physiotherapy robot service system based on multimodal identity recognition and streaming large-scale model dialogue according to claim 1, characterized in that, The physiotherapy service is divided into multiple business stages according to the execution process. These multiple business stages include at least the identity verification stage, the needs confirmation stage, the bed guidance stage, the physiotherapy service stage, the experience summary stage, and the recommendation stage.

8. The physiotherapy robot service system based on multimodal identity recognition and streaming large-scale dialogue according to claim 7, characterized in that, A watershed mechanism is set between the demand confirmation stage and the bed-on guidance stage. The watershed mechanism takes the user's confirmation of starting the physiotherapy service as the watershed marker. Before the watershed marker, the absence of a detected user's intention to discontinue the physiotherapy service is used as a constraint condition. After the watershed marker, the absence of a detected user's intention to discontinue the physiotherapy service is not used as a constraint condition.

9. The physiotherapy robot service system based on multimodal identity recognition and streaming large-scale model dialogue according to claim 7, characterized in that, The process orchestrator is also used for: When entering the next business phase, obtain the phase identifier and structured information of the next business phase; The structured information is updated into the user profile, and the interactive audio of the current business stage is cleared, so that the service engine corresponding to the large model dialogue service and / or speech synthesis service can form context information based on the interactive audio of the next business stage, the stage identifier of the next business stage, and the updated user profile.

10. The physiotherapy robot service system based on multimodal identity recognition and streaming large-scale model dialogue according to claim 7, characterized in that, During the physiotherapy service phase, the process orchestrator is also used to call the service layer for proactive interaction based on the service progress of the physiotherapy service. The proactive interaction includes multiple proactive interaction phases, which include at least a sensory landing interaction phase, a deep companionship interaction phase, and a warm closing interaction phase. The sensory landing interaction stage is used to guide the interaction before the progress of the physiotherapy service reaches the first progress threshold in order to help the user experience the physiotherapy service. The deep companionship and interaction phase is used to conduct emotional interaction with the user when the progress of the physical therapy service is between the first progress threshold and the second progress threshold, so as to establish an emotional connection with the user. The warm closing interaction phase is used to conduct experiential interaction after the physiotherapy service progress reaches the second progress threshold in order to understand the user's physiotherapy service experience.

11. The physiotherapy robot service system based on multimodal identity recognition and streaming large-scale dialogue according to claim 10, characterized in that, The service layer actively interacts based on the progress of the physiotherapy service, specifically by invoking the service layer: Based on the monitoring phase trigger point of the physiotherapy service progress; When a stage trigger point is detected, obtain the large model instruction template corresponding to the detected stage trigger point; Based on the large model instruction template, obtain the information required for active interaction, and call the speech synthesis service to form active interaction speech based on the information required for active interaction, so as to trigger the active interaction stage corresponding to the trigger point based on the active interaction speech.

12. The physiotherapy robot service system based on multimodal identity recognition and streaming large-scale model dialogue according to claim 7, characterized in that, The process orchestrator is also used for: Acquire facial images captured by an external device layer, and perform identity recognition based on the facial images; When the user's identity is recognized, the voice recognition service is invoked to verify the voiceprint. When the user's identity is not recognized, the speech recognition service is invoked to guide the user to register and build a voiceprint template. During the process of guiding the user to register, the user's name is obtained through interaction with the user.

13. The physiotherapy robot service system based on multimodal identity recognition and streaming large-scale model dialogue according to claim 1, characterized in that, The aforementioned service processes also include a user data management service, and the engine layer further includes a user data management engine; the user data management service is used to manage user profiles, voiceprint template banks, and facial images.

14. A method for providing physiotherapy robot services based on multimodal identity recognition and streaming large-scale model dialogue, characterized in that, The physiotherapy robot service system based on multimodal identity recognition and streaming large-scale model dialogue as described in any one of claims 1-13 is used, wherein the physiotherapy robot service method based on multimodal identity recognition and streaming large-scale model dialogue specifically includes: Receive data information collected by external devices through the external device layer; The process orchestrator invokes service processes in the service layer based on the data information, and the service processes invoke their corresponding service engines to perform interactive services.

15. A physiotherapy robot, characterized in that, The physiotherapy robot is equipped with a physiotherapy robot service system based on multimodal identity recognition and streaming large model dialogue as described in any one of claims 1-13.