User information processing method and device based on XR equipment, equipment and medium
By deploying facial recognition and voice recognition models in XR devices and combining them with large language models, the problem of XR devices failing to effectively utilize image and voice information in hospital scenarios has been solved. This has enabled rapid acquisition of user information and guidance prompts, and improved the communication efficiency between doctors and patients.
Patent Information
- Application Number
- CN202510784509.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-12
AI Technical Summary
In existing technologies, XR devices fail to effectively integrate image acquisition, voice acquisition, information display, voice recognition, and image recognition functions in hospital scenarios, resulting in doctors and nurses being unable to obtain patient information in a timely manner for communication and guidance.
By deploying facial recognition models, speech recognition models, and natural language processing models in XR devices, user image and voice data are obtained, identified and processed, user prompt information is generated, and displayed in the display module, combined with a large language model to provide guidance prompts.
It enables users wearing XR devices to quickly obtain user information of the communication object and timely display guidance prompts, improving the communication efficiency between doctors and patients.
Smart Images

Figure CN120636402A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of mixed reality technology, and in particular to a method, apparatus, device, and medium for processing user information based on XR devices. Background Art
[0002] At present, XR devices (XR stands for Extended Reality, which means extended reality and is a hybrid technology of virtual reality, augmented reality, mixed reality and other technologies) have been increasingly used in daily life. For example, if a user wears an XR device (such as XR glasses), it can be used as an auxiliary device for translation when communicating with others; for example, the user can display navigation information through the lenses of XR glasses to guide the user to the destination. However, at present, XR devices have not yet been specifically applied in the hospital scenario, and the image acquisition, voice acquisition, information display, voice recognition and image recognition functions of XR devices have not been effectively integrated. Doctors, nurses and other users cannot obtain relevant information about patients in a timely manner for subsequent communication, guidance and other operations. Summary of the Invention
[0003] The embodiments of the present invention provide a user information processing method, apparatus, device and medium based on XR devices, aiming to solve the problem in the prior art that XR devices have not yet been specifically applied in the hospital scenario, and the image acquisition, voice acquisition, information display, voice recognition and image recognition functions of XR devices have not been effectively and comprehensively utilized, so that doctors, nurses and other users cannot obtain relevant information about patients in a timely manner for subsequent communication, guidance and other operations.
[0004] In a first aspect, an embodiment of the present invention provides a user information processing method based on an XR device, which is applied to the XR device and includes:
[0005] In response to a user identification instruction, acquiring a current user image corresponding to the user identification instruction;
[0006] Recognize the current user image based on the deployed face recognition model to obtain a current user recognition result;
[0007] If it is determined that the current user identification result belongs to the preset specific user information set, obtaining current user prompt information corresponding to the current user identification result and displaying it on a display module corresponding to the XR device;
[0008] Acquire user communication voice data collected based on the current user prompt information;
[0009] Performing speech recognition on the current communication speech data based on the deployed speech recognition model to obtain a current speech recognition result;
[0010] If a text processing instruction is detected, a text summary or automatic summary processing is performed on the current speech recognition result based on a pre-trained natural language processing model to obtain a current text processing result.
[0011] In a second aspect, an embodiment of the present invention further provides a user information processing apparatus based on an XR device, configured on the XR device, comprising:
[0012] A user image acquisition unit, configured to acquire, in response to a user identification instruction, a current user image corresponding to the user identification instruction;
[0013] A user face recognition unit is used to recognize the current user image based on the deployed face recognition model to obtain a current user recognition result;
[0014] a current prompt information generating unit, configured to, if it is determined that the current user recognition result belongs to a preset specific user information set, obtain current user prompt information corresponding to the current user recognition result and display it on a display module corresponding to the XR device;
[0015] A voice data collection unit, configured to obtain user communication voice data collected based on the current user prompt information;
[0016] A speech recognition unit, configured to perform speech recognition on the current communication speech data based on a deployed speech recognition model to obtain a current speech recognition result;
[0017] The text intelligent processing unit is used to perform text summarization or automatic summary processing on the current speech recognition result based on a pre-trained natural language processing model if a text processing instruction is detected to obtain the current text processing result.
[0018] In a third aspect, an embodiment of the present invention further provides a computer device comprising a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the method described in the first aspect is implemented.
[0019] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the method described in the first aspect can be implemented.
[0020] Embodiments of the present invention provide a user information processing method, apparatus, device, and medium based on an XR device. The method includes: in response to a user recognition instruction, obtaining a current user image corresponding to the user recognition instruction; recognizing the current user image based on a deployed face recognition model to obtain a current user recognition result; if it is determined that the current user recognition result belongs to a preset specific user information set, obtaining current user prompt information corresponding to the current user recognition result and displaying it in a display module corresponding to the XR device; obtaining user communication voice data collected based on the current user prompt information; performing voice recognition on the current communication voice data based on a deployed voice recognition model to obtain a current voice recognition result; if a text processing instruction is detected, performing text summarization or automatic summary processing on the current voice recognition result based on a pre-trained natural language processing model to obtain a current text processing result. The embodiments of the present invention enable the user wearing the XR device to perform user image recognition and voice recognition on the communication partner in a timely manner, and can not only quickly obtain the user information of the communication partner, but also quickly obtain and intuitively display the guidance prompt information required during the communication process. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0022] Figure 1 A schematic diagram of an application scenario of the user information processing method based on an XR device provided by an embodiment of the present invention;
[0023] Figure 2 A flowchart of a method for processing user information based on an XR device provided in an embodiment of the present invention;
[0024] Figure 3 A schematic diagram of a sub-flow diagram of a user information processing method based on an XR device provided in an embodiment of the present invention;
[0025] Figure 4 A schematic block diagram of a user information processing apparatus based on an XR device provided in an embodiment of the present invention;
[0026] Figure 5 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0028] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0029] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0030] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0031] Please also refer to Figure 1 and Figure 2 ,in Figure 1 Schematic diagram of a scenario of a method for processing user information based on an XR device according to an embodiment of the present invention. Figure 2 FIG is a flow chart of a method for processing user information based on an XR device according to an embodiment of the present invention. Figure 1 As shown, the user information processing method based on the XR device provided in an embodiment of the present invention is applied to the XR device 10, the XR device 10 is communicatively connected to the server 20, and the XR device 10 can also be communicatively connected to the receiving terminal 30 (such as the smart terminal used by the user).
[0032] like Figure 2 As shown, the method includes the following steps S110-S160.
[0033] S110 . In response to a user identification instruction, obtain a current user image corresponding to the user identification instruction.
[0034] In this embodiment, the technical solution is described with the server as the execution entity. The core components such as the processor, voice acquisition module, image acquisition module, display module (generally integrated on the lens of the XR device), storage module, power module, etc. are integrated on the XR device. The voice acquisition module, image acquisition module, display module, storage module, power module, etc. are all connected to the processor, and lightweight face recognition model, voice recognition model, natural language processing model and other artificial intelligence models are deployed therein. When a doctor user or a nurse user wears the XR device, they can guide communication with the patient in the daily communication process. For example, taking a doctor user wearing the XR device to communicate with the patient as an example, when the XR device detects that a user enters the image acquisition field of view corresponding to the image acquisition module, it can automatically trigger the generation of a user recognition instruction, thereby obtaining the current user image through the image acquisition module. Of course, the XR device can also communicate with the hospital's server. When the calling system in the server prompts the doctor user with the patient user information of the next patient (such as the desensitized name, gender, etc.), it can also generate patient-related user identification instructions in the server's calling system, and after the patient enters the doctor's office, the current user image can be obtained through the image acquisition module of the XR device.
[0035] Among them, when a patient uses the appointment registration system corresponding to the call system in the server, he or she can log in to the appointment registration system in the server through the patient's user terminal and make a user appointment by using the patient's real-name authentication and desensitized user identity information. After the user appointment is completed, the patient's user identity information is added to the call system, and the queue is called during the corresponding appointment time period.
[0036] S120: Recognize the current user image based on the deployed face recognition model to obtain a current user recognition result.
[0037] In this embodiment, since a pre-trained and lightweight face recognition model has been deployed in the XR device, after obtaining the current user image, face image feature extraction and face image feature comparison can be performed to obtain the current user recognition result.
[0038] For example, when a patient has completed user registration in advance through the appointment registration system corresponding to the call system in the server, the patient can authorize the upload of his or her user avatar image and the real-name authenticated and desensitized user identity information to the registration system of the server. The registration system of the server can extract the facial image features of the patient's user face image and store it in a designated storage area in the server, such as a trusted execution area. Afterwards, when the patient actually arrives at the corresponding doctor's office in the hospital according to the registration information, the doctor's XR device can perform facial recognition processing after collecting the patient's current user image in combination with the local facial recognition model of the XR device and the facial image features in the designated storage area in the server to quickly obtain the current user recognition result, and the current user recognition result at least includes user identity information, etc.
[0039] S130: If it is determined that the current user identification result belongs to a preset specific user information set, current user prompt information corresponding to the current user identification result is obtained and displayed in a display module corresponding to the XR device.
[0040] In this embodiment, when it is determined that the current user identification result belongs to a preset specific user information set, it means that it is necessary to remind the doctor user of the relevant users that he is concerned about. At this time, the user identity information corresponding to the current user identification result can be collected to generate the current user prompt information and displayed in the display module corresponding to the XR device. After viewing the current user prompt information, the doctor user can refer to the current user prompt information to communicate with the patient, so as to obtain the required key information more quickly. Among them, the specific user information set can be regularly updated by the hospital staff, and the regularly updated specific user information set can also be regularly sent to the XR device for local storage.
[0041] Of course, if it is determined that the current user recognition result does not belong to the preset specific user information set, the XR device directly collects the user communication voice data and executes subsequent steps S150 to S160.
[0042] In one embodiment, if Figure 3 As shown, step S130 includes:
[0043] S131. If it is determined that the user unique identification code corresponding to the current user identification result is the same as the user unique identification code of one of the specific user information in the specific user information set, then determining that the current user identification result belongs to the preset specific user information set, and obtaining target specific user information corresponding to the current user identification result from the specific user information set;
[0044] S132. Obtain the desensitized user name, user tag, user's last visit information, disease-specific prompt information, historical chat record summary information, current user process node information and current communication prompt information from the target specific user information as the current user prompt information.
[0045] In this embodiment, in addition to the user identity information, the current user identification result may also include a unique user identification code. If it is determined that the unique user identification code corresponding to the current user identification result is the same as the unique user identification code of one of the specific user information in the specific user information set, then the patient's user data information belongs to the specific user information set, and the current user identification result may be determined to belong to the preset specific user information set.
[0046] Afterwards, combined with the local user database of the XR device and the comprehensive user database in the server, the desensitized user name corresponding to the user unique identification code of the current user identification result (such as displaying the desensitized user name in the form of Li**, etc.), user tags (such as emotionally sensitive users, talkative users, silent users, etc. tags), user's last visit information (such as the time of the patient user's last visit to the hospital and a brief introduction to related handling matters), special disease prompt information (such as whether the patient has a special disease that needs special attention), historical chat record summary information (such as key extraction information of the chat record of the patient user's last visit to the hospital to communicate with the doctor), current user process node information (for example, when the patient user sees a doctor in the hospital that day, in addition to entering the doctor's clinic, he also needs to go to multiple other places for testing, payment or medication, etc. Every time the patient user enters the doctor's clinic, the doctor user can quickly determine the current user process node information corresponding to the patient user, such as the current user process node information is to get the order for re-examination, payment or medication after testing, etc.) and current communication prompt information. The current communication prompt information can be generated by combining user tags, disease-specific prompt information, historical chat record summary information, and current user process node information. More specifically, the user tags, disease-specific prompt information, historical chat record summary information, and current user process node information can be input into the large language model as context information in combination with the lightweight large language model deployed in the XR device to obtain the current communication prompt information that guides the doctor user to communicate with the patient.
[0047] After obtaining the above information of the user corresponding to the unique user identification code of the current user identification result, the current user prompt information is directly composed of the desensitized user name, user label, user's last visit information, special disease prompt information, historical chat record summary information, current user process node information and current communication prompt information. Because the above-mentioned current user prompt information includes multiple contents, the current user prompt information can be displayed in pages. For example, at least 3 page pages can be displayed in the display module corresponding to the XR device, where the first page page displays the desensitized user name, user label, user's last visit information, special disease prompt information, the second page page displays the historical chat record summary information, and the third page page displays the current user process node information and current communication prompt information. Different page pages require the doctor user to click the relevant physical buttons on the XR device (such as the page up button, the page down button, etc.) to turn the page and view the page.
[0048] S140: Acquire user communication voice data collected based on the current user prompt information.
[0049] In this embodiment, the doctor user can communicate with the patient on-site in combination with the current user prompt information, and on the premise that the patient authorizes the acquisition of voice data, the user communication voice data during the communication process is collected legally and compliantly to obtain the user communication voice data.
[0050] In one embodiment, before step S140, the method further includes:
[0051] Based on the current communication prompt information in the current user prompt information, a plurality of communication prompt texts are generated in sequence, and the plurality of communication prompt texts are displayed in a display module corresponding to the XR device.
[0052] In this embodiment, in order to guide the doctor user to communicate more efficiently with the patient user on-site, the communication text can be split based on the current communication prompt information in the current user prompt information to generate multiple communication prompt texts. For example, the generated multiple communication prompt texts include at least N1 communication prompt texts (N1 is a positive integer), each of the above communication prompt texts has a sorting order, and the multiple communication prompt texts are displayed in the display module of the XR device in sequence according to the corresponding sorting order.
[0053] When the doctor user communicates with the patient with reference to the multiple communication prompt texts, the voice data during the communication process will be collected by the voice collection module to obtain the user communication voice data.
[0054] S150: Perform speech recognition on the current communication speech data based on the deployed speech recognition model to obtain a current speech recognition result.
[0055] In this embodiment, after obtaining the current communication voice data, voice recognition may not be performed in real time. Instead, after the communication between the doctor user and the patient user is completed, the data is uploaded to the server as communication history voice data for voice recognition. Of course, in order to obtain the voice communication content in real time, after the communication between the doctor user and the patient user is completed, voice recognition and text extraction are promptly performed on the XR device based on the voice recognition model to obtain the current voice recognition result. The current voice recognition result obtained can be used as the communication minutes information of the communication between the doctor user and the patient user, and updated to the historical chat record summary information corresponding to the patient user.
[0056] Of course, a speech recognition model with dialect recognition or foreign language translation functions can also be pre-deployed in the XR device. When the patient communicates in a dialect or foreign language that the doctor user cannot understand, the doctor user can also obtain the information expressed by the patient in a timely manner, and can also translate the reply information spoken by the doctor user into the corresponding dialect or foreign language for voice playback.
[0057] In one embodiment, after step S150, the method further includes:
[0058] If it is detected that the current speech recognition result belongs to other communication texts that are not included in the multiple communication prompt texts, the other communication texts are input into the locally pre-deployed large language model to obtain the model output prompt text, and the model output prompt text information is displayed in the display module corresponding to the XR device.
[0059] In this embodiment, when a doctor user communicates with a patient user on-site, in addition to the doctor user communicating based on the guidance of the current communication prompt information, the patient user may also have some additional questions. If this additional question is about other professional fields unknown to the doctor user, the patient user's other communication text can also be input into the locally pre-deployed large language model to obtain the model output prompt text, and the model output prompt text information is displayed in the display module corresponding to the XR device. At this time, the doctor user can refer to the model output prompt text information displayed in the display module to further communicate with the patient user.
[0060] In one embodiment, after step S150, the method further includes:
[0061] If a user process node consultation request is detected, user process node information to be processed is generated based on the current user process node information in the current user prompt information, and the user process node information to be processed is sent to the receiving terminal selected on the XR device.
[0062] In this embodiment, after the doctor user completes this round of communication with the patient user, if the patient user needs to ask further questions to obtain the next user process node, the current user process node information in the current user prompt information that has been obtained in the doctor user's XR device generates the user process node information to be processed (the user process node information to be processed can prompt the patient user to process the user process nodes that need to be processed in the next few steps, such as blood test, test result collection, return to the clinic to review the blood test result sheet, and wait for the user process node information to be processed), and after the doctor user selects the receiving terminal of the patient user, the doctor user's XR device sends the user process node information to be processed to the receiving terminal. By viewing the user process node information to be processed, the patient user can more clearly understand the subsequent user process nodes and perform corresponding processing.
[0063] S160: If a text processing instruction is detected, a text summary or automatic summary processing is performed on the current speech recognition result based on a pre-trained natural language processing model to obtain a current text processing result.
[0064] In this embodiment, when the doctor user confirms that the entire communication process with the patient user is completed, he or she can operate to select to end the communication, and further select the text processing button on the user interaction interface displayed by the XR device. At this time, the XR device can perform text summary or automatic summary processing on the current speech recognition result based on the natural language processing model to obtain the current text processing result. The obtained current text processing result and the doctor user identity information and the user identity information corresponding to the current user recognition result are packaged and compressed, and uploaded to the server for data archiving. Of course, in order to ensure data security, the current text processing result and the doctor user identity information and the user identity information corresponding to the current user recognition result can also be encrypted, and then the encrypted data can be uploaded to the server for information storage.
[0065] In one embodiment, after step S160, the method further includes:
[0066] Acquire current user detection data of a user corresponding to the current user image;
[0067] Obtaining the current text processing result and the current classification result of the current user detection data based on a pre-trained classification model;
[0068] If it is determined that the current classification result belongs to the first preset classification type, the current text processing result and the current user detection data are stored in a storage space corresponding to the current text processing result and the current user detection data.
[0069] In this embodiment, after the doctor user synchronously obtains the current user detection data from the server after performing relevant indicator detection (such as blood indicator detection) on the patient user, the current text processing result and the current user detection data can be combined into comprehensive input data, and the comprehensive input data can be input into a classification model (such as a convolutional neural network, etc.) to obtain the current classification result.
[0070] Afterwards, the current classification result can be matched with a first preset classification type. If the current classification result is determined to belong to the first preset classification type, for example, if the patient's blood type is determined to be a rare blood type based on the current user's test data, the current text processing result and the current user's test data can be stored in a storage space corresponding to the current text processing result and the current user's test data. This storage space can then be used to search for users of this type and conduct further user communication.
[0071] In one embodiment, after step S150, the method further includes:
[0072] If a schedule prompt request is detected, the corresponding current schedule prompt information is obtained, and the schedule time information and schedule navigation information corresponding to the current schedule prompt information are pre-stored in the corresponding pre-storage space.
[0073] In this embodiment, in addition to daily medical work, the doctor user may also have other schedules such as academic conferences, daily training, etc. When the XR device obtains the schedule reminder information pre-entered by the doctor user (including multiple schedule items, and each schedule item corresponds to an effective time interval), the current schedule reminder information of the day can be extracted in combination with the date. The schedule time information and schedule navigation information corresponding to the current schedule reminder information can be pre-stored in the corresponding pre-storage space, and before the current system time reaches the corresponding schedule effective time interval, the schedule items to be effective and their schedule time information and schedule navigation information are extracted from the pre-storage space and displayed in a timely manner on the display model of the XR device.
[0074] It can be seen that the implementation of the embodiment of this method can enable the user wearing the XR device to perform user image recognition and voice recognition of the communication object in a timely manner, not only to quickly obtain the user information of the communication object, but also to quickly obtain and intuitively display the guidance prompt information required during the communication process.
[0075] Figure 4 : is a schematic block diagram of a user information processing device based on an XR device provided by an embodiment of the present invention. Figure 4 As shown, corresponding to the above-mentioned user information processing method based on XR device, the present invention also provides a user information processing device 100 based on XR device. The user information processing device 100 based on XR device includes a unit for executing the above-mentioned user information processing method based on XR device. Figure 4 The user information processing device 100 based on the XR device includes: a user image acquisition unit 110, a user face recognition unit 120, a current prompt information generation unit 130, a voice data collection unit 140, a voice recognition unit 150 and a text intelligent processing unit 160.
[0076] The user image acquisition unit 110 is configured to, in response to a user identification instruction, acquire a current user image corresponding to the user identification instruction.
[0077] In this embodiment, the technical solution is described with the server as the execution entity. The core components such as the processor, voice acquisition module, image acquisition module, display module (generally integrated on the lens of the XR device), storage module, power module, etc. are integrated on the XR device. The voice acquisition module, image acquisition module, display module, storage module, power module, etc. are all connected to the processor, and lightweight face recognition model, voice recognition model, natural language processing model and other artificial intelligence models are deployed therein. When a doctor user or a nurse user wears the XR device, they can guide communication with the patient in the daily communication process. For example, taking a doctor user wearing the XR device to communicate with the patient as an example, when the XR device detects that a user enters the image acquisition field of view corresponding to the image acquisition module, it can automatically trigger the generation of a user recognition instruction, thereby obtaining the current user image through the image acquisition module. Of course, the XR device can also communicate with the hospital's server. When the calling system in the server prompts the doctor user with the patient user information of the next patient (such as the desensitized name, gender, etc.), it can also generate patient-related user identification instructions in the server's calling system, and after the patient enters the doctor's office, the current user image can be obtained through the image acquisition module of the XR device.
[0078] Among them, when a patient uses the appointment registration system corresponding to the call system in the server, he or she can log in to the appointment registration system in the server through the patient's user terminal and make a user appointment by using the patient's real-name authentication and desensitized user identity information. After the user appointment is completed, the patient's user identity information is added to the call system, and the queue is called during the corresponding appointment time period.
[0079] The user face recognition unit 120 is configured to recognize the current user image based on a deployed face recognition model to obtain a current user recognition result.
[0080] In this embodiment, since a pre-trained and lightweight face recognition model has been deployed in the XR device, after obtaining the current user image, face image feature extraction and face image feature comparison can be performed to obtain the current user recognition result.
[0081] For example, when a patient has completed user registration in advance through the appointment registration system corresponding to the call system in the server, the patient can authorize the upload of his or her user avatar image and the real-name authenticated and desensitized user identity information to the registration system of the server. The registration system of the server can extract the facial image features of the patient's user face image and store it in a designated storage area in the server, such as a trusted execution area. Afterwards, when the patient actually arrives at the corresponding doctor's office in the hospital according to the registration information, the doctor's XR device can perform facial recognition processing after collecting the patient's current user image in combination with the local facial recognition model of the XR device and the facial image features in the designated storage area in the server to quickly obtain the current user recognition result, and the current user recognition result at least includes user identity information, etc.
[0082] The current prompt information generating unit 130 is configured to obtain current user prompt information corresponding to the current user recognition result if it is determined that the current user recognition result belongs to a preset specific user information set, and display the current user prompt information on a display module corresponding to the XR device.
[0083] In this embodiment, when it is determined that the current user identification result belongs to a preset specific user information set, it means that it is necessary to remind the doctor user of the relevant users that he is concerned about. At this time, the user identity information corresponding to the current user identification result can be collected to generate the current user prompt information and displayed in the display module corresponding to the XR device. After viewing the current user prompt information, the doctor user can refer to the current user prompt information to communicate with the patient, so as to obtain the required key information more quickly. Among them, the specific user information set can be regularly updated by the hospital staff, and the regularly updated specific user information set can also be regularly sent to the XR device for local storage.
[0084] In one embodiment, the current prompt information generating unit 130 is specifically configured to:
[0085] If it is determined that the user unique identification code corresponding to the current user identification result is the same as the user unique identification code of one of the specific user information in the specific user information set, then the current user identification result is determined to belong to the preset specific user information set, and the target specific user information corresponding to the current user identification result is obtained from the specific user information set;
[0086] The desensitized user name, user tag, user's last visit information, special disease prompt information, historical chat record summary information, current user process node information and current communication prompt information are obtained from the target specific user information as the current user prompt information.
[0087] In this embodiment, in addition to the user identity information, the current user identification result may also include a unique user identification code. If it is determined that the unique user identification code corresponding to the current user identification result is the same as the unique user identification code of one of the specific user information in the specific user information set, then the patient's user data information belongs to the specific user information set, and the current user identification result may be determined to belong to the preset specific user information set.
[0088] Afterwards, combined with the local user database of the XR device and the comprehensive user database in the server, the desensitized user name corresponding to the user unique identification code of the current user identification result (such as displaying the desensitized user name in the form of Li**, etc.), user tags (such as emotionally sensitive users, talkative users, silent users, etc. tags), user's last visit information (such as the time of the patient user's last visit to the hospital and a brief introduction to related handling matters), special disease prompt information (such as whether the patient has a special disease that needs special attention), historical chat record summary information (such as key extraction information of the chat record of the patient user's last visit to the hospital to communicate with the doctor), current user process node information (for example, when the patient user sees a doctor in the hospital that day, in addition to entering the doctor's clinic, he also needs to go to multiple other places for testing, payment or medication, etc. Every time the patient user enters the doctor's clinic, the doctor user can quickly determine the current user process node information corresponding to the patient user, such as the current user process node information is to get the order for re-examination, payment or medication after testing, etc.) and current communication prompt information. The current communication prompt information can be generated by combining user tags, disease-specific prompt information, historical chat record summary information, and current user process node information. More specifically, the user tags, disease-specific prompt information, historical chat record summary information, and current user process node information can be input into the large language model as context information in combination with the lightweight large language model deployed in the XR device to obtain the current communication prompt information that guides the doctor user to communicate with the patient.
[0089] After obtaining the above information of the user corresponding to the unique user identification code of the current user identification result, the current user prompt information is directly composed of the desensitized user name, user label, user's last visit information, special disease prompt information, historical chat record summary information, current user process node information and current communication prompt information. Because the above-mentioned current user prompt information includes multiple contents, the current user prompt information can be displayed in pages. For example, at least 3 page pages can be displayed in the display module corresponding to the XR device, where the first page page displays the desensitized user name, user label, user's last visit information, special disease prompt information, the second page page displays the historical chat record summary information, and the third page page displays the current user process node information and current communication prompt information. Different page pages require the doctor user to click the relevant physical buttons on the XR device (such as the page up button, the page down button, etc.) to turn the page and view the page.
[0090] The voice data collecting unit 140 is configured to obtain user communication voice data collected based on the current user prompt information.
[0091] In this embodiment, the doctor user can communicate with the patient on-site in combination with the current user prompt information, and on the premise that the patient authorizes the acquisition of voice data, the user communication voice data during the communication process is collected legally and compliantly to obtain the user communication voice data.
[0092] In one embodiment, the user information processing apparatus 100 based on the XR device further includes:
[0093] A communication prompt text generation unit is used to sequentially generate multiple communication prompt texts based on the current communication prompt information in the current user prompt information, and display the multiple communication prompt texts in a display module corresponding to the XR device.
[0094] In this embodiment, in order to guide the doctor user to communicate more efficiently with the patient user on-site, the communication text can be split based on the current communication prompt information in the current user prompt information to generate multiple communication prompt texts. For example, the generated multiple communication prompt texts include at least N1 communication prompt texts (N1 is a positive integer), each of the above communication prompt texts has a sorting order, and the multiple communication prompt texts are displayed in the display module of the XR device in sequence according to the corresponding sorting order.
[0095] When the doctor user communicates with the patient with reference to the multiple communication prompt texts, the voice data during the communication process will be collected by the voice collection module to obtain the user communication voice data.
[0096] The speech recognition unit 150 is configured to perform speech recognition on the current communication speech data based on the deployed speech recognition model to obtain a current speech recognition result.
[0097] In this embodiment, after obtaining the current communication voice data, voice recognition may not be performed in real time. Instead, after the communication between the doctor user and the patient user is completed, the data is uploaded to the server as communication history voice data for voice recognition. Of course, in order to obtain the voice communication content in real time, after the communication between the doctor user and the patient user is completed, voice recognition and text extraction are promptly performed on the XR device based on the voice recognition model to obtain the current voice recognition result. The current voice recognition result obtained can be used as the communication minutes information of the communication between the doctor user and the patient user, and updated to the historical chat record summary information corresponding to the patient user.
[0098] Of course, a speech recognition model with dialect recognition or foreign language translation functions can also be pre-deployed in the XR device. When the patient communicates in a dialect or foreign language that the doctor user cannot understand, the doctor user can also obtain the information expressed by the patient in a timely manner, and can also translate the reply information spoken by the doctor user into the corresponding dialect or foreign language for voice playback.
[0099] In one embodiment, the user information processing apparatus 100 based on the XR device further includes:
[0100] The model output prompt text acquisition unit is used to input the other communication text into the locally pre-deployed large language model if it is detected that the current speech recognition result belongs to other communication text that is not included in the multiple communication prompt texts, obtain the model output prompt text, and display the model output prompt text information in the display module corresponding to the XR device.
[0101] In this embodiment, when a doctor user communicates with a patient user on-site, in addition to the doctor user communicating based on the guidance of the current communication prompt information, the patient user may also have some additional questions. If this additional question is about other professional fields unknown to the doctor user, the patient user's other communication text can also be input into the locally pre-deployed large language model to obtain the model output prompt text, and the model output prompt text information is displayed in the display module corresponding to the XR device. At this time, the doctor user can refer to the model output prompt text information displayed in the display module to further communicate with the patient user.
[0102] In one embodiment, the user information processing apparatus 100 based on the XR device further includes:
[0103] The unit for obtaining information on user process nodes to be processed is used to generate information on user process nodes to be processed based on the current user process node information in the current user prompt information if a user process node consultation request is detected, and send the information on user process node to be processed to the receiving terminal selected on the XR device.
[0104] In this embodiment, after the doctor user completes this round of communication with the patient user, if the patient user needs to ask further questions to obtain the next user process node, the current user process node information in the current user prompt information that has been obtained in the doctor user's XR device generates the user process node information to be processed (the user process node information to be processed can prompt the patient user to process the user process nodes that need to be processed in the next few steps, such as blood test, test result collection, return to the clinic to review the blood test result sheet, and wait for the user process node information to be processed), and after the doctor user selects the receiving terminal of the patient user, the doctor user's XR device sends the user process node information to be processed to the receiving terminal. By viewing the user process node information to be processed, the patient user can more clearly understand the subsequent user process nodes and perform corresponding processing.
[0105] The text intelligent processing unit 160 is configured to perform text summarization or automatic summary processing on the current speech recognition result based on a pre-trained natural language processing model to obtain a current text processing result if a text processing instruction is detected.
[0106] In this embodiment, when the doctor user confirms that the entire communication process with the patient user is completed, he or she can operate to select to end the communication, and further select the text processing button on the user interaction interface displayed by the XR device. At this time, the XR device can perform text summary or automatic summary processing on the current speech recognition result based on the natural language processing model to obtain the current text processing result. The obtained current text processing result and the doctor user identity information and the user identity information corresponding to the current user recognition result are packaged and compressed, and uploaded to the server for data archiving. Of course, in order to ensure data security, the current text processing result and the doctor user identity information and the user identity information corresponding to the current user recognition result can also be encrypted, and then the encrypted data can be uploaded to the server for information storage.
[0107] In one embodiment, the user information processing apparatus 100 based on the XR device further includes:
[0108] a current user detection data acquisition unit, configured to acquire current user detection data of a user corresponding to the current user image;
[0109] a current classification result acquisition unit, configured to acquire the current text processing result and the current classification result of the current user detection data based on a pre-trained classification model;
[0110] The classification result related data storage unit is used to store the current text processing result and the current user detection data in a storage space corresponding to the current text processing result and the current user detection data if it is determined that the current classification result belongs to the first preset classification type.
[0111] In this embodiment, after the doctor user synchronously obtains the current user detection data from the server after performing relevant indicator detection (such as blood indicator detection) on the patient user, the current text processing result and the current user detection data can be combined into comprehensive input data, and the comprehensive input data can be input into a classification model (such as a convolutional neural network, etc.) to obtain the current classification result.
[0112] Afterwards, the current classification result can be matched with a first preset classification type. If the current classification result is determined to belong to the first preset classification type, for example, if the patient's blood type is determined to be a rare blood type based on the current user's test data, the current text processing result and the current user's test data can be stored in a storage space corresponding to the current text processing result and the current user's test data. This storage space can then be used to search for users of this type and conduct further user communication.
[0113] In one embodiment, the user information processing apparatus 100 based on the XR device further includes:
[0114] The current schedule prompt information acquiring unit is configured to acquire the corresponding current schedule prompt information if a schedule prompt request is detected, and pre-store the schedule time information and schedule navigation information corresponding to the current schedule prompt information in the corresponding pre-storage space.
[0115] In this embodiment, in addition to daily medical work, the doctor user may also have other schedules such as academic conferences, daily training, etc. When the XR device obtains the schedule reminder information pre-entered by the doctor user (including multiple schedule items, and each schedule item corresponds to an effective time interval), the current schedule reminder information of the day can be extracted in combination with the date. The schedule time information and schedule navigation information corresponding to the current schedule reminder information can be pre-stored in the corresponding pre-storage space, and before the current system time reaches the corresponding schedule effective time interval, the schedule items to be effective and their schedule time information and schedule navigation information are extracted from the pre-storage space and displayed in a timely manner on the display model of the XR device.
[0116] It can be seen that the implementation of the embodiment of the device can enable the user wearing the XR device to perform user image recognition and voice recognition of the communication object in a timely manner, not only to quickly obtain the user information of the communication object, but also to quickly obtain and intuitively display the guidance prompt information required during the communication process.
[0117] The above-mentioned user information processing device based on XR device can be implemented in the form of a computer program. The computer program can be used in Figure 5 Runs on the computer equipment shown.
[0118] See also Figure 5 , Figure 5 This is a schematic block diagram of a computer device provided by an embodiment of the present invention. The computer device integrates any of the user information processing devices based on XR devices provided by an embodiment of the present invention.
[0119] See Figure 5 The computer device 400 includes a processor 402 , a memory, and a network interface 405 connected via a system bus 401 , wherein the memory may include a storage medium 403 and an internal memory 404 .
[0120] The storage medium 403 may store an operating system 4031 and a computer program 4032. The computer program 4032 includes program instructions, which, when executed, may enable the processor 402 to execute a user information processing method based on an XR device.
[0121] The processor 402 is used to provide computing and control capabilities to support the operation of the entire computer device.
[0122] The internal memory 404 provides an environment for the operation of the computer program 4032 in the storage medium 403. When the computer program 4032 is executed by the processor 402, the processor 402 can execute the above-mentioned user information processing method based on the XR device.
[0123] The network interface 405 is used to communicate with other devices through the network. Figure 5 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0124] The processor 402 is configured to run a computer program 4032 stored in the memory to implement the above-mentioned method for processing user information based on an XR device.
[0125] It should be understood that in the embodiment of the present invention, the processor 402 may be a central processing unit (CPU), and the processor 402 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0126] Those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program includes program instructions, which can be stored in a storage medium that is computer-readable. The program instructions are executed by at least one processor in the computer system to implement the steps in the method of the above-described embodiment.
[0127] Therefore, the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the processor executes the user information processing method based on the XR device as described above.
[0128] The storage medium may be any computer-readable storage medium that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.
[0129] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0130] In the several embodiments provided herein, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the various units is merely a logical functional division, and actual implementation may employ other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be omitted or not implemented.
[0131] The steps in the methods of the embodiments of the present invention may be adjusted in order, combined, or deleted as needed. The units in the devices of the embodiments of the present invention may be combined, divided, or deleted as needed. Furthermore, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit.
[0132] If this integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, terminal, or network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present invention.
[0133] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A user information processing method based on an XR device, applied to an XR device, characterized in that: include: In response to a user identification instruction, acquiring a current user image corresponding to the user identification instruction; Recognize the current user image based on the deployed face recognition model to obtain a current user recognition result; If it is determined that the current user identification result belongs to the preset specific user information set, obtaining current user prompt information corresponding to the current user identification result and displaying it on a display module corresponding to the XR device; Acquire user communication voice data collected based on the current user prompt information; Performing speech recognition on the current communication speech data based on the deployed speech recognition model to obtain a current speech recognition result; If a text processing instruction is detected, a text summary or automatic summary processing is performed on the current speech recognition result based on a pre-trained natural language processing model to obtain a current text processing result.
2. The method according to claim 1, characterized in that If it is determined that the current user identification result belongs to the preset specific user information set, obtaining current user prompt information corresponding to the current user identification result includes: If it is determined that the user unique identification code corresponding to the current user identification result is the same as the user unique identification code of one of the specific user information in the specific user information set, then the current user identification result is determined to belong to the preset specific user information set, and the target specific user information corresponding to the current user identification result is obtained from the specific user information set; The desensitized user name, user tag, user's last visit information, disease-specific prompt information, historical chat record summary information, current user process node information and current communication prompt information are obtained from the target specific user information as the current user prompt information.
3. The method according to claim 2, characterized in that Before the step of obtaining the user communication voice data collected based on the current user prompt information, the method further includes: Based on the current communication prompt information in the current user prompt information, a plurality of communication prompt texts are generated in sequence, and the plurality of communication prompt texts are displayed in a display module corresponding to the XR device.
4. The method according to claim 3, characterized in that After the step of performing speech recognition on the current communication speech data based on the deployed speech recognition model to obtain a current speech recognition result, the method further includes: If it is detected that the current speech recognition result belongs to other communication texts that are not included in the multiple communication prompt texts, the other communication texts are input into the locally pre-deployed large language model to obtain the model output prompt text, and the model output prompt text information is displayed in the display module corresponding to the XR device.
5. The method according to claim 2, characterized in that After the step of performing speech recognition on the current communication speech data based on the deployed speech recognition model to obtain a current speech recognition result, the method further includes: If a user process node consultation request is detected, user process node information to be processed is generated based on the current user process node information in the current user prompt information, and the user process node information to be processed is sent to the receiving terminal selected on the XR device.
6. The method according to claim 1, characterized in that After the step of performing text summarization or automatic summarization processing on the current speech recognition result based on a pre-trained natural language processing model to obtain the current text processing result if a text processing instruction is detected, the method further includes: Acquire current user detection data of a user corresponding to the current user image; Obtaining the current text processing result and the current classification result of the current user detection data based on a pre-trained classification model; If it is determined that the current classification result belongs to the first preset classification type, the current text processing result and the current user detection data are stored in a storage space corresponding to the current text processing result and the current user detection data.
7. The method according to claim 1, characterized in that After the step of performing speech recognition on the current communication speech data based on the deployed speech recognition model to obtain a current speech recognition result, the method further includes: If a schedule prompt request is detected, the corresponding current schedule prompt information is obtained, and the schedule time information and schedule navigation information corresponding to the current schedule prompt information are pre-stored in the corresponding pre-storage space.
8. A user information processing device based on an XR device, configured on the XR device, characterized in that: include: A user image acquisition unit, configured to acquire, in response to a user identification instruction, a current user image corresponding to the user identification instruction; A user face recognition unit is used to recognize the current user image based on the deployed face recognition model to obtain a current user recognition result; a current prompt information generating unit, configured to, if it is determined that the current user recognition result belongs to a preset specific user information set, obtain current user prompt information corresponding to the current user recognition result and display it on a display module corresponding to the XR device; A voice data collection unit, configured to obtain user communication voice data collected based on the current user prompt information; A speech recognition unit, configured to perform speech recognition on the current communication speech data based on a deployed speech recognition model to obtain a current speech recognition result; The text intelligent processing unit is used to perform text summarization or automatic summary processing on the current speech recognition result based on a pre-trained natural language processing model if a text processing instruction is detected to obtain the current text processing result.
9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the user information processing method based on the XR device according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which includes program instructions. When the program instructions are executed by a processor, the method for processing user information based on an XR device according to any one of claims 1 to 7 can be implemented.