A cross-platform medical registration interaction method based on AI intelligent agents and multimodal recognition

By using AI intelligent agents to simulate human operation and multimodal recognition technology, the incompatibility issue between the online registration system and the HIS system was resolved, enabling flexible, secure, and low-cost cross-platform registration system integration, thus improving registration efficiency and system stability.

CN120412144BActive Publication Date: 2025-12-02SHAANXI LIXING ZHITONG TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510504467.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-12-02
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

The existing online registration system for hospitals suffers from incompatible interfaces, difficulties in upgrading, and high costs in data interaction with the HIS system, lacking flexibility and universal adaptability.

Method used

It employs an AI agent to simulate human operation, and uses image recognition and control tree recognition technologies to achieve cross-platform medical registration interaction, avoiding reliance on the standardized interface of the HIS system. It combines a deep learning framework and a multimodal recognition model to adapt to the interface differences of different HIS systems.

Benefits of technology

It improved the versatility and applicability of the online registration system, reduced system integration costs, enhanced the security and flexibility of data interaction, reduced integration time and manpower investment, and ensured the stability and reliability of the HIS system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412144B_ABST
    Figure CN120412144B_ABST
Patent Text Reader

Abstract

This application relates to the field of medical information system technology and discloses a cross-platform medical registration interaction method based on AI intelligent agents and multimodal recognition. Applied to an AI intelligent agent deployed on an intranet server, the method includes: accessing a message queue built on a front-end machine via the intranet to obtain registration request data sent by the online registration system; recognizing the login interface of the HIS terminal installed on the same intranet server, simulating manual operation by entering the pre-configured login account and password into the corresponding input boxes, and clicking the login button to complete the login; locating the registration information page, simulating manual operation by entering the registration request data into the corresponding input boxes or selection boxes; recognizing the registration confirmation interface, simulating manual operation to complete the registration; parsing the registration result information from the registration result display interface, sending it back to the message queue, and then having the message queue feed back to the online registration system. This method enhances the versatility and universality of the online registration system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this application relate to the field of medical information system technology, and in particular to a cross-platform medical registration interaction method based on AI intelligent agents and multimodal recognition. Background Technology

[0002] Currently, hospital online registration systems typically interact with the Hospital Information System (HIS) through standardized interfaces. However, this data interaction method has certain limitations. On the one hand, some older HIS systems may not provide standardized interfaces, making it impossible for online registration systems to connect with them. On the other hand, even if the HIS system has standardized interfaces, problems may arise in interface updates, maintenance, or compatibility with different systems, thus affecting the normal operation of registration services.

[0003] Traditional data interaction methods rely on predefined interface specifications between systems, lacking flexibility and universal adaptability. When a HIS system is upgraded, modified, or replaced, the online registration system may need to redevelop its interface to interact with the HIS system, which is costly and time-consuming. Therefore, there is an urgent need for a more universal and flexible data interaction method to meet the diverse needs of different hospital HIS systems. Summary of the Invention

[0004] In view of this, embodiments of this application propose a cross-platform medical registration interaction method based on AI intelligent agents and multimodal recognition. It does not rely on the standardized interface of the HIS system, but uses AI intelligent agents to simulate human operation to achieve effective data interaction between the online registration system and the HIS system. The data interaction process is highly secure, reduces the cost of system integration, and enhances the versatility and universal applicability of the online registration system.

[0005] To achieve the above objectives, embodiments of this application propose a cross-platform medical registration interaction method based on AI agents and multimodal recognition, applied to an AI agent deployed on an intranet server. The method includes the following steps: accessing a message queue built on a front-end machine via the intranet to obtain registration request data sent by the online registration system; wherein, the registration request data includes at least patient information, the registration department, and the registration time; recognizing the login interface of the HIS system on the HIS terminal installed on the same intranet server, locating the positions of each input box and login button, simulating manual operation of the mouse and keyboard, and inputting the pre-configured login account and password into the corresponding input boxes. The user clicks the login button to log in to the HIS system; identifies the information entry interface of the HIS system, locates the registration information page, positions each input box and selection box, and simulates manual operation of the mouse and keyboard to enter the registration request data into the corresponding input boxes or selection boxes; identifies the registration confirmation interface of the HIS system, locates the registration confirmation button, simulates manual operation of the mouse to click the registration confirmation button, triggers the registration process, and completes the registration; identifies the registration result display interface of the HIS system, parses the registration result information, sends the registration result information back to the message queue, and then the message queue feeds back the registration result information to the online registration system.

[0006] To achieve the above objectives, embodiments of this application also propose a cross-platform medical registration interaction system based on AI intelligent agents and multimodal recognition, used for data interaction with the HIS system. The system includes: a front-end interaction module deployed on the user end, a message queue module deployed on a front-end server, and an AI intelligent agent and an HIS terminal deployed on an intranet server. The front-end interaction module allows users to enter registration request data through the online registration system and send it to the message queue module, as well as receive registration result information from the message queue module. The message queue module receives and stores the number of registration requests. The system provides access to and reads information from the AI ​​agent, and receives registration result information from the AI ​​agent, then feeds it back to the front-end interaction module. The HIS terminal, or HIS system, is a user-operated terminal program or page that provides an interface for registration services, including a login interface, information entry interface, registration confirmation interface, and registration result interface. The AI ​​agent accesses the message queue module via the intranet to read registration request data, simulates manual operation to log in to the HIS terminal, enter registration request data, execute registration operations, parse registration result information, and send the registration result information back to the message queue module.

[0007] To achieve the above objectives, embodiments of this application also propose an AI agent. This AI agent is equipped with at least the deep learning framework TensorFlow and the YOLO algorithm model, and integrates an image recognition module, a control tree plugin, and scripts for simulating mouse and keyboard operations. It can accurately locate elements on various pages of the HIS system, including input boxes, selection boxes, and buttons, and also possesses the ability to simulate operation execution. The AI ​​agent can access the message queue module via the intranet to read registration request data, simulate manual keyboard and mouse operation to log in to the HIS system, enter registration request data, execute registration operations, parse registration result information, and send the registration result information back to the message queue module. The AI ​​agent also possesses dynamic decision-making capabilities, and can adjust the operation process according to interface changes.

[0008] In some optional embodiments, the AI ​​agent uses image recognition technology and process monitoring technology to detect the running status of the HIS terminal program on the intranet server. It prioritizes using control tree recognition technology to log in and enter registration information to obtain registration results. If the HIS system does not support control tree recognition technology, it triggers image recognition technology to locate elements including input boxes, selection boxes, and buttons. Then, it uses a script to control the mouse and keyboard to simulate manual login and entry of registration information to complete the registration task.

[0009] In some optional embodiments, image recognition technology is implemented based on an image recognition model. The AI ​​agent pre-trains, evaluates, and optimizes the image recognition model through the following steps: collecting a large number of screenshots of login interfaces, information entry interfaces, registration confirmation interfaces, and registration result display interfaces of different HIS systems, covering different resolutions, interface layouts, and color styles; labeling key elements, including input boxes, selection boxes, and buttons, and clarifying the location and category of key elements; performing data augmentation on each screenshot, including random cropping and scaling, to construct training sample sets, test sample sets, and validation sample sets; and selecting a deep learning framework and image recognition model; wherein, the selected deep learning framework includes at least PyTorch and MX. The system supports .NET, JAX, PaddlePaddle, and TensorFlow. Selectable image recognition models include at least YOLO, SSD, NanoDet, and DETR models. Based on the training sample set and the selected deep learning framework, the chosen image recognition model is iteratively trained until convergence, according to the set learning rate, number of iterations, batch size, and momentum. The trained image recognition model is then tested on a test sample set. Based on the validation sample set, the image recognition model that passes the test is evaluated for accuracy, recall, and F1 score. The model parameters are adjusted based on the evaluation results until all three metrics meet the preset standards. The image recognition model is pre-trained using screenshots of numerous HIS system interfaces, making it more suitable for interface recognition and localization within HIS systems.

[0010] In some optional embodiments, the AI ​​agent pre-configures control tree recognition rules through the following steps: HTML DOM-based rule configuration, applicable to the interface of a HIS system built on Web technology. This HTML DOM-based rule configuration includes: element location rules, locating elements using their tag name, ID, class name, and attribute values; if an element does not have an explicit ID or class name, it is located based on its hierarchical relationship in the DOM tree and the characteristics of adjacent elements; attribute matching rules, determining the function of special elements, including buttons, by matching attribute values; ignoring case differences in text; dynamic element processing rules, using an event listener mechanism for dynamically loaded elements, re-parses the DOM tree to locate newly appearing elements when specific events, including clicks and scrolling, occur on the page; and Windows API-based rule configuration, applicable to the interface of a HIS system built on Windows desktop application technology. This Windows API-based rule configuration includes: window handle acquisition, using Windows API functions to obtain the handle of windows in the HIS system, searching by window title and class name; and control enumeration, using Enum Child... Windows functions enumerate all controls within a window, obtaining the handle and basic information of each control; control properties are retrieved using the Get Window Text or Get Class Name functions to obtain the text content and class name attribute values ​​of the control, determining the control's type and function based on the attribute values; control operations are performed using the corresponding Windows API functions for different types of controls. The control tree identification configuration includes rules applicable to interfaces of HIS systems built with different technologies, further enhancing the universality of the AI ​​agent, enabling it to adapt to various complex HIS systems and improving its versatility.

[0011] In some optional embodiments, before each element recognition, the AI ​​agent first obtains a screenshot of the current interface and compares it with a pre-saved screenshot of a standard interface. It calculates the positional difference between corresponding elements in the two screenshots. If the positional difference exceeds a preset threshold, the element is determined to have shifted. For shifted elements, the AI ​​agent needs to calculate the shift ratio in both the horizontal and vertical directions. The formula for calculating the shift ratio is: p x = (x2-x1) / x1; p y = (y2-y1) / y1; where (x1,y1) represents the coordinates of the top-left corner of an element in the screenshot of the standard interface, (x2,y2) represents the coordinates of the top-left corner of an element in the screenshot of the current interface, and p x p represents the horizontal offset ratio. yThis represents the offset ratio in the vertical direction; based on the calculated offset ratios in the horizontal and vertical directions, the position of the element to be identified is adjusted using the following formula: Where (x0, y0) represents the click position of the element to be identified before adjustment. This indicates the adjusted click position of the element to be identified. During the identification process, it is not uncommon for element positions in the HIS system interface to shift. Without adjustment, accurate identification and subsequent operations cannot be performed. Therefore, the AI ​​agent incorporates an element offset detection and correction mechanism, which can reposition elements through image recognition and offset calculation, thereby improving efficiency and accuracy in the interaction process.

[0012] In some optional embodiments, the AI ​​agent is pre-written using Python's PyAutoGUI library to simulate human mouse and keyboard operations, including login, information entry, and registration, based on the HIS system's operating procedures. The script sets a randomized range for the intervals between these simulated mouse and keyboard operations, randomly selecting intervals between 0.5 and 1 second to avoid detection of automated operations by the HIS system. This randomized operation interval design simulates real human operation as closely as possible, effectively reducing the likelihood of the HIS system detecting automated operations.

[0013] In some optional embodiments, the message queue is pre-built on the front-end machine as a message queue module. RabbitMQ message queue software is selected for installation and configuration to ensure stable reception of registration request data sent by the online registration system and to ensure data interaction with the AI ​​agent deployed on the intranet server. Firewall rules are configured to allow only specific IP addresses to access the message queue, and HTTPS encrypted transmission is enabled. The AI ​​agent and HIS terminal are pre-deployed on the intranet server to ensure stable operation of the AI ​​agent and HIS terminal. The AI ​​agent is initialized and configured, including loading the image recognition model and setting control tree recognition rules. The front-end interaction module of the online registration system adopts the form of a web page or mobile application to interact with the message queue module on the front-end machine through the network.

[0014] The embodiments of this application propose a cross-platform medical registration interaction method based on AI intelligent agents and multimodal recognition, which has at least the following advantages compared with existing methods that simulate manual operation of HIS systems.

[0015] First, it has strong versatility. This application does not rely on the standardized interface of the HIS system, and can also achieve effective integration with HIS systems that do not provide standardized interfaces or whose interfaces are incompatible, greatly improving the applicability of the online registration system. By combining multimodal interaction methods of image recognition and control tree recognition, it can adapt to the interface differences of HIS systems developed by different manufacturers.

[0016] Secondly, it offers high flexibility. When the HIS system undergoes interface updates or upgrades, the AI ​​agent can quickly adapt to the new interface layout by retraining the image recognition model or adjusting the control tree recognition rules, without requiring large-scale modifications to the entire system. Its dynamic decision-making capability can adjust the operation process in real time according to interface changes, ensuring the stable operation of the system.

[0017] Third, it offers high security. The message queue is hosted on a front-end server, while the AI ​​agent and HIS terminal are installed on an intranet server. The AI ​​agent accesses data on the front-end server via the intranet, effectively preventing direct access from the external network to the hospital's intranet HIS system and reducing the risk of data leakage and attacks. Simultaneously, HTTPS is used to encrypt data transmission, and firewall rules and IP whitelists are implemented to ensure the security of the hospital's information system. Furthermore, the randomization of the AI ​​agent's simulated operations and the simulation of behavioral patterns reduce the risk of detection by the HIS system's anti-scraping strategies.

[0018] Fourth, low cost. This application avoids the high costs of developing and maintaining standardized interfaces for different HIS systems, reducing the time and manpower required for system integration.

[0019] Fifth, it can simulate real human operation. The AI ​​agent interacts with the HIS system by simulating human operation of a mouse and keyboard, without affecting the original architecture and data security of the HIS system, thus ensuring the stability and reliability of the HIS system. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies of this application will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart of a cross-platform medical registration interaction method based on AI intelligent agent and multimodal recognition provided in one embodiment of this application;

[0022] Figure 2This is a schematic diagram illustrating the information flow when an AI agent interacts with an HIS system, as provided in one embodiment of this application.

[0023] Figure 3 This is a schematic diagram of the structure of a cross-platform medical registration interaction system based on AI intelligent agents and multimodal recognition, provided in another embodiment of this application. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the various embodiments of this application will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details have been provided in the various embodiments of this application to help readers better understand this application. However, the technical solutions claimed in this application can be implemented even without these technical details and various changes and modifications based on the following embodiments. The division of the various embodiments below is for the convenience of description and should not constitute any limitation on the specific implementation of this application. The various embodiments can be combined with and referenced by each other without contradiction.

[0025] One embodiment of this application proposes a cross-platform medical registration interaction method based on AI intelligent agents and multimodal recognition, which is applied to AI intelligent agents deployed on an intranet server. The implementation details of the cross-platform medical registration interaction method based on AI intelligent agents and multimodal recognition proposed in this embodiment are described in detail below. The following implementation details are provided for ease of understanding and are not necessary for implementing this solution.

[0026] The specific process of the cross-platform medical registration interaction method based on AI intelligent agent and multimodal recognition proposed in this embodiment can be described as follows: Figure 1 As shown, it includes:

[0027] Step 101: Access the message queue built on the front-end machine via the intranet to obtain the registration request data sent by the online registration system.

[0028] This section first introduces the deployment of the AI ​​agent and the various modules and components that interact with it. The AI ​​agent and HIS terminal are pre-deployed (installed) on the same intranet server. It is necessary to ensure that the intranet server has good performance and security, and can stably run the AI ​​agent and HIS terminal. After the AI ​​agent is deployed, it needs to be initialized and configured, including loading the image recognition model and setting the control tree recognition rules.

[0029] The message queue is pre-built on the front-end machine as a message queue module. RabbitMQ message queue software is selected for installation and configuration (RocketMQ, ActiveMQ, RabbitMQ, Kafka, and other message queue software can also be hovered over) to ensure that it can stably receive registration request data sent by the online registration system and ensure that it can interact with the AI ​​agent deployed on the intranet server. At the same time, firewall rules are configured to allow only specific IP addresses (such as the IP of the intranet server and the IP of the front-end server of the online registration system) to access the message queue, and HTTPS encrypted transmission is enabled.

[0030] The front-end interaction module of the online registration system uses a web page or mobile application (such as H5, APP, mini-program, etc.) to interact with the message queue module on the front-end server via the network. The front-end interaction module needs to be configured accordingly to ensure that it can correctly send registration request data and receive registration result information.

[0031] In the implementation, when a user needs to register for an appointment, they can submit a registration request through the front-end interaction module (i.e., the online registration system). The front-end interaction module sends the registration request data to the message queue module on the front-end server. The sent registration request data is encrypted to ensure its security during transmission. The AI ​​agent accesses the message queue module on the front-end server via the intranet to obtain the registration request data sent by the online registration system. This registration request data includes, but is not limited to, patient information, the department to be registered for, and the registration time.

[0032] Step 102: Identify the login interface of the HIS system on the HIS terminal installed on the same intranet server, locate the positions of each input box and login button, simulate manual operation of mouse and keyboard, enter the pre-configured login account and password into the corresponding input boxes, and click the login button to complete the login to the HIS system.

[0033] In the specific implementation, the system first determines whether the HIS system is open. The AI ​​agent uses image recognition and process monitoring to detect the running status of the HIS terminal program on the intranet server. The AI ​​agent monitors the HIS terminal's operation by comparing a pre-stored screenshot of a prominent area of ​​the running HIS terminal with a screenshot of the current interface to determine if it is running normally. If the HIS terminal program is not running, image recognition is used to locate the HIS terminal icon, and the script controls the mouse to double-click the icon to open the HIS terminal. During the login phase, the AI ​​agent uses a control tree traversal algorithm to locate the username and password input boxes and uses the SendMessage function to simulate keyboard input to complete authentication. If the control tree structure is unavailable due to HIS system limitations, an image matching algorithm is triggered. A pixel-level comparison is performed between a pre-stored login interface screenshot and the current interface to locate the input area. After that, the script controls the input area to simulate manual mouse and keyboard operation, inputting the pre-configured login account and password into the corresponding input boxes and clicking the login button to complete the HIS system login.

[0034] In one example, the AI ​​agent is pre-configured with two tools: an image recognition model and control tree recognition rules, to complete the recognition task of various interfaces of the HIS system. When recognizing the login interface of the HIS system, the AI ​​agent first uses image recognition technology to locate the positions of each input box and the login button, and then uses control tree recognition technology to further confirm the positions of the located input boxes and the login button. The control tree recognition technology employs a combination of HTML DOM parsing and the Windows API.

[0035] The AI ​​agent combines multimodal interaction methods such as image recognition and control tree recognition, and can adjust the operation process in real time according to the interface changes. It also adopts a dynamic decision-making mechanism, which greatly improves the adaptability and stability of the system.

[0036] In one example, if the AI ​​agent encounters an abnormality in the recognition of interface elements or a login failure during the process of logging into the HIS terminal, it can retry according to the preset retry mechanism and record the abnormal information for subsequent analysis.

[0037] In one example, image recognition technology is based on an image recognition model, and the AI ​​agent needs to train, evaluate, and optimize the image recognition model in advance.

[0038] First, the AI ​​agent will collect a large number of screenshots of login interfaces, information entry interfaces, registration confirmation interfaces, and registration result display interfaces of different HIS systems, covering different resolutions, interface layouts, and color styles. It will also label key elements, including input boxes, selection boxes, and buttons, to clarify the location and category of the elements.

[0039] Next, the AI ​​agent needs to perform data augmentation on each screenshot, including random cropping and scaling, to construct training, testing, and validation sample sets. Data augmentation can significantly improve the model's generalization ability. Specifically, random cropping involves randomly selecting a region from the original screenshot, with the cropping ratio randomly chosen between 0.7 and 0.9, and the scaling ratio randomly adjusted between 0.8 and 1.2.

[0040] After constructing the sample set, the AI ​​agent needs to select a deep learning framework and an image recognition model. The selectable deep learning frameworks include at least PyTorch, MXNet, JAX, PaddlePaddle, and TensorFlow, while the selectable image recognition models include at least YOLO, SSD, NanoDet, and DETR. This embodiment selects the TensorFlow deep learning framework to train the YOLO model.

[0041] After selecting a deep learning framework and image recognition model, the AI ​​agent can iteratively train the chosen image recognition model until convergence based on the training sample set and the selected deep learning framework, according to the set learning rate, number of iterations, batch size, and momentum. The performance of the trained image recognition model is then tested based on the test sample set. The initial learning rate is set to 0.001, employing an exponential decay strategy, decreasing every 100 iterations at a decay rate of 0.9. The number of iterations is set to 500 to ensure the model fully learns the data features. The batch size is set to 16, meaning 16 images are input for each training iteration. The momentum is set to 0.9, a setting that helps accelerate model convergence.

[0042] After completing model training, the AI ​​agent also needs to evaluate the accuracy, recall, and F1 score of the tested image recognition model based on the validation sample set, and adjust the model parameters of the image recognition model based on the evaluation results until the accuracy, recall, and F1 score all meet the preset standards. The image recognition model that meets the standards is then deployed into the AI ​​agent.

[0043] The image recognition model is pre-trained based on a large number of screenshots of different interfaces of the HIS system, making it more suitable for the recognition and localization of HIS system interfaces.

[0044] In one example, the AI ​​agent needs to pre-configure the control tree recognition rules.

[0045] The first step is to configure rules based on HTML DOM parsing. This rule configuration is applicable to the interface of the HIS system built on Web technology and includes three parts: element location rules, attribute matching rules, and dynamic element processing rules.

[0046] Element location rules utilize information such as the element's tag name, ID, class name, and attribute values ​​for location. If an element does not have an explicit ID or class name, it is located based on its hierarchical relationship in the DOM tree and the characteristics of adjacent elements. For example, for an input box element, a precise search can be performed using the tag name "input" combined with its ID or class name.

[0047] The attribute matching rules determine the function of special elements, including buttons (e.g., the "value" attribute of a button represents the text displayed on the button), by matching attribute values. Case sensitivity and fuzzy matching are considered; for text with differences in capitalization, case-insensitive matching is performed.

[0048] For dynamically loaded elements (such as dropdown options dynamically generated by JavaScript), an event listener mechanism is required. When specific events such as clicks and scrolling occur on the page, the DOM tree needs to be re-parsed to locate the newly appeared element.

[0049] The next step is to configure rules based on the Windows API. This rule configuration applies to the interface of the HIS system built on Windows desktop application technology. It includes four parts: window handle acquisition, control enumeration, control property acquisition, and control operation.

[0050] Window handle acquisition: Use Windows API functions (such as FindWindow, FindWindowEx, etc.) to obtain the handle of a window in the HIS system, and search for it by window title or class name.

[0051] Control enumeration: Use the Enum Child Windows function to enumerate all controls within the window and obtain the handle and basic information of each control.

[0052] To obtain control properties, use the Get Window Text function or the Get Class Name function to get the text content and class name property values ​​of the control, and determine the type and function of the control based on the property values.

[0053] For control manipulation, use the corresponding Windows API functions for different types of controls. For example, use the SendMessage function to send a click message for buttons, and use the SendMessage function to input text for input fields.

[0054] The control tree identifies and configures rules for interfaces of HIS systems built with different technologies, further enhancing the universality of AI agents and enabling them to adapt to various complex HIS systems, thus improving their versatility.

[0055] In one example, before each element recognition, the AI ​​agent needs to obtain a screenshot of the current interface and compare it with a pre-saved screenshot of the standard interface. It calculates the positional difference between corresponding elements in the two screenshots. If the positional difference exceeds a preset threshold (such as 5 pixels), it is determined that the element has shifted.

[0056] For elements that have shifted, the AI ​​agent calculates the percentage shift in the horizontal direction and the percentage shift in the vertical direction. The formula for calculating the percentage shift is:

[0057] p x = (x2-x1) / x1;

[0058] p y = (y2-y1) / y1;

[0059] Where (x1, y1) represents the coordinates of the top-left corner of an element in the screenshot of the standard interface, (x2, y2) represents the coordinates of the top-left corner of an element in the screenshot of the current interface, and p x p represents the horizontal offset ratio. y This indicates the percentage of offset in the vertical direction.

[0060] Next, the AI ​​agent adjusts the position of the element to be located based on the calculated offset ratios in the horizontal and vertical directions. The position adjustment formula is as follows:

[0061]

[0062] Where (x0, y0) represents the click position of the element to be positioned before adjustment. This indicates the clickable position of the element that needs to be positioned.

[0063] During the recognition process, it is not uncommon for the position of elements in the HIS system interface to shift. Without adjustment, accurate recognition and subsequent operations cannot be performed. Based on this, the AI ​​agent has a built-in element shift detection and correction mechanism, which can reposition the element through image recognition and shift calculation, thereby improving the efficiency and accuracy of the interaction process.

[0064] In one example, the AI ​​agent simulates human mouse and keyboard operations via script. The AI ​​agent is pre-written using Python's PyAutoGUI library, based on the HIS system's operational procedures, to simulate human mouse and keyboard operations, including login, information entry, and registration. The script sets a randomized range for the intervals between these simulated operations, randomly selecting intervals between 0.5 and 1 second to avoid detection of automated operations by the HIS system.

[0065] The design of random operation intervals can simulate real human operation as much as possible, thereby effectively reducing the chance of automated operation being detected by the HIS system.

[0066] Step 103: Identify the information entry interface of the HIS system, locate the registration information page, pinpoint the positions of each input box and selection box, simulate manual operation of the mouse and keyboard, and enter the registration request data into the corresponding input box or selection box.

[0067] In its implementation, after logging into the HIS terminal, the AI ​​agent prioritizes parsing the hierarchical structure of the registration menu (e.g., "Outpatient Management → Registration") using a control tree, then expands the menu nodes level by level and activates the target page. When entering registration data, the AI ​​agent parses the control tree structure of the registration page, using a triple positioning method based on control type (e.g., Edit), window class name (e.g., THOspinalEdit), and position coordinates to accurately locate input boxes for patient name, ID number, and registration department. It then uses the SetDlgItemText function to directly write the data, avoiding the latency issues associated with keyboard simulation. When control tree parsing fails, the image recognition module is activated to identify the registration information page, locate the positions of each input box and selection box, and simulate manual mouse and keyboard operation to enter the registration request data into the corresponding input or selection boxes.

[0068] In one example, during the data entry process for a registration request, if a drop-down selection box is encountered, the AI ​​agent will simulate clicking the drop-down selection box with a mouse, and then locate the target option and make a selection based on image recognition or control tree recognition.

[0069] Step 104: Identify the registration confirmation interface of the HIS system, locate the registration confirmation button, simulate manual operation by clicking the registration confirmation button with a mouse, trigger the registration process, and complete the registration.

[0070] In practice, after the AI ​​agent completes the entry of the registration request data, it needs to identify the registration confirmation interface of the HIS system, locate the registration confirmation button, simulate manual operation by clicking the registration confirmation button, trigger the registration process, and complete the registration.

[0071] In one example, for some HIS systems, the information entry interface and the registration confirmation interface are the same. In this case, the AI ​​agent does not need to recognize repeatedly. It only needs to simulate human operation by clicking the registration confirmation button after completing the entry of the registration request data.

[0072] In one example, when the AI ​​agent is recognizing the registration confirmation interface of the HIS system, it can use either image recognition technology or control tree recognition technology.

[0073] Step 105: Locate the registration result display interface of the HIS system, parse out the registration result information, send the registration result information back to the message queue, and then the message queue feeds back the registration result information to the online registration system.

[0074] In practice, after the AI ​​agent completes the registration, it needs to identify the registration result display interface of the HIS system, parse the registration result information, send the registration result information back to the message queue, and then the message queue feeds back the registration result information to the online registration system. The online registration system displays the registration result information to the user through the front-end interaction module.

[0075] In one example, for an interface built using a control tree, the AI ​​agent can read the attribute values ​​of relevant controls to obtain the registration result. For instance, if a label control's text attribute stores the information "Registration successful," the agent can obtain the result by parsing this control's attribute. Some controls have different states that reflect the registration result. For example, a button control changes to the "Registered" state when registration is successful; the AI ​​agent can determine the result by judging the control's state.

[0076] In one example, if the HIS system displays registration results in text format on the interface, the AI ​​agent can use image recognition technology to convert the text information on the interface into computer-processable text. For instance, when the result displays "Registration successful, order number 123456," OCR can accurately identify the key information. To improve recognition accuracy, the image can be preprocessed, such as grayscale conversion, noise reduction, and binarization. If the registration results are presented using image features such as icons or colors, the AI ​​agent can analyze these features to determine the result. For example, a successful registration is represented by a green checkmark icon, and a failed registration by a red cross icon; the AI ​​agent determines the result by recognizing the shape, color, and other features of the icons.

[0077] In one example, the AI ​​agent needs to verify the registration results multiple times during the process of parsing the registration information to improve accuracy.

[0078] In one example, the information flow when the AI ​​agent interacts with the HIS system is as follows: Figure 2 As shown.

[0079] The cross-platform medical registration interaction method proposed in this embodiment, based on AI intelligent agents and multimodal recognition, has at least the following advantages compared with existing methods that simulate manual operation of HIS systems.

[0080] First, it has strong versatility. This embodiment does not rely on the standardized interface of the HIS system, and can also achieve effective integration with HIS systems that do not provide a standardized interface or whose interfaces are incompatible, greatly improving the applicability of the online registration system. By combining multimodal interaction methods of image recognition and control tree recognition, it can adapt to the interface differences of HIS systems developed by different manufacturers.

[0081] Secondly, it offers high flexibility. When the HIS system undergoes interface updates or upgrades, the AI ​​agent can quickly adapt to the new interface layout by retraining the image recognition model or adjusting the control tree recognition rules, without requiring large-scale modifications to the entire system. Its dynamic decision-making capability can adjust the operation process in real time according to interface changes, ensuring the stable operation of the system.

[0082] Third, it offers high security. The message queue is hosted on a front-end server, while the AI ​​agent and HIS terminal are installed on an intranet server. The AI ​​agent accesses data on the front-end server via the intranet, effectively preventing direct access from the external network to the hospital's intranet HIS system and reducing the risk of data leakage and attacks. Simultaneously, HTTPS is used to encrypt data transmission, and firewall rules and IP whitelists are implemented to ensure the security of the hospital's information system. Furthermore, the randomization of the AI ​​agent's simulated operations and the simulation of behavioral patterns reduce the risk of detection by the HIS system's anti-scraping strategies.

[0083] Fourth, low cost. This embodiment avoids the high costs of developing and maintaining standardized interfaces for different HIS systems, reducing the time and manpower required for system integration.

[0084] Fifth, it can simulate real human operation. The AI ​​agent interacts with the HIS system by simulating human operation of a mouse and keyboard, without affecting the original architecture and data security of the HIS system, thus ensuring the stability and reliability of the HIS system.

[0085] The steps of the various methods described above are only for clarity. In implementation, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are all within the protection scope of this application. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, but without changing the core design of the algorithm and process, are also within the protection scope of this application.

[0086] Another embodiment of this application proposes a cross-platform medical registration interaction system based on AI intelligent agents and multimodal recognition, which is used to interact with the HIS system. The implementation details of the cross-platform medical registration interaction system based on AI intelligent agents and multimodal recognition proposed in this embodiment are described in detail below. The following implementation details are provided for ease of understanding and are not necessary for implementing this solution.

[0087] The structural diagram of the cross-platform medical registration interaction system based on AI intelligent agent and multimodal recognition proposed in this embodiment can be seen as follows: Figure 3 As shown, it includes: a front-end interaction module 301 deployed on the user terminal 21, a message queue module 302 deployed on the front-end machine 22, and an AI intelligent agent 303 and an HIS terminal 304 deployed on the intranet server 23.

[0088] The front-end interaction module 301 is used for users to enter registration request data through the online registration system and send it to the message queue module 302, as well as to receive registration result information from the message queue module 302.

[0089] The message queue module 302 is used to receive and store registration request data for the AI ​​agent 303 to access and read, and to receive registration result information from the AI ​​agent 303 and feed it back to the front-end interaction module 301.

[0090] HIS terminal 304 refers to the user-operated terminal program or page of the HIS system, which provides an operation interface for registration services, including a login interface, information entry interface, registration confirmation interface, and registration result interface.

[0091] AI agent 303 is used to access message queue module 302 via intranet to read registration request data, simulate manual operation to log in to HIS terminal 304, enter registration request data, execute registration operation, parse registration result information, and send the registration result information back to message queue module 302.

[0092] It is worth mentioning that all modules and units involved in this embodiment are logical modules. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. Furthermore, to highlight the innovative aspects of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed in this application; however, this does not mean that other units do not exist in this embodiment.

[0093] It is not difficult to see that this embodiment is a system embodiment corresponding to the above method embodiment. This embodiment can be implemented in conjunction with the above method embodiment. The relevant technical details and technical effects mentioned in the above method embodiment are still effective in this embodiment. In order to reduce repetition, they will not be repeated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above method embodiment.

[0094] In another embodiment, we conducted performance tests on the cross-platform medical registration interaction method based on AI intelligent agents and multimodal recognition in a laboratory environment. Test results showed that the registration success rate of this application reached over 98%, and the average response time was reduced from 150 seconds for traditional manual operation to approximately 30 seconds. Tests were conducted on HIS system interfaces with different resolutions and layouts, and the accuracy rates for image recognition and control tree recognition both reached over 95%. By deploying the method proposed in this application across various systems, effective integration between the online registration system and the HIS system was achieved, improving registration efficiency, reducing patient waiting time, and simultaneously reducing the hospital's system integration costs and maintenance workload. The hospital reported that the system operated stably, and no HIS system anomalies occurred due to automated operation.

[0095] Another embodiment of this application proposes an AI agent, which is equipped with at least the deep learning framework TensorFlow and the YOLO algorithm model, and integrates an image recognition module, a control tree plugin, and a script for simulating mouse and keyboard operations. It can accurately locate elements on each page of the HIS system, including input boxes, selection boxes, and buttons, and has the ability to simulate operation execution. The AI ​​agent can read registration request data through the intranet access message queue module, simulate manual operation of keyboard and mouse to log in to the HIS system, enter registration request data, execute registration operations, parse registration result information, and send the registration result information back to the message queue module. The AI ​​agent also has dynamic decision-making capabilities and can adjust the operation process according to interface changes.

[0096] It will be understood by those skilled in the art that the above embodiments are specific embodiments for implementing this application, and in practical applications, various changes in form and detail may be made without departing from the spirit and scope of this application. For those skilled in the art, several improvements and modifications can be made without departing from the principles of this application, and these improvements and modifications are also considered to be within the scope of protection of this application.

Claims

1. A cross-platform medical appointment interaction method based on AI intelligent agents and multimodal recognition, applied to AI intelligent agents deployed on intranet servers, characterized in that, include: Access the message queue built on the front-end machine via the intranet to obtain registration request data sent by the online registration system; the registration request data includes at least patient information, registration department, and registration time; The login interface of the HIS system on the HIS terminal installed on the same intranet server is identified, the positions of each input box and login button are located, the mouse and keyboard are simulated to input the pre-configured login account and password into the corresponding input boxes, and the login button is clicked to complete the login to the HIS system. Identify the information entry interface of the HIS system, locate the registration information page, pinpoint the positions of each input box and selection box, simulate manual operation of mouse and keyboard, and enter the registration request data into the corresponding input box or selection box. The registration confirmation interface of the HIS system is identified, the location of the registration confirmation button is located, and the manual operation of clicking the registration confirmation button is simulated to trigger the registration process and complete the registration. The system identifies the registration result display interface of the HIS system, parses the registration result information, sends the registration result information back to the message queue, and then the message queue feeds back the registration result information to the online registration system.

2. The cross-platform medical registration interaction method based on AI intelligent agent and multimodal recognition according to claim 1, characterized in that, The AI ​​agent uses image recognition and process monitoring technologies to detect the running status of HIS terminal programs on the intranet server, and prioritizes the use of control tree recognition technology to log in and enter registration information to obtain registration results. If the HIS system does not support control tree recognition technology, image recognition technology is triggered. The image recognition technology identifies elements including input boxes, selection boxes, and buttons. Then, the script controls the mouse and keyboard to simulate manual login and entry of registration information to complete the registration task.

3. The cross-platform medical registration interaction method based on AI intelligent agent and multimodal recognition according to claim 2, characterized in that, Image recognition technology is based on image recognition models. The AI ​​agent trains, evaluates, and optimizes the image recognition model in advance through the following steps: We collected a large number of screenshots of login interfaces, information entry interfaces, registration confirmation interfaces, and registration result display interfaces from various HIS systems, covering different resolutions, interface layouts, and color styles. We labeled key elements, including input boxes, selection boxes, and buttons, and clarified the location and category of key elements. Data augmentation, including random cropping and scaling, is performed on each screenshot to construct a training sample set, a test sample set, and a validation sample set. Choose a deep learning framework and an image recognition model; the deep learning frameworks that can be selected include at least PyTorch, MXNet, JAX, PaddlePaddle, and TensorFlow, and the image recognition models that can be selected include at least YOLO, SSD, NanoDet, and DETR. Based on the training sample set and the selected deep learning framework, the selected image recognition model is iteratively trained until convergence according to the set learning rate, number of iterations, batch size and momentum, and the performance of the trained image recognition model is tested based on the test sample set. Based on the validation sample set, the accuracy, recall, and F1 score of the tested image recognition model are evaluated. The model parameters of the image recognition model are then adjusted based on the evaluation results until the accuracy, recall, and F1 score all meet the preset standards.

4. The cross-platform medical registration interaction method based on AI intelligent agent and multimodal recognition according to claim 2, characterized in that, The AI ​​agent pre-configures the control tree recognition rules through the following steps: Configure rules based on HTML DOM parsing, applicable to the interface of HIS systems built on web technologies. The configuration of rules based on HTML DOM parsing includes: Element identification rules are based on information such as the element's tag name, ID, class name, and attribute value. If an element does not have an explicit ID or class name, it is identified based on the element's hierarchical relationship in the DOM tree and the characteristics of adjacent elements. The attribute matching rules determine the function of special elements, including buttons, by matching attribute values. For text with differences in capitalization, case-insensitive matching is performed. The rules for handling dynamic elements are as follows: For dynamically loaded elements, an event listener mechanism is used. When specific events, including clicks and scrolling, occur on the page, the DOM tree is re-parsed to identify the newly appearing elements. Configure rules based on the Windows API, applicable to the interface of HIS systems built on Windows desktop application technology. The rules configuration based on the Windows API include: Window handle acquisition: Use Windows API functions to obtain the handle of a window in the HIS system, by searching for the window title or class name; Control enumeration: Enumerate all controls within the window using the Enum Child Windows function to obtain the handle and basic information of each control; To obtain control properties, use the Get Window Text function or the Get Class Name function to get the text content and class name property values ​​of the control, and determine the type and function of the control based on the property values; For control manipulation, the corresponding Windows API functions are used to operate on different types of controls.

5. The cross-platform medical registration interaction method based on AI intelligent agent and multimodal recognition according to claim 2, characterized in that, Before each element location operation, the AI ​​agent first takes a screenshot of the current interface and compares it with a pre-saved screenshot of the standard interface. It calculates the positional difference between corresponding elements in the two screenshots. If the positional difference exceeds a preset threshold, it determines that the element has shifted. For elements that have shifted, the AI ​​agent needs to calculate the shift ratio in both the horizontal and vertical directions. The formula for calculating the shift ratio is: p x =(x2-x1) / x1; p y =(y2-y1) / y1; Where (x1, y1) represents the coordinates of the top-left corner of an element in the screenshot of the standard interface, (x2, y2) represents the coordinates of the top-left corner of an element in the screenshot of the current interface, and p x p represents the horizontal offset ratio. y Indicates the offset ratio in the vertical direction; Based on the calculated offset ratios in the horizontal and vertical directions, the positions of the elements requiring positioning are adjusted using the following formula: Where (x0, y0) represents the click position of the element to be positioned before adjustment. This indicates the clickable position of the element that needs to be positioned.

6. The cross-platform medical registration interaction method based on AI intelligent agent and multimodal recognition according to claim 1, characterized in that, The AI ​​agent was pre-written using Python's PyAutoGUI library to simulate human mouse and keyboard operations based on the HIS system's operating procedures, including login, information entry, and registration operations. The script sets a randomized range for the interval between simulating human mouse and keyboard operations. The interval is randomly selected between 0.5 seconds and 1 second to avoid being detected as automated by the HIS system.

7. A cross-platform medical registration interaction method based on AI intelligent agent and multimodal recognition according to any one of claims 2 to 6, characterized in that, The message queue is pre-built on the front-end machine in the form of a message queue module. RabbitMQ message queue software is selected for installation and configuration to ensure that it can stably receive registration request data sent by the online registration system, ensure that it can interact with the AI ​​agent deployed on the intranet server, and configure firewall rules to only allow specific IP addresses to access the message queue and enable HTTPS encrypted transmission. The AI ​​agent and HIS terminal are pre-deployed on the intranet server to ensure that the intranet server can run the AI ​​agent and HIS terminal stably. The AI ​​agent is initialized and configured, including loading the image recognition model and setting the control tree recognition rules. The front-end interaction module of the online registration system adopts the form of a web page or mobile application, and interacts with the message queue module on the front-end machine through the network.

8. A cross-platform medical registration and interaction system based on AI intelligent agents and multimodal recognition, used for data interaction with the HIS system, characterized in that, include: The front-end interaction module is deployed on the user end, the message queue module is deployed on the front-end machine, and the AI ​​intelligent agent and HIS terminal are deployed on the intranet server. The front-end interaction module is used to allow users to enter registration request data through the online registration system and send it to the message queue module, as well as to receive registration result information from the message queue module. The message queue module is used to receive and store registration request data for the AI ​​agent to access and read, as well as to receive registration result information from the AI ​​agent and feed it back to the front-end interaction module. HIS terminal, or HIS system terminal program or page, is used to provide the operation interface for registration services, including login interface, information entry interface, registration confirmation interface and registration result interface. The AI ​​agent is used to access the message queue module via the intranet to read registration request data, simulate manual operation to log in to the HIS terminal, enter registration request data, execute registration operation, parse registration result information, and send the registration result information back to the message queue module.

9. An AI intelligent agent, characterized in that, The AI ​​agent is equipped with at least the deep learning framework TensorFlow and YOLO algorithm model, and integrates an image recognition module, a control tree plugin, and a script for simulating mouse and keyboard operations. It can accurately locate elements on each page of the HIS system, including input boxes, selection boxes, and buttons, and also has the ability to simulate operation execution. The AI ​​agent can access the message queue module via the intranet to read registration request data, simulate manual operation of the keyboard and mouse to log in to the HIS system, enter registration request data, execute registration operations, parse registration result information, and send the registration result information back to the message queue module. The AI ​​agent also has dynamic decision-making capabilities, and can adjust the operation process according to changes in the interface.

Citation Information

Patent Citations

  • AI port platform, application method thereof and AI application system

    CN107610763A

  • Registration method and apparatus, storage medium and electronic equipment

    CN108053046A