Intelligent medical care visual voice call control method and system
By real-time monitoring of patients' vital signs and environmental data, combined with multimodal input to automatically identify patients' intentions, the system solves the problems of slow response and low accuracy of existing intelligent voice call systems, and achieves fast and accurate allocation of medical resources and patient safety.
Patent Information
- Application Number
- CN202510907343.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-09-19
AI Technical Summary
Existing intelligent voice call systems in the medical field have problems such as slow response speed, frequent manual intervention, and inability to accurately identify patient needs. Especially in emergency scenarios, it is difficult to respond to patient needs quickly and accurately.
By real-time monitoring of patients' vital signs and environmental data, combined with multimodal inputs such as voice, text, and images, it automatically identifies patients' intentions, sets thresholds to trigger alarms, and allocates medical resources through an intelligent call control center to achieve rapid response without human intervention.
It significantly improves emergency response speed and judgment accuracy, reduces patients' dependence on medical equipment operations, supports home and community medical scenarios, rationally allocates medical resources, and ensures patient health and safety.
Smart Images

Figure CN120676093A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent medical care technology, and in particular to an intelligent medical care visual voice call control method and system. Background Art
[0002] Intelligent voice call systems are increasingly being used in the medical field, primarily through speech recognition, natural language processing, IoT technology, and intelligent decision-making algorithms to improve the efficiency and quality of medical services. Traditional medical call systems typically rely on manual operation or simple button presses to initiate requests for help, resulting in slow response times, frequent manual intervention, and inability to accurately identify patient needs. Especially in emergency settings, quickly and accurately responding to patient needs has become a major challenge in improving medical quality and patient safety.
[0003] The emergence of intelligent voice call systems aims to address the challenges of traditional medical call systems through voice recognition and voice interaction. Through voice input, patients can make calls without physical contact, which is particularly important for those who are physically unable to operate equipment or in emergency situations. Intelligent voice call systems not only recognize basic patient voice commands but also incorporate artificial intelligence technologies for semantic understanding, sentiment analysis, and intent recognition, automatically adjusting the priority and type of call response. For example, the system can determine whether the patient's voice command is an emergency and immediately initiate emergency response.
[0004] However, existing intelligent voice call systems still face challenges in many areas. For example, speech recognition accuracy, the ability to identify different patient needs, interference from ambient noise, and adaptability in different scenarios remain key issues that require improvement and advancement. Therefore, the challenge is to improve the efficiency and reliability of intelligent voice call systems through more precise algorithms, more data support, and more efficient decision-making mechanisms. Summary of the Invention
[0005] In response to the above-mentioned defects, the technical problem solved by the present invention is to provide an intelligent medical visual voice call control method system, which automatically monitors changes in multiple key parameters including oxygen concentration, heart rate, respiratory rate, blood pressure, body temperature, falls, CO2 concentration, etc., and automatically triggers an alarm. It also combines voice calls and visual information to promptly transmit alarm information to relevant medical staff, thereby improving treatment efficiency, reducing human intervention, ensuring the health and safety of patients (monitored objects), and reducing the risk of danger.
[0006] A first aspect of the present invention provides an intelligent medical visual voice call control method, comprising: S1: Real-time monitoring acquires first data representing the vital characteristics of the monitored subject, second data about the monitored subject's environment, and the monitored subject's expressed intentions. S2: Based on changes in the first and second data, and the monitored subject's expressed intentions, intelligently calls a control center.
[0007] According to an embodiment of the present invention, the expression form of the monitored subject's expression intention includes: voice, text, code, picture, and machine recognition language converted by button.
[0008] According to one embodiment of the present invention, the intelligent call control center based on the changes in the first data, the second data, and the expression intention of the monitored subject includes: According to the first data, the intelligent call control center is configured, including setting a first threshold value related to the first data, monitoring the change of the first data in real time and comparing whether the first data is not less than the first threshold value. If the first data is not less than the first threshold value, the intelligent call control center is configured.
[0009] According to the changes in the second data, the intelligent call control center is configured to set a second threshold related to the second data, monitor the changes in the second data in real time and compare whether the second data is not less than the second threshold. If the second data is not less than the second threshold, the intelligent call control center.
[0010] The intelligent call control center uses machine-recognized language based on voice, text, code / command, image, and button conversion, including: Voice is collected through a microphone, sampled into digital signals, an audio stream is generated, sound features are extracted, and a sound matrix is formed; the pronunciation probability distribution at different time points is output based on the sound features and converted into text; the text is directly collected; the code is parsed to extract the corresponding intention; the intention code corresponding to each button is set; the image area is bound to its corresponding label; the text, intention, and label data are preprocessed to perform intent recognition and semantic analysis, and are output in a unified format, which are packaged and sent together with the monitoring device number, the identity information of the monitored subject, the environmental location information of the monitored subject, and the call timestamp data.
[0011] A second aspect of the present invention provides an intelligent medical visual voice call control method for responding to a call request from a monitored subject, comprising: receiving the call request, parsing the call request, and allocating medical resources corresponding to the call request.
[0012] According to an embodiment of the present invention, parsing the call request includes: obtaining the identity information of the monitored subject, the monitoring device number, the environmental location information of the monitored subject, the timestamp of the request, and the request intention information.
[0013] According to one embodiment of the present invention, the request intention information is calibrated with different priority levels, including: collecting the medical records and first request intention information of the monitored subject, performing intelligent model training, obtaining the priority level of the first request intention information related to the medical records of the monitored subject, calibrating the first request call information according to the priority level, and obtaining the second call request information.
[0014] According to one embodiment of the present invention, allocating medical resources corresponding to the call request according to the call request includes: Allocating medical resources corresponding to the monitored subject based on the monitored subject's identity information, the timestamp of the request, and the second request intention information: including: arranging the second request intention information in sequence according to the second request intention information and the timestamp of the request. Searching for the monitored subject's identity information corresponding to the timestamp of the request of the arranged second request intention information to complete the arrangement order of the monitored subject's request for allocation of medical resources. Allocating medical resources to the monitored subject in sequence according to the arrangement order of the monitored subject's request for allocation of medical resources.
[0015] A third aspect of the present invention discloses an intelligent medical visual voice call control system, comprising: a collection module for real-time monitoring and acquiring first data representing vital characteristics of a monitored subject, second data representing the monitored subject's environment, and the monitored subject's expressed intentions; and a call module for intelligently calling a control center based on changes in the first and second data and the monitored subject's expressed intentions.
[0016] According to an embodiment of the present invention, the expression form of the monitored subject's expression intention includes: voice, text, code, picture, and machine recognition language converted by button.
[0017] According to one embodiment of the present invention, the call module includes: a first call unit, used to intelligently call a control center based on the first data, including setting a first threshold related to the first data, monitoring the changes in the first data in real time and comparing whether the first data is not less than the first threshold, and if the first data is not less than the first threshold, intelligently call the control center.
[0018] The second call unit intelligently calls the control center according to the second data change, including setting a second threshold related to the second data, monitoring the second data change in real time and comparing whether the second data is not less than the second threshold. If the second data is not less than the second threshold, the intelligent call control center.
[0019] The third call unit is used to convert the machine-recognized language into an intelligent call control center based on voice, text, code / instruction, picture, and button, including: an acquisition unit, used to collect voice through a microphone, sample it into a digital signal, generate an audio stream, extract sound features, and form a sound matrix; output the pronunciation probability distribution at different time points according to the sound features and convert it into text; directly collect the text; parse the code and extract the corresponding intention; set the intention code corresponding to each button; and bind the image area to its corresponding label.
[0020] The data preprocessing unit is used to preprocess the text, intent, and tag data, perform intent recognition and semantic analysis, and output them in a unified format.
[0021] The sending unit is used to package and send the data from the data pre-processing unit, the monitoring device number, the identity information of the monitored subject, the environmental location information of the monitored subject, and the call time.
[0022] A fourth aspect of the present invention discloses an intelligent medical visual voice call control system, which is used to respond to a call request from a monitored subject and includes: a receiving module for receiving and parsing the call request.
[0023] The response module is used to allocate medical resources corresponding to the call request according to the call request.
[0024] According to one embodiment of the present invention, the receiving module includes: a receiving unit for receiving the call request; a parsing unit for parsing the call request, including: parsing the call request into identity information of the monitored subject, a monitoring device number, information about the environmental location of the monitored subject, a timestamp of the request, and information about the request intent.
[0025] According to one embodiment of the present invention, the system also includes a training module for training the priority level of the medical record relationship between the call request and the identity information of the monitored subject; and includes: a collection unit for collecting the medical record of the monitored subject and the first request intention information.
[0026] A training unit is used to take the first request intention information and the corresponding medical record information of the monitored subject as input for training, obtain the priority level of the first request intention information related to the medical record of the monitored subject, and calibrate the first request call information as the second request call information according to the priority level.
[0027] According to one embodiment of the present invention, the response module includes: a first arrangement unit for sequentially arranging the second request intent information according to the second request intent information and the timestamp of the request. A second arrangement unit for searching for the identity information of the monitored subject corresponding to the timestamp of the request of the arranged second request intent information to complete the arrangement order of the monitored subject's request for allocation of medical resources. An allocation unit for sequentially allocating medical resources to the monitored subject according to the arrangement order of the monitored subject's request for allocation of medical resources.
[0028] The fifth aspect of the present invention provides an intelligent device, including a transmitter, a receiver, a memory and a processor; the memory is used to store computer instructions; the processor is used to run the computer instructions stored in the memory to implement the above intelligent medical visual voice call control method. A sixth aspect of the present invention provides a storage medium, comprising: a readable storage medium and computer instructions, wherein the computer instructions are stored in the readable storage medium; the computer instructions are used to implement the above-mentioned intelligent medical visual voice call control method.
[0029] The beneficial effects provided by the present invention are as follows: First, a multimodal call request control method is adopted to significantly improve the emergency response speed and judgment accuracy, while reducing the patient's dependence on the operation of medical equipment and improving the self-service efficiency; at the same time, it supports the wide deployment of home scenarios, community medical care, and smart terminals, and the adaptation scenarios are not restricted. Secondly, through intelligent model pre-training, different intention information is matched with the patient's condition. In the case of limited medical resources or a large number of patients, medical resources are allocated fairly and orderly according to the severity of each patient's condition, ensuring efficient and orderly response to requests, and rationalizing and equitably dividing medical resources to provide patients with fair and reasonable medical treatment. Furthermore, through iterative training and learning of the patient's request intention and the patient's condition data by the intelligent model, the trend of the condition is monitored and predicted in real time, and a better and more predictable treatment plan is provided to the monitored subject in advance, completing the medical and health ecosystem. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0031] Figure 1 This is a flow chart of the intelligent medical visual voice call control method disclosed in an embodiment of the present invention; Figure 2 This is a flow chart of another intelligent medical visual voice call control method disclosed in an embodiment of the present invention; Figure 3 This is a block diagram of the intelligent medical visual voice call control system disclosed in an embodiment of the present invention; Figure 4 This is a block diagram of another intelligent medical visual voice call control system disclosed in an embodiment of the present invention.
[0032] The above drawings illustrate specific embodiments of the present disclosure, which will be described in more detail below. These drawings and textual descriptions are not intended to limit the scope of the present disclosure in any way, but rather to illustrate the concepts of the present disclosure to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0033] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.
[0034] The intelligent medical visual voice call system integrates voice recognition, visual interfaces, and artificial intelligence analysis technologies to enable intelligent interaction, task management, and dynamic scheduling between patients and nurses. It supports patients initiating voice requests, automatically identifies intent, determines priority, and completes AI-powered medical resource scheduling. A visual platform uniformly displays, records, and tracks task status.
[0035] The first aspect of the present invention provides an intelligent medical visualization voice call control method, such as Figure 1 As shown, S1: real-time monitoring obtains first data representing the life characteristics of the monitored subject, second data of the environment where the monitored subject is located, and the expression intention of the monitored subject.
[0036] The data representing the vital characteristics of the monitored subject here include heart rate, blood oxygen saturation, respiratory rate, body temperature, blood pressure, abnormal electrocardiogram waveform, body position, skin electrical response / skin temperature, etc.
[0037] The sensor module is used to monitor and obtain data representing the vital characteristics of the monitored subject in real time, as well as data about the environment in which the monitored subject is located. This module transmits the monitoring data to the data processing module via wireless or wired means.
[0038] Specifically, taking oxygen concentration as an example, oxygen sensors are installed in wards or on patients to continuously monitor the patient's oxygen concentration. The data collected by the oxygen sensors is transmitted to a data processing center via wireless communication.
[0039] In addition, the real-time oxygen concentration data is analyzed by the data processing module, which determines whether the oxygen concentration is within the safe range based on preset thresholds. If the oxygen concentration falls below the set safety threshold, the system will trigger an alarm. When the oxygen concentration falls below the safety threshold, the system immediately activates the alarm system and alerts medical staff through a voice broadcast function. The voice broadcast content includes the prompt "Oxygen concentration is too low, please address it immediately."
[0040] The system displays real-time oxygen concentration data, alarm information, and corresponding treatment steps to medical staff through a visual interface. The monitor will have a graphical display, including the oxygen concentration change curve, alarm time, treatment suggestions, and other information.
[0041] While making the voice broadcast, the system will send an emergency call to medical staff or hospital rescue center through the built-in communication module (such as Wi-Fi, GSM, etc.) to ensure timely response.
[0042] For example, heart rate, detected by electrocardiogram or photoelectric method, reflects heart health. If it is too fast or too slow, emergency calls will be automatically made. Blood oxygen saturation, monitored together with oxygen concentration, is an important indicator. If it is below 90%, an alarm will be triggered. Respiratory rate, which indicates the speed of breathing, can be detected by chest movement, and alarms will be triggered if it stops or becomes short of breath. Body temperature, commonly used to detect fever, infection, etc., will trigger a warning if it exceeds 38°C or is below 35°C. Blood pressure, abnormal upper and lower pressure (such as hypertensive crisis), will automatically contact medical staff. Abnormal electrocardiogram (ECG) waveforms, such as arrhythmia and atrial fibrillation, are visualized in real time with voice broadcast. Body position status (fall detection), using acceleration sensors to detect falls, and calling immediately after the elderly fall. Skin galvanic response / skin temperature, used to assess stress, shock and other conditions, indicating hypothermia, overheating, and tension.
[0043] Environmental data for the monitored subject includes: ambient oxygen concentration, carbon dioxide concentration, carbon monoxide concentration, smoke / fire alarm, temperature and humidity, light intensity, etc. Especially in confined spaces or when using oxygen cylinders, an ambient oxygen concentration below 18% may cause suffocation, triggering an automatic alarm; poor indoor ventilation or respiratory arrest can lead to an increase in carbon dioxide concentration, prompting windows to be opened or exhaust ventilation to be activated; carbon monoxide concentrations pose a high risk of poisoning, requiring immediate alarm, linked ventilation, or automatic emergency calls. Smoke / fire alarms are also part of safety infrastructure. If smoke is present, a direct voice notification will be given and 119 will be automatically dialed. Maintain a comfortable nursing environment, and notify the adjustment equipment of abnormal high temperatures or humidity. Insufficient light prompts the user to illuminate the room at night, which can be used for nighttime lighting automation.
[0044] The expressions of the monitored subject's intention include: voice, text, code, picture, and machine recognition language converted by buttons.
[0045] S2: Based on changes in the first and second data and the monitored subject's expressed intention, the health and safety of the monitored subject is monitored in real time, and a medical worker is intelligently called. This includes: intelligently calling a medical worker based on the first data, including setting a first threshold related to the first data, monitoring changes in the first data in real time, and comparing whether the first data is not less than the first threshold. If the first data is not less than the first threshold, intelligently calling a medical worker.
[0046] Intelligently calling a medical worker based on the change of the second data includes: setting a second threshold related to the second data, monitoring the change of the second data in real time and comparing whether the second data is not less than the second threshold, and intelligently calling a medical worker if the second data is not less than the second threshold.
[0047] The first threshold or the second threshold can be set according to the patient's condition. For example, some patients focus on blood pressure, some patients focus on respiratory rate, and some patients focus on electrocardiogram, etc. Different vital sign parameters or combinations of parameters can set thresholds for important observation data for different patients in turn, which is convenient for targeted and enhanced care of different patients' conditions.
[0048] An intelligent module can be used here, taking the big data patient's vital characteristics data as input. Pre-training learning includes simulation learning of different patients' conditions, and outputting potential risk values and condition predictions corresponding to different patients' conditions; using its predicted data as a reference standard for the patient's condition, and comparing it with the patient's current vital characteristics data, the patient's condition and health status are tracked in real time, making it convenient for subsequent hospitals to observe the condition at different time periods and make reasonable treatment plans at the moment.
[0049] The monitored subject's expressed intent can be voice, text, or machine language converted from keystrokes. For example, if a patient says "Nurse, I feel dizzy," the system will recognize "I feel dizzy" as the request. The monitored subject's expressed intent is identified using speech recognition engines such as Whisper, iFlyTek, and Google ASR.
[0050] First, capture the subject's voice using a microphone or other voice acquisition device. Keep the environment quiet and avoid background noise. Perform pre-processing on the audio signal using denoising and echo cancellation. For different engines, perform sampling rate conversion and format conversion (for example, converting the audio to WAV or FLAC).
[0051] The microphone collects a continuous analog audio signal, which means , where t represents time. To make it suitable for computer processing, it is converted into a digital signal and sampled.
[0052] (1) Where T is the sampling period (usually 1 / 16KHZ or 1 / 44.1KHZ), is a discrete signal.
[0053] The sampled digital signal is converted into acoustic features suitable for speech recognition, and the audio signal is subjected to short-time Fourier transform using Mel-frequency cepstral coefficients (MFCC) to obtain a spectrogram, which is then processed through a Mel filter.
[0054] (2) in, For signal In the time-frequency domain, is the window function. The spectrum obtained by short-time Fourier transform is processed by the Mel filter bank to obtain the Mel spectrum: (3) in, The response of the mth Mel filter in the Mel spectrum when The weights of the Mel filter bank are Time is the frequency component in the short-time Fourier transform.
[0055] Mel spectrum After logarithmic transformation and discrete cosine transformation, the Mel frequency cepstrum coefficients are obtained: (4) in, is the Mel frequency cepstral coefficient, which represents the characteristics of the audio signal on the Mel frequency scale at that moment. M is the number of coefficients, is the kth frequency response of the Mel spectrum.
[0056] After modeling the sound features, the pronunciation frequency distribution at each time point is output using a deep neural network or a long short-term memory network.
[0057] For a given feature matrix , the acoustic model outputs the pronunciation probability distribution at each time step t ,in is the label of the factor or word corresponding to the time point.
[0058] (5) in, Represents the i-th Mel-frequency cepstral coefficient at time t. The phoneme, word or phonetic symbol at time t; Represents the feature vector at time t (such as MFCC, treble, etc.). Indicates that the features are processed by the neural network model The mapping output of is a parameter of the model, and the Softmax function normalizes the output probability to the probability distribution of each word or factor. Represents the probability of the phoneme or word corresponding to the current time t.
[0059] It is the feature vector of the t-th frame, which consists of multiple features, including multiple Mel-frequency cepstral coefficients, which contain temporal context information. The feature vector extracted from the tth frame of the audio signal. It is a representation of the audio signal after converting the time domain data into frequency domain features, usually containing Mel-frequency cepstral coefficients (MFCCs) and other acoustic features (such as zero-crossing rate and pitch).
[0060] It is the feature data input into the acoustic model (such as neural network, HMM, etc.) to predict the probability distribution of pronunciation.
[0061] are the individual Mel-frequency cepstral coefficients, It is the feature vector at time t, which contains multiple Mel-frequency cepstral coefficients (and other possible features). In the feature extraction process, the Mel-frequency cepstral coefficients of multiple frames are firstly taken and then combined into a feature vector. , this vector contains all relevant features at the tth moment.
[0062] The decoding process typically involves mapping the probability output of the acoustic model to actual text, using a language model (such as an N-gram model or a neural network language model) to ensure that the output text is grammatically reasonable.
[0063] The probability output by the acoustic model and the probability provided by the language model play different roles in the speech recognition system, and they work together to improve the accuracy of speech-to-text conversion.
[0064] Use recursive neural networks to model the context of previous words and output the probability distribution of the current word. Similarly, long short-term memory networks can be used in language models to capture long-range dependent context information. Given a large amount of text data, the neural network learns the dependency relationship between words and outputs the conditional probability of each word. .
[0065] Language models describe the relationships between words and assess the plausibility of a word sequence within a specific language. For example, they measure the probability of a word appearing in a given context. During speech recognition, language models help predict which words are more likely to appear in a given context. Their primary purpose is to smooth speech recognition output and generate text that better conforms to linguistic rules.
[0066] In speech recognition, the language model provides conditional probabilities for different word sequences. Indicates the probability of the current word t given the previous (t-1) words in the context. Helps adjust the speech recognition output. For example, the speech recognition system can recognize both "to" and "too", but if the context should be "to", then the language will choose which one to give a higher probability. represents the output vocabulary at time t, Represents the vocabulary at the previous (t-1) time. The acoustic model provides the pronunciation probability, and the language model provides the probability of language rationality.
[0067] If the probability output by the acoustic model is , the probability provided by the language model is , then we need to get the optimal word sequence by maximizing the joint probability: (6) in, Represents the optimal word sequence output by the decoder. Finally, the speech recognition system converts the audio signal into the corresponding text output through the decoding process: (7) in, is the final text output.
[0068] Use speech recognition engines such as Whisper, iFlyTek, and Google ASR to convert speech signals into text. For example, use an ASR model (such as Whisper) to convert speech to text. Text input can be done manually or through web input boxes, directly capturing text without conversion. Each button corresponds to an intent code, such as selecting an image or binding different labels to corresponding areas, such as clicking the "lower abdomen" area. For code / command input, perform formal parsing on the input code or command to extract key intent.
[0069] Process different types of input (voice, text, code, images, buttons, etc.) uniformly and standardize the output. Use appropriate technologies (such as ASR, NLP, regular expressions, and computer vision) to preprocess different types of input and convert the raw input into processable features.
[0070] Intent recognition is performed through the intent recognition module. For example, for text input, intent classification is performed using an NLP model. Code / command input is classified using regular expression matching and rule mapping. Image input is classified using image recognition technology. Button input is analyzed using event listeners and preset button mappings.
[0071] Convert the parsed intent information into a unified standard format for structured output, ensuring that intents of different input types can be processed uniformly. Regardless of the input type, they are ultimately converted into a unified intent format.
[0072] For example, the voice input "I have a headache, please give me some painkillers." is converted to: { "type": "speech", "intent": "request_medicine", "details": { "body_part": "head", "urgency": "medium" } } Text input: "Please call the nurse for me." is converted to: { "type": "text", "intent": "call_nurse", "details": {} } Code / instruction input: #call_nurse. Translated to: { "type": "command", "intent": "call_nurse", "details": {} } Image input: Click on the "abdomen" image to convert to: { "type": "image_click", "intent": "pain_report", "details": { "body_part": "abdomen" } } Button input: Click the "Call Nurse" button. Translates to: { "type": "button_click", "intent": "call_nurse", "details": {} } Through multimodal input, it is compatible with the actual conditions of more patients, and also provides patients with multiple ways to call medical workers, ensuring that patients can call the control center for timely medical treatment.
[0073] When a call request is transmitted, it carries the device number, patient name, ward number, and call requirement information. These data are encapsulated in a unified structure, which is standardized in the JSON format in this invention: { "device_id": "device12345", / / device number "patient_info": { "name": "Zhang San", / / patient's name "room_number": "A202" / / Ward number }, "call_request": { "intent": "call_nurse", / / Call request information, such as: calling a nurse "urgency": "high", / / Urgency, 'high', 'medium', 'low' "details": { "body_part": "abdomen", / / The relevant body part (if applicable) "additional_notes": "abdominal pain" / / Other supplementary information } }, "timestamp": "2025-06-26T15:30:00" / / The timestamp of the request to ensure the timeliness of the data } Among them, device_id: Each device has a unique identifier used to distinguish different devices. patient_info: Contains basic patient information, such as name and ward number. call_request: Contains detailed information about the call request, such as the request type (such as calling a nurse, requesting medication, etc.), urgency (such as high, low, or medium), and possible additional instructions. timestamp: Ensures the recording of time, so that data can be sorted or analyzed over time.
[0074] Whenever a device is activated (e.g., a bedside call button, a voice assistant, or a mobile device), it first collects relevant data, including patient information and call request. The device obtains patient information (device ID, patient name, ward number) from local data or user input. The device can display patient information through a user interface, obtain user confirmation through voice interaction, or perform fingerprint entry to confirm the user's identity.
[0075] The above information is encapsulated into a unified JSON format. To ensure the security of data during transmission, sensitive information can be encrypted, such as encrypting the patient's name and / or ward number. The JSON data is encrypted using the AES encryption algorithm, and the data packet is digitally signed to ensure that the data is not tampered with.
[0076] Data is sent to the control center API via POST requests. For applications requiring real-time responses, data can be sent via WebSockets. Asynchronous transmission can also be achieved through message queues like Kafka or RabbitMQ, ensuring efficient transmission of large amounts of data.
[0077] For sensitive information such as patient names or medical conditions, strong encryption algorithms (such as AES-256) are used to encrypt sensitive data to ensure that the data cannot be stolen during transmission.
[0078] The second aspect of the present invention provides an intelligent medical visual voice call control method for responding to a call request from a monitored subject, such as Figure 2 As shown, the method includes: receiving the call request, parsing the call request, and allocating medical staff according to the call request.
[0079] The control center receives HTTP requests or WebSocket messages from clients (such as bedside call buttons or mobile devices), including detailed request information, and passes them to the intent recognition and dispatch system for processing.
[0080] Extract request data, including the identity confidence of the monitored subject, such as the patient's name, ward number or equipment number, and call request text information, such as "call a nurse" or "request medicine" and other specific intentions.
[0081] Based on the JSON representation of the client's call request, the server receives the HTTP POST request through a framework such as Flask or FastAPI. The server listens to a specified interface (such as / api / request) and waits for the request. Alternatively, the client can use a combination of GET, POST, PUT, PATCH, DELETE, OPTIONS, and HEAD request methods. The server can classify the requests and return different responses based on the classification results.
[0082] Get the request body and extract the JSON data from the HTTP request. Get the following fields from the data dictionary: the identity of the monitored subject, the monitoring device ID, the monitored subject's location, the timestamp of the request, and the request intent (e.g., "I want to change my dressing," "I want to take my temperature," "I feel chest tightness and need a physical examination," etc.). The request intention information is marked with different priority levels, including: collecting the medical records of the monitored subject and the first request intention information, and performing intelligent training. The first request intention information here represents the big data model or clinical practice data, and performing model pre-training to obtain the priority level of the request intention information related to the medical records of the monitored subject, and obtaining the second call request information according to the priority level; then the second call request information and its corresponding medical records are used as input to this intelligent model and input into the training.
[0083] For example, the request intent for patients with hypertension and diabetes is: "I have a headache", "I have lung discomfort and difficulty breathing", "I am vomiting", "I feel uncomfortable in my heart", which are first-level response requests, high priority, and emergency requests.
[0084] For example, for patients with hypertension and diabetes, the request intents of "I want to go to the bathroom" and "I want to change my medicine" are second-level response requests. They are medium priority and are routine diagnostic requests.
[0085] For example, requests for patients with hypertension and diabetes such as "When can I be discharged from the hospital?" and "I would like to make an appointment for a check-up the day after tomorrow" are of low priority.
[0086] All requests for different medical records are prioritized, rigorously calibrated based on clinical medical cases and triage within the big data model. Furthermore, when changes occur to the triage system or clinical medical cases, the intelligent model's parameters and standards are promptly revised. The correlations between newly added conditions or cases and patient needs are continuously refined through intelligent model training and learning.
[0087] At the same time, the intelligent model also has the function of pre-judgment. According to the input of request intention data, it can understand the development and changes of the patient's condition or predict the possible development direction and risks of the condition, so that when the patient has hidden risks of the condition, it can be discovered and treated early, and the direction of the condition can be monitored and predicted in real time, providing the monitored subject with better predictable treatment plans in advance, and completing the medical and health ecosystem.
[0088] The calibrated medical record request priorities are sorted by timestamp. For example, in the advanced medical record request for diabetic patients, the request intent with a timestamp of 22:30, "I have lung discomfort and difficulty breathing," and the request intent with a timestamp of 22:29, "I vomit," are arranged in order of timestamp and corresponding medical resources are allocated to them.
[0089] Typically, servers receive HTTP POST requests through frameworks like Flask or FastAPI, and they listen on a specific endpoint (e.g., / api / request) for incoming client requests. Flask is a lightweight web framework suitable for building small web applications or API services. FastAPI is a modern, high-performance web framework suitable for building RESTful APIs, offering superior performance and automatic documentation generation. Clients send data to the server via HTTP POST requests, and the server receives and processes the data.
[0090] An endpoint is a URL exposed by a server to clients. Clients interact with the server by accessing these endpoints. The server, through a framework such as Flask or FastAPI, listens for and processes client requests at a specific URL path. Once a client sends a request to this path, the server receives and parses the request and returns a response based on the request content.
[0091] In a hospital's intensive care unit (ICU), the system monitors the oxygen levels of multiple patients within the ward and transmits the data in real time to a centralized processing system. By setting an oxygen concentration threshold (for example, 18%), the system automatically triggers an alarm if a patient's oxygen concentration drops below that threshold.
[0092] When an alarm is triggered, the system will announce the patient's oxygen level is low, please check immediately. A visual display will also show the patient's current oxygen level and historical data trends. If a response is not immediate, the system will also initiate an emergency call to the nurses and attending physician in the ward to ensure prompt action.
[0093] In home care, an intelligent medical visual voice call control system can be used for daily monitoring of the elderly or patients with chronic diseases. An oxygen sensor is installed in the home environment and connected to a smartphone via Wi-Fi. When the patient's oxygen concentration falls below a set threshold (for example, 19%), the system issues a voice alarm and sends a notification via a mobile app.
[0094] The system automatically adjusts oxygen supply equipment (e.g., oxygen cylinders, oxygen machines, etc.) based on real-time data and notifies family members through voice notifications: "The patient's oxygen concentration is too low and the oxygen supply has been activated." If the situation is serious, the system will automatically send an emergency call to the family, reminding them to take further action as soon as possible.
[0095] Monitoring oxygen concentration is crucial in high-risk environments like mines and chemical plants. If the system detects oxygen concentration below a preset threshold (e.g., 18%), it will immediately announce "Warning: Low oxygen concentration, evacuate immediately" and display the current oxygen concentration and safe evacuation routes.
[0096] The system will also send emergency call information to the wireless communication equipment of on-site staff through the emergency communication module and automatically activate the emergency evacuation plan.
[0097] The third aspect of the present invention discloses an intelligent medical visual voice call control system 30, such as Figure 3 As shown, the system includes: a collection module 301 for real-time monitoring and acquiring first data representing the vital characteristics of the monitored subject, second data representing the environment in which the monitored subject resides, and the expressed intention of the monitored subject; and a call module 302 for intelligently calling the control center based on changes in the first and second data and the expressed intention of the monitored subject.
[0098] According to an embodiment of the present invention, the expression form of the monitored subject's expression intention includes: voice, text, code, picture, and machine recognition language converted by button.
[0099] According to one embodiment of the present invention, the call module includes: a first call unit, used to intelligently call a control center based on the first data, including setting a first threshold related to the first data, monitoring the changes in the first data in real time and comparing whether the first data is not less than the first threshold, and if the first data is not less than the first threshold, intelligently call the control center.
[0100] The second call unit intelligently calls the control center according to the second data change, including setting a second threshold related to the second data, monitoring the second data change in real time and comparing whether the second data is not less than the second threshold. If the second data is not less than the second threshold, the intelligent call control center.
[0101] The third call unit is used to convert the machine-recognized language into an intelligent call control center based on voice, text, code / instruction, picture, and button, including: an acquisition unit, used to collect voice through a microphone, sample it into a digital signal, generate an audio stream, extract sound features, and form a sound matrix; output the pronunciation probability distribution at different time points according to the sound features and convert it into text; directly collect the text; parse the code and extract the corresponding intention; set the intention code corresponding to each button; and bind the image area to its corresponding label.
[0102] The data preprocessing unit is used to preprocess the text, intent, and tag data, perform intent recognition and semantic analysis, and output them in a unified format.
[0103] The sending unit is used to package and send the data from the data pre-processing unit, the monitoring device number, the identity information of the monitored subject, the environmental location information of the monitored subject, and the call time.
[0104] The fourth aspect of the present invention discloses an intelligent medical visual voice call control system, such as Figure 4 As shown, the system 40 is used to respond to a call request from a monitored subject, and includes: a receiving module 401 for receiving and analyzing the call request; and a responding module 402 for allocating medical resources corresponding to the call request.
[0105] The receiving module includes: a receiving unit for receiving the call request; a parsing unit for parsing the call request, including: parsing the call request into the identity information of the monitored subject, the monitoring device number, the location information of the monitored subject, the timestamp of the request, and the request intention information.
[0106] The system further includes a training module for training the priority level of the medical record relationship between the call request and the monitored subject's identity information; and includes: a collection unit for collecting the monitored subject's medical record and first request intention information.
[0107] A training unit is used to take the first request intention information and the corresponding medical record information of the monitored subject as input for training, obtain the priority level of the first request intention information related to the medical record of the monitored subject, and calibrate the first request call information as the second request call information according to the priority level.
[0108] The response module includes: a first arrangement unit for sequentially arranging the second request intent information according to the second request intent information and the timestamp of the request occurrence; a second arrangement unit for searching for the identity information of the monitored subject corresponding to the timestamp of the request occurrence of the arranged second request intent information to complete the arrangement order of the monitored subject's request for allocation of medical resources; and an allocation unit for sequentially allocating medical resources to the monitored subject according to the arrangement order of the medical resource allocation requests.
[0109] The fifth aspect of the present invention provides an intelligent device, including a transmitter, a receiver, a memory and a processor; the memory is used to store computer instructions; the processor is used to run the computer instructions stored in the memory to implement the above intelligent medical visual voice call control method. A sixth aspect of the present invention provides a storage medium, comprising: a readable storage medium and computer instructions, wherein the computer instructions are stored in the readable storage medium; the computer instructions are used to implement the above-mentioned intelligent medical visual voice call control method.
[0110] The beneficial effects provided by the present invention are as follows: First, a multimodal call request control method is adopted to significantly improve the emergency response speed and judgment accuracy, while reducing the patient's dependence on the operation of medical equipment and improving the self-service efficiency; at the same time, it supports the wide deployment of home scenarios, community medical care, and smart terminals, and the adaptation scenarios are not restricted. Secondly, through intelligent model pre-training, different intention information is matched with the patient's condition. In the case of limited medical resources or a large number of patients, medical resources are allocated fairly and orderly according to the severity of each patient's condition, ensuring efficient and orderly response to requests, and rationalizing and equitably dividing medical resources to provide patients with fair and reasonable medical treatment. Furthermore, through iterative training and learning of the patient's request intention and the patient's condition data by the intelligent model, the trend of the condition is monitored and predicted in real time, and a better and more predictable treatment plan is provided to the monitored subject in advance, completing the medical and health ecosystem.
[0111] Obviously, the above specific implementation cases are merely examples for illustrating the application of the present method, and are not intended to limit the implementation methods. A person skilled in the art can make other variations and modifications based on the above description to study other related issues. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
[0112] Those skilled in the art will understand that all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk, etc. Various media that can store program codes.
[0113] The embodiments of electronic devices and the like described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the embodiments. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0114] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention, rather than to limit them. Although the embodiments of the present invention have been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
[0116] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.
[0117] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. Intelligent medical visual voice call control method, characterized in that: The method comprises: S1: Real-time monitoring to obtain first data representing the life characteristics of the monitored subject, second data of the environment in which the monitored subject is located, and the expression intention of the monitored subject; S2: intelligently call a control center based on changes in the first data, the second data, and the expression intention of the monitored subject.
2. The method according to claim 1, characterized in that The expression forms of the monitored subject's expression intention include: voice, text, code, picture, and machine recognition language converted by button.
3. The method according to claim 2, characterized in that The intelligent call control center according to the changes of the first data and the second data and the expression intention of the monitored subject includes: The intelligent call control center is configured to set a first threshold value related to the first data according to the first data, monitor the change of the first data in real time and compare whether the first data is not less than the first threshold value, and if the first data is not less than the first threshold value, the intelligent call control center; the intelligent call control center according to the change of the second data, including setting a second threshold related to the second data, monitoring the change of the second data in real time and comparing whether the second data is not less than the second threshold, and if the second data is not less than the second threshold, the intelligent call control center; The intelligent call control center uses machine-recognized language based on voice, text, code / command, image, and button conversion, including: Voice is collected through a microphone, sampled into digital signals, an audio stream is generated, sound features are extracted, and a sound matrix is formed; the pronunciation probability distribution at different time points is output based on the sound features and converted into text; the text is directly collected; the code is parsed to extract the corresponding intention; the intention code corresponding to each button is set; the image area is bound to its corresponding label; the text, intention, and label data are preprocessed to perform intent recognition and semantic analysis, and are output in a unified format, which are packaged and sent together with the monitoring device number, the identity information of the monitored subject, the environmental location information of the monitored subject, and the call timestamp data.
4. Intelligent medical visual voice call control method, characterized in that: The method is used to respond to a call request from a monitored subject, comprising: receiving the call request and parsing the call request; Allocate corresponding medical resources according to the call request.
5. The method according to claim 4, characterized in that The parsing of the call request includes: Obtain the identity information of the monitored subject, the monitoring device number, the location information of the monitored subject, the timestamp of the request, and the request intent information.
6. The method according to claim 4, characterized in that The request intent information is marked with different priority levels, including: Collect the medical records and first request intention information of the monitored subject, perform intelligent model training, obtain the priority level of the first request intention information related to the medical records of the monitored subject, calibrate the first request call information according to the priority level, and obtain the second call request information.
7. The method according to claim 6, characterized in that Allocating medical resources corresponding to the call request according to the call request includes: Allocate corresponding medical resources based on the monitored subject's identity information, the timestamp of the request, and the second request intent information, including: Arrange the second request intent information in sequence according to the second request intent information and the timestamp of the request occurrence; Find the identity information of the monitored subject corresponding to the timestamp of the second request intention information request, and complete the arrangement order of the monitored subject's request for allocation of medical resources; Medical resources are allocated to the monitored subjects in sequence according to the order in which they request allocation of medical resources.
8. Intelligent medical visual voice call control system, characterized by: The system comprises: An acquisition module, configured to monitor and acquire in real time first data representing the vital characteristics of a monitored subject, second data representing the environment in which the monitored subject is located, and the expression intention of the monitored subject; The calling module is used to intelligently call the control center according to the changes of the first data and the second data and the expression intention of the monitored subject.
9. The system according to claim 8, characterized in that The expression forms of the monitored subject's expression intention include: voice, text, code, picture, and machine recognition language converted by button.
10. The system according to claim 9, characterized in that The call module includes: a first calling unit, configured to intelligently call a control center based on the first data, including setting a first threshold value related to the first data, monitoring changes in the first data in real time and comparing whether the first data is not less than the first threshold value, and intelligently calling the control center if the first data is not less than the first threshold value; a second calling unit, based on a change in the second data, intelligently calling the control center, including setting a second threshold related to the second data, monitoring the change in the second data in real time and comparing whether the second data is not less than the second threshold, and if the second data is not less than the second threshold, intelligently calling the control center; The third call unit is an intelligent call control center that uses machine-recognized language based on voice, text, code / command, image, and button conversion, including: The acquisition unit is configured to collect voice through a microphone, sample it into a digital signal, generate an audio stream, extract sound features, and form a sound matrix; output the pronunciation probability distribution at different time points based on the sound features, and convert it into text; directly collect the text; parse the code to extract the corresponding intent; set the intent code corresponding to each button; and bind the image area to its corresponding label; A data preprocessing unit, configured to preprocess the text, intent, and tag data, perform intent recognition and semantic analysis, and output the data in a unified format; The sending unit is used to package and send the data from the data pre-processing unit, the monitoring device number, the identity information of the monitored subject, the environmental location information of the monitored subject, and the call time.
Citation Information
Patent Citations
Intelligent ward calling method and system based on Internet of Things
CN118118600A
Priority ranking method, system, apparatus and program in response to patient calls
CN118658597A
Intelligent ward communication calling system based on Internet of Things
CN118969214A
Medical intercom system and method based on audio processing
CN118972716A