Pain assessment method and system based on visual big language model
By automatically analyzing children's facial and limb behaviors through a visual large language model, the accuracy and real-time limitations of existing child pain assessment methods are addressed, achieving highly accurate and consistent pain assessment, adapting to the complexity of the clinical environment, and providing explainable assessment results.
Patent Information
- Application Number
- CN202510651485.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-09-26
AI Technical Summary
Existing methods for assessing children's pain rely on manual subjective judgment, which lacks accuracy and real-time performance. It is difficult to accurately and objectively assess the degree of children's pain in a clinical setting, affecting the effectiveness of medical operations and the psychological state of children.
A pain assessment method based on a visual large language model is adopted. Through multi-perspective video acquisition, spatiotemporal sampling processing and multi-step guided instruction sequence, the visual large language model is used to automatically analyze children's facial expressions and body behaviors, generate pain level scores, and combine multimodal data for dynamic adjustment and standardized evaluation.
It achieves high-accuracy and consistency in pain assessment, exceeds expert-level judgment capabilities, monitors pain changes in real time, reduces human bias, improves the objectivity and efficiency of assessment, adapts to the complexity of the clinical environment, and provides interpretable assessment results.
Smart Images

Figure CN120708887A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pain assessment, and in particular to a pain assessment method and system based on a visual large language model. Background Art
[0002] Pain assessment in children is a key link in pediatric clinical care and pain management. Failure to correctly assess and treat pain may lead to excessive fear and resistance in children during diagnosis and treatment, and in severe cases, cause operation failure or medical complications. For example, during venipuncture, persistent pain will increase children's fear and resistance, affecting the success rate of puncture; in the long run, childhood pain experience may lead to anxiety and fear of medical treatment in the future. Therefore, there is an urgent need for an objective and accurate pain assessment method in clinical practice, so that medical staff can take timely intervention measures (such as psychological comfort, analgesic adjuvants or local anesthesia, etc.) according to the degree of pain to relieve children's pain.
[0003] The FLACC scale is widely used in clinical practice to assess pain in children. The FLACC scale observes and scores pain in infants and children based on five dimensions: facial expression (Face), leg movement (Legs), activity (Activity), crying (Cry), and soothing (Consolability). Each dimension is scored as 0, 1, or 2, with a total score of 0–10 for all five dimensions. This scale is simple and easy to use and is currently one of the gold standards for pediatric pain assessment. However, the FLACC scale requires observational judgment by experienced healthcare professionals and is somewhat subjective: different observers may have different interpretations of the same behavior. Furthermore, in a busy clinical environment, continuously monitoring and accurately scoring each child's subtle pain manifestations can be challenging. Furthermore, due to experience and empathy, assessors may sometimes tend to overestimate pain (e.g., anticipating pain during a puncture and defaulting to a score of 1), which can slightly deviate from the child's actual pain experience. In summary, existing pain assessment methods primarily rely on manual and subjective methods, leaving room for improvement in terms of accuracy and real-time objectivity. Summary of the Invention
[0004] In order to solve the above-mentioned defects, the present invention proposes a pain assessment method and system based on a visual large language model, providing an objective, non-contact child pain assessment method based on a visual large language model, which can automatically analyze the facial expressions and body behaviors of children in the video, accurately determine their pain level, and achieve an assessment effect that exceeds the level of senior pain doctors, thereby helping medical staff to understand the child's pain situation in a timely and accurate manner and take corresponding measures.
[0005] The technical solution adopted by the present invention is a pain assessment method based on a visual large language model, comprising the following steps:
[0006] S100, synchronously collecting video data of the target object's face and limbs;
[0007] S200, performing spatiotemporal sampling processing on the video data to obtain image data;
[0008] S300, generating a multi-step guidance instruction sequence based on preset pain assessment dimensions;
[0009] S400, inputting the image data and the instruction sequence into a visual large language model, and the visual large language model performs step-by-step reasoning according to a preset pain assessment dimension;
[0010] S500. The visual large language model generates scoring basis, scoring results and comprehensive pain level for each dimension in sequence according to the reasoning results of each dimension.
[0011] Preferably, the S100 further includes: obtaining the pain category of the target object;
[0012] The spatiotemporal sampling processing of the video data in S200 specifically includes: when the pain category of the target object is acute, extracting key frames from the video data; when the pain category of the target object is persistent, intercepting the video data at a fixed frame rate.
[0013] Preferably, the input method of inputting the image data into the visual large language model in S400 is selected based on the occlusion rate of the target object, and the input method includes sequential input, spliced input and simultaneous input.
[0014] Preferably, the S100 further includes: obtaining attribute parameters of the target object, the attribute parameters including at least one of age, disease classification or clinical environment type; the method for generating the preset evaluation dimension of the S300 includes: matching an applicable pain scoring system as the preset evaluation dimension according to the attribute parameters, the pain scoring system including at least one of the FLACC scale, the CHEOPS scale, or the NRPS scale.
[0015] Preferably, the S100 also includes obtaining at least one of the target object's audio data, physiological sensor data, historical medical data, or current medication status; the S200 also includes adjusting the weight of the audio data, physiological sensor data, historical medical data, or current medication status according to the data quality; the S400 also includes inputting the audio data, physiological sensor data, historical medical data, or current medication status after adjusting the weight into the visual large language model.
[0016] Preferably, the method for generating the multi-step boot instruction sequence includes:
[0017] S310, generating an instruction template, binding the model function to pain assessment, and requiring the model to output results according to the process of visual feature extraction, dimension score calculation, and comprehensive grade generation;
[0018] S320: Inject dynamic interactive instructions to decompose the assessment task into dimension-by-dimension inquiries according to the preset pain assessment dimensions, and guide the model to first output the observation basis and then give a score.
[0019] Preferably, the visual large language model is trained using multimodal training samples, and the method for constructing the multimodal training samples includes:
[0020] Acquire video data of multiple target objects from different perspectives;
[0021] Experts conduct multi-dimensional evaluation text on the video data;
[0022] Construct an alignment relationship between text and video data to form multimodal training samples.
[0023] Preferably, the training of the large visual language model adopts a parameter optimization strategy, which uses low-rank adaptation technology to insert into the model to train small-scale adaptation parameters; and uses DeepSpeed's ZeRO optimization algorithm to perform distributed training in a multi-GPU environment.
[0024] Preferably, the visual language model adopts a cross-validation strategy, and when the validation loss is the lowest, the corresponding model weight is used as the final model.
[0025] The present invention also discloses a pain assessment system based on a visual large language model, which can execute the pain assessment method according to any one of claims 1 to 9, and is characterized by comprising:
[0026] An acquisition module, comprising at least two image / video acquisition devices for acquiring video data of a target object from multiple perspectives;
[0027] An evaluation module that deploys a large visual language model and a multi-step guided instruction sequence;
[0028] The output module includes a display terminal, an early warning device and a data interface.
[0029] Compared with the prior art, the present invention has the following beneficial effects:
[0030] (1) Eliminate human bias and improve objectivity: Utilizing a large visual model to automatically interpret children's pain behavioral signals avoids interference from subjective factors and ensures a unified and objective scoring standard. Specifically, the present invention uses the standardized FLACC evaluation index, which allows the model's output to directly correspond to the universal human standard, enabling comparison and verification.
[0031] (2) Exceeding expert-level accuracy: By introducing the powerful cognitive and reasoning capabilities of the visual language model and combining it with large-scale real-world data training, the model's judgment of pain levels exceeds the consensus level of senior pain management experts, thus outperforming existing solutions in terms of accuracy.
[0032] (3) Real-time continuous monitoring: The method of the present invention can instantly analyze the video stream and realize real-time monitoring of the child's pain status. This means that throughout the medical procedure, the model can continuously assess pain changes and immediately issue an early warning if pain increases, assisting medical staff in timely intervention.
[0033] (4) Multi-view information fusion: By utilizing the images from multiple surveillance cameras, the present invention can obtain children’s facial expressions and behavioral data from different angles, overcome the information loss caused by single-angle occlusion, and ensure the integrity and robustness of the assessment.
[0034] (5) Enhanced interpretability and acceptability: The method of the present invention allows the model to provide a basis for judgment (such as a description of observed facial expressions and body movements) while giving a pain score, similar to the thinking process of an expert assessment. This chain thinking output makes the results more transparent, easier for medical staff to understand and trust, and also helps to improve the model later.
[0035] (6) Ease of deployment and integration: This method can be implemented using the hospital's existing surveillance camera equipment and general-purpose computing platforms, without requiring additional specialized sensors or invasive equipment, and without burdening the patient. Through software and hardware integration, the present invention can be easily integrated into clinical workflows, improving efficiency while not impacting normal operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The present invention is described in detail below with reference to the embodiments and accompanying drawings, in which:
[0037] Figure 1 It is a flow chart of a pain assessment method based on a visual large language model;
[0038] Figure 2 This is the overall flow chart of training and fine-tuning of the large visual language model;
[0039] Figure 3 This is a comparison chart of the consistency between different models / methods and the evaluation results of pain management experts;
[0040] Figure 4 This is a comparison chart of the model of the present invention and the average evaluation results of pain management experts;
[0041] Figure 5 It is a schematic diagram of the zero-sample thought chain prompt project;
[0042] Figure 6This is a comparison chart before and after fine-tuning the visual language model. DETAILED DESCRIPTION
[0043] To make the objectives, technical solutions, and advantages of the present invention more apparent, embodiments of the present invention will be described in further detail below with reference to the accompanying drawings. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar components or components having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.
[0044] In one embodiment, a pain assessment method based on a visual large language model is described. Figure 1 , including the following steps:
[0045] S100: Synchronously collect video data of the target object's face and limbs.
[0046] Specifically, in the data collection process, a three-dimensional collection system is constructed using multiple camera devices. For example, in scenarios such as hospital wards or rehabilitation centers, high-definition cameras are deployed at different angles, such as in front of and to the side of the target object, to ensure that the target object's facial expressions and body movements can be captured in all directions and without blind spots. At the same time, to ensure data synchronization, all cameras use high-precision clock synchronization technologies such as the Network Time Protocol (NTP) or the Precision Time Protocol (PTP) to ensure that the collected video data remains consistent in the time dimension. In addition, considering the security and privacy of the data, the collected video data will be encrypted using mature encryption algorithms such as AES (Advanced Encryption Standard) to prevent data leakage.
[0047] S200: Perform spatiotemporal sampling processing on the video data to obtain image data.
[0048] Specifically, spatiotemporal sampling processing includes two dimensions: temporal sampling and spatial sampling. For temporal sampling, keyframe extraction or fixed frame rate capture can be used. For spatial sampling, image segmentation techniques can be used to extract facial and limb regions from video frames. Semantic segmentation models from deep learning, such as U-Net, YOLO, and segment anything, can be used to train on a large amount of labeled facial and limb image data to accurately identify and segment the corresponding regions, thereby obtaining high-quality image data.
[0049] S300: Generate a multi-step guidance instruction sequence based on preset pain assessment dimensions.
[0050] Specifically, the preset pain assessment dimensions may include facial expressions, body movements, behavioral performance and other aspects. Taking facial expressions as an example, they can be subdivided into sub-dimensions such as the degree of frowning, the degree of eye closure, and the degree of mouth distortion; body movements may include the degree of body curling, the frequency of limb tremors, etc. For each dimension, a detailed multi-step guidance instruction sequence is designed. For example, for the facial expression dimension, the instruction sequence may be: "First, identify the position and shape of the eyebrows in the image to determine whether there is a frown; second, observe the degree of eye opening and closing to determine the degree of eye closure; finally, analyze the changes in the mouth contour and evaluate the distortion of the mouth." Through this step-by-step guidance approach, the visual language model can perform analysis in a more organized manner.
[0051] S400: Input the image data and instruction sequence into a visual large language model, and the visual large language model performs step-by-step reasoning according to a preset pain assessment dimension.
[0052] Specifically, after receiving the image data and instruction sequence, the visual large language model first uses its natural language processing capabilities to understand the meaning of the instruction sequence. For image data, it is necessary to use image-text cross-modal processing technology to convert the image information into a textual representation that the visual large language model can understand. Then, based on the instruction sequence and the preset pain assessment dimensions, the visual large language model analyzes and infers various aspects of the image in sequence. During the inference process, the visual large language model combines the knowledge and patterns it has learned from a large amount of text data to judge and infer the pain-related features reflected in the image.
[0053] S500. The visual large language model generates scoring basis, scoring results and comprehensive pain level for each dimension in sequence according to the reasoning results of each dimension.
[0054] Specifically, after completing the reasoning of each dimension, the visual large language model will generate a detailed scoring basis for each dimension according to the pre-set scoring rules. For example, in the facial expression dimension, if it is detected that the eyebrows are severely wrinkled, the eyes are closed, and the mouth is obviously distorted, the visual large language model will explain the degree of pain corresponding to these features, and thus give a score for that dimension. After obtaining the scores of each dimension, the comprehensive pain level is calculated to finally obtain a comprehensive pain level that fully reflects the pain level of the target object. The model is required to provide a basis for judgment (such as describing the observed facial expressions, body movements, etc.) while giving the pain score, which is similar to the thinking process during expert evaluation. The output of this chain thinking makes the results more transparent, easier for medical staff to understand and trust, and also helps to improve the model in the future.
[0055] Improved objectivity and consistency: Compared to manual scoring using the FLACC scale, this embodiment eliminates subjective bias by utilizing the visual language model. The model uses a unified standard for evaluation, resulting in more consistent results regardless of evaluator differences or fatigue. This means the system maintains consistent pain assessment criteria, day or night, ensuring high reliability.
[0056] Accuracy exceeding that of experts: In clinical tests, the evaluation accuracy of the method of this embodiment reached more than 86%, which is higher than the average level of human experts (about 84.7%). In particular, in certain dimensions (such as facial expressions, crying level, and soothing ability), the model's accuracy in identifying pain is about 5 percentage points higher than that of experts. This shows that the present invention not only exceeds the expert level, but is even more sensitive in distinguishing subtle pain and can detect details that the human eye may overlook. Compared with the best existing automatic algorithms, the accuracy of this embodiment is more significantly improved. For example, compared with deep learning methods such as CPANN, the accuracy of the model of the present invention on the same data is about 15% to 25% higher, achieving a qualitative leap. This high accuracy that exceeds the expert level is difficult to achieve with existing technologies.
[0057] Real-time and automatic, reducing the burden of manpower: The method of this embodiment can continuously monitor the pain status of children and output the results immediately. In contrast, manual assessment requires nurses to observe the entire process, which makes it difficult for one person to focus for a long time and it is difficult for one person to take care of multiple children. With this method, nurses can focus on processing the early warnings issued by the model and intervening in high-risk children, greatly improving work efficiency. Other existing automatic algorithms cannot achieve true real-time performance if they require complex preprocessing or offline analysis. This embodiment has obvious advantages in this regard.
[0058] Stronger robustness and adaptability: Through the use of multi-view cameras and the generalization ability of large models, this method can adapt to various complex situations in real medical scenarios, including occlusion, light changes, behavioral differences among children of different ages, etc. In comparison, some traditional algorithms work well on laboratory data, but are unstable in complex clinical environments. This method has been trained with real-world data and has shown good robustness. In addition, the method of this embodiment has a certain degree of universality and can be extended to similar behavioral assessment tasks without the need to develop new algorithms from scratch. Its adaptability is better than existing systems with single functions. The method of this embodiment can be implemented using the hospital's existing surveillance camera equipment and general computing platforms, without the need for additional specialized sensors or invasive equipment, and without increasing the burden on children.
[0059] Strong interpretability facilitates clinical acceptance: Unlike many black-box deep learning algorithms that simply output a single score or grade, the method of this invention provides a rationale for each score (e.g., "the child frowned, so the facial score was 1"). This interpretability allows medical staff to quickly understand the scoring basis and make judgments accordingly, increasing trust in AI results and helping to advance the clinical application of this technology.
[0060] In one embodiment, the step S100 further includes: obtaining the pain category of the target subject. The step S200 of performing spatiotemporal sampling processing on the video data specifically includes: extracting key frames from the video data when the pain category of the target subject is acute; and capturing the video data at a fixed frame rate when the pain category of the target subject is persistent.
[0061] In step S100, the pain category of the target object is obtained by integrating multi-source information. For example, the pain category annotated by the medical diagnosis is extracted from the electronic medical record, and the description of the patient or family member is analyzed using natural language processing technology. It can also be combined with the physiological indicators monitored by wearable devices (such as heart rate variability and skin electrical response) to comprehensively judge whether it is acute pain or persistent pain. In step S200, differentiated spatiotemporal sampling is implemented according to the pain category. In acute pain, a key frame extraction algorithm is used to quickly lock the key frame at the moment of pain onset by analyzing the motion and texture changes of the video frame; for persistent pain, video data is captured at a fixed frame rate to ensure the completeness of the record while reasonably controlling the amount of data.
[0062] The operating method of this embodiment achieves efficient data processing and accurate assessment. Key frame extraction for acute pain allows for rapid screening of key information, reducing the amount of data to be processed; fixed frame rate capture for persistent pain avoids data redundancy and improves data processing efficiency. Simultaneously, the two sampling methods match different pain characteristics, providing high-quality data for the visual language model and enhancing assessment accuracy. Furthermore, this reduces data storage and computing resource requirements, optimizes resource utilization, and can adapt to different pain types, expanding the scope of application of the assessment method.
[0063] In one embodiment, the input method of inputting the image data into the visual language model in S400 is selected based on the occlusion rate of the target object, and the input method includes sequential input, spliced input, and simultaneous input.
[0064] In step S400, computer vision algorithms analyze the occlusion rate of the target object (e.g., clothing obstruction, device interference) in real time. When the occlusion rate is low, a sequential or spliced input method is used; when the occlusion rate is high, a simultaneous input method is used to enhance feature complementarity. This strategy significantly improves assessment accuracy while optimizing computing resources, achieving a balance between accurate assessment and efficient computation.
[0065] In one embodiment, the S100 further includes: obtaining attribute parameters of the target object, wherein the attribute parameters include age, disease classification, or clinical environment type, etc.; the method for generating the preset evaluation dimension of the S300 includes: matching an applicable pain scoring system as the preset evaluation dimension according to the attribute parameters, wherein the pain scoring system includes a FLACC scale, a CHEOPS scale, or an NRPS scale, etc.
[0066] For example, infants and young children are matched with the FLACC scale (facial expressions, leg movements, and other dimensions), pediatric patients are matched with the CHEOPS scale (crying, facial expressions, and other dimensions), and adult patients are matched with the NRPS scale (numeric rating dimensions). Dimension weights can also be adjusted based on disease classification and clinical setting type (e.g., postoperative intensive care unit, outpatient clinic), for example, increasing the weight of body movement dimensions for postoperative patients.
[0067] This strategy significantly improves assessment accuracy, capturing the pain characteristics of different groups (e.g., behavioral manifestations in children, subjective experiences in adults) through a targeted scoring system. It also enhances clinical practicality. Automatic adaptation reduces the workload of medical staff in manually selecting scales, avoiding assessment bias caused by inappropriate scale selection. Furthermore, the scoring system can be dynamically updated according to the latest clinical guidelines, ensuring the scientific and timely nature of the assessment criteria.
[0068] In one embodiment, the S100 also includes obtaining at least one of the target object's audio data, physiological sensor data, historical medical data, or current medication status; the S200 also includes adjusting the weight of the audio data, physiological sensor data, historical medical data, or current medication status according to the data quality; the S400 also includes inputting the weighted audio data, physiological sensor data, historical medical data, or current medication status into the visual large language model.
[0069] In S100, the data collection dimension is expanded to simultaneously acquire audio data (such as screams), physiological sensor data (heart rate, blood oxygen), historical medical data (surgical records), and current medication status. In S200, the weight of each modality is dynamically adjusted through the data quality assessment model. For example, a high weight is assigned to physiological sensor data with stable signals, and audio data affected by environmental noise is subjected to noise reduction processing and then downgraded as needed. In S400, the weighted multimodal data is input into the visual language model, and feature complementarity is achieved through the cross-modal attention mechanism.
[0070] This strategy integrates multidimensional features: audio data captures the characteristics of patients' pain cries, physiological data reflects autonomic nervous system responses, historical data indicates pain tolerance, and medication usage information aids in determining drug effectiveness. Dynamic weight adjustment improves robustness, reducing the weight of audio in noisy environments and increasing the weight of other data in the event of sensor failure. The method of this embodiment significantly improves model inference accuracy.
[0071] In one embodiment, the method for generating the multi-step boot instruction sequence includes:
[0072] S310. Generate an instruction template, bind the model function to pain assessment, and require the model to output results according to the process of visual feature extraction, dimension score calculation, and comprehensive grade generation. Construct a three-level process template. The first step requires the model to extract visual features such as facial expressions (such as the degree of frowning) and body movements (such as the angle of forward leaning); the second step is to calculate the score of each dimension (such as the five dimensions of the FLACC scale) based on the extracted features; the third step is to combine the weights of each dimension to generate the final pain grade. The template embeds domain knowledge constraints (such as "a score of 2 is scored if the corners of the mouth are pulled down more than 3mm").
[0073] S320: Inject dynamic interactive instructions to break down the assessment task into dimension-by-dimension queries based on the preset pain assessment dimensions, guiding the model to first output observational evidence before providing a score. Break down the assessment task into a sequence of dimension-level queries, for example: "1. Observe the facial image and list the three most significant pain-related features; 2. Based on these features, provide a score and basis for the facial expression dimension; 3. Combine the body movement dimension scores to calculate the overall pain level." Each subtask forces the model to output observational evidence before providing a conclusion.
[0074] Instruction templates ensure that the model outputs step-by-step according to clinical logic, making the scoring basis traceable and the reasoning process transparent. Structured output can directly generate standardized assessment reports, reducing the cost of secondary interpretation by medical staff.
[0075] In one embodiment, a pain assessment system based on a visual large language model, capable of implementing the aforementioned pain assessment method, comprises an acquisition module, an evaluation module, and an output module. The acquisition module comprises at least two video acquisition devices for acquiring video data of a target object from multiple perspectives; the evaluation module deploys the visual large language model and a multi-step guidance instruction sequence; and the output module comprises a display terminal, an early warning device, and a data interface.
[0076] Specifically, the acquisition module deploys a multi-view camera array (such as the front, side, and top) to synchronously collect facial micro-expression and body movement data, and can also support the access of audio and physiological sensor data. The evaluation module integrates an image preprocessing engine (spatiotemporal sampling, feature extraction) and a visual language model to achieve step-by-step reasoning of pain dimensions through a multi-step guided instruction sequence. The display terminal of the output module presents the pain level and visual analysis report in real time; the early warning device triggers an audible and visual alarm when the pain exceeds the threshold; and the data interface supports seamless connection with the hospital information system.
[0077] By utilizing images from multiple surveillance cameras, the present invention can capture facial expressions and behavioral data from different angles of a child, overcoming information loss caused by occlusion from a single angle and ensuring the integrity and robustness of the assessment. The system utilizes existing standard camera equipment for non-contact data acquisition and achieves objective assessment through intelligent analysis. The entire process can be automatically and efficiently executed in a clinical environment, significantly improving the quality and efficiency of child pain assessments. Through software and hardware integration, the present invention can be easily integrated into clinical workflows, improving efficiency while not affecting normal operations.
[0078] This system can be implemented using only the hospital's existing video surveillance equipment, eliminating the need for attaching any sensors to children and making it completely non-invasive. Compared to solutions that require wearable devices (such as electrocardiogram (ECG) monitors) or dedicated cameras, this system is low-cost, easy to deploy, and friendly to the hospital environment. Furthermore, the system's software can be integrated into existing hospital information systems, with assessment results automatically entered into electronic medical records, achieving a high level of informatization. While many existing studies remain at the laboratory prototype stage, without considering product integration, this system offers significant practical advantages.
[0079] Overall, this system outperforms existing best-in-class solutions in terms of accuracy, robustness, interpretability, and ease of use. It not only improves the quality and efficiency of assessments, but also has the potential to reduce medical errors and improve the patient experience.
[0080] In a specific embodiment, a method and system for assessing pain in children based on a visual large language model includes the following steps and modules:
[0081] First, multiple video capture devices (such as surveillance cameras) are deployed at the medical procedure site to capture real-time images of the child's facial expressions and body movements from different angles. The captured video frames are then fed into a pre-trained and optimized visual language model for analysis. This model, based on a large multimodal model architecture that combines computer vision and natural language processing capabilities, can understand image content and generate text output. Using a chain-of-thought strategy, the model is guided to sequentially reason and assess the child's pain expression along the five dimensions of the FLACC scale: facial expression, leg movement, overall activity, crying intensity, and soothingness. For each dimension, the model first extracts relevant features from the image (such as whether the face is frowning or the legs are curled up), provides corresponding judgment criteria, and then outputs a score for that dimension. After evaluating all five indicators, the model combines the scores from each dimension to generate a quantitative score or grade corresponding to the child's pain intensity at that moment (e.g., a pain score of 0-10 or a scale of "no pain, mild, moderate, severe"). Finally, the assessment results are stored in the system database and displayed on the medical staff's terminal, triggering alarms or prompting nursing interventions when necessary. The entire technical solution is centered on a pre-trained large-scale visual language model. Through specialized dataset fine-tuning and prompt engineering, it has expert-like pain assessment capabilities. This solution utilizes existing standard cameras for contactless data acquisition and achieves objective assessment through intelligent analysis. The entire process can be run automatically and efficiently in a clinical setting, significantly improving the quality and efficiency of pediatric pain assessment.
[0082] The following is a more specific embodiment.
[0083] (1) Overall system architecture: The system of this embodiment mainly consists of a data acquisition module, an AI evaluation model module and a result output module. The data acquisition module includes multiple cameras installed in the consulting room or operating room, which are used to capture real-time videos of child patients during medical operations from different perspectives. Typically, one camera is aimed at the top of the child's face, and the other camera shoots the whole body from the side to ensure that comprehensive behavioral information is obtained. The AI evaluation model module is deployed on a computing device (such as a GPU server), and its core is the trained visual large language model PA-VLLF (Pain Assessment Vision Large Language Framework). The result output module includes a display screen or mobile terminal, as well as an interface with the hospital information system, which is used to convey the pain score generated by the model to medical staff or record and archive it.
[0084] (2) Model selection and capabilities: The PA-VLLF model is built based on OpenAI's ChatGPT-4o and Alibaba's open-source Qwen2-VL multimodal large language model (Chinese name "Tongyi Qianwen 2-VL"). The Qwen2-VL model has visual and language processing capabilities, and can encode image content and generate corresponding text descriptions. The model has different scale versions, such as the 7B model with approximately 7 billion parameters and the 72B model with 72 billion parameters. The two are similar in architecture but have different performance and resource requirements. The present invention uses the powerful image understanding and reasoning capabilities of such large models to enable the model to act as an "expert assistant for pain assessment."
[0085] (3) Chain-of-Thought Prompt Engineering: To enable the visual language model to be systematically evaluated according to the FLACC metric, this embodiment designs a prompt word scheme, introducing a chain-of-thought prompt and reflection mechanism to guide the model to reason step by step like a human expert. The prompt word consists of two parts: system prompt and user prompt.
[0086] System prompts: This section prompts the model to act as a pain assessment assistant and specifies the format and reasoning for responses. For example, system prompts instruct the model to carefully observe and describe the basis for each response before assigning a score, encouraging self-reflection to ensure the reliability of the conclusion. System prompts effectively establish "thinking rules" for the model, ensuring that the model understands the need to complete tasks in a chain of steps rather than simply providing a conclusion.
[0087] User prompts: This section combines specific tasks and asks the model to evaluate each item based on the FLACC scale. The user prompt breaks down the complex pain assessment task into five subtasks, asking the model one by one about the child's "facial expression", "leg activity", "activity status", "crying performance" and "soothing ability". For example, for "facial expression", the prompt may describe: "Please judge whether the child's face shows signs of pain, such as frowning, closing eyes or trembling jaw, and score accordingly (0 = relaxed or smiling, 1 = occasional painful expression, 2 = continuous obvious painful expression)." After receiving this prompt, the model will first analyze the image content, for example, answering: "The child's brows are furrowed, there are tears in the corners of the eyes, and the expression is painful, so the facial expression shows pain." The model then gives a score for the item based on the analysis, for example: "Facial score = 2". In a similar manner, user prompts continue to instruct the model to assess "legs": for example, noting whether the child's legs are curled or kicking; assessing "activity": observing whether the child is writhing or attempting to break free; assessing "crying": determining whether the child cries or moans; and assessing "soothingness": observing whether the child calms down when soothed. For each prompt, the model first outputs the observed evidence and then the score. Through this chained questioning and answering process, the model effectively executes a step-by-step FLACC evaluation process. The entire prompting process ensures that the model's output follows a fixed format and sequence. For example, each dimension first states the evidence and then presents the "score = X" result. This design not only allows the model to systematically consider information from each dimension, avoiding omissions, but also improves the interpretability and consistency of the results. In actual testing, this chained prompting significantly improved evaluation accuracy compared to simple prompts that simply output a total score.
[0088] (4) Multi-view video processing: In practical applications, images from different camera perspectives can complement each other and improve the comprehensiveness of the evaluation. The present invention supports fusion analysis of multi-view data. The implementation method includes: inputting key video frames from multiple perspectives into the model for evaluation in sequence, or using the context memory function of the model to connect the images and analysis of each perspective into a conversation process. For example, the system can first input the image captured by the front camera and let the model perform a FLACC chain evaluation for that perspective; then input the image of the side camera and prompt the model to analyze again or verify the previous conclusion in combination with the new perspective. Because the model has conversation context memory, it can adjust the score based on the information of the front and back perspectives, thereby obtaining more accurate results using multi-view data. Another alternative implementation is to splice or input the multi-view images into the model at the same time (provided that the model architecture allows multiple image inputs), so that the model can take into account multiple picture features in a single answer. In either case, the present invention overcomes the problem of occlusion or field of view limitation that may occur with a single camera through multi-view fusion. For example, when the front perspective fails to observe the child's lower limb movements due to obstruction by family members, the image from the side perspective can provide effective information, so that the model can still correctly evaluate the "leg activity" indicator. Experimental data also show that multi-perspective input can improve the consistency and stability of evaluation results.
[0089] (5) Model training and fine-tuning: see Figure 2 To adapt the visual-language model to the task of pediatric pain assessment, we performed supervised fine-tuning on a large amount of labeled data. We constructed a new Clinical Pain Assessment (CPA) dataset, collected from real-world hospital settings. This dataset contains 1,248 video clips of 104 children undergoing procedures such as venipuncture, each lasting from a few seconds to over ten seconds. Each video clip was captured simultaneously by two cameras at different angles, covering both facial and full-body views. For each video, we invited five experienced pain management experts to independently score the video using the FLACC scale. The resulting five sets of scores served as the foundational annotations for model training. Because individual experts may provide slightly different scores, we used the most frequently occurring score among the five experts as the "ground truth" (equivalent to the consensus of the expert majority vote, i.e., the maximum agreement score). We then calculated the mean and standard deviation of all expert scores to assess the bias of the model output. During the fine-tuning data preparation process, we converted each video clip and its expert ratings into a multimodal dialogue format that the model can learn from. The specific approach is to construct a sample that interacts with the model, where the input includes a prompt describing the task (simulating the content of the system prompts and user prompts mentioned above) and a video frame image, and the output is the full evaluation process text generated by the expected model (including the basis description and correct score for each FLACC dimension). In order to obtain high-quality output text, we combined the scores given by experts and medical common sense to compile the corresponding ideal answers, see Figure 6 For example, if experts unanimously agree that the child in a video has a total score of 5 (where Face = 1, Legs = 1, Activity = 1, Cry = 2, and Consolability = 0), we write a description of what the model should output based on the video content (e.g., "The face has a brief frown, score 1; the legs are slightly curled, score 1;...", etc.). Such question-answer pairs serve as training samples, guiding the model to learn how to analyze the image step by step and produce a score agreed upon by the experts. When generating these training dialogues on a large scale, we referenced the format specifications of multimodal datasets such as Alpaca to ensure that each data entry contains clear instructions and expected answers. After preparing the training corpus, we performed supervised fine-tuning (SFT) on the Qwen2-VL model. During training, we employed an efficient parameter optimization strategy: Low-rank adaptation (LoRA) was inserted into the model to train only small-scale adaptation parameters, reducing the memory and data requirements of large-scale model fine-tuning; DeepSpeed's ZeRO optimization algorithm was used for distributed training in a multi-GPU environment to accelerate model convergence.
[0090] In this example, the model was fine-tuned and trained on two NVIDIA A800 GPUs. For the 7B-parameter Qwen2-VL model, each round of fine-tuning (on the cross-validation partitioned dataset) took approximately 40 minutes to complete. For the 72B-parameter model, due to its larger scale, each round of fine-tuning took approximately 3 hours.
[0091] In order to fully evaluate the generalization performance of the model, a 7-fold cross-validation strategy was adopted: the dataset was divided into 7 parts according to the acquisition time, 6 parts were selected as training sets and 1 part was selected as validation sets in turn, and the training and evaluation were repeated. We adjusted the hyperparameters such as optimizer and learning rate for the 7B and 72B models respectively to achieve the best results, and monitored the validation set loss curve during training. When the validation loss is the lowest, we save the corresponding model weights (merging the LoRA adaptation module into the main model) as the final model. The fine-tuned models are named PA-VLLF-Qwen-7B and PA-VLLF-Qwen-72B. In addition, the technical solution of the present invention is also applicable to large closed-source models: without fine-tuning models such as ChatGPT, we can directly send carefully designed chain prompts and images through API calls to let ChatGPT perform pain assessment. This zero-sample prompt strategy has also achieved comparable results in experiments, see Figure 5 Therefore, the method of the present invention includes both fine-tuning open source models to implement dedicated evaluators and utilizing prompt engineering to perform pain assessment on existing general models.
[0092] (6) Reasoning and Output: After training, the PA-VLLF model can be used for actual evaluation. When in use, the system obtains a real-time video stream from the camera and extracts key frames or captures images at a fixed frame rate to input into the model. During the reasoning phase, the model follows the previously defined prompts and gradually outputs various criteria and scores. Typical output examples:
[0093] Face: The child's face is relaxed most of the time, but occasionally frowns. Score = 1
[0094] Legs: Legs are naturally relaxed, not tense or curled up. Score = 0
[0095] Activity: The child twists slightly but generally cooperates and does not struggle violently. Score = 1
[0096] Crying: The child cries briefly and then becomes quiet, with only soft sobs. Score = 1
[0097] Soothing ability: Calms down quickly when soothed by parents. Rating = 0
[0098] Total score: 1+0+1+1+0=3 points (mild pain)
[0099] The above output clearly lists the model's judgment and score for each dimension, as well as the final calculated total score. In this way, medical staff do not need to watch the entire video, but only need to read the analysis report given by the model to understand the child's pain manifestations and degree. If the pain score determined by the model is high (for example, close to the full score of 10, indicating severe pain), the system can automatically issue an early warning, prompting doctors and nurses to take quick measures (such as stopping the operation, comforting or additional analgesic interventions). The entire reasoning and output process is usually completed within a few seconds, which can meet the clinical real-time requirements. It is worth mentioning that since the output format of the model corresponds to the manual FLACC evaluation form, the hospital can easily record the results in the electronic medical record or use it for subsequent pain management data analysis.
[0100] In summary, the technical solution of this embodiment realizes the automated, expert-level assessment of children's pain through the organic combination of "multimodal large model + chain thinking prompts + dedicated data training". The model input and output are compatible with existing clinical processes: the input is ordinary video images, and the output is standard FLACC scale items and total scores. The detailed implementation steps, model structure and training methods ensure the effectiveness and robustness of this solution in complex real-world scenarios. It can be seen from the above explanation that the present invention provides not only an algorithm model, but also a complete system solution that can guide the implementation of the entire process from data acquisition to result application, and has strong practical value.
[0101] The advantages of this embodiment include:
[0102] (1) Introducing large models for pain assessment: We propose a solution to apply large-scale visual-language pre-training models to the task of children's pain assessment, which fully utilizes the powerful image understanding and reasoning capabilities of large visual-language models and achieves assessment accuracy that is difficult to achieve in traditional computer vision fields.
[0103] (2) FLACC-based chained thinking prompting strategy: A chained thinking prompting project was designed to guide the model to analyze the image step by step and give a score according to the five dimensions of the FLACC scale. The model outputs the observation basis and scoring results at each step, making the evaluation process transparent and traceable. This prompting strategy effectively improves the reliability and interpretability of the model evaluation.
[0104] (3) Multi-view video fusion assessment: This system supports the use of multi-angle camera data to jointly assess pain. By sequentially or concurrently introducing images from different perspectives into the model, the model improves its adaptability to complex scenarios. Multi-view fusion analysis improves the model's ability to handle occlusion and interference, ensuring consistent and accurate pain assessment.
[0105] (4) Expert-annotated training data and model fine-tuning: A dataset containing detailed expert rating annotations was constructed, and supervised fine-tuning techniques were used to adapt the large model to the specific task. The model learned the rules of expert evaluation, making the output highly consistent with clinical experts. This method of dataset construction and large model fine-tuning for specific medical evaluation tasks ensured that the model performance reached a practical level.
[0106] (5) Formatting and application of output results: The model output is directly aligned with the FLACC standard and includes scores for each dimension and explanations of their basis, making it easy to integrate into existing medical records and processes. This formatted and interpretable output improves the acceptability of AI assessment results in clinical practice.
[0107] (6) Non-contact real-time assessment system: The entire solution uses non-invasive video acquisition to detect pain, which is more suitable for children than solutions that require wearing sensors. Real-time automated assessment reduces the workload of nurses and the risk of human omissions.
[0108] In one embodiment, the feasibility and superior performance of the present invention have been verified through rigorous experiments. During the research process, we conducted in-depth tests on the CPA clinical dataset constructed above to evaluate the accuracy and stability of the model. Using a 7-fold cross-validation method, the dataset was divided into training and test sets, and the experiments were repeated multiple times to verify the reliability of the model effect. Key experimental results are as follows:
[0109] (1) The overall accuracy rate exceeds the expert level: see Figure 3Using the Max Agreement metric, which uses the majority opinion of five experts as the standard, the PA-VLLF model (based on ChatGPT) achieved an overall accuracy of 86.36% on the test set. A PA-VLLF fine-tuned from the Qwen2-VL-72B model achieved an accuracy of 84.97%. In comparison, the average accuracy of five pain management experts on the same dataset was approximately 84.69%. This means that the model's judgment matches the expert consensus better than the average level of experts, surpassing expert-level performance. The model also significantly outperformed other machine learning baselines we implemented: Traditional methods such as the SVM classifier, ResNet-50 neural network, and the CPANN model achieved significantly lower accuracy on the same data, while PA-VLLF achieved significantly higher accuracy by 17.48% to 26.22%, respectively, a significant improvement.
[0110] (2) Z-score evaluation also performed well: The Z-score indicator was also used to evaluate the closeness of the model score to the average expert score. If the model score is within one standard deviation of the average score of the five experts, the evaluation is considered reasonable. The results show that the accuracy of PA-VLLF (ChatGPT version) under the Z-score indicator is about 88.81%, and PA-VLLF (Qwen2-VL-72B version) is about 83.92%, which are significantly ahead of the baseline model (an increase of about 11.9% to 13.7%). This shows that the pain score given by the model is very close to the comprehensive judgment of the experts, with little deviation.
[0111] (3) Accuracy analysis of indicators in each dimension: see Figure 4The model achieved high accuracy across all five FLACC dimensions. In the "Facial Expression" dimension, PA-VLLF (ChatGPT) achieved an accuracy rate of approximately 89.16%, in the "Degree of Crying" dimension, as high as 98.95%, and in the "Soothing" dimension, approximately 93.71%. These accuracy rates were 4.97% to 7.41% higher than the average accuracy of individual expert assessments. The AI model was particularly accurate and reliable compared to human assessments of whether a child's crying was painless and whether they could be soothed. This may be because the model's judgments were strictly based on observable behavior, while experts are sometimes influenced by subjective impressions. For example, when a child is experiencing no pain but is undergoing a puncture, experts tend to assign a score of 1 to avoid missing any pain, while the model more objectively assigns a score of 0. Statistics show that in the overall score assessment, the PA-VLLF model rated "no pain" (0) approximately 10% higher than the expert average. This suggests that the model is more confident in judging "no pain," reducing the likelihood of humans over-scoring due to empathy, thereby reducing the false positive rate. This is demonstrated by the confusion matrix analysis: the model is more confident than experts in distinguishing between 0 (no pain) and 1 (mild pain), significantly reducing the misclassification of 1 as a score when it should be 0. This more refined distinction makes the assessment results more consistent with the child's actual condition and demonstrates the objectivity and fairness of AI-based assessments.
[0112] (4) Model scale and performance: The performance of models of different scales was compared. The accuracy of the fine-tuned Qwen2-VL-7B model in the maximum consistency index also exceeded 80% (about 82%). Although it is slightly lower than the 72B large model, it is still much higher than the traditional method. This shows that even with a smaller model, good results can be achieved, which is conducive to actual deployment. The 72B model is slightly better than the 7B model in all indicators, especially more robust in subtle behavior recognition. The two are consistent in trend, which proves the scalability of the method of the present invention: larger models and more data can further improve performance. On the other hand, the prompt-based ChatGPT model achieved the highest accuracy of more than 86% without training, demonstrating the versatility of the solution of the present invention.
[0113] (5) Simulation and actual use: In order to test the adaptability of the model in the actual environment, we deployed the PA-VLLF system in the hospital experimental ward for trial operation. Without disturbing normal nursing, the system monitored multiple puncture operations in real time, and the evaluation results were basically consistent with the FLACC scores recorded by the nurses on duty. In several cases with large differences in evaluation, subsequent analysis found that the model identified some details that the nurses did not notice (such as the child's brief frown), resulting in slightly higher scores, which suggests that the model can help reduce missed evaluations. Overall, the trial operation proved that the system of the present invention has good practical feasibility, the output results can be used as a reference for medical staff, and provide a second judgment in terms of objectivity, which is conducive to improving the reliability of pain assessment.
[0114] Through the above experiments and preliminary clinical use, the effectiveness of the method of the present invention has been verified. The model can consistently assess pain in real-world video data, achieving accuracy that matches and partially exceeds that of human experts. Furthermore, the output interpretation information was also found helpful by clinical staff. These results fully demonstrate the feasibility and superiority of the present invention and lay the foundation for its wider application.
[0115] The technical solution provided by the present invention has good scalability and flexibility, and can be modified in many aspects or applied to other related purposes as needed.
[0116] (1) Model replacement and upgrade: The present invention is not limited to the use of the Qwen2-VL model. Any large-scale model with visual and language multimodal capabilities can be used to achieve the purpose of the present invention. For example, other open source multimodal large models (such as Meta's LLaVA model series or OpenAI's GPT-4V model) can be used to replace Qwen2-VL as the core analysis engine. With the development of large model technology, new higher-performance models can also be integrated into this solution to improve the evaluation effect. For example, models with larger parameter scales in the future or large models optimized for the medical field can be applied to the present invention to make pain recognition more accurate. Accordingly, the prompt strategy can also be slightly adjusted to adapt to the interaction format of different models. In addition, in some embodiments, the model can even be deployed instead of locally, and cloud services can be used: for example, the camera image is sent to the ChatGPT-4 model in the cloud through a secure interface, and the evaluation results are returned by the cloud. This cloud deployment as an alternative can reduce local computing pressure, although data transmission delay and privacy security need to be considered.
[0117] (2) Adjustment of prompt engineering: Chain thinking prompts play an important role in the present invention. In different application scenarios, the prompt strategy can be flexibly modified. For example, for children who are older or able to communicate, a language inquiry link can be added to the prompt (such as asking the model to analyze the child's speech or shouting). For other types of medical operation pain assessment, special dimension prompts can be added, such as adding assessment prompts for posture changes or sleeping and eating status in postoperative pain monitoring. You can even change the scoring criteria and replace FLACC with other scoring systems (such as the CHEOPS Children's Postoperative Pain Scale or the Adult NRPS Pain Score) to let the model output results according to the new criteria. The language of the prompt can also be localized into English or other languages to adapt to the clinical environment of different countries. The core of the method of the present invention lies in chain multi-step reasoning, and this idea can be generalized and is not limited to FLACC.
[0118] (3) Multimodal data fusion: Although the present invention mainly uses video image information to assess pain, its design concept allows the introduction of other data sources as auxiliary, thus forming a richer multimodal assessment scheme. Alternative solutions include: adding audio input, allowing the system to analyze the child's crying intensity, pitch changes and other sound signals to more accurately assess the "crying degree" dimension; introducing physiological sensor data, such as heart rate, blood pressure, blood oxygen and other vital signs, as additional indicators to prompt the model reference, because pain is often accompanied by physiological changes, this information can improve the credibility of the assessment. It can even be combined with the child's historical medical data or current medication status, and the model can comprehensively judge the possibility and degree of pain. These additional modalities can be used as extensions to further enhance system performance.
[0119] (4) Hardware deployment flexibility: In some resource-constrained situations, the method steps of the present invention can be partially adjusted to ensure usability. For example, if only a single camera can be installed on site, the present invention can also rely solely on single-view video for evaluation, still covering the five FLACC indicators (but the occluded information may be lacking). This single-camera simplified solution, as a sub-solution of the present invention, is suitable for small medical points with limited equipment. Similarly, if the GPU computing power is limited, you can choose to use a simplified model (such as the 7B model fine-tuned by the present invention) for inference in exchange for faster response speed and lower hardware requirements. Although the performance may be slightly reduced, it can still provide valuable evaluation results in most cases. In addition, the system's video frame sampling rate can also be adjusted according to demand: in scenarios where continuous monitoring is not required, the frame rate can be reduced or even analyzed based on static images to reduce the amount of calculation. In short, the present invention can flexibly adapt to different hardware environments and usage requirements to ensure its practicality.
[0120] (5) Other application areas: The method of the present invention is not limited to children's pain assessment. Its core idea (i.e., chain analysis using a large visual model) can be extended to many related fields. For example, when the method is applied to adult pain assessment, new prompts and scoring rules can be defined based on the adult's facial expressions and limb reactions, which can be used to detect postoperative pain, chronic pain attacks, etc.; when applied to neonatal or infant pain monitoring, model prompts can be customized in combination with the unique behavioral characteristics of neonates (such as special facial muscle movements and crying patterns); when extended to the field of emotion and expression recognition, such as assessing the patient's anxiety level, sedation level, etc., by simply changing the content of the chain prompt, the model can perform similar item-by-item judgments (such as eye contact, speech, and movement) to output a comprehensive emotion score. In addition, in the veterinary field, the idea of the present invention can also be used for animal pain recognition (the model needs to be trained to recognize the facial and behavioral characteristics of animals). In non-medical scenarios, this method can also be used for emotional computing in human-computer interaction, abnormal behavior analysis in security monitoring, etc. By replacing the corresponding evaluation criteria and training data, the chain large model evaluation framework of the present invention can become a universal intelligent behavior analysis toolbox with broad application prospects.
[0121] In summary, the technical solution of the present invention is highly adaptable and versatile. Various design variations revolve around the same inventive concept, namely, utilizing a multimodal artificial intelligence model for step-by-step reasoning and evaluation. Therefore, whether it involves replacing models, adjusting prompts, adding or removing modules, or applying the technology to other evaluation tasks, these variations fall within the scope of the present invention and are covered by the patent protection requirements.
[0122] In this specification, the use of terms such as "Embodiment 1," "this embodiment," and "in one embodiment" indicates that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in the invention or at least one embodiment or example of the invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example; furthermore, the specific features, structures, materials, or characteristics described may be appropriately combined in any one or more embodiments or examples.
[0123] In the description of this specification, the terms "connect," "install," "fix," "dispose," and "have" are to be understood in a broad sense. For example, "connect" can mean a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, or it can be internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in this application based on specific circumstances.
[0124] In the description of this specification, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises", "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.
[0125] The above description of the embodiments is to facilitate ordinary technicians in this technical field to understand and apply the technology of this case. People familiar with the technology in this field can obviously make various modifications to these examples easily and apply the general principles described here to other embodiments without having to go through creative work. Therefore, this case is not limited to the above embodiments. Modifications to the following situations should all be within the scope of protection of this case: ① A new technical solution implemented based on the technical solution of the present invention and combined with existing common knowledge, the technical effect produced by the new technical solution does not exceed the technical effect of the present invention; ② The equivalent replacement of some features of the technical solution of the present invention with the known technology, the technical effect produced is the same as the technical effect of the present invention; ③ The technical solution of the present invention is expandable, and the substantive content of the expanded technical solution does not exceed the technical solution of the present invention; ④ The equivalent transformation made by the content of the description and drawings of the present invention is directly or indirectly applied to other related technical fields.
Claims
1. A pain assessment method based on a visual large language model, characterized in that: The following steps are involved: S100, synchronously collecting video data of the target object's face and limbs; S200, performing spatiotemporal sampling processing on the video data to obtain image data; S300, generating a multi-step guidance instruction sequence based on preset pain assessment dimensions; S400, inputting the image data and the instruction sequence into a visual large language model, and the visual large language model performs step-by-step reasoning according to a preset pain assessment dimension; S500. The visual large language model generates scoring basis, scoring results and comprehensive pain level for each dimension in sequence according to the reasoning results of each dimension.
2. The pain assessment method according to claim 1, wherein: The S100 further includes: obtaining the pain category of the target object; The spatiotemporal sampling processing of the video data in S200 specifically includes: when the pain category of the target object is acute, extracting key frames from the video data; when the pain category of the target object is persistent, intercepting the video data at a fixed frame rate.
3. The pain assessment method according to claim 1, wherein: The input method of inputting the image data into the visual large language model in S400 is selected based on the occlusion rate of the target object, and the input method includes sequential input, spliced input and simultaneous input.
4. The pain assessment method according to claim 1, wherein: The S100 also includes: obtaining attribute parameters of the target object, the attribute parameters including at least one of age, disease classification, or clinical environment type; the method for generating the preset evaluation dimension of the S300 includes: matching an applicable pain scoring system as the preset evaluation dimension according to the attribute parameters, the pain scoring system including at least one of the FLACC scale, the CHEOPS scale, or the NRPS scale.
5. The pain assessment method according to claim 1, wherein: The S100 also includes obtaining at least one of the target object's audio data, physiological sensor data, historical medical data, or current medication status; the S200 also includes adjusting the weight of the audio data, physiological sensor data, historical medical data, or current medication status according to the data quality; the S400 also includes inputting the audio data, physiological sensor data, historical medical data, or current medication status after adjusting the weight into the visual large language model.
6. The pain assessment method according to claim 1, wherein: The method for generating the multi-step boot instruction sequence includes: S310, generating an instruction template, binding the model function to pain assessment, and requiring the model to output results according to the process of visual feature extraction, dimension score calculation, and comprehensive grade generation; S320: Inject dynamic interactive instructions to decompose the assessment task into dimension-by-dimension inquiries according to the preset pain assessment dimensions, and guide the model to first output the observation basis and then give a score.
7. The pain assessment method according to claim 1, wherein: The visual language model is trained using multimodal training samples, and the method for constructing the multimodal training samples includes: Acquire video data of multiple target objects from different perspectives; Experts conduct multi-dimensional evaluation text on the video data; Construct an alignment relationship between text and video data to form multimodal training samples.
8. The pain assessment method according to claim 7, wherein: The training of the large visual language model adopts a parameter optimization strategy. The parameter optimization strategy uses low-rank adaptation technology to insert into the model to train small-scale adaptation parameters; and uses DeepSpeed's ZeRO optimization algorithm to perform distributed training in a multi-GPU environment.
9. The pain assessment method according to claim 8, characterized in that: The visual language model adopts a cross-validation strategy. When the validation loss is the lowest, the corresponding model weight is used as the final model.
10. A pain assessment system based on a visual large language model, capable of executing the pain assessment method according to any one of claims 1 to 9, characterized in that: include: An acquisition module, comprising at least two video acquisition devices, for acquiring video data of a target object from multiple perspectives; An evaluation module that deploys a large visual language model and a multi-step guided instruction sequence; The output module includes a display terminal, an early warning device and a data interface.
Citation Information
Cited By
Model training method, video generation method, electronic equipment and storage medium
CN120953453A
Pain assessment model construction method and non-contact continuous pain assessment method
CN121439247A
A pain assessment model construction method and a non-contact continuous pain assessment method
CN121439247B