Field first-aid data processing method and device based on multi-modal intelligent algorithm and medium
By adopting multimodal intelligent algorithms in field first aid scenarios, fusing wound images and voice data, and using multimodal large models and retrieval enhancement generation methods, the problem that a single data processing technology cannot integrate multiple data sources is solved, and efficient and accurate first aid data processing and independent decision-making are achieved.
Patent Information
- Application Number
- CN202510061100.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
AI Technical Summary
In the field first aid scenario, a single natural language processing technology or computer vision technology cannot integrate multiple data sources at the same time, resulting in the inability to fully understand the on-site situation and the inability to fully respond quickly and make accurate decisions.
The data processing method based on multimodal intelligent algorithm is adopted, and the wound image and wound description speech data are obtained, and the wound classification model and large language model are used for data analysis and information fusion, combining the multimodal large model and the retrieval enhancement generation method to output the processing results.
It realizes multi-source data fusion and independent decision-making for field first aid data, improves the efficiency and accuracy of first aid, and provides immediate diagnostic advice without waiting for the doctor to go online in real time.
Smart Images

Figure CN119993455A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and more specifically to a field emergency data processing method, device and medium based on a multimodal intelligent algorithm. Background Art
[0002] In the wilderness first aid scenario, single natural language processing technology or computer vision technology has significant limitations. Because they can only process one type of data (such as text or images), and cannot simultaneously integrate multiple data sources to fully understand the situation on the scene. For example, although natural language processing technology can parse the descriptions of patients or witnesses, it cannot directly identify and analyze the patient's physical condition or the on-site environment; and although computer vision technology can capture and identify information in images, it may not be able to accurately judge complex environments and symptoms. Therefore, this single data processing method cannot achieve sufficiently rapid response and accurate decision-making in an environment such as wilderness first aid that requires quick and accurate judgment.
[0003] Therefore, how to provide a field emergency data processing method based on a multimodal intelligent algorithm is a problem that technical personnel in this field urgently need to solve. Summary of the invention
[0004] In view of this, the present invention provides a field first aid data processing method, device and medium based on a multimodal intelligent algorithm, which can integrate multiple data sources and have autonomous decision-making capabilities, and is of great significance for improving the efficiency and accuracy of field first aid.
[0005] In order to achieve the above object, the present invention adopts the following technical solution:
[0006] The method for processing field emergency data based on multimodal intelligent algorithm includes:
[0007] Obtain field first aid data, including wound images and voice data describing the injury;
[0008] The wound image is analyzed through the wound classification model to determine the wound image classification result, and the voice data describing the injury is converted into text and the text is corrected through the large language model;
[0009] Compare the corrected text with the wound image classification results, and perform information fusion when the results are consistent;
[0010] The fused information is used as the input of the multimodal large model, and the retrieval enhancement generation method is integrated to accurately output the processing results.
[0011] Preferably, the wound image is analyzed by a wound classification model to determine the wound image classification result, specifically including:
[0012] The wound image is input into the Resnet model and processed as follows:
[0013] receiving the wound image through an input layer;
[0014] Extract basic features of wound images through convolutional layers;
[0015] Enhance feature learning capabilities through multiple residual blocks;
[0016] After passing the batch normalization layer, the spatial dimension of the feature map is reduced through the transition layer;
[0017] The feature map is compressed into a vector of fixed length through a global average pooling layer;
[0018] Map the learned features to the final category scores through the fully connected layer;
[0019] Use the Softmax function to convert the category score into a probability distribution and map the probability distribution to the [0,1] interval;
[0020] It is determined whether the mapped score is less than a threshold value. If so, the collected wound image data is processed by super-resolution technology to obtain a refined wound image, thereby performing secondary classification.
[0021] Preferably, the fused information is used as the input of the multimodal large model, and the retrieval enhancement generation method is integrated to accurately output the processing results, specifically including:
[0022] The fused information is preprocessed. The wound image is divided into small blocks based on the Vision Transformer visual encoder. Each small block is converted into an image feature vector, and the image feature vector is compressed to a fixed length through a single-layer cross-attention module.
[0023] The compressed image is fused with the text information to form a multimodal feature vector;
[0024] The multimodal feature vector is input into the multi-layer Transformer network, and the model's ability to understand and process multimodal information is enhanced through the self-attention mechanism and multi-layer perceptron layer;
[0025] The relevant information of the fusion information is queried in the pre-built vector database by using the retrieval enhancement generation method;
[0026] Combine the relevant information with the processed multimodal feature vector and output the processed result through the decoder.
[0027] Preferably, a multimodal large model training phase is also included:
[0028] Obtain various types of wound image data, injury description data, corresponding drug information and wound types as training data;
[0029] The wound image data, injury description data, corresponding drug information and wound type are used as the input of the multimodal large model, and the retrieval enhancement generation method is integrated into the multimodal large model for fine-tuning;
[0030] The prediction results of the multimodal large model are obtained through forward propagation, and the difference between the prediction results and the actual labels, i.e. the loss value, is calculated;
[0031] The back-propagation algorithm is used to calculate the gradient of the loss value to the multimodal large model parameters, and the model parameters are updated according to the hyperparameters until the performance of the model on the validation set reaches the preset standard or reaches the specified number of iterations.
[0032] Preferably, it also includes:
[0033] The wound classification model is trained based on the wound image data and wound types in the training data.
[0034] A computer device comprises: a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the steps of a field emergency data processing method based on a multimodal intelligent algorithm are implemented.
[0035] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a field emergency data processing method based on a multimodal intelligent algorithm.
[0036] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a field first aid data processing method, equipment and medium based on a multimodal intelligent algorithm. Injured persons can reflect their current conditions by inputting data in various forms such as voice and taking photos. By using intelligent diagnosis technology, there is no need to wait for doctors to come online in real time, and instant advice can be provided to users quickly and accurately. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0038] Figure 1 A flow chart of the field emergency data processing method based on a multimodal intelligent algorithm provided by the present invention.
[0039] Figure 2 This is a flow chart of the model training phase provided by the present invention.
[0040] Figure 3 This is a schematic diagram of the wound classification model structure provided by the present invention. DETAILED DESCRIPTION
[0041] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0042] The embodiment of the present invention discloses a method for processing field emergency data based on a multimodal intelligent algorithm, such as Figure 1 As shown, including:
[0043] Obtain field first aid data, including wound image data and injury description voice data;
[0044] The wound image data is analyzed through the wound classification model to determine the wound image classification results. At the same time, the voice data describing the injury is converted into text and the text is corrected through the large language model.
[0045] Compare the corrected text with the wound image classification results, and perform information fusion when the results are consistent;
[0046] The fused information is used as the input of the multimodal large model, and the retrieval enhancement generation method is integrated to accurately output the processing results.
[0047] In this embodiment, when the wound image data taken by the injured person is transmitted to the wound classification model, the model analyzes the image to accurately determine the type of injury to which the wound belongs, and gives a corresponding classification and score. To ensure the accuracy and reliability of the results, the classification model uses a preset threshold as the judgment standard. The score output by the classification model is mapped by the softmax function, and the value is between [0,1], so the threshold is set according to this range.
[0048] If the score is lower than this threshold, the system will not pass the preliminary classification results to the downstream process for information fusion, but will start super-resolution processing technology. Super-resolution processing aims to significantly improve image quality through advanced algorithms, thereby capturing more wound details to support more accurate secondary classification. If the re-classification result after image quality enhancement still does not meet the standard value, a request for re-shooting will be issued to the user.
[0049] Therefore, the wound image data is analyzed through the wound classification model to determine the wound image classification results, in which Resnet is used as the classification model, such as Figure 3 As shown in the figure, the ResNet model consists of an input layer, a convolutional layer, multiple core residual blocks, a batch normalization layer, a ReLU activation function, a global average pooling layer, a fully connected layer, and an output layer. Batch normalization and ReLU activation functions are used to accelerate training and introduce nonlinearity. The global average pooling layer compresses the feature map into a vector of fixed length. Finally, the fully connected layer and the output layer map the learned features to the final category output. The residual block is used to solve the gradient vanishing problem and enhance the feature learning ability. Each residual block of the model contains two paths: one is the main path, which contains multiple convolutional layers and nonlinear activation functions; the other is a shortcut path that directly connects the input and output of the block. The specific process includes:
[0050] The wound image is input into the Resnet model and processed as follows:
[0051] receiving the wound image through an input layer;
[0052] Extract basic features of wound images through convolutional layers;
[0053] Enhance feature learning capabilities through multiple residual blocks;
[0054] The spatial dimension of the feature map is reduced through a transition layer after a batch normalization layer and a ReLU activation function;
[0055] The feature map is compressed into a vector of fixed length through a global average pooling layer;
[0056] Map the learned features to the final category scores through the fully connected layer;
[0057] Use the Softmax function to convert the category score into a probability distribution and map the probability distribution to the [0,1] interval;
[0058] It is determined whether the mapped score is less than a threshold value. If so, the collected wound image data is processed by super-resolution technology to obtain a refined wound image, thereby performing secondary classification.
[0059] While processing wound image data, Automatic Speech Recognition (ASR) technology begins to efficiently convert the user's input voice into text. Subsequently, the text information is transmitted to the text correction submodule, which accurately corrects the text converted by ASR by calling the advanced language large model API. In order to speed up this process and ensure accuracy, the VLLM (Virtual Large Language Model) framework is used for inference acceleration, thereby achieving fast and accurate text conversion of user voice information. Specifically, it includes:
[0060] Large language models usually use a multi-layer neural network structure to process text data through a self-attention mechanism. When processing text, the model first divides the input text into vocabulary or sub-word units, and then converts these units into high-dimensional vectors through an embedding layer. Next, these vectors are fed into multiple encoder layers, each of which contains self-attention and feedforward neural networks. The self-attention mechanism enables the model to capture long-distance dependencies in the text, while the feedforward network further extracts features. After multiple layers of encoding, the text representation is fed into the decoder, which outputs the corresponding output. When using a large language model to correct a piece of text, first input the original text into the model. The model will identify the wrong or unnatural parts of the text based on its pre-trained huge language knowledge base, and then replace or modify these parts by generating more appropriate words or phrases, and output the corrected text to make the sentence more fluent, accurate and in line with the context.
[0061] Before each modality data is input into the multimodal large model, a sophisticated calibration process is required to ensure data consistency and accuracy. This process aims to ensure that there is no disagreement between the wound classification results and the speech conversion results, as inconsistencies may cause ambiguity in the large model when processing data. To achieve this goal, text retrieval methods in natural language processing are used to detect the occurrence of specific text and compare it with the wound image classification results. Only when the results of the two are consistent will the integrated data be input into the multimodal large model for further online diagnosis.
[0062] If the image quality is insufficient, super-resolution technology is used to improve it. If the text information does not match, the user is required to re-upload the input voice information. Finally, the integrated data is input into the multimodal large model for online diagnosis.
[0063] In this embodiment, in order to improve the efficiency of reasoning, the VLLM framework is used to accelerate the reasoning of the multimodal large model. On this basis, the retrieval-augmented generation technology (RAG) is introduced to overcome the possible limitations of the model, such as hallucination phenomena and problems limited by limited knowledge, thereby improving the accuracy and reliability of text generation. RAG technology also significantly enhances the model's negative rejection ability. When the retrieved document does not contain the answer to the question, the multimodal large model can wisely refuse to answer and avoid generating answers based on wrong or irrelevant information. This feature is particularly critical in online diagnosis scenarios, which can effectively prevent the injured from misleading to take wrong wound treatment measures due to model misjudgment. In the reasoning stage, the large model using the RAG (Retrieval Augmented Generation) method will receive the user's diagnostic query and use the pre-built vector database to find the most relevant medical information to the query. This retrieved information will be integrated into the model's reasoning process, and work together with other components of the large model (such as encoders and decoders) to generate accurate and comprehensive diagnostic results. In this way, the RAG method significantly enhances the performance and reliability of the multimodal large model in the reasoning stage. QwenVL is used as the multimodal large model module, which includes:
[0064] Preprocess the fused information. Based on the Vision Transformer visual encoder, the image is divided into small blocks of 14x14 pixels. Each small block is converted into an image feature vector. The image feature vector is compressed to a fixed length of 256 through a single-layer cross-attention module, while retaining the image's position information.
[0065] The compressed image is fused with the text information to form a multimodal feature vector;
[0066] The multimodal feature vectors are input together into a large language model QwenLM, which is based on the pre-trained Qwen-7B model, through a multi-layer Transformer network, and the self-attention mechanism and multi-layer perceptron layer enhance the model's understanding and processing capabilities of multimodal information;
[0067] Query the pre-built vector database for information related to text features through the retrieval enhancement generation method;
[0068] Combine the relevant information with the processed multimodal feature vector and output the processed result through the decoder.
[0069] In this embodiment, if Figure 2 As shown, before the data processing is formally carried out, the model needs to be trained, as follows:
[0070] Collect various types of wound data, including but not limited to wound image data, injury description data (such as shape, size, color, bleeding, signs of infection, etc.), basic information of the injured (such as age, gender, previous medical history, etc.), and drug information (such as drug name, purpose, usage, etc.).
[0071] At the same time, the type, status and possible complications of the wound are carefully and accurately judged and the corresponding information is marked. The drug information is also strictly reviewed and annotated as a mark of the corresponding wound to ensure that appropriate drug recommendations can be provided to the injured in subsequent model applications.
[0072] Based on the above data, the wounds are divided into several categories, such as cuts, insect bites, snake bites, and fractures. After preprocessing the images, the training of the wound classification model begins.
[0073] The output type of the wound classification model, wound image data, injury description data, and drug information are used as input to construct a vector database. The process of constructing the vector database involves converting the output type of the wound classification model, wound image data, injury description data, and drug information into points in a high-dimensional vector space, and then converting these data into numerical vectors through specific encoding techniques. These vectors are then stored in the database for efficient similarity search and data analysis. Specifically, first, feature extraction is performed on the wound image and injury description, the image is converted into a vector of pixel values, and the text description is converted into a word embedding vector; then, the output type of the wound classification model is also converted into a corresponding numerical vector. Finally, these vectors are stored in the vector database together with the vector of drug information, so that similar wound cases can be quickly retrieved or corresponding drugs can be recommended through distance metrics between vectors (such as cosine similarity). After that, the Retrieval Augmented Generation (RAG) component is integrated into the multimodal large model, and combined with fine-tuning technology, the performance of the model can be effectively optimized.
[0074] During the fine-tuning process, first load the pre-trained large model, which is usually trained on a large amount of general data and has a certain generalization ability. Next, input the training set data into the multimodal large model, obtain the model's prediction results through forward propagation, and calculate the difference between the prediction results and the actual labels, that is, the loss value. Then, use the backpropagation algorithm to calculate the gradient of the loss value to the model parameters, and update the model parameters according to hyperparameters such as the learning rate. This process will be iterated continuously until the performance of the model on the validation set reaches the preset standard or reaches the specified number of iterations.
[0075] Fine-tuning a large multimodal model can quickly adapt to new tasks and data distributions, improving the model's performance on new tasks. At the same time, since only some parameters are adjusted during fine-tuning, the computing resources and time cost required for fine-tuning are lower and more efficient than training a new model from scratch.
[0076] The present invention realizes efficient and accurate multimodal data processing by training wound classification models and medical multimodal large models, combined with fine-tuning technology, ensuring accurate drug recommendations and treatment plan decision-making support in the reasoning stage.
[0077] This embodiment provides a computer device, including: a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the steps of a field emergency data processing method based on a multimodal intelligent algorithm are implemented.
[0078] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of a field emergency data processing method based on a multimodal intelligent algorithm are implemented.
[0079] A person skilled in the art can understand that: all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc. Various media that can store program codes.
[0080] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0081] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for processing field emergency data based on a multimodal intelligent algorithm, characterized in that: include: Obtain field first aid data, including wound images and voice data describing the injury; The wound image is analyzed through the wound classification model to determine the wound image classification result, and the voice data describing the injury is converted into text and the text is corrected through the large language model; Compare the corrected text with the wound image classification results, and perform information fusion when the results are consistent; The fused information is used as the input of the multimodal large model, and the retrieval enhancement generation method is integrated to accurately output the processing results.
2. The method for processing field emergency data based on a multimodal intelligent algorithm according to claim 1, characterized in that: The wound image is analyzed through the wound classification model to determine the wound image classification results, including: The wound image is input into the Resnet model and processed as follows: receiving the wound image through an input layer; Extract basic features of wound images through convolutional layers; Enhance feature learning capabilities through multiple residual blocks; The spatial dimension of the feature map is reduced through a transition layer after a batch normalization layer and a ReLU activation function; The feature map is compressed into a vector of fixed length through a global average pooling layer; Map the learned features to the final category scores through the fully connected layer; Use the Softmax function to convert the category score into a probability distribution and map the probability distribution to the [0,1] interval; It is determined whether the mapped score is less than a threshold value. If so, the collected wound image data is processed by super-resolution technology to obtain a refined wound image, thereby performing secondary classification.
3. The method for processing field emergency data based on multimodal intelligent algorithm according to claim 1, characterized in that: The fused information is used as the input of the multimodal large model, and the retrieval enhancement generation method is integrated to accurately output the processing results, including: The fused information is preprocessed. The wound image is divided into small blocks based on the Vision Transformer visual encoder. Each small block is converted into an image feature vector, and the image feature vector is compressed to a fixed length through a single-layer cross-attention module. The compressed image is fused with the text information to form a multimodal feature vector; The multimodal feature vector is input into the multi-layer Transformer network, and the model's ability to understand and process multimodal information is enhanced through the self-attention mechanism and multi-layer perceptron layer; The relevant information of the fusion information is queried in the pre-built vector database by using the retrieval enhancement generation method; Combine the relevant information with the processed multimodal feature vector and output the processed result through the decoder.
4. The method for processing field emergency data based on multimodal intelligent algorithm according to claim 1, characterized in that: It also includes the multimodal large model training phase: Obtain various types of wound image data, injury description data, corresponding drug information and wound types as training data; The wound image data, injury description data, corresponding drug information and wound type are used as the input of the multimodal large model, and the retrieval enhancement generation method is integrated into the multimodal large model for fine-tuning; The prediction results of the multimodal large model are obtained through forward propagation, and the difference between the prediction results and the actual labels, i.e. the loss value, is calculated; The back-propagation algorithm is used to calculate the gradient of the loss value to the parameters of the multimodal large model, and the model parameters are updated according to the hyperparameters until the performance of the model on the validation set reaches the preset standard or reaches the specified number of iterations.
5. The method for processing field emergency data based on multimodal intelligent algorithm according to claim 4 is characterized in that: Also includes: The wound classification model is trained based on the wound images and wound types in the training data.
6. A computer device, characterized in that: include: A memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
7. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, which, when executed by a processor, implements the steps of the method described in any one of claims 1 to 5.
Citation Information
Cited By
PMCT image acquisition and multi-modal data processing system and method for forensic traumatology deep learning
CN120564225A
Emergency team real-time voice instruction analysis and collaborative optimization method and system
CN120783747A