Marine first aid auxiliary diagnosis method and system based on artificial intelligence

By adopting multimodal perception fusion algorithm based on artificial intelligence and deep learning technology in maritime first aid, the problems of slow response speed and low accuracy are solved, and faster and more accurate first aid response is achieved.

CN120021940APending Publication Date: 2025-05-23CSSC HAISHEN MEDICAL TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411987475.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing technology has slow response speed and low accuracy in maritime first aid, and cannot effectively use artificial intelligence for in-depth analysis, and lacks professional considerations for environmental factors.

Method used

Using artificial intelligence-based maritime first aid assisted diagnosis method, we use a multimodal perception fusion algorithm and a deep learning framework for joint analysis, locate injuries and key environmental factors, predict the development trends of potential risk factors, and provide professional advice through semantic similarity calculation and reinforcement learning optimization search strategies.

Benefits of technology

It significantly improves the speed and accuracy of maritime emergency response, provides immediate, context-specific professional advice, ensuring the safety of the injured and the effectiveness of rescue operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120021940A_ABST
    Figure CN120021940A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence auxiliary medical treatment, and provides an artificial intelligence-based marine first aid auxiliary diagnosis method and system. The method comprises the following steps: receiving a real-time multimedia data stream from an offshore first-aid site; carrying out conjoint analysis on the image and audio information, positioning key elements of the involved injured individuals and the surrounding environment, and obtaining an emergency scene description after enhanced analysis; extracting professional suggestions suitable for the current first-aid situation in combination with correction parameters of influences of specific variables of the marine environment on injuries and patients; activating a remote expert system, interacting with field first-aid personnel through two-way video communication, and generating personalized guidance information; and combining the personalized guidance information with the calibrated visual field of the first-aid personnel, and displaying decision support information provided by artificial intelligence in the visual field of the first-aid personnel. According to the technical scheme, the first-aid response speed and accuracy are improved, and it is ensured that the wounded patient obtains the most appropriate medical treatment in the first time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence-assisted medical technology, and in particular, to an artificial intelligence-based auxiliary diagnosis method and system for marine emergency treatment. Background Art

[0002] In marine environments, such as ship accidents and diving accidents, rapid and accurate emergency response is crucial. Due to the complex and changeable rescue environment at sea and the distance from land-based medical resources, high requirements are placed on the analysis of real-time multimedia data streams (images and audio information). The method needs to be able to handle unstable data transmission and provide reliable assessment and guidance of the injured in the absence of direct medical assistance.

[0003] Currently, first aid at sea usually relies on radio communication or satellite phone to contact doctors on shore for advice. Some more advanced systems have begun to use video calls to enhance communication, but the functions of these systems are mainly limited to the transmission of visual information, fail to make full use of artificial intelligence for in-depth analysis, and lack professional consideration of environmental factors.

[0004] However, existing solutions often fail to provide immediate, context-based professional advice, especially when faced with complex emergencies, where the guidance provided may not be specific or timely enough. In addition, traditional telemedicine methods cannot effectively simulate the on-site environment, making it difficult for remote experts to fully understand the actual situation, thus affecting the quality of guidance. At the same time, they also lack the ability to adjust strategies based on real-time changes, which limits their application efficiency in dynamic environments. By introducing advanced technologies such as multimodal perception fusion algorithms, spatiotemporal context model predictions, semantic similarity calculations, and reinforcement learning optimized retrieval strategies, the effectiveness of maritime first aid auxiliary diagnosis can be significantly improved. Summary of the invention

[0005] The embodiments of the present application provide an artificial intelligence-based maritime emergency diagnosis auxiliary method and system to solve the problems of slow emergency response and low accuracy in the prior art.

[0006] In a first aspect, an embodiment of the present application provides an artificial intelligence-based marine emergency diagnosis method, comprising:

[0007] Receiving a real-time multimedia data stream from a marine emergency scene, wherein the multimedia data stream includes image and audio information;

[0008] Based on the multimedia data stream, a multimodal perception fusion algorithm is applied to jointly analyze the image and audio information, and at the same time, the target detection and tracking technology in the deep learning framework is used to locate the key elements of the injured individuals and the surrounding environment involved, and the development trend of potential risk factors is predicted using a spatiotemporal context model to obtain an enhanced parsed emergency scene description;

[0009] Based on the key elements extracted from the emergency scene description after the enhanced analysis, a semantic similarity calculation algorithm is used in combination with correction parameters of the impact of variables unique to the marine environment on the injured, and historical cases with high relevance are retrieved from the cloud medical knowledge graph. The retrieval strategy is dynamically optimized through a reinforcement learning algorithm to extract professional advice applicable to the current emergency situation.

[0010] Based on the professional advice and the emergency scenario description after the enhanced analysis, a remote expert system is activated to interact with the on-site emergency personnel through two-way video communication, natural language processing technology is used to understand and generate dialogue content for specific situations, and a virtual reality or augmented reality simulation training module is used to reproduce the emergency scenario, so that the remote expert can provide guidelines based on the emergency scenario description and generate personalized guidance information;

[0011] The personalized guidance information is combined with the calibrated first aid personnel's perspective, and the decision support information provided by artificial intelligence is displayed in the first aid personnel's field of view. The decision support information includes customized injury status assessment, recommended first aid steps and risk warnings.

[0012] Optionally, based on the multimedia data stream, a multimodal perception fusion algorithm is applied to jointly analyze the image and audio information, and target detection and tracking technology in a deep learning framework is used to locate key elements of the injured individuals and the surrounding environment involved, and a spatiotemporal context model is used to predict the development trend of potential risk factors, so as to obtain an enhanced parsed emergency scene description, including:

[0013] Using the received and synchronized multimedia data streams, the image and audio information received from multiple sensors is preliminarily processed, including noise reduction and format conversion operations, to ensure that the data of different modalities are aligned in timestamps, and obtain preprocessed multimedia data;

[0014] Based on the preprocessed multimedia data, a multimodal perception fusion algorithm is applied to jointly analyze the image and audio information, extract cross-modal features, and identify and associate objects in the image with events in the audio through an attention mechanism or a similarity matching algorithm, so as to construct a complete and semantically rich scene representation and obtain a fusion analysis result;

[0015] According to the fusion analysis results, a deep learning framework is used to detect targets in the image, automatically identify the injured individuals and rescuers, rescue equipment, and obstacles in the image, and track them in real time, continuously monitor position changes, and obtain target detection and tracking data;

[0016] Based on the target detection and tracking data, a spatiotemporal context model is constructed, and the changes in time series and spatial distribution characteristics are considered. The historical data of the same or similar situations in the past period of time are combined to predict the development trend of potential risk factors. According to the variables unique to the marine environment and their impact on the status of the injured and the rescue operation, the prediction results are adjusted to obtain the spatiotemporal context prediction results;

[0017] According to the spatiotemporal context prediction results, a comprehensive description including detailed location, action, and potential risk factors is formed to obtain an emergency scene description after enhanced analysis.

[0018] Optionally, based on the preprocessed multimedia data, a multimodal perception fusion algorithm is applied to jointly analyze the image and audio information, extract cross-modal features, and identify and associate objects in the image with events in the audio through an attention mechanism or a similarity matching algorithm to construct a complete and semantically rich scene representation, and obtain a fusion analysis result, including:

[0019] Using the preprocessed multimedia data, respectively extracting features from the image and the audio information, extracting visual features from the image, and extracting acoustic features from the audio, to obtain a preliminary modal feature set;

[0020] Based on the preliminary modal feature set, a multimodal perceptual fusion algorithm is applied to map the feature vectors of the image and audio into a shared feature space to facilitate direct comparison and association between different modalities, wherein the multimodal perceptual fusion algorithm is a deep learning-based model that generates a comprehensive feature representation;

[0021] According to the comprehensive feature representation, using an attention mechanism algorithm, a correlation score between the object detected in the image and the audio clip is calculated, the object in the image and the event in the audio are identified and associated, and a cross-modal association result is obtained;

[0022] Based on the cross-modal association results, a dynamic scene graph containing time series information is further constructed, which describes the position and action state of objects at the current moment, records the changing trends of entities over time and the interactive relationships between them, and forms a complete and semantically rich scene representation;

[0023] Based on the complete scene representation and combined with contextual information, the key elements in the scene are semantically interpreted to obtain fusion analysis results.

[0024] Optionally, the spatiotemporal context model is constructed based on the target detection and tracking data, taking into account the changes in time series and spatial distribution characteristics, combining historical data in the same or similar situations over a period of time in the past, predicting the development trend of potential risk factors, and adjusting the prediction results according to the variables unique to the marine environment and their impact on the status of the injured and the rescue operation, to obtain the spatiotemporal context prediction results, including:

[0025] Using the target detection and tracking data, the position and motion status of the injured individuals, rescuers, rescue equipment, and obstacles identified in the image are continuously monitored, and the changing trajectories of the injured individuals, rescuers, rescue equipment, and obstacles over time are recorded to form time series data, and dynamic position and motion status records are obtained;

[0026] Based on the dynamic position and action state records, a spatiotemporal context model is constructed, wherein the spatiotemporal context model considers the evolution of the position and action state of the object in the time dimension, analyzes the distribution characteristics in space, so as to capture the complete information of the dynamic scene and generate the spatiotemporal context model;

[0027] Based on the spatiotemporal context model, combined with historical data in the same or similar contexts over a period of time, typical patterns and development trends are identified through deep learning algorithms to obtain historical pattern and development trend analysis results;

[0028] Based on the analysis results of the historical patterns and development trends, and according to the variables unique to the marine environment, the impact of these variables on the current state of the injured and the rescue operations is evaluated, and environmental correction parameters are introduced to adjust the parameter settings of the prediction model to ensure that the prediction results are more in line with the actual environment, thereby obtaining an adjusted prediction model;

[0029] According to the adjusted prediction model, the development trend of potential risk factors is predicted, including the deterioration of the health condition of the injured, the increased difficulty of rescue, and the aggravation of environmental risks, and a spatiotemporal context prediction result including detailed location, action, and potential risk factors is generated.

[0030] Optionally, based on the key elements extracted from the emergency scene description after the enhanced analysis, a semantic similarity calculation algorithm is used in combination with correction parameters of the impact of variables unique to the marine environment on the injured, and historical cases with high relevance are retrieved from the cloud medical knowledge graph, and the retrieval strategy is dynamically optimized through a reinforcement learning algorithm to extract professional suggestions suitable for the current emergency situation, including:

[0031] By using the emergency scene description after enhanced parsing, key elements are extracted, wherein the key elements include the status of the injured individual, the condition of the surrounding environment, and the development trend of any potential risk factors, and a key element set is obtained;

[0032] Based on the key element set, a semantic similarity calculation algorithm is applied to compare the key elements of the current scenario with the historical cases in the cloud medical knowledge graph. The semantic similarity calculation algorithm can identify and quantify the semantic similarities between different cases, find the historical case set that best matches the current situation, and generate a preliminary matching historical case list;

[0033] In combination with variables unique to the marine environment, environmental parameters such as seawater temperature, salinity, and wind and wave levels are introduced as correction factors to evaluate the impact of the variables on the patient status and rescue operations, and the results calculated by the semantic similarity calculation algorithm are adjusted to ensure that the retrieved historical cases have similarities in medical treatment and reference value in environmental adaptability, and the preliminary matching historical case list is corrected to obtain a corrected matching historical case list;

[0034] According to the revised matching historical case list, highly relevant historical cases are retrieved from the cloud medical knowledge graph, wherein the historical cases contain similar medical treatment methods, and take into account successful first aid experience and failed lessons in a specific marine environment, so as to generate a highly relevant historical case set;

[0035] Based on the highly relevant historical case set, the retrieval strategy is dynamically optimized through a reinforcement learning algorithm, and the reinforcement learning algorithm continuously improves itself according to the newly added data to improve the accuracy and efficiency of future case retrieval to obtain an optimized retrieval strategy;

[0036] Based on the optimized search strategy, the highly relevant historical cases retrieved are further analyzed to extract professional advice applicable to the current emergency situation. The professional advice integrates the successful experience of historical cases, best practices and the latest medical research results.

[0037] Optionally, based on the key element set, a semantic similarity calculation algorithm is applied to compare the key elements of the current scene with historical cases in the cloud medical knowledge graph. The semantic similarity calculation algorithm can identify and quantify the semantic similarities between different cases, find the historical case set that best matches the current situation, and generate a preliminary matching historical case list, including:

[0038] Using the extracted key element set, structurally encoding the key elements to ensure that each element is accurately represented, thereby obtaining structured key element data;

[0039] Based on the structured key element data, a multi-dimensional feature vector is constructed, wherein the multi-dimensional feature vector includes information directly describing the patient status and environmental conditions, time series information, and spatial distribution characteristics, so as to fully reflect the characteristics of the current emergency scene and generate a multi-dimensional feature vector;

[0040] Based on the multi-dimensional feature vector, a semantic similarity calculation algorithm is applied to compare the multi-dimensional feature vector of the current scene with the feature vector of the historical case in the cloud medical knowledge graph, and the semantic distance or similarity score between the two is calculated through natural language processing technology and machine learning models to identify historical cases with high similarity and obtain a list of candidate cases;

[0041] Based on the candidate case list, combined with the semantic similarity score, the candidate cases are screened and sorted, and according to the set threshold or ranking rule, several historical cases that best match the current emergency scenario are selected to form a preliminary matching historical case list.

[0042] Optionally, based on the professional advice and the emergency scenario description after enhanced analysis, a remote expert system is activated to interact with on-site emergency personnel through two-way video communication, natural language processing technology is used to understand and generate dialogue content for specific situations, and a virtual reality or augmented reality simulation training module is used to reproduce the emergency scenario, so that the remote expert can provide guidelines based on the emergency scenario description and generate personalized guidance information, including:

[0043] Utilizing the professional advice and the emergency scene description after enhanced analysis, a remote expert system is activated; the remote expert system is connected to remote medical experts via a high-speed network to ensure that they obtain the latest emergency scene information in real time and obtain an activated remote expert system;

[0044] Based on the activated remote expert system, a two-way video communication link is established to realize real-time interaction between the remote expert and the on-site emergency personnel. Through high-definition video streaming, the remote expert can directly observe the emergency scene and communicate with the emergency personnel in real time, thus obtaining a two-way video communication link and a real-time interactive environment;

[0045] Based on the two-way video communication link and real-time interactive environment, natural language processing technology is applied to analyze the questions and statements of the emergency personnel, and accurate and easy-to-understand responses are automatically generated to ensure the efficiency and accuracy of communication and obtain optimized conversation content;

[0046] In combination with the optimized dialogue content and emergency scenario description, a virtual reality or augmented reality simulation training module is used to reproduce the emergency scenario, wherein the training module helps the remote expert to have a deeper understanding of the on-site environment and simulate different emergency operations, so as to provide intuitive guidance to on-site emergency personnel and generate a simulated emergency scenario;

[0047] Based on the simulated emergency scenario, the remote expert adjusts the guidelines according to the actual scenario and provides precise operational suggestions and decision support. The operational suggestions take into account the current patient status, environmental conditions and lessons learned from historical cases to obtain optimized guidelines.

[0048] According to the optimized guidelines, all relevant information is integrated, including the professional advice, emergency scenario descriptions, optimized dialogue content and simulated first aid scenarios, to generate personalized guidance information.

[0049] In a second aspect, an embodiment of the present application provides an artificial intelligence-based marine emergency diagnosis system, which is used to execute the above method, including:

[0050] A multimedia data acquisition device configured to receive a real-time multimedia data stream from a marine emergency scene, wherein the multimedia data stream includes image and audio information;

[0051] An information processing unit is configured to jointly analyze the image and audio information based on the multimedia data stream, locate the key elements of the injured individuals and the surrounding environment involved, predict the development trend of potential risk factors, combine the correction parameters of the impact of variables unique to the marine environment on the injured, retrieve and extract professional suggestions applicable to the current emergency situation from historical cases in the cloud medical knowledge map, activate the remote expert system to provide guidelines, and generate personalized guidance information;

[0052] The cloud medical knowledge graph is configured to store historical cases and current emergency scene processing information for subsequent retrieval at other emergency scenes;

[0053] A remote expert system configured to interact with first responders on site through a video communication display device, using a virtual reality or augmented reality simulation training module for reproducing first responder scenarios, providing guidelines based on emergency scenario descriptions, and generating personalized guidance information;

[0054] The video communication display device is configured for interaction between a remote expert and first responders on site, and for displaying to the remote expert the conversation content for a specific situation understood and generated using natural language processing technology, and for displaying to the first responders on site personalized guidance information given by the remote expert.

[0055] Optionally, the information processing unit includes:

[0056] An analysis and prediction module is used to jointly analyze the image and audio information based on the multimedia data stream using a multimodal perception fusion algorithm, and to locate the key elements of the injured individuals and the surrounding environment using target detection and tracking technology in a deep learning framework, and to predict the development trend of potential risk factors using a spatiotemporal context model to obtain an enhanced parsed emergency scene description;

[0057] A correction and optimization module is used to retrieve highly relevant historical cases from the cloud medical knowledge graph based on the key elements extracted from the emergency scene description after the enhanced analysis, using a semantic similarity calculation algorithm combined with correction parameters of the impact of variables unique to the marine environment on the injured, and dynamically optimize the retrieval strategy through a reinforcement learning algorithm to extract professional suggestions suitable for the current emergency situation;

[0058] An activation providing module is used to activate a remote expert system based on the professional advice and the emergency scenario description after the enhanced analysis.

[0059] In an embodiment of the present application, a real-time multimedia data stream is received from a marine emergency rescue site, wherein the multimedia data stream includes image and audio information; based on the multimedia data stream, a multimodal perception fusion algorithm is applied to jointly analyze the image and audio information, and at the same time, the target detection and tracking technology in the deep learning framework is used to locate the key elements of the injured individuals and the surrounding environment involved, and the spatiotemporal context model is used to predict the development trend of potential risk factors to obtain an enhanced parsed emergency scene description; based on the key elements extracted from the enhanced parsed emergency scene description, a semantic similarity calculation algorithm is used in combination with correction parameters for the impact of variables unique to the marine environment on the injured, and historical cases with high relevance are retrieved from the cloud-based medical knowledge graph, and through strong The artificial intelligence learning algorithm dynamically optimizes the retrieval strategy and extracts professional advice applicable to the current emergency situation. Based on the professional advice and the emergency scene description after enhanced analysis, the remote expert system is activated to interact with the on-site emergency personnel through two-way video communication. Natural language processing technology is used to understand and generate dialogue content for specific situations. Virtual reality or augmented reality simulation training modules are used to reproduce the emergency scene, so that remote experts can provide guidelines based on the emergency scene description and generate personalized guidance information. The personalized guidance information is combined with the calibrated emergency personnel's perspective, and the decision support information provided by artificial intelligence is displayed in the emergency personnel's field of view, and the decision support information includes customized injury status assessment, recommended first aid steps and risk warnings.

[0060] The technical solution of this application has the following beneficial effects:

[0061] (1) This application receives and analyzes multimedia data streams from the marine emergency scene in real time, combines multimodal perception fusion algorithms with target detection and tracking technology in a deep learning framework, and the system can quickly locate the key elements of the injured individual and the surrounding environment, and predict the development trend of potential risk factors. This enables emergency personnel to obtain a comprehensive and accurate description of the emergency scene in the shortest time, so as to make decisions quickly.

[0062] (2) Using the semantic similarity calculation algorithm combined with the correction parameters of the impact of marine environment-specific variables on the injured, highly relevant historical cases were retrieved from the cloud medical knowledge graph. The retrieval strategy was dynamically optimized through the reinforcement learning algorithm, which further improved the accuracy of historical case matching, provided more targeted professional advice for the current emergency situation, and ensured the effectiveness of the treatment plan.

[0063] (3) Activate the remote expert system to interact with on-site emergency personnel through two-way video communication. Natural language processing technology is used to understand and generate conversation content for specific situations. Virtual reality or augmented reality simulation training modules reproduce emergency scenarios. Remote experts provide specific guidelines based on the emergency scenario description after enhanced analysis and generate personalized guidance information, which greatly enhances the effectiveness and efficiency of remote medical support.

[0064] (4) Combine personalized guidance information with the calibrated first aid personnel’s perspective, and display decision support information provided by artificial intelligence in the first aid personnel’s field of view, including customized patient status assessment, recommended first aid steps, and risk warnings. This instant feedback mechanism not only improves the accuracy of first aid operations, but also reduces the possibility of human error and ensures the safety of the injured.

[0065] (5) The system continuously optimizes its own search strategy and suggestion generation mechanism by learning and analyzing historical data, helping medical institutions to better predict possible emergencies, rationally allocate resources, and improve overall emergency response capabilities. At the same time, it also accumulates valuable experience and data for future emergency operations.

[0066] (6) By introducing virtual reality or augmented reality technology, remote experts can understand the on-site situation more intuitively and provide more detailed and professional guidance. This immersive interactive method not only improves the operating experience of emergency personnel, but also enhances their trust in the system and promotes the application and popularization of technology.

[0067] Furthermore, the embodiment of the present application also applies multimodal perception fusion algorithms and target detection and tracking technology in a deep learning framework to jointly analyze image and audio information from multiple sensors, extract cross-modal features, and construct a complete and semantically rich scene representation. The system not only automatically identifies injured individuals, rescuers, rescue equipment, and obstacles and tracks their position changes in real time, but also predicts the development trend of potential risk factors through a spatiotemporal context model to form an enhanced parsed emergency scene description. Furthermore, a semantic similarity calculation algorithm is used in combination with variables unique to the marine environment (such as seawater temperature, salinity, and wind and wave levels) to retrieve highly relevant historical cases from the cloud-based medical knowledge graph, and a reinforcement learning algorithm is used to dynamically optimize the retrieval strategy to extract professional advice suitable for the current emergency situation.

[0068] Through the above method, the speed and accuracy of maritime emergency response have been significantly improved, ensuring that the injured receive the most appropriate medical treatment at the first time. Intelligent means have enhanced the relevance and practicality of case retrieval, provided personalized remote expert guidance, improved the operational accuracy of emergency personnel, and optimized resource allocation and emergency response capabilities. Especially in complex and changing marine environments, this method has greatly improved the support level and success rate of emergency decision-making, and achieved more efficient and safer maritime emergency services. In addition, the system continuously optimizes the retrieval strategy through learning and self-improvement of historical data, further improving the accuracy and efficiency of future case retrieval, and accumulating valuable experience and data for future emergency operations.

[0069] These and other aspects of the present application will become more clearly understood in the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0071] Figure 1 A flowchart of an artificial intelligence-based maritime emergency diagnosis auxiliary method provided in an embodiment of the present application;

[0072] Figure 2 A schematic diagram of the structure of an artificial intelligence-based marine emergency diagnosis system provided in an embodiment of the present application;

[0073] Figure 3 A schematic diagram of the structure of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0074] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0075] In some of the processes described in the specification and claims of this application and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent the order of precedence, and do not limit the "first" and "second" to be different types.

[0076] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0077] Figure 1 A flowchart of an artificial intelligence-based marine emergency diagnosis method is provided for an embodiment of the present application, such as Figure 1 As shown, the method includes:

[0078] 101. Receive a real-time multimedia data stream from a marine emergency scene, wherein the multimedia data stream includes image and audio information;

[0079] It involves receiving real-time multimedia data streams containing image and audio information from the marine emergency scene. These data streams include but are not limited to video surveillance, drone photography, images recorded by handheld devices, as well as radio communications and audio collected by microphones. Multimedia data streams are used to capture visual and auditory information at the emergency scene and provide raw materials for subsequent analysis.

[0080] The system collects on-site image and audio data through a variety of sensors (such as cameras and microphones) and transmits them to the central processing unit. In order to ensure the consistency and availability of the data, the system performs preliminary processing on the received data, including noise reduction, format conversion and other operations to ensure that the data of different modes are aligned in timestamps to form preprocessed multimedia data. This process ensures the basic data quality for subsequent analysis and improves the response speed and accuracy of the overall system.

[0081] For example, in a maritime ship collision, the camera on the rescue helicopter and the ship-borne camera simultaneously recorded images of the accident scene, while the microphones on the ship and the helicopter captured cries for help and other environmental sounds. All of this data is transmitted to the command center on shore in real time, and after preliminary processing, a unified time-synchronized data stream is formed, providing a reliable foundation for the subsequent multimodal perception fusion.

[0082] 102. Based on the multimedia data stream, a multimodal perception fusion algorithm is applied to jointly analyze the image and audio information, and target detection and tracking technology in a deep learning framework is used to locate the key elements of the injured individuals and the surrounding environment involved, and a spatiotemporal context model is used to predict the development trend of potential risk factors, so as to obtain an enhanced parsed emergency scene description;

[0083] The pre-processed image and audio information are jointly analyzed using a multimodal perception fusion algorithm to extract cross-modal features, and the objects in the image and the events in the audio are identified and associated through an attention mechanism or similarity matching algorithm to construct a complete and semantically rich scene representation. Target detection and tracking technology is used to locate the key elements of the injured individual and the surrounding environment, and the spatiotemporal context model predicts the development trend of potential risk factors, ultimately obtaining an enhanced parsed emergency scene description.

[0084] Based on the pre-processed multimedia data, the system applies a multimodal perception fusion algorithm to extract visual features (such as color, texture, shape) in images and acoustic features (such as frequency, pitch, loudness) in audio, and uses target detection and tracking technology in a deep learning framework to identify and track injured individuals, rescuers, rescue equipment, and obstacles. Through the spatiotemporal context model, combined with historical data in the same or similar situations over a period of time in the past, the development trend of potential risk factors is predicted, thereby forming a detailed description of the emergency scene.

[0085] In the above-mentioned ship collision accident, the system not only identified the locations of multiple injured people, but also tracked their movement paths, while detecting floating debris and other obstacles that may hinder rescue operations. Through the spatiotemporal context model, the system predicted the changing trend of waves and their impact on rescue operations, and generated a detailed description of the emergency scene, helping the rescue team better plan the route of action and response measures.

[0086] 103. Based on the key elements extracted from the emergency scene description after the enhanced analysis, the semantic similarity calculation algorithm is combined with the correction parameters of the impact of the variables unique to the marine environment on the injured, and historical cases with high relevance are retrieved from the cloud medical knowledge graph, and the retrieval strategy is dynamically optimized through the reinforcement learning algorithm to extract professional suggestions suitable for the current emergency situation;

[0087] Based on the key elements extracted from the emergency scene description after enhanced analysis, the semantic similarity calculation algorithm is combined with the correction parameters of the impact of marine environment-specific variables (such as sea temperature, salinity, wind and wave level) on the injured, and highly relevant historical cases are retrieved from the cloud medical knowledge graph. The retrieval strategy is dynamically optimized through the reinforcement learning algorithm to extract professional suggestions suitable for the current emergency situation.

[0088] The system first extracts key elements, such as the status of the injured individual, the state of the surrounding environment, and the development trend of potential risk factors, to form a set of key elements. Then, the semantic similarity calculation algorithm is used to compare the key elements of the current scene with the historical cases in the cloud medical knowledge graph to find the most matching set of historical cases. The similarity calculation results are adjusted by combining the variables unique to the marine environment as correction factors to ensure that the retrieved historical cases have similarities in medical treatment and have reference value in environmental adaptability. Finally, the retrieval strategy is continuously optimized through the reinforcement learning algorithm to improve the accuracy and efficiency of future case retrieval.

[0089] Continuing with the example of the above-mentioned ship collision accident, the system extracted key factors such as the victim's drowning status, body temperature drop, and wind and wave levels in the nearby sea area. Through the semantic similarity calculation algorithm, the system found several similar drowning rescue cases in history and made corrections based on the marine environmental conditions at the time. In the end, the system put forward targeted first aid suggestions, such as the use of specific insulation equipment and anti-cold drugs, to ensure the effectiveness of the rescue operation.

[0090] 104. Based on the professional advice and the emergency scenario description after the enhanced analysis, activate the remote expert system to interact with the on-site emergency personnel through two-way video communication, use natural language processing technology to understand and generate dialogue content for specific situations, and use virtual reality or augmented reality simulation training modules to reproduce the emergency scenario, so that the remote expert can provide guidelines based on the emergency scenario description and generate personalized guidance information;

[0091] Based on professional advice and enhanced parsed emergency scenario descriptions, the remote expert system is activated to interact with on-site first responders through two-way video communication, and natural language processing technology is used to understand and generate conversation content for specific situations. Virtual reality or augmented reality simulation training modules are used to reproduce emergency scenarios, enabling remote experts to provide guidelines based on the emergency scenario descriptions and generate personalized guidance information.

[0092] The system activates the remote expert system, connecting on-site first responders and remote experts through two-way video communication. Natural language processing technology is used to understand and generate conversation content to ensure smooth and efficient communication. Virtual reality or augmented reality simulation training modules reproduce the first aid scene, allowing remote experts to intuitively understand the on-site situation and provide specific guidelines. The personalized guidance information generated by the system includes detailed operation steps, precautions and risk warnings, helping first responders perform their tasks more effectively.

[0093] In the case of a ship collision, remote experts observed the situation on site through video communication and recreated the drowning rescue scene using a virtual reality module. Experts recommended the use of specific rescue tools and techniques based on the personalized guidance information provided by the system, and provided real-time guidance to first responders on site through video communication. This interactive method greatly improved the accuracy and efficiency of rescue and ensured the safety of the injured.

[0094] 105. Combine the personalized guidance information with the calibrated first aid personnel's perspective, and display the decision support information provided by artificial intelligence in the first aid personnel's field of view, wherein the decision support information includes customized patient status assessment, recommended first aid steps and risk warnings.

[0095] The personalized guidance information is combined with the calibrated first aid personnel's perspective, and the decision support information provided by artificial intelligence is displayed in the first aid personnel's field of view, including customized injury status assessment, recommended first aid steps and risk warning. This process ensures that first aid personnel can obtain the most important information at the first time and make the most appropriate decision.

[0096] The system combines personalized guidance information with the first responders' field of view, and displays customized decision support information directly in the first responders' field of view through AR glasses or other display devices. This information includes assessment of the patient's status, recommended first aid steps, and risk warnings, ensuring that first responders can take action quickly and accurately. In addition, the system will dynamically update information based on actual conditions to provide continuous support.

[0097] In a ship collision accident, the AR glasses worn by the first responders displayed the decision support information provided by the system in real time, such as the specific location of the injured, recommended rescue steps and potential risk points. This information helped the first responders quickly determine the best rescue plan, avoiding unnecessary delays and erroneous operations, and significantly improving the success rate of rescue.

[0098] Through the implementation of steps 101 to 105, from receiving real-time multimedia data streams to finally displaying decision support information provided by artificial intelligence, the entire process significantly improves the speed and accuracy of maritime emergency response, ensuring that the injured receive the most appropriate medical treatment as soon as possible. Intelligent means enhance the relevance and practicality of case retrieval, provide personalized remote expert guidance, improve the operational accuracy of emergency personnel, and optimize resource allocation and emergency response capabilities. Especially in complex and changing marine environments, this method greatly improves the support level and success rate of emergency decision-making, and realizes more efficient and safer maritime emergency services. By learning from historical data and self-improvement, the system continuously optimizes the retrieval strategy, further improves the accuracy and efficiency of future case retrieval, and accumulates valuable experience and data for future emergency operations.

[0099] In order to further improve the accuracy and real-time performance of the analysis of maritime emergency scenarios, in some embodiments, the multimodal perception fusion algorithm is used to jointly analyze the image and audio information based on the multimedia data stream in step 102, and the target detection and tracking technology in the deep learning framework is used to locate the key elements of the injured individuals and the surrounding environment involved, and the spatiotemporal context model is used to predict the development trend of potential risk factors, so as to obtain an enhanced analysis of the emergency scenario description, including:

[0100] Using the received and synchronized multimedia data streams, the image and audio information received from multiple sensors are preliminarily processed, including noise reduction and format conversion operations, to ensure that data of different modes are aligned in timestamps, and obtain preprocessed multimedia data; based on the preprocessed multimedia data, a multimodal perception fusion algorithm is applied to jointly analyze the image and audio information, extract cross-modal features, and identify and associate objects in the image with events in the audio through an attention mechanism or a similarity matching algorithm, to construct a complete and semantically rich scene representation, and obtain a fusion analysis result; based on the fusion analysis result, a deep learning framework is used to perform target detection on the image, and automatically identify The injured individuals and rescuers, rescue equipment, and obstacles in the image are tracked in real time, and the position changes are continuously monitored to obtain target detection and tracking data; based on the target detection and tracking data, a spatiotemporal context model is constructed, considering the changes in time series and spatial distribution characteristics, combined with historical data in the same or similar situations over a period of time in the past, to predict the development trend of potential risk factors, and according to the variables unique to the marine environment and their impact on the status of the injured and the rescue operations, the prediction results are adjusted to obtain the spatiotemporal context prediction results; based on the spatiotemporal context prediction results, a comprehensive description of factors including detailed location, action, and potential risks is formed to obtain an enhanced parsed emergency scene description.

[0101] In this embodiment, the preliminary processing and timestamp alignment are to ensure the temporal consistency of data of different modalities, and the system performs preliminary processing on the images and audio information received from multiple sensors (such as cameras and microphones). These processes include noise reduction, format conversion and other operations to eliminate unnecessary interference and ensure data compatibility. By synchronizing the timestamp, all data are ensured to be aligned at the same time point, thereby providing a reliable basis for subsequent joint analysis. The multimodal perception fusion algorithm is used to jointly analyze the preprocessed image and audio information and extract cross-modal features. For example, visual features may include color, texture, shape, etc., while acoustic features cover frequency, pitch, loudness, etc. Through the attention mechanism or similarity matching algorithm, objects in the image and events in the audio are identified and associated to construct a complete and semantically rich scene representation. This process not only improves the utilization efficiency of single modality information, but also enhances the ability to understand the overall scene. Target detection and tracking technology is a target detection and tracking technology based on a deep learning framework that can automatically identify key objects in the image, such as injured individuals, rescuers, rescue equipment and obstacles, and track their position changes in real time. This step is crucial for continuously monitoring position changes in a dynamic environment, which helps to adjust rescue strategies in a timely manner. The spatiotemporal context model combines the changes in time series and the spatial distribution characteristics, considers historical data under the same or similar situations over a period of time, and predicts the development trend of potential risk factors. In particular, for variables unique to the marine environment (such as sea temperature, salinity, and wind and wave levels), the prediction results are adjusted to make the prediction closer to the actual situation. Ultimately, a comprehensive description of factors such as detailed location, action, and potential risks is formed, and an enhanced parsed emergency scene description is obtained.

[0102] In the embodiment of the present application, first, the image and audio information from multiple sensors are received and synchronized, and preliminary processing (such as noise reduction and format conversion) is performed to ensure that the data of different modes are aligned in the timestamp to obtain the preprocessed multimedia data. Secondly, based on the preprocessed multimedia data, the multimodal perception fusion algorithm is applied to jointly analyze the image and audio information, extract cross-modal features, and identify and associate objects in the image with events in the audio through the attention mechanism or similarity matching algorithm, build a complete and semantically rich scene representation, and obtain the fusion analysis result. Then, according to the fusion analysis result, the image is detected by the deep learning framework, and the injured individuals, rescuers, rescue equipment and obstacles in the image are automatically identified, and their position changes are tracked in real time, and the position changes are continuously monitored to obtain target detection and tracking data. Further, based on the target detection and tracking data, a spatiotemporal context model is constructed, considering the changes in the time series and the spatial distribution characteristics, and combining the historical data in the same or similar situations in the past period of time to predict the development trend of potential risk factors. According to the variables unique to the marine environment and their impact on the state of the injured and the rescue operation, the prediction results are adjusted to obtain the spatiotemporal context prediction results. Finally, based on the spatiotemporal context prediction results, a comprehensive description including detailed location, action, potential risks and other factors is formed to obtain the emergency scene description after enhanced analysis.

[0103] Here is a specific example:

[0104] In a maritime ship collision, the system first receives image and audio information from multiple sensors (such as ship-borne cameras, drone cameras, and shipboard microphones). After preliminary processing (noise reduction, format conversion), all data are synchronized to the same timestamp, forming a unified time-synchronized data stream.

[0105] Next, the system applied a multimodal perceptual fusion algorithm to extract visual features in the image (such as the color and shape of the ship debris) and acoustic features in the audio (such as the frequency and loudness of the distress call). Through the attention mechanism, the system successfully associated the floating wreckage in the image with the distress call in the audio, and constructed a complete representation of the emergency scene.

[0106] The system then used object detection and tracking technology in a deep learning framework to automatically identify the locations of multiple drowning people and track their movement paths in real time. At the same time, the system also detected floating debris and other obstacles that might hinder rescue operations.

[0107] In order to predict the development trend of potential risk factors, the system built a spatiotemporal context model, combined with historical data of similar situations in the past period of time, and predicted the trend of sea waves and their impact on rescue operations. In particular, the system introduced marine environmental variables such as sea temperature, salinity and wind and wave levels as correction factors to adjust the prediction results to make them closer to the actual situation.

[0108] Finally, the system generated a detailed description of the emergency scene, including the specific location of the drowning person, the distribution of floating debris, and the changing trend of the waves, helping the rescue team to better plan the route of action and response measures. This comprehensive and accurate scene analysis greatly improves the success rate and efficiency of rescue.

[0109] In order to solve the accuracy and real-time problems of cross-modal data fusion, in some embodiments, the method in step 102 is based on the pre-processed multimedia data, and a multimodal perception fusion algorithm is applied to jointly analyze the image and audio information, extract cross-modal features, and identify and associate objects in the image with events in the audio through an attention mechanism or a similarity matching algorithm, so as to construct a complete and semantically rich scene representation, and obtain a fusion analysis result, including:

[0110] Using the preprocessed multimedia data, feature extraction is performed on the image and audio information respectively, visual features are extracted from the image, and acoustic features are extracted from the audio to obtain a preliminary modal feature set; based on the preliminary modal feature set, a multimodal perception fusion algorithm is applied to map the feature vectors of the image and audio into a shared feature space to facilitate direct comparison and association between different modalities. The multimodal perception fusion algorithm is a model based on deep learning and generates a comprehensive feature representation; based on the comprehensive feature representation, an attention mechanism algorithm is used to calculate the correlation score between the object detected in the image and the audio clip, identify and associate the object in the image with the event in the audio, and obtain a cross-modal association result; based on the cross-modal association result, a dynamic scene graph containing time series information is further constructed, the dynamic scene graph describes the object position and action state at the current moment, records the change trend of entities over time and the interaction relationship between them, and forms a complete and semantically rich scene representation; based on the complete scene representation, combined with contextual information, the key elements in the scene are semantically interpreted to obtain a fusion analysis result.

[0111] In this embodiment, the preliminary modal feature set is to achieve effective joint analysis of image and audio information. The system first extracts visual features and acoustic features from the preprocessed multimedia data. Visual features include color, texture, shape, etc., which are used to describe objects in the image; acoustic features include frequency, pitch, loudness, etc., which are used to describe events in the audio. These features constitute the preliminary modal feature set, which provides a basis for subsequent multimodal perception fusion. The shared feature space mapping is to facilitate direct comparison and association between different modalities. The system maps the feature vectors of images and audio to a shared feature space. This process enables data originally belonging to different modalities to be compared and associated in the same space, thereby improving the effect of feature fusion. The construction of a shared feature space usually relies on a deep learning model, such as a multimodal embedding network or a multi-task learning model. The comprehensive feature representation is a comprehensive feature representation generated in a shared feature space, which integrates multiple information from images and audio to form a richer and more comprehensive feature representation. This comprehensive feature representation not only contains the information of the original modality, but also enhances the complementarity and correlation between different modalities, and improves the accuracy of overall scene understanding. The attention mechanism and correlation score calculation are used to identify and associate objects in images with events in audio. The system uses the attention mechanism algorithm to calculate the correlation score between objects detected in images and audio clips. By measuring the similarity and correlation strength between the two, the system can effectively identify the correspondence between objects in images and events in audio, and form cross-modal correlation results. The dynamic scene graph is based on the cross-modal correlation results. The system further constructs a dynamic scene graph containing time series information. The graph describes the position and action state of objects at the current moment, and records the changing trends of entities over time and their interactions. The dynamic scene graph not only provides a static scene description, but also captures the dynamic changes of the scene, enhancing the understanding and prediction capabilities of emergency situations. The semantic interpretation and fusion analysis results are that the system semantically interprets the key elements in the scene based on the complete scene representation and combined with contextual information (such as environmental background, historical data, etc.). This process generates the final fusion analysis results through natural language processing technology and pre-trained deep learning models to provide support for subsequent decision-making.

[0112] In the embodiment of the present application, first, the pre-processed multimedia data is used to extract features of the image and audio information respectively, and visual features (such as color, texture, shape) are extracted from the image, and acoustic features (such as frequency, pitch, loudness) are extracted from the audio to obtain a preliminary modal feature set. Secondly, based on the preliminary modal feature set, a multimodal perception fusion algorithm is applied to map the feature vectors of the image and audio into a shared feature space to facilitate direct comparison and association between different modalities. The multimodal perception fusion algorithm is a model based on deep learning to generate a comprehensive feature representation. Then, based on the comprehensive feature representation, the attention mechanism algorithm is used to calculate the correlation score between the object detected in the image and the audio clip, identify and associate the object in the image with the event in the audio, and obtain a cross-modal association result. Furthermore, based on the cross-modal association result, a dynamic scene graph containing time series information is further constructed, which describes the position and action state of the object at the current moment, records the trend of the entity over time and the interaction between them, and forms a complete and semantically rich scene representation. Finally, based on the complete scene representation and combined with contextual information (such as environmental background, historical data, etc.), the key elements in the scene are semantically interpreted to generate the final fusion analysis results to provide support for subsequent decision-making.

[0113] Here is a specific example:

[0114] In a maritime ship collision, the system first receives image and audio information from multiple sensors (such as ship-borne cameras, drone cameras, and shipboard microphones). After preliminary processing (noise reduction, format conversion), all data are synchronized to the same timestamp, forming a unified time-synchronized data stream.

[0115] Next, the system extracts features from the image and audio information respectively, extracting visual features from the image (such as the color and shape of the ship debris) and acoustic features from the audio (such as the frequency and loudness of the call for help). These features constitute a preliminary modal feature set, which provides a basis for subsequent multimodal perception fusion.

[0116] The system then maps the feature vectors of the image and audio into a shared feature space, allowing data originally belonging to different modalities to be compared and associated in the same space. Through deep learning models (such as multimodal embedding networks), the system generates a comprehensive feature representation that incorporates multiple information from images and audio to form a richer and more comprehensive feature representation.

[0117] Next, the system uses an attention mechanism algorithm to calculate the correlation score between the objects detected in the image (such as drowning people, floating debris) and the audio clips (such as cries for help), identify and associate the objects in the image with the events in the audio, and obtain cross-modal association results. For example, the system successfully associated the drowning person in the image with the cries for help in the audio, confirming their location and status.

[0118] Based on the cross-modal association results, the system further constructs a dynamic scene graph containing time series information, which describes the location and action status of objects at the current moment, and records the changing trends of entities over time and their interactions. The dynamic scene graph not only provides a static scene description, but also captures the dynamic changes of the scene, helping the rescue team to better plan action routes and response measures.

[0119] Finally, the system semantically interprets the key elements in the scene based on the complete scene representation and contextual information (such as marine environmental conditions and historical rescue data), generating the final fusion analysis results. For example, the system recommends the use of specific insulation equipment and anti-cold drugs to ensure the effectiveness of the rescue operation and the safety of the injured.

[0120] In this way, the system not only improves its ability to understand complex maritime emergency scenarios, but also significantly improves the speed and accuracy of rescue responses.

[0121] In order to solve the real-time prediction and response capabilities of potential risk factors in complex rescue scenarios, in some embodiments, the spatiotemporal context model is constructed based on the target detection and tracking data in step 102, and the changes in time series and spatial distribution characteristics are considered. The historical data of the same or similar situations in the past period of time are combined to predict the development trend of potential risk factors, and the prediction results are adjusted according to the variables unique to the marine environment and their impact on the status of the injured and the rescue operation, so as to obtain the spatiotemporal context prediction results, including:

[0122] Using the target detection and tracking data, the position and motion status of the injured individuals, rescuers, rescue equipment, and obstacles identified in the image are continuously monitored, and the changing trajectories of the injured individuals, rescuers, rescue equipment, and obstacles over time are recorded to form time series data, and dynamic position and motion status records are obtained; based on the dynamic position and motion status records, a spatiotemporal context model is constructed, which takes into account the evolution of the position and motion status of the object in the time dimension, analyzes the distribution characteristics in space, so as to capture the complete information of the dynamic scene, and generates a spatiotemporal context model; according to the spatiotemporal context model, combined with the same or similar The historical data in the context of the situation are used to identify typical patterns and development trends through deep learning algorithms to obtain historical pattern and development trend analysis results; based on the historical pattern and development trend analysis results, according to the variables unique to the marine environment, the impact of these variables on the current status of the injured and rescue operations is evaluated, and environmental correction parameters are introduced to adjust the parameter settings of the prediction model to ensure that the prediction results are more in line with the actual environment, and an adjusted prediction model is obtained; based on the adjusted prediction model, the development trend of potential risk factors is predicted, including the deterioration of the health status of the injured, increased difficulty in rescue, and intensified environmental risks, and a spatiotemporal context prediction result containing detailed location, action, and potential risk factors is generated.

[0123] In this embodiment, target detection and tracking data refers to image and video information obtained and processed by sensors (such as cameras, radars, etc.), which contains information about the position and motion state of specific targets (such as injured individuals, rescuers, rescue equipment, obstacles, etc.). These data are used to continuously monitor the changing trajectories of these objects over time to form time series data. The spatiotemporal context model aims to capture the dynamic evolution of objects in time and space dimensions. It not only records the position changes of objects, but also analyzes the evolution of their motion states over time and the distribution characteristics in space. This model can provide a more complete and dynamic scene description, which is crucial for understanding the development of events in complex environments. The results of historical pattern and development trend analysis are that through deep learning algorithms, the system identifies typical patterns and development trends in the same or similar situations over a period of time in the past. This helps to predict what may happen in the current situation, so as to prepare countermeasures in advance. Historical pattern and development trend analysis is based on the learning results of a large amount of historical data, which provides a scientific basis for prediction. Environmental correction parameters take into account variables unique to the marine environment (such as wind speed, wave height, water flow direction, etc.), which may significantly affect the state of the injured and the effectiveness of rescue operations. Therefore, introducing environmental correction parameters when constructing a prediction model can ensure that the prediction results are more in line with the actual environment and improve the accuracy and reliability of the prediction.

[0124] In the embodiment of the present application, firstly, the target detection and tracking data is used to continuously monitor the position and motion state of the injured individuals and rescuers, rescue equipment, and obstacles identified in the image, and the change trajectory of these objects over time is recorded to form time series data, and the dynamic position and motion state records are obtained. Secondly, based on the dynamic position and motion state records, a spatiotemporal context model is constructed, which takes into account the evolution of the position and motion state of the object in the time dimension, analyzes the distribution characteristics in space, and captures the complete information of the dynamic scene to generate a spatiotemporal context model. Then, according to the spatiotemporal context model, combined with historical data in the same or similar situations over a period of time in the past, a deep learning algorithm is used to identify typical patterns and development trends, and historical pattern and development trend analysis results are obtained. Further, based on the historical pattern and development trend analysis results, according to the variables unique to the marine environment, the impact of these variables on the current state of the injured and the rescue operation is evaluated, and environmental correction parameters are introduced to adjust the parameter settings of the prediction model to ensure that the prediction results are more in line with the actual environment, and an adjusted prediction model is obtained. Finally, based on the adjusted prediction model, the development trend of potential risk factors is predicted, including the deterioration of the health status of the injured, the increased difficulty of rescue, and the intensification of environmental risks, and a spatiotemporal context prediction result containing detailed location, action, and potential risk factors is generated.

[0125] Here is a specific example:

[0126] In a marine disaster rescue scenario, the system first receives target detection and tracking data from drones, satellite images, and other ground sensors. These data contain information about the location and motion status of injured individuals, rescuers, rescue equipment, and obstacles (such as floating debris).

[0127] Next, the system continuously monitors these targets and records their changing trajectories over time to form time series data. For example, the system detects that a casualty is gradually pushed away from the shore by the current, while rescuers are trying to approach him. At the same time, the system also records the movement of nearby rescue ships and life rafts, as well as obstacles in the surrounding sea area.

[0128] Then, based on these dynamic position and action state records, the system builds a spatiotemporal context model that not only captures the evolution of objects in the time dimension, but also analyzes their distribution characteristics in space. For example, the model shows the changes in the distance between the injured and the rescuers over time, as well as the potential obstruction of the rescue path by obstacles.

[0129] Furthermore, the system combines historical data from similar rescue missions in the past and uses deep learning algorithms to identify typical patterns and trends. For example, under similar conditions in history, certain types of obstacles often lead to rescue delays, or under certain weather conditions, the physical condition of the injured may deteriorate rapidly.

[0130] Taking into account variables unique to the marine environment, such as wind speed and wave height, the system introduces environmental correction parameters to adjust the prediction model. For example, if strong winds are forecast to come, the system will predict how this will affect the status of the injured and rescue operations, and then adjust its prediction model to adapt to the new environmental conditions.

[0131] Finally, based on the adjusted prediction model, the system predicts the development trend of potential risk factors, including the possibility of worsening health conditions of the injured, increased difficulty in rescue, and increased environmental risks. The system generates detailed spatiotemporal contextual prediction results, indicating which areas are at higher risk and when and where special precautions should be taken. For example, the system recommends giving priority to dispatching rescue equipment with wind resistance and reminding rescue teams to pay attention to upcoming severe weather so that rescue strategies can be adjusted in time to ensure the safety and efficiency of rescue operations.

[0132] In this way, the system not only improves the understanding and prediction of potential dangerous factors in complex rescue scenarios, but also significantly enhances the speed and accuracy of emergency response, providing strong support for successful rescue.

[0133] In order to solve the problem of accurate matching of historical cases and extraction of professional advice in complex marine emergency scenarios, in some embodiments, the key elements extracted from the emergency scenario description after enhanced analysis in step 103 are used to retrieve historical cases with high relevance from the cloud medical knowledge graph using a semantic similarity calculation algorithm combined with correction parameters for the impact of marine environment-specific variables on the injured, and dynamically optimize the retrieval strategy through a reinforcement learning algorithm to extract professional advice suitable for the current emergency scenario, including:

[0134] By using the emergency scene description after enhanced analysis, key elements are extracted, and the key elements include the status of the injured individual, the condition of the surrounding environment, and the development trend of any potential risk factors, so as to obtain a set of key elements; based on the set of key elements, a semantic similarity calculation algorithm is applied to compare the key elements of the current scene with the historical cases in the cloud medical knowledge graph. The semantic similarity calculation algorithm can identify and quantify the semantic similarity between different cases, find the historical case set that best matches the current situation, and generate a preliminary matching list of historical cases; combined with variables unique to the marine environment, environmental parameters such as seawater temperature, salinity, and wind and wave level are introduced as correction factors to evaluate the impact of the variables on the status of the injured and the rescue operation, and adjust the results calculated by the semantic similarity calculation algorithm to ensure that the retrieved historical cases have similarities in medical treatment and have reference in environmental adaptability. value, the preliminary matching historical case list is revised to obtain a revised matching historical case list; according to the revised matching historical case list, highly relevant historical cases are retrieved from the cloud medical knowledge graph, the historical cases contain similar medical treatment methods, and take into account successful first aid experiences and failed lessons in specific marine environments, to generate a highly relevant historical case set; based on the highly relevant historical case set, the retrieval strategy is dynamically optimized through a reinforcement learning algorithm, the reinforcement learning algorithm continuously improves itself according to the newly added data, improves the accuracy and efficiency of future case retrieval, and obtains an optimized retrieval strategy; according to the optimized retrieval strategy, the retrieved highly relevant historical cases are further analyzed to extract professional suggestions suitable for the current first aid scenario, the professional suggestions integrate the successful experience of historical cases, best practices and the latest medical research results.

[0135] In this embodiment, the key element set refers to the key information extracted from the emergency scene description after enhanced parsing, including the status of the injured individual (such as injury type, physiological state), the condition of the surrounding environment (such as weather conditions, sea characteristics) and the development trend of any potential risk factors (such as wave changes, wind speed increase). These key elements are used to build a comprehensive description framework to ensure the accuracy of subsequent analysis and retrieval. The semantic similarity calculation algorithm aims to identify and quantify the semantic similarity between different cases. Through natural language processing technology and machine learning models, the system can compare the key elements of the current scene with historical cases in the cloud medical knowledge graph to find the most matching historical case set. This algorithm not only considers the similarity of text content, but also combines multimodal data (such as images, audio) for comprehensive evaluation, which improves the accuracy and reliability of matching. The environmental correction factor takes into account the variables unique to the marine environment (such as sea water temperature, salinity, wind and wave level), which may significantly affect the status of the injured and the effectiveness of the rescue operation. Therefore, the introduction of the environmental correction factor when applying the semantic similarity calculation algorithm can ensure that the retrieved historical cases have similarities in medical treatment and are also of reference value in environmental adaptability. A highly relevant historical case set refers to a set of highly relevant historical cases retrieved from the cloud medical knowledge graph after environmental correction. These cases not only contain similar medical treatment methods, but also take into account successful first aid experiences and failed lessons in specific marine environments, providing valuable experience for current first aid scenarios. Reinforcement learning algorithms are used to dynamically optimize retrieval strategies, continuously self-improve based on newly added data, and improve the accuracy and efficiency of future case retrieval. Through reinforcement learning, the system can better adapt to changing first aid needs and provide more personalized and effective professional advice.

[0136] In the embodiment of the present application, first, the emergency scene description after enhanced parsing is used to extract key elements, which include the state of the injured individual, the state of the surrounding environment, and the development trend of any potential risk factors, to obtain a set of key elements. Secondly, based on the set of key elements, a semantic similarity calculation algorithm is applied to compare the key elements of the current scene with the historical cases in the cloud medical knowledge map, identify and quantify the semantic similarity between different cases, find the historical case set that best matches the current situation, and generate a preliminary matching historical case list. Then, combined with the variables unique to the marine environment (such as sea water temperature, salinity, wind and wave level), these environmental parameters are introduced as correction factors to evaluate their impact on the state of the injured and the rescue operation, adjust the results of the semantic similarity calculation algorithm, ensure that the retrieved historical cases have similarities in medical treatment and have reference value in environmental adaptability, and correct the preliminary matching historical case list to obtain a corrected matching historical case list. Further, according to the corrected matching historical case list, highly relevant historical cases are retrieved from the cloud medical knowledge map, which not only contain similar medical treatment methods, but also consider successful first aid experience and failure lessons in a specific marine environment, and generate a highly relevant historical case set. Furthermore, based on a highly relevant historical case set, the retrieval strategy is dynamically optimized through a reinforcement learning algorithm. The algorithm continuously improves itself based on the newly added data to improve the accuracy and efficiency of future case retrieval and obtain an optimized retrieval strategy. Finally, based on the optimized retrieval strategy, the retrieved highly relevant historical cases are further analyzed to extract professional advice applicable to the current emergency situation. These professional advices combine the successful experience of historical cases, best practices, and the latest medical research results to provide reliable guidance and support for on-site emergency personnel.

[0137] Here is a specific example:

[0138] In a complex maritime ship collision accident, the system first received an emergency scenario description after enhanced analysis, which included a drowning person's hypothermia, difficulty breathing and other injury conditions, environmental conditions such as strong winds and waves and low sea temperature in the nearby sea area, and potential dangerous factors such as the gradually intensifying waves.

[0139] Next, the system extracted these key elements and formed a set of key elements. Then, it applied a semantic similarity calculation algorithm to compare the key elements of the current scene with historical cases in the cloud medical knowledge graph and found a list of preliminary matching historical cases. For example, the system found some similar drowning rescue cases in history, which contained similar physical conditions and environmental conditions.

[0140] In order to ensure that the retrieved historical cases have similar medical treatments and reference value in terms of environmental adaptability, the system introduces marine environment-specific variables such as seawater temperature, salinity, and wind and wave levels as correction factors. By evaluating the impact of these variables, the system adjusts the results of the semantic similarity calculation algorithm, corrects the preliminary matching historical case list, and obtains the corrected matching historical case list.

[0141] Based on the revised matching historical case list, the system retrieved highly relevant historical cases from the cloud medical knowledge graph. These cases not only include similar medical treatment methods (such as the use of thermal insulation equipment and anti-cold drugs), but also take into account successful first aid experiences and failed lessons in specific marine environments. For example, the system found a historical case in which the rescue team successfully used a specific type of thermal insulation equipment, which significantly improved the survival rate of drowning victims.

[0142] To further optimize the search strategy, the system applied a reinforcement learning algorithm to continuously improve itself based on newly added data, thereby improving the accuracy and efficiency of future case searches. For example, the system learned how to more effectively identify and utilize certain specific environmental variables, thereby improving the accuracy of searches.

[0143] Finally, the system further analyzed the retrieved highly relevant historical cases based on the optimized search strategy and extracted professional suggestions applicable to the current emergency situation. These suggestions combine the successful experience of historical cases, best practices and the latest medical research results to provide detailed guidance for first responders on the scene. For example, the system recommends the immediate use of specific insulation equipment and the adoption of specific first aid measures to cope with the impact of low temperature environments on the injured.

[0144] In this way, the system not only improves the ability to accurately match historical cases in complex maritime emergency scenarios, but also significantly enhances the practicality and reliability of professional advice, providing strong support for successful rescue.

[0145] In order to solve the problem of accurate matching of historical cases in complex emergency scenarios, in some embodiments, the key elements of the current scenario are compared with the historical cases in the cloud medical knowledge graph based on the key element set in step 103 by applying a semantic similarity calculation algorithm, identifying and quantifying the semantic similarity between different cases, finding the historical case set that best matches the current situation, and generating a preliminary matching historical case list, including:

[0146] Using the extracted key element set, the key elements are structuredly encoded to ensure that each element is accurately represented, and structured key element data is obtained; based on the structured key element data, a multidimensional feature vector is constructed, and the multidimensional feature vector contains information that directly describes the injury status and environmental conditions, time series information, and spatial distribution characteristics, so as to comprehensively reflect the characteristics of the current emergency scene and generate a multidimensional feature vector; based on the multidimensional feature vector, a semantic similarity calculation algorithm is applied to compare the multidimensional feature vector of the current scene with the feature vector of historical cases in the cloud medical knowledge graph, and the semantic distance or similarity score between the two is calculated through natural language processing technology and machine learning models, and historical cases with high similarity are identified to obtain a list of candidate cases; based on the candidate case list, combined with the semantic similarity score, the candidate cases are screened and sorted, and according to the set threshold or ranking rule, a number of historical cases that best match the current emergency scene are selected to form a preliminary matching list of historical cases.

[0147] In this embodiment, structured key element data refers to ensuring that each key element (such as patient status, environmental conditions, potential risk factors, etc.) extracted from the emergency scene description is accurately represented by structured coding. This coding method is not only convenient for subsequent processing, but also improves the consistency and comparability of information. The multidimensional feature vector contains multiple dimensions, directly describing the information of the patient status and environmental conditions, time series information (such as the change of events over time), and spatial distribution characteristics (such as the position relationship of objects in space). These dimensions together constitute a data structure that fully reflects the characteristics of the current emergency scene, enabling the system to more accurately capture and analyze the dynamic changes of the scene. The semantic distance or similarity score is calculated by natural language processing technology and machine learning models. The semantic distance or similarity score between the multidimensional feature vector of the current scene and the feature vector of the historical case in the cloud medical knowledge graph. This step is used to quantify the similarity between different cases, so as to identify historical cases with high similarity. The candidate case list refers to a group of historical cases that may be highly similar to the current emergency scene that are initially screened out after semantic similarity calculation. These cases are selected based on the similarity score between their feature vectors and the feature vector of the current scene, providing a basis for further screening. The preliminary matching historical case list is a final list of several historical cases that best match the current emergency scenario. These cases are further screened and sorted to ensure that they have similarities in medical treatment and reference value in environmental adaptability.

[0148] In the embodiment of the present application, first, the key elements are structured encoded using the extracted key element set to ensure that each element is accurately represented and obtain structured key element data. For example, the state of the injured (such as body temperature, heart rate), environmental conditions (such as wind speed, wave height), potential risk factors (such as current direction) and other information are converted into a standardized format for subsequent processing. Secondly, based on the structured key element data, a multidimensional feature vector is constructed. This vector contains information that directly describes the state of the injured and the environmental conditions, as well as time series information (such as the change of events over time) and spatial distribution characteristics (such as the position relationship of objects in space) to fully reflect the characteristics of the current emergency scene and generate a multidimensional feature vector. Then, according to the multidimensional feature vector, a semantic similarity calculation algorithm is applied to compare the multidimensional feature vector of the current scene with the feature vector of the historical case in the cloud medical knowledge graph. The semantic distance or similarity score between the two is calculated by natural language processing technology and machine learning models, and historical cases with high similarity are identified to obtain a list of candidate cases. Finally, based on the list of candidate cases, the candidate cases are screened and sorted in combination with the semantic similarity score. According to the set threshold or ranking rules, several historical cases that best match the current emergency scenario are selected to form a preliminary matching historical case list. For example, the system may set a minimum similarity threshold, and only cases exceeding this threshold will be selected into the preliminary matching list.

[0149] Here is a specific example:

[0150] In a complex maritime ship collision accident, the system first received an emergency scenario description after enhanced analysis, which included a drowning person's hypothermia, difficulty breathing and other injury conditions, environmental conditions such as strong winds and waves and low sea temperature in the nearby sea area, and potential dangerous factors such as the gradually intensifying waves.

[0151] Next, the system extracted these key elements and encoded them in a structured manner. For example, the system converted the patient's physiological indicators such as body temperature and heart rate, as well as environmental conditions (such as wind speed and wave height) into a standardized format to ensure that each element was accurately represented and structured key element data was obtained.

[0152] Then, based on these structured key element data, the system constructs a multi-dimensional feature vector. This vector not only contains information that directly describes the patient's status and environmental conditions, but also includes time series information (such as changes in the patient's status over time) and spatial distribution characteristics (such as the positional relationship between rescue equipment and obstacles), which fully reflects the characteristics of the current emergency scene.

[0153] In order to find the historical cases that best match the current scenario, the system applies a semantic similarity calculation algorithm to compare the multi-dimensional feature vector of the current scenario with the feature vector of the historical cases in the cloud medical knowledge graph. Through natural language processing technology and machine learning models, the system calculates the semantic distance or similarity score between the two, identifies historical cases with high similarity, and forms a list of candidate cases.

[0154] Finally, the system screens and sorts the candidate cases based on the candidate case list and the semantic similarity score. According to the set threshold or ranking rules, the system selects several historical cases that best match the current emergency scenario and forms a preliminary matching list of historical cases. For example, the system found some similar drowning rescue cases in history, which included similar physical conditions and environmental conditions. These cases not only provide valuable first aid experience and lessons, but also provide specific guidance and suggestions for on-site first aid personnel, such as the use of specific insulation equipment and cold-resistant drugs.

[0155] In this way, the system not only improves the ability to accurately match historical cases in complex emergency scenarios, but also significantly enhances the practicality and reliability of professional advice, providing strong support for successful rescue. In addition, the system continuously optimizes retrieval strategies through learning and self-improvement of historical data, improving the accuracy and efficiency of future case retrieval.

[0156] In order to solve the problem of efficient communication and accurate guidance between remote experts and on-site emergency personnel in complex emergency scenarios, in some embodiments, the remote expert system is activated based on the professional advice and the emergency scenario description after enhanced analysis in step 104, and interacts with the on-site emergency personnel through two-way video communication. Natural language processing technology is used to understand and generate dialogue content for specific situations, and a virtual reality or augmented reality simulation training module is used to reproduce the emergency scenario, so that the remote expert can provide guidelines based on the emergency scenario description and generate personalized guidance information, including:

[0157] The remote expert system is activated by using the professional advice and the enhanced parsed emergency scenario description; the remote expert system is connected to remote medical experts via a high-speed network to ensure that they obtain the latest emergency scene information in real time and obtain an activated remote expert system; based on the activated remote expert system, a two-way video communication link is established to achieve real-time interaction between remote experts and on-site emergency personnel. Through high-definition video streaming, remote experts can intuitively observe the emergency scene and communicate with emergency personnel in real time, obtaining a two-way video communication link and a real-time interactive environment; based on the two-way video communication link and the real-time interactive environment, natural language processing technology is applied to analyze the questions and statements of emergency personnel, and accurate and easy-to-understand responses are automatically generated to ensure the efficiency and accuracy of communication. Optimized dialogue content is obtained; combining the optimized dialogue content and emergency scenario description, a virtual reality or augmented reality simulation training module is used to reproduce the first aid scenario, wherein the training module helps the remote expert to have a deeper understanding of the on-site environment and simulate different first aid operations, so as to provide intuitive guidance to on-site first aid personnel and generate a simulated first aid scenario; based on the simulated first aid scenario, the remote expert adjusts the guidelines according to the actual scenario, provides precise operational suggestions and decision support, and the operational suggestions take into account the current state of the injured, environmental conditions and lessons learned from historical cases, and obtains optimized guidelines; according to the optimized guidelines, all relevant information is integrated, including the professional advice, emergency scenario description, optimized dialogue content and simulated first aid scenario, to generate personalized guidance information.

[0158] In this embodiment, the remote expert system refers to a platform that connects remote medical experts through a high-speed network to ensure that they can obtain the latest emergency scene information in real time. The system not only provides the function of data transmission, but also integrates a variety of technical support tools, such as high-definition video streaming, natural language processing, etc., to achieve efficient remote collaboration. The two-way video communication link refers to a real-time video communication channel based on a high-speed network, allowing remote experts to interact instantly with on-site emergency personnel. The high-definition video stream enables remote experts to intuitively observe the emergency scene and communicate face-to-face with emergency personnel, improving the efficiency and accuracy of communication. The optimized dialogue content is through natural language processing technology, the system analyzes the questions and statements of emergency personnel, and automatically generates accurate and easy-to-understand responses. This optimized dialogue content not only improves the efficiency of communication, but also ensures the accuracy of information transmission, avoiding misunderstandings and delays. The virtual reality or augmented reality simulation training module uses advanced graphics rendering technology and interactive design to help remote experts understand the on-site environment more deeply and simulate different emergency operations. Through virtual or augmented reality technology, remote experts can test different rescue plans in a realistic environment, thereby providing more intuitive and specific guidance to on-site emergency personnel. Personalized guidance information refers to the final guidance information generated after integrating all relevant information (including professional advice, emergency scenario descriptions, optimized dialogue content, and simulated emergency scenarios). This information not only includes precise operational advice and decision support, but also takes into account the current state of the patient, environmental conditions, and lessons learned from historical cases, ensuring the practicality and reliability of the guidance.

[0159] In the embodiment of the present application, first, the remote expert system is activated using professional advice and enhanced parsed emergency scene descriptions. The system connects remote medical experts through a high-speed network to ensure that they obtain the latest emergency scene information in real time and obtain an activated remote expert system. Secondly, based on the activated remote expert system, a two-way video communication link is established to achieve real-time interaction between remote experts and on-site emergency personnel. Through high-definition video streaming, remote experts can intuitively observe the emergency scene and communicate with emergency personnel in real time to obtain a two-way video communication link and a real-time interactive environment. Then, based on the two-way video communication link and the real-time interactive environment, natural language processing technology is applied to analyze the questions and statements of emergency personnel, automatically generate accurate and easy-to-understand responses, ensure the efficiency and accuracy of communication, and obtain optimized dialogue content. Further, combined with the optimized dialogue content and emergency scene description, virtual reality or augmented reality simulation training modules are used to reproduce the emergency scene. These training modules help remote experts to have a deeper understanding of the on-site environment and simulate different emergency operations in order to provide intuitive guidance to on-site emergency personnel and generate simulated emergency scenes. Furthermore, based on the simulated emergency scene, remote experts adjust guidelines according to actual scenes to provide precise operational advice and decision support. These operational recommendations take into account the current state of the injured, environmental conditions, and lessons learned from historical cases to obtain optimized guidelines. Finally, based on the optimized guidelines, all relevant information, including professional advice, emergency scenario descriptions, optimized dialogue content, and simulated first aid scenarios, is integrated to generate personalized guidance information. This information provides comprehensive and specific guidance for on-site first aid personnel to ensure the effectiveness and safety of rescue operations.

[0160] Here is a specific example:

[0161] In a complex maritime ship collision accident, the system first received professional advice and an enhanced analysis of the emergency scenario description, which included a drowning person's hypothermia, breathing difficulties and other injury conditions, environmental conditions such as strong winds and waves and low sea temperature in the nearby sea area, and potential dangerous factors such as the gradually intensifying waves.

[0162] Next, the system activated the remote expert system and connected a marine medicine expert through a high-speed network. The expert watched the accident scene in real time through the system's high-definition video stream and communicated with the first responders on the scene in real time, establishing a two-way video communication link and a real-time interactive environment.

[0163] To ensure efficient and accurate communication, the system applies natural language processing technology to analyze the questions and statements raised by first responders and automatically generates accurate and easy-to-understand responses. For example, when first responders asked how to deal with hypothermia in drowning victims, the system quickly generated detailed suggestions for warming measures.

[0164] Furthermore, the system combines optimized dialogue content and emergency scenario descriptions with a virtual reality simulation training module to reproduce the first aid scenario. This module helps remote experts gain a deeper understanding of the on-site environment and simulates different types of first aid operations, such as cardiopulmonary resuscitation and the use of thermal insulation equipment. Through these simulations, remote experts can intuitively see the effects of various operations, thereby providing more specific guidance to on-site first aid personnel.

[0165] Based on the simulated emergency scenarios, remote experts adjusted the guidelines according to the actual situation and provided precise operational advice and decision support. For example, experts recommended the immediate use of specific insulation equipment and specific first aid measures to cope with the impact of low temperature environments on the injured. In addition, experts also reminded emergency personnel to pay attention to the upcoming severe weather and take precautions in advance.

[0166] Finally, the system integrates all relevant information, including professional advice, emergency scenario descriptions, optimized conversation content, and simulated first aid scenarios, to generate personalized guidance information. This information provides comprehensive and specific guidance to first responders on the scene to ensure the effectiveness and safety of rescue operations. For example, the system generates a detailed rescue plan that lists the specific operating methods and precautions for each step, helping first responders to implement rescue quickly and effectively.

[0167] In this way, the system not only improves the communication efficiency between remote experts and on-site emergency personnel, but also significantly enhances the accuracy and success rate of rescue operations, providing strong support for successful rescue.

[0168] This application takes into account that in complex rescue and emergency response scenarios, especially in environments such as maritime first aid, an accurate understanding of the on-site situation is crucial for the rapid and effective implementation of rescue. Traditional single-modal data analysis methods (such as relying solely on images or audio) are difficult to fully capture the complexity and dynamic changes of emergency scenes. In order to overcome these limitations, a method based on a multimodal perception fusion algorithm has been developed, which aims to extract cross-modal features and construct a complete and semantically rich scene representation by jointly analyzing image and audio information. This method not only improves the ability to understand complex environments, but also significantly enhances the accuracy and real-time performance of decision support systems, providing reliable technical support for rescue operations. Therefore, a new optional solution is proposed, which includes:

[0169] Based on the preprocessed multimedia data, a multimodal perception fusion algorithm is applied to jointly analyze the image and audio information, extract cross-modal features, and identify and associate objects in the image with events in the audio through an attention mechanism or a similarity matching algorithm to construct a complete and semantically rich scene representation, and obtain fusion analysis results, including:

[0170] Using the preprocessed multimedia data, feature extraction is performed on the image and audio information respectively, visual features are extracted from the image, and acoustic features are extracted from the audio to obtain a preliminary modal feature set;

[0171] The comprehensive feature representation F is generated by the following formula combined :

[0172] F combined =W fusion ·(W image ·F image +W audio ·F audio +b)

[0173] Among them, the feature set F image and F audio is a set of feature vectors extracted from image and audio data respectively; W image and W audio are the weight matrices of image and audio features respectively; W fusion is the weight matrix of the fusion layer; b is the bias term;

[0174] Based on the comprehensive feature representation, the attention mechanism algorithm is used to calculate the correlation score between the objects detected in the image and the audio clips, identify and associate the objects in the image with the events in the audio, and obtain cross-modal association results;

[0175] The relevance score S(i, j) is calculated by the following formula:

[0176]

[0177] Where S(i, j) is a correlation score that measures the strength of association between image object i and audio segment j; i and j represent the index of the object in the image and the audio segment, respectively; score is a function that measures the similarity between the two; α and β are weight and exponent parameters for adjusting the similarity score; F combined (i) and F combined (j) are the comprehensive feature representations of the i-th image object and the j-th audio segment respectively; k represents the index of all audio segments, which is used to normalize the similarity score; ∑ k is a summation symbol that traverses all audio segment indexes k.

[0178] Based on the cross-modal association results, a dynamic scene graph containing time series information is further constructed. The dynamic scene graph describes the position and action state of the object at the current moment, records the change trend of the entities over time and the interaction between them, and forms a complete and semantically rich scene representation. The dynamic scene graph is updated by combining the state of the previous moment and the current cross-modal association results, and applying a set of pre-trained weights and bias items for optimization to generate a new graph.

[0179] Based on the complete scene representation and combined with contextual information, the key elements in the scene are semantically interpreted to obtain the fusion analysis result R; the semantic interpretation is optimized by combining the dynamic scene graph and contextual information and applying a set of pre-trained weights and bias items to generate the final fusion analysis result.

[0180] The following is a detailed explanation of each parameter:

[0181] F image : A set of feature vectors extracted from image data. Feature extraction is performed on images using a trained convolutional neural network or other deep learning model to obtain feature vectors that can represent the image content.

[0182] F audio : A set of feature vectors extracted from audio data. The audio signal is processed through a trained recurrent neural network model to capture its time series characteristics and generate feature vectors.

[0183] W image : The weight matrix of image features. This weight matrix is ​​automatically learned through the training process of the multimodal fusion model. It reflects the relative importance of image features in the comprehensive feature representation.

[0184] W audio : The weight matrix of audio features. This weight matrix is ​​also obtained through training, and it is responsible for further fusing the weighted image and audio features into a higher-level comprehensive feature representation.

[0185] b: is the bias term, which is a parameter automatically optimized during model training. It is used to adjust the baseline value of feature representation to ensure the rationality of model output.

[0186] S(i, j): Relevance score, which measures the strength of association between image object i and audio segment j. Calculated by the formula, this score reflects the degree of semantic similarity between the two.

[0187] F combined (i) and F combined(j): The comprehensive feature representation of the i-th image object and the j-th audio clip, respectively. These feature vectors are obtained through the multimodal perception fusion algorithm.

[0188] score: A function that measures the similarity between two comprehensive feature representations. A variety of similarity metrics can be used, such as cosine similarity, Euclidean distance, etc. The specific choice depends on the needs of the application scenario and the characteristics of the data.

[0189] α: The weight parameter for adjusting the similarity score. Usually determined by cross-validation or hyperparameter optimization methods. It is used to control the sensitivity of the similarity score so that the model can better adapt to the similarity measurement in different scenarios.

[0190] β: Exponential parameter used to adjust the distribution of similarity scores. Also determined by hyperparameter optimization method. Its role is to change the nonlinearity of similarity scores, thus affecting the final relevance score.

[0191] ∑ k : A summation symbol that traverses all audio segment indexes k and is used to normalize the similarity scores. This is done automatically during the calculation process to ensure that the sum of all relevance scores is 1, thus achieving the effect of probability distribution.

[0192] The following is an introduction to the design reasons of each sub-item:

[0193] W image ·F image and W audio ·F audio : These two items represent the results of image and audio features after being transformed by their respective weight matrices. They are designed to ensure that the features of each modality can be properly represented in the comprehensive feature representation, while taking into account that different modalities may contribute differently to specific tasks.

[0194] The weighted sum of image and audio features is used to create a unified feature space that includes both visual and auditory information, allowing the model to use both types of input to gain a more comprehensive understanding.

[0195] The purpose of the whole formula is to achieve the effective fusion of two different types of data, image and audio, to produce a richer and more accurate feature representation F combined The model first extracts features from images and audio independently. Then, by introducing the weight matrix W image and W audio, the model can adjust the features of different modalities according to their importance. Then, by simply adding the weighted features, the model constructs a common feature space to represent the visual and auditory information in the current scene. Finally, through an additional layer of transformation (W fusion ) and the bias term b, the model generates the final comprehensive feature representation F combined This representation combines all available information and aims to improve the performance of downstream tasks such as recognition, classification or understanding complex multimedia environments.

[0196] α·score(F combined (i), F combined (j)) β :This part calculates the adjusted similarity score. The weight α and exponent β are introduced to flexibly adjust the range and distribution of the similarity score to make it more in line with actual needs.

[0197] exp(α·score(F combined (i), F combined (j)) β ): The exponential function exp is used to convert the adjusted similarity score into a positive number, and the influence of high scores can be significantly amplified, making it easier for the model to identify the most relevant audio clips.

[0198] The purpose of the entire formula is to calculate the correlation score S(i, j) between the image object i and each audio clip j, so as to identify and associate the two and construct a complete cross-modal association result. Specifically, the original similarity score between the image object and the audio clip is first calculated by the similarity measurement function score. Then, by introducing the weight α and the exponent β, and applying the exponential function exp, the model adjusts and emphasizes the original similarity score, highlighting the importance of high similarity scores. Finally, by summing and normalizing the similarity scores of all candidate audio clips, the model generates a valid probability distribution in which each score represents the strength of association between the corresponding audio clip and the image object.

[0199] This approach not only improves the ability to understand complex multimodal scenes, but also significantly enhances the model's accuracy in identifying and associating image and audio information. In this way, the system can process multimedia data more efficiently and provide more accurate and meaningful analysis results.

[0200] Here is a specific example:

[0201] Assume that in a complex marine ship collision accident, the system receives image and audio information transmitted by multiple sensors (such as cameras and microphones). After preliminary processing, a synchronized time-stamped data stream is formed. The system needs to use a multimodal perception fusion algorithm to jointly analyze this data to construct a complete emergency scene representation and provide effective decision support.

[0202] The following is the numerical substitution and calculation process:

[0203] Feature extraction and generation of preliminary modal feature sets:

[0204] Assume that the feature vector set F extracted from the image image Contains 5 dimensions, each of which corresponds to a different visual feature (such as color, texture, shape, etc.);

[0205] For example, F image = [0.8, 0.6, 0.7, 0.9, 0.5];

[0206] Similarly, the feature vector set F extracted from the audio audio Contains 3 dimensions, each corresponding to a different acoustic feature (such as frequency, pitch, loudness);

[0207] For example, F audio = [0.4, 0.7, 0.6];

[0208] Comprehensive feature representation generation:

[0209] Set the weight matrix W image and W audio , and the bias term b;

[0210] For example, W image =[0.5 0.5 0.5 0.5 0.5] T , W audio =[0.6 0.6 0.6] T , b = 0.1;

[0211] Fusion layer weight matrix W fusion Set to [0.8 0.8] T ;

[0212] According to formula F combined =W fusion ·(W image ·F image +W audio ·F audio + b);

[0213] Compute the comprehensive feature representation:

[0214]

[0215] The calculation results are:

[0216]

[0217] Relevance score calculation:

[0218] Set the parameters α = 0.5, β = 2;

[0219] Assume that there are two image objects i and j, and their corresponding comprehensive feature representations are F combined (i) = [2.296, 2.296] and F combined (j) = [2.3, 2.3];

[0220] Using the formula Where k represents the index of all audio segments;

[0221] Assuming there is only one audio clip j, then ∑ k exp(α·score(F combined (i), F combined (k)) β )=exp(α·score(F combined (i), F combined (j)) β );

[0222] Calculate score(F combined (i), F combined (j)), here we simply use the Euclidean distance as the similarity function score (F combined (i), F combined (j))=∥F combined (i)-F combined (j)∥ 2 ;

[0223] Calculate the relevance score S(i, j):

[0224]

[0225] The dynamic scene graph is updated by combining the state of the previous moment and the current cross-modal association results, and applying a set of pre-trained weights and bias items to optimize and generate a new graph.

[0226] For example, suppose the graph state at the previous moment is G t-1 , the current cross-modal association result is A t , the pre-trained weights and bias terms are W graph and b graph, then the new graph G t It can be updated by the following formula:

[0227] G t =W graph ·(G t-1 +A t )+b graph

[0228] Assume G t-1 =[0.7, 0.8], A t =[0.2, 0.3], W graph =0.9, b graph =0.1, then:

[0229] G t = 0.9·([0.7, 0.8]+[0.2, 0.3])+0.1=0.9·[0.9, 1.1]+0.1=[0.9, 1.09]

[0230] By calculating the correlation score S(i, j), the system can accurately identify the association between objects in the image and the audio clips, thereby constructing a complete scene representation. By combining the state of the previous moment and the current cross-modal association results, the system can update the dynamic scene map in real time to ensure continuous monitoring and understanding of emergency scenes. The final fusion analysis result R provides a semantic interpretation of the key elements in the scene, providing reliable decision support for rescuers and ensuring the effectiveness and safety of rescue operations.

[0231] In this way, the system not only improves its ability to understand complex first aid scenarios, but also significantly enhances the practicality and reliability of professional advice, providing strong support for successful rescue.

[0232] This application takes into account that in medical emergency and emergency response scenarios, rapid and accurate diagnosis of the condition and formulation of treatment plans are crucial to improving rescue efficiency and patient survival rates. Traditional diagnostic methods rely on the doctor's experience and limited information on the scene, and it is difficult to fully consider the valuable experience in historical cases. In order to improve this situation, a method based on a semantic similarity calculation algorithm has been developed, which aims to identify and quantify the semantic similarities between different cases by comparing the key elements of the current emergency scenario with historical cases in the cloud medical knowledge graph, so as to find the most matching historical cases and provide strong support for decision-making. This method not only improves the ability to understand complex medical scenarios, but also significantly enhances the accuracy and real-time performance of the decision support system, providing reliable technical support for emergency operations. Therefore, a new optional solution is proposed, which includes:

[0233] Based on the key element set, a semantic similarity calculation algorithm is applied to compare the key elements of the current scenario with the historical cases in the cloud medical knowledge graph. The semantic similarity calculation algorithm can identify and quantify the semantic similarities between different cases, find the historical case set that best matches the current situation, and generate a preliminary matching historical case list, including:

[0234] Using the extracted key element set, structurally encoding the key elements to ensure that each element is accurately represented, thereby obtaining structured key element data;

[0235] Based on the structured key element data, a multi-dimensional feature vector is constructed, wherein the multi-dimensional feature vector includes information directly describing the patient status and environmental conditions, time series information, and spatial distribution characteristics, so as to fully reflect the characteristics of the current emergency scene and generate a multi-dimensional feature vector;

[0236] The multi-dimensional feature vector V of the current scene is constructed by the following formula: current :

[0237] V current =W feature ·(W transform ·D key +b transform )+b feature

[0238] Among them, W feature and W transform is the feature transformation weight matrix; b transform and b feature is the bias term; V current is the multi-dimensional feature vector of the current scene; D key It is structured key element data;

[0239] Based on the multi-dimensional feature vector, a semantic similarity calculation algorithm is applied to compare the multi-dimensional feature vector of the current scene with the feature vector of the historical case in the cloud medical knowledge graph, and the semantic distance or similarity score between the two is calculated through natural language processing technology and machine learning models to identify historical cases with high similarity and obtain a list of candidate cases;

[0240] The similarity score S is calculated by the following formula: similarity (i, j):

[0241]

[0242] Among them, S similarity(i, j) is a score measuring the semantic similarity between the current scene i and the historical case j; i represents the current scene; j and k represent the indexes of the historical cases, respectively; α and β are weight and exponential parameters for adjusting the similarity score; γ is a scaling factor; W align is the alignment weight matrix, which is used to adjust the degree of alignment between feature vectors; is the transpose of the feature vector of the current scene; V history (j) is the feature vector of historical case j; V history (k) represents the multi-dimensional feature vector of the kth historical case in the cloud medical knowledge graph;

[0243] Based on the candidate case list, combined with the semantic similarity score, the candidate cases are screened and sorted, and according to the set threshold or ranking rule, several historical cases that best match the current emergency scenario are selected to form a preliminary matching historical case list;

[0244] The following formula is used to form a preliminary matching historical case list L initial :

[0245]

[0246] Among them, L initial is a list of historical cases that are initially matched; T is a pre-set similarity threshold; T ′ is the time window length; w t is the time weight coefficient; θ is the comprehensive score threshold; j represents the index of the historical case; is a comprehensive score that takes into account the time decay effect; j| is used to limit which historical cases will be selected into the preliminary matching historical case list L initial middle.

[0247] The following is a detailed explanation of each parameter:

[0248] V current : The multi-dimensional feature vector of the current scene. Through the above formula, from the structured key element data D key Extracted and converted.

[0249] W feature and W transform : They are feature transformation weight matrices, which are used to adjust and transform feature representations. These weight matrices are automatically learned during model training and are designed to optimize the quality and effectiveness of feature representations.

[0250] b transform and b feature: Bias term, used to adjust the base value of feature representation. It is also automatically adjusted by the back propagation algorithm during model training to help the model better fit the data.

[0251] D key : Structured key element data, that is, a set of encoded key elements. It is obtained by preprocessing and structured encoding the original description to ensure that each key element is accurately represented.

[0252] S similarity (i, j): A score that measures the semantic similarity between the current scenario i and the historical case j. It is calculated using the formula and reflects the strength of the association between the two.

[0253] α and β: weight and exponent parameters that adjust the similarity score. Usually determined by cross-validation or hyperparameter optimization methods. Their role is to adjust the sensitivity and distribution of the similarity score.

[0254] γ: Scaling factor used to adjust the alignment between feature vectors. Determined by experiments or hyperparameter optimization methods to ensure that the comparison between feature vectors is more reasonable.

[0255] W align : Alignment weight matrix, used to adjust the degree of alignment between feature vectors. Automatically learned during model training to optimize the matching effect between feature vectors.

[0256] T: A pre-set similarity threshold used to screen candidate cases. It is set based on the needs and experience of the application scenario to ensure that the screened related cases have sufficient similarity.

[0257] T′: Time window length, used to consider the time decay effect. It is set according to the time characteristics of the specific task to ensure that the comprehensive score can reflect the impact of different time periods.

[0258] w t : Time weight coefficient, used to adjust the importance of similarity scores in different time periods. It can be set according to actual needs, and can be linearly decreasing or in other forms to simulate the time decay effect.

[0259] θ: Comprehensive score threshold, used to further screen candidate cases. It is set according to the needs and experience of the application scenario to ensure that the historical cases finally selected have a high comprehensive score.

[0260] The feature vector of the current scene is transposed and used to perform a dot product operation with the feature vector of the historical case. current Transpose to get.

[0261] V history(j): Feature vector of historical case j. Obtained from the cloud medical knowledge graph through the same feature extraction and transformation process.

[0262] V history (k): Multi-dimensional feature vector of the kth historical case in the cloud medical knowledge graph. Obtained through feature extraction and conversion process.

[0263] The comprehensive score takes into account the time decay effect and is used to evaluate the overall similarity of candidate cases over a period of time. It is calculated by weighted summation, where w t is the time weight coefficient, S similarity (i, j, t) is the similarity score at a specific time point.

[0264] L initial : A preliminary matching historical case list, including several historical cases that best match the current emergency scenario. The candidate cases are selected and sorted by the set threshold and comprehensive scoring rules.

[0265] Here is a specific example:

[0266] Assume that in a complex traffic accident rescue scenario, the system receives image and audio information transmitted by multiple sensors (such as cameras and microphones), and combines it with the descriptions provided by the first responders to form a set of key elements about the patient's status and environmental conditions. After preliminary processing, a synchronized timestamp data stream is formed. The system needs to use the above-mentioned semantic similarity calculation algorithm to jointly analyze these data to find the most matching historical cases and provide effective decision support.

[0267] The following is the numerical substitution and calculation process:

[0268] Assume that the extracted key elements include "injured bleeding", "confused consciousness", "location of the accident", etc., which are converted into structured data D key

[0269] Set the weight matrix W feature and W transform , and the bias term b transform and b feature ;

[0270]

[0271] Structured key element data D key It may be a vector containing multiple dimensions, such as D key =[0.7, 0.8, 0.9].

[0272] According to the formula

[0273] V current =W feature ·(W transform ·D key +b transform )+b feature ;

[0274] Calculate the multi-dimensional feature vector V of the current scene current :

[0275]

[0276] The calculation results are:

[0277]

[0278] Semantic similarity calculation:

[0279] Set parameters α=0.5, β=2, γ=0.9.

[0280] Assume that there are two historical cases j and k, and their corresponding multi-dimensional feature vectors are V history (j) = [3.8, 3.8] and V history (k) = [3.9, 3.9].

[0281] Using the formula Where i represents the current scene; j and k represent the indexes of historical cases respectively; W align is the alignment weight matrix, which is used to adjust the degree of alignment between feature vectors. For example,

[0282] calculate and

[0283]

[0284] Calculate the similarity score S similarity (i, j) and S similarity (i, k):

[0285]

[0286] Since the result of the exponential function can be very large, only the calculation steps are shown here. The actual value may cause overflow, so appropriate numerical stabilization techniques should be used in actual implementation to avoid this problem.

[0287] Based on the similarity score calculated above, the threshold T is set to 0.5 and the time window length is T ′ =5, comprehensive score threshold θ = 0.7, time weight coefficient w t =[0.9, 0.8, 0.7, 0.6, 0.5].

[0288] If S similarity (i, j)>T and Then case j will be selected into the preliminary matching historical case list L initial middle.

[0289] For example, if S similarity (i, j) = 0.6 satisfies S similarity (i, j)>T, and satisfy Then case j will be added to L initial .

[0290] By calculating the similarity score S similarity (i, j), the system can accurately identify the historical case that best matches the current situation, thus providing a reference for emergency personnel. Preliminary matching historical case list L initial It provides effective guidance for current emergency scenarios, helping emergency personnel make correct decisions more quickly. By continuously accumulating new cases and updating the cloud medical knowledge graph, the system can continuously improve its judgment ability and accuracy, and further support medical emergency work.

[0291] In this way, the system not only improves its ability to understand complex first aid scenarios, but also significantly enhances the practicality and reliability of professional advice, providing strong support for successful rescue.

[0292] Figure 2 The present application provides a schematic diagram of a structure of an artificial intelligence-based marine emergency diagnosis system for the above method, such as Figure 2 As shown, the system includes:

[0293] The multimedia data acquisition device 21 is configured to receive a real-time multimedia data stream from a marine emergency scene, wherein the multimedia data stream includes image and audio information.

[0294] The information processing unit 22 is configured to jointly analyze the image and audio information based on the multimedia data stream, locate the key elements of the injured individuals and the surrounding environment involved, predict the development trend of potential risk factors, combine the correction parameters of the impact of the variables unique to the marine environment on the injured, retrieve and extract professional suggestions applicable to the current emergency situation from historical cases in the cloud medical knowledge map, activate the remote expert system to provide guidelines, and generate personalized guidance information:

[0295] The analysis and prediction module 221 is used to apply a multimodal perception fusion algorithm to jointly analyze the image and audio information based on the multimedia data stream, and use the target detection and tracking technology in the deep learning framework to locate the key elements of the injured individuals and the surrounding environment, and use the spatiotemporal context model to predict the development trend of potential risk factors to obtain an enhanced parsed emergency scene description;

[0296] The correction and optimization module 222 is used to retrieve highly relevant historical cases from the cloud medical knowledge graph based on the key elements extracted from the emergency scene description after the enhanced analysis, using a semantic similarity calculation algorithm combined with correction parameters of the impact of variables unique to the marine environment on the injured, and dynamically optimize the retrieval strategy through a reinforcement learning algorithm to extract professional suggestions suitable for the current emergency situation;

[0297] The activation providing module 223 is used to activate the remote expert system based on the professional advice and the emergency scenario description after the enhanced analysis.

[0298] The cloud medical knowledge graph 23 is configured to store historical cases and current emergency scene processing information for subsequent retrieval at other emergency scenes.

[0299] The remote expert system 24 is configured to interact with on-site emergency personnel through a video communication display device, use a virtual reality or augmented reality simulation training module to reproduce the emergency scene, provide guidelines based on the emergency scene description, and generate personalized guidance information.

[0300] The video communication display device 25 is configured for interaction between the remote expert and the on-site emergency personnel, and for displaying the conversation content for a specific situation understood and generated using natural language processing technology to the remote expert, and for displaying personalized guidance information given by the remote expert to the on-site emergency personnel.

[0301] Figure 2 The artificial intelligence-based marine emergency diagnosis system can perform Figure 1 The implementation principle and technical effect of the artificial intelligence-based marine emergency diagnosis method described in the illustrated embodiment will not be described in detail. The specific manner in which each module and unit performs operations in the artificial intelligence-based marine emergency diagnosis system in the above embodiment has been described in detail in the embodiment of the method, and will not be described in detail here.

[0302] In one possible design, Figure 2 The artificial intelligence-based marine emergency diagnosis system of the embodiment shown can be implemented as a computing device, such as Figure 3 As shown, the computing device may include a storage component 31 and a processing component 32;

[0303] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 32 .

[0304] The processing component 32 is used to: receive a real-time multimedia data stream from a marine emergency rescue site, the multimedia data stream containing image and audio information; based on the multimedia data stream, apply a multimodal perception fusion algorithm to jointly analyze the image and audio information, and use the target detection and tracking technology in the deep learning framework to locate the key elements of the injured individuals and the surrounding environment involved, and use the spatiotemporal context model to predict the development trend of potential risk factors to obtain an enhanced parsed emergency scene description; based on the key elements extracted from the enhanced parsed emergency scene description, use a semantic similarity calculation algorithm combined with correction parameters for the impact of variables unique to the marine environment on the injured, retrieve highly relevant historical cases from the cloud medical knowledge map, and pass The retrieval strategy is dynamically optimized through a reinforcement learning algorithm to extract professional advice applicable to the current emergency situation. Based on the professional advice and the emergency scene description after the enhanced analysis, the remote expert system is activated to interact with the on-site emergency personnel through two-way video communication. Natural language processing technology is used to understand and generate dialogue content for specific situations. Virtual reality or augmented reality simulation training modules are used to reproduce the emergency scene, so that remote experts can provide guidelines based on the emergency scene description and generate personalized guidance information. The personalized guidance information is combined with the calibrated emergency personnel's perspective, and the decision support information provided by artificial intelligence is displayed in the emergency personnel's field of view, and the decision support information includes customized injury status assessment, recommended first aid steps and risk warnings.

[0305] The processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to perform the above method.

[0306] The storage component 31 is configured to store various types of data to support operations at the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0307] Of course, the computing device may also include other components, such as input / output interfaces, display components, communication components, etc.

[0308] The input / output interface provides an interface between the processing component and the peripheral interface module, which may be an output device, an input device, etc.

[0309] The communication component is configured to facilitate, among other things, wired or wireless communications between the computing device and other devices.

[0310] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. In this case, the computing device can refer to a cloud server, and the above-mentioned processing components, storage components, etc. can be basic server resources rented or purchased from the cloud computing platform.

[0311] The present application also provides a computer storage medium storing a computer program, wherein the computer program can achieve the above-mentioned Figure 1 The embodiment shown is an artificial intelligence-based auxiliary diagnosis method for marine first aid.

[0312] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0313] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0314] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0315] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An artificial intelligence-based marine emergency aid diagnosis method, characterized in that: include: Receiving a real-time multimedia data stream from a marine emergency scene, wherein the multimedia data stream includes image and audio information; Based on the multimedia data stream, a multimodal perception fusion algorithm is applied to jointly analyze the image and audio information, and at the same time, the target detection and tracking technology in the deep learning framework is used to locate the key elements of the injured individuals and the surrounding environment involved, and the development trend of potential risk factors is predicted using a spatiotemporal context model to obtain an enhanced parsed emergency scene description; Based on the key elements extracted from the emergency scene description after the enhanced analysis, a semantic similarity calculation algorithm is used in combination with correction parameters of the impact of variables unique to the marine environment on the injured, and historical cases with high relevance are retrieved from the cloud medical knowledge graph. The retrieval strategy is dynamically optimized through a reinforcement learning algorithm to extract professional advice applicable to the current emergency situation. Based on the professional advice and the emergency scenario description after the enhanced analysis, a remote expert system is activated to interact with the on-site emergency personnel through two-way video communication, natural language processing technology is used to understand and generate dialogue content for specific situations, and a virtual reality or augmented reality simulation training module is used to reproduce the emergency scenario, so that the remote expert can provide guidelines based on the emergency scenario description and generate personalized guidance information; The personalized guidance information is combined with the calibrated first aid personnel's perspective, and the decision support information provided by artificial intelligence is displayed in the first aid personnel's field of view. The decision support information includes customized injury status assessment, recommended first aid steps and risk warnings.

2. The method according to claim 1, characterized in that Based on the multimedia data stream, a multimodal perception fusion algorithm is applied to jointly analyze the image and audio information, and the target detection and tracking technology in the deep learning framework is used to locate the key elements of the injured individuals and the surrounding environment involved, and the spatiotemporal context model is used to predict the development trend of potential risk factors, so as to obtain an enhanced parsed emergency scene description, including: Using the received and synchronized multimedia data streams, the image and audio information received from multiple sensors is preliminarily processed, including noise reduction and format conversion operations, to ensure that the data of different modalities are aligned in timestamps, and obtain preprocessed multimedia data; Based on the preprocessed multimedia data, a multimodal perception fusion algorithm is applied to jointly analyze the image and audio information, extract cross-modal features, and identify and associate objects in the image with events in the audio through an attention mechanism or a similarity matching algorithm, so as to construct a complete and semantically rich scene representation and obtain a fusion analysis result; According to the fusion analysis results, a deep learning framework is used to detect targets in the image, automatically identify the injured individuals and rescuers, rescue equipment, and obstacles in the image, and track them in real time, continuously monitor position changes, and obtain target detection and tracking data; Based on the target detection and tracking data, a spatiotemporal context model is constructed, and the changes in time series and spatial distribution characteristics are considered. The historical data of the same or similar situations in the past period of time are combined to predict the development trend of potential risk factors. According to the variables unique to the marine environment and their impact on the status of the injured and the rescue operation, the prediction results are adjusted to obtain the spatiotemporal context prediction results; According to the spatiotemporal context prediction results, a comprehensive description including detailed location, action, and potential risk factors is formed to obtain an emergency scene description after enhanced analysis.

3. The method according to claim 2, characterized in that Based on the preprocessed multimedia data, a multimodal perception fusion algorithm is applied to jointly analyze the image and audio information, extract cross-modal features, and identify and associate objects in the image with events in the audio through an attention mechanism or a similarity matching algorithm to construct a complete and semantically rich scene representation, and obtain fusion analysis results, including: Using the preprocessed multimedia data, respectively extracting features from the image and the audio information, extracting visual features from the image, and extracting acoustic features from the audio, to obtain a preliminary modal feature set; Based on the preliminary modal feature set, a multimodal perceptual fusion algorithm is applied to map the feature vectors of the image and audio into a shared feature space to facilitate direct comparison and association between different modalities, wherein the multimodal perceptual fusion algorithm is a deep learning-based model that generates a comprehensive feature representation; According to the comprehensive feature representation, using an attention mechanism algorithm, a correlation score between the object detected in the image and the audio clip is calculated, the object in the image and the event in the audio are identified and associated, and a cross-modal association result is obtained; Based on the cross-modal association results, a dynamic scene graph containing time series information is further constructed, which describes the position and action state of objects at the current moment, records the changing trends of entities over time and the interactive relationships between them, and forms a complete and semantically rich scene representation; Based on the complete scene representation and combined with contextual information, the key elements in the scene are semantically interpreted to obtain fusion analysis results.

4. The method according to claim 2, characterized in that: Based on the target detection and tracking data, a spatiotemporal context model is constructed, and the changes in time series and spatial distribution characteristics are considered. The historical data under the same or similar situations in the past period of time are combined to predict the development trend of potential risk factors. According to the variables unique to the marine environment and their impact on the status of the injured and the rescue operation, the prediction results are adjusted to obtain the spatiotemporal context prediction results, including: Using the target detection and tracking data, the position and motion status of the injured individuals, rescuers, rescue equipment, and obstacles identified in the image are continuously monitored, and the changing trajectories of the injured individuals, rescuers, rescue equipment, and obstacles over time are recorded to form time series data, and dynamic position and motion status records are obtained; Based on the dynamic position and action state records, a spatiotemporal context model is constructed, wherein the spatiotemporal context model considers the evolution of the position and action state of the object in the time dimension, analyzes the distribution characteristics in space, so as to capture the complete information of the dynamic scene and generate the spatiotemporal context model; Based on the spatiotemporal context model, combined with historical data in the same or similar contexts over a period of time, typical patterns and development trends are identified through deep learning algorithms to obtain historical pattern and development trend analysis results; Based on the analysis results of the historical patterns and development trends, and according to the variables unique to the marine environment, the impact of these variables on the current state of the injured and the rescue operations is evaluated, and environmental correction parameters are introduced to adjust the parameter settings of the prediction model to ensure that the prediction results are more in line with the actual environment, thereby obtaining an adjusted prediction model; According to the adjusted prediction model, the development trend of potential risk factors is predicted, including the deterioration of the health condition of the injured, the increased difficulty of rescue, and the aggravation of environmental risks, and a spatiotemporal context prediction result including detailed location, action, and potential risk factors is generated.

5. The method according to claim 1, characterized in that According to the key elements extracted from the emergency scene description after the enhanced analysis, the semantic similarity calculation algorithm is combined with the correction parameters of the impact of the variables unique to the marine environment on the injured, and historical cases with high relevance are retrieved from the cloud medical knowledge graph. The retrieval strategy is dynamically optimized through the reinforcement learning algorithm to extract professional suggestions suitable for the current emergency situation, including: By using the emergency scene description after enhanced parsing, key elements are extracted, wherein the key elements include the status of the injured individual, the condition of the surrounding environment, and the development trend of any potential risk factors, and a key element set is obtained; Based on the key element set, a semantic similarity calculation algorithm is applied to compare the key elements of the current scenario with the historical cases in the cloud medical knowledge graph. The semantic similarity calculation algorithm can identify and quantify the semantic similarities between different cases, find the historical case set that best matches the current situation, and generate a preliminary matching historical case list; In combination with variables unique to the marine environment, environmental parameters such as seawater temperature, salinity, and wind and wave levels are introduced as correction factors to evaluate the impact of the variables on the patient status and rescue operations, and the results calculated by the semantic similarity calculation algorithm are adjusted to ensure that the retrieved historical cases have similarities in medical treatment and reference value in environmental adaptability, and the preliminary matching historical case list is corrected to obtain a corrected matching historical case list; According to the revised matching historical case list, highly relevant historical cases are retrieved from the cloud medical knowledge graph, wherein the historical cases contain similar medical treatment methods, and take into account successful first aid experience and failed lessons in a specific marine environment, so as to generate a highly relevant historical case set; Based on the highly relevant historical case set, the retrieval strategy is dynamically optimized through a reinforcement learning algorithm, and the reinforcement learning algorithm continuously improves itself according to the newly added data to improve the accuracy and efficiency of future case retrieval to obtain an optimized retrieval strategy; Based on the optimized search strategy, the highly relevant historical cases retrieved are further analyzed to extract professional advice applicable to the current emergency situation. The professional advice integrates the successful experience of historical cases, best practices and the latest medical research results.

6. The method according to claim 5, characterized in that Based on the key element set, the semantic similarity calculation algorithm is applied to compare the key elements of the current scene with the historical cases in the cloud medical knowledge graph. The semantic similarity calculation algorithm can identify and quantify the semantic similarities between different cases, find the historical case set that best matches the current situation, and generate a preliminary matching historical case list, including: Using the extracted key element set, structurally encoding the key elements to ensure that each element is accurately represented, thereby obtaining structured key element data; Based on the structured key element data, a multi-dimensional feature vector is constructed, wherein the multi-dimensional feature vector includes information directly describing the patient status and environmental conditions, time series information, and spatial distribution characteristics, so as to fully reflect the characteristics of the current emergency scene and generate a multi-dimensional feature vector; Based on the multi-dimensional feature vector, a semantic similarity calculation algorithm is applied to compare the multi-dimensional feature vector of the current scene with the feature vector of the historical case in the cloud medical knowledge graph, and the semantic distance or similarity score between the two is calculated through natural language processing technology and machine learning models to identify historical cases with high similarity and obtain a list of candidate cases; Based on the candidate case list, combined with the semantic similarity score, the candidate cases are screened and sorted, and according to the set threshold or ranking rule, several historical cases that best match the current emergency scenario are selected to form a preliminary matching historical case list.

7. The method according to claim 1, characterized in that Based on the professional advice and the emergency scenario description after enhanced analysis, the remote expert system is activated to interact with the on-site emergency personnel through two-way video communication, natural language processing technology is used to understand and generate dialogue content for specific situations, and a virtual reality or augmented reality simulation training module is used to reproduce the emergency scenario, so that the remote expert can provide guidelines based on the emergency scenario description and generate personalized guidance information, including: Utilizing the professional advice and the emergency scene description after enhanced analysis, a remote expert system is activated; the remote expert system is connected to remote medical experts via a high-speed network to ensure that they obtain the latest emergency scene information in real time and obtain an activated remote expert system; Based on the activated remote expert system, a two-way video communication link is established to realize real-time interaction between the remote expert and the on-site emergency personnel. Through high-definition video streaming, the remote expert can directly observe the emergency scene and communicate with the emergency personnel in real time, thus obtaining a two-way video communication link and a real-time interactive environment; Based on the two-way video communication link and real-time interactive environment, natural language processing technology is applied to analyze the questions and statements of the emergency personnel, and accurate and easy-to-understand responses are automatically generated to ensure the efficiency and accuracy of communication and obtain optimized conversation content; In combination with the optimized dialogue content and emergency scenario description, a virtual reality or augmented reality simulation training module is used to reproduce the emergency scenario, wherein the training module helps the remote expert to have a deeper understanding of the on-site environment and simulate different emergency operations, so as to provide intuitive guidance to on-site emergency personnel and generate a simulated emergency scenario; Based on the simulated emergency scenario, the remote expert adjusts the guidelines according to the actual scenario and provides precise operational suggestions and decision support. The operational suggestions take into account the current patient status, environmental conditions and lessons learned from historical cases to obtain optimized guidelines. According to the optimized guidelines, all relevant information is integrated, including the professional advice, emergency scenario descriptions, optimized dialogue content and simulated first aid scenarios, to generate personalized guidance information.

8. An artificial intelligence-based marine emergency diagnosis system, used to execute the method according to any one of claims 1 to 7, characterized in that: include: A multimedia data acquisition device configured to receive a real-time multimedia data stream from a marine emergency scene, wherein the multimedia data stream includes image and audio information; An information processing unit is configured to jointly analyze the image and audio information based on the multimedia data stream, locate the key elements of the injured individuals and the surrounding environment involved, predict the development trend of potential risk factors, combine the correction parameters of the impact of variables unique to the marine environment on the injured, retrieve and extract professional suggestions applicable to the current emergency situation from historical cases in the cloud medical knowledge map, activate the remote expert system to provide guidelines, and generate personalized guidance information; The cloud medical knowledge graph is configured to store historical cases and current emergency scene processing information for subsequent retrieval at other emergency scenes; A remote expert system configured to interact with first responders on site through a video communication display device, using a virtual reality or augmented reality simulation training module for reproducing first responder scenarios, providing guidelines based on emergency scenario descriptions, and generating personalized guidance information; The video communication display device is configured for interaction between a remote expert and first responders on site, and for displaying to the remote expert the conversation content for a specific situation understood and generated using natural language processing technology, and for displaying to the first responders on site personalized guidance information given by the remote expert.

9. The system according to claim 8, characterized in that The information processing unit comprises: An analysis and prediction module is used to jointly analyze the image and audio information based on the multimedia data stream using a multimodal perception fusion algorithm, and to locate the key elements of the injured individuals and the surrounding environment using target detection and tracking technology in a deep learning framework, and to predict the development trend of potential risk factors using a spatiotemporal context model to obtain an enhanced parsed emergency scene description; A correction and optimization module is used to retrieve highly relevant historical cases from the cloud medical knowledge graph based on the key elements extracted from the emergency scene description after the enhanced analysis, using a semantic similarity calculation algorithm combined with correction parameters of the impact of variables unique to the marine environment on the injured, and dynamically optimize the retrieval strategy through a reinforcement learning algorithm to extract professional suggestions suitable for the current emergency situation; An activation providing module is used to activate a remote expert system based on the professional advice and the emergency scenario description after the enhanced analysis.

Citation Information

Cited By

  • Emergency command rescue data distributed storage method and system based on artificial intelligence

    CN120492546A

  • Distributed storage method and system for emergency command and rescue data based on artificial intelligence

    CN120492546B

  • Emergency team real-time voice instruction analysis and collaborative optimization method and system

    CN120783747A