A method and system for image-text question answering and person swimming detection

By combining visual and text information for deep semantic reasoning through a graphic question-and-answer method, the limitations of existing image processing systems in abnormal event detection are overcome, intelligent question-and-answer and anomaly detection for swimming personnel are realized, the accuracy and flexibility of detection are improved, and the cost of manual participation is reduced.

CN120356139BActive Publication Date: 2025-09-16CHINA TOWER CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510805757.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-16
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

Existing image processing systems lack comprehensive image understanding and reasoning capabilities when detecting abnormal events in complex images, especially people swimming in illegal areas. They have difficulty coping with the diversity and suddenness of abnormal behaviors, and rely on large amounts of labeled data and high-cost manual annotation, and are unable to provide intelligent question-answering and rich contextual explanations.

Method used

It adopts a picture-text question-answering method, collects swimming scene images, historical records and user input questions, uses a large language model to generate behavioral hypotheses and features, combines visual and text information for deep semantic reasoning, generates reasoning chains and feeds back results, supports multimodal information processing and modular design, and can independently upgrade or replace functional modules to adapt to different environments.

Benefits of technology

It achieves more accurate anomaly detection and intelligent question-answering in human swimming detection, reduces the cost of manual participation, improves the flexibility and adaptability of the system, can identify anomalies that are difficult to detect with visual models, and optimizes model updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356139B_ABST
    Figure CN120356139B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image processing technology, and specifically discloses a method and system for image-text question-answering and swimming detection, comprising the following steps: collecting swimming scene images, scene history records, and reasoning chains of scene history records, collecting user input questions; generating associated behavioral features; judging whether the associated behavioral features are identical to the behavioral features of historical scene records; if it is found that the historical scene records are lacking and judgment cannot be made, generating a behavior feature question to be confirmed for the behavior feature that cannot be judged, searching for relevant behavior features in the swimming scene image, reasoning the behavior feature question to be confirmed, generating a reasoning chain to be confirmed, and outputting a reasoning chain result based on the reasoning chain of the scene history record with the greatest overlap between the reasoning chain to be confirmed and the scene history record; and feeding back the reasoning chain and the reasoning chain result to the user. This method can handle complex behavior analysis, realize automatic monitoring, intelligent question-answering, and anomaly detection in swimming scenes, and significantly improve the accuracy and flexibility of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a method and system for image-text question answering and person swimming detection. Background Art

[0002] Existing image processing systems typically focus on a single task, such as target detection, object recognition, or anomaly detection, and lack comprehensive image understanding and reasoning capabilities. They are particularly limited in detecting abnormal events in complex images, such as people swimming in illegal areas.

[0003] Automatically detecting abnormal events in images is one of the most valuable problems in image analysis. Traditional abnormal event detection methods require pre-defining the type of abnormal event to be detected and collecting a large number of manually labeled samples to build an abnormal event detection model. However, due to the broad and vague definition of abnormal behavior, it is difficult to pre-list all possible abnormal behavior types. Furthermore, most behaviors in daily life are normal, making it difficult to collect sufficient abnormal behavior samples for modeling and analysis.

[0004] Traditional image anomaly detection methods often rely on specific image features or labeled datasets. For example, methods based on convolutional neural networks (CNNs) use feature extraction to determine whether an image contains unusual swimming events. However, these methods rely on a large number of labeled normal and abnormal samples. In practice, unusual swimming events are often rare and unpredictable. Existing methods typically model normal behavior patterns by analyzing a large number of normal behavior samples. When the behavior pattern in an image cannot be explained by the established normal behavior model, the image is considered to contain abnormal behavior.

[0005] Furthermore, many current anomaly detection systems rely on learning models from expert-labeled object detection datasets to identify known anomaly types. While this approach can effectively detect certain specific anomalous behaviors, it struggles to cover all possible anomalies, especially when integrating cross-scenario and multimodal information. Furthermore, manual labeling is costly and struggles to cope with the diversity and suddenness of abnormal behaviors.

[0006] When applied to online image detection tasks, these anomaly detection methods typically require large amounts of labeled data to distinguish between normal and abnormal situations. However, image anomaly detection tasks are relatively simple, typically only able to indicate whether an anomaly exists in an image or mark the location of the anomaly, without being able to further intelligently answer questions about specific image details and provide richer context and explanation. Summary of the Invention

[0007] The purpose of the present invention is to provide a method for image-text question-answering and swimming detection, which overcomes the limitations of the existing technology in coping with complex behavior analysis, realizes automatic monitoring, intelligent question-answering and anomaly detection in swimming scenarios, and significantly improves the accuracy and flexibility of the system.

[0008] The purpose of the present invention can be achieved through the following technical solutions:

[0009] A method for image-text question answering and person swimming detection includes the following steps:

[0010] Collect swimming scene images, scene history records, and reasoning chains of scene history records, and collect user input questions;

[0011] Generate behavioral hypotheses based on user input questions, and generate associated behavioral features based on the behavioral hypotheses;

[0012] According to the swimming scene image, historical scene record and the reasoning chain of the historical scene record, determine whether the associated behavioral features are the same as the behavioral features of the historical scene record. If they are the same, output the reasoning chain result that is the same as the reasoning chain of the scene history record. If they are different, regenerate a new behavioral hypothesis until a judgment can be made and output the corresponding reasoning chain result. If it is found that there is a lack of historical scene records and it is impossible to make a judgment, generate a behavior feature problem to be confirmed for the behavior feature that cannot be judged; according to the behavior feature problem to be confirmed, search for relevant behavior features in the swimming scene image, and based on the search results and the behavior feature problem to be confirmed, infer the behavior feature problem to be confirmed and generate the reasoning chain to be confirmed; detect the degree of overlap between the reasoning chain to be confirmed and the reasoning chain of the scene history record, take the reasoning chain of the scene history record with the largest overlap and output the reasoning chain result;

[0013] Feedback the inference chain and the inference chain results to the user to complete the detection.

[0014] In a further embodiment, the method of collecting swimming scene images, scene history records, and reasoning chains of scene history records, and collecting user input questions includes the following steps:

[0015] The swimming scene images are collected through the camera; the scene history records and the reasoning chain of the scene history records are called by calling the model; and the user input questions are collected through the input model.

[0016] In a further solution, the calling model and the input model both adopt large language models.

[0017] In a further solution, the scene history record includes historical image records and historical conversation records, and the historical conversation records express the historical images through text records; the reasoning chain of the scene history record is a reasoning process related to the historical conversation through text records.

[0018] In a further solution, a method for generating a behavior hypothesis based on a user input question and generating associated behavior features based on the behavior hypothesis includes:

[0019] User questions are input into the large language model, and behavioral hypotheses are generated through the large language model, and associated behavioral features are generated based on the behavioral hypotheses.

[0020] In a further embodiment, the method for determining whether the associated behavior characteristics are the same as the behavior characteristics of the historical scene record based on the swimming scene image, the historical scene record, and the inference chain of the historical scene record includes:

[0021] The swimming scene image, historical scene record and the reasoning chain of the historical scene record are input into the large language model to generate the behavioral features of the historical scene record. The large language model is used to determine whether the associated behavioral features are the same as the behavioral features of the historical scene record.

[0022] In a further embodiment, the method further comprises:

[0023] For the problem of confirming behavioral characteristics, the inference chain results are judged manually by watching the real-time scene monitoring;

[0024] Confirm the correctness of the human judgment reasoning chain result and the reasoning chain result judged by coincidence, select the correct reasoning chain result and store it as the scene history record and the reasoning chain of the scene history record.

[0025] The present invention also provides a graphic question-and-answer and swimming detection system, comprising:

[0026] An acquisition module, used to acquire swimming scene images, scene history records, and reasoning chains of scene history records, and to acquire user input questions;

[0027] The associated behavior feature generation module is used to generate behavior hypotheses based on user input questions, and generate associated behavior features based on the behavior hypotheses;

[0028] A judgment module is used to judge whether the associated behavioral features are the same as the behavioral features of the historical scene records based on the swimming scene image, the historical scene record and the reasoning chain of the historical scene record; if they are the same, an inference chain result that is the same as the inference chain of the scene history record is output; if they are different, a new behavioral hypothesis is regenerated until a judgment can be made and the corresponding inference chain result is output; if it is found that there is a lack of historical scene records and it is impossible to make a judgment, a behavior feature question to be confirmed is generated for the behavior feature that cannot be judged; based on the behavior feature question to be confirmed, relevant behavior features in the swimming scene image are searched, and based on the search results and the behavior feature question to be confirmed, the behavior feature question to be confirmed is inferred to generate an inference chain to be confirmed; the degree of overlap between the inference chain to be confirmed and the inference chain of the scene history record is detected, and the inference chain of the scene history record with the largest overlap is taken to output the inference chain result;

[0029] The feedback module is used to feed back the reasoning chain and the reasoning chain results to the user to complete the detection.

[0030] In a further solution, the acquisition module is a large language model acquisition module, and the associated behavior feature generation module is a large language model feature generation module; the judgment module is a large language model judgment module, a multimodal question-answering model judgment module and a target detection model judgment module. The large language model judgment module is used to judge whether the associated behavior features are the same as the behavior features of the historical scene records based on the swimming scene image, the historical scene record and the reasoning chain of the historical scene record. If they are the same, the reasoning chain result that is the same as the reasoning chain of the scene history record is output. If they are different, a new behavior hypothesis is regenerated until a judgment can be made and the corresponding reasoning chain result is output. If a missing result is found, the result is returned. If there is a lack of historical scene records and judgment cannot be made, then a question about behavioral features to be confirmed is generated for the behavioral features that cannot be judged. The multimodal question-answering model judgment module is used to input the question about behavioral features to be confirmed, and according to the question about behavioral features to be confirmed, search for relevant behavioral features in the swimming scene image, and infer the question about behavioral features to be confirmed based on the search results and the question about behavioral features to be confirmed, and output the reasoning chain to be confirmed; the target detection model judgment module is used to input the reasoning chain to be confirmed, detect the degree of overlap between the reasoning chain to be confirmed and the reasoning chain of the scene historical records, take the reasoning chain of the scene historical records with the largest degree of overlap and output the reasoning chain result; the feedback module is a large language model feedback module.

[0031] In a further solution, the system further includes a selection and recording module, wherein the selection and recording module is used to manually determine the result of the reasoning chain by watching the real-time scene monitoring for the problem of confirming the behavioral characteristics;

[0032] Confirm the correctness of the human judgment reasoning chain result and the reasoning chain result judged by coincidence, select the correct reasoning chain result and store it as the scene history record and the reasoning chain of the scene history record.

[0033] Beneficial effects of the present invention:

[0034] This system collects user input questions and then infers scene images based on them. Because the input questions can include real-time questions related to swimming training or preset questions about anomaly detection, it can simultaneously perform swimming image-text question-and-answer (Q&A) and image anomaly detection, enabling more accurate anomaly detection and intelligent Q&A during the analysis of swimming images. By combining image and text information for analysis, the system not only relies on visual features but also incorporates deep semantic reasoning based on text descriptions, effectively improving the accuracy of anomaly detection. Compared to traditional single-visual detection, the system can more comprehensively understand the content of swimming images and identify anomalies that would be difficult to detect using visual models alone. Through intelligent Q&A, the system automatically answers specific questions about the content of swimming images, reducing the need for manual image analysis. This Q&A mechanism explains detected anomalies, helping operators quickly understand the details of the anomaly in the image, significantly reducing the burden of manual intervention. Anomaly detection is integrated with inference chaining: through inference chaining, the system can gradually answer complex questions in the image and identify potential anomalies during the inference process. This method can screen out image samples that are valuable for model updates, especially those abnormal scenes that are difficult for the current model to identify. These abnormal scene samples can also be used to further optimize historical scene records.

[0035] The system of the present invention adopts a modular design, which independently separates functional modules such as image-text question and answer, multimodal information processing, target detection and abnormal event detection. Each functional module can be independently upgraded or replaced according to different application scenarios, thereby enhancing the adaptability of the system in different environments. In addition, this modular design is convenient for integration with other existing systems, greatly improving the maintainability and scalability of the system, and can better meet the complex and diverse image analysis needs. In summary, the present invention provides a method for image-text question and answer of swimming personnel and image anomaly detection, which can effectively improve the efficiency of anomaly detection and significantly reduce the cost of manual participation. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0037] Figure 1 It is a structural schematic diagram of the present invention. DETAILED DESCRIPTION

[0038] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0039] A method for image-text question answering and person swimming detection includes the following steps:

[0040] S1, collect swimming scene images, scene history records and reasoning chains of scene history records, and collect user input questions;

[0041] S2. Generate behavioral hypotheses based on user input questions, and generate associated behavioral features based on the behavioral hypotheses;

[0042] S3. Based on the swimming scene image, the historical scene record, and the reasoning chain of the historical scene record, determine whether the associated behavioral features are the same as the behavioral features of the historical scene record. If they are the same, output the reasoning chain result that is the same as the reasoning chain of the scene history record. If they are different, regenerate a new behavioral hypothesis until a judgment can be made and output the corresponding reasoning chain result. If it is found that there is a lack of historical scene records and it is impossible to make a judgment, generate a behavior feature question to be confirmed for the behavior feature that cannot be judged; based on the behavior feature question to be confirmed, search for relevant behavior features in the swimming scene image, and based on the search results and the behavior feature question to be confirmed, infer the behavior feature question to be confirmed and generate the reasoning chain to be confirmed; detect the overlap between the reasoning chain to be confirmed and the reasoning chain of the scene history record, take the reasoning chain of the scene history record with the largest overlap and output the reasoning chain result;

[0043] S4. Feedback the inference chain and the inference chain results to the user to complete the detection.

[0044] In some embodiments, step S1 includes the following steps:

[0045] The swimming scene images are collected through the camera; the scene history records and the reasoning chain of the scene history records are called by calling the model; and the user input questions are collected through the input model.

[0046] The calling model and the input model both adopt a large language model.

[0047] The scene history record includes a historical image record and a historical conversation record, wherein the historical conversation record expresses the historical image through a text record; and the reasoning chain of the scene history record is a reasoning process related to the historical conversation through the text record.

[0048] The input swimming scene image can also first go through a series of image preprocessing steps, including image size normalization, denoising, and adaptive brightness and contrast adjustment. The result of preprocessing is high-quality, standardized image data, which is convenient for subsequent model analysis.

[0049] At the same time, the historical conversations and reasoning chains related to the swimming scene are converted into a structured text format. For example, a historical conversation might be recorded as "On May 1, 2024, swimmer A was detected as immobile in shallow water for over one minute. The system issued an alert, and the lifeguard confirmed that there was no anomaly." The reasoning chain can be organized as "Detecting prolonged inactivity → Querying historical records → Finding similar cases → Recommending manual review." This text data is organized in a unified JSON format for easy reading and processing. This type of anomaly detection can be performed in real time based on user-entered questions, improving anomaly detection and discovery efficiency. Users can also directly enter relevant anomaly detection questions. For example, after collecting user-entered questions from the large language model, the large language model generates abnormal behavior hypotheses, such as immobility, flapping, or disappearance. Taking immobility as an example, behavioral features are generated, such as the unchanged shape of the person's detection box or the minimal change in the position of the head detection box, the duration of inactivity, body posture (supine or prone), the presence of spontaneous breathing movements, and the presence of struggling or movement of limbs.

[0050] The standardized image data and structured text descriptions are fed into the judgment module, providing rich context for subsequent behavior recognition, anomaly reasoning, and intelligent question-answering. The judgment module determines whether the historical scenario contains the same behavioral characteristics. If so, for example, if the historical scenario depicts a person head-down underwater for a certain period of time, the inference chain for the immobility behavior is output as drowning, along with the associated inference chain. However, for other behavioral characteristics, such as head-up, the inference chain may contain uncertainty, potentially including results that indicate resting or drowning. Then a new behavior hypothesis is regenerated, such as whether there is a behavior of asking for help or whether there are other people approaching to help, and relevant behavior features are generated. Continue to judge. If the behavior features are the same as the historical scene records of swimmers who are resting around, the reasoning chain result of normal rest is output. If there is no behavior around and the same historical scene records cannot be matched, it is impossible to judge whether it is a normal rest or a prohibited state of waiting for rescue. Then a problem of unconfirmed behavior features is generated, such as whether the behavior feature of a swimmer standing still with his head up is a normal rest or waiting for rescue. Find the relevant behavior features in the swimming scene image and generate a reasoning chain to be confirmed. For example, the generated reasoning chain to be confirmed is: long-term stillness is detected → historical records are queried. → Found to be a similar case of waiting for rescue → Manual review is recommended; or detected to be stationary for a long time → query historical records → Found to be a similar case of waiting for rescue → Manual review is recommended; then the overlap between the reasoning chain to be confirmed and the recorded reasoning chain is judged. The overlap detection method can be to judge the overlap of relevant behavioral features. If there are many identical relevant behavioral features, it is judged that the overlap is high. The behavioral features can be used to determine whether they overlap through the target detection model, such as setting a confidence threshold such as 0.5, and a non-maximum suppression (NMS) parameter to remove detection frames with too high overlap, generate an overlap anomaly probability, etc., select the one with the lowest overlap anomaly probability as having a high overlap, and select the one with a high overlap anomaly probability as having no overlap, so as to ensure that the detection results are accurate and reliable.

[0051] In some embodiments, a method for generating a behavior hypothesis based on a user input question and generating associated behavior features based on the behavior hypothesis includes:

[0052] User questions are input into the large language model, and behavioral hypotheses are generated through the large language model, and associated behavioral features are generated based on the behavioral hypotheses.

[0053] In some embodiments, the method of determining whether the associated behavior characteristics are the same as the behavior characteristics of the historical scene records based on the swimming scene image, the historical scene record, and the inference chain of the historical scene records includes:

[0054] The swimming scene image, historical scene record and the reasoning chain of the historical scene record are input into the large language model to generate the behavioral features of the historical scene record. The large language model is used to determine whether the associated behavioral features are the same as the behavioral features of the historical scene record.

[0055] In some embodiments, the method further includes step S5: for the problem of confirming the behavioral characteristics, manually determining the result of the reasoning chain by watching the real-time scene monitoring;

[0056] Confirm the correctness of the human judgment reasoning chain result and the reasoning chain result judged by coincidence, select the correct reasoning chain result and store it as the scene history record and the reasoning chain of the scene history record.

[0057] The present invention also provides an embodiment of a graphic question-and-answer and swimming detection system, which includes:

[0058] An acquisition module, used to acquire swimming scene images, scene history records, and reasoning chains of scene history records, and to acquire user input questions;

[0059] The associated behavior feature generation module is used to generate behavior hypotheses based on user input questions, and generate associated behavior features based on the behavior hypotheses;

[0060] A judgment module is used to judge whether the associated behavioral features are the same as the behavioral features of the historical scene records based on the swimming scene image, the historical scene record and the reasoning chain of the historical scene record; if they are the same, an inference chain result that is the same as the inference chain of the scene history record is output; if they are different, a new behavioral hypothesis is regenerated until a judgment can be made and the corresponding inference chain result is output; if it is found that there is a lack of historical scene records and it is impossible to make a judgment, a behavior feature question to be confirmed is generated for the behavior feature that cannot be judged; based on the behavior feature question to be confirmed, relevant behavior features in the swimming scene image are searched, and based on the search results and the behavior feature question to be confirmed, the behavior feature question to be confirmed is inferred to generate an inference chain to be confirmed; the degree of overlap between the inference chain to be confirmed and the inference chain of the scene history record is detected, and the inference chain of the scene history record with the largest overlap is taken to output the inference chain result;

[0061] The feedback module is used to feed back the reasoning chain and the reasoning chain results to the user to complete the detection.

[0062] In some embodiments, the acquisition module is a large language model acquisition module, and the associated behavior feature generation module is a large language model feature generation module; the judgment module is a large language model judgment module, a multimodal question-answering model judgment module and a target detection model judgment module. The large language model judgment module is used to judge whether the associated behavior features are the same as the behavior features of the historical scene records based on the swimming scene image, the historical scene record and the reasoning chain of the historical scene record. If they are the same, the reasoning chain result that is the same as the reasoning chain of the scene history record is output. If they are different, a new behavior hypothesis is regenerated until a judgment can be made and the corresponding reasoning chain result is output. If a missing result is found, the result of the corresponding reasoning chain is output. If there is a lack of historical scene records and judgment cannot be made, then a question about behavioral features to be confirmed is generated for the behavioral features that cannot be judged. The multimodal question-answering model judgment module is used to input the question about behavioral features to be confirmed, and according to the question about behavioral features to be confirmed, search for relevant behavioral features in the swimming scene image, and infer the question about behavioral features to be confirmed based on the search results and the question about behavioral features to be confirmed, and output the reasoning chain to be confirmed; the target detection model judgment module is used to input the reasoning chain to be confirmed, detect the degree of overlap between the reasoning chain to be confirmed and the reasoning chain of the scene historical records, take the reasoning chain of the scene historical records with the largest degree of overlap and output the reasoning chain result; the feedback module is a large language model feedback module.

[0063] It also includes a selection and recording module, which is used to manually determine the result of the reasoning chain by watching the real-time scene monitoring when confirming the behavior characteristics problem;

[0064] Confirm the correctness of the human judgment reasoning chain result and the reasoning chain result judged by coincidence, select the correct reasoning chain result and store it as the scene history record and the reasoning chain of the scene history record.

[0065] In some embodiments, the system can combine large-scale language models with multimodal question-answering models to jointly analyze the swimmer's behavior, support automatic question-answering, and generate reasoning chains to gradually answer the user's input questions. For example, when a user asks "Is this swimmer normal?" the system can input this multimodal data into the question-answering processing module based on historical behavior data, current posture, and whether any abnormal movements have occurred, and perform the following joint analysis and reasoning:

[0066] The system first calls the visual recognition module to analyze the current posture, movement, position and other information of the target swimmer in the monitoring image, and at the same time retrieves the swimmer's historical activity data. The large language model works in conjunction with the multimodal question-answering model:

[0067] The language model is responsible for understanding user questions and integrating the analysis process. The multimodal model extracts features from unstructured data such as video frames and action sequences and makes a preliminary judgment on whether the swimmer's movements are abnormal.

[0068] The system automatically generates a chain of reasoning, breaking down and answering the user's question step by step. The first step analyzes "Is the current posture normal for swimming?" The second step compares it to "Is there any difference compared to historical swimming data?" The third step checks for abnormal behavior. The results of each analysis step are fed back to the language model, and after integrating all the evidence, the system generates a final answer. The system provides the user with natural language feedback on the analysis process and results.

[0069] The target detection model is used to identify water images and generate target detection results. After input into the large language model, text descriptions and visually annotated images are obtained.

[0070] The system integrates the reasoning chains and reasoning chain results obtained in the previous steps, which can be used to optimize the system and enrich the system information.

[0071] The newly added video data will first be used to train and fine-tune the target detection model, enabling it to more accurately identify the posture, movements and specific performance in the water of swimmers of different body shapes and ages.

[0072] Retrain the multimodal question-answering model using newly collected question-answer pairs (such as questions asked by users for new scenarios and their standard answers).

[0073] The system inputs the reasoning chain in new scenarios, user feedback, expert-annotated analysis processes, etc. into the large language model to further optimize its reasoning and comprehensive analysis capabilities.

[0074] It supports the independent replacement or upgrade of any model module without major adjustments to other modules, enabling flexible maintenance and rapid response to new demands.

[0075] For example, after the target detection model is updated, the system can be seamlessly integrated into the existing process to ensure the stable operation of the overall function. How the large language model works:

[0076] The input to the large-scale language model includes system prompts, user conversation content, and the user input as the end of the conversation. The output is the new conversation content generated by the language model. The conversation is generated using an iterative "user input - model response" model.

[0077] The system has added a batch of video data shot in waters where swimming is prohibited (such as reservoirs and deep-water construction areas). After the above process, the update and coordination of each model are achieved as follows:

[0078] By learning from the newly added videos, the object detection model has improved its ability to recognize different water signs, warning signs, and environmental features. For example, it can accurately detect "No Swimming" signs and specific water boundaries.

[0079] The multimodal question-answering model uses newly collected question-answering data (such as "Is swimming allowed in this area?", "Is there any risk in swimming here?" and other questions and standard answers) to improve its understanding of scenario compliance and question-answering capabilities.

[0080] The large-scale language model optimizes the generation of reasoning chains and comprehensive judgment capabilities by learning how experts reason step by step through analytical processes such as "Is this water area a swimming-allowed area?" and "Does the current behavior violate regulations?"

[0081] Ultimately, the system can automatically identify swimming behavior in prohibited areas and, through multimodal analysis and reasoning chains, clearly conclude: "Swimming in prohibited areas is detected, which is abnormal behavior." This promptly triggers an alarm mechanism, notifying relevant management personnel to intervene. This process fully demonstrates the coordinated updates of various model modules and the improvement of the overall system capabilities.

[0082] Generating an inference chain involves the following steps:

[0083] Step 1: User input and scene collection.

[0084] Input content: The user enters questions through the system interface, such as: "Is there any abnormal swimming behavior in the current pool?"

[0085] Image acquisition: The system automatically acquires the current pool monitoring image or video frame as multimodal input.

[0086] Step 2: Preliminary analysis of large language models.

[0087] Input: The system inputs the user question and the current scene description (such as time, place, and historical dialogue) into a large language model.

[0088] Processing: Large language models try to directly infer and give answers based on the existing information.

[0089] Branching decision:

[0090] If there is sufficient historical information (e.g., the pool has been monitored continuously before and there are no abnormalities), the model can directly output: "No abnormal behavior is currently detected."

[0091] If the information is insufficient, proceed to the next step.

[0092] Step 3: Generate multimodal questions.

[0093] Input: After analyzing the large language model, it found that key visual evidence (such as the status of the swimmer) was missing, so it automatically generated further confirmation questions, such as "Is there any swimmer who is motionless for a long time?"

[0094] Output: Questions that need further confirmation.

[0095] Step 4: Multimodal question-answering model analysis.

[0096] Input: The system inputs the current image / video frame and the above questions that need further confirmation into the multimodal question answering model.

[0097] Processing: Multimodal question answering models combine visual content and text questions to provide answers, such as: "A swimmer has been detected to be stationary for more than 2 minutes."

[0098] Output: The answer is returned to the large language model to enrich the reasoning chain information.

[0099] Step 5: Target detection model auxiliary verification.

[0100] Trigger condition: If the answer of the multimodal question-answering model is still uncertain (such as "still" but the specific action cannot be determined), the large language model will further call the object detection model.

[0101] Input: current image / video frame.

[0102] Processing: The object detection model identifies the position, posture, and movements (such as floating, struggling, sinking, etc.) of all swimmers.

[0103] Output: Detection results (such as “Swimmer A is floating face-down with no autonomous movements”) are fed back to the large language model.

[0104] Step 6: Comprehensive reasoning and reasoning chain generation.

[0105] Input: Large language models integrate user input, historical information, multimodal question-answering results, and object detection results.

[0106] Reasoning chain generation: User asks “Is there any unusual behavior?”

[0107] Check the historical information and find no abnormalities.

[0108] Generate the question: "Has anyone been still for a long time?".

[0109] The multimodal question-answering model answers: "One person stood still for 2 minutes."

[0110] The target detection model was further called and it was found that the person was floating face down and not moving.

[0111] Comprehensive judgment: "Suspected drowning, abnormal behavior."

[0112] Feedback module output and response: The system displays the complete reasoning chain and final conclusion to the user, such as, "Analysis detected a swimmer who was stationary and floating face-down for an extended period of time. Drowning is suspected, and an alarm has been automatically triggered." When the multimodal question-answering model is used, the large language model generates questions related to the swimming scenario, such as, "Is the swimmer abnormal?" This question and image are input into the multimodal question-answering model to obtain a system-generated answer. When the multimodal question-answering model is called upon and object detection results are required, the object detection model is used for precise positioning.

[0113] When the target detection model is used: the large language model generates key objects related to the swimming scene (such as "swimmer" and "water surface status") and inputs them into the target detection model to obtain detection results, such as the target's location information, whether the head is above the water, and whether there is any struggling movement.

[0114] The results of the target detection model are fed into a large language model to obtain a text description of the historical scene record, which may include:

[0115] 1. Number of swimmers detected.

[0116] 2. Screen size.

[0117] 3. Label, category, and confidence score of each instance.

[0118] 4. Detection frame position and abnormal behavior identification (such as drowning alarm).

[0119] How the object detection model works:

[0120] 1. Preprocess the input swimming scene image and detection objects (such as swimmers) to obtain standardized image data suitable for the input model, and perform inference through a neural network based on the Transformer architecture.

[0121] 2. Set the confidence threshold to 0.5 and the non-maximum suppression (NMS) parameter to remove detection frames with excessive overlap to ensure accurate and reliable detection results.

[0122] 3. Generate structured detection results, including the number of people, behavior, behavioral characteristics, abnormality probability, etc.

[0123] For example: Person B: detection box [400, 500, 470, 600], behavior is struggling, behavior feature is struggling posture, water area is "no swimming", abnormal probability = 0.65 (abnormal alarm).

[0124] How the multimodal question answering model works:

[0125] 1. Preprocess the input swimming scene images and related issues.

[0126] 2. Input the pre-processed image and formatted related questions into the multimodal question answering model, and use a neural network based on the Transformer architecture for reasoning.

[0127] 3. Generate an answer, such as: "The swimmer is swimming normally." or "The swimmer has not moved for a long time, which may be abnormal."

[0128] The large language model, target detection model, and multimodal question-answering model of the present invention can be replaced by other models with visual-language understanding capabilities, such as CLIP, BLIP, and MiniGPT-4, to adapt to different computing resources and application requirements.

[0129] Throughout this specification, references to terms such as "one embodiment," "example," or "specific example" indicate that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0130] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention, and such changes and modifications fall within the scope of the invention as claimed.

Claims

1. A method for image-text question-answering and person swimming detection, characterized in that: The following steps are involved: Collect swimming scene images, scene history records, and reasoning chains of scene history records, and collect user input questions; Generate behavioral hypotheses based on user input questions, and generate associated behavioral features based on the behavioral hypotheses; According to the swimming scene image, historical scene record and the reasoning chain of the historical scene record, determine whether the associated behavioral features are the same as the behavioral features of the historical scene record. If they are the same, output the reasoning chain result that is the same as the reasoning chain of the scene history record. If they are different, regenerate a new behavioral hypothesis until a judgment can be made and output the corresponding reasoning chain result. If it is found that there is a lack of historical scene records and it is impossible to make a judgment, generate a behavior feature problem to be confirmed for the behavior feature that cannot be judged; according to the behavior feature problem to be confirmed, search for relevant behavior features in the swimming scene image, and based on the search results and the behavior feature problem to be confirmed, infer the behavior feature problem to be confirmed and generate the reasoning chain to be confirmed; detect the degree of overlap between the reasoning chain to be confirmed and the reasoning chain of the scene history record, take the reasoning chain of the scene history record with the largest overlap and output the reasoning chain result; Feedback the inference chain and the inference chain results to the user to complete the detection.

2. A method for image-text question-answering and swimming detection according to claim 1, characterized in that: The method for collecting swimming scene images, scene history records, and reasoning chains of scene history records, and collecting user input questions comprises the following steps: The swimming scene images are collected through the camera; the scene history records and the reasoning chain of the scene history records are called by calling the model; and the user input questions are collected through the input model.

3. A method for image-text question-answering and swimming detection according to claim 2, characterized in that: The calling model and the input model both adopt a large language model.

4. A method for image-text question-answering and swimming detection according to claim 2, characterized in that: The scene history record includes a historical image record and a historical conversation record, wherein the historical conversation record expresses the historical image through a text record; and the reasoning chain of the scene history record is a reasoning process related to the historical conversation through the text record.

5. The method for image-text question-answering and swimming detection according to claim 1, characterized in that: Based on the user input question, a behavior hypothesis is generated. The method of generating associated behavior features based on the behavior hypothesis includes: User questions are input into the large language model, and behavioral hypotheses are generated through the large language model, and associated behavioral features are generated based on the behavioral hypotheses.

6. The method for image-text question-answering and swimming detection according to claim 1, characterized in that: The method for determining whether the associated behavior characteristics are the same as the behavior characteristics of the historical scene records based on the swimming scene image, the historical scene record, and the inference chain of the historical scene records includes: The swimming scene image, historical scene record and the reasoning chain of the historical scene record are input into the large language model to generate the behavioral features of the historical scene record. The large language model is used to determine whether the associated behavioral features are the same as the behavioral features of the historical scene record.

7. The method for image-text question-answering and swimming detection according to claim 1, characterized in that: The method further comprises: For the problem of confirming behavioral characteristics, the inference chain results are judged manually by watching the real-time scene monitoring; Confirm the correctness of the human judgment reasoning chain result and the reasoning chain result judged by coincidence, select the correct reasoning chain result and store it as the scene history record and the reasoning chain of the scene history record.

8. A graphic question-and-answer and swimming detection system, characterized in that: include: An acquisition module, used to acquire swimming scene images, scene history records, and reasoning chains of scene history records, and to acquire user input questions; The associated behavior feature generation module is used to generate behavior hypotheses based on user input questions, and generate associated behavior features based on the behavior hypotheses; A judgment module is used to judge whether the associated behavioral features are the same as the behavioral features of the historical scene records based on the swimming scene image, the historical scene record and the reasoning chain of the historical scene record; if they are the same, an inference chain result that is the same as the inference chain of the scene history record is output; if they are different, a new behavioral hypothesis is regenerated until a judgment can be made and the corresponding inference chain result is output; if it is found that there is a lack of historical scene records and it is impossible to make a judgment, a behavior feature question to be confirmed is generated for the behavior feature that cannot be judged; based on the behavior feature question to be confirmed, relevant behavior features in the swimming scene image are searched, and based on the search results and the behavior feature question to be confirmed, the behavior feature question to be confirmed is inferred to generate an inference chain to be confirmed; the degree of overlap between the inference chain to be confirmed and the inference chain of the scene history record is detected, and the inference chain of the scene history record with the largest overlap is taken to output the inference chain result; The feedback module is used to feed back the reasoning chain and the reasoning chain results to the user to complete the detection.

9. The graphic question-and-answer and swimming detection system according to claim 8, characterized in that: The acquisition module is a large language model acquisition module, and the associated behavior feature generation module is a large language model feature generation module; the judgment module is a large language model judgment module, a multimodal question-answering model judgment module and a target detection model judgment module. The large language model judgment module is used to judge whether the associated behavior features are the same as the behavior features of the historical scene records based on the swimming scene image, the historical scene record and the reasoning chain of the historical scene record. If they are the same, the reasoning chain result that is the same as the reasoning chain of the scene history record is output. If they are different, a new behavior hypothesis is regenerated until a judgment can be made and the corresponding reasoning chain result is output. If it is found that there is a lack of historical scene records, the result of the reasoning chain is generated. If the scene record cannot be judged, a behavior feature question to be confirmed is generated for the behavior feature that cannot be judged. The multimodal question-answering model judgment module is used to input the behavior feature question to be confirmed, and according to the behavior feature question to be confirmed, search for relevant behavior features in the swimming scene image, and infer the behavior feature question to be confirmed based on the search result and the behavior feature question to be confirmed, and output the reasoning chain to be confirmed; the target detection model judgment module is used to input the reasoning chain to be confirmed, detect the overlap between the reasoning chain to be confirmed and the reasoning chain of the scene historical record, take the reasoning chain of the scene historical record with the largest overlap and output the reasoning chain result; the feedback module is a large language model feedback module.

10. The graphic question-and-answer and swimming detection system according to claim 8, characterized in that: It also includes a selection and recording module, which is used to manually determine the result of the reasoning chain by watching the real-time scene monitoring when confirming the behavior characteristics problem; Confirm the correctness of the human judgment reasoning chain result and the reasoning chain result judged by coincidence, select the correct reasoning chain result and store it as the scene history record and the reasoning chain of the scene history record.

Citation Information

Patent Citations

  • Visual dialogue answer generation method and device based on graph perception

    CN115129839A

  • Intelligent question and answer method and device and electronic equipment

    CN117828017A