Image-text question and answer and person swimming detection method and system
Through the graphic and text Q&A method combined with image and text information for deep semantic reasoning, the limitations of image processing systems in the prior art in detecting complex image abnormal events are solved, automatic monitoring and intelligent Q&A are realized, the accuracy of abnormal detection and system adaptability are improved, and the cost of manual participation is reduced.
Patent Information
- Application Number
- CN202510805757.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-17
AI Technical Summary
The existing image processing system lacks comprehensive image understanding and reasoning capabilities when detecting abnormal events in complex images, especially when people swimming in illegal areas, and is difficult to cope with the diversity and suddenness of abnormal behaviors. It relies on a large amount of labeling data and high-cost manual labeling, and cannot provide intelligent question-and-answer and rich contextual explanations.
The graphic and text Q&A method is adopted to collect swimming scene images, history records and user input problems, and use large language models to generate behavioral assumptions and features, combine image and text information for deep semantic reasoning, generate inference chains and feedback results, realize automatic monitoring and intelligent Q&A, and support modular design to enhance system adaptability and maintenance.
It improves the accuracy and flexibility of abnormal detection, reduces the cost of manual participation, can understand the image content more comprehensively, identify abnormal situations that are difficult to detect in the visual model, and interprets detection details through intelligent question and answers, and filters valuable image sample optimization models.
Smart Images

Figure CN120356139A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to a method and system for image-text question answering and personnel swimming detection. Background Art
[0002] Existing image processing systems usually focus on single tasks, such as only implementing object detection, object recognition, or anomaly detection, lacking comprehensive image understanding and reasoning capabilities, and especially having limitations in detecting abnormal events in complex images, such as personnel swimming in illegal areas.
[0003] How to automatically detect abnormal events in images is one of the most valuable problems in the field of image analysis. Traditional abnormal event detection methods require predefining the types of abnormal events to be detected and collecting a large number of manually labeled samples to establish an abnormal event detection model. However, due to the broad and vague definition of abnormal behaviors, it is difficult to list all possible types of abnormal behaviors in advance. In addition, most behaviors in daily life are normal behaviors, so it is difficult to collect enough abnormal behavior samples for modeling and analysis.
[0004] Traditional image anomaly detection methods often rely on specific image features or labeled datasets. For example, methods based on convolutional neural networks (CNNs) judge whether an image contains an abnormal personnel swimming event through feature extraction. However, these methods rely on a large number of labeled normal and abnormal samples. In practical applications, abnormal personnel swimming events are usually rare and unpredictable. Existing methods usually model normal behavior patterns by analyzing a large number of normal behavior samples. When the behavior pattern in an image cannot be explained by the established normal behavior model, it is considered that there is an abnormal behavior in the image.
[0005] In addition, many current anomaly detection systems rely on learning models from manually labeled object detection datasets by experts to identify known anomaly types. Although this method can effectively detect certain specific abnormal behaviors, it is difficult to cover all possible abnormal situations, especially in scenarios of cross-scene and multi-modal information fusion. In addition, the cost of manual annotation is high, and it is difficult to cope with the diversity and suddenness of abnormal behaviors.
[0006] When applied to the online detection task of images, these anomaly detection methods usually require a large amount of labeled data to distinguish normal and abnormal situations. And the form of the image anomaly detection task is relatively single, usually only able to indicate whether there is an anomaly in the image or mark the location of the anomaly, and unable to further conduct intelligent question answering on the specific details in the image, providing richer context and explanations. Summary of the Invention
[0007] The object of the present invention is to provide a method for graphic and text Q&A and personnel swimming detection, to overcome the limitations in the prior art that are difficult to cope with complex behavior analysis, to achieve automatic monitoring, intelligent Q&A and anomaly detection in the swimming scenario, and to significantly improve the accuracy and flexibility of the system.
[0008] The object of the present invention can be achieved by the following technical solutions: A method for graphic and text Q&A and personnel swimming detection, comprising the following steps: Collect swimming scenario images, scenario historical records and the inference chain of the scenario historical records, and collect user input questions; Generate behavior hypotheses according to the user input questions, and generate associated behavior characteristics according to the behavior hypotheses; According to the swimming scenario images, historical scenario records and the inference chain of the historical scenario records, judge whether the associated behavior characteristics are the same as the behavior characteristics of the historical scenario records. If they are the same, output the inference chain result that is the same as the inference chain of the scenario historical records. If they are different, regenerate new behavior hypotheses until a judgment can be made and then output the corresponding inference chain result. If it is found that there is a lack of historical scenario records and no judgment can be made, generate a problem of to-be-confirmed behavior characteristics for the behavior characteristics that cannot be judged; according to the problem of to-be-confirmed behavior characteristics, search for relevant behavior characteristics in the swimming scenario images, and reason about the problem of to-be-confirmed behavior characteristics based on the search results and the problem of to-be-confirmed behavior characteristics to generate an inference chain to be confirmed; detect the coincidence degree between the inference chain to be confirmed and the inference chain of the scenario historical records, and output the inference chain result of the scenario historical record with the largest coincidence degree; Feed back the inference chain and the inference chain result to the user to complete the detection.
[0009] In a further solution, the method for collecting swimming scenario images, scenario historical records and the inference chain of the scenario historical records, and collecting user input questions includes the following steps: Collect swimming scenario images through a camera; call scenario historical records and the inference chain of the scenario historical records through a model call; collect user input questions through an input model.
[0010] In a further solution, both the model call and the input model adopt large language models.
[0011] In a further solution, the scenario historical records include historical image records and historical conversation records, the historical conversation records describe the historical images through text records; the inference chain of the scenario historical records records the inference process related to the historical conversation through text records.
[0012] In a further solution, the method for generating behavior hypotheses according to the user input questions and generating associated behavior characteristics according to the behavior hypotheses includes: Input the user's question into the large language model, generate behavioral hypotheses through the large language model, and generate associated behavioral characteristics based on the behavioral hypotheses.
[0013] In a further solution, the method for determining whether the associated behavioral characteristics are the same as the behavioral characteristics of the historical scene record according to the swimming scene image, the historical scene record, and the reasoning chain of the historical scene record includes: Input the swimming scene image, the historical scene record, and the reasoning chain of the historical scene record into the large language model to generate the behavioral characteristics of the historical scene record, and determine whether the associated behavioral characteristics are the same as the behavioral characteristics of the historical scene record through the large language model.
[0014] In a further solution, the method further includes: For the problem of behavioral characteristics to be confirmed, manually judge the reasoning chain result by watching the real-time scene monitoring; Confirm the correctness of the manually judged reasoning chain result and the reasoning chain result judged by the coincidence degree, and select the correct reasoning chain result to be stored as the scene historical record and the reasoning chain of the scene historical record.
[0015] The present invention also proposes a graphic and text Q&A and personnel swimming detection system, including: An acquisition module, configured to acquire swimming scene images, scene historical records, and reasoning chains of scene historical records, and acquire user input questions; An associated behavioral characteristic generation module, configured to generate behavioral hypotheses according to the user input question, and generate associated behavioral characteristics according to the behavioral hypotheses; A judgment module, configured to determine whether the associated behavioral characteristics are the same as the behavioral characteristics of the historical scene record according to the swimming scene image, the historical scene record, and the reasoning chain of the historical scene record. If they are the same, output a reasoning chain result that is the same as the reasoning chain of the scene historical record. If they are different, regenerate new behavioral hypotheses until a judgment can be made and then output the corresponding reasoning chain result. If it is found that there is a lack of historical scene records and no judgment can be made, generate a problem of behavioral characteristics to be confirmed for the unjudgeable behavioral characteristics; according to the problem of behavioral characteristics to be confirmed, search for relevant behavioral characteristics in the swimming scene image, and reason about the problem of behavioral characteristics to be confirmed according to the search result and the problem of behavioral characteristics to be confirmed, and generate an unconfirmed reasoning chain; detect the coincidence degree between the unconfirmed reasoning chain and the reasoning chain of the scene historical record, and output the reasoning chain result of the scene historical record with the largest coincidence degree; A feedback module, configured to feedback the reasoning chain and the reasoning chain result to the user to complete the detection.
[0016] In a further solution, the acquisition module is a large language model acquisition module, and the associated behavior feature generation module is a large language model feature generation module; the judgment module includes a large language model judgment module, a multimodal question and answer model judgment module, and an object detection model judgment module. The large language model judgment module is used to determine whether the associated behavior features are the same as the behavior features of the historical scene record based on the swimming scene image, the historical scene record, and the reasoning chain of the historical scene record. If they are the same, it outputs the reasoning chain result that is the same as the reasoning chain of the scene historical record. If they are different, it regenerates a new behavior hypothesis until a judgment can be made and then outputs the corresponding reasoning chain result. If it is found that there is a lack of historical scene record and no judgment can be made, it generates a problem of behavior features to be confirmed for the behavior features that cannot be judged. The multimodal question and answer model judgment module is used to input the problem of behavior features to be confirmed, search for relevant behavior features in the swimming scene image according to the problem of behavior features to be confirmed, reason about the problem of behavior features to be confirmed based on the search result and the problem of behavior features to be confirmed, and output the reasoning chain to be confirmed. The object detection model judgment module is used to input the reasoning chain to be confirmed, detect the coincidence degree between the reasoning chain to be confirmed and the reasoning chain of the scene historical record, and output the reasoning chain result by taking the reasoning chain of the scene historical record with the largest coincidence degree. The feedback module is a large language model feedback module.
[0017] In a further solution, the above system further includes a selection and recording module, which is used to, for the problem of behavior features to be confirmed, manually judge the reasoning chain result by watching the real-time scene monitoring. Confirm the correctness of the manually judged reasoning chain result and the reasoning chain result judged by coincidence degree, and select the correct reasoning chain result to be stored as the scene historical record and the reasoning chain of the scene historical record.
[0018] Advantages of the present invention: The present invention collects the questions input by the user and infers the scene images based on the questions. Since the questions for input can include real-time questions related to swimming training itself or preset questions about anomaly detection, it can simultaneously perform person swimming graphic and text Q&A and image anomaly detection, and can achieve more accurate anomaly detection and intelligent Q&A in the process of analyzing the input person swimming images. By combining image and text information for analysis, the detection of the present invention not only depends on the visual features in the image, but also can perform in-depth semantic reasoning by combining text descriptions, thereby effectively improving the accuracy of anomaly event detection. Compared with traditional single-visual detection, the present invention can more comprehensively understand the content of person swimming images and identify abnormal situations that are difficult to discover only relying on visual models; through the intelligent Q&A function, the present invention can automatically answer specific questions about the content of person swimming images, thus reducing the need for manual image analysis. This Q&A mechanism can explain the detected abnormal events, help the operator quickly understand the abnormal details in the image, and greatly reduce the burden of manual participation; combination of anomaly detection and inference chain: through inference chain judgment, the system can gradually answer complex questions in the image and discover potential abnormal events during the inference process. This method can screen out image samples valuable for model update, especially those abnormal scenarios that are difficult for the current model to distinguish, and these abnormal scenario samples can be used to further optimize the historical scenario records; The system of the present invention adopts a modular design, independently separates functional modules such as graphic and text Q&A, multi-modal information processing, target detection, and abnormal event detection. Each functional module can be independently upgraded or replaced according to different application scenarios, thereby enhancing the adaptability of the system in different environments. In addition, this modular design is convenient for integration with other existing systems, greatly improving the maintainability and scalability of the system, and being able to better meet complex and diverse image analysis requirements. In summary, the present invention provides a method for person swimming graphic and text Q&A and image anomaly detection, which can effectively improve the efficiency of anomaly detection and significantly reduce the cost of manual participation. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 It is a schematic structural diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0022] A method for graphic and text Q&A and personnel swimming detection includes the following steps: S1. Collect swimming scene images, scene historical records, and the inference chain of the scene historical records, and collect the questions input by the user; S2. Generate behavior hypotheses according to the questions input by the user, and generate associated behavior features according to the behavior hypotheses; S3. According to the swimming scene images, historical scene records, and the inference chain of the historical scene records, judge whether the associated behavior features are the same as the behavior features of the historical scene records. If they are the same, output the inference chain result that is the same as the inference chain of the scene historical record. If they are different, regenerate new behavior hypotheses until a judgment can be made and then output the corresponding inference chain result. If it is found that there is a lack of historical scene records and no judgment can be made, generate a problem of behavior features to be confirmed for the behavior features that cannot be judged; according to the problem of behavior features to be confirmed, search for relevant behavior features in the swimming scene images, and reason about the problem of behavior features to be confirmed based on the search results and the problem of behavior features to be confirmed, and generate an inference chain to be confirmed; detect the coincidence degree between the inference chain to be confirmed and the inference chain of the scene historical record, and output the inference chain result of the scene historical record with the largest coincidence degree; S4. Feed back the inference chain and the inference chain result to the user to complete the detection.
[0023] In some embodiments, step S1 includes the following steps: Collect swimming scene images through a camera; call the scene historical records and the inference chain of the scene historical records through a calling model; collect the questions input by the user through an input model.
[0024] Both the calling model and the input model adopt large language models.
[0025] The scene historical records include historical image records and historical conversation records. The historical conversation records describe the historical images through text records; the inference chain of the scene historical records records the inference process related to the historical conversation through text.
[0026] For the input swimming scene images, a series of image preprocessing steps can also be performed first, including image size normalization, noise removal, and adaptive adjustment of brightness and contrast. The result of the preprocessing is high-quality and standardized image data, which is convenient for subsequent model analysis.
[0027] Meanwhile, convert the historical conversations and reasoning chains related to the swimming scenario into a structured text format. For example, the historical conversation can be recorded as "On May 1, 2024, swimmer A was detected to be stationary in the shallow water area for more than 1 minute, the system issued a warning, and the lifeguard confirmed no abnormalities"; the reasoning chain can be sorted out as "Detect long-term stillness → Query historical records → Find similar cases → Suggest manual review". These text data are uniformly organized in JSON format for easy reading and processing. This anomaly detection can be judged in real time through questions pre-entered by the user, which can improve the efficiency of anomaly detection and discovery. Of course, the user can also directly input relevant anomaly detection questions. For example, after collecting the user's input questions from the large language model, the large language model generates hypotheses about abnormal behaviors, such as stationary behavior, thrashing behavior, disappearance behavior, etc. Taking the stationary behavior as an example, generate the behavioral characteristics of stationary behavior, such as the shape of the detection box of the person does not change or the position of the detection box of the head does not change much, the stationary time, the body posture (face up or face down), whether there is an autonomous breathing action, and whether the limbs struggle or move.
[0028] The above standardized image data and structured text description are input into the judgment module together, providing rich context information for subsequent behavior recognition, anomaly reasoning, and intelligent question answering. The judgment module determines whether there are the same behavior characteristics in the historical scenario. If so, for example, if the behavior characteristics are the same as those of the behavior of the head facing underwater for more than a certain period of time in the historical scenario, the reasoning chain result of this motionless behavior is output as a drowning behavior, and the relevant reasoning chain is output. However, for behavior characteristics such as the head facing up, the result of the reasoning chain may be uncertain. For example, the result of the reasoning chain includes motionless during normal rest or motionless in the case of drowning. Then, new behavior hypotheses are generated, such as generating whether there is a behavior of calling for help or whether there is a behavior of other people approaching to help, generating relevant behavior characteristics, and continuing to judge. If it is the same as the behavior characteristics recorded in the historical scenario of a swimming person taking a rest with someone beside, the reasoning chain result of normal rest is output. If there is no one beside, and the same historical scenario record cannot be matched, it cannot be determined whether it is normal rest or prohibited from moving and waiting for rescue. Then, a problem of behavior characteristics to be confirmed is generated, such as whether the behavior characteristics of the motionless swimming person with the head facing up are normal rest or waiting for rescue. The relevant behavior characteristics in the swimming scenario image are found, and a reasoning chain to be confirmed is generated. For example, the generated reasoning chain to be confirmed is: detecting long-term immobility → querying historical records → finding similar cases of waiting for rescue → suggesting manual review; or detecting long-term immobility → querying historical records → finding similar cases of waiting for rescue → suggesting manual review. Then, the coincidence degree between the reasoning chain to be confirmed and the recorded reasoning chain is judged. The detection method of the coincidence degree can be to judge the coincidence degree of relevant behavior characteristics. If the number of the same relevant behavior characteristics is large, it is judged that the coincidence degree is high. Among them, the behavior characteristics can be judged whether they coincide through an object detection model. For example, a confidence threshold such as 0.5 and non-maximum suppression (NMS) parameters are set to remove detection frames with too high overlap, generating a coincidence anomaly probability, etc. The one with the lowest coincidence anomaly probability is selected as having a high coincidence degree, and the one with a high coincidence anomaly probability is selected as not coincident to ensure the accuracy and reliability of the detection results.
[0029] In some embodiments, the method for generating behavior hypotheses according to user input questions and generating associated behavior characteristics according to the behavior hypotheses includes: Input the user question into the large language model, and generate behavior hypotheses through the large language model, and generate associated behavior characteristics based on the behavior hypotheses.
[0030] In some embodiments, the method for determining whether the associated behavior characteristics are the same as the behavior characteristics of the historical scenario record according to the swimming scenario image, the historical scenario record, and the reasoning chain of the historical scenario record includes: Input the swimming scenario image, the historical scenario record, and the reasoning chain of the historical scenario record into the large language model to generate the behavior characteristics of the historical scenario record, and judge whether the associated behavior characteristics are the same as the behavior characteristics of the historical scenario record through the large language model.
[0031] In some embodiments, it further includes step S5: for the problem of behavior characteristics to be confirmed, the result of the inference chain is judged manually by watching the real-time scene monitoring; Confirm the correctness of the result of the manual judgment inference chain and the result of the inference chain judged by the coincidence degree, and select the correct result of the inference chain to be stored as the scene history record and the inference chain of the scene history record.
[0032] The present invention also proposes an embodiment of a graphic and text Q&A and personnel swimming detection system, and the system includes: An acquisition module, configured to acquire swimming scene images, scene history records and inference chains of scene history records, and acquire user input questions; An associated behavior feature generation module, configured to generate behavior hypotheses according to user input questions, and generate associated behavior features according to the behavior hypotheses; A judgment module, configured to judge whether the associated behavior features are the same as the behavior features of the historical scene record according to the swimming scene image, the historical scene record and the inference chain of the historical scene record. If they are the same, output the inference chain result that is the same as the inference chain of the scene history record. If they are different, regenerate new behavior hypotheses until a judgment can be made and then output the corresponding inference chain result. If it is found that there is a lack of historical scene records and no judgment can be made, generate a problem of behavior characteristics to be confirmed for the behavior characteristics that cannot be judged; according to the problem of behavior characteristics to be confirmed, search for relevant behavior characteristics in the swimming scene image, and reason about the problem of behavior characteristics to be confirmed according to the search result and the problem of behavior characteristics to be confirmed, and generate an inference chain to be confirmed; detect the coincidence degree between the inference chain to be confirmed and the inference chain of the scene history record, and output the inference chain result of the scene history record with the largest coincidence degree; A feedback module, configured to feedback the inference chain and the inference chain result to the user to complete the detection.
[0033] In some embodiments, the acquisition module is a large language model acquisition module, and the associated behavior feature generation module is a large language model feature generation module; the judgment module includes a large language model judgment module, a multimodal question and answer model judgment module, and an object detection model judgment module. The large language model judgment module is used to determine whether the associated behavior features are the same as the behavior features of the historical scene record according to the swimming scene image, the historical scene record, and the reasoning chain of the historical scene record. If they are the same, it outputs the reasoning chain result that is the same as the reasoning chain of the scene historical record. If they are different, it regenerates a new behavior hypothesis until a judgment can be made and then outputs the corresponding reasoning chain result. If it is found that there is a lack of historical scene record and no judgment can be made, it generates a problem of behavior features to be confirmed for the behavior features that cannot be judged. The multimodal question and answer model judgment module is used to input the problem of behavior features to be confirmed, search for relevant behavior features in the swimming scene image according to the problem of behavior features to be confirmed, reason about the problem of behavior features to be confirmed based on the search result and the problem of behavior features to be confirmed, and output the reasoning chain to be confirmed; the object detection model judgment module is used to input the reasoning chain to be confirmed, detect the coincidence degree between the reasoning chain to be confirmed and the reasoning chain of the scene historical record, and output the reasoning chain result of the scene historical record with the largest coincidence degree; the feedback module is a large language model feedback module.
[0034] It further includes a selection and recording module, which is used to, for the problem of behavior features to be confirmed, manually judge the reasoning chain result by watching the real-time scene monitoring. Confirm the correctness of the manually judged reasoning chain result and the reasoning chain result judged by the coincidence degree, and select the correct reasoning chain result to be stored as the scene historical record and the reasoning chain of the scene historical record.
[0035] In some embodiments, in this system, a large language model and a multimodal question and answer model can be combined to jointly analyze the behavior of swimmers, support automatic question answering, generate reasoning chains, and gradually answer the questions input by users. For example, when the user asks "Is this swimmer normal?", the system can input these multimodal data such as historical behavior data, current posture, and whether there are abnormal actions into the question answering processing module according to the historical behavior data, current posture, and whether there are abnormal actions, and perform the following joint analysis and reasoning: The system first calls the visual recognition module to analyze information such as the current posture, actions, and position of the target swimmer in the monitoring screen, and at the same time retrieves the historical activity data of the swimmer. The large language model and the multimodal question and answer model work together: The language model is responsible for understanding the user's question and integrating the analysis process. The multimodal model extracts features from unstructured data such as video frames and action sequences and preliminarily judges whether there are abnormal behaviors in the swimmer's actions.
[0036] The system automatically generates an inference chain, gradually decomposes and answers the user's questions. The first step: Analyze "Is the current posture a normal swimming posture?" The second step: Compare "Is there any difference compared with the historical swimming data?" The third step: Check "Are there any abnormal behaviors?" The analysis results of each step will be fed back to the language model. After synthesizing all the evidence, the system generates a final answer. The system feeds back the analysis process and results to the user in the form of natural language.
[0037] The target detection model is used to identify the water area pictures, generate target detection results, and after inputting them into the large language model, obtain text descriptions and visualized annotation pictures.
[0038] The system integrates the inference chain and the inference chain results obtained in the foregoing steps, which can be used to optimize the system and enrich the information of the system.
[0039] The newly added video data is first used to train and fine-tune the target detection model, enabling it to more accurately identify the postures, movements, and specific performances in the water of swimmers of different body types and age groups.
[0040] Use the newly collected question-and-answer pairs (such as the questions and standard answers raised by users for new scenarios) to retrain the multi-modal question-and-answer model.
[0041] The system inputs the inference chain, user feedback, analysis processes annotated by experts, etc. in the new scenario into the large language model to further optimize its inference and comprehensive analysis capabilities.
[0042] Supports replacing or upgrading any model module individually, and other modules do not need to be significantly adjusted, realizing flexible maintenance and quick response to new requirements.
[0043] For example, after the target detection model is updated, the system can be seamlessly integrated into the existing process to ensure the stable operation of the overall function. The operation mode of the large language model: The inputs of the large language model include system prompts, user conversation content, and end with the user input. The output is the new conversation content generated by the language model. The generation of the conversation adopts an iterative mode of "user input - model response".
[0044] The system has newly added a batch of video data taken in waters where swimming is prohibited (such as reservoirs, deep water construction areas, etc.). Through the above processes, the update and coordination of each model are as follows: By learning the newly added videos, the target detection model has improved its ability to identify different water area signs, warning signs, and environmental features. For example, it can accurately detect "No Swimming" signs and specific water area boundaries.
[0045] The multi-modal question answering model utilizes newly collected question and answer data (such as questions like "Is swimming allowed in this area?" and "Is there any risk of swimming here?" and their corresponding standard answers), enhancing the understanding of scene compliance and question answering capabilities.
[0046] The large language model optimizes the generation of reasoning chains and comprehensive judgment capabilities by learning how experts reason step by step in analyzing processes such as "Whether this water area belongs to the permitted swimming area" and "Whether the current behavior violates the regulations".
[0047] Finally, the system can automatically identify swimming behaviors occurring in the prohibited swimming waters, and through multi-modal analysis and reasoning chains, clearly give the conclusion: "It is detected that someone is swimming in the prohibited swimming waters, which is an abnormal behavior", thus triggering the alarm mechanism in a timely manner and notifying relevant management personnel for intervention. This process fully demonstrates the collaborative update of each model module and the improvement of the overall system capabilities.
[0048] The generation of the reasoning chain includes the following steps: Step 1: User input and scene collection.
[0049] Input content: The user inputs a question through the system interface, such as: "Is there any abnormal swimming behavior in the current swimming pool?".
[0050] Image collection: The system automatically collects the current swimming pool surveillance images or video frames as multi-modal input.
[0051] Step 2: Preliminary analysis by the large language model.
[0052] Input: The system inputs the user's question, the current scene description (such as time, location, historical conversation) into the large language model.
[0053] Processing: The large language model attempts to directly reason and give an answer based on the existing information.
[0054] Branch decision: If the historical information is sufficient (such as the swimming pool has been continuously monitored before and there is no abnormality), the model can directly output: "No abnormal behavior is detected currently".
[0055] If the information is insufficient, proceed to the next step.
[0056] Step 3: Generate multi-modal questions.
[0057] Input: After analysis by the large language model, it is found that key visual evidence (such as the state of the swimmer) is lacking, so it automatically generates further questions that need to be confirmed, such as "Is there a swimmer who has been stationary for a long time?". Output: Further questions that need to be confirmed.
[0058] Step 4: Analysis by the multi-modal question answering model.
[0059] Input: The system inputs the current image / video frame and the above-mentioned questions that further need to be confirmed into the multi-modal question-answering model.
[0060] Processing: The multi-modal question-answering model combines the visual content and the text question to give an answer, such as: "It is detected that a swimmer has been stationary for more than 2 minutes."
[0061] Output: This answer is returned to the large language model to enrich the inference chain information.
[0062] Step Five: Assisted verification by the object detection model.
[0063] Trigger condition: If the answer of the multi-modal question-answering model still has uncertainty (such as "stationary" but the specific action cannot be judged), the large language model will further call the object detection model.
[0064] Input: The current image / video frame.
[0065] Processing: The object detection model identifies the positions, postures, and actions of all swimmers (such as floating, struggling, sinking, etc.).
[0066] Output: The detection result (such as "Swimmer A is in a prone floating position and has no autonomous movement") is fed back to the large language model.
[0067] Step Six: Comprehensive reasoning and inference chain generation.
[0068] Input: The large language model integrates the user input, historical information, multi-modal question-answering results, and object detection results.
[0069] Inference chain generation: The user asks "Is there any abnormal behavior?" Check the historical information, no abnormality.
[0070] Generate a question: "Is there anyone stationary for a long time?"
[0071] The multi-modal question-answering model answers: "There is one person who has been stationary for 2 minutes."
[0072] Further call the object detection model and find that the person is in a prone floating position and has no movement.
[0073] Comprehensive judgment: "Suspicious of drowning, which belongs to abnormal behavior."
[0074] Feedback Module Output and Response: The output content is that the system displays the complete reasoning chain and the final conclusion to the user, such as: "After analysis, a swimmer was detected to be stationary for a long time and in a face-down floating state, suspected of drowning, and an automatic alarm has been triggered." When adopting the multi-modal Q&A model behavior: The large language model will generate questions related to the swimming scenario, such as "Is there anything abnormal with this swimmer?" and input this question and the image into the multi-modal Q&A model to obtain the answer generated by the system. When the target detection result is required when calling the multi-modal Q&A model, the target detection model will be called for accurate positioning.
[0075] When adopting the target detection model behavior: The large language model will generate key objects related to the swimming scenario (such as "swimming personnel", "water surface state"), and input them into the target detection model to obtain detection results, such as the position information of the target, whether the head is above the water surface, and whether there are struggling movements.
[0076] After the result of the target detection model is given to the large language model, the text description of the historical scene record can include: 1. The number of swimming personnel detected.
[0077] 2. The picture size.
[0078] 3. The labels, categories, and confidence scores of each instance.
[0079] 4. The detection box position and abnormal behavior identification (such as drowning alarm).
[0080] Operation mode of the target detection model: 1. Preprocess the input swimming scenario image and detection object (such as swimming personnel) to obtain standardized image data suitable for input into the model, and perform reasoning through a neural network based on the Transformer architecture.
[0081] 2. Set confidence thresholds such as 0.5 and non-maximum suppression (NMS) parameters to remove detection boxes with too high overlap to ensure the accuracy and reliability of the detection results.
[0082] 3. Generate structured detection results, including the number of people, behaviors, behavior characteristics, abnormal probabilities, etc.
[0083] For example: Person B: Detection box [400,500,470,600], behavior is struggling, behavior characteristic is the struggling posture, water area is "swimming prohibited", abnormal probability = 0.65 (abnormal alarm).
[0084] Operation mode of the multi-modal Q&A model: 1. Preprocess the input swimming scenario image and related questions.
[0085] 2. Input the preprocessed image and the formatted relevant questions into the multi-modal question-answering model, and perform inference using a neural network based on the Transformer architecture.
[0086] 3. Generate answers, such as: "The swimmer is swimming normally." or "The swimmer has not moved for a long time and may be abnormal."
[0087] The large language model, object detection model, and multi-modal question-answering model of the present invention can be replaced by other models with visual-language understanding capabilities, such as CLIP, BLIP, MiniGPT-4, to adapt to different computing resources and application requirements.
[0088] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0089] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and the above embodiments and the descriptions in the specification only illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.
Claims
1. A method for graphic and text Q&A and personnel swimming detection, characterized in that, It includes the following steps: Collect swimming scene images, scene historical records, and the inference chain of the scene historical records, and collect user input questions; Generate behavior hypotheses based on the user input questions, and generate associated behavior characteristics based on the behavior hypotheses; Based on the swimming scene images, historical scene records, and the inference chain of the historical scene records, determine whether the associated behavior characteristics are the same as the behavior characteristics of the historical scene records. If they are the same, output the inference chain result that is the same as the inference chain of the scene historical records. If they are different, regenerate new behavior hypotheses until a judgment can be made and then output the corresponding inference chain result. If it is found that there is a lack of historical scene records and a judgment cannot be made, generate a problem of behavior characteristics to be confirmed for the behavior characteristics that cannot be judged; Based on the problem of behavior characteristics to be confirmed, search for relevant behavior characteristics in the swimming scene images, and based on the search results and the problem of behavior characteristics to be confirmed, reason about the problem of behavior characteristics to be confirmed and generate an inference chain to be confirmed; Detect the overlap degree between the inference chain to be confirmed and the inference chain of the scene historical records, and output the inference chain result of the scene historical record with the largest overlap degree; Feed back the inference chain and the inference chain result to the user to complete the detection.
2. The method for graphic and text Q&A and personnel swimming detection according to claim 1, characterized in that, The method for collecting swimming scene images, scene historical records, and the inference chain of the scene historical records, and collecting user input questions includes the following steps: Collect swimming scene images through a camera; Call the scene historical records and the inference chain of the scene historical records through a calling model; Collect user input questions through an input model.
3. The method for graphic text Q&A and personnel swimming detection according to claim 2, characterized in that, Both the calling model and the input model use large language models.
4. A method for graphic text Q&A and personnel swimming detection according to claim 2, characterized in that, The scene historical records include historical image records and historical conversation records. The historical conversation records describe the historical images through text records; The inference chain of the scene historical records reasons about the inference process related to the historical conversation through text records.
5. A method for graphic text Q&A and personnel swimming detection according to claim 1, characterized in that, The method for generating behavior hypotheses based on the user input questions and generating associated behavior characteristics based on the behavior hypotheses includes: Input the user questions into the large language model, and generate behavior hypotheses through the large language model, and generate associated behavior characteristics based on the behavior hypotheses.
6. A method for graphic text Q&A and personnel swimming detection according to claim 1, characterized in that, The method for determining whether the associated behavior characteristics are the same as the behavior characteristics of the historical scene records based on the swimming scene images, historical scene records, and the inference chain of the historical scene records includes: Input the swimming scene images, historical scene records, and the inference chain of the historical scene records into the large language model to generate the behavior characteristics of the historical scene records, and determine whether the associated behavior characteristics are the same as the behavior characteristics of the historical scene records through the large language model.
7. A method for graphic text Q&A and personnel swimming detection according to claim 1, characterized in that, The method further includes: For the problem of behavior characteristics to be confirmed, manually judge the inference chain result by watching the real-time scene monitoring; Confirm the correctness of the manually judged inference chain result and the inference chain result judged by the overlap degree, and select the correct inference chain result to be stored as the scene historical records and the inference chain of the scene historical records.
8. A graphic text Q&A and personnel swimming detection system, characterized in that, It includes: A collection module for collecting swimming scene images, scene historical records, and the inference chain of the scene historical records, and collecting user input questions; An associated behavior characteristic generation module for generating behavior hypotheses based on the user input questions and generating associated behavior characteristics based on the behavior hypotheses; A judgment module, which is used to judge whether the associated behavior features are the same as the behavior features of the historical scene record according to the swimming scene image, the historical scene record, and the reasoning chain of the historical scene record. If they are the same, it outputs the reasoning chain result that is the same as the reasoning chain of the scene historical record. If they are different, it regenerates a new behavior hypothesis until a judgment can be made and then outputs the corresponding reasoning chain result. If it is found that there is a lack of historical scene record and no judgment can be made, it generates a problem of behavior features to be confirmed for the behavior features that cannot be judged; according to the problem of behavior features to be confirmed, it searches for relevant behavior features in the swimming scene image, and based on the search result and the problem of behavior features to be confirmed, it reasons about the problem of behavior features to be confirmed and generates a reasoning chain to be confirmed; it detects the coincidence degree between the reasoning chain to be confirmed and the reasoning chain of the scene historical record, and outputs the reasoning chain result of the scene historical record with the largest coincidence degree. A feedback module, which is used to feedback the reasoning chain and the reasoning chain result to the user to complete the detection.
9. A graphic text Q&A and personnel swimming detection system according to claim 8, characterized in that, The acquisition module is a large language model acquisition module, and the associated behavior feature generation module is a large language model feature generation module; the judgment module is a large language model judgment module, a multimodal question and answer model judgment module, and an object detection model judgment module. The large language model judgment module is used to judge whether the associated behavior features are the same as the behavior features of the historical scene record according to the swimming scene image, the historical scene record, and the reasoning chain of the historical scene record. If they are the same, it outputs the reasoning chain result that is the same as the reasoning chain of the scene historical record. If they are different, it regenerates a new behavior hypothesis until a judgment can be made and then outputs the corresponding reasoning chain result. If it is found that there is a lack of historical scene record and no judgment can be made, it generates a problem of behavior features to be confirmed for the behavior features that cannot be judged. The multimodal question and answer model judgment module is used to input the problem of behavior features to be confirmed, search for relevant behavior features in the swimming scene image according to the problem of behavior features to be confirmed, and based on the search result and the problem of behavior features to be confirmed, it reasons about the problem of behavior features to be confirmed and outputs the reasoning chain to be confirmed; the object detection model judgment module is used to input the reasoning chain to be confirmed, detect the coincidence degree between the reasoning chain to be confirmed and the reasoning chain of the scene historical record, and output the reasoning chain result of the scene historical record with the largest coincidence degree; the feedback module is a large language model feedback module.
10. A graphic text Q&A and personnel swimming detection system according to claim 8, characterized in that, It also includes a selection and recording module, which is used to, for the problem of behavior features to be confirmed, manually judge the reasoning chain result by watching the real-time scene monitoring. Confirm the correctness of the manually judged reasoning chain result and the reasoning chain result judged by coincidence degree, and select the correct reasoning chain result to be stored as the scene historical record and the reasoning chain of the scene historical record.
Citation Information
Patent Citations
Visual dialogue answer generation method and device based on graph perception
CN115129839A
Intelligent question and answer method and device and electronic equipment
CN117828017A
Problem assignment method based on large language model
CN118410876A
Question answering method, system and equipment based on domain-specific knowledge graph and medium
CN118897886A
Multi-source and multi-mode fused knowledge reasoning method, system and device and medium
CN119005340A