Video monitoring terminal and action detection method and equipment based on multi-modal large model
By using multimodal large models for action detection on the video surveillance terminal, the problem of low accuracy of action detection in the prior art is solved, and more efficient and accurate action recognition is achieved.
Patent Information
- Application Number
- CN202411794243.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the accuracy of action detection is low, and it is impossible to effectively detect whether the turntable patrol movement is carried out accurately.
Using video surveillance terminal and action detection method based on multimodal large model, we use patrol videos, determine the target image frame, input suspected video segments and prompt words into the multimodal large model for action recognition.
It improves the accuracy and efficiency of action detection, can accurately identify whether there are preset actions in the inspection video, and reduces the number of inferences of multimodal large models.
Smart Images

Figure CN119942632A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a video surveillance terminal, a motion detection method and a device based on a multimodal large model. Background Art
[0002] Inspection video inspection is a business under the railway maintenance section. Its purpose is to supervise railway inspectors to conduct inspections in accordance with the railway turnout inspection operation instructions and to conduct full inspections on each turnout inspection item. Currently, railway maintenance section staff manually inspect the inspection videos to check whether the railway inspectors in each inspection video have completed qualified inspections on each turnout inspection item, such as whether they squat down to check the height of the rails, whether they use a small hammer to knock and check the rail guard bolts, etc. The existing small model solution can only detect specific targets, such as detecting the action of the limiter bolt. The small model can detect the appearance of the limiter in the video, but it cannot determine whether the limiter bolt has been knocked with a small hammer.
[0003] Therefore, how to accurately detect the turnout inspection action has become a problem that needs to be solved urgently. Summary of the invention
[0004] The embodiments of the present application provide a video surveillance terminal, a motion detection method and a device based on a multimodal large model, so as to solve the problem of low motion detection accuracy in the prior art.
[0005] The present application provides a video monitoring terminal, which includes a display and a processor:
[0006] The processor is configured to:
[0007] Obtain inspection videos recording the inspection process of on-duty personnel;
[0008] Determine that a target image frame containing a preset action exists in the inspection video, and determine a video segment containing the target image frame as a suspected video segment;
[0009] Inputting the suspected video segment and the prompt word corresponding to the preset action contained in the suspected video segment into the multimodal large model to obtain a check result, wherein the check result is used to describe whether the preset action exists in the suspected video segment;
[0010] The display is configured to display the inspection results of the inspection video.
[0011] The above technical solution has the following advantages or beneficial effects: obtaining an inspection video recording the inspection process of the on-duty personnel, determining the target image frame with a preset action in the inspection video, determining the video segment containing the target image frame as a suspected video segment, inputting the suspected video segment and the prompt word corresponding to the preset action contained in the suspected video segment into the multimodal large model, and obtaining an inspection result for describing whether the preset action exists in the suspected video segment. First, determining the suspected video segment with a preset action in the inspection video, and using the multimodal large model to identify the preset action in the suspected video segment, improves the accuracy of action detection, and the multimodal large model only identifies one action based on the prompt word corresponding to the preset action, which improves the efficiency of action detection.
[0012] Furthermore, the processor is specifically configured to:
[0013] Get the sample images saved for each preset action;
[0014] For each example image, the similarity between the example image and each image frame to be detected in the inspection video is determined, and the image frame to be detected corresponding to the highest similarity is determined as the target image frame, and the target image frame contains the preset action corresponding to the example image.
[0015] The above technical solution has the following advantages or beneficial effects: a sample image is saved in advance for each preset action, and then a target image frame similar to the sample image is searched for in all the image frames to be detected, thereby improving the accuracy of determining the target image frame and further improving the accuracy of action detection.
[0016] Furthermore, if a plurality of example images are stored for each preset action, the processor is further configured to:
[0017] For each preset action, deduplication processing is performed on each target image frame corresponding to the preset action.
[0018] The above technical solution has the following advantages or beneficial effects: deduplication processing is performed on the target image frame corresponding to each preset action, which reduces the calculation amount of action detection and improves the efficiency of target detection while ensuring the accuracy of action detection.
[0019] Furthermore, if a plurality of example images are stored for each preset action, the processor is further configured to:
[0020] For each preset action, a preset number of target image frames are retained according to the similarity corresponding to each target image frame containing the preset action and the similarity threshold saved for the preset action.
[0021] The above technical solution has the following advantages or beneficial effects: according to the similarity corresponding to each target image frame, only a preset number of target image frames are retained for each preset action. Under the premise of ensuring the accuracy of action detection, the calculation amount of action detection is reduced and the efficiency of action detection is improved.
[0022] Furthermore, the processor is also configured to:
[0023] In the inspection video, one frame is selected at every set interval as the image frame to be detected.
[0024] The above technical solution has the following advantages or beneficial effects: after the inspection video is acquired, the inspection video is subjected to frame extraction processing, thereby reducing the number of image frames to be processed later and improving the efficiency of dynamic inspection.
[0025] Furthermore, the suspected video segment includes at most one target image frame and other non-target image frames; the processor is specifically configured to:
[0026] For each preset action, sorting each suspected video segment according to the similarity corresponding to the target image frame included in each suspected video segment corresponding to the preset action;
[0027] According to the order of the sorted suspected video segments, the suspected video segments and the prompt words pre-saved for the preset action are input into the multimodal large model in sequence until the inspection result output by the multimodal large model is that the preset action exists in the suspected video segment.
[0028] The above technical solution has the following advantages or beneficial effects: since the higher the similarity corresponding to the target image frame, the greater the possibility of the presence of the preset action, therefore, when using the multimodal large model to process the suspected video segment, each suspected video segment is input into the multimodal large model in turn according to the size of the corresponding similarity, until the inspection result output by the multimodal large model is that the preset action exists in the suspected video segment, thereby reducing the number of multimodal large model inferences as much as possible and improving the accuracy of action detection.
[0029] Furthermore, the processor is also configured to:
[0030] If it is determined that the preset action does not exist according to the inspection result of the last suspected video segment after sorting, it is determined that the executor has not performed the preset action.
[0031] Furthermore, the fine-tuning training process of the multimodal large model includes:
[0032] Acquire an instruction data set, wherein the instruction data set includes target recognition training data, video description training data, and video action recognition training data;
[0033] The original multimodal large model is fine-tuned and trained based on the training data included in the instruction data set to obtain the multimodal large model.
[0034] The above technical solution has the following advantages or beneficial effects: the original multimodal large model is fine-tuned and trained using target recognition training data, video description training data and video action recognition training data, so that the obtained multimodal large model has the ability of target recognition, video description and video action recognition, thereby improving the accuracy of action detection.
[0035] The present application provides an action detection method based on a multimodal large model, the method comprising:
[0036] Obtain inspection videos recording the inspection process of on-duty personnel;
[0037] Determine that a target image frame containing a preset action exists in the inspection video, and determine a video segment containing the target image frame as a suspected video segment;
[0038] The suspected video segment and the prompt word corresponding to the preset action contained in the suspected video segment are input into the multimodal large model to obtain a check result, and the check result is used to describe whether the preset action exists in the suspected video segment.
[0039] The present application provides an action detection device based on a multimodal large model, the device comprising:
[0040] The acquisition module is used to acquire the inspection video recording the inspection process of the on-duty personnel;
[0041] A determination module, used for determining a target image frame containing a preset action in the inspection video, and determining a video segment containing the target image frame as a suspected video segment;
[0042] The detection module is used to input the suspected video segment and the prompt word corresponding to the preset action contained in the suspected video segment into the multimodal large model to obtain an inspection result, and the inspection result is used to describe whether the preset action exists in the suspected video segment.
[0043] The present application also provides an electronic device, which includes a processor, and the processor is used to implement the steps of any of the above-mentioned action detection methods based on a multimodal large model when executing a computer program stored in a memory.
[0044] The present application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of any of the above-mentioned multimodal large model-based action detection methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solution of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0046] Figure 1 A schematic diagram of the structure of a video surveillance terminal provided in an embodiment of the present application;
[0047] Figure 2 A schematic diagram of a sample image provided in an embodiment of the present application;
[0048] Figure 3 A schematic diagram of a portion of image frames in a video provided in an embodiment of the present application;
[0049] Figure 4 A schematic diagram of a portion of image frames in a video provided in an embodiment of the present application;
[0050] Figure 5 A schematic diagram of an action detection process provided in an embodiment of the present application;
[0051] Figure 6 A schematic diagram of an action detection process based on a multimodal large model provided in an embodiment of the present application;
[0052] Figure 7 A schematic diagram of the structure of a motion detection device based on a multimodal large model provided in an embodiment of the present application;
[0053] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0054] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the embodiment of the present application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiment is only a part of the embodiment of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field belong to the scope of protection of this application.
[0055] The existing multimodal large model has the ability to understand videos. In the turnout inspection scenario, there are more than ten actions that need to be detected. The existing solution has two disadvantages: 1) When using the multimodal large model for judgment, if there is no prior knowledge, that is, it is known which action needs to be judged in this video, the multimodal large model reasoning requires a long prompt word or multiple rounds of questions and answers, so the question and answer reasoning time for a video is too long. 2) The inspection video is about 10 minutes long, and most of the existing multimodal large models cannot support the input of the entire video. If accuracy needs to be guaranteed, each video should not exceed 10 seconds. Therefore, the entire video needs to be split, and then the multimodal large model is used to judge in turn. As a result, for a 10-minute inspection video, about 60 inferences will be made, which is too many inferences. In order to improve these two drawbacks, the present application provides a video surveillance terminal, a motion detection method and a device based on a multimodal large model. The video surveillance terminal includes a display and a processor. The display is configured to: display the inspection result of the inspection video; the processor is configured to: obtain the inspection video that records the inspection process of the on-duty personnel; determine the target image frame containing the preset action in the inspection video, and determine the video segment containing the target image frame as a suspected video segment; input the suspected video segment and the prompt word corresponding to the preset action contained in the suspected video segment into the multimodal large model to obtain the inspection result, and the inspection result is used to describe whether there is a preset action in the suspected video segment.
[0056] The video monitoring terminal 10 provided in this application can be used for motion detection. Figure 1 A structural schematic diagram of a video surveillance terminal provided in an embodiment of the present application, the video surveillance terminal 10 includes a display 101 and a processor 102, wherein the processor 102 is configured to obtain an inspection video that records the inspection process of an on-duty personnel; determine that a target image frame with a preset action exists in the inspection video, and determine a video segment containing the target image frame as a suspected video segment; input the suspected video segment and a prompt word corresponding to the preset action contained in the suspected video segment into a multimodal large model to obtain an inspection result, wherein the inspection result is used to describe whether the preset action exists in the suspected video segment, and the display 101 is configured to display the inspection result of the inspection of the inspection video.
[0057] In order to check the on-duty process of the on-duty personnel, in the embodiment of the present application, the inspection video of the on-duty personnel's inspection process can be obtained. The inspection video can be shot by the terminal carried by the on-duty personnel. The inspection video can be a real-time file or an offline file.
[0058] After the inspection video is acquired, it is possible to determine the target image frames in the inspection video where preset actions exist. In an embodiment of the present application, it is possible to determine whether there are preset actions in each image frame in the inspection video based on a pre-trained target detection model. Among them, the preset action can be to check the switch number shooting, check the height direction of the switch, check the insulating joint, check the close fit between the point rail and the base rail, check the close fit between the point rail and the top iron, check the gap between the point rail and the slide bed, check the top iron bolts, check the limiter (or spacer iron) bolts, check the guard rail bolts, check the filling of the work record, etc. It should be noted that the preset action can be one or more, and technicians in this field can configure the preset action as needed.
[0059] After determining that there is a target image frame of a preset action, the video segment containing the target image frame can be determined as a suspected video segment. Exemplarily, the same preset action can correspond to a suspected video segment, and the target image frame included in the suspected video segment is the same preset action, and the suspected video segment can include multiple target image frames of the same preset action. Of course, the suspected video segment can also include multiple target image frames of preset actions.
[0060] After the suspected video segment is determined, the multimodal large model can be used to check the action in the suspected video segment. In order for the multimodal large model to conduct targeted analysis when checking the suspected video, in an embodiment of the present application, a corresponding prompt word can be saved for each preset action. For example, the prompt word "Please analyze whether there is action 1 in the video" can be saved for action 1. In an embodiment of the present application, the suspected video segment and the prompt word corresponding to the preset action contained in the suspected video segment can be input into the multimodal large model to obtain an inspection result. The inspection result is a text used to describe whether there is a preset action in the suspected video segment.
[0061] After obtaining each inspection result output by the multi-modal large model, the processor 102 can control the display 101 to display each inspection result so that the user can view the inspection result. Of course, the processor 102 can also feed back the inspection result to the terminal that sends the inspection video.
[0062] In order to further improve the accuracy of action detection, based on the above embodiment, in the embodiment of the present application, the processor 102 is specifically configured as follows:
[0063] Get the sample images saved for each preset action;
[0064] For each example image, the similarity between the example image and each image frame to be detected in the inspection video is determined, and the image frame to be detected corresponding to the highest similarity is determined as the target image frame, and the target image frame contains the preset action corresponding to the example image.
[0065] In order to accurately detect whether there is a preset action in the inspection video, in an embodiment of the present application, an example image can be saved in advance for each preset action, and the example image saved for each preset action can be one or more. When determining the target image frame in which the preset action exists in the inspection video, the example image saved for each preset action can be obtained, and for each example image, the similarity between each image frame to be detected in the inspection video of the example image is determined. After determining each similarity, the image frame to be detected corresponding to the highest similarity can be determined as the target image frame, and the target image frame contains the preset action corresponding to the example image. In other words, for each example image, the image frame to be detected with the highest similarity to the example image is found in the image frames to be detected as the target image frame.
[0066] In an embodiment of the present application, example images are saved in advance for each preset action, and target image frames similar to the example images are subsequently searched for in all image frames to be detected, thereby improving the accuracy of determining the target image frames and thereby improving the accuracy of action detection.
[0067] In order to further improve the accuracy of action detection, based on the above embodiments, in the embodiment of the present application, if multiple example images are saved for each preset action, the processor 102 is further configured to:
[0068] Since different people have different postures when doing the same action, and when a certain action needs to be completed with the help of tools, different people use different tools. In order to further improve the accuracy of action detection, in an embodiment of the present application, multiple example images can be saved for each preset action, so that more and more accurate target image frames can be determined based on the multiple example images. When multiple example images are saved for a preset action, a target image frame can be determined based on an example image. However, the target image frames determined based on different example images may be the same image frame. For example, the image frame to be detected with the highest similarity corresponding to example image 1 is image frame to be detected 3, the image frame to be detected with the highest similarity corresponding to example image 3 is image frame to be detected 3, and the image frame to be detected with the highest similarity corresponding to example image 5 is image frame to be detected 3. Since the target image frames corresponding to different example images are the same image frame to be detected, in order to avoid repeated operations later, after determining the target image frame corresponding to each preset action. For each preset action, each target image frame corresponding to the preset action can be deduplicated.
[0069] Specifically, assuming that the image frame to be detected with the highest similarity corresponding to example image 1 corresponding to a preset action is image frame 3 to be detected, that is, the target image frame is image frame 3 to be detected; the image frame to be detected with the highest similarity corresponding to example image 3 is image frame 3 to be detected, that is, the target image frame is image frame 3 to be detected; the image frame to be detected with the highest similarity corresponding to example image 25 is image frame 8 to be detected, that is, the target image frame is image frame 8 to be detected; the image frame to be detected with the highest similarity corresponding to example image 5 is image frame 3 to be detected, that is, the target image frame is image frame 3 to be detected. As can be seen from the above content, image frame 3 to be detected appears repeatedly as the target image frame, so the repeated target image frames can be deduplicated and only one target image frame is retained.
[0070] Of course, when deduplication is performed on each target image frame, the similarity between each target image frame can also be calculated separately. When the similarity is higher than a threshold, it can be considered that the corresponding two target image frames are highly similar, and only one target image frame can be retained.
[0071] In the embodiment of the present application, deduplication processing is performed on the target image frames corresponding to each preset action, which reduces the computational complexity of action detection and improves the efficiency of action detection while ensuring the accuracy of action detection.
[0072] In order to further improve the accuracy of action detection, based on the above embodiments, in the embodiment of the present application, if multiple example images are saved for each preset action, the processor 102 is further configured to:
[0073] For each preset action, a preset number of target image frames are retained according to the similarity corresponding to each target image frame containing the preset action and the similarity threshold saved for the preset action.
[0074] Since a large number of example images may be saved for a preset action, a large number of target image frames will be determined in the image frames to be detected, and some of these target image frames may not correspond to the preset action. Therefore, in the embodiment of the present application, for each preset action, a preset number of target image frames can be retained according to the similarity corresponding to each target image frame containing the preset action and the similarity threshold saved for the preset action. In other words, only the preset number of target image frames whose similarity exceeds the preset threshold are retained.
[0075] Specifically, the similarity threshold saved for each preset action can be obtained. During the target image frame screening process, for each preset action, target image frames with a similarity threshold smaller than the preset action are filtered out, and the first three target image frames with the highest similarity are taken from the remaining target image frames as the suspected action locations.
[0076] In an embodiment of the present application, based on the similarity corresponding to each target image frame, only a preset number of target image frames are retained for each preset action. While ensuring the accuracy of action detection, the computational complexity of action detection is reduced and the efficiency of action detection is improved.
[0077] In order to further improve the efficiency of action detection, based on the above embodiments, in the embodiment of the present application, the processor 102 is further configured as follows:
[0078] In the inspection video, one frame is selected at every set interval as the image frame to be detected.
[0079] In order to further improve the efficiency of motion detection, after obtaining the inspection video recording the inspection process of the on-duty personnel, before determining whether there is a target image frame of a preset action in the inspection video, in the embodiment of the present application, the inspection video can also be subjected to frame extraction processing, for example, one frame is selected as the image frame to be detected at every set interval in the inspection video. Exemplarily, the frame extraction frequency can be set to extract one frame group as the image frame to be detected every 10 frames, and all the image frames to be detected in the entire video can be obtained.
[0080] In an embodiment of the present application, after the inspection video is acquired, frame extraction is performed on the inspection video, thereby reducing the number of image frames to be processed subsequently and improving the efficiency of dynamic inspection.
[0081] In order to further improve the accuracy of action detection, based on the above embodiments, in the embodiment of the present application, the suspected video segment includes at most one target image frame and other non-target image frames; the processor 102 is specifically configured as follows:
[0082] For each preset action, sorting each suspected video segment according to the similarity corresponding to the target image frame included in each suspected video segment corresponding to the preset action;
[0083] According to the order of the sorted suspected video segments, the suspected video segments and the prompt words pre-saved for the preset action are input into the multimodal large model in sequence until the inspection result output by the multimodal large model is that the preset action exists in the suspected video segment.
[0084] In the embodiment of the present application, after each target image frame is determined, when determining the suspected video segment containing the target image frame, each suspected video segment may include at most one target image frame and other non-target image frames. In other words, one suspected video segment is determined based on one target image frame, and the number of target image frames is consistent with the number of suspected video segments determined later.
[0085] When the processor 102 inputs the suspected video segments and the prompt words corresponding to the preset actions contained in the suspected video segments into the multimodal large model, the processor 102 may sort each suspected video segment according to the similarity corresponding to the target image frame included in each suspected video segment corresponding to the preset action for each preset action. For example, each suspected video segment is sorted in descending order of similarity.
[0086] After the sorting is completed, the suspected video segments and the prompt words pre-saved for the preset action are input into the multimodal large model in the order of the sorted suspected video segments, until the inspection result output by the multimodal large model is that the preset action exists in the suspected video segment.
[0087] Since the suspected video segments are limited, if it is determined that the preset action does not exist according to the inspection result of the last suspected video segment after sorting, it can be determined that the executor did not perform the preset action.
[0088] Specifically, suppose that for action 1, there are 3 suspected video segments that are suspected to have taken place, and they are sorted by similarity and input into the multimodal large model in sequence. Because it is known that the action to be judged in the suspected video segment is whether action 1 has taken place, the prompt word input into the multimodal large model can be: Is there action 1 in the video? If the answer given by the multimodal large model is yes, the result can be directly output, otherwise the next suspected video of the action is input.
[0089] In order to further improve the accuracy of action detection, based on the above embodiments, in the embodiment of the present application, the fine-tuning training process of the multimodal large model includes:
[0090] Acquire an instruction data set, wherein the instruction data set includes target recognition training data, video description training data, and video action recognition training data;
[0091] The original multimodal large model is fine-tuned and trained based on the training data included in the instruction data set to obtain the multimodal large model.
[0092] Since the open source general multimodal large model does not have professional knowledge in a specific field, for example, professional knowledge in the direction of turnout inspection, it is necessary to construct training data to fine-tune the open source general multimodal large model (hereinafter referred to as the original multimodal large model) and embed the professional knowledge in the specific field into the original multimodal large model, for example, embed the corresponding knowledge of turnouts and the knowledge of corresponding inspection items for turnout inspection into the original multimodal large model. Therefore, in the embodiments of the present application, an instruction data set is designed to allow the original multimodal large model to learn professional knowledge in a specific field. After fine-tuning the original multimodal large model using the instruction data set, the obtained multimodal large model has the ability to recognize professional knowledge in a specific field, for example, the ability to recognize turnout inspection actions. It should be noted that the equipment used for fine-tuning training of the multi-model may be the same as the equipment involved in the above-mentioned embodiments, or it may be different.
[0093] In an embodiment of the present application, the original multimodal large model can be fine-tuned based on the instruction data set to obtain a multimodal large model for action recognition. When performing fine-tuning training, an instruction data set can be obtained. Since the multimodal large model needs to have the ability to recognize specific targets, the ability to analyze video overviews, and the ability to recognize specific actions in videos, in an embodiment of the present application, the instruction data set may include target recognition training data, video description training data, and video action recognition training data. After obtaining the instruction data set, the original multimodal large model can be fine-tuned based on the training data included in the instruction data set to obtain a multimodal large model.
[0094] Specifically, the instruction data set may be constructed according to the following example. Of course, those skilled in the art may also construct it as needed.
[0095] 1) Object recognition training data.
[0096] The original multimodal large model is trained based on the target recognition training data in order to train the multimodal large model to have the perception ability of specific targets. The turnout specific target data can be constructed to enable the multimodal large model to recognize the corresponding special objects on the turnout. Special objects include: limiters, spacers, insulating joints, guardrails, turnout numbers, slide bed plates, etc.
[0097] Exemplarily, the training data corresponding to the stopper target may be:
[0098] {
[0099] "id":0,
[0100] "image":"path / to / image.jpg",
[0101] "width":1080,
[0102] "height":1920,
[0103] "conversations":[
[0104] {"from":"human","value":" \nWhat is in the red box in the picture"},
[0105] {"from":"gpt","text":"limiter"} ]
[0107] }.
[0108] Among them, "id":0 means that the identification of the training data is 0; "image":"path / to / image.jpg" means that the storage location of the image corresponding to the training data is path / to / image.jpg. For example, the image is Figure 2 As shown, Figure 2 A schematic diagram of a sample image provided in an embodiment of the present application is shown in FIG. Figure 2 The object in the rectangular box is the limiter; "width":1080 means the width of the sample image is 1080; "height":1920 means the height of the sample image is 1920; {"from":"human","value":" \nWhat is in the red box in the picture"} indicates the prompt word corresponding to the sample image; {"from":"gpt","text":"limiter"} indicates the standard answer corresponding to the sample image.
[0109] 2) Video description training data.
[0110] The original multimodal large model is trained based on the video description training data in order to train the multimodal large model to have the ability to describe the video in detail.
[0111] Exemplarily, the video description training data may be:
[0112] {
[0113] "id":2,
[0114] "video":"path / to / video.mp4",
[0115] "conversations":[
[0116] {"from":"human","value":" <video>\nPlease describe the video in detail"},
[0117] {"from":"gpt","text":"This video records a turnout inspection and maintenance worker checking the tightness of the point rail and the base rail. The worker wears protective gloves and holds a ruler in his hand. He then leans over and inserts the ruler between the point rail and the base rail. This inspection method is used to determine whether there is a gap between the point rail and the base rail, thereby ensuring the safety and normal operation of the track. This video records the complete action of checking the tightness of the point rail and the base rail, indicating that the maintenance worker's actions are standardized and the inspection of this point is in place"} ]
[0119] }.
[0120] Among them, "id":2 indicates that the identifier of the training data is 2; "video":"path / to / video.mp4" indicates that the storage location of the video corresponding to the training data is path / to / video.mp4. For example, some image frames included in the video are Figure 3 As shown, Figure 3 A schematic diagram of a portion of image frames in a video provided in an embodiment of the present application; {"from":"human","value":" <video>\nPlease describe the video in detail"} indicates the prompt word corresponding to the video; {"from":"gpt","text":"This video records a turnout inspection and maintenance worker checking the tightness of the point rail and the base rail. The worker wears protective gloves and holds a ruler in his hand. He then leans over and inserts the ruler between the point rail and the base rail. This inspection method is used to determine whether there is a gap between the point rail and the base rail, thereby ensuring the safety and normal operation of the track. This video records the complete action of checking the tightness of the point rail and the base rail, indicating that the maintenance worker's actions are standardized and the inspection of this point is in place"} indicates the standard answer corresponding to the video. The standard answer is in the description Figure 3 Contents in the image shown.
[0121] 3) Video action recognition training data.
[0122] The original multimodal large model is trained based on video action recognition training data in order to train the multimodal large model to have the ability to answer video action questions.
[0123] Exemplarily, the video action recognition training data may be:
[0124] {
[0125] "id":1,
[0126] "video":"path / to / video.mp4",
[0127] "conversations":[
[0128] {"from":"human","value":" <video>\nDid the maintenance worker in the video check the limiter bolts? "},
[0129] {"from":"gpt","text":"Yes. In the video, the maintenance worker holds a small hammer and knocks on the bolts on both sides of the limiter to determine whether the limiter bolts are tightened. The video clearly records the maintenance worker's action of checking the limiter bolts with a small hammer."} ]
[0131] }.
[0132] Among them, "id":1 indicates that the identifier of the training data is 1; "video":"path / to / video.mp4" indicates that the storage location of the video corresponding to the training data is path / to / video.mp4. For example, some image frames included in the video are Figure 4 As shown, Figure 4 A schematic diagram of a portion of image frames in a video provided in an embodiment of the present application; {"from":"human","value":" <video>\nDid the maintenance worker check the limiter bolts in the video? "} indicates the prompt word corresponding to the video; {"from":"gpt","text":"Yes. In the video, the maintenance worker holds a small hammer and knocks on the bolts on both sides of the limiter to determine whether the limiter bolts are tightened. The video clearly records the maintenance worker's action of checking the limiter bolts with a small hammer. "} indicates the standard answer corresponding to the video. The standard answer is in the description Figure 4 Contents in the image shown.
[0133] The following describes a specific embodiment of the process of performing motion detection on a video surveillance terminal. Figure 5 A schematic diagram of an action detection process provided in an embodiment of the present application is shown in FIG. Figure 5 As shown in the figure, the action inspection process mainly includes three parts: acquiring video frames, image retrieval and large model reasoning.
[0134] 1) Obtain video frames. First, the inspection video is subjected to frame extraction, with a frame extraction frequency of 1 frame per 10 frames to obtain the image frame to be detected.
[0135] 2) Image retrieval.
[0136] a) 100 sample images are used as queries for each preset action. For example, the current algorithm has 8 preset actions and 800 sample images. In the initialization stage of the algorithm, the features of the 800 sample images are first extracted to obtain a matrix of (1, 800, 2048), which is stored in the memory for later use. Specifically, each sample image obtains a 1*2048 feature matrix, and then the 800 images are combined row by row to obtain 800*2048. Finally, for the convenience of calculation, one dimension is added to become a 1*800*2048 matrix.
[0137] b) All the image frames to be detected are subjected to feature extraction through the image retrieval model to obtain a matrix of (1, n, 2048), where n is the number of image frames to be detected. The image retrieval model can be a backbone network of ResNet in the Treid library.
[0138] c) Calculate the similarity between the feature map matrix of the sample image and the feature map matrix of the image frame to be detected, and then take out the rank 1 corresponding to each sample image, that is, the image frame to be detected with the highest similarity. The image frame with detection is the target image frame.
[0139] d) The 800 target image frames with the highest similarity are sorted and deduplicated for each preset action. Each action is given a list of 1-100 containing the frame number and similarity information of the target image frame. Each preset action is set with a corresponding preset threshold, and the target image frames with a value less than the corresponding preset threshold are filtered out. Then the first three target image frames with the highest similarity are taken as the suspected action locations.
[0140] e) According to the suspected action positions taken in the previous step, the corresponding video segments are cropped to obtain a maximum of 24 suspected video segments.
[0141] 3) Large model reasoning.
[0142] The suspected video segments filtered out by the image retrieval part are input into the multimodal large model for judgment.
[0143] For example, for preset action 1, there are 3 suspected video segments, which are input into the multimodal large model in sequence according to the similarity. Because it is known that the suspected video needs to determine whether preset action 1 occurs, the question input into the multimodal large model can be: Did action 1 occur in the video? If the answer of the multimodal large model is yes, the detection result is directly output, otherwise the next suspected video segment of the preset action is input.
[0144] In the prior art, it is impossible to detect action in time sequence based on a small model, and the understanding ability based on a multimodal large model can identify the action in the video. However, due to the long video length and the variety of actions, the solution based on a single multimodal large model needs to segment all videos and then infer each segment. In addition, if there is no prior knowledge, the accuracy of letting the large model identify what action is in the video is low. In an embodiment of the present application, each action only needs to be inferred up to 3 times, and each recognition of the multimodal large model only needs to identify whether it is a specified action. In other words, for each video reasoning, it is known which action the video needs to judge, and the reasoning of a video segment is up to nx3 times, where n is the number of actions. The accuracy and efficiency of action recognition are effectively improved.
[0145] The action detection method based on the combination of large and small models proposed in this application first locates the target action position based on the image retrieval model, and then performs video understanding and recognition of the action based on the multimodal large model, thereby achieving accurate recognition and judgment of different actions in long videos in inspection scenarios, while effectively reducing the number of inference times of the multimodal large model.
[0146] It should be noted that the video surveillance terminal provided in this application is not only used for motion detection in turnout inspection scenarios, but is also applicable to other scenarios, and technicians in this field can configure it as needed.
[0147] On the basis of the above embodiments, an embodiment of the present application further provides a motion detection method based on a multimodal large model. The motion detection method based on a multimodal large model is consistent with the methods in the above embodiments and will not be repeated here. Figure 6 A schematic diagram of an action detection process based on a multimodal large model provided in an embodiment of the present application, the process includes:
[0148] S601: Obtaining an inspection video recording the inspection process of the on-duty personnel.
[0149] S602: Determine whether a target image frame containing a preset action exists in the inspection video, and determine a video segment containing the target image frame as a suspected video segment.
[0150] S603: Input the suspected video segment and the prompt word corresponding to the preset action contained in the suspected video segment into the multimodal large model to obtain a check result, where the check result is used to describe whether the preset action exists in the suspected video segment.
[0151] In a possible implementation manner, determining that a target image frame having a preset action in the inspection video includes:
[0152] Get the sample images saved for each preset action;
[0153] For each example image, the similarity between the example image and each image frame to be detected in the inspection video is determined, and the image frame to be detected corresponding to the highest similarity is determined as the target image frame, and the target image frame contains the preset action corresponding to the example image.
[0154] In a possible implementation manner, if a plurality of example images are stored for each preset action, the method further includes:
[0155] For each preset action, deduplication processing is performed on each target image frame corresponding to the preset action.
[0156] In a possible implementation manner, if a plurality of example images are stored for each preset action, the method further includes:
[0157] For each preset action, a preset number of target image frames are retained according to the similarity corresponding to each target image frame containing the preset action and the similarity threshold saved for the preset action.
[0158] In a possible implementation, after acquiring the inspection video recording the inspection process of the on-duty personnel and before determining that there is a target image frame of a preset action in the inspection video, the method further includes:
[0159] In the inspection video, one frame is selected at every set interval as the image frame to be detected.
[0160] In a possible implementation manner, the suspected video segment includes at most one target image frame and other non-target image frames;
[0161] The step of inputting the suspected video segment and the prompt word corresponding to the preset action contained in the suspected video segment into the multimodal large model to obtain the inspection result includes:
[0162] For each preset action, sorting each suspected video segment according to the similarity corresponding to the target image frame included in each suspected video segment corresponding to the preset action;
[0163] According to the order of the sorted suspected video segments, the suspected video segments and the prompt words pre-saved for the preset action are input into the multimodal large model in sequence until the inspection result output by the multimodal large model is that the preset action exists in the suspected video segment.
[0164] In a possible implementation, the method further includes:
[0165] If it is determined that the preset action does not exist according to the inspection result of the last suspected video segment after sorting, it is determined that the executor has not performed the preset action.
[0166] In a possible implementation, the fine-tuning training process of the multimodal large model includes:
[0167] Acquire an instruction data set, wherein the instruction data set includes target recognition training data, video description training data, and video action recognition training data;
[0168] The original multimodal large model is fine-tuned and trained based on the training data included in the instruction data set to obtain the multimodal large model.
[0169] Based on the above embodiments, Figure 7 A schematic diagram of the structure of a motion detection device based on a multimodal large model provided in an embodiment of the present application, the device comprising:
[0170] The acquisition module 701 is used to acquire the inspection video recording the inspection process of the on-duty personnel;
[0171] A determination module 702 is used to determine that a target image frame having a preset action exists in the inspection video, and determine a video segment containing the target image frame as a suspected video segment;
[0172] The detection module 703 is used to input the suspected video segment and the prompt word corresponding to the preset action contained in the suspected video segment into the multimodal large model to obtain a check result, and the check result is used to describe whether the preset action exists in the suspected video segment.
[0173] In a possible implementation, the acquisition module 701 is further used to acquire a sample image saved for each preset action;
[0174] The determination module 702 is specifically used to determine, for each example image, the similarity between the example image and each image frame to be detected in the inspection video, and determine the image frame to be detected corresponding to the highest similarity as the target image frame, wherein the target image frame contains the preset action corresponding to the example image.
[0175] In a possible implementation manner, if a plurality of example images are stored for each preset action, the device further includes:
[0176] The screening module 704 is used to perform deduplication processing on each target image frame corresponding to each preset action.
[0177] In a possible implementation, if multiple example images are saved for each preset action, the screening module 704 is further used to retain a preset number of target image frames for each preset action based on the similarity corresponding to each target image frame containing the preset action and the similarity threshold saved for the preset action.
[0178] In a possible implementation manner, the screening module 704 is further configured to select one frame as the image frame to be detected at every set number of intervals in the inspection video.
[0179] In a possible implementation manner, the suspected video segment includes at most one target image frame and other non-target image frames;
[0180] The detection module 703 is specifically used to sort each suspected video segment for each preset action according to the similarity corresponding to the target image frame included in each suspected video segment corresponding to the preset action; and input the suspected video segments and the prompt words pre-saved for the preset action into the multimodal large model in sequence according to the order of the sorted suspected video segments, until the inspection result output by the multimodal large model is that the preset action exists in the suspected video segment.
[0181] In a possible implementation, the detection module 703 is further configured to determine that the executor has not performed the preset action if it is determined that the preset action does not exist according to the inspection result of the last suspected video segment after sorting.
[0182] In a possible implementation, the acquisition module 701 is further used to acquire an instruction data set, wherein the instruction data set includes target recognition training data, video description training data, and video action recognition training data;
[0183] The training module 705 is used to fine-tune the original multimodal large model based on the training data included in the instruction data set to obtain the multimodal large model.
[0184] Based on the same inventive concept, an embodiment of the present application further provides an electronic device, Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown in FIG. Figure 8 As shown, it includes: a processor 801, a communication interface 802, a memory 803 and a communication bus 804, wherein the processor 801, the communication interface 802, and the memory 803 communicate with each other through the communication bus 804;
[0185] The memory 803 stores a computer program. When the program is executed by the processor 801, the processor 801 executes the steps of the action detection method based on the multimodal large model described in the above embodiments.
[0186] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface 802 is used for communication between the above electronic device and other devices. The memory may include a random access memory (RAM) and may also include a non-volatile memory (NVM), such as at least one disk storage. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0187] The above-mentioned processor can be a general-purpose processor, including a central processing unit, a network processor (Network Processor, NP), etc.; it can also be a digital signal processing processor (Digital Signal Processing, DSP), an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, etc.
[0188] On the basis of the above embodiments, an embodiment of the present invention further provides a computer-readable storage medium, which stores a computer program executable by a processor. When the program runs on the processor, the processor implements the steps of the action detection method based on the multimodal large model described in the above embodiments.
[0189] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0190] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0191] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0192] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0193] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.< / video> < / video> < / video> < / video>
Claims
1. A video monitoring terminal, characterized in that: The video surveillance terminal includes a display and a processor: The processor is configured to: Obtain inspection videos recording the inspection process of on-duty personnel; Determine that a target image frame containing a preset action exists in the inspection video, and determine a video segment containing the target image frame as a suspected video segment; Inputting the suspected video segment and the prompt word corresponding to the preset action contained in the suspected video segment into the multimodal large model to obtain a check result, wherein the check result is used to describe whether the preset action exists in the suspected video segment; The display is configured to display the inspection results of the inspection video.
2. The video surveillance terminal according to claim 1, characterized in that: The processor is specifically configured to: Get the sample images saved for each preset action; For each example image, the similarity between the example image and each image frame to be detected in the inspection video is determined, and the image frame to be detected corresponding to the highest similarity is determined as the target image frame, and the target image frame contains the preset action corresponding to the example image.
3. The video surveillance terminal according to claim 2, characterized in that: If a plurality of example images are stored for each preset action, the processor is further configured to: For each preset action, deduplication processing is performed on each target image frame corresponding to the preset action.
4. The video surveillance terminal according to claim 2, characterized in that: If a plurality of example images are stored for each preset action, the processor is further configured to: For each preset action, a preset number of target image frames are retained according to the similarity corresponding to each target image frame containing the preset action and the similarity threshold saved for the preset action.
5. The video surveillance terminal according to claim 2, characterized in that: The processor is further configured to: In the inspection video, one frame is selected at every set interval as the image frame to be detected.
6. The video surveillance terminal according to claim 2, characterized in that: The suspected video segment includes at most one target image frame and other non-target image frames; the processor is specifically configured to: For each preset action, sorting each suspected video segment according to the similarity corresponding to the target image frame included in each suspected video segment corresponding to the preset action; According to the order of the sorted suspected video segments, the suspected video segments and the prompt words pre-saved for the preset action are input into the multimodal large model in sequence until the inspection result output by the multimodal large model is that the preset action exists in the suspected video segment.
7. The video surveillance terminal according to claim 6, characterized in that: The processor is further configured to: If it is determined that the preset action does not exist according to the inspection result of the last suspected video segment after sorting, it is determined that the executor has not performed the preset action.
8. The video surveillance terminal according to claim 1, characterized in that: The fine-tuning training process of the multimodal large model includes: Acquire an instruction data set, wherein the instruction data set includes target recognition training data, video description training data, and video action recognition training data; The original multimodal large model is fine-tuned and trained based on the training data included in the instruction data set to obtain the multimodal large model.
9. A method for motion detection based on a multimodal large model, characterized in that: The method comprises: Obtain inspection videos recording the inspection process of on-duty personnel; Determine that a target image frame containing a preset action exists in the inspection video, and determine a video segment containing the target image frame as a suspected video segment; The suspected video segment and the prompt word corresponding to the preset action contained in the suspected video segment are input into the multimodal large model to obtain a check result, and the check result is used to describe whether the preset action exists in the suspected video segment.
10. An electronic device, characterized in that: The electronic device comprises a processor, and the processor is used to implement the steps of the action detection method based on a multimodal large model as claimed in claim 9 when executing a computer program stored in a memory.