Video defense point location batch drawing method based on multi-modal large model, storage medium, equipment and computer program product
Through the multi-modal large model-based video point batch drawing method, the problem of automatic drawing of video surveillance area in the prior art is solved, and the problem of time-consuming and labor-intensive drawing of video surveillance area and image quality affecting the model training effect is achieved, and high-precision drawing of video surveillance target area and multiple requirements are achieved.
Patent Information
- Application Number
- CN202510279761.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
In the automatic drawing of video surveillance area, the data volume is large, manual labeling is time-consuming and labor-intensive, the image quality is uneven and the model training effect is poor, and the scalability is not enough to meet the drawing of multiple needs.
The video defense point batch drawing method based on multimodal large model is adopted. By collecting video surveillance images for manual annotation, prompt words and question-and-answer pairs of multimodal large model are set, and the model is trained until the IOU loss function converges, so that high-precision drawing of the video surveillance target area is achieved.
The amount of manual annotation in the video surveillance target area is reduced, the accuracy and robustness of prediction are improved, and the prediction and drawing of the video surveillance target area with high accuracy can be achieved in different preset requirements and monitoring environments.
Smart Images

Figure CN120220020A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence. Specifically, it relates to a method for batch drawing of video defense points based on a multimodal large model, a storage medium, a device, and a computer program product. Background Art
[0002] With the acceleration of the urbanization process, traffic accidents, public security incidents, etc. occurring on each street and region are increasing day by day. The deployment of video surveillance is particularly important. Traditional video deployment tasks are numerous and the types of video surveillance areas are relatively complex. Using manual methods to draw video surveillance areas, although the accuracy is high, the labor and time costs are high, time-consuming and laborious, and it is impossible to achieve the deployment of large-scale video surveillance area drawing.
[0003] Currently, the automatic drawing of video surveillance areas is achieved through deep learning technology. Specifically, by collecting a large number of video surveillance images and manually annotating the corresponding surveillance area information, a deep learning network model is trained to achieve the prediction of surveillance areas in video surveillance images. Generally speaking, the data volume of video surveillance images is very large, and each video surveillance image needs to be annotated with a surveillance area, which will consume a lot of manpower and time. In addition, video surveillance images are affected by factors such as lighting, occlusion, and angle, resulting in uneven image quality, which affects the training effect of the deep learning network model. Moreover, the deep learning network model can only draw surveillance areas of preset categories and cannot recognize situations outside the preset categories, with poor scalability and unable to meet the multi-demand drawing of video surveillance areas. Summary of the Invention
[0004] Aiming at the problems existing in the prior art, the present invention provides a method for batch drawing of video defense points based on a multimodal large model, a storage medium, a device, and a computer program product, which realizes the high-precision drawing of the video surveillance target area through the multimodal large model and reduces the human and material costs.
[0005] To achieve the above technical objectives, the present invention adopts the following technical solutions: A method for batch drawing of video defense points based on a multimodal large model specifically includes the following steps:
[0006] Step S1: Collect video surveillance images with different surveillance target areas and manually annotate the target surveillance areas on each video surveillance image;
[0007] Step S2: Set the prompt of the multimodal large model and set the question-and-answer pairs of the multimodal large model according to the collected video surveillance images;
[0008] Step S3: Input the collected video surveillance images into the multi-modal large model in sequence for training until the IOU loss function converges, and complete the training of the multi-modal large model;
[0009] Step S4: Input the video surveillance images to be batch-drawn with the monitored target areas into the trained multi-modal large model, determine the preset requirements for the user's target monitored area in the question-answer pair, and predict the graphic vertex coordinate sets of the normalized target monitored areas in each video surveillance image;
[0010] Step S5: Draw the monitored target areas according to the predicted graphic vertex coordinate sets of the target monitored areas.
[0011] Further, in Step S1, manually mark the graphic vertex coordinate set T i ={(x i1 ,y i1 ),…,(xi j ,yi j ),…,(xi n ,yi n}, where i represents the index of the video surveillance image, n represents the number of graphic vertices collected for the target monitored area on the i-th video surveillance image, j represents the index of n, and (xi j ,yi j ) represents the coordinates of the j-th vertex of the target monitored area on the i-th video surveillance image after normalization.
[0012] Further, the prompt of the multi-modal large model is specifically: As a video surveillance area division system, it is necessary to give the graphic vertex set covering the monitored target area and represent it in a normalized manner according to the preset requirements of the user's target monitored area and the input video surveillance images.
[0013] Further, the question-answer pair of the multi-modal large model is saved in the json file format and consists of the storage path information of the collected video surveillance images, the annotation information of the target monitored area in the corresponding video surveillance images, and the preset requirements of the user's target monitored area. The specific content is:
[0014] {
[0015] "image":"path / to / imagei.jpg",
[0016] "question":"The preset requirements for the user's target monitored area",
[0017] "label":[T i ={(xi1,yi1),…,(xi j ,yij ),…,(xi n ,yi n )}]
[0018] }。
[0019] Furthermore, the specific process of step S3 is as follows:
[0020] Step S3.1: Initialize the weight parameters of the multi-modal large model;
[0021] Step S3.2: Input the collected video surveillance images into the multi-modal large model, and predict the set of graphic vertex coordinates of the normalized target surveillance area in the corresponding video surveillance images according to the corresponding Q&A pairs and the prompt word prompt;
[0022] Step S3.3: Calculate the IOU loss function according to the set of graphic vertex coordinates of the predicted surveillance area and the set of graphic vertex coordinates of the manually labeled surveillance area;
[0023] Step S3.4: If the IOU loss function converges, complete the training of the multi-modal large model; otherwise, update the weight parameters of the multi-modal large model according to the gradient of the IOU loss function, and repeat steps S3.2 - S3.4.
[0024] Furthermore, the calculation process of the IOU loss function in step S3.3 is as follows:
[0025]
[0026] where θ represents the weight parameters of the multi-modal large model, S predict represents the area of the graph enclosed by the set of graphic vertex coordinates of the predicted surveillance area on the video surveillance image, and S label represents the area of the graph enclosed by the set of graphic vertex coordinates of the manually labeled surveillance area on the corresponding video surveillance image.
[0027] Furthermore, the process of updating the weight parameters of the multi-modal large model in step S3.4 is as follows:
[0028]
[0029] where, represents the updated weight parameters of the multi-modal large model, η represents the learning rate, represents the gradient of the IOU loss function.
[0030] Furthermore, the present invention also provides a computer-readable storage medium storing a computer program, and the computer program enables a computer to execute the method for batch drawing of video defense points based on a multi-modal large model.
[0031] Furthermore, the present invention also provides an electronic device, comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the method for batch drawing of video defense points based on a multi-modal large model is implemented.
[0032] Furthermore, the present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method for batch drawing of video defense points based on a multi-modal large model is implemented.
[0033] Compared with the prior art, the present invention has the following beneficial effects: The method for batch drawing of video defense points based on a multi-modal large model of the present invention sets up question-and-answer pairs of the multi-modal large model according to the collected video surveillance images, and sets up the prompt word prompt of the multi-modal large model. With only a small number of video surveillance images, the multi-modal large model can more accurately and quickly capture the user's intentions and needs, so as to more precisely predict the video surveillance target area. Compared with deep learning technology, it can greatly reduce the amount of manual annotation of the video surveillance target area; at the same time, using the huge memory reserve of the multi-modal large model, even for preset requirements and surveillance environments that have not been trained, high-accuracy prediction and drawing of the video surveillance target area can be achieved, with good robustness and generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is a flowchart of the method for batch drawing of video defense points based on a multi-modal large model of the present invention;
[0035] Figure 2 is a comparison diagram of the video surveillance target area drawn by the method for batch drawing of video defense points of the present invention and the real video surveillance target area. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] The technical solutions of the present invention will be further explained below with reference to the accompanying drawings.
[0037] As Figure 1 is a flowchart of the method for batch drawing of video defense points based on a multi-modal large model of the present invention, and the method for batch drawing of video defense points specifically includes the following steps:
[0038] Step S1, collect video surveillance images with different surveillance target areas, and manually annotate the target surveillance areas on each video surveillance image. It is necessary to manually annotate the graphic vertex coordinate set T of the normalized target surveillance area i ={(x i1 , y i1 ),…,(xi j , yi j ),…,(xin , yi n )}, where i represents the index of the video surveillance image, n represents the number of graphic vertices collected in the target surveillance area of the i-th video surveillance image, j represents the index of n, and (xi j , yi j ) represents the coordinates of the j-th vertex in the target surveillance area of the i-th video surveillance image after normalization processing.
[0039] Step S2: Set the prompt of the multi-modal large model to guide the multi-modal large model to complete the prediction and drawing of the surveillance target area in the video surveillance image, and set the question-and-answer pairs of the multi-modal large model according to the collected video surveillance images. With only a small number of video surveillance images, the multi-modal large model can more accurately and quickly capture the user's intentions and needs, thereby more precisely predicting the video surveillance target area. Compared with deep learning technology, it can greatly reduce the amount of manual annotation of the video surveillance target area; among them, the prompt of the multi-modal large model is specifically: as a video surveillance area division system, it is necessary to give a set of graphic vertices covering the surveillance target area according to the preset requirements of the user's target surveillance area and the input video surveillance image, and represent it in a normalized manner; the question-and-answer pairs of the multi-modal large model are saved in the json file format and consist of the storage path information of the collected video surveillance images, the annotation information of the target surveillance area in the corresponding video surveillance images, and the preset requirements of the user's target surveillance area. The preset requirements of the user's target surveillance area can be set as construction areas, highway areas, etc. By integrating information from different sources through the question-and-answer pairs, the diversity of information helps the multi-modal large model to more comprehensively understand the situation of the surveillance target area, thereby improving the accuracy of prediction. The specific content of the question-and-answer pairs is:
[0040] {
[0041] "image": "path / to / imagei.jpg",
[0042] "question": "The preset requirements of the user's target surveillance area",
[0043] "label": [T i = {(xi1, yi1), …, (xi j , yi j ), …, (xi n , yi n )}]
[0044] }。
[0045] Step S3: Input the collected video surveillance images into the multi-modal large model for training in sequence until the IOU loss function converges, and complete the training of the multi-modal large model; specifically including the following sub-steps:
[0046] Step S3.1: Initialize the weight parameters of the multi-modal large model;
[0047] Step S3.2: Input the collected video surveillance images into the multi-modal large model, and according to the corresponding Q&A pairs and the prompt, predict the set of graphic vertex coordinates of the normalized target surveillance area in the corresponding video surveillance image;
[0048] Step S3.3: Calculate the IOU loss function according to the set of graphic vertex coordinates of the predicted surveillance area and the set of graphic vertex coordinates of the manually annotated surveillance area corresponding to it:
[0049]
[0050] where, θ represents the weight parameters of the multi-modal large model, S predict represents the area of the graphic enclosed by the set of graphic vertex coordinates of the predicted surveillance area on the video surveillance image, and S label represents the area of the graphic enclosed by the set of graphic vertex coordinates of the manually annotated surveillance area on the corresponding video surveillance image.
[0051] Step S3.4: If the IOU loss function converges, complete the training of the multi-modal large model; otherwise, update the weight parameters of the multi-modal large model according to the gradient of the IOU loss function, and repeat steps S3.2 - S3.4. In the present invention, updating the weight parameters of the multi-modal large model through the gradient of the IOU loss function can better learn the common features and differences between different modal data, thereby improving the generalization ability on different preset requirements and data sets; at the same time, the computational efficiency of the gradient update of the IOU loss function is relatively high, which helps to reduce the computational time in the training and inference processes of the multi-modal large model and improve the overall efficiency.
[0052] The process of updating the weight parameters of the multi-modal large model in the present invention is as follows:
[0053]
[0054] where, represents the updated weight parameters of the multi-modal large model, η represents the learning rate, which controls the update step of the weight parameters; represents the gradient of the IOU loss function.
[0055] Step S4: Input the video surveillance images of the monitoring target areas to be batch-drawn into the trained multi-modal large model, determine the preset requirements of the user's target monitoring areas in the Q&A pairs, and predict the graphic vertex coordinate sets of the normalized target monitoring areas in each video surveillance image.
[0056] Step S5: Draw the monitoring target areas according to the predicted graphic vertex coordinate sets of the target monitoring areas. Since the multi-modal large model has a huge memory reserve, it can achieve high-accuracy prediction and drawing of video surveillance target areas even for preset requirements and monitoring environments that have not been trained, and has good robustness and generalization ability.
[0057] As Figure 2 shown, the green box shows the monitoring target areas drawn by the video defense point batch drawing method based on the multi-modal large model of the present invention, and the red box shows the real monitoring target areas. It can be seen that the monitoring target areas drawn by the present invention basically cover the entire range of the real monitoring target areas, indicating that the video defense point batch drawing method of the present invention realizes high-accuracy drawing of video surveillance target areas and reduces the human and material costs.
[0058] In a technical solution of the present invention, there is also provided a computer-readable storage medium storing a computer program, and the computer program causes a computer to execute the video defense point batch drawing method based on the multi-modal large model.
[0059] In a technical solution of the present invention, there is also provided an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the video defense point batch drawing method based on the multi-modal large model is implemented.
[0060] In a technical solution of the present invention, there is also provided a computer program product including a computer program, and when the computer program is executed by a processor, the video defense point batch drawing method based on the multi-modal large model is implemented.
[0061] In the embodiments disclosed in the present application, the computer storage medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The computer storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the computer storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0062] Those of ordinary skill in the art will recognize that the units and algorithm steps of the examples described in connection with the embodiments disclosed in the present application can be implemented in electronic hardware or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled artisans may use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.
[0063] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art of this technology, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.
Claims
1. A method for batch drawing of video defense points based on a multimodal large model, characterized in that: The specific steps include: Step S1, collecting video surveillance images of different surveillance target areas, and manually marking the target surveillance area on each video surveillance image; Step S2: setting a prompt word prompt of the multimodal large model, and setting a question-answer pair of the multimodal large model according to the collected video surveillance images; Step S3: input the collected video surveillance images into the multimodal large model in sequence for training until the IOU loss function converges, thus completing the training of the multimodal large model; Step S4: input the video surveillance images of the target surveillance area to be drawn in batches into the trained multimodal large model, determine the preset requirements of the user's target surveillance area in the question-answer pair, and predict the normalized graphic vertex coordinate set of the target surveillance area in each video surveillance image; Step S5: Draw the monitoring target area according to the predicted graphic vertex coordinate set of the target monitoring area.
2. According to the method for batch drawing of video defense points based on a multimodal large model according to claim 1, it is characterized in that: Step S1: manually mark the normalized target monitoring area's vertex coordinate set T i ={(xi1,yi1),…,(xi j ,yi j ),…,(xi n ,yi n )}, where i represents the index of the video surveillance image, n represents the number of graph vertices collected in the target surveillance area on the i-th video surveillance image, j represents the index of n, (xi j ,yi j ) represents the normalized coordinates of the j-th vertex of the target surveillance area on the i-th video surveillance image.
3. According to the method for batch drawing of video defense points based on a multimodal large model according to claim 2, it is characterized in that: The prompt word prompt of the multimodal large model is specifically: as a video surveillance area division system, it is necessary to provide a set of graphic vertices covering the surveillance target area according to the preset requirements of the user's target surveillance area and the input video surveillance image, and use normalized representation.
4. According to the method for batch drawing of video defense points based on a multimodal large model according to claim 3, it is characterized in that: The question-answer pairs of the multimodal large model are saved in the json file format, which consists of the storage path information of the collected video surveillance images, the annotation information of the target surveillance area in the corresponding video surveillance images, and the preset requirements of the user's target surveillance area. The specific contents are as follows: { "image":"path / to / imagei.jpg", "question":"Preset requirements for user target monitoring areas", "label":[T i ={(xi1,yi1),…,(xi j ,yi j ),…,xi n ,yi n )}] }。 5. According to the method for batch drawing of video defense points based on multimodal large model according to claim 4, it is characterized in that: The specific process of step S3 is: Step S3.1, initializing the weight parameters of the multimodal large model; Step S3.2: input the collected video surveillance images into the multimodal large model, and predict the normalized graphic vertex coordinate set of the target surveillance area in the corresponding video surveillance image according to the corresponding question-answer pair and the prompt word prompt; Step S3.3, calculating the IOU loss function according to the predicted graphic vertex coordinate set of the monitoring area and the corresponding manually annotated graphic vertex coordinate set of the monitoring area; Step S3.4: If the IOU loss function converges, the training of the multimodal large model is completed; otherwise, the weight parameters of the multimodal large model are updated according to the gradient of the IOU loss function, and steps S3.2-S3.4 are repeated.
6. The method for batch drawing of video defense points based on a multimodal large model according to claim 5 is characterized in that: The calculation process of the IOU loss function in step S3.3 is: Among them, θ represents the weight parameter of the multimodal large model, S predict represents the area of the graph enclosed by the coordinates of the graph vertices of the predicted monitoring area on the video surveillance image, S label Represents the graphic area enclosed by the coordinates of the graphic vertices of the manually marked surveillance area on the corresponding video surveillance image.
7. The method for batch drawing of video defense points based on a multimodal large model according to claim 6 is characterized in that: The process of updating the weight parameters of the multimodal large model in step S3.4 is: in, represents the weight parameter of the updated multimodal large model, η represents the learning rate, Represents the gradient of the IOU loss function.
8. A computer-readable storage medium storing a computer program, characterized in that: The computer program enables the computer to execute the method for batch drawing of video deployment points based on a multimodal large model as described in any one of claims 1-7.
9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for batch drawing of video defense points based on a multimodal large model as described in any one of claims 1 to 7 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for batch drawing of video defense points based on a multimodal large model described in any one of claims 1 to 7 is implemented.