Setting behavior recognition method and device based on image data
By performing periodic sampling and multimodal analysis of video frames in video surveillance technology, and using visual processing models and natural language processing technology to identify and set behaviors, the problem of low recognition accuracy in complex scenarios is solved in the existing technology, achieving higher recognition accuracy and resource savings.
Patent Information
- Application Number
- CN202411972653.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-12-30
AI Technical Summary
The existing video surveillance technology has poor detection effect in complex scenarios, high false alarm rate, and it is difficult to identify uncommon or unconventional action combinations, resulting in a low recognition accuracy rate.
Using a setting behavior recognition method based on image data, video frames are periodically sampled, and visual processing models and natural language processing technology are used to describe and analyze the video data in text, determine whether there is a setting behavior, and calculate its probability.
It improves the accuracy of identification of specific behaviors in video frames, reduces misjudgment, reduces dependence on on-site supervisors, and saves processing resources.
Smart Images

Figure CN119964233A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to intelligent video surveillance technology based on large model technology, and in particular to a method and device for identifying set behaviors based on image data. Background Art
[0002] At present, public places install cameras to collect corresponding video streams, and use human key point detection, behavior analysis and other methods in the field of computer vision to detect public security incidents. Although such methods can achieve certain detection effects, their detection effects are not good in complex scenes. Specifically, due to the diversity of human actions, a high false alarm rate is caused. In addition, if the camera is blocked, the possibility of misjudgment will also increase. In addition, although the camera can capture the behavior of public security incidents well, its recognition accuracy will be greatly reduced for those uncommon or unconventional action combinations. Therefore, the video detection technology of existing cameras does not perform well in complex scenes, resulting in a low accuracy rate of video recognition. Summary of the invention
[0003] The present application provides a method and device for identifying set behaviors based on image data, so as to at least solve the above technical problems existing in the prior art.
[0004] According to a first aspect of the present application, a method for identifying a set behavior based on image data is provided, comprising:
[0005] Periodically sampling the video frames acquired by the image acquisition unit to obtain sampled video data;
[0006] Performing object recognition on the sampled video data, and taking the video data containing the set objects in the sampled video data, and the number of the set objects is greater than or equal to a set threshold, as valid video data;
[0007] Selecting a corresponding visual processing model and a natural language processing model; generating a prompt statement for multimodal content analysis for the visual processing model;
[0008] Using the visual processing model to perform text description on the valid video data, and generate image captions for the valid video data;
[0009] The image subtitles are input into a corresponding processing model to determine whether a set behavior exists in the valid video data, determine the probability of the set behavior existing, and output the determined probability.
[0010] In some executable embodiments, determining the probability of the existence of a set behavior includes:
[0011] Generate image captions for the valid video data frame by frame, trigger the natural language processing type to perform multimodal content analysis on the image captions of each frame, and determine whether the frame image has a set behavior and a first probability of the set behavior;
[0012] And, inputting the image caption and the video data frame into the multimodal large model, and determining whether the frame image has a set behavior and a second probability of the set behavior through the multimodal large model;
[0013] The average value of the first probability and the second probability is taken as the probability of the existence of the setting behavior of the valid video data.
[0014] In some executable embodiments, the performing of object recognition on the sampled video data and taking the video data containing the set objects in the sampled video data and the number of the set objects being greater than or equal to a set threshold as the valid video data includes:
[0015] For the sampled video data, the YOLO in the target detection algorithm is used to detect the objects in the video frames, count the number of objects in the video frames, and remove the video frames with the number of objects less than the set threshold in the sampled video data as valid video data.
[0016] In some executable embodiments, the detecting of objects in video frames by using YOLO in the target detection algorithm includes:
[0017] The image information features of each video frame of the sampled video data are extracted through the backbone network Backbone, and the image information features are convoluted to obtain the local feature information of each video frame;
[0018] Normalize the local feature information of each video frame to reduce the covariate shift of the local feature information;
[0019] The normalized local feature information is input into the activation function to determine the object in the video frame.
[0020] In some executable embodiments, the object is a person; and the set behavior includes excessive physical contact behavior.
[0021] According to a second aspect of the present application, there is provided a device for identifying a set behavior based on image data, comprising:
[0022] A sampling unit, used for periodically sampling the video frames acquired by the image acquisition unit to obtain sampled video data;
[0023] The recognition unit is used to perform object recognition on the sampled video data, and to take the video data containing the set objects in the sampled video data, and the number of the set objects is greater than or equal to a set threshold, as the valid video data;
[0024] A selection unit, used to select a corresponding visual processing model and a natural language processing model;
[0025] A first generating unit, configured to generate a prompt statement for multimodal content analysis for the visual processing model;
[0026] A second generating unit is used to perform text description on the effective video data by using the visual processing model to generate image subtitles for the effective video data;
[0027] The determination unit is used to input the image subtitle into a corresponding processing model to determine whether a set behavior exists in the valid video data, determine the probability of the set behavior existing, and output the determined probability.
[0028] In some executable embodiments, the determining unit is further configured to:
[0029] Generate image captions for the valid video data frame by frame, trigger the natural language processing type to perform multimodal content analysis on the image captions of each frame, and determine whether the frame image has a set behavior and a first probability of the set behavior;
[0030] And, inputting the image caption and the video data frame into the multimodal large model, and determining whether the frame image has a set behavior and a second probability of the set behavior through the multimodal large model;
[0031] The average value of the first probability and the second probability is taken as the probability of the existence of the setting behavior of the valid video data.
[0032] In some executable embodiments, the identification unit is further used to:
[0033] For the sampled video data, the YOLO in the target detection algorithm is used to detect the objects in the video frames, count the number of objects in the video frames, and remove the video frames with the number of objects less than the set threshold in the sampled video data as valid video data.
[0034] In some executable embodiments, the identification unit is further used to:
[0035] The image information features of each video frame of the sampled video data are extracted through the backbone network Backbone, and the image information features are convoluted to obtain the local feature information of each video frame;
[0036] Normalize the local feature information of each video frame to reduce the covariate shift of the local feature information;
[0037] The normalized local feature information is input into the activation function to determine the object in the video frame.
[0038] In some executable embodiments, the object is a person; and the set behavior includes excessive physical contact behavior.
[0039] According to a third aspect of the present application, an electronic device is provided, including:
[0040] at least one processor; and
[0041] a memory communicatively connected to the at least one processor; wherein,
[0042] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method for identifying a set behavior based on image data.
[0043] According to a fourth aspect of the present application, a non-temporary computer-readable storage medium is provided. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the steps of the method for identifying set behaviors based on image data.
[0044] The method and device, electronic device, and storage medium for identifying set behaviors based on image data of the present application, by introducing a visual processing model and a natural language processing model, the visual processing model provides a stronger image understanding and guiding sentence processing capability, and reduces the misjudgment of specific behaviors in video frames through precise behavior analysis and natural language processing. The technical solution of the embodiment of the present application reduces the reliance on on-site supervisors, reduces unnecessary calculations by performing interval sampling of video frames and filtering invalid frames, and saves a large amount of processing resources.
[0045] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] By reading the detailed description below with reference to the accompanying drawings, the above and other purposes, features and advantages of the exemplary embodiments of the present application will become readily understood. In the accompanying drawings, several embodiments of the present application are shown in an exemplary and non-limiting manner, wherein:
[0047] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts.
[0048] Figure 1 A schematic diagram of a process flow of a method for identifying a set behavior based on image data according to an embodiment of the present application is shown;
[0049] Figure 2A schematic diagram of the structure of a device for identifying a set behavior based on image data according to an embodiment of the present application is shown;
[0050] Figure 3 A schematic diagram of the structure of an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0051] In order to make the purpose, features, and advantages of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.
[0052] Figure 1 FIG. 1 is a flow chart showing a method for identifying a set behavior based on image data according to an embodiment of the present application. Figure 1 As shown, the method for identifying a set behavior based on image data in the embodiment of the present application includes the following processing steps:
[0053] Step 101: Periodically sample the video frames acquired by the image acquisition unit to obtain sampled video data.
[0054] In an embodiment of the present application, the image acquisition unit may be a camera, such as a video surveillance camera, or a depth camera such as a binocular camera, etc., which can acquire images or videos of the environment to monitor a set area.
[0055] In the embodiments of the present application, some public security incidents are mainly monitored, mainly some fighting incidents, so that public security incidents can be discovered and curbed in time to ensure the stability of social security and protect people's life safety.
[0056] In the embodiment of the present application, if image recognition is performed on all video frames, this will result in a very large amount of data processing, resulting in a slow data processing speed. However, since public security incidents are low-probability events, they are generally difficult to occur, and only occasionally occur in some remote or hidden areas. Moreover, even if a public security incident occurs, its occurrence will be continuous. The embodiment of the present application is aimed at the acquisition of video frames by the camera, and a periodic sampling method is used to obtain video frames for a period of time. By performing image analysis on the periodically used video frames, it can be determined whether a public security incident has occurred.
[0057] In an embodiment of the present application, the sampling period can be 30 seconds, or it can be 1 minute, 2 minutes, etc. For example, 30 seconds of video frames are sampled every two minutes to analyze the sampled video frames to determine whether a corresponding public security incident has occurred.
[0058] Taking fighting as an example of a public security incident, generally speaking, fighting is a continuous action, and its duration can be as short as tens of seconds, or as long as several minutes or even longer. Even if it lasts only a dozen seconds, the corresponding video frame spans dozens or even hundreds of frames in the video stream. If fighting is detected for every frame in the video stream, it will consume huge computing resources. And to some extent, it is unnecessary, because fighting spans many frames in the video stream. Selecting a part of the video frames from the video stream to detect fighting by sampling with a fixed frame interval can represent the results of fighting detection for all video frames in the video stream to a certain extent. Based on the above analysis, it is selected to extract the corresponding video frames from the video stream by sampling with a fixed frame interval, and then combine the subsequent related algorithms to detect fighting.
[0059] Step 102 , performing object recognition on the sampled video data, and taking the video data containing the set objects in the sampled video data and the number of the set objects being greater than or equal to a set threshold as valid video data.
[0060] In the embodiment of the present application, the object to be identified in the video image frame is a person.
[0061] For the sampled video data, the Yolo in the target detection algorithm is used to detect the objects in the video frame, count the number of objects in the video frame, and remove the video frames with the number of objects less than the set threshold in the sampled video data as valid video data. The Yolo network structure can use Yolov8 to identify the objects and their number in the video image frame. Specifically, the image information features of each video frame of the sampled video data are extracted through the backbone network (Backbone), and the image information features are convoluted to obtain the local feature information of each video frame; the local feature information of each video frame is normalized to reduce the covariate shift of the local feature information; the normalized local feature information is input into the activation function to determine the object in the video frame.
[0062] Taking the setting behavior of fighting as an example, the participants of this behavior are people, and at least two people participate. Therefore, the prerequisite for the existence of fighting behavior in a video image frame is that there are at least two people in the video image frame. If the number of people in a video frame is less than 2, then there is generally no fighting behavior in the video frame. Taking this into account, for any video frame, first use yolov8 in the target detection algorithm to detect the people in the video frame, and then count the number of people people_num. If the value of people_num is less than 2, it means that there is generally no fighting behavior in the video frame, and such video frames are defined as invalid video frames, and no subsequent judgment is required. Choose to discard the video frame and no longer input it into the visual large model and the multimodal large model for multimodal analysis, which can save computing resources and reduce unnecessary algorithm time. If the value of people_num is greater than or equal to 2, it means that there may be fighting behavior in the video frame, and such video frames are defined as valid video frames. In this case, choose to input the video frame into the visual large model and the multimodal large model for subsequent analysis and judgment.
[0063] Step 103, selecting a corresponding visual processing model and a natural language processing type; generating a prompt statement for multimodal content analysis for the visual processing model.
[0064] In the embodiment of the present application, the visual processing model can be selected from at least one of the following three categories:
[0065] Textually Prompted Models: These include contrastive, generative, hybrid, and conversational models. As an example, CLIP (Contrastive Language-Image Pre-training) is a typical model. CLIP is a multimodal model based on contrastive learning. It pre-trains on large-scale image-text datasets to learn the matching relationship between images and texts. It encodes images and texts into the same vector space, making similar images and texts closer in space, thereby achieving cross-modal semantic understanding and retrieval.
[0066] Visually Prompted Models: These include SAM (Segment Anything Model) and SegGPT, which use visual prompts to perform tasks such as image segmentation. SAM segments specific objects from images through user prompts (such as clicks, picture frames, masks, text, etc.). SAM has the ability of zero-sample generalization, that is, it can segment visual objects on the image even if these objects have not appeared in the training set. The SAM model consists of an image encoder, a prompt encoder, and a mask decoder, and can quickly (about 50 milliseconds) predict masks based on prompts in the browser.
[0067] Heterogeneous Modalities-Based Models: These include ImageBind and Valley, which are designed to handle multiple types of input data and achieve cross-modal learning. As an example, Stable Diffusion is a typical model that generates high-quality images from text descriptions through text embedding, latent space sampling, step-by-step denoising generation of the U-Net network, and image decoding of the VAE decoder.
[0068] Before using the visual big model to describe the acquired video frames, in order to enable the visual big model to focus on the behavior of the people in the image, it is necessary to write prompts for the visual big model. The prompt content of the visual big model can be written as: "Analyze and describe the specific actions of the people in the image, and try to explain the possible purpose or emotional state behind their behavior."
[0069] Step 104: Use the visual processing model to perform text description on the valid video data to generate image subtitles for the valid video data.
[0070] Then, the visual big model is used to describe the acquired video frames and generate "image subtitles", which are the descriptions of the image by the visual big model. After obtaining the image subtitles generated by the visual big model, the image subtitles are input into the natural language processing big model, and the powerful natural language processing capabilities of the natural language processing big model are used to analyze the above image subtitles. In order for the natural language processing big model to better analyze whether there is fighting behavior from the image subtitles, it is necessary to write a prompt for the natural language processing big model in advance. The content of the prompt is as follows:
[0071] "You are a professional image analysis expert. Your task is to determine whether there is fighting in the image based on the image captions provided, and give the probability of such behavior. Please read the following image captions carefully and answer the following questions:
[0072] 1. Image captions: [insert image captions here]
[0073] 2. Question:
[0074] Please give the probability P1 (0%-100%) that there is fighting in the image.
[0075] Please answer the question explicitly using a percentage."
[0076] Step 105, input the image subtitles into a corresponding processing model to determine whether a set behavior exists in the valid video data, determine the probability of the set behavior existing, and output the determined probability.
[0077] In an embodiment of the present application, image subtitles are generated frame by frame for valid video data, and the natural language processing type is triggered to perform multimodal content analysis on the image subtitles of each frame to determine the probability of the existence of a set behavior in the frame image; the average value of the probabilities of all frame images is taken as the probability of the existence of the set behavior in the valid video data.
[0078] For valid video frames, after determining that the number of people therein is not less than 2, the video frame is selected to be input into the visual large model, and the corresponding image subtitles are generated, and input into the multimodal large model together with the image subtitles, and the powerful cross-modal processing capabilities of the multimodal large model are used to determine whether there is a set behavior in the video frame, and the set behavior includes excessive physical contact, such as fighting. The embodiment of the present application outputs the probability P2 of the existence of fighting through the multimodal large model. In order to make the focus of the multimodal large model more targeted and output the desired results, it is also necessary to write a prompt for the multimodal large model. The content of the prompt of the multimodal large model is as follows:
[0079] “You are a professional image and text analysis expert, and your task is to extract relevant information from the provided images and their captions. Please carefully look at the following images and their captions and answer the following questions:
[0080] 1. Image: [insert video frame here]
[0081] 2. Image subtitles: [Insert the image subtitles corresponding to the video frame here]
[0082] 3. Question:
[0083] -Please give the probability P2 (0%-100%) that there is fighting in the image.
[0084] Please answer the question explicitly using a percentage."
[0085] For any video frame, the corresponding probabilities P1 and P2 are weighted and summed, and recorded as probability P. For example, the weighted arithmetic mean of probabilities P1 and P2 is calculated as the basis for judging whether the set behavior occurs in the video frame. At this time, P = P1 / 2+P2 / 2, which is the probability that a fight occurs in the video frame. At this time, the weights of the two probabilities are both 1. If the value of probability P is greater than the set threshold value such as 60%, it is determined that a fight occurs in the video frame. Otherwise, it is determined that there is no fight in the video frame. The probability P in the video frame can also be the geometric mean of probabilities P1 and P2, as the probability that a fight occurs in the video frame.
[0086] Figure 2 The schematic diagram of the structure of the device for identifying a set behavior based on image data in the embodiment of the present application is shown as follows: Figure 2 As shown, the setting behavior recognition device based on image data in the embodiment of the present application includes:
[0087] The sampling unit 20 is used to periodically sample the video frames collected by the image acquisition unit to obtain sampled video data;
[0088] The identification unit 21 is used to perform object identification on the sampled video data, and to take the video data containing the set objects in the sampled video data, and the number of the set objects is greater than or equal to a set threshold, as the valid video data;
[0089] A selection unit 22, used for selecting a corresponding visual processing model and a natural language processing type;
[0090] A first generating unit 23, configured to generate a prompt statement for multimodal content analysis for the visual processing model;
[0091] A second generating unit 24 is used to perform text description on the effective video data by using the visual processing model to generate image subtitles for the effective video data;
[0092] The determination unit 25 is used to input the image subtitles into the natural language processing type, trigger the natural language processing type to perform multimodal content analysis on the image subtitles to determine whether there is a set behavior in the valid video data, determine the probability of the set behavior, and output the determined probability.
[0093] In some executable embodiments, the determining unit 25 is further configured to:
[0094] Generate image captions for the valid video data frame by frame, trigger the natural language processing type to perform multimodal content analysis on the image captions of each frame, and determine whether the frame image has a set behavior and a first probability of the set behavior;
[0095] And, inputting the image caption and the video data frame into the multimodal large model, and determining whether the frame image has a set behavior and a second probability of the set behavior through the multimodal large model;
[0096] The average value of the first probability and the second probability is taken as the probability of the existence of the setting behavior of the valid video data.
[0097] In some executable embodiments, the identification unit 21 is further configured to:
[0098] For the sampled video data, the YOLO in the target detection algorithm is used to detect the objects in the video frames, count the number of objects in the video frames, and remove the video frames with the number of objects less than the set threshold in the sampled video data as valid video data.
[0099] In some executable embodiments, the identification unit 21 is further configured to:
[0100] The image information features of each video frame of the sampled video data are extracted through the backbone network Backbone, and the image information features are convoluted to obtain the local feature information of each video frame;
[0101] Normalize the local feature information of each video frame to reduce the covariate shift of the local feature information;
[0102] The normalized local feature information is input into the activation function to determine the object in the video frame.
[0103] In some executable embodiments, the object is a person; and the set behavior includes excessive physical contact behavior.
[0104] In an exemplary embodiment, the aforementioned units and the like may be implemented by one or more central processing units (CPU), graphics processing units (GPU), application specific integrated circuits (ASIC), DSPs, programmable logic devices (PLD), complex programmable logic devices (CPLD), field programmable gate arrays (FPGA), general-purpose processors, controllers, microcontrollers (MCU), microprocessors, or other electronic components.
[0105] Regarding the device in the above embodiment, the specific manner in which each module and unit performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0106] According to an embodiment of the present application, the present application also provides an electronic device and a readable storage medium.
[0107] Figure 3 8 is a schematic block diagram of an example electronic device 800 that can be used to implement an embodiment of the present application. Figure 3 As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0108] A number of components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a data processing transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0109] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as a set behavior recognition method based on image data. For example, in some embodiments, the set behavior recognition method based on image data may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the set behavior recognition method based on image data described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute the set behavior recognition method based on image data in any other appropriate manner (for example, by means of firmware).
[0110] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0111] The program code for implementing the method of the present application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, implements the functions / operations specified in the flow chart and / or block diagram. The program code can be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0112] In the context of the present application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0113] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0114] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0115] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0116] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution disclosed in this application can be achieved, and this document is not limited here.
[0117] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0118] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A method for identifying a set behavior based on image data, characterized in that: The method comprises: Periodically sampling the video frames acquired by the image acquisition unit to obtain sampled video data; Performing object recognition on the sampled video data, and taking the video data containing the set objects in the sampled video data, and the number of the set objects is greater than or equal to a set threshold, as valid video data; Selecting a corresponding visual processing model and a natural language processing model; generating a prompt sentence for multimodal content analysis for the visual processing model; Using the visual processing model to perform text description on the valid video data, and generate image captions for the valid video data; The image subtitles are input into a corresponding processing model to determine whether a set behavior exists in the valid video data, determine the probability of the set behavior existing, and output the determined probability.
2. The method for identifying a set behavior based on image data according to claim 1, characterized in that: The determining of the probability of the existence of the set behavior includes: Generate image captions for the valid video data frame by frame, trigger the natural language processing type to perform multimodal content analysis on the image captions of each frame, and determine whether the frame image has a set behavior and a first probability of the set behavior; And, inputting the image caption and the video data frame into the multimodal large model, and determining whether the frame image has a set behavior and a second probability of the set behavior through the multimodal large model; The average value of the first probability and the second probability is taken as the probability of the existence of the setting behavior of the valid video data.
3. The method for identifying a set behavior based on image data according to claim 1, characterized in that: The object recognition is performed on the sampled video data, and the video data containing the set objects in the sampled video data, and the number of the set objects is greater than or equal to the set threshold, is used as the valid video data, including: For the sampled video data, the YOLO in the target detection algorithm is used to detect the objects in the video frames, count the number of objects in the video frames, and remove the video frames with the number of objects less than the set threshold in the sampled video data as valid video data.
4. The method for identifying a set behavior based on image data according to claim 3, characterized in that: The object detection in the video frame using the YOLO in the target detection algorithm includes: The image information features of each video frame of the sampled video data are extracted through the backbone network Backbone, and the image information features are convoluted to obtain the local feature information of each video frame; Normalize the local feature information of each video frame to reduce the covariate shift of the local feature information; The normalized local feature information is input into the activation function to determine the object in the video frame.
5. The method for identifying a set behavior based on image data according to any one of claims 1 to 4, characterized in that: The object is a person; the set behavior includes excessive physical contact.
6. A device for identifying a set behavior based on image data, characterized in that: The device comprises: A sampling unit, used for periodically sampling the video frames acquired by the image acquisition unit to obtain sampled video data; The recognition unit is used to perform object recognition on the sampled video data, and to take the video data containing the set objects in the sampled video data, and the number of the set objects is greater than or equal to a set threshold, as the valid video data; A selection unit, used to select a corresponding visual processing model and a natural language processing model; A first generating unit, configured to generate a prompt statement for multimodal content analysis for the visual processing model; A second generating unit is used to perform text description on the effective video data by using the visual processing model to generate image subtitles for the effective video data; The determination unit is used to input the image subtitle into a corresponding processing model to determine whether a set behavior exists in the valid video data, determine the probability of the set behavior existing, and output the determined probability.
7. The device for identifying set behaviors based on image data according to claim 6, characterized in that: The determining unit is further configured to: Generate image captions for the valid video data frame by frame, trigger the natural language processing type to perform multimodal content analysis on the image captions of each frame, and determine whether the frame image has a set behavior and a first probability of the set behavior; And, inputting the image caption and the video data frame into the multimodal large model, and determining whether the frame image has a set behavior and a second probability of the set behavior through the multimodal large model; The average value of the first probability and the second probability is taken as the probability of the existence of the setting behavior of the valid video data.
8. The device for identifying set behaviors based on image data according to claim 6, characterized in that: The identification unit is further used for: For the sampled video data, the YOLO in the target detection algorithm is used to detect the objects in the video frames, count the number of objects in the video frames, and remove the video frames with the number of objects less than the set threshold in the sampled video data as valid video data.
9. The device for identifying a set behavior based on image data according to claim 8, characterized in that: The identification unit is further used for: The image information features of each video frame of the sampled video data are extracted through the backbone network Backbone, and the image information features are convoluted to obtain the local feature information of each video frame; Normalize the local feature information of each video frame to reduce the covariate shift of the local feature information; The normalized local feature information is input into the activation function to determine the object in the video frame.
10. The device for identifying set behaviors based on image data according to any one of claims 6 to 9, characterized in that: The object is a person; the set behavior includes excessive physical contact.
Citation Information
Patent Citations
Visual model-based large language model video time sequence positioning method and product
CN117851638A
Video processing method and device, electronic equipment and storage medium
CN117851639A
Airport boundary detection method and system based on multiple modes
CN119152445A
Learning to Personalize Vision-Language Models through Meta-Personalization
US20240419726A1
Method for video action recognition
WO2024071836A1