Information sending method and device, electronic equipment and computer readable medium
By performing cross-frame extraction processing on security videos and utilizing a distributed model set, accurate object intent information is generated, solving the problems of limited video features and inaccurate model output in existing object intent detection models, and realizing efficient object anomaly monitoring and alarm functions.
Patent Information
- Application Number
- CN202411344175.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-25
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-09-25
AI Technical Summary
Existing object intent detection models have limited video features when generating object intent information, resulting in inaccurate outputs and poor coordination between the outputs of different models in the model set.
By cross-sampling security videos, a sequence of frame images is generated. Then, using a multimodal distributed model set, including a master node and multiple child nodes, an object intent recognition information generation model is used. Combined with video, audio, and image sequences, accurate object intent information is generated, and the target object is labeled in the abnormal intent information table.
It enables accurate and efficient monitoring of anomalies in security videos, generates more precise information about the intent of objects, and can promptly trigger alarms for the target objects.
Smart Images

Figure CN119314078B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the technical field of computer technology, and specifically to an information sending method and device, an electronic device and a computer readable medium. BACKGROUND
[0002] At present, the current security awareness has been widely rooted in people's thinking. For object anomaly detection in security video, the commonly used way is: through an object intention detection model based on video form, to generate object intention information for security video.
[0003] However, when the above method is used to generate object intention information, the following technical problems often exist:
[0004] The video features learned by the object intention detection model are limited, resulting in inaccurate object intention information.
[0005] In the process of using the technical solution to solve the above technical problem one, the following technical problem often accompanies: the training of each model in the model set is also crucial. For the above technical problem, the conventional solution is generally: by training each model in the model set to generate a model set. However, the above conventional solution still has the following problem two: the output cooperation between each model in the obtained model set is not accurate enough, resulting in inaccurate subsequent output object intention information.
[0006] The above information disclosed in the background section is only intended to enhance the understanding of the background of the present inventive concept, and therefore, it can include information that does not form the prior art known to those of ordinary skill in the art in the country. SUMMARY
[0007] The summary section of the present disclosure is used to introduce the concept in a brief form, which will be described in detail in the specific embodiments section. The summary section of the present disclosure is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0008] Some embodiments of the present disclosure propose an information sending method and device, an electronic device and a computer readable medium to solve one or more of the technical problems mentioned in the background section.
[0009] In a first aspect, some embodiments of the present disclosure provide an information sending method, including: obtaining a security video for a target monitoring area; performing cross-frame extraction processing on the security video to generate at least one frame image sequence; generating a video range for a target object according to the at least one frame image sequence; performing image sequence segmentation on each frame image sequence in the at least one frame image sequence according to the video range to generate segmented image sequences, to obtain at least one segmented image sequence; obtaining a video voice in the security video within the video range and a distributed-based model set, wherein the model set includes: a multi-modal-based master node object intention recognition information generation model, and a plurality of multi-modal-based sub-node object intention recognition information generation models; generating object intention information for the target object according to the model set, the at least one segmented image sequence, and the video voice; in response to determining that the object intention information is in an abnormal intention information table, performing target object labeling on a security sub-video corresponding to the video range to generate a labeled video; and sending the labeled video and the object intention information to an alarm terminal.
[0010] In a second aspect, some embodiments of the present disclosure provide an information sending device, including: a first obtaining unit configured to obtain a security video for a target monitoring area; a cross-frame extraction processing unit configured to perform cross-frame extraction processing on the security video to generate at least one frame image sequence; a first generating unit configured to generate a video range for a target object according to the at least one frame image sequence; a segmentation unit configured to perform image sequence segmentation on each frame image sequence in the at least one frame image sequence according to the video range to generate segmented image sequences, to obtain at least one segmented image sequence; a second obtaining unit configured to obtain a video voice in the security video within the video range and a distributed-based model set, wherein the model set includes: a multi-modal-based master node object intention recognition information generation model, and a plurality of multi-modal-based sub-node object intention recognition information generation models; a second generating unit configured to generate object intention information for the target object according to the model set, the at least one segmented image sequence, and the video voice; a labeling unit configured to, in response to determining that the object intention information is in an abnormal intention information table, perform target object labeling on a security sub-video corresponding to the video range to generate a labeled video; and a sending unit configured to send the labeled video and the object intention information to an alarm terminal.
[0011] In a third aspect, some embodiments of the present disclosure provide an electronic device, comprising: one or more processors; a storage device having stored thereon one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the method as described in any implementation manner of the first aspect.
[0012] In a fourth aspect, some embodiments of the present disclosure provide a computer readable medium having stored thereon a computer program, wherein the program, when executed by a processor, implements the method as described in any implementation manner of the first aspect.
[0013] The above various embodiments of the present disclosure have the following beneficial effects: the information sending method of some embodiments of the present disclosure can accurately and efficiently supervise the object anomaly of the security video by generating object intention information. Specifically, the reason why the related object anomaly supervision is not accurate is that the video features learned by the object intention detection model are limited, resulting in inaccurate generated object intention information. Based on this, the information sending method of some embodiments of the present disclosure first acquires the security video for the target monitoring area as the data basis to facilitate subsequent identification of the object intention information corresponding to the target object. Then, the security video is subjected to cross-frame extraction processing to generate at least one frame image sequence, which can obtain at least one frame image sequence with more characteristic details and different characteristic details corresponding to each frame image sequence, and can subsequently accurately generate object intention information. Next, according to the at least one frame image sequence, a video range for the target object is generated to facilitate subsequent generation of at least one segmentation image sequence to determine at least one segmentation image sequence associated with the target object and only containing valid features. Furthermore, according to the video range, image sequence segmentation is performed on each frame image sequence in the at least one frame image sequence to generate a segmentation image sequence, obtaining at least one segmentation image sequence to facilitate subsequent generation of accurate object intention information. Then, the video voice in the security video in the video range and the distributed-based model set are acquired, wherein the model set includes a multi-modal main node object intention recognition information generation model and a plurality of multi-modal sub-node object intention recognition information generation models. Here, by using the distributed-based model set, subsequent object intention information can be accurately generated with less computational complexity. In addition, the video voice as a modal data assisting the subsequent generation of object intention information facilitates the subsequent extraction of more feature information about the target object to generate more accurate object intention information. Furthermore, according to the model set, the at least one segmentation image sequence, and the video voice, the object intention information for the target object can be accurately generated. Secondly, in response to determining that the object intention information is in the abnormal intention information table, the target object is labeled in the security sub-video corresponding to the video range to generate a labeled video to facilitate subsequent implementation of corresponding alarm operations for the target object. Finally, the labeled video and the object intention information are sent to the alarm terminal. In summary, by cross-frame extraction processing of the security video and using the distributed-based model set, the object intention information corresponding to the target object can be accurately and efficiently generated. BRIEF DESCRIPTION OF DRAWINGS
[0014] The above and other features, aspects and advantages of various embodiments of the present disclosure will become more apparent with reference to the following specific embodiments described exemplary and illustrated in the accompanying drawings. Identical or similar components shown throughout the figures are identified by identical or similar reference numerals. It is to be understood that the drawings are schematic and elements and features are not necessarily to scale.
[0015] Figure 1 is a flowchart of some embodiments of the information sending method according to the present disclosure;
[0016] Figure 2 is a structural schematic diagram of some embodiments of the information sending apparatus according to the present disclosure;
[0017] Figure 3 is a structural schematic diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0018] Embodiments of the present disclosure will be described in detail with reference to the drawings, wherein the same or similar components are denoted by the same or similar reference numerals. It is to be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure will be more thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are only for illustrative purposes and should not be construed as limiting the scope of protection of the present disclosure.
[0019] It is further noted that, for the sake of brevity, only some of the relevant parts of the drawings are shown in the figures. The embodiments and features of the present disclosure can be combined with each other as long as they do not conflict with each other.
[0020] It is to be noted that the terms “first”, “second”, and the like in the present disclosure are used only to distinguish different devices, modules or units, and do not imply the order or sequence of the functions performed by the devices, modules or units.
[0021] It is to be noted that the terms “one”, “multiple” in the present disclosure are illustrative and not restrictive, and those skilled in the art should understand that “one” or “multiple” should be understood as “one or more” unless otherwise explicitly stated in the context.
[0022] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0023] The present disclosure will be described in detail with reference to the accompanying drawings and embodiments.
[0024] Reference Figure 1FIG. 10 shows a flow 100 of some embodiments of the information sending method according to the present disclosure. The information sending method comprises the following steps:
[0025] Step 101, obtaining a security video for a target monitoring area.
[0026] In some embodiments, the subject performing the above information sending method can obtain the security video for the target monitoring area through wired connection or wireless connection. The target monitoring area can be an area for security monitoring. The security video can be a regionally captured video.
[0027] Step 102, performing cross-frame extraction processing on the security video to generate at least one frame image sequence.
[0028] In some embodiments, the subject performing the above information sending method can perform cross-frame extraction processing on the security video to generate at least one frame image sequence.
[0029] In some optional implementations of some embodiments, the cross-frame extraction processing on the security video to generate at least one frame image sequence can comprise the following steps:
[0030] First, obtaining a pre-set frame extraction duration sequence. The frame extraction durations in the frame extraction duration sequence decrease from left to right. For example, the frame extraction duration sequence is [10, 9, 7, 5, 3, 2].
[0031] Second, performing video division on the security video according to the maximum frame extraction duration in the frame extraction duration sequence to generate a security sub-video sequence. Each video duration corresponding to a security sub-video is the same as the maximum frame extraction duration.
[0032] As an example, the subject performing the above information sending method can perform video division on the security video according to the maximum frame extraction duration as the video division duration to generate a security sub-video sequence.
[0033] Third, for a target frame extraction duration in the frame extraction duration sequence, the following frame extraction steps are performed:
[0034] Sub-step 1, determining a previous frame extraction duration corresponding to the target frame extraction duration. The previous frame extraction duration can be the frame extraction duration adjacent to the left of the target frame extraction duration in the frame extraction duration sequence.
[0035] Sub-step 2, determining a plurality of frame image sub-sequences corresponding to the previous frame extraction duration as a plurality of first frame image sub-sequences. The plurality of first frame image sub-sequences can be a plurality of image sub-sequences generated based on the frame extraction steps.
[0036] Sub-step 3, extract a target number of frame images from each of the security sub-videos in the security sub-video sequence to generate a frame image sub-sequence as a second frame image sub-sequence, to obtain a plurality of second frame image sub-sequences. Wherein, the target number is the same as the target frame extraction duration value. Wherein, there is a time one-to-one correspondence relationship between the second frame image sub-sequence in the plurality of second frame image sub-sequences and the first frame image sub-sequence in the plurality of first frame image sub-sequences. The frame image coincidence degree between each frame image included in each second frame image sub-sequence and each frame image included in the corresponding first frame image sub-sequence is greater than a preset coincidence degree. For example, the preset coincidence degree can be 30%.
[0037] Sub-step 4, in response to determining that the target frame extraction duration is the minimum frame extraction duration, determining the plurality of second frame image sub-sequences as the at least one frame image sequence. Wherein, the minimum frame extraction duration can be the frame extraction duration with the smallest duration in the frame extraction duration sequence.
[0038] Fourth step, in response to determining that the target frame extraction duration is not the minimum frame extraction duration, taking the next frame extraction duration corresponding to the target frame extraction duration as the target frame extraction duration, and continuing to execute the frame extraction step.
[0039] Step 103, generating a video range for the target object according to the at least one frame image sequence.
[0040] In some embodiments, the execution subject can generate a video range for the target object according to the at least one frame image sequence. Wherein, the video range can be a picture range in the security video that has an association relationship with the target object. The video range can be represented by a video time range.
[0041] In some optional implementations of some embodiments, the generating a video range for the target object according to the at least one frame image sequence comprises:
[0042] First step, for each frame image sequence in the at least one frame image sequence, executing the following first generation step:
[0043] Sub-step 1, input each frame image in the frame image sequence into a target detection model for the target object to generate target detection information, to obtain a target detection information sequence. Wherein, the target detection information can be detection information of whether there is target object corresponding object information in the frame image. In practice, the target detection model can be a YOLO model.
[0044] Sub-step 2, remove the frame images in the frame image sequence that are represented by the corresponding target detection information as not having the target object, to obtain a removed frame image sequence.
[0045] Sub-step 3, the above-mentioned removed frame image sequence is supplemented with front and rear continuous frame images to generate a supplemented frame image sequence. Among them, the supplemented two image sequences in the supplemented frame image are respectively the image sequence with time continuity corresponding to the previous direction of the removed frame image and the image sequence with time continuity corresponding to the next direction of the removed frame image.
[0046] Sub-step 4, determine the initial video range corresponding to the above-mentioned supplemented frame image sequence.
[0047] As an example, the above-mentioned execution subject can determine the video time range corresponding to the supplemented frame image sequence as the initial video range.
[0048] Second step, take the union of the obtained at least one initial video range to generate a union video range as the above-mentioned video range. Among them, the union processing can be video time range union processing.
[0049] Step 104, according to the above-mentioned video range, the image sequence segmentation is performed on each frame image sequence in the above-mentioned at least one frame image sequence to generate a segmented image sequence, and at least one segmented image sequence is obtained.
[0050] In some embodiments, the above-mentioned execution subject can perform image sequence segmentation on each frame image sequence in the above-mentioned at least one frame image sequence according to the above-mentioned video range to generate a segmented image sequence, and at least one segmented image sequence is obtained. Among them, the video time range corresponding to the segmented image sequence is the same as the video time range corresponding to the above-mentioned video range.
[0051] As an example, the above-mentioned execution subject can perform image sequence segmentation on each frame image sequence in the above-mentioned at least one frame image sequence according to the video time range corresponding to the video range to generate a segmented image sequence, and at least one segmented image sequence is obtained.
[0052] Step 105, obtain the video voice in the above-mentioned security video and the distributed-based model set within the above-mentioned video range.
[0053] In some embodiments, the execution subject can obtain a video voice in the security video within the video range and a distributed model set. The video voice can be a voice occurring in the security video within the video range. The model set includes a multi-modal main node object intent recognition information generation model and a plurality of multi-modal sub-node object intent recognition information generation models. The main node object intent recognition information generation model can be a main neural network model for generating object intent recognition information based on multi-modal input data. The plurality of multi-modal sub-node object intent recognition information generation models can be a plurality of slave neural network models for generating object intent recognition information based on multi-modal input data. The accuracy of the object intent recognition information generated by the main node object intent recognition information generation model is higher than that of the object intent recognition information generated by the sub-node object intent recognition information generation model. The computational complexity of the object intent recognition information generated by the main node object intent recognition information generation model is much higher than that of the object intent recognition information generated by the sub-node object intent recognition information generation model. In practice, the main node object intent recognition information generation model can be a classification model that determines the actual intent recognition information from a plurality of intent recognition information based on an input data set. In practice, the main node object intent recognition information generation model can be a Transformer model.
[0054] Step 106, generating object intent information for the target object according to the model set, the at least one segmented image sequence, and the video voice.
[0055] In some embodiments, the execution subject can generate object intent information for the target object according to the model set, the at least one segmented image sequence, and the video voice. The object intent information can be the intent information of the behavior intent of the target object in the security video. For example, the object intent information can be the intent information of the illegal intrusion of the target object.
[0056] In some optional implementations of some embodiments, the generation of object intent information for the target object according to the model set, the at least one segmented image sequence, and the video voice can include the following steps:
[0057] First, for each segmented image sequence in the at least one segmented image sequence, the following second generation step is performed:
[0058] Sub-step 1, generating an image feature information sequence for the segmented image sequence.
[0059] As an example, the execution subject can input each segmented image in the segmented image sequence to the feature extraction model to generate image feature information, to obtain an image feature information sequence. The image feature information can represent the image feature semantic content corresponding to the segmented image. The feature extraction model can be a convolutional neural network model connected in multiple layers.
[0060] Sub-step 2, according to the image set corresponding to each image feature information in the image feature information sequence, the video voice is divided into sub-voice to generate a sub-voice sequence. The sub-voice in the sub-voice sequence has a time corresponding relationship with the image feature information in the image feature information sequence. That is, the sub-voice corresponding voice content and time length has a one-to-one corresponding relationship with the image feature information corresponding content and time length.
[0061] As an example, for each image feature information in the image feature information sequence, first, determine the segmented image sub-sequence corresponding to the image feature information. Then, determine the image time period corresponding to the segmented image sub-sequence. Finally, cut out the voice corresponding to the image time period from the video voice as the sub-voice corresponding to the image feature information.
[0062] Sub-step 3, for each sub-voice in the sub-voice sequence, the following third generation step is performed:
[0063] First sub-step, audio pre-processing is performed on the sub-voice to generate processed audio. The audio pre-processing can be audio noise reduction processing.
[0064] Second sub-step, fast Fourier transform is performed on the processed audio to generate a frequency domain signal.
[0065] Third sub-step, using a mel filter bank, a processed frequency domain signal corresponding to the frequency domain signal is generated.
[0066] Fourth sub-step, logarithmic processing is performed on the processed frequency domain signal to generate a linearly enhanced signal.
[0067] Fifth sub-step, discrete cosine transformation is performed on the linearly enhanced signal to generate a transformed audio feature information.
[0068] Sub-step 4, according to the number of segmented images corresponding to the segmented image sequence, select the corresponding sub-node object intention recognition information generation model from the plurality of sub-node object intention recognition information generation models as the target sub-node object intention recognition information generation model, wherein the output weight proportion corresponding to the target sub-node object intention recognition information generation model has a corresponding relationship with the number of segmented images. The number of segmented images can be the number of images of each segmented image included in the segmented image sequence. Each of the plurality of sub-node object intention recognition information generation models has a corresponding output weight proportion. Specifically, according to the complexity of the model structure of the corresponding sub-node object intention recognition information generation model (i.e. the number of network units included), the corresponding output weight proportion is different. The higher the complexity of the model structure, the higher the corresponding output weight proportion, indicating that the output of the corresponding model is more accurate and more important.
[0069] Sub-step 5, according to the image feature information sequence and the obtained changed audio feature information sequence, using the target sub-node object intention recognition information generation model, generate the initial object intention information. Wherein, the initial object intention information can represent the preliminary judgment of the target object in the security video.
[0070] As an example, the execution subject can directly input the image feature information sequence and the obtained changed audio feature information sequence into the target sub-node object intention recognition information generation model to generate the initial object intention information.
[0071] Second step, according to the obtained at least one initial object intention information, generate the object intention information.
[0072] As an example, the execution subject can perform intention information summarization on at least one initial object intention information to generate a summarized intention information as the object intention information.
[0073] In some optional implementations of some embodiments, the target sub-node object intention recognition information generation model includes a plurality of feature extraction units, a plurality of feature fusion units and an identification information generation layer. Wherein, the feature extraction unit can be a neural network unit for feature depth extraction. The feature fusion unit can be a unit for multi-modal feature information fusion. The identification information generation layer can be a network layer for generating identification information. In practice, the feature fusion unit can be an attention mechanism model. The identification information generation layer can be a plurality of fully connected layers in series.
[0074] It should be noted that according to the number of feature extraction units and feature fusion units included in each sub-node object intention recognition information generation model, the corresponding output weight proportion is set for the corresponding model.
[0075] Optionally, the generating, according to the image feature information sequence and the obtained changed audio feature information sequence, the initial object intention information by using the target sub-node object intention recognition information generation model, can include the following steps:
[0076] In the first step, a time step sequence corresponding to the image feature information sequence is determined. The image feature information in the image feature information sequence has a time correspondence with the time steps in the time step sequence.
[0077] In the second step, a target time step is selected from the time step sequence.
[0078] As an example, the execution subject can randomly select the target time step from the time step sequence.
[0079] In the third step, for the target time step, the following third generation step is performed:
[0080] Sub-step 1: obtaining the image feature information in the image feature information sequence that has a time correspondence with the target time step as target image feature information, and obtaining the changed audio feature information in the changed audio feature information sequence that has a time correspondence with the target time step as target changed audio feature information.
[0081] Sub-step 2: determining the feature extraction unit corresponding to the target time step as a target feature extraction unit, so as to determine the feature fusion unit corresponding to the target time step as a target feature fusion unit.
[0082] Sub-step 3: determining the previous time step corresponding to the target time step. The previous time step can be the time step in the time step sequence that is located at the left adjacent position of the target time step.
[0083] Sub-step 4: determining the fusion feature information corresponding to the previous time step as target fusion feature information. The fusion feature information can be multi-modal feature fusion information representing each input data before the previous time step.
[0084] Sub-step 5: inputting the target changed audio feature information into the audio feature weight information generation model included in the target feature extraction unit to generate audio feature weight information. The audio feature weight information generation model can be a neural network model for generating audio feature weight information. The audio feature weight information can represent the importance and richness of each audio feature included in the target changed audio feature information. The higher the corresponding numerical value of the audio feature weight information, the higher the importance and richness of each audio feature included. In practice, the audio feature weight information generation model can be a multi-layer convolutional neural network model.
[0085] Sub-step 6, multiplying the audio feature weight information and the target changed audio feature information to generate multiplied audio feature information.
[0086] Sub-step 7, inputting the multiplied audio feature information and the target image feature information into the attention mechanism model included in the target feature extraction unit to generate first attention weight information for the multiplied audio feature information and second attention weight information for the target image feature information. The first attention weight information can represent the importance of the multiplied audio feature information. The second attention weight information can represent the importance of the target image feature information.
[0087] Sub-step 8, inputting the first attention weight information, the second attention weight information, the multiplied audio feature information and the target image feature information into the target feature fusion unit to generate fusion feature information.
[0088] Sub-step 9, performing feature information splicing on the fusion feature information and the target fusion feature information to generate spliced feature information.
[0089] Sub-step 10, in response to determining that the target time step is the latest time step in the time step sequence, inputting the spliced feature information corresponding to the latest time step into the recognition information generation layer to generate initial object intent information.
[0090] Fourth step, in response to determining that the target time step is not the latest time step in the time step sequence, taking the next time step corresponding to the target time step as the target time step, and continuing to execute the third generation step.
[0091] In view of the problems of the above conventional solutions, the output coordination between the models in the model set obtained in the above technical problem two is not accurate enough, resulting in the object intent information output subsequently is not accurate enough. Combined with the advantages / technical status owned, it can be decided to adopt the following solutions.
[0092] In some optional implementations of some embodiments, the distributed-based model set is trained by the following steps:
[0093] First step, obtaining a training data set, wherein the training data includes: target video and intent recognition label.
[0094] Second step, selecting target training data from the training data set.
[0095] Third step, for the target training data, the following training steps are executed:
[0096] Sub-step 1, determine the target video and intent recognition label corresponding to the target training data.
[0097] Sub-step 2, generate at least one target segmentation image sequence and target video speech for the target video.
[0098] Sub-step 3, generate target initial object intent information corresponding to each initial sub-node object intent recognition information generation model in the plurality of initial sub-node object intent recognition information generation models according to the at least one target segmentation image sequence and the target video speech, to obtain at least one target initial object intent information.
[0099] Sub-step 4, generate target candidate object intent information for the at least one target initial object intent information using the initial main node object intent recognition information generation model.
[0100] Sub-step 5, determine the difference information between the target candidate object intent information and the intent recognition label.
[0101] Sub-step 6, generate first loss information for the difference information.
[0102] Sub-step 7, determine the initial sub-node object intent recognition information generation model set that is different from the corresponding output and the intent recognition label.
[0103] Sub-step 8, for each initial sub-node object intent recognition information generation model in the initial sub-node object intent recognition information generation model set, perform the following information generation steps:
[0104] First sub-step, determine the target initial object intent information corresponding to the initial sub-node object intent recognition information generation model as first initial object intent information.
[0105] Second sub-step, determine the second loss information between the first initial object intent information and the intent recognition label.
[0106] As an example, the execution subject can use a cross-entropy loss function to determine the second loss information between the first initial object intent information and the intent recognition label.
[0107] Sub-step 9, according to the obtained second loss information set, update the model parameters of each model in the initial sub-node object intent recognition information generation model set to generate a sub-node object intent recognition information generation model set.
[0108] In response to determining that the first loss information is less than the first predetermined value, the obtained child node object intention recognition information generation model set, the remaining child node object intention recognition information generation model set, and the initial main node object intention recognition information generation model are combined to generate a model set. The remaining child node object intention recognition information generation model set can be a model set obtained by removing the initial child node object intention recognition information generation model set from the plurality of initial child node object intention recognition information generation models.
[0109] In response to determining that the first loss information is greater than or equal to the first predetermined value, the initial main node object intention recognition information generation model is updated according to the first loss information to obtain an updated main node object intention recognition information generation model.
[0110] In the fifth step, the target training data is removed from the training data set to obtain a removed training data set.
[0111] In the sixth step, training data is reselected from the removed training data set as candidate training data.
[0112] In the seventh step, the updated main node object intention recognition information generation model, the child node object intention recognition information generation model set, and the remaining child node object intention recognition information generation model set are combined to generate an initial combined model set.
[0113] In the eighth step, the initial combined model set is used as an initial model set, and the candidate training data is used as target training data, and the training steps are continued. The initial model set includes the plurality of initial child node object intention recognition information generation models and the initial main node object intention recognition information generation model.
[0114] The content in the above "in some optional implementations of some embodiments" is an application point of the present disclosure, which solves the technical problem of "the output coordination between the models in the obtained model set is not accurate enough, resulting in less accurate output object intention information" mentioned in the background. Based on this, the present disclosure can obtain more accurate output object intention information by sequentially training the main node object intention recognition information generation model and the plurality of child node object intention recognition information generation models, and more accurately updating the plurality of child node object intention recognition information generation models under the output limitation of the main node corresponding model.
[0115] In some optional implementations of some embodiments, the object intention information can be generated according to the obtained at least one initial object intention information, which can include the following steps:
[0116] A first step is to determine at least one output weight proportion corresponding to the at least one initial object intent information.
[0117] A second step is to classify each initial object intent information in the at least one initial object intent information to generate a candidate object intent information set. Each candidate object intent information has a corresponding object intent category. The object intent category can be a category to which the object intent belongs.
[0118] A third step is to determine each candidate object intent information in the candidate object intent information set to perform the following fourth generation step:
[0119] Sub-step 1, determine the output weight proportion set corresponding to the candidate object intent information.
[0120] As an example, first, the execution subject can determine the sub-node object intent recognition information generation model corresponding to the candidate object intent information as the target sub-node object intent recognition information generation model. Then, determine the output weight proportion corresponding to the target sub-node object intent recognition information generation model.
[0121] Sub-step 2, sum each output weight proportion in the output weight proportion set to generate an output weight proportion sum.
[0122] A fourth step is to determine whether there is a candidate object intent information in the candidate object intent information set whose corresponding output weight proportion sum is greater than a first weight proportion. The first weight proportion can be a pre-set proportion value. For example, the first weight proportion can be 0.8.
[0123] A fifth step is to determine the candidate object intent information whose corresponding output weight proportion sum is greater than the first weight proportion as the object intent information in response to determining that there is.
[0124] Optionally, the object intent information is generated according to the at least one initial object intent information obtained, and further includes:
[0125] A first step is to determine at least one candidate object intent information whose corresponding output weight proportion sum is greater than a second weight proportion from the candidate object intent information set as at least one target candidate object intent information in response to determining that there is no. The second weight proportion is greater than the first weight proportion.
[0126] A second step is to select the object intent information from the at least one target candidate object intent information by using the main node object intent recognition information generation model.
[0127] As an example, first, the execution subject can perform deduplication fusion processing on the at least one frame image sequence to generate a frame image set. Then, the frame image set and the at least one target candidate object intention information are input into the master node object intention recognition information generation model to generate object intention information.
[0128] Thirdly, at least one sub-node object intention recognition information generation model corresponding to output intention information that is not the object intention information is determined from the plurality of sub-node object intention recognition information generation models.
[0129] Fourthly, model retraining is performed on the at least one sub-node object intention recognition information generation model.
[0130] Step 107, in response to determining that the object intention information is in the abnormal intention information table, performing target object labeling on the security sub-video corresponding to the video range to generate a labeled video.
[0131] In some embodiments, in response to determining that the object intention information is in the abnormal intention information table, the execution subject can perform target object labeling on the security sub-video corresponding to the video range to generate a labeled video. The abnormal intention information table can be a table including various abnormal intention information for objects. The abnormal intention information represents intention information that is dangerous for security.
[0132] As an example, the execution subject can use a bounding box labeling model to perform target object labeling on the security sub-video corresponding to the video range to generate a labeled video. The bounding box labeling model can be a sub-model of a target detection model.
[0133] Step 108, sending the labeled video and the object intention information to an alarm terminal.
[0134] In some embodiments, the execution subject can send the labeled video and the object intention information to an alarm terminal. The alarm terminal can be a terminal that performs alarm processing for security risks.
[0135] The above various embodiments of the present disclosure have the following beneficial effects: the information sending method of some embodiments of the present disclosure can accurately and efficiently supervise the object anomaly of the security video by generating object intention information. Specifically, the reason why the related object anomaly supervision is not accurate is that the video features learned by the object intention detection model are limited, resulting in inaccurate generated object intention information. Based on this, the information sending method of some embodiments of the present disclosure first acquires the security video for the target monitoring area as the data basis to facilitate subsequent identification of the object intention information corresponding to the target object. Then, the security video is subjected to cross-frame extraction processing to generate at least one frame image sequence, which can obtain at least one frame image sequence with more characteristic details and different characteristic details corresponding to each frame image sequence, and can subsequently accurately generate object intention information. Next, according to the at least one frame image sequence, a video range for the target object is generated to facilitate subsequent generation of at least one segmentation image sequence to determine at least one segmentation image sequence associated with the target object and containing only valid features. Furthermore, according to the video range, image sequence segmentation is performed on each frame image sequence in the at least one frame image sequence to generate a segmentation image sequence, obtaining at least one segmentation image sequence to facilitate subsequent generation of accurate object intention information. Then, the video voice in the security video within the video range and the distributed model set are acquired, wherein the model set includes a multi-modal master node object intention recognition information generation model and a plurality of multi-modal sub-node object intention recognition information generation models. Here, by using the distributed model set, subsequent object intention information can be accurately generated with less computational amount. In addition, the video voice is a modal data assisting the subsequent generation of object intention information, which facilitates the subsequent extraction of more feature information about the target object to generate more accurate object intention information. Furthermore, according to the model set, the at least one segmentation image sequence and the video voice, the object intention information for the target object can be accurately generated. Secondly, in response to determining that the object intention information is in the abnormal intention information table, the target object is labeled in the security sub-video corresponding to the video range to generate a labeled video to facilitate subsequent implementation of corresponding alarm operations for the target object. Finally, the labeled video and the object intention information are sent to the alarm terminal. In summary, by cross-frame extraction processing of the security video and using the distributed model set, the object intention information corresponding to the target object can be accurately and efficiently generated.
[0136] Further reference Figure 2 As an implementation of the method shown in the above figures, the present disclosure provides some embodiments of an information sending device, which device embodiments are used to implement the method of Figure 1The information sending apparatus can be applied to various electronic devices corresponding to the method embodiments shown.
[0137] As shown in Figure 2 An information sending apparatus 200 includes a first obtaining unit 201, a cross-frame processing unit 202, a first generating unit 203, a dividing unit 204, a second obtaining unit 205, a second generating unit 206, a labeling unit 207, and a sending unit 208. The first obtaining unit 201 is configured to obtain a security video for a target monitoring area. The cross-frame processing unit 202 is configured to perform cross-frame processing on the security video to generate at least one frame image sequence. The first generating unit 203 is configured to generate a video range for a target object according to the at least one frame image sequence. The dividing unit 204 is configured to perform image sequence division on each frame image sequence in the at least one frame image sequence according to the video range to generate a divided image sequence, thereby obtaining at least one divided image sequence. The second obtaining unit 205 is configured to obtain a video voice in the security video within the video range and a distributed-based model set, wherein the model set includes a multi-modal-based master node object intention recognition information generation model and a plurality of multi-modal-based sub-node object intention recognition information generation models. The second generating unit 206 is configured to generate object intention information for the target object according to the model set, the at least one divided image sequence, and the video voice. The labeling unit 207 is configured to perform target object labeling on a security sub-video corresponding to the video range to generate a labeled video in response to determining that the object intention information is in an abnormal intention information table. The sending unit 208 is configured to send the labeled video and the object intention information to an alarm terminal.
[0138] It can be understood that the units described in the information sending apparatus 200 correspond to the respective steps in the method described with reference to Figure 1 Therefore, the operations, features, and beneficial effects described above for the method also apply to the information sending apparatus 200 and the units included therein, which will not be described here again.
[0139] Reference is made below to Figure 3 which shows a structural schematic diagram of an electronic device (e.g., an electronic device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.
[0140] As shown in Figure 3As shown, the electronic device 300 can include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301 that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 302 or loaded into a random access memory (RAM) 303 from a storage device 308. Various programs and data required for the operation of the electronic device 300 are also stored in the RAM 303. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0141] In general, the following devices can be connected to the I / O interface 305: input devices 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 308 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 309. The communication devices 309 can allow the electronic device 300 to communicate wirelessly or wired with other devices to exchange data. Although Figure 3 The electronic device 300 is shown with various devices, but it should be understood that all of the illustrated devices are not required, and more or fewer devices can alternatively be implemented. Figure 3 Each block shown in the flowcharts can represent a device, or multiple devices, as necessary.
[0142] In particular, processes described above with reference to the flowcharts can be implemented as a computer software program according to some embodiments of the present disclosure. For example, some embodiments of the present disclosure include a computer program product including a computer program carried on a computer readable medium, the computer program containing program code for performing the methods illustrated by the flowcharts. In some such embodiments, the computer program can be downloaded and installed from a network through the communication devices 309, or installed from the storage devices 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-described functions defined in the methods of some embodiments of the present disclosure are performed.
[0143] Note that the computer-readable medium or media used to provide the computer program sequence to the computer system can be embedded in a computer program product, which comprises all the respective features, which are provided with the computer program sequence, and which are enumerated above. It is understood that the computer-readable medium or media described herein are included in the computer program product, or are a component of the computer program product. In some embodiments of the disclosure, the computer-readable storage medium can be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In some embodiments of the disclosure, a computer-readable storage medium can be any tangible medium that contains, or stores a program for use by or in connection with an instruction execution system, apparatus, or device. In some embodiments of the disclosure, a computer-readable signal medium can include a computer-readable storage medium in baseband or propagated as a carrier wave in a propagated data signal, which contains a computer-readable program code. Such a propagated signal can take a wide variety of forms, including, but not limited to, electro-magnetic, optical, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium that is not a computer-readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wire line, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0144] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0145] The computer readable medium can be included in the electronic device, or can exist separately from the electronic device. The computer readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire a security video for a target monitoring area; perform cross-frame extraction processing on the security video to generate at least one frame image sequence; generate a video range for a target object according to the at least one frame image sequence; perform image sequence segmentation on each frame image sequence in the at least one frame image sequence according to the video range to generate segmented image sequences, to obtain at least one segmented image sequence; acquire a video voice in the security video within the video range and a distributed-based model set, wherein the model set includes: a multi-modal-based master node object intention recognition information generation model, and a plurality of multi-modal-based sub-node object intention recognition information generation models; generate object intention information for the target object according to the model set, the at least one segmented image sequence, and the video voice; in response to determining that the object intention information is in an abnormal intention information table, perform target object labeling on a security sub-video corresponding to the video range to generate a labeled video; and send the labeled video and the object intention information to an alarm terminal.
[0146] Computer program code for carrying out operations of some embodiments of the disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0147] The computer program product of the first aspect can include a computer readable storage medium. The computer readable storage medium can include instructions. The instructions can include one or both of: instructions for causing a computer to perform the operations of the method of the first aspect; and instructions for causing a computer to perform the operations of the method of the second aspect. The computer readable storage medium can include a computer readable storage medium as defined above. The computer readable storage medium can include one or both of: a computer readable storage medium that stores the instructions; and a computer readable storage medium that transmits the instructions. The computer readable storage medium can include one or both of: a computer readable storage medium that is further configured to, working with the computer program, cause the computer to perform the operations of the method of the first aspect; and a computer readable storage medium that is further configured to, working with the computer program, cause the computer to perform the operations of the method of the second aspect.
[0148] The units described in some embodiments of the present disclosure can be implemented by means of software, or can be implemented by hardware. The described units can also be arranged in a processor, for example, a processor can be described as: a processor includes a first acquisition unit, a cross-frame processing unit, a first generation unit, a segmentation unit, a second acquisition unit, a second generation unit, a labeling unit and a sending unit. Among them, the names of these units do not constitute a limitation to the units themselves in some cases, for example, the sending unit can also be described as: a unit for sending the labeled video and the object intention information to the alarm terminal.
[0149] The functions described above in the present document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, example types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0150] The above description is merely some of the preferred embodiments of the present disclosure and a description of the principles of the technology used. Those skilled in the art should understand that the scope of the application involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or equivalent features without departing from the above inventive concept. For example, the above features can be replaced with technical features disclosed in the embodiments of the present disclosure (but not limited to) having similar functions to form technical solutions.
Claims
1. A method for sending information, comprising: Acquire security video of the target monitoring area; The security video is divided into video segments based on the maximum frame extraction duration in the frame extraction duration sequence to generate a security sub-video sequence. For the target frame extraction duration in the frame extraction duration sequence, perform the following frame extraction steps: Determine multiple frame image subsequences corresponding to the previous frame extraction duration corresponding to the target frame extraction duration, and use them as multiple first frame image subsequences. Extract a target number of frame images from each security sub-video in the security sub-video sequence to generate a frame image sub-sequence, thus obtaining multiple second frame image sub-sequences; In response to determining that the target frame extraction duration is the minimum frame extraction duration, the plurality of second frame image subsequences are determined as at least one frame image sequence; In response to determining that the target frame extraction duration is not the minimum frame extraction duration, the next frame extraction duration corresponding to the target frame extraction duration is taken as the target frame extraction duration, and the frame extraction step is continued. Based on the at least one frame image sequence, generate a video range for the target object; Based on the video range, each frame image sequence in the at least one frame image sequence is segmented to generate a segmented image sequence, thereby obtaining at least one segmented image sequence; Acquire video and audio data from the security video within the video range and a distributed model set, wherein the model set includes: a multimodal master node object intent recognition information generation model and a multimodal multiple child node object intent recognition information generation model; Based on the model set, the at least one segmented image sequence, and the video audio, generate object intent information for the target object; In response to determining that the object intent information is in the abnormal intent information table, the security sub-video corresponding to the video range is labeled with the target object to generate a labeled video; The labeled video and the object's intent information are sent to the alarm terminal.
2. The method according to claim 1, wherein, The step of generating a video range for the target object based on the at least one frame image sequence includes: For each frame image sequence in the at least one frame image sequence, the following first generation step is performed: Each frame image in the frame image sequence is input into the target detection model for the target object to generate target detection information, thus obtaining a target detection information sequence; Remove the corresponding target detection information from the frame image sequence to represent the absence of the target object, and obtain the frame image sequence after removal; The removed image sequence is supplemented with consecutive frames before and after it to generate a supplemented image sequence. Determine the initial video range corresponding to the supplemented frame image sequence; The obtained initial video ranges are subjected to a union process to generate a union video range, which is used as the video range.
3. The method according to claim 1, wherein, The step of generating object intent information for the target object based on the model set, the at least one segmented image sequence, and the video audio includes: For each segmented image sequence in the at least one segmented image sequence, the following second generation step is performed: Generate a sequence of image feature information for the segmented image sequence; Based on the image set corresponding to each image feature information in the image feature information sequence, the video speech is divided into speech segments to generate sub-speech sequences, wherein the sub-speech segments in the sub-speech sequences have a time correspondence with the image feature information in the image feature information sequence; For each sub-speech in the sub-speech sequence, perform the following third generation step: The sub-speech is preprocessed to generate processed audio; The processed audio is subjected to a Fast Fourier Transform to generate a frequency domain signal; Using a Mel filter bank, a processed frequency domain signal corresponding to the frequency domain signal is generated; The processed frequency domain signal is subjected to logarithmic processing to generate a linearly enhanced signal; The linear enhancement signal is subjected to discrete cosine transformation to generate transformed audio feature information; Based on the number of segmented images corresponding to the segmented image sequence, a corresponding child node object intent recognition information generation model is selected from the plurality of child node object intent recognition information generation models as the target child node object intent recognition information generation model, wherein the output weight ratio of the target child node object intent recognition information generation model is related to the number of segmented images; Based on the image feature information sequence and the obtained changed audio feature information sequence, the target child node object intent recognition information generation model is used to generate initial object intent information; The object intent information is generated based on at least one initial object intent information obtained.
4. The method according to claim 3, wherein, The target sub-node object intent recognition information generation model includes: multiple feature extraction units, multiple feature fusion units, and a recognition information generation layer; and The step of generating initial object intent information based on the image feature information sequence and the obtained changed audio feature information sequence, using the target child node object intent recognition information generation model, includes: Determine the time step sequence corresponding to the image feature information sequence; Select the target time step from the time step sequence; For the target time step, perform the following third generation step: Image feature information that has a time correspondence with the target time step is obtained from the image feature information sequence and used as target image feature information; and changed audio feature information that has a time correspondence with the target time step is obtained from the changed audio feature information sequence and used as target changed audio feature information. The feature extraction unit corresponding to the target time step is determined as the target feature extraction unit, and the feature fusion unit corresponding to the target time step is determined as the target feature fusion unit. Determine the previous time step corresponding to the target time step; The fusion feature information corresponding to the previous time step is determined as the target fusion feature information; The changed audio feature information of the target is input into the audio feature weight information generation model included in the target feature extraction unit to generate audio feature weight information; The audio feature weight information and the target changed audio feature information are multiplied together to generate multiplied audio feature information; The multiplied audio feature information and the target image feature information are input into the attention mechanism model included in the target feature extraction unit to generate first attention weight information for the multiplied audio feature information and second attention weight information for the target image feature information; The first attention weight information, the second attention weight information, the multiplied audio feature information, and the target image feature information are input into the target feature fusion unit to generate fused feature information; The fused feature information and the target fused feature information are concatenated to generate concatenated feature information; In response to determining that the target time step is the latest time step in the time step sequence, the splicing feature information corresponding to the latest time step is input to the recognition information generation layer to generate initial object intent information; In response to determining that the target time step is not the latest time step in the time step sequence, the next time step corresponding to the target time step is taken as the target time step, and the third generation step is continued.
5. The method according to claim 3, wherein, The step of generating the object intent information based on at least one obtained initial object intent information includes: Determine at least one output weight percentage corresponding to the at least one initial object intent information; The initial object intent information in the at least one initial object intent information is classified to generate a candidate object intent information set, wherein each candidate object intent information has a corresponding object intent category; Determine the intent information of each candidate object in the candidate object intent information set, and perform the following fourth generation step: Determine the output weight proportion set corresponding to the intent information of the candidate objects; The output weight percentages in the output weight percentage set are summed to generate the total output weight percentage. Determine whether there exists any candidate object intent information in the candidate object intent information set whose total corresponding output weight percentage is greater than the first weight percentage; In response to the determination of existence, the candidate object intent information whose sum of corresponding output weight percentages is greater than the first weight percentage is determined as object intent information.
6. The method according to claim 5, wherein, The step of generating the object intent information based on at least one obtained initial object intent information further includes: In response to the determination that it does not exist, at least one candidate object intent information whose corresponding output weight ratio is greater than the second weight ratio is determined from the candidate object intent information set, and is used as at least one target candidate object intent information, wherein the second weight ratio is greater than the first weight ratio. Using the master node object intent recognition information to generate a model, the object intent information is selected from the at least one target candidate object intent information; From the plurality of child node object intent recognition information generation models, determine at least one child node object intent recognition information generation model whose corresponding output intent information is not the object intent information; Perform model retraining for the intent recognition information generation model for the at least one child node object.
7. An information transmitting device, comprising: The first acquisition unit is configured to acquire security video of the target monitoring area. The apparatus further includes: dividing the security video into video segments based on the maximum frame extraction duration in the frame extraction duration sequence to generate a security sub-video sequence; for a target frame extraction duration in the frame extraction duration sequence, performing the following frame extraction steps: determining multiple frame image sub-sequences corresponding to the previous frame extraction duration corresponding to the target frame extraction duration, as multiple first frame image sub-sequences; extracting a target number of frame images from each security sub-video in the security sub-video sequence to generate a frame image sub-sequence, obtaining multiple second frame image sub-sequences; in response to determining that the target frame extraction duration is the minimum frame extraction duration, determining the multiple second frame image sub-sequences as at least one frame image sequence; in response to determining that the target frame extraction duration is not the minimum frame extraction duration, taking the next frame extraction duration corresponding to the target frame extraction duration as the target frame extraction duration, and continuing to perform the frame extraction steps; The first generation unit is configured to generate a video range for a target object based on the at least one frame image sequence; The segmentation unit is configured to perform image sequence segmentation on each frame image sequence in the at least one frame image sequence according to the video range to generate a segmented image sequence, thereby obtaining at least one segmented image sequence; The second acquisition unit is configured to acquire video and audio from the security video within the video range and a distributed model set, wherein the model set includes: a multimodal master node object intent recognition information generation model and a multimodal multiple child node object intent recognition information generation model. The second generation unit is configured to generate object intent information for the target object based on the model set, the at least one segmented image sequence, and the video audio. The annotation unit is configured to, in response to determining that the object intent information is in the abnormal intent information table, annotate the security sub-video corresponding to the video range with target objects to generate an annotated video; The sending unit is configured to send the labeled video and the object intent information to the alarm terminal.
8. An electronic device, comprising: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-6.
9. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Behavior recognition method and device
CN112820071A
Self-adaptive frame extraction method and system for human body detection algorithm of edge AI equipment
CN118018667A