Object behavior prompting method and device, processor and electronic equipment

By extracting the target characteristics of the object from the video data and performing behavior prediction, the problem of low behavior detection accuracy in the prior art is solved, and higher behavior detection accuracy and stability are achieved.

CN119942404APending Publication Date: 2025-05-06CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411999233.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The behavior detection methods used in existing detection devices are easily affected by the external environment, resulting in low accuracy of behavior detection.

Method used

By determining the target area where the object's geometric outline is located from the video data, the target features are extracted, and the current behavior is predicted based on these features to obtain the prediction results. If the prediction result represents a similarity greater than the similarity threshold, prompt information is output, allowing the current behavior to be determined as the target behavior.

Benefits of technology

This method can reduce the difficulty of behavior detection, improve the accuracy of behavior detection, and avoid the influence of external environment, thereby effectively solving the problem of low behavior detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942404A_ABST
    Figure CN119942404A_ABST
Patent Text Reader

Abstract

The invention discloses an object behavior prompting method and device, a processor and electronic equipment. The method comprises the steps of determining a target area where a geometric contour of at least one object is located from video data; extracting target features of the object from a target image where the target area is located; based on the target feature, predicting a current behavior executed by the object to obtain a prediction result, the prediction result being used for representing the similarity between the current behavior and a target behavior associated with the target feature; in response to the fact that the similarity expressed by the prediction result is larger than a similarity threshold value, prompt information is output, and the prompt information is used for prompting that the current behavior is allowed to be determined as the target behavior. The technical problem of low accuracy of behavior detection is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and in particular to a method, device, processor and electronic device for prompting object behavior. Background Art

[0002] In some places, some behaviors need to be prohibited. In order to ensure that some behaviors can be effectively prohibited in the above places, detection devices are often used to detect the above behaviors. However, the detection methods used in existing detection devices are easily affected by the external environment, which makes it difficult to detect behaviors, thereby leading to the technical problem of low accuracy of behavior detection.

[0003] Currently, no effective solution has been proposed to the technical problem of low accuracy of the above-mentioned behavior detection. Summary of the invention

[0004] The embodiments of the present invention provide a method, device, processor and electronic device for prompting an object behavior, so as to at least solve the technical problem of low accuracy of behavior detection.

[0005] According to one aspect of an embodiment of the present invention, a method for prompting an object behavior is provided, the method comprising: determining a target area where a geometric outline of at least one object is located from video data; extracting a target feature of the object from a target image where the target area is located; based on the target feature, predicting a current behavior performed by the object to obtain a prediction result, wherein the prediction result is used to represent the similarity between the current behavior and a target behavior associated with the target feature; in response to the similarity represented by the prediction result being greater than a similarity threshold, outputting a prompt message, wherein the prompt message is used to prompt that the current behavior is allowed to be determined as the target behavior.

[0006] Optionally, determining a target area where a geometric outline of at least one object is located from the video data includes: determining at least one video frame data including the geometric outline from the video data; and determining at least one target area from the video frame data.

[0007] Optionally, determining at least one target area from the video frame data includes: generating a bounding box matching the geometric contour from the video frame data; and determining an area where the bounding box is located in the video frame data as the target area.

[0008] Optionally, extracting target features of the object from a target image where the target area is located includes: intercepting image data of the target area from video frame data to obtain the target image; scaling the target image; and extracting target features from the scaled target image.

[0009] Optionally, based on the target feature, predicting the current behavior performed by the object to obtain a prediction result includes: standardizing the target feature; and predicting the current behavior based on the standardized target feature to obtain a prediction result.

[0010] Optionally, based on the standardized sub-goal features, the current behavior of the object corresponding to the sub-goal features is predicted to obtain a prediction result, including: inputting the standardized sub-goal features into a target prediction model, using the target prediction model to predict the current behavior of the object corresponding to the sub-goal features to obtain a prediction result, wherein the target prediction model is obtained by training an initial prediction model using target feature samples of object samples and corresponding prediction result samples, and the prediction result samples are used to represent the similarity between the historical behavior performed by the object samples and the target behavior.

[0011] According to one aspect of an embodiment of the present invention, a device for prompting an object behavior is provided, and the device may include: a determination unit, used to determine, from video data, a target area where a geometric outline of at least one object is located; an extraction unit, used to extract a target feature of the object from a target image where the target area is located; a prediction unit, used to predict a current behavior performed by the object based on the target feature, and obtain a prediction result, wherein the prediction result is used to represent the similarity between the current behavior and the target behavior associated with the target feature; an output unit, used to output prompt information in response to the similarity represented by the prediction result being greater than a similarity threshold, wherein the prompt information is used to prompt that the current behavior is allowed to be determined as the target behavior.

[0012] According to another aspect of an embodiment of the present invention, a processor is provided, which is used to run a program, wherein when the program is run by the processor, the object behavior prompt method in the embodiment of the present invention is executed.

[0013] According to another aspect of an embodiment of the present invention, an electronic device is provided, including: a memory storing an executable program; and a processor for running the program, wherein when the program is running, the object behavior prompting method in each embodiment of the present invention is executed.

[0014] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the object behavior prompting method in the present invention.

[0015] According to another aspect of an embodiment of the present invention, a computer program product is further provided. The computer program product includes a computer program. When the computer program is executed by a processor, the object behavior prompting method in the embodiment of the present invention is implemented.

[0016] According to another aspect of an embodiment of the present invention, a computer program product is also provided, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the object behavior prompting method in the embodiment of the present invention is implemented.

[0017] According to another aspect of an embodiment of the present invention, an embodiment of the present application further provides a computer program, which, when executed by a processor, implements the object behavior prompting method in the above-mentioned embodiment of the present invention.

[0018] In an embodiment of the present invention, when prompting an object behavior, a target area where the geometric outline of at least one object is located can be determined from video data. From the target image where the determined target area is located, the target features of the object can be extracted, and the current behavior performed by the object can be predicted based on the extracted target features, and a prediction result can be obtained, and in response to the obtained prediction result indicating that the similarity between the current behavior and the target behavior is greater than a similarity threshold, prompt information can be output. Since the current behavior being performed is predicted, the similarity between the current behavior and the target behavior associated with the target feature can be obtained, avoiding the influence of the external environment, thereby achieving the purpose of reducing the difficulty of behavior detection, thereby solving the technical problem of low accuracy of behavior detection, and further achieving the technical effect of improving the accuracy of behavior detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0020] Figure 1 is a flow chart of a method for prompting an object behavior according to an embodiment of the present invention;

[0021] FIG2( a ) is a flow chart of a method for identifying smoking behavior based on a key point sequence according to an embodiment of the present invention;

[0022] FIG2( b ) is a schematic diagram of a smoking region of interest according to an embodiment of the present invention;

[0023] FIG2( c ) is a flow chart of a model training method according to an embodiment of the present invention;

[0024] Figure 3 2 is a schematic diagram of a device for prompting an object behavior according to an embodiment of the present invention. DETAILED DESCRIPTION

[0025] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only embodiments of a part of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0027] According to an embodiment of the present invention, a method for prompting object behavior is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0028] Figure 1 1 is a flowchart of a method for prompting an object behavior according to an embodiment of the present invention. The method may include the following steps:

[0029] Step S101, determining a target area where a geometric outline of at least one object is located from video data.

[0030] In the technical solution provided in the above step S101 of the present invention, the above video data may include at least one video frame data, the above object may include sub-objects, the sub-objects may include faces, hands, cigarettes, etc., wherein the face may be a human face.

[0031] In this embodiment, the target region may be a preset region of interest (ROI). The size of the ROI may be an initial default size or a size set according to user requirements, and the shape of the ROI may be an initial default shape or a shape set according to user requirements. For example, if the size of the ROI is limited to 400*400 in user requirements, and the shape of the ROI is limited to a square, then the size of the ROI is 400*400, and the shape of the ROI is a square. This is only an example and is not specifically limited.

[0032] In this embodiment, a target area where the geometric outline of at least one object is located is determined from the video data. Optionally, this embodiment performs frame extraction on the collected video data, and the target area where the geometric outline of at least one object is located can be determined from the video data after frame extraction. The frame number that needs to be satisfied by the frame extraction can be, but is not limited to, 10fps.

[0033] Step S102: extracting target features of the object from the target image where the target area is located.

[0034] In the technical solution provided in the above step S102 of the present invention, the above target features may include features of a face, features of a hand, features of a cigarette, and the like.

[0035] In this embodiment, after determining the target area where the geometric outline of at least one object is located from the video data, the target features of the object are extracted from the target image where the target area is located. Optionally, this embodiment extracts features of the target image where the target area is located based on the determination of the target area, and the features of the target image can be obtained. From the features of the above-obtained target image, the target features of the object can be determined, for example, the features of the face, the features of the hand, the features of the cigarette, etc. can be determined.

[0036] Optionally, the target features of the object can be determined from the features of the target image obtained. That is, the features of the target image obtained can be divided according to the sub-objects present in the object to obtain the features of the sub-objects, and the target features of the object can be determined based on the features of the sub-objects obtained.

[0037] Step S103: predict the current behavior performed by the object based on the target features to obtain a prediction result.

[0038] In the technical solution provided in step S103 of the present invention, the prediction result can be used to indicate the similarity between the current behavior and the target behavior associated with the target feature. The target behavior can be a pre-set behavior, for example, the pre-set behavior can be a smoking behavior, which is only used as an example and is not specifically limited.

[0039] In this embodiment, after extracting the target features of the object from the target image where the target area is located, the current behavior performed by the object is predicted based on the target features to obtain a prediction result. Optionally, this embodiment performs feature enhancement on the extracted target features on the basis of extracting the target features of the object, and predicts the current behavior performed by the object using the enhanced target features to obtain a prediction result, that is, the similarity between the current behavior and the target behavior associated with the target features can be obtained.

[0040] Optionally, effective features are screened out from the extracted target features, and the effective features are enhanced, wherein the effective features can be used to represent features without error items and without missing items. The enhanced effective features are used to predict the current behavior performed by the object, and the similarity between the current behavior and the target behavior associated with the target features can be obtained.

[0041] Step S104: in response to the similarity indicated by the prediction result being greater than the similarity threshold, outputting prompt information.

[0042] In the technical solution provided in the above step S104 of the present invention, the above prompt information can be used to prompt permission to determine the current behavior as the target behavior.

[0043] In this embodiment, the above-mentioned similarity threshold may be a critical threshold for determining whether the current behavior is close to the target behavior. The critical threshold may be an initialization threshold or a threshold set according to different types of target behaviors. For example, if the target behavior is smoking behavior, the similarity threshold may be 98%. Optionally, the similarity represented by the prediction result is 100%, that is, the current behavior is consistent with the target behavior. Optionally, in the case where a certain prediction error is allowed in the prediction result, even if the similarity represented by the prediction result is not 100%, for example, the similarity represented by the prediction result is greater than 98% and less than 100%, it means that the current behavior is close to the target behavior. In this case, the current behavior may also be determined as the target behavior. This is only an example and is not specifically limited here.

[0044] In this embodiment, after the current behavior performed by the object is predicted based on the target feature and the prediction result is obtained, in response to the similarity represented by the prediction result being greater than the similarity threshold, prompt information is output. Optionally, based on the prediction result, the embodiment judges the relationship between the similarity represented by the prediction result and the similarity threshold, and a judgment result can be obtained. If the judgment result obtained is that the similarity is greater than the similarity threshold, prompt information can be output, thereby prompting to allow the current behavior to be determined as the target behavior.

[0045] It should be noted that the object behavior prompt method in the present application can be applied to at least one of the following places: medical places, construction sites, restaurants, warehouses, offices, special parks, etc. These are only examples and are not specifically limited.

[0046] In the above steps S101 to S104 of the present application, when prompting the behavior of an object, a target area where the geometric outline of at least one object is located can be determined from the video data. From the target image where the determined target area is located, the target features of the object can be extracted, and the current behavior performed by the object can be predicted based on the extracted target features, and a prediction result can be obtained, and in response to the obtained prediction result indicating that the similarity between the current behavior and the target behavior is greater than the similarity threshold, a prompt message can be output. Since the current behavior being performed is predicted, the similarity between the current behavior and the target behavior associated with the target feature can be obtained, avoiding the influence of the external environment, thereby achieving the purpose of reducing the difficulty of behavior detection, thereby solving the technical problem of low accuracy of behavior detection, and further achieving the technical effect of improving the accuracy of behavior detection.

[0047] The above method of this embodiment is further introduced below.

[0048] As an optional implementation method, step S101, determining a target area where a geometric outline of at least one object is located from video data, includes: determining at least one video frame data including a geometric outline from the video data; determining at least one target area from the video frame data.

[0049] In this embodiment, at least one video frame data including the geometric outline is determined from the video data. Optionally, this embodiment extracts frames from the collected video data to obtain at least one video frame data including the geometric outline.

[0050] In this embodiment, after determining at least one video frame data including a geometric outline from the video data, at least one target area is determined from the video frame data. Optionally, this embodiment can generate a geometric outline of an object in the determined video frame data based on the at least one video frame data. The initial area where the generated geometric outline is located is expanded, and the initial area obtained by the expansion is determined as the target area.

[0051] As an optional implementation manner, determining at least one target area from the video frame data includes: generating a bounding box matching the geometric contour from the video frame data; and determining the area where the bounding box is located in the video frame data as the target area.

[0052] In this embodiment, the above-mentioned boundary box may also be referred to as an expanded box.

[0053] In this embodiment, after determining at least one video frame data from the video data, a bounding box matching the geometric outline is generated from the video frame data. Optionally, this embodiment can generate a geometric outline of the object in the determined video frame data based on the at least one video frame data. The initial area where the generated geometric outline is located is expanded to obtain a bounding box matching the geometric outline.

[0054] In this embodiment, after generating a bounding box matching the geometric contour from the video frame data, the area where the bounding box is located in the video frame data is determined as the target area. Optionally, based on the bounding box, this embodiment uses the initial area obtained by expansion as the area where the bounding box is located in the video frame data, and determines the area where the bounding box is located in the video frame data as the target area.

[0055] As an optional implementation method, step S102 extracts target features of the object from the target image where the target area is located, including: intercepting image data of the target area from the video frame data to obtain the target image; scaling the target image; and extracting target features from the scaled target image.

[0056] In this embodiment, after determining the target area where the geometric outline of at least one object is located from the video data, the image data of the target area is intercepted from the video frame data to obtain the target image. Optionally, based on the determination of the target area, the embodiment intercepts the image data of the target area from the video frame data according to the size of the target area to obtain the target image. For example, if the size of the target area is 500*500, the image data of the target area is intercepted according to 500*500, and a target image with a size of 500*500 can be obtained. The numerical values ​​here are only for illustration and are not specifically limited.

[0057] In this embodiment, after the image data of the target area is intercepted from the video frame data and the target image is obtained, the target image is scaled; and the target features are extracted from the scaled target image. Optionally, on the basis of obtaining the target image, the embodiment scales the obtained target image according to the scaling ratio. The features of the scaled target image are extracted by extracting the features. From the features of the scaled target image, the target features of the object can be determined, for example, the features of a face, the features of a hand, and the features of a cigarette, etc. can be determined.

[0058] Optionally, the target features of the object can be determined from the features of the scaled target image. That is, the features of the scaled target image are divided according to the sub-objects present in the object to obtain the features of the scaled sub-objects, and the features of the scaled sub-objects can be combined into the target features of the object.

[0059] As an optional implementation method, step S103 predicts the current behavior performed by the object based on the target features to obtain a prediction result, including: standardizing the target features; and predicting the current behavior based on the standardized target features to obtain a prediction result.

[0060] In this embodiment, the above-mentioned standardization process may include missing item supplementation and normalization process.

[0061] In this embodiment, after extracting the target features of the object from the target image where the target area is located, the target features are standardized. Optionally, on the basis of extracting the target features of the object, this embodiment supplements the missing items of the extracted target features and normalizes the supplemented target features, thereby achieving standardized processing of the extracted target features.

[0062] In this embodiment, after the target feature is standardized, the current behavior is predicted based on the standardized target feature to obtain a prediction result. Optionally, this embodiment enhances the target feature after the standardized process on the basis of the standardized process, and uses the enhanced target feature to predict the current behavior performed by the object, so as to obtain the similarity between the current behavior and the target behavior associated with the target feature.

[0063] As an optional implementation method, based on the standardized sub-goal features, the current behavior of the object corresponding to the sub-goal features is predicted to obtain a prediction result, including: inputting the standardized sub-goal features into a target prediction model, and using the target prediction model to predict the current behavior of the object corresponding to the sub-goal features to obtain a prediction result.

[0064] In this embodiment, the target prediction model can be obtained by training the initial prediction model using the target feature samples of the object samples and the corresponding prediction result samples. The prediction result samples can be used to indicate the similarity between the historical behavior performed by the object samples and the target behavior. The target prediction model can be, but is not limited to, a smoking behavior recognition reasoning model, an ignition behavior recognition reasoning model, and a littering behavior recognition reasoning model.

[0065] In this embodiment, after the target feature is standardized, the standardized sub-target feature is input into the target prediction model, and the target prediction model is used to predict the current behavior of the object corresponding to the sub-target feature to obtain a prediction result. Optionally, this embodiment performs feature enhancement on the target feature after the standardized process on the basis of the target feature being standardized, and the enhanced target feature is input into the target prediction model, and the prediction layer in the target prediction model is used to predict the current behavior of the object corresponding to the sub-target feature to obtain a prediction result.

[0066] In an embodiment of the present invention, when prompting an object behavior, a target area where the geometric outline of at least one object is located can be determined from video data. From the target image where the determined target area is located, the target features of the object can be extracted, and the current behavior performed by the object can be predicted based on the extracted target features, and a prediction result can be obtained, and in response to the obtained prediction result indicating that the similarity between the current behavior and the target behavior is greater than a similarity threshold, prompt information can be output, thereby achieving the purpose of reducing the difficulty of behavior detection, thereby solving the technical problem of low accuracy of behavior detection, and further achieving the technical effect of improving the accuracy of behavior detection.

[0067] The technical solution of the embodiment of the present invention is illustrated below in conjunction with preferred implementation modes.

[0068] In some places, some behaviors need to be prohibited. In order to ensure that some behaviors can be effectively prohibited in the above places, detection devices are often used to detect the above behaviors. However, the detection methods used in existing detection devices are easily affected by the external environment, which makes it difficult to detect behaviors, thereby leading to the technical problem of low accuracy of behavior detection.

[0069] In order to solve the above technical problems, an embodiment of the present invention proposes a method for prompting object behavior. When prompting object behavior, a target area where the geometric outline of at least one object is located can be determined from video data. From the target image where the determined target area is located, the target features of the object can be extracted, and the current behavior performed by the object can be predicted based on the extracted target features, and a prediction result can be obtained, and in response to the obtained prediction result indicating that the similarity between the current behavior and the target behavior is greater than a similarity threshold, a prompt information can be output, thereby achieving the purpose of reducing the difficulty of behavior detection, thereby solving the technical problem of low accuracy of behavior detection, and further achieving the technical effect of improving the accuracy of behavior detection.

[0070] In this embodiment, the smoking behavior recognition method based on the key point sequence is executed to determine whether the current behavior performed by the object is the target behavior. For example, FIG2(a) is a flowchart of a smoking behavior recognition method based on the key point sequence according to an embodiment of the present invention. As shown in FIG2(a), the method may include the following steps:

[0071] Step S201, obtaining video stream data collected by a collection device.

[0072] In the technical solution provided in the above step S201 of the present invention, the above acquisition equipment may include: a camera and a drone, etc.

[0073] After acquiring the video stream data collected by the collection device, the process proceeds to step S202 to perform frame extraction processing on the video stream data to obtain video frame data.

[0074] In the technical solution provided in the above step S202 of the present invention, the camera image sequence is collected and sampled, wherein the sampling frequency may be above 10 fps, and the numerical value here is only used as an example and is not specifically limited.

[0075] After the video stream data is subjected to frame extraction processing to obtain the video frame data, the process proceeds to step S203 to perform target detection on the video frame data.

[0076] In the technical solution provided in the above step S203 of the present invention, the sampled video frame data is input into the target detection model for target detection, thereby detecting faces, cigarettes, hands, etc. from the video frame data.

[0077] After performing target detection on the video frame data, the process proceeds to step S204 and step S205, where the detected object is preprocessed using the target detection frame, and the detected face is tracked.

[0078] In the technical solution provided in step S205 of the present invention, the detected face is tracked by a tracker (ByteTrack). In the video frame data, the face, cigarette and hand in the ROI with the same tracking identification (ID) are regarded as a whole object, and the ROI can be an expanded frame based on the face frame or an expanded frame based on a certain position on the face.

[0079] After the detected face is tracked, the process proceeds to step S206 and step S207 to extract the smoking region of interest using the face frame, and to extract key points from the smoking region of interest.

[0080] In the technical solution provided in the above step S207 of the present invention, the content within the ROI is intercepted, the intercepted target image is scaled to an image of a fixed size (for example but not limited to, scaled to 384*384, 256*256, etc.), and key points are extracted from the scaled image, so that key points such as the mouth, cigarette, and hand, as well as the confidence levels corresponding to the key points, can be obtained.

[0081] In this embodiment, the key point information is judged as follows in combination with the detected face frame, cigarette frame and hand frame. The specific judgment strategies are as follows: Strategy 1, judge whether the rectangular frame circumscribed by the mouth key point is in the lower half of the face frame, and whether the intersection over union (IOU) with the face frame is greater than 80%. If the IOU is greater than 80%, the original information of the face key point is retained, otherwise the face key point is set to an invalid value; Strategy 2, judge whether the IOU of the rectangular frame circumscribed by the cigarette key point and the cigarette detection frame is greater than 50%. If it is greater than 50%, the original information of the cigarette key point is retained, otherwise the cigarette key point is set to an invalid value; Strategy 3, judge whether the IOU of the rectangular frame circumscribed by the hand key point and the hand detection frame is greater than 50%. If it is greater than 50%, the original information of the hand key point is retained, otherwise the hand key point is set to an invalid value.

[0082] After performing target detection on the video frame data and extracting key points from the smoking region of interest, the process proceeds to step S204.

[0083] In this embodiment, the optimized key point information of N consecutive frames is divided according to the tracking ID, and then the missing key point information of the same ID is supplemented by linear interpolation, and the supplemented key point information is normalized to obtain the normalized key point information, that is, the target feature can be obtained.

[0084] After preprocessing the detected object using the target detection frame, proceed to step S208 and step S209, input the target features into the smoking behavior recognition inference model for prediction, obtain the prediction result, and output prompt information in response to the similarity represented by the prediction result being greater than the similarity threshold.

[0085] In the technical solution provided in the above step S209 of the present invention, the normalized key point information is input into the smoking behavior recognition inference model based on time series for recognition, so as to determine whether the current behavior is smoking behavior.

[0086] In this embodiment, the smoking region of interest may be as shown in FIG2(b). For example, FIG2(b) is a schematic diagram of a smoking region of interest according to an embodiment of the present invention. As shown in FIG2(b), the smoking region of interest may be an expanded frame S' based on the expansion of the face frame S, for example, a smoking region of interest frame based on the expansion of the face frame.

[0087] In this embodiment, the overlap between the candidate frame and the target frame can be expressed by an intersection-to-union ratio. For example, the overlap can be expressed by the ratio between the intersection of the candidate frame A and the target frame B, and the union of the candidate frame A and the target frame B.

[0088] Optionally, the candidate frame A may be a rectangular frame circumscribed by the mouth key point, a rectangular frame circumscribed by the cigarette key point, or a rectangular frame circumscribed by the hand key point. The target frame B may be a face frame, a cigarette detection frame, or a hand detection frame.

[0089] In this embodiment, the model training method is executed to train a target detection model, a key point detection model, and a smoking behavior recognition inference model. For example, FIG2(c) is a flow chart of a model training method according to an embodiment of the present invention. As shown in FIG2(c), the method may include the following steps:

[0090] Step S231, collecting positive and negative sample video sequences of smoking behavior.

[0091] In the technical solution provided in step S231 of the present invention, the positive sample video sequence may include the entire continuous action of smoking, such as holding a cigarette in the hand and holding the cigarette in the mouth. The negative sample video sequence may include the entire continuous action of holding a rod-shaped object in the hand, different lighting and false smoking action.

[0092] After collecting the positive and negative sample video sequences of smoking behavior, the process proceeds to step S232 to perform frame extraction processing on the positive and negative sample video sequences.

[0093] After the positive and negative sample video sequences are subjected to frame extraction processing, the process proceeds to step S233 to perform data labeling on the video frame sequence.

[0094] In the technical solution provided in the above step S233 of the present invention, data annotation of the video frame sequence may include: data annotation of the face ID, face frame, cigarette frame, 8 mouth key points, 2 cigarette key points and hand frame, etc. For non-existent content, it is marked that there is no such attribute. The number of the above key points is only for example and is not specifically limited.

[0095] After data annotation of the video frame sequence, proceed to step S234, step S235 and step S236 to pre-process the training data of the target detection model, pre-process the training data of the key point detection model, and pre-process the training data of the smoking behavior recognition inference model.

[0096] In the technical solution provided in the above step S234 of the present invention, a non-continuous frame image set including annotation information of a face, a cigarette and a hand can be extracted from the annotation results.

[0097] In the technical solution provided in step S235 of the present invention, a data set for training a key point detection model can be extracted from the labeled results. The video data in the data set for training the key point detection model needs to include at least one of the following information: 8 mouth key points, 2 cigarette key points, and a hand detection frame, etc. The number of key points here is only for example and is not specifically limited.

[0098] In the technical solution provided in the above step S236 of the present invention, a data set for training a smoking behavior recognition inference model can be extracted from the results obtained by annotation. Among them, the data set for training a smoking behavior recognition inference model may include M subsets of continuous frame images, each subset includes N frames of images; each subset provides a label of whether the subset is smoking; each subset of N frames of images needs to meet a scene of more than t seconds, and includes at least 10 frames per second (that is, the frame rate of the original video needs to meet more than 10fps); the N frames of images in each subset are obtained by uniform sampling. In each subset, the face of the same ID, the cigarette and the hand in the ROI are treated as a whole object. If a sub-object does not exist, the key point information corresponding to the sub-object is set to an invalid value.

[0099] After preprocessing the training data of the target detection model, proceed to step S237 to train the target detection model.

[0100] In the technical solution provided in the above step S237 of the present invention, the above target detection model can be used to detect sub-objects such as faces, cigarettes and hands included in the video frame sequence.

[0101] After preprocessing the training data of the key point detection model, proceed to step S238 to train the key point detection model.

[0102] After preprocessing the training data of the smoking behavior recognition inference model, the process proceeds to step S239 to train the smoking behavior recognition inference model.

[0103] After training the target detection model, training the key point detection model, and training the smoking behavior recognition inference model, proceed to step S240 to export the trained target detection model, key point detection model, and smoking behavior recognition inference model.

[0104] In the technical solution provided in the above step S240 of the present invention, operations such as model conversion and model quantization are performed on the trained target detection model, key point detection model and smoking behavior recognition inference model for deployment and implementation.

[0105] In this embodiment, when prompting an object behavior, a target area where the geometric outline of at least one object is located can be determined from the video data. From the target image where the determined target area is located, the target features of the object can be extracted, and the current behavior performed by the object can be predicted based on the extracted target features, and a prediction result can be obtained, and in response to the obtained prediction result indicating that the similarity between the current behavior and the target behavior is greater than the similarity threshold, prompt information can be output, thereby achieving the purpose of reducing the difficulty of behavior detection, thereby solving the technical problem of low accuracy of behavior detection, and further achieving the technical effect of improving the accuracy of behavior detection.

[0106] According to an embodiment of the present invention, a device for prompting an object behavior is also provided. It should be noted that the device for prompting an object behavior can be used to execute a method for prompting an object behavior in an embodiment.

[0107] Figure 3 is a schematic diagram of a device for prompting object behavior according to an embodiment of the present invention. Figure 3 As shown, the object behavior prompting device 300 may include: a determination unit 301 , an extraction unit 302 , a prediction unit 303 and an output unit 304 .

[0108] The determination unit 301 is used to determine a target area where a geometric outline of at least one object is located from video data.

[0109] The extraction unit 302 is used to extract target features of the object from the target image where the target area is located.

[0110] The prediction unit 303 is used to predict the current behavior performed by the object based on the target feature to obtain a prediction result, wherein the prediction result is used to represent the similarity between the current behavior and the target behavior associated with the target feature.

[0111] The output unit 304 is configured to output prompt information in response to the similarity indicated by the prediction result being greater than a similarity threshold, wherein the prompt information is used to prompt that the current behavior is allowed to be determined as the target behavior.

[0112] Optionally, the determination unit 301 may include: a first determination module, used to determine at least one video frame data including a geometric contour from the video data; and a second determination module, used to determine at least one target area from the video frame data.

[0113] Optionally, the second determination module may include: a generation submodule for generating a bounding box matching the geometric contour from the video frame data; and a determination submodule for determining the area where the bounding box is located in the video frame data as the target area.

[0114] Optionally, the extraction unit 302 may include: a capture module for capturing image data of a target area from video frame data to obtain a target image; a scaling module for scaling the target image; and an extraction module for extracting target features from the scaled target image.

[0115] Optionally, the prediction unit 303 may include: a processing module, used for performing standardization processing on the target feature; and a prediction module, used for predicting the current behavior based on the target feature after the standardization processing to obtain a prediction result.

[0116] Optionally, the prediction module may include: a prediction sub-module, used to input the standardized sub-target features into the target prediction model, and use the target prediction model to predict the current behavior of the object corresponding to the sub-target features to obtain a prediction result, wherein the target prediction model is obtained by training an initial prediction model using the target feature samples of the object samples and the corresponding prediction result samples, and the prediction result samples are used to represent the similarity between the historical behavior executed by the object samples and the target behavior.

[0117] In this embodiment, a determination unit is used to determine, from video data, a target area where a geometric outline of at least one object is located; an extraction unit is used to extract a target feature of the object from a target image where the target area is located; a prediction unit is used to predict a current behavior performed by the object based on the target feature to obtain a prediction result, wherein the prediction result is used to represent the similarity between the current behavior and the target behavior associated with the target feature; an output unit is used to output prompt information in response to the similarity represented by the prediction result being greater than a similarity threshold, wherein the prompt information is used to prompt that the current behavior is allowed to be determined as the target behavior, thereby achieving the purpose of reducing the difficulty of behavior detection, thereby solving the technical problem of low accuracy of behavior detection, and further achieving the technical effect of improving the accuracy of behavior detection.

[0118] According to an embodiment of the present invention, a processor is further provided, the processor being used to run a program, wherein the program, when run by the processor, executes the object behavior prompting method in the embodiment.

[0119] According to an embodiment of the present invention, there is further provided an electronic device, comprising: a memory storing an executable program; and a processor for running the program, wherein the method for prompting the object behavior in the embodiment is executed when the program is running.

[0120] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the object behavior prompting method in the embodiment.

[0121] According to an embodiment of the present invention, a computer program product is further provided. The computer program product includes a computer program. When the computer program is executed by a processor, the method for prompting the object behavior in the embodiment is implemented.

[0122] According to an embodiment of the present invention, a computer program product is also provided, including a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the object behavior prompting method in the embodiment is implemented.

[0123] According to an embodiment of the present invention, a computer program is further provided. When the computer program is executed by a processor, the object behavior prompting method in the embodiment is implemented.

[0124] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0125] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0126] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0127] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed over multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0128] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0129] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the relevant technology or the whole or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, referred to as Read-Only Memory), random access memory (RAM, referred to as Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program codes.

[0130] The above are only preferred embodiments of the present invention. It should be pointed out that, for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A method for prompting an object behavior, characterized in that: include: Determine, from the video data, a target region where a geometric outline of at least one object is located; Extracting target features of the object from a target image where the target area is located; Based on the target feature, predicting a current behavior performed by the object to obtain a prediction result, wherein the prediction result is used to represent the similarity between the current behavior and the target behavior associated with the target feature; In response to the similarity represented by the prediction result being greater than a similarity threshold, prompt information is output, wherein the prompt information is used to prompt permission to determine the current behavior as the target behavior.

2. The method according to claim 1, characterized in that Determining a target area where a geometric outline of at least one object is located from the video data includes: Determine at least one video frame data including the geometric contour from the video data; At least one target area is determined from the video frame data.

3. The method according to claim 2, characterized in that Determining at least one target area from the video frame data includes: Generating a bounding box matching the geometric contour from the video frame data; The area where the bounding box is located in the video frame data is determined as the target area.

4. The method according to claim 2, characterized in that: Extracting target features of the object from a target image where the target area is located includes: intercepting image data of the target area from the video frame data to obtain the target image; Scaling the target image; The target features are extracted from the scaled target image.

5. The method according to claim 1, characterized in that Based on the target feature, a current behavior performed by the object is predicted to obtain a prediction result, including: Performing standardization on the target features; Based on the target features after the normalization process, the current behavior is predicted to obtain the prediction result.

6. The method according to claim 5, characterized in that Predicting the current behavior based on the standardized target feature to obtain the prediction result includes: The target feature after standardization is input into the target prediction model, and the current behavior is predicted using the target prediction model to obtain the prediction result, wherein the target prediction model is obtained by training an initial prediction model using the target feature samples of the object samples and the corresponding prediction result samples, and the prediction result samples are used to represent the similarity between the historical behavior performed by the object sample and the target behavior.

7. A device for prompting object behavior, characterized in that: include: A determination unit, used to determine a target area where a geometric outline of at least one object is located from the video data; An extraction unit, used for extracting target features of the object from the target image where the target area is located; A prediction unit, configured to predict a current behavior performed by the object based on the target feature to obtain a prediction result, wherein the prediction result is used to indicate a similarity between the current behavior and a target behavior associated with the target feature; The output unit is used to output prompt information in response to the similarity represented by the prediction result being greater than a similarity threshold, wherein the prompt information is used to prompt permission to determine the current behavior as the target behavior.

8. A processor, characterized in that: The processor is used to run a program, wherein the program, when run by the processor, executes the object behavior prompting method described in any one of claims 1 to 6.

9. An electronic device, characterized in that: include: A memory storing an executable program; A processor is used to run the program, wherein the program executes the object behavior prompting method described in any one of claims 1 to 6 when running.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored executable program, wherein when the executable program is run, the device where the storage medium is located is controlled to execute the object behavior prompting method described in any one of claims 1 to 6.