Method and device for detecting behavior of disseminating advertisement, electronic equipment and storage medium
By performing target detection and positional relationship analysis on video frames, the system automatically identifies advertising distribution behavior, solving the problem of low efficiency in manual identification in existing technologies and improving detection efficiency and accuracy.
Patent Information
- Application Number
- CN202211170644.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-23
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2042-09-23
AI Technical Summary
In existing technologies, the identification of street advertising distribution mainly relies on manual identification, which results in a large investment of human resources, incomplete coverage, low identification efficiency, and inability to effectively identify objects closely related to advertising distribution.
By performing target detection on each video frame in the video to be detected, the location information of candidate objects and advertisements is obtained, the objects to be detected are filtered out, and the presence of advertising behavior is determined based on the positional relationship. The overlapping information of video frames is used to determine the occurrence of the behavior.
It enables automatic identification of advertising distribution behavior, improves detection efficiency and accuracy, reduces reliance on large amounts of behavioral video, and reduces the investment of human resources.
Smart Images

Figure CN115471872B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, and particularly relates to a behavior detection method and device for distributing advertisements, an electronic device and a storage medium. BACKGROUND
[0002] In the modern commercial society, most of the commodity and service information is delivered through advertisements. Merchants can improve their profits through advertising. However, some merchants take improper channels to carry out advertising for economic interests, and the phenomenon of randomly distributing advertising leaflets occurs frequently. Distributing advertising leaflets in streets with heavy traffic not only seriously affects the traffic order, but also causes environmental pollution and affects the city appearance because the advertising leaflets are easily discarded.
[0003] In recent years, the behavior of distributing advertisements on the street is severely cracked down. At present, the behavior of distributing advertisements is mainly identified by manual recognition, and then the relevant illegal personnel are cracked down. However, the above-mentioned method for identifying the behavior of distributing advertisements needs to invest huge human resources, and is not comprehensive and has low identification efficiency. SUMMARY
[0004] The embodiment of the present application provides a behavior detection method and device for distributing advertisements, an electronic device and a storage medium, so as to improve the behavior detection efficiency.
[0005] The embodiment of the present application provides a behavior detection method for distributing advertisements, comprising:
[0006] Respectively performing target detection on each video frame in a to-be-detected video to obtain candidate position information of a candidate object and advertising position information of an advertisement in each video frame;
[0007] Respectively determining the position relationship between the candidate object and the advertisement in each video frame based on the candidate position information and the advertising position information corresponding to each video frame;
[0008] Based on the determined position relationship, screening a to-be-detected object from the candidate object, and taking the remaining candidate object as a distributing object;
[0009] Based on the position relationship between the to-be-detected object and the corresponding advertisement, determining whether the to-be-detected object has the behavior of distributing the advertisement.
[0010] In an optional implementation, the screening of the to-be-detected object from the candidate object based on the determined position relationship and the taking of the remaining candidate object as a distributing object comprise:
[0011] For each candidate object, the following operations are respectively performed:
[0012] If the number of target video frames corresponding to the one candidate object is greater than the preset frame threshold, the one candidate object is determined as the to-be-detected object, wherein the position relationship between the candidate object and the advertisement in the target video frame is overlapping.
[0013] Otherwise, the one candidate object is divided into the scattering object.
[0014] In an optional implementation, the position relationship between the candidate object and the advertisement in each video frame is determined based on the candidate position information and the advertisement position information corresponding to each video frame, including:
[0015] The following operations are respectively performed on the candidate object in each video frame:
[0016] For one video frame, object key point detection is performed on a video frame region containing the candidate object in the video frame to obtain position information of each target key point in the video frame region, and the video frame region is determined according to the candidate position information;
[0017] The position information of a target part of the candidate object is determined based on the position information of each target key point.
[0018] The position relationship between the candidate object and the advertisement is determined based on the position information of the target part and the advertisement position information.
[0019] In an optional implementation, if it is determined that the to-be-detected object has the behavior of scattering the advertisement, the method further includes:
[0020] The number of times that the to-be-detected object scatters the advertisement is determined based on the number information of the scattering objects whose position relationship with the advertisement is overlapping.
[0021] In an optional implementation, if it is determined that the to-be-detected object has the behavior of scattering the advertisement, the method further includes:
[0022] The to-be-associated advertisements in the video frames are obtained based on the determined position relationships, and the to-be-associated advertisements are the advertisements whose position relationship with the candidate object in the corresponding video frame is non-overlapping.
[0023] The to-be-associated advertisements are associated with the to-be-detected object, and the advertisement abandonment level corresponding to the to-be-detected object is determined based on the number information of the to-be-associated advertisements associated with the to-be-detected object.
[0024] In an optional implementation, the target detection is performed on each video frame in the video to be detected to obtain candidate position information of a candidate object and advertisement position information of an advertisement in the video frame, and the target detection comprises:
[0025] The following operations are performed on each video frame respectively:
[0026] For a video frame, target detection is performed on the video frame, each detected target is taken as a first target, and first position information of each first target is obtained; each first target is an object or an advertisement in the corresponding video frame;
[0027] Based on the first position information and second position information of each second target in a current target set in the corresponding video frame, the first target and the second target are matched, the current target set comprises targets detected in a video frame before the video frame;
[0028] If the matching is successful and the first target is an object, the first target is taken as a candidate object, and the first position information is taken as candidate position information of the candidate object;
[0029] If the matching is successful and the first target is an advertisement, the first target is taken as an advertisement, and the first position information is taken as advertisement position information of the advertisement.
[0030] In an optional implementation, the method further comprises:
[0031] The first position information of the first target matched successfully is taken as second position information of the corresponding second target in the video frame; and / or,
[0032] The first target matched unsuccessfully is added to the object set as a new second target.
[0033] In an optional implementation, the method further comprises:
[0034] For a second target, if there is no first target matched successfully with the second target within a preset number of frames, the second target is removed from the object set.
[0035] In an optional implementation, before the position relationship between the candidate object and the advertisement in the corresponding video frame is determined based on the candidate position information and the advertisement position information of each video frame respectively, the method further comprises:
[0036] Target detection is performed on each video frame to obtain a first confidence of the candidate object in the video frame, and the candidate object with a first confidence lower than a first confidence threshold is removed; and / or,
[0037] respectively, to obtain a second confidence of the advertisement in each video frame, and remove the advertisement with a second confidence lower than a second confidence threshold.
[0038] The embodiment of the application provides a behavior detection device for spreading an advertisement, which comprises:
[0039] A detection unit is configured to perform target detection on each video frame in a video to be detected respectively, to obtain candidate position information of a candidate object and advertisement position information of an advertisement in each video frame.
[0040] A first determination unit is configured to determine a position relationship between the candidate object and the advertisement in a corresponding video frame based on the candidate position information and the advertisement position information corresponding to each video frame respectively.
[0041] A screening unit is configured to screen a to-be-detected object from the candidate object based on the determined position relationship, and take the remaining candidate object as a spreading object.
[0042] A second determination unit is configured to determine whether the to-be-detected object has a behavior of spreading the advertisement based on the position relationship between the to-be-detected object and the corresponding advertisement.
[0043] Optionally, the screening unit is specifically configured to:
[0044] The following operations are performed respectively for each candidate object:
[0045] If the number of target video frames corresponding to one candidate object is greater than a preset frame number threshold, the one candidate object is taken as the to-be-detected object, wherein the position relationship between the candidate object and the advertisement in the target video frame is overlapping.
[0046] Otherwise, the one candidate object is divided into the spreading object.
[0047] Optionally, the first determination unit is specifically configured to:
[0048] The following operations are performed respectively for each candidate object in each video frame:
[0049] For one video frame, object key point detection is performed on a video frame region containing the candidate object in the video frame, to obtain position information of each target key point in the video frame region, wherein the video frame region is determined according to the candidate position information.
[0050] The position information of a target part of the candidate object is determined based on the position information of each target key point.
[0051] determine a position relationship between the candidate object and the advertisement based on the position information of the object and the position information of the advertisement.
[0052] Optionally, if it is determined that the to-be-detected object has the behavior of distributing the advertisement, the apparatus further includes a third determination unit configured to:
[0053] determine the number of times that the to-be-detected object distributes the advertisement based on the position relationship between the to-be-detected object and the advertisement.
[0054] Optionally, if it is determined that the to-be-detected object has the behavior of distributing the advertisement, the apparatus further includes an association unit configured to:
[0055] obtain a to-be-associated advertisement in each video frame based on the determined position relationship, the to-be-associated advertisement being an advertisement that has no position relationship with the candidate object in the corresponding video frame;
[0056] associate the to-be-associated advertisement with the to-be-detected object, and determine an advertisement abandonment level corresponding to the to-be-detected object based on the number of to-be-associated advertisements associated with the to-be-detected object.
[0057] Optionally, the detection unit is specifically configured to:
[0058] perform the following operations on each video frame respectively:
[0059] for one video frame, perform target detection on the one video frame, take each detected target as a first target, and obtain first position information of each first target; each first target is an object or an advertisement in the corresponding video frame;
[0060] perform matching on the first targets and second targets based on the first position information and second position information of the second targets in the corresponding video frame in a current target set, the current target set including targets detected in video frames before the one video frame;
[0061] if the matching is successful and the first target is an object, take the first target as a candidate object, and take the first position information as candidate position information of the candidate object;
[0062] if the matching is successful and the first target is an advertisement, take the first target as an advertisement, and take the first position information as advertisement position information of the advertisement.
[0063] Optionally, the detection unit is further configured to:
[0064] add the first target whose matching fails to the object set as a new second target.
[0065] add the first target whose matching fails to the object set as a new second target.
[0066] Optionally, the detection unit is further configured to:
[0067] For one second target, if there is no first target that matches the one second target successfully within a preset number of frames, the one second target is removed from the object set.
[0068] Optionally, the apparatus further comprises a removal unit configured to:
[0069] perform target detection on each of the video frames to obtain a first confidence of a candidate object in each of the video frames, and remove a candidate object whose first confidence is lower than a first confidence threshold; and / or,
[0070] perform target detection on each of the video frames to obtain a second confidence of an advertisement in each of the video frames, and remove an advertisement whose second confidence is lower than a second confidence threshold.
[0071] An electronic device provided by an embodiment of the present application includes a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of any of the above-mentioned behavior detection methods of spreading advertisements.
[0072] An embodiment of the present application provides a computer readable storage medium, which includes a computer program, and when the computer program runs on an electronic device, the computer program is used to make the electronic device execute the steps of any of the above-mentioned behavior detection methods of spreading advertisements.
[0073] An embodiment of the present application provides a computer program product, which includes a computer program, and the computer program is stored in a computer readable storage medium; when a processor of an electronic device reads the computer program from the computer readable storage medium, the processor executes the computer program, so that the electronic device executes the steps of any of the above-mentioned behavior detection methods of spreading advertisements.
[0074] The present application has the following beneficial effects:
[0075] The application embodiment provides a behavior detection method and device for distributing advertisements, an electronic device and a storage medium. The method comprises the following steps: performing target detection on each video frame in a to-be-detected video to obtain candidate position information of a candidate object and advertisement position information of an advertisement in each video frame; then, based on the candidate position information and the advertisement position information corresponding to each video frame, screening a to-be-detected object from the candidate object, and taking the remaining candidate object as a distributing object; finally, determining whether the to-be-detected object has a behavior of distributing advertisements based on the number information of the distributing objects that have an overlapping relationship with the advertisement. In this way, the behavior of distributing advertisements can be automatically identified, the detection efficiency of the behavior of distributing advertisements is improved, and the accuracy of the behavior detection is improved.
[0076] Other features and advantages of the present application will be set forth in the following description, and in part will become apparent from the description, or can be learned by practice of the present application. The objects and other advantages of the present application will be realized and attained by the structure particularly pointed out in the written description and claims hereof as well as the appended drawings. BRIEF DESCRIPTION OF DRAWINGS
[0077] The accompanying drawings, which are included to provide a further understanding of the present application and are incorporated in and constitute a part of this application, illustrate embodiments of the present application and serve to explain the present application. In the drawings:
[0078] Figure 1 An optional schematic diagram of an application scenario in the embodiments of the present application;
[0079] Figure 2 An implementation flowchart of a behavior detection method for distributing advertisements in the embodiments of the present application;
[0080] Figure 3 A schematic diagram of a video frame in a to-be-detected video in the embodiments of the present application;
[0081] Figure 4 A flowchart of a position relationship determination method in the embodiments of the present application;
[0082] Figure 5 A schematic diagram of a calculation principle of a hand region in the embodiments of the present application;
[0083] Figure 6 A flowchart of a position relationship determination method in the embodiments of the present application;
[0084] Figure 7 A structural schematic diagram of a behavior detection system in the embodiments of the present application;
[0085] Figure 8A work flow diagram of an alarm logic judgment module in an embodiment of the present application;
[0086] Figure 9 A structural diagram of a behavior detection device for distributing advertisements in an embodiment of the present application;
[0087] Figure 10 A hardware component structural diagram of an electronic device applying an embodiment of the present application;
[0088] Figure 11 A hardware component structural diagram of another electronic device applying an embodiment of the present application. DETAILED DESCRIPTION
[0089] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be described below in connection with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments described in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0090] Some concepts involved in the embodiments of the present application will be introduced below.
[0091] Candidate object: refers to an object obtained by target detection on each video frame in a to-be-detected video, candidate position information of the candidate object includes position information of the candidate object in each video frame, in the embodiments of the present application, the candidate object is mainly described as a person in the to-be-detected video.
[0092] To-be-detected object: refers to a candidate object carrying an advertisement determined according to a position relationship between the candidate object and the advertisement, that is, the to-be-detected object may have a behavior of distributing the advertisement, after the to-be-detected object is screened out, whether the to-be-detected object has the behavior of distributing the advertisement needs to be determined according to a position relationship between the advertisement corresponding to the to-be-detected object and a distributing object.
[0093] Distributing object: refers to an object in the candidate object other than the to-be-detected object, in a video frame, there may be multiple candidate objects, the distributing object can be a candidate object having a position relationship with the advertisement as no overlap, or a candidate object having a position relationship with the advertisement as having overlap, but the number of video frames having overlap is less than a preset frame threshold, that is, the candidate object does not carry the advertisement for a long time, the candidate object can be considered as not having the behavior of distributing the advertisement, and the candidate object can be an object of distributing the advertisement by the to-be-detected object, so it is called the distributing object.
[0094] To-be-associated advertisement: refers to an advertisement whose position relationship with a candidate object in a video frame is non-overlapping, i.e., there is no overlapping between the to-be-associated advertisement and the position of the candidate object, for example, the to-be-associated advertisement is an abandoned advertisement.
[0095] Advertisement abandonment level: when a to-be-detected object spreads an advertisement, part of the spreading object will abandon the advertisement on the ground, thereby causing environmental pollution. The environmental pollution caused is due to the spreading behavior of the to-be-detected object, which can be represented by the advertisement abandonment level to form a complete snapshot evidence chain of the spreading behavior of the to-be-detected object.
[0096] In the embodiments of the present application, the term "and / or" describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone. The character " / " generally represents an "or" relationship between the front and rear associated objects.
[0097] The design idea of the embodiments of the present application will be briefly introduced as follows:
[0098] In the modern commercial society, most of the commodity and service information is delivered through advertisements, and merchants can improve their profits through advertisement publicity. However, some merchants take unfair means for economic interests, and the phenomenon of randomly distributing advertising leaflets is rampant. Distributing advertising leaflets in streets with heavy traffic not only seriously affects the traffic order, but also causes environmental pollution and affects the city appearance because the advertising leaflets are easily discarded at will.
[0099] In recent years, relevant departments in various places have been cracking down on the behavior of distributing advertisements on the street. In related technologies, the behavior detection is mainly performed in the following ways:
[0100] Method one: responsible personnel patrol the place where the behavior of distributing advertisements may exist, and manually identify the behavior of distributing advertisements, and then crack down on the relevant illegal personnel.
[0101] However, based on the method one to identify the behavior of distributing advertisements, a huge amount of human resources needs to be invested, and the coverage is not comprehensive, and the identification efficiency is low.
[0102] Method two: obtaining a behavior observation video of a target object, training a target behavior recognition model using the behavior observation video, and identifying a target behavior based on the target behavior recognition model.
[0103] However, based on the method two to identify the behavior of distributing advertisements, the internal characteristics of the behavior of distributing advertisements are not fully utilized to predict human behavior, and objects closely related to the behavior of distributing advertisements cannot be identified, and this method excessively relies on a large number of behavior video clips for training samples, and the accuracy is low for small sample video materials.
[0104] The embodiment of the present application provides a behavior detection method and device for distributing advertisements, an electronic device and a storage medium. The candidate position information of candidate objects and the advertisement position information of advertisements are obtained by performing target detection on each video frame in a to-be-detected video respectively; then, the to-be-detected objects are selected from the candidate objects based on the candidate position information and the advertisement position information corresponding to each video frame, and the remaining candidate objects are taken as distribution objects; finally, whether the to-be-detected objects have the behavior of distributing advertisements is determined based on the position relationship between the to-be-detected objects and the advertisements, that is, the number information of the distribution objects that have the overlapping relationship with the advertisements. In the foregoing manner, the behavior of distributing advertisements can be automatically recognized, the detection efficiency of the behavior of distributing advertisements is improved, the accuracy of behavior detection is improved, and a large number of behavior videos do not need to be used as training samples, so that the difficulty of behavior detection is reduced.
[0105] The preferred embodiments of the present application are described below in combination with the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application, and the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0106] As shown in FIG. 1, it is an application scenario diagram of the embodiment of the present application. The application scenario diagram includes two terminal devices 110 and one server 120. Figure 1
[0107] In the embodiment of the present application, the terminal device 110 includes but is not limited to a mobile phone, a tablet computer, a notebook computer, a desktop computer, an electronic book reader, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal and the like. The behavior detection related client can be installed on the terminal device, which can be software (such as a browser, a behavior detection software and the like), a webpage, an applet and the like. The server 120 is a background server corresponding to the software or the webpage, the applet and the like, or a server specially used for behavior detection, which is not limited in the present application. The server 120 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks (CDN), and big data and artificial intelligence platforms and the like basic cloud computing services.
[0108] It should be noted that the behavior detection method for spreading advertisements in the embodiments of the present application can be executed by an electronic device, which can be the server 120 or the terminal device 110, that is, the method can be executed by the server 120 or the terminal device 110 alone, or by the server 120 and the terminal device 110 together. For example, when executed by the server 120 and the terminal device 110 together, the terminal device 110 acquires the to-be-detected video and sends the to-be-detected video to the server 120, the server 120 respectively performs target detection on each video frame in the to-be-detected video to obtain candidate position information of a candidate object and advertisement position information of an advertisement in each video frame; respectively based on the candidate position information and the advertisement position information corresponding to each video frame, determines a position relationship between the candidate object and the advertisement in the corresponding video frame; based on the determined position relationships, filters out a to-be-detected object from the candidate objects, and takes the remaining candidate objects as spreading objects; based on the position relationship between the to-be-detected object and the corresponding advertisement, determines whether the to-be-detected object has the behavior of spreading advertisements, and the server 120 sends information of the to-be-detected object having the behavior of spreading advertisements to the terminal device 110, so that the user of the terminal device 110 takes corresponding measures to manage the behavior of spreading advertisements.
[0109] In an optional implementation, the terminal device 110 and the server 120 can communicate through a communication network.
[0110] In an optional implementation, the communication network is a wired network or a wireless network.
[0111] It should be noted that, Figure 1 It should be noted that,
[0112] In the embodiments of the present application, when the number of servers is multiple, the multiple servers can form a blockchain, and the servers are nodes on the blockchain; the behavior detection method for spreading advertisements disclosed in the embodiments of the present application, wherein the to-be-detected video involved can be saved on the blockchain.
[0113] In addition, the embodiments of the present application can be applied to various scenes, not only including behavior detection scenes, but also including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving and the like.
[0114] The behavior detection method for spreading advertisements provided by the exemplary embodiments of the present application will be described below in combination with the application scenarios described above and with reference to the accompanying drawings. It should be noted that the above-mentioned application scenarios are only shown for the purpose of facilitating understanding of the spirit and principles of the present application, and the embodiments of the present application are not limited in this respect.
[0115] Referring toFigure 2 Fig. 8 shows an embodiment of a method for detecting a behavior of distributing an advertisement provided by the present application, and the embodiment is illustrated as a flowchart of the method for detecting a behavior of distributing an advertisement. The embodiment is taken as an example in which the execution subject is a server. The embodiment includes the following steps S21-S24.
[0116] S21: The server detects a target in each video frame in the video to be detected to obtain candidate position information of a candidate object and advertisement position information of an advertisement in each video frame.
[0117] The video to be detected is captured by installing a camera in an area in which a behavior of distributing an advertisement needs to be detected, for example, a camera is installed near a city sidewalk, and the camera needs to clearly capture a real-time image of a passerby.
[0118] Referring to Fig. 2, a schematic diagram of a video frame in the video to be detected is shown. A closed road detection rule area is set in a monitoring image, as shown by a black thick line frame in Fig. 2. The road detection rule area in the video frame is detected. After the video to be detected is captured, the video to be detected is decoded by using a video coding and decoding technology to detect a target based on real-time image data. Figure 3 Figure 3
[0119] The target in the video to be detected is detected to obtain candidate position information of a candidate object. The candidate position information of the candidate object refers to position information of the candidate object in each video frame, for example, candidate position information of a candidate object 1 includes position information of the candidate object 1 in the first to fifth video frames. The candidate position information can be represented by upper left corner coordinates (x1, y1) and lower right corner coordinates (x2, y2) of a bounding box (coordinate box) of the candidate object in the video frame, or can be represented by center point coordinates of the bounding box and width and height of the bounding box. The present application does not make a specific limitation here. In the embodiment of the present application, the candidate object is mainly taken as a person in the video to be detected, and a black thin solid line frame in Fig. 3 represents a bounding box of a human body in a video frame. Figure 3
[0120] The advertisement is obtained by detecting a target in each video frame in the video to be detected. The advertisement includes any form of advertisement such as a leaflet, a book, and a manual, and the present application does not make a specific limitation here. The advertisement position information of the advertisement refers to position information of the advertisement in each video frame, and a dashed line frame in Fig. 4 represents a bounding box of a human body in a video frame. Figure 3
[0121] When performing target detection on the to-be-detected video, the target detection model A trained in the following manner can be used: first, collect video sequences of urban sidewalks, obtain pictures as training data sets by acquiring video frames, then prepare training data sets for the target detection network A, and label the human body and advertisements as positive samples, so that the target detection model A can detect both the human body and the advertisements. After obtaining the trained target detection model A, the target detection module A detects the human body and the advertisements in the video frame to obtain the target coordinate frame Rect (position information) including the upper left corner coordinate point (x1, y1), the lower right corner coordinate (x2, y2), the confidence score ConfOD (a decimal between 0 and 1), and the target type TypeOD (category: 1 human body, 2 advertisement).
[0122] After obtaining the first confidence of the candidate object in each video frame and the second confidence of the advertisement in each video frame by performing target detection on each video frame respectively, the target can be filtered based on the confidence score of the target. When the target is a candidate object, the candidate object with a first confidence lower than a first confidence threshold is removed; when the target is an advertisement, the advertisement with a second confidence lower than a second confidence threshold is removed.
[0123] Specifically, for example, the confidence score of the candidate object 2 in the first video frame is 0.8, the confidence score in the second video frame is 0.2, and the confidence score in the third video frame is 0.7. The first confidence threshold is 0.5, indicating that the candidate object 2 detected in the second video frame may be a false detection, and the candidate object 2 in the second frame is removed. The candidate position information of the candidate object 2 only includes the position information of the candidate object 2 in the first frame and the third frame; correspondingly, the confidence score of the to-be-detected object 1 in the first video frame is 0.6, the confidence score in the second video frame is 0.7, and the confidence score in the third video frame is 0.3. The second confidence threshold is 0.5, indicating that the to-be-detected object 1 detected in the third video frame may be a false detection, and the to-be-detected object 1 in the third frame is removed. The candidate position information of the to-be-detected object 1 only includes the position information of the to-be-detected object 1 in the first frame and the second frame.
[0124] Based on the above manner, candidate objects with low confidence and advertisements with low confidence can be removed, false detection targets can be filtered, and the accuracy of behavior detection can be improved.
[0125] When performing target detection on each video frame, only the position information of the target contained in each video frame can be obtained, and it is impossible to distinguish the position information of which target, so the position information of the same target in different video frames can be associated in the following manner.
[0126] In an alternative embodiment, the following operations can be performed on each video frame in step S21:
[0127] For a video frame, target detection is performed on the video frame, each detected target is taken as a first target, and first position information of each first target is obtained; based on the first position information and second position information of each second target in the current target set in the corresponding video frame, each first target and each second target are matched;
[0128] If the matching is successful and the first target is an object, the first target is taken as a candidate object, and the first position information is taken as candidate position information of the candidate object;
[0129] If the matching is successful and the first target is an advertisement, the first target is taken as an advertisement, and the first position information is taken as advertisement position information of the advertisement.
[0130] In the method, each first target is an object or an advertisement in a corresponding video frame; the current target set includes targets detected in a video frame before the video frame; the first position information of the first target represents a bounding box (detection coordinate box) of the first target in the video frame; the second position information of the second target represents a bounding box (tracking coordinate box) of the second target in the video frame; when the first target and the second target are matched, an intersection over union (IOU) of the detection coordinate box of the first target in the current frame and the tracking coordinate box of the second target in the previous frame is calculated; when the IOU of the first target and a certain second target is greater than an intersection over union threshold, the matching is successful; when the first target can be matched with multiple second targets, the first target is matched with the second target with the greatest IOU.
[0131] Taking a video frame as the second video frame in the video to be detected as an example, target 1 and target 2 are detected, at this time, the current target set includes targets detected in the first video frame, including target A, target B and target C, the IOU of target 1 and target A, target B and target C is calculated respectively, it is determined that the IOU of target 1 and target B is greater than the intersection over union threshold, target 1 and target B are matched successfully, target 1 is an object, target 1 is taken as candidate object 1, and the first position information of target 1 is taken as the candidate position information of candidate object 1 in the second video frame, that is, the candidate position information of candidate object 1 in the first frame is the second position information of target B, and the candidate position information of candidate object 1 in the second frame is the first position information of target 1, the tracking trajectory of candidate object 1 in the video to be detected is tracked through the above method, and each candidate object has position information in at least two video frames. The matching process of target 2 is the same as above, and will not be described here.
[0132] Based on the above manner, the target obtained by the target detection of the current frame is matched with the existing tracking target, so that the false detection target appearing only in one frame can be filtered, and the accuracy of behavior detection is improved.
[0133] In an optional embodiment, if the first target can be successfully matched with the second target, the first position information of the first target matched successfully is taken as the second position information of the corresponding second target in a video frame; if the first target fails to be matched with each second target, the first target matched unsuccessfully is added to the object set as a new second target.
[0134] Specifically, if the one video frame is the third video frame, the first target 2 is successfully matched with the second target 3, the first position information of the first target 2 is taken as the second position information of the second target 3 in the third video frame, and if the first target 3 fails to be matched with each second target, the first target 3 is added to the object set as a new second target.
[0135] In addition, in the object set, for one second target, if there is no first target matched successfully with the one second target within a preset frame number, the one second target is removed from the object set.
[0136] The preset frame number can be 12 frames, that is, if a second target cannot be successfully matched with a first target for more than 12 frames, the second target is removed from the object set.
[0137] Specifically, the manner of tracking by detection can be adopted to obtain the candidate position information of the candidate object and the advertisement position information of the advertisement in each video frame, the target in the current frame is matched with the tracking trajectory of the existing tracking target through a matching algorithm, and a new tracking trajectory of the tracking target is formed, wherein the tracking target has four state bits: Create, Update, Lost, and Delete.
[0138] The matching process is to calculate the IOU of the detection coordinate frame of the first target in the current frame and the tracking coordinate frame of the previous frame. When the IOU is greater than a threshold, it is considered as a successful match. When the detection coordinate frame of the first target can be successfully matched with multiple tracking coordinate frames, the tracking frame with the largest IOU is taken. In the matching process, the following situations may occur: when the first target can be matched with an existing tracking target, the tracking trajectory of the tracking target is updated, and the state bit of the tracking target is Update; the current first target cannot find a tracking target to match, which means that the target is a newly appearing tracking target, a tracking trajectory is created for the target, the state bit of the tracking target is Create, and the target is numbered (i.e., ID); if a tracking target (i.e., the second target) does not have a detection target (i.e., the first target) in the current frame to match, it means that the tracking target is lost in the video, and the state bit of the tracking target is Lost; when the state bit of a tracking target is Lost for more than 12 frames, the state bit of the tracking target is updated to Delete, and the tracking trajectory of the tracking target is deleted. When the state bit of the tracking target is Update, it is used as a candidate object or a detected advertisement, that is, the candidate object or the advertisement has position information in at least two frames. Through the above method, false detection targets, such as targets appearing in only one frame, can be filtered out, the calculation amount is reduced, and the accuracy of behavior detection is improved.
[0139] S22: The server determines the position relationship between the candidate object and the advertisement in the corresponding video frame based on the candidate position information and the advertisement position information corresponding to each video frame, respectively;
[0140] The candidate position information of at least one candidate object and the advertisement position information of at least one advertisement corresponding to each video frame can be used to determine the position relationship between the candidate object and the advertisement in the video frame. For example, the boundary box of the candidate object in the video frame can be determined according to the candidate position information, the boundary box of the advertisement in the video frame can be determined according to the advertisement position information, and whether there is an IOU intersection between the boundary boxes can be calculated. If there is, the position relationship between the corresponding candidate object and the advertisement is overlapping; if not, the position relationship between the corresponding candidate object and the advertisement is non-overlapping.
[0141] When the candidate object is a human body, the behavior of distributing an advertisement is closely related to the key points of the human arm. Therefore, whether there is an intersection (i.e., overlapping) between the boundary box of the human hand region and the advertisement can also be used to determine the position relationship between the candidate object and the advertisement. In an optional embodiment, as shown in FIG. 8, step S22 can be implemented as the following steps: Figure 4
[0142] Step S221: For a video frame, object key point detection is performed on a video frame region containing a candidate object in the video frame to obtain position information of each target key point in the video frame region, the video frame region being determined according to candidate position information;
[0143] Step S222: Based on the position information of each target key point, position information of a target part of the candidate object is determined.
[0144] Step S223: Based on the part position information and the advertisement position information, a position relationship between the candidate object and the advertisement is determined.
[0145] Specifically, when the target category is a human body, the human body frame (i.e., the video frame region containing the candidate object) is sent into the trained human body skeleton key point network B to obtain position information of 17 key points of the human body. Since the handout advertisement behavior is closely related to the key points related to the arms of the human body, only four key points, i.e., the left wrist joint, the left elbow joint, the right wrist joint, and the right elbow joint, are taken as target key points, and the center point coordinates (x, y) of each target key point are obtained. The palm center (left and right) coordinates are predicted through the wrist and elbow center point positions. The left elbow joint coordinates are (x1, y1), the left wrist joint coordinates are (x2, y2), the left palm center coordinates (x3, y3) are obtained, and the calculation formula is as follows:
[0146]
[0147]
[0148] d2 = 0.2 * 2 * d1
[0149] Δx = 0.2 * d1 * cos θ
[0150] Δy = 0.2 * d1 * sin θ
[0151] x3 = Δx + x2
[0152] y3 = Δv + y2
[0153] wherein θ is the included angle between the wrist and the elbow, d1 is the distance between the wrist and the elbow, 0.2 is a prediction coefficient, and finally the position information (i.e., the part position information) of the left hand region H is obtained: a square region with the palm (x3, y3) as the center point and d2 as the side length, as shown in Figure 5As shown, it is a schematic diagram of a calculation principle of a hand region in an embodiment of the present application. Similarly, the position information of the right hand region can also be obtained based on the above manner, which will not be described herein. After obtaining the position information of the left hand region and the position information of the right hand region of the human body, whether the left hand region and the right hand region of the candidate object intersect with the advertisement can be calculated respectively to determine the positional relationship between the candidate object and the to-be-detected object.
[0154] The human body skeleton key point network B can be trained in the following manner: a video sequence of a city sidewalk is collected, a picture is obtained by acquiring a video frame as a training data set. The training data set for the human body skeleton key point network B is prepared, and 17 key points of the human body are labeled, so that the trained human body skeleton key point network B can predict 17 key points of the human body. As shown in Figure 3 As shown, the black points are the key points related to the human arm (elbow, wrist), i.e. the target key points.
[0155] If the bounding box of the advertisement exists the IOU intersection with the human hand region H (left, right), the flag of the person carrying the advertisement leaflet is marked as 1, and then the detection result is integrated. When the target type is a human body, the information of the human body target includes: a human body coordinate frame, a confidence score of the human body coordinate frame, a current frame carrying an advertisement leaflet flag, and a carried advertisement coordinate frame. When the target type is an advertisement leaflet, the information of the advertisement target includes: an advertisement coordinate frame and a confidence score of the advertisement coordinate frame.
[0156] Referring to Figure 6 It is a flowchart of a positional relationship determination method in an embodiment of the present application, which includes the following steps:
[0157] S61: input a video frame image of a current frame;
[0158] S62: a target detection network A performs target detection on the video frame image to obtain the position information of the target;
[0159] S63: determine whether the type of the target is a human body, if yes, execute step S64, if not, execute step S66;
[0160] S64: a human body skeleton key point network B performs key point detection on a region containing the human body to obtain the position information of the target key point;
[0161] S65: predict the position information of the hand region of the human body target based on the position information of the target key point;
[0162] S66: associate the human body target with the advertisement target with the intersection.
[0163] S23: The server screens the to-be-detected object from the candidate objects based on the determined position relationship, and takes the remaining candidate objects as the distributing objects;
[0164] Specifically, the screening is performed according to the position relationship between the candidate objects and the advertisement, the to-be-detected object is an object in which the candidate object is likely to have the behavior of distributing the advertisement, and the distributing object is an object in which the to-be-detected object distributes the advertisement.
[0165] In the embodiments of the present application, the human body target and the advertisement target are detected by using the target detection network, the wrist and elbow key points are recognized by using the human body skeleton key point network, and the hand region is predicted. The association with the advertisement target can effectively identify the behavior of distributing the advertisement, and a large number of behavior video clips do not need to be used as a training set, and the practicability is strong.
[0166] In an optional implementation, step S23 can be implemented as the following steps:
[0167] For a candidate object, if the number of target video frames corresponding to the candidate object is greater than a preset frame number threshold, the candidate object is taken as the to-be-detected object, wherein the position relationship between the candidate object and the advertisement in the target video frame is overlapping; otherwise, the candidate object is divided into the distributing object.
[0168] Specifically, when the candidate objects are screened, if the number of video frames in which the position relationship between a candidate object and the advertisement is overlapping is greater than a preset frame number threshold, the candidate object is taken as the to-be-detected object. The position relationship between the candidate object and the advertisement being overlapping can represent that the candidate object carries the advertisement in the corresponding video frame, and the number of target video frames represents the cumulative time in which the candidate object carries the advertisement. If the cumulative time is greater than the preset frame number threshold, it is considered that the candidate object carries the advertisement for a long time, and the behavior of distributing the advertisement is likely to exist.
[0169] For example, when the human body target (candidate object) carries the advertisement leaflet flag flag is 1, that is, the position relationship between the human body target and the advertisement in the current frame is overlapping, the cumulative time t in which the human body target carries the advertisement is increased by 1, and it is determined whether t is greater than a threshold t1. If yes, it is considered that the human body target carries the advertisement leaflet for a long time, and it is necessary to further determine whether the distributing behavior exists.
[0170] S24: The server determines whether the to-be-detected object has the behavior of distributing the advertisement based on the number information of the distributing objects in which the position relationship with the corresponding advertisement is overlapping.
[0171] Specifically, after the to-be-detected object is determined in step S23, the number of times of distributing the advertisement by the to-be-detected object can be determined according to the positional relationship between the advertisement intersecting with the to-be-detected object and the distribution object, and if the number of times of distribution is greater than a number threshold, it can be determined that the to-be-detected object has the behavior of distributing the advertisement.
[0172] For example, the coordinate frame of the advertisement corresponding to the to-be-detected object is subjected to IOU intersection calculation with the coordinate frames of the other respective distribution objects, if there is a distribution object with intersection, the number information is increased by 1, and when the number information is greater than a set number threshold n1, it is determined that the to-be-detected object has the behavior of distributing the advertisement.
[0173] In the embodiments of the present application, the candidate position information of the candidate objects and the advertisement position information of the advertisements in each video frame in the to-be-detected video are obtained by respectively performing target detection on each video frame; then, based on the candidate position information and the advertisement position information corresponding to each video frame, the to-be-detected object is selected from the candidate objects, and the remaining candidate objects are taken as distribution objects; finally, based on the number information of the distribution objects overlapping with the advertisements, it is determined whether the to-be-detected object has the behavior of distributing the advertisement, and through the above-mentioned manner, the behavior of distributing the advertisement can be automatically recognized, the detection efficiency of the behavior of distributing the advertisement is improved, and the accuracy of behavior detection is improved.
[0174] In an optional implementation, if it is determined that the to-be-detected object has the behavior of distributing the advertisement, the number of times of distributing the advertisement by the to-be-detected object is determined based on the number information of the distribution objects overlapping with the advertisements.
[0175] Specifically, the number information of the distribution objects overlapping with the advertisements is equal to the number of times of distributing the advertisement by the to-be-detected object, that is, if the number information of the distribution objects overlapping with the advertisements is 10, it indicates that the to-be-detected object distributes the advertisement to 10 distribution objects, and the number of times of distributing the advertisement is 10.
[0176] In an optional implementation, if it is determined that the to-be-detected object has the behavior of distributing the advertisement, based on the determined positional relationship, the to-be-associated advertisements in each video frame are obtained, and the to-be-associated advertisements are associated with the to-be-detected object, and based on the number information of the to-be-associated advertisements associated with the to-be-detected object, the advertisement abandonment level corresponding to the to-be-detected object is determined.
[0177] The to-be-associated advertisements are the advertisements without overlapping with the candidate objects in the corresponding video frames, and the to-be-associated advertisements are the abandoned advertisements. Since the to-be-associated advertisements are abandoned, it will cause environmental pollution, so the number information of the to-be-associated advertisements associated with the to-be-detected object needs to be counted to determine the advertisement abandonment level (i.e., the flyer abandonment level), and the advertisement abandonment level needs to be reported as the evidence of the behavior of distributing the advertisement by the to-be-detected object.
[0178] Based on the above manner, after determining that the to-be-detected object exists in the behavior of distributing advertisements, the number of times of distributing advertisements and the level of abandoning advertisements of the to-be-detected object are identified, a complete snapshot evidence chain is formed, manual identification is not needed, the detection of the behavior of distributing advertisements and the collection of evidence are completed, and the behavior detection efficiency is improved.
[0179] The behavior detection method of distributing advertisements in the application is used for the detection of the behavior of distributing advertisements on the street, and the application is referred to Figure 7 FIG. 1 is a structural schematic diagram of a behavior detection system in an embodiment of the application. The behavior detection system takes a deep learning method as a basic method and includes the following modules.
[0180] The data acquisition module sets a monitoring area and acquires city sidewalk monitoring data.
[0181] The deep network identification module detects a human target and an advertisement target from the city sidewalk monitoring data acquired by the data acquisition module, identifies human skeleton key points of the human target, obtains the key points of the human target, predicts a hand region of the human target and associates the hand region with the carried advertisement, and sends the integrated information to the multi-target tracking module.
[0182] The multi-target tracking module actively tracks the human target and the advertisement target detected by the deep network identification module, obtains a motion trajectory and an ID of the human target and a motion trajectory and an ID of the advertisement target, and then sends the tracking result to the alarm logic judgment module.
[0183] The alarm logic judgment module judges whether the human target exists in the behavior of distributing advertisements according to the motion trajectory and the ID of the human target and the motion trajectory and the ID of the advertisement target, and outputs a behavior judgment result of distributing advertisements, a number of times of distributing advertisements, and a level of abandoning advertisements.
[0184] In the embodiment of the application, not only the behavior of distributing advertisements can be accurately identified, but also the level of abandoning advertisements and the number of times of distributing advertisements can be identified, a reported evidence chain is complete, and can be used as a law enforcement basis. The alarm logic judgment module is identified through cumulative counting, position association, and other judgment logics, and can avoid the influence of individual frame false detection results on the overall alarm accuracy.
[0185] The specific working process of the alarm logic judgment module is shown in FIG. 8, which is a working process schematic diagram of an alarm logic judgment module in an embodiment of the application and includes the following steps. Figure 8
[0186] S801: Loop the tracking result.
[0187] S802: Obtain the confidence score of the human target in the current frame, the position information in the video frame, the flag of carrying the advertising leaflet, the position information of the carried advertising, and the confidence score and position information of the advertising target in the video frame;
[0188] S803: Determine whether the category of the current target is a human target. If yes, execute step S807; if no, execute step S804;
[0189] S804: Determine whether the confidence score of the current advertising target is greater than a threshold th1. If yes, execute step S805; if no, execute step S801;
[0190] S805: Determine whether the center point coordinate of the current advertising target is within a set monitoring area. If yes, execute step S806; if no, execute step S801;
[0191] S806: Determine whether there is an IOU intersection between the current advertising target and other human targets. If yes, execute step S801; if no, execute step S813;
[0192] S807: Determine whether the confidence score of the current human target is greater than a threshold th2. If yes, execute step S808; if no, execute step S801;
[0193] S808: Determine whether the flag of carrying the advertising leaflet of the current human target is 1. If yes, execute step S809; if no, execute step S801;
[0194] S809: Add 1 to the accumulated time t of the advertising carried by the current human target;
[0195] S810: Determine whether t is greater than a threshold t1. If yes, execute step S811; if no, execute step S801;
[0196] S811: Perform IOU intersection calculation on the advertising carried by the current human target and other human targets to obtain the number of human targets with intersection as the number of times of distribution of the current human target;
[0197] S812: Determine whether the number of times of distribution is greater than a set threshold n1. If yes, execute step S813; if no, execute step S801;
[0198] S813: Determine the number of advertising targets without intersection with other human targets as the leaflet abandonment level of the current human target;
[0199] S814: Upload the current human target, the leaflet abandonment level, and the number of times of distribution to the upper platform as the alarm result and notify the law enforcement personnel.
[0200] The alarm logic judgment module in the application has progressive steps, strict logic and strong practicability. By setting the confidence score threshold, the false detection target is filtered out. By setting the cumulative time of carrying the advertisement, the alarm accuracy of distributing the advertisement is improved; by calculating the position relationship between the human body and the carried advertisement, the behavior of distributing the advertisement can be accurately identified. The human body target of distributing the advertisement, the number of distributing the advertisement and the level of abandoning the leaflet are reported to form a complete snapshot evidence chain.
[0201] Based on the same inventive concept, the application also provides a behavior detection device for distributing an advertisement. As shown in Figure 9 The behavior detection device 900 for distributing an advertisement can include:
[0202] The detection unit 901 is configured to perform target detection on each video frame in the video to be detected respectively, and obtain candidate position information of a candidate object and advertisement position information of an advertisement in each video frame;
[0203] The first determination unit 902 is configured to determine the position relationship between the candidate object and the advertisement in the corresponding video frame based on the candidate position information and the advertisement position information corresponding to each video frame respectively;
[0204] The screening unit 903 is configured to screen the to-be-detected object from the candidate object based on the determined position relationship, and take the remaining candidate object as a distribution object;
[0205] The second determination unit 904 is configured to determine whether the to-be-detected object has the behavior of distributing the advertisement based on the position relationship between the to-be-detected object and the corresponding advertisement.
[0206] Optionally, the screening unit 903 is specifically configured to:
[0207] For each candidate object, the following operations are performed respectively:
[0208] For a candidate object, if the number of target video frames corresponding to the candidate object is greater than a preset frame number threshold, the candidate object is taken as the to-be-detected object, wherein the position relationship between the candidate object and the advertisement in the target video frame is overlapping;
[0209] Otherwise, the candidate object is divided into a distribution object.
[0210] Optionally, the first determination unit 902 is specifically configured to:
[0211] For each candidate object in each video frame, the following operations are performed respectively:
[0212] For a video frame, performing object key point detection on a video frame region containing a candidate object in the video frame to obtain position information of each target key point in the video frame region, the video frame region being determined according to candidate position information;
[0213] Based on the position information of each target key point, determining part position information of a target part of the candidate object;
[0214] Based on the part position information and the advertisement position information, determining a position relationship between the candidate object and the advertisement.
[0215] Optionally, if it is determined that the to-be-detected object has the behavior of distributing the advertisement, the apparatus further includes a third determination unit 905, configured to:
[0216] Based on the position relationship between the to-be-detected object and the advertisement, determining the number of times of distributing the advertisement by the to-be-detected object.
[0217] Optionally, if it is determined that the to-be-detected object has the behavior of distributing the advertisement, the apparatus further includes an association unit 906, configured to:
[0218] Based on the determined position relationships, obtaining a to-be-associated advertisement in each video frame, the to-be-associated advertisement being an advertisement that has no overlapping with the candidate object in a corresponding video frame;
[0219] Associating the to-be-associated advertisement with the to-be-detected object, and determining a discarded advertisement level corresponding to the to-be-detected object based on the number of to-be-associated advertisements associated with the to-be-detected object.
[0220] Optionally, the detection unit 901 is specifically configured to:
[0221] Respectively performing the following operations on each video frame:
[0222] For a video frame, performing target detection on the video frame, taking each detected target as a first target, and obtaining first position information of each first target; each first target being an object or an advertisement in a corresponding video frame;
[0223] Based on the first position information and second position information of each second target in a current target set in the corresponding video frame, matching each first target and each second target, the current target set including targets detected in a video frame before the video frame;
[0224] If the matching is successful and the first target is an object, taking the first target as a candidate object, and taking the first position information as candidate position information of the candidate object;
[0225] If the matching is successful and the first target is an advertisement, taking the first target as the advertisement, and taking the first position information as advertisement position information of the advertisement.
[0226] Optionally, the detection unit 901 is further configured to:
[0227] add the first target whose matching is successful to the object set as a new second target.
[0228] add the first target whose matching is successful to the object set as a new second target.
[0229] Optionally, the detection unit 901 is further configured to:
[0230] remove a second target from the object set if there is no first target matching the second target successfully within a preset number of frames.
[0231] Optionally, the apparatus further includes a removal unit 907 configured to:
[0232] perform target detection on each video frame respectively to obtain a first confidence of a candidate object in each video frame, and remove the candidate object whose first confidence is lower than a first confidence threshold; and / or
[0233] perform target detection on each video frame respectively to obtain a second confidence of an advertisement in each video frame, and remove the advertisement whose second confidence is lower than a second confidence threshold.
[0234] For ease of description, each part is described as a module (or unit) according to function. Of course, the functions of each module (or unit) can be implemented in the same or multiple software or hardware in the implementation of the present application.
[0235] Those skilled in the art can understand that each aspect of the present application can be implemented as a system, a method or a program product. Therefore, each aspect of the present application can be specifically implemented as follows: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software aspects, which can be collectively referred to as "circuitry", "module" or "system".
[0236] Based on the same inventive concept as the method embodiments described above, an electronic device is also provided in the embodiments of the present application. In one embodiment, the electronic device can be a server, such as the server 120 as shown in Figure 1 In this embodiment, the structure of the electronic device can be as shown in Figure 10 , which includes a memory 1001, a communication module 1003, and one or more processors 1002.
[0237] The memory 1001 is configured to store a computer program executed by the processor 1002. The memory 1001 can mainly include a program storage area and a data storage area. The program storage area can store an operating system and programs required for running an instant messaging function, etc. The data storage area can store various instant messaging information and operation instruction sets, etc.
[0238] The memory 1001 can be a volatile memory such as a random-access memory (RAM), or a non-volatile memory such as a read-only memory, a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD), or any other medium capable of carrying or storing desired computer programs in the form of instructions or data structures and capable of being accessed by a computer, but is not limited to this. The memory 1001 can be a combination of the above memories.
[0239] The processor 1002 can include one or more central processing units (CPUs) or digital processing units, etc. The processor 1002 is configured to implement the above-mentioned behavior detection method for spreading advertisements when invoking the computer program stored in the memory 1001.
[0240] The communication module 1003 is configured to communicate with terminal devices and other servers.
[0241] The specific connection medium between the above-mentioned memory 1001, communication module 1003 and processor 1002 is not limited in the embodiments of the present application. In the embodiments of the present application, the memory 1001 and the processor 1002 are connected through a bus 1004, and the bus 1004 is described by a thick line in the embodiments of the present application. The connection mode between other components is only schematically described, and is not limited. The bus 1004 can be divided into an address bus, a data bus, a control bus, etc. For the convenience of description, only one thick line is used to describe the bus 1004, but it does not mean that there is only one bus or only one type of bus. Figure 10 Figure 10 Figure 10
[0242] The memory 1001 stores a computer storage medium, and the computer storage medium stores computer executable instructions. The computer executable instructions are used to implement the behavior detection method for spreading advertisements in the embodiments of the present application. The processor 1002 is configured to execute the above-mentioned behavior detection method for spreading advertisements, as shown in FIG. 8. Figure 2
[0243] In another embodiment, the electronic device can also be other electronic devices, such as Figure 1 The terminal device 110 as shown. In this embodiment, the structure of the electronic device can be as shown, including: a communication component 1110, a memory 1120, a display unit 1130, a camera 1140, a sensor 1150, an audio circuit 1160, a Bluetooth module 1170, a processor 1180 and the like. Figure 11
[0244] The communication component 1110 is used for communication with the server. In some embodiments, a wireless fidelity (WiFi) module can be included, which belongs to a short-range wireless transmission technology. The electronic device can help users to send and receive information through the WiFi module.
[0245] The memory 1120 can be used to store software programs and data. The processor 1180 executes various functions and data processing of the terminal device 110 by running the software programs or data stored in the memory 1120. The memory 1120 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device. The memory 1120 stores an operating system that enables the terminal device 110 to operate. In this application, the memory 1120 can store the operating system and various application programs, and can also store the computer programs for executing the behavior detection method for distributing advertisements in the embodiments of the application.
[0246] The display unit 1130 can also be used to display information input by the user or information provided to the user, as well as the graphical user interface (GUI) of various menus of the terminal device 110. Specifically, the display unit 1130 can include a display screen 1132 arranged on the front of the terminal device 110. The display screen 1132 can be configured in the form of a liquid crystal display, a light-emitting diode, etc. The display unit 1130 can be used to display the behavior detection user interface and the like in the embodiments of the application.
[0247] The display unit 1130 can also be used to receive input digital or character information, and generate signal input related to user settings and function control of the terminal device 110. Specifically, the display unit 1130 can include a touch screen 1131 arranged on the front of the terminal device 110, which can collect touch operations of the user thereon or therearound, such as clicking buttons, dragging scroll boxes, etc.
[0248] The touch screen 1131 can be overlaid on the display screen 1132, or the touch screen 1131 can be integrated with the display screen 1132 to realize the input and output functions of the terminal device 110. After integration, the touch screen 1131 can be referred to as a touch display screen. The display unit 1130 can display an application program and corresponding operation steps.
[0249] The camera 1140 can be used to capture still images, and a user can post comments on images captured by the camera 1140 through an application. The camera 1140 can be one or multiple. An object generates an optical image through a lens and projects the optical image onto a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts an optical signal into an electrical signal, and then transmits the electrical signal to the processor 1180 to convert the electrical signal into a digital image signal.
[0250] The terminal device can also include at least one sensor 1150, such as an acceleration sensor 1151, a distance sensor 1152, a fingerprint sensor 1153, and a temperature sensor 1154. The terminal device can also be configured with a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, a light sensor, a motion sensor, and other sensors.
[0251] The audio circuit 1160, the speaker 1161, and the microphone 1162 can provide an audio interface between a user and the terminal device 110. The audio circuit 1160 can convert received audio data into an electrical signal and transmit the electrical signal to the speaker 1161, which converts the electrical signal into a sound signal for output. The terminal device 110 can also be configured with a volume button for adjusting the volume of the sound signal. On the other hand, the microphone 1162 converts a sound signal collected into an electrical signal, which is received by the audio circuit 1160 and converted into audio data. The audio data is then output to the communication component 1110 for transmission to, for example, another terminal device 110, or to the memory 1120 for further processing.
[0252] The Bluetooth module 1170 is used to interact with other Bluetooth devices having a Bluetooth module through a Bluetooth protocol. For example, the terminal device can establish a Bluetooth connection with a wearable electronic device (e.g., a smart watch) having a Bluetooth module through the Bluetooth module 1170, and thus interact with the wearable electronic device.
[0253] The processor 1180 is a control center of the terminal device, which connects all parts of the terminal through various interfaces and lines, and performs various functions of the terminal device and processes data by running or executing software programs stored in the memory 1120 and calling data stored in the memory 1120. In some embodiments, the processor 1180 can include one or more processing units; the processor 1180 can also integrate an application processor and a baseband processor, wherein the application processor mainly processes operating systems, user interfaces, and application programs, and the baseband processor mainly processes wireless communication. It can be understood that the above-mentioned baseband processor can also not be integrated into the processor 1180. In the present application, the processor 1180 can run an operating system, an application program, a user interface display and a touch response, and a behavior detection method for spreading advertisements of the embodiments of the present application. In addition, the processor 1180 is coupled with the display unit 1130.
[0254] In some possible implementation manners, various aspects of the behavior detection method for spreading advertisements provided by the present application can also be implemented in the form of a program product, which includes a computer program for causing an electronic device to perform the steps in the behavior detection method for spreading advertisements according to various exemplary embodiments of the present application described above in the specification when the program product is run on the electronic device, for example, the steps shown in FIG. 8. Figure 2
[0255] The program product can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0256] The program product of the embodiments of the present application can adopt a portable compact disk read-only memory (CD-ROM) and include a computer program, and can be run on an electronic device. However, the program product of the present application is not limited to this, and in this document, the readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with a command execution system, device or apparatus.
[0257] A readable signal medium can include a computer program readable computer program code, data signals, or any other medium of the like, tangibly embodying computer program code and / or computer instructions that are executable by a computer program execution system, apparatus, or device. The computer program code and / or computer instructions can be transmitted in a computer program product, modulated onto a carrier wave, and / or otherwise be transmitted from one place to another via a network.
[0258] The computer program code and / or computer instructions can be transmitted in a computer program product, modulated onto a carrier wave, and / or otherwise be transmitted from one place to another via a network.
[0259] The computer program code and / or computer instructions can be transmitted in a computer program product, modulated onto a carrier wave, and / or otherwise be transmitted from one place to another via a network.
[0260] It should be noted that although the above detailed description refers to several units or sub-units of the apparatus, such division is merely exemplary and not mandatory. Indeed, according to an embodiment of the application, features and functions of two or more units described above can be embodied in one unit. Conversely, features and functions of one unit described above can be further divided into several units.
[0261] Moreover, while operations of the method of the present application are described in a particular order in the figures, this is not required or implied in any manner, and one can perform the operations in any order, or perform some operations but not others, or perform them all, and still be in accordance with the present application. Additionally or alternatively, some steps can be combined into one step, and / or one step can be divided into multiple steps.
[0262] Those skilled in the art will appreciate that embodiments of the present application can be readily used as software, hardware, or a combination of software and hardware. In one embodiment, the present application can be implemented in software and can be stored on a computer readable medium, which can include random access memory (RAM), read only memory (ROM), magnetic disk or optical disk, or the like. The software implementation can comprise one or more computer program components embodied on one or more computer readable medium(s).
[0263] The present application is described in reference to the flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program commands. These computer program commands can be provided to a processor of a general purpose computer, special purpose computer, an embedded processor or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart illustrations and / or block diagrams block or blocks. Figure 1 one or more flowcharts and / or blocks Figure 1 means for carrying out the function specified by the block(s) of the flowchart and / or block diagram.
[0264] These computer program commands can also be stored in a computer readable medium that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable medium produce an article of manufacture including instructions which implement the function specified in the flowchart illustrations and / or block diagrams block or blocks. Figure 1 one or more flowcharts and / or blocks Figure 1 means for carrying out the function specified by the block(s) of the flowchart and / or block diagram.
[0265] These computer program commands can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart illustrations and / or block diagrams block or blocks. Figure 1 one or more flowcharts and / or blocks Figure 1 means for carrying out the function specified by the block(s) of the flowchart and / or block diagram.
[0266] While preferred embodiments of the application have been described, modifications and variations can be apparent to those skilled in the art once aware of the general underlying concepts. Accordingly, the appended claims are intended to embrace all such modifications and variations as fall within the scope of the application.
[0267] Obviously, many modifications and variations of the present application are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the appended claims and their equivalents, the application can be practiced otherwise than as specifically described.
Claims
1. A method for detecting the behavior of distributing advertisements, characterized in that, The method includes: Target detection is performed on each video frame in the video to be detected to obtain candidate location information of candidate objects and advertisement location information of advertisements in each video frame; Based on the candidate position information and the advertisement position information corresponding to each video frame, the positional relationship between the candidate object and the advertisement in the corresponding video frame is determined. For each candidate object, the following operations are performed: For a candidate object, if the number of target video frames corresponding to the candidate object is greater than a preset frame number threshold, then the candidate object is regarded as a detection object; otherwise, the candidate object is classified as a distribution object; wherein, the positional relationship between the candidate object and the advertisement in the target video frame is that there is overlap. Based on the number of overlapping distribution objects with the corresponding advertisement, it is determined whether the object to be detected is distributing the advertisement.
2. The method as described in claim 1, characterized in that, The step of determining the positional relationship between the candidate object and the advertisement in the corresponding video frame based on the candidate position information and the advertisement position information corresponding to each video frame includes: Perform the following operations on the candidate objects in each video frame: For a video frame, object keypoint detection is performed on the video frame region containing the candidate object to obtain the position information of each target keypoint in the video frame region. The video frame region is determined based on the candidate position information. Based on the location information of each target key point, the location information of the target part of the candidate object is determined; Based on the location information of the object and the location information of the advertisement, the positional relationship between the candidate object and the advertisement is determined.
3. The method as described in claim 1 or 2, characterized in that, If it is determined that the object to be detected is distributing the advertisement, the method further includes: Based on the number of overlapping distribution targets with respect to the advertisement, the number of times the target object distributes the advertisement is determined.
4. The method as described in claim 1 or 2, characterized in that, If it is determined that the object to be detected is distributing the advertisement, the method further includes: Based on the determined positional relationships, advertisements to be associated in each video frame are obtained. The advertisements to be associated are those whose positional relationship with the candidate object in the corresponding video frame does not overlap. The advertisements to be associated are associated with the object to be detected, and the advertisement abandonment level corresponding to the object to be detected is determined based on the number of advertisements associated with the object to be detected.
5. The method as described in claim 1 or 2, characterized in that, The step of performing target detection on each video frame in the video to be detected, and obtaining candidate location information of candidate objects and ad location information of ads in each video frame, includes: Perform the following operations on each video frame: For a video frame, target detection is performed on the video frame, and each detected target is taken as a first target, and the first position information of each first target is obtained; each first target is an object or advertisement in the corresponding video frame. Based on each first location information, and the second location information of each second target in the current target set in the corresponding video frame, the first target and the second target are matched. The current target set includes targets detected in video frames before the first video frame. If a match is successful, and the first target is an object, then the first target is taken as a candidate object, and the first location information is taken as the candidate location information of the candidate object; If a match is successful, and the first target is an advertisement, then the first target is used as the advertisement, and the first location information is used as the advertisement location information of the advertisement.
6. The method as described in claim 5, characterized in that, The method further includes: The first location information of the successfully matched first target is used as the second location information of the corresponding second target in the video frame; and / or, The first target that failed to match is added as a new second target to the object collection.
7. The method as described in claim 5, characterized in that, The method further includes: If, for a second target, there is no first target that successfully matches the second target within a preset number of frames, then the second target is removed from the object set.
8. The method as described in claim 1 or 2, characterized in that, Before determining the positional relationship between the candidate object and the advertisement in the corresponding video frame based on the candidate position information and the advertisement position information corresponding to each video frame, the method further includes: Target detection is performed on each video frame to obtain the first confidence score of candidate objects in each video frame, and candidate objects with a first confidence score lower than a first confidence threshold are removed; and / or, Target detection is performed on each video frame to obtain the second confidence level of the advertisements in each video frame, and advertisements with a second confidence level lower than the second confidence threshold are removed.
9. A device for detecting the distribution of advertisements, characterized in that, include: The detection unit is used to perform target detection on each video frame in the video to be detected, and to obtain the candidate position information of the candidate object and the advertisement position information of the advertisement in each video frame. The first determining unit is used to determine the positional relationship between the candidate object and the advertisement in the corresponding video frame based on the candidate position information and the advertisement position information corresponding to each video frame, respectively. The filtering unit is used to perform the following operations for each candidate object: for a candidate object, if the number of target video frames corresponding to the candidate object is greater than a preset frame number threshold, then the candidate object is regarded as a detection object; otherwise, the candidate object is classified as a distribution object; wherein, the positional relationship between the candidate object and the advertisement in the target video frame is that there is overlap. The second determining unit is used to determine whether the object to be detected has engaged in distributing the advertisement based on the number of distribution objects that overlap with the corresponding advertisement in terms of their positional relationship.
10. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program that, when executed by the processor, causes the processor to perform the steps of any of the methods described in claims 1 to 8.
11. A computer-readable storage medium, characterized in that, It includes a computer program that, when run on an electronic device, causes the electronic device to perform the steps of any of the methods described in claims 1 to 8.
12. A computer program product, characterized in that, The method includes a computer program stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, causing the electronic device to perform the steps of any one of claims 1 to 8.
Citation Information
Patent Citations
Detection method and detection system for leaflet issuing behaviors
CN110533011A
Video detection method and device, electronic equipment and storage medium
CN113255625A