Theft intention identification method and device
By determining the degree of overlap between the detection boxes of the object image and the person image in the target image, theft discrimination information is generated, which solves the problems of low efficiency and low accuracy of theft intent recognition in the existing technology, and achieves more efficient and accurate theft intent recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANKER INNOVATIONS TECH CO LTD
- Filing Date
- 2024-10-31
- Publication Date
- 2026-05-01
AI Technical Summary
In existing technologies, theft intent recognition is inefficient and inaccurate, especially when processing a large number of video frames, which results in low recognition efficiency due to the large amount of data.
By acquiring images of objects and people in the target image, the degree of overlap of the detection boxes of the object images and people images is determined, theft discrimination information is generated, and the degree of overlap and preset threshold are used to determine whether there is a theft intention. This can be further confirmed by associating video frame sequences.
It improves the efficiency and accuracy of identifying theft intent, reduces false alarms, minimizes interference from frequent alarms, and enhances security.
Smart Images

Figure CN121963067A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of security, and more particularly to a method and apparatus for identifying theft intent. Background Technology
[0002] Theft intent identification is an important security measure that aims to determine whether an individual has a tendency or plan to commit theft by observing, analyzing, and evaluating their behavior, speech, facial expressions, and related environmental factors.
[0003] For example, if someone frequently focuses on a particular item, their gaze wanders, and they deliberately avoid eye contact, this could be a sign of potential theft. Similarly, someone displaying excessive interest in another person's belongings, appearing physically close yet tense and uneasy, may also have theft intent. Furthermore, some people's words can also reveal their theft intent. For instance, discussing in private how to bypass security systems and obtain other people's valuables.
[0004] In related technologies, it is common to analyze a large number of video frames captured by a camera to determine whether a person intends to steal an item. However, this method requires processing a large amount of data, resulting in low efficiency in identifying theft intent. Furthermore, schemes that identify theft intent using single video frames often have low accuracy.
[0005] It is evident that improving the efficiency and accuracy of identifying theft intent is a technical issue worthy of attention. Summary of the Invention
[0006] In view of this, in order to solve some or all of the above-mentioned technical problems, embodiments of this application provide a method and apparatus for identifying theft intent.
[0007] In a first aspect, embodiments of this application provide a method for identifying theft intent, the method comprising:
[0008] Acquire a target image, wherein the target image includes an object image and a person image, the object image representing a target object and the person image representing a target person;
[0009] A first detection box and a second detection box are determined in the target image, wherein the first detection box is the detection box of the object image and the second detection box is the detection box of the person image;
[0010] Determine the degree of overlap between the first detection box and the second detection box;
[0011] Based on the degree of overlap, theft detection information is generated, wherein theft detection information indicates whether the target person has the intention to steal the target item.
[0012] In one possible implementation, generating theft detection information based on the degree of overlap includes:
[0013] Determine whether the degree of overlap is greater than or equal to a preset threshold;
[0014] If the degree of overlap is greater than or equal to the preset threshold, determine whether the behavior of the target person represented by the personnel image is theft, so as to obtain a first determination result;
[0015] Based on the first determination result, theft detection information is generated.
[0016] In one possible implementation, the target image is a video frame from a target video; and the generation of theft detection information based on the degree of overlap includes:
[0017] Determine whether the degree of overlap is greater than or equal to a preset threshold;
[0018] If the degree of overlap is greater than or equal to the preset threshold, extract the associated video frame sequence of the target image from the target video;
[0019] Theft detection information is generated based on the associated video frame sequence.
[0020] In one possible implementation, the method is applied to a first device, where the data processing amount of the target image is less than the data processing amount of the video frames in the associated video frame sequence; and
[0021] The process of generating theft detection information based on the associated video frame sequence includes:
[0022] The associated video frame sequence is sent to a second device; wherein the second device is used to generate theft detection information based on the associated video frame sequence; the computing power of the second device is greater than that of the first device;
[0023] The theft identification information returned by the second device is received to generate the theft identification information.
[0024] In one possible implementation, the associated video frame sequence includes: a preceding video frame and a following video frame of the target image; the preceding video frame is a video frame in the target video that precedes the target image, and the following video frame is a video frame in the target video that follows the target image; and
[0025] The process of generating theft detection information based on the associated video frame sequence includes:
[0026] Determine the first state information of the target item in the previous video frame;
[0027] Determine the second state information of the target item in the subsequent video frame;
[0028] Based on the first state information, the second state information, and the associated video frame sequence, theft detection information is generated.
[0029] In one possible implementation, the first state information indicates whether the target item represented by the item image in the previous video frame is located in the first region, and the second state information indicates whether the target item represented by the item image in the subsequent video frame is located in the first region; and
[0030] The process of generating theft detection information based on the first state information, the second state information, and the associated video frame sequence includes:
[0031] If the first state information indicates that the target item represented by the item image in the previous video frame is located in the first area, initial theft detection information is generated based on the associated video frame sequence.
[0032] If the initial theft identification information indicates that the target person has the intention to steal the target item, and the second state information indicates that the target item represented by the item image in the subsequent video frame is not located in the first area, then final theft identification information indicating that the target person has the intention to steal the target item is generated.
[0033] In one possible implementation, generating theft detection information based on the degree of overlap includes:
[0034] Determine whether both the target person and the target item are located in a preset area to obtain a second determination result;
[0035] Based on the second determination result and the degree of overlap, theft detection information is generated.
[0036] In one possible implementation, before determining the degree of overlap between the first detection box and the second detection box, the method further includes:
[0037] Retrieve the pre-entered set of personnel information;
[0038] Determine whether the personnel information set includes target personnel information representing the target personnel; and
[0039] Determining the degree of overlap between the first detection box and the second detection box includes:
[0040] If the target personnel information is not included in the personnel information set, the degree of overlap between the first detection box and the second detection box is determined.
[0041] In one possible implementation, after generating the theft detection information, the method further includes at least one of the following:
[0042] If the theft detection information indicates that the target person has the intent to steal, the expulsion device is controlled to perform an expulsion operation.
[0043] If the theft detection information indicates that the target person has the intent to steal, a prompt message is sent to a preset terminal.
[0044] Secondly, embodiments of this application provide a device for identifying intent to steal, the device comprising:
[0045] The first acquisition unit is used to acquire a target image, wherein the target image includes an object image and a person image, the object image representing a target object and the person image representing a target person;
[0046] The first determining unit is used to determine a first detection box and a second detection box in the target image, wherein the first detection box is the detection box of the object image and the second detection box is the detection box of the person image.
[0047] The second determining unit is used to determine the degree of overlap between the first detection box and the second detection box;
[0048] The generation unit is used to generate theft identification information based on the degree of overlap, wherein theft identification information indicates whether the target person has the intention to steal the target item.
[0049] In one possible implementation, generating theft detection information based on the degree of overlap includes:
[0050] Determine whether the degree of overlap is greater than or equal to a preset threshold;
[0051] If the degree of overlap is greater than or equal to the preset threshold, determine whether the behavior of the target person represented by the personnel image is theft, so as to obtain a first determination result;
[0052] Based on the first determination result, theft detection information is generated.
[0053] In one possible implementation, the target image is a video frame from a target video; and the generation of theft detection information based on the degree of overlap includes:
[0054] Determine whether the degree of overlap is greater than or equal to a preset threshold;
[0055] If the degree of overlap is greater than or equal to the preset threshold, extract the associated video frame sequence of the target image from the target video;
[0056] Theft detection information is generated based on the associated video frame sequence.
[0057] In one possible implementation, the apparatus is applied to a first device, wherein the data processing amount of the target image is less than the data processing amount of the video frames in the associated video frame sequence; and
[0058] The process of generating theft detection information based on the associated video frame sequence includes:
[0059] The associated video frame sequence is sent to a second device; wherein the second device is used to generate theft detection information based on the associated video frame sequence; the computing power of the second device is greater than that of the first device;
[0060] The theft identification information returned by the second device is received to generate the theft identification information.
[0061] In one possible implementation, the associated video frame sequence includes: a preceding video frame and a following video frame of the target image; the preceding video frame is a video frame in the target video that precedes the target image, and the following video frame is a video frame in the target video that follows the target image; and
[0062] The process of generating theft detection information based on the associated video frame sequence includes:
[0063] Determine the first state information of the target item in the previous video frame;
[0064] Determine the second state information of the target item in the subsequent video frame;
[0065] Based on the first state information, the second state information, and the associated video frame sequence, theft detection information is generated.
[0066] In one possible implementation, the first state information indicates whether the target item represented by the item image in the previous video frame is located in the first region, and the second state information indicates whether the target item represented by the item image in the subsequent video frame is located in the first region; and
[0067] The process of generating theft detection information based on the first state information, the second state information, and the associated video frame sequence includes:
[0068] If the first state information indicates that the target item represented by the item image in the previous video frame is located in the first area, initial theft detection information is generated based on the associated video frame sequence.
[0069] If the initial theft identification information indicates that the target person has the intention to steal the target item, and the second state information indicates that the target item represented by the item image in the subsequent video frame is not located in the first area, then final theft identification information indicating that the target person has the intention to steal the target item is generated.
[0070] In one possible implementation, generating theft detection information based on the degree of overlap includes:
[0071] Determine whether both the target person and the target item are located in a preset area to obtain a second determination result;
[0072] Based on the second determination result and the degree of overlap, theft detection information is generated.
[0073] In one possible implementation, before determining the degree of overlap between the first detection frame and the second detection frame, the device further includes:
[0074] The second acquisition unit is used to acquire a pre-entered set of personnel information;
[0075] The third determining unit is configured to determine whether the personnel information set includes target personnel information representing the target personnel; and
[0076] Determining the degree of overlap between the first detection box and the second detection box includes:
[0077] If the target personnel information is not included in the personnel information set, the degree of overlap between the first detection box and the second detection box is determined.
[0078] In one possible implementation, after generating the theft detection information, the device further includes at least one of the following:
[0079] The control unit is used to control the expulsion device to perform an expulsion operation when the theft identification information indicates that the target person has the intent to steal;
[0080] The sending unit is used to send a prompt message to a preset terminal when the theft identification information indicates that the target person has the intent to steal.
[0081] Thirdly, embodiments of this application provide an electronic device, including:
[0082] Memory, used to store computer programs;
[0083] A processor is configured to execute a computer program stored in the memory, and when the computer program is executed, to implement any embodiment of the method for identifying theft intent according to the first aspect of this application.
[0084] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, it implements the method of any embodiment of the method for identifying theft intent as described in the first aspect above.
[0085] Fifthly, embodiments of this application provide a computer program product comprising computer-readable code that, when executed on a device, causes a processor in the device to implement the method of any embodiment of the theft intent identification method of the first aspect described above.
[0086] The theft intent identification method provided in this application can acquire a target image, wherein the target image includes an object image and a person image, the object image representing the target object, and the person image representing the target person. Then, a first detection box and a second detection box are determined in the target image, wherein the first detection box is the detection box of the object image, and the second detection box is the detection box of the person image. Next, the degree of overlap between the first and second detection boxes is determined. Finally, based on the degree of overlap, theft discrimination information is generated, wherein theft discrimination information indicates whether the target person has the intent to steal the target object. Therefore, in some cases, the degree of overlap between the detection boxes of the object image and the person image in a single image can be used to determine whether the target person has the intent to steal the target object, thus improving the efficiency and accuracy of theft intent identification. Attached Figure Description
[0087] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0088] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0089] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0090] Figure 1 A flowchart illustrating a method for identifying theft intent provided in an embodiment of this application;
[0091] Figure 2 A flowchart illustrating another method for identifying theft intent provided in an embodiment of this application;
[0092] Figure 3 A flowchart illustrating another method for identifying theft intent provided in an embodiment of this application;
[0093] Figure 4 A schematic diagram of the structure of a theft intent identification device provided in an embodiment of this application;
[0094] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0095] Various exemplary embodiments of this application will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of this application.
[0096] Those skilled in the art will understand that the terms "first" and "second" in the embodiments of this application are only used to distinguish different steps, devices or modules, and do not represent any specific technical meaning, nor do they indicate the logical order between them.
[0097] It should also be understood that in this embodiment, "multiple" can refer to two or more, and "at least one" can refer to one, two or more.
[0098] It should also be understood that any component, data or structure mentioned in the embodiments of this application can generally be understood as one or more unless explicitly defined or given contrary guidance in the context.
[0099] Furthermore, the term "and / or" in this application is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this application generally indicates that the preceding and following related objects have an "or" relationship.
[0100] It should also be understood that the description of the various embodiments in this application emphasizes the differences between the various embodiments, and the similarities or similarities can be referred to each other. For the sake of brevity, they will not be described in detail.
[0101] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the scope of this application and its application or use.
[0102] Techniques, methods, and equipment known to those skilled in the art in the relevant field may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0103] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0104] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. To facilitate understanding of the embodiments of this application, the application will be described in detail below with reference to the accompanying drawings and embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0105] To address the technical problem of how to improve the efficiency and accuracy of identifying theft intent in the prior art, this application provides a method and apparatus for identifying theft intent, which can improve the efficiency and accuracy of identifying theft intent.
[0106] Figure 1 This is a flowchart illustrating a method for identifying theft intent provided in an embodiment of this application. This method can be applied to one or more electronic devices such as smartphones, laptops, desktop computers, portable computers, and servers. Furthermore, the execution entity of this method can be hardware or software. When the execution entity is hardware, it can be one or more of the aforementioned electronic devices. For example, a single electronic device can execute this method, or multiple electronic devices can cooperate with each other to execute this method. When the execution entity is software, this method can be implemented as multiple software programs or software modules, or as a single software program or software module. No specific limitations are imposed here.
[0107] like Figure 1 As shown, the method specifically includes:
[0108] Step 101: Obtain the target image, wherein the target image includes an object image and a person image, the object image representing the target object and the person image representing the target person.
[0109] In this embodiment, the target image can be any image that includes both object and person images. As an example, the target image can be a video frame extracted from a video that includes both object and person images. As yet another example, the target image can also be a video frame extracted from a video that includes both object and person images, where the target object represented by the object image and the target person represented by the person image are both located within a preset area.
[0110] The aforementioned preset area can be set by a user or other object, or it can be an area determined by the aforementioned executing entity or other electronic device that meets preset conditions. For example, the preset condition could be that the area includes a preset item. In this case, the preset item can refer to the same item as the target item.
[0111] The preset area can be a fixed area or an area whose position changes. For example, if the preset condition is that the area includes a preset item, and if the preset item (such as a robot vacuum cleaner, a pet, etc.) can move, then the preset area can be an area whose position changes.
[0112] The target item can be an item represented by an image, and the target person can be a person represented by an image.
[0113] Step 102: Determine a first detection box and a second detection box in the target image, wherein the first detection box is the detection box of the object image and the second detection box is the detection box of the person image.
[0114] In this embodiment, a target detection algorithm can be used to detect targets in the target image, thereby determining the first detection box and the second detection box in the target image.
[0115] Among them, object detection algorithms are algorithms in the field of computer vision used to identify and locate specific targets (such as target objects or people) in images (including the aforementioned target images). These algorithms are able to determine the location of the target (usually by drawing a rectangle or more complex shapes) and may include classifying the target.
[0116] Here, OVOD (Open Vocabulary Object Detection) can be used for object detection. OVOD allows the model to detect and recognize new object categories that have not been seen during the training phase, thereby achieving generalized object detection.
[0117] Step 103: Determine the degree of overlap between the first detection box and the second detection box.
[0118] In this embodiment, the degree of overlap can represent the proportion or level of the overlapping portion of the first and second detection boxes in the whole. As an example, the degree of overlap can be represented by at least one of the following: the number of pixels overlapping between the image regions corresponding to the first and second detection boxes, the ratio of the overlapping area between the image regions corresponding to the first and second detection boxes, the intersection-union ratio (IUGR) of the image regions corresponding to the first and second detection boxes, etc.
[0119] Step 104: Based on the degree of overlap, generate theft detection information, wherein theft detection information indicates whether the target person has the intention to steal the target item.
[0120] In this embodiment, theft detection information can be generated in various ways based on the degree of overlap.
[0121] As an example, when the degree of overlap is greater than or equal to a preset threshold, theft identification information indicating that the target person has the intent to steal the target item can be generated; when the degree of overlap is less than the preset threshold, theft identification information indicating that the target person does not have the intent to steal the target item can be generated. The preset threshold can be set by a user or other object, or it can be determined by statistically analyzing the correspondence between the degree of overlap and theft identification information.
[0122] As another example, when the degree of overlap is greater than or equal to a preset threshold, and both the target person and the target item are located in a preset area, theft identification information indicating that the target person has the intent to steal the target item can be generated; when the degree of overlap is less than the preset threshold, or when at least one of the target person and the target item is not located in the preset area, theft identification information indicating that the target person does not have the intent to steal the target item can be generated.
[0123] In some optional implementations of this embodiment, theft detection information can be generated based on the degree of overlap in the following manner:
[0124] The first step is to determine whether the degree of overlap is greater than or equal to a preset threshold.
[0125] The aforementioned preset threshold can be set by users or other objects, or it can be determined by statistically analyzing the correspondence between the degree of overlap and theft identification information.
[0126] The second step is to determine whether the behavior of the target person represented by the personnel image is theft if the degree of overlap is greater than or equal to the preset threshold, so as to obtain the first determination result.
[0127] The first determination result can indicate whether the behavior of the target person represented by the personnel image is theft.
[0128] Theft can be one or more acts that constitute theft. For example, theft could be bending down, reaching out and glancing sideways, or reaching out and glancing sideways while walking or running quickly.
[0129] The third step is to generate theft detection information based on the first determination result.
[0130] Here, various methods can be used to generate theft detection information based on the first determination result.
[0131] As an example, if the first determination result indicates that the behavior of the target person represented by the personnel image is theft, theft discrimination information indicating that the target person has the intent to steal the target item can be generated; if the first determination result indicates that the behavior of the target person represented by the personnel image is not theft, theft discrimination information indicating that the target person does not have the intent to steal the target item can be generated.
[0132] In addition, other methods can be used to generate theft identification information based on the first determination result. Please refer to the following description for details, which will not be elaborated here.
[0133] It is understandable that, among the above-mentioned optional implementation methods, theft discrimination information can be generated by determining whether the behavior of the target person represented by the personnel image constitutes theft. This can improve the accuracy of theft intent recognition.
[0134] In some optional implementations of this embodiment, theft detection information can be generated based on the degree of overlap in the following manner:
[0135] The first step is to determine whether the target personnel and the target items are both located in a preset area, in order to obtain a second determination result.
[0136] The second determination result mentioned above can indicate whether the target person and the target item are both located in the preset area.
[0137] The aforementioned preset area can be set by a user or other object, or it can be an area determined by the aforementioned executing entity or other electronic device that meets preset conditions. For example, the aforementioned preset condition could be that the area includes a preset item. In this case, the aforementioned preset item can refer to the same item as the aforementioned target item.
[0138] The preset area can be a fixed area or an area whose position changes. For example, if the preset condition is that the area includes a preset item, and if the preset item (such as a robot vacuum cleaner, a pet, etc.) can move, then the preset area can be an area whose position changes.
[0139] The second step involves generating theft detection information based on the second determination result and the degree of overlap.
[0140] Here, theft detection information can be generated in various ways based on the second determination result and the degree of overlap.
[0141] As an example, if the second determination result indicates that both the target person and the target item are located in a preset area, and the degree of overlap is greater than or equal to the preset threshold, theft discrimination information indicating that the target person has the intention to steal the target item can be generated; if the second determination result indicates that at least one of the target person and the target item is not located in the preset area, or the degree of overlap is less than the preset threshold, theft discrimination information indicating that the target person does not have the intention to steal the target item can be generated.
[0142] As another example, if the second determination result indicates that both the target person and the target item are located in a preset area, and the degree of overlap is greater than or equal to the preset threshold, it can be further determined whether the behavior of the target person represented by the personnel image is theft, thus obtaining a first determination result. The first determination result indicates whether the behavior of the target person represented by the personnel image is theft. Subsequently, based on the first determination result, theft discrimination information is generated.
[0143] In addition, other methods can be used to generate theft detection information based on the second determination result and the degree of overlap. Please refer to the description below for details, which will not be elaborated here.
[0144] It is understood that, in the above-mentioned optional implementation methods, theft discrimination information can be generated based on both the second determination result and the degree of overlap. This can further improve the accuracy of theft intent identification.
[0145] In some optional implementations of this embodiment, before determining the degree of overlap between the first detection box and the second detection box, the following steps may also be performed:
[0146] The first step is to obtain a pre-entered set of personnel information.
[0147] The personnel information in the above-mentioned personnel information set can represent a family member and their relatives and friends.
[0148] In practice, personnel information can be entered by collecting images of relevant personnel.
[0149] The second step is to determine whether the personnel information set includes target personnel information representing the target personnel.
[0150] Based on this, the degree of overlap between the first detection box and the second detection box can be determined even if the target personnel information is not included in the personnel information set.
[0151] Optionally, if the target person's information is included in the personnel information set, it is not necessary to determine the degree of overlap between the first and second detection boxes. Furthermore, theft discrimination information indicating that the target person does not have the intent to steal the target item can be generated.
[0152] It is understood that, in the above-mentioned optional implementation methods, if the target person's information is not included in the personnel information set, the degree of overlap between the detection boxes of the item image and the personnel image in the image can be used to determine whether the target person has the intention to steal the target item. This can improve the accuracy of theft intent recognition. Furthermore, in scenarios where it is determined that the target person has the intention to steal the target item and an alarm is required, the disturbance caused to users and other parties by frequent alarm prompts can be reduced.
[0153] In some optional implementations of this embodiment, after generating theft identification information, if the theft identification information indicates that the target person has the intent to steal, the expulsion device can be further controlled to perform an expulsion operation, and / or a prompt message can be sent to a preset terminal.
[0154] The expulsion device can be an audio output device, a mobile robot, etc.
[0155] When the ejection device is an audio output device, the ejection operation can output an alarm prompt audio. This audio can be set by a user or other entity.
[0156] When the expulsion device is a mobile robot, the expulsion operation can be carried out by moving towards the location of the target person.
[0157] The default terminal can be a device pre-associated with the aforementioned execution entity. As an example, the default terminal can be a device logged in with an administrator account.
[0158] It is understandable that, among the above-mentioned optional implementation methods, controlling the expulsion device to perform the expulsion operation can reduce the probability of the target item being stolen, and sending a prompt message to the preset terminal can promptly inform the user of the preset terminal that the target item may or is about to be stolen.
[0159] It should be noted that, where there is no conflict, the technical features described in different alternative implementations can be included in the same technical solution. For the sake of brevity, they will not be elaborated here.
[0160] The theft intent identification method provided in this application can acquire a target image, wherein the target image includes an object image and a person image, the object image representing the target object, and the person image representing the target person. Then, a first detection box and a second detection box are determined in the target image, wherein the first detection box is the detection box of the object image, and the second detection box is the detection box of the person image. Next, the degree of overlap between the first and second detection boxes is determined. Finally, based on the degree of overlap, theft discrimination information is generated, wherein theft discrimination information indicates whether the target person has the intent to steal the target object. Therefore, in some cases, the degree of overlap between the detection boxes of the object image and the person image in a single image can be used to determine whether the target person has the intent to steal the target object, thus improving the efficiency and accuracy of theft intent identification.
[0161] Figure 2 This is a flowchart illustrating another method for identifying theft intent provided in an embodiment of this application. Figure 2 As shown, the method specifically includes:
[0162] Step 201: Obtain a target image, wherein the target image includes an object image and a person image, the object image represents a target object, the person image represents a target person, and the target image is a video frame in a target video.
[0163] In this embodiment, the target image is a video frame from the target video. Specifically, the target image can be a video frame extracted from the video that includes images of objects and people. In addition, step 201 and... Figure 1 Step 101 in the corresponding embodiment is basically the same, and will not be repeated here.
[0164] Step 202: Determine a first detection box and a second detection box in the target image, wherein the first detection box is the detection box of the object image and the second detection box is the detection box of the person image.
[0165] In this embodiment, step 202 and Figure 1 Step 102 in the corresponding embodiment is basically the same, and will not be repeated here.
[0166] Step 203: Determine the degree of overlap between the first detection box and the second detection box.
[0167] In this embodiment, step 203 and Figure 1 Step 103 in the corresponding embodiment is basically the same, and will not be repeated here.
[0168] Step 204: Determine whether the degree of overlap is greater than or equal to a preset threshold.
[0169] In this embodiment, the aforementioned preset threshold can be set by an object such as a user, or it can be determined by statistically analyzing the correspondence between the degree of overlap and theft detection information.
[0170] Step 205: If the degree of overlap is greater than or equal to the preset threshold, extract the associated video frame sequence of the target image from the target video.
[0171] In this embodiment, the associated video frame sequence can be composed of video frames in the target video that are associated with the target image.
[0172] As an example, the associated video frame sequence may include: the target image, N video frames of the target image preceding it, and M video frames of the target image following it.
[0173] As yet another example, the associated video frame sequence may also include: N video frames of the target image preceding the target image, and M video frames of the target image following the target image.
[0174] In the above examples, N and M are positive integers, and the values of N and M can be equal or unequal. The preceding video frame is the video frame in the target video that precedes the target image, and the following video frame is the video frame in the target video that follows the target image.
[0175] As another example, the associated video frame sequence may also include: video frames containing images of the target item in the target video, and / or video frames containing images of the target person in the target video.
[0176] Step 206: Based on the associated video frame sequence, generate theft detection information, wherein theft detection information indicates whether the target person has the intent to steal the target item.
[0177] In this embodiment, theft detection information can be generated based on the associated video frame sequence using various methods.
[0178] As an example, the aforementioned associated video frame sequence can be input into a pre-trained Large Language Model (LLM) to generate theft detection information. The LLM can represent the correspondence between the prompt words, the associated video frame sequence, and theft detection information.
[0179] Among them, large-scale language models are natural language processing models based on deep learning technology, which have a very high number of parameters and powerful language understanding and generation capabilities.
[0180] As an example, the aforementioned large language model could be MLLM (Multimodal Large Language Models). MLLM builds upon LLM's ability to understand language by incorporating the ability to understand other modalities, enabling the understanding and generation of content involving multiple data types. Here, "modality" refers to different types of data input, such as text, images, audio, and video. Through training on massive amounts of data, multimodal large models can learn the complementarity and correlation between different modalities. For example, MLLM's input data could include the aforementioned associated video frame sequences and cue words, while its output data could be theft detection information.
[0181] In addition, other methods can be used to generate theft detection information based on the associated video frame sequence. Please refer to the description below for details, which will not be elaborated here.
[0182] In some optional implementations of this embodiment, the method is applied to a first device. The amount of data processing required for the target image is less than the amount of data processing required for the video frames in the associated video frame sequence.
[0183] Based on this, theft detection information can be generated using the associated video frame sequence in the following manner:
[0184] The first step is to send the associated video frame sequence to the second device.
[0185] The second device is used to generate theft detection information based on the associated video frame sequence. The computing power of the second device is greater than that of the first device.
[0186] As an example, the first device described above could be an edge computing device. The first device can process video data (such as the target image mentioned above) acquired from a smart camera. This first device has a certain computing power, enabling real-time customized target property detection and human detection. Furthermore, the device includes a microphone and some audio-visual equipment, capable of repelling threats to property security and playing welcome messages to family members or those on a whitelist.
[0187] The second device mentioned above can be a home intelligent central control system (server). The second device can serve as a computing power center and intelligent center, equipped with a high-performance computing chip, capable of building multiple video streams, and capable of real-time processing of video stream behavior.
[0188] Here, the second device can input the associated video frame sequence into a pre-trained large-scale language model to generate theft detection information. Alternatively, the second device can also generate theft detection information based on the associated video frame sequence, the state information of the target item in the preceding video frame, and the state information of the target item in the subsequent video frame.
[0189] The second step is to receive the theft detection information returned by the second device in order to generate the theft detection information.
[0190] It is understood that in the above-mentioned optional implementation methods, a first device with lower computing power can process video frames with smaller data processing volume, and a second device with higher computing power can process multiple video frames with larger data processing volume. Thus, the efficiency of identifying theft intent can be improved by the combination of the two.
[0191] In some optional implementations of this embodiment, the associated video frame sequence includes: a preceding video frame and a following video frame of the target image. The preceding video frame is the video frame in the target video that precedes the target image. The following video frame is the video frame in the target video that follows the target image.
[0192] Based on this, theft detection information can be generated using the associated video frame sequence in the following manner:
[0193] The first step is to determine the first state information of the target item in the previous video frame.
[0194] The first state information can indicate the state of the target item in the preceding video frame. For example, the first state information can indicate the position of the target item in the preceding video frame, or it can indicate whether the target item represented by the image of the item in the preceding video frame is located in a preset area.
[0195] The second step is to determine the second state information of the target item in the subsequent video frame.
[0196] The second state information can indicate the state of the target item in the subsequent video frame. For example, the second state information can indicate the position of the target item in the subsequent video frame, or it can indicate whether the target item represented by the image in the subsequent video frame is located in a preset area.
[0197] The third step is to generate theft detection information based on the first state information, the second state information, and the associated video frame sequence.
[0198] Here, theft detection information can be generated in various ways based on the first state information, the second state information, and the associated video frame sequence.
[0199] As an example, if the first state information indicates the position of the target item in the preceding video frame, and the second state information indicates the position of the target item in the following video frame, then if the distance between the position indicated by the first state information and the position indicated by the second state information is greater than or equal to a preset distance threshold, then theft identification information can be further generated based on the associated video frame sequence; if the distance between the position indicated by the first state information and the position indicated by the second state information is less than the preset distance threshold, then theft identification information indicating that the target person does not have the intent to steal the target item can be generated.
[0200] In addition, other methods can be used to generate theft detection information based on the first state information, the second state information, and the associated video frame sequence. Please refer to the description below for details, which will not be elaborated upon here.
[0201] It is understood that, in the above-mentioned optional implementation methods, theft discrimination information can be generated based on the first state information, the second state information and the associated video frame sequence, thereby further improving the accuracy of theft intent recognition.
[0202] In some application scenarios of the above-mentioned optional implementation methods, the first state information indicates whether the target item represented by the item image in the previous video frame is located in the first region, and the second state information indicates whether the target item represented by the item image in the subsequent video frame is located in the first region.
[0203] The first region can represent the aforementioned preset region, or the first region can be a region with a preset area and shape, and whose position is variable.
[0204] Based on this, theft detection information can be generated using the following method, based on the first state information, the second state information, and the associated video frame sequence:
[0205] The first step is to generate initial theft detection information based on the associated video frame sequence, when the first state information indicates that the target item represented by the item image in the previous video frame is located in the first area.
[0206] Here, various methods can be used to generate initial theft detection information based on the associated video frame sequence.
[0207] As an example, the aforementioned associated video frame sequence can be input into a pre-trained large-scale language model to generate initial theft detection information. This large-scale language model can represent the correspondence between the prompt words, the associated video frame sequence, and the initial theft detection information.
[0208] As another example, initial theft detection information can also be generated based on the associated video frame sequence and whether the target person and the target item are both located in a preset area.
[0209] The initial state information indicates whether the target person has the intention to steal the target item.
[0210] The second step involves generating final theft identification information indicating that the target person has the intent to steal the target item when the initial theft identification information indicates that the target person has the intent to steal the target item, and the second state information indicates that the target item represented by the item image in the later video frame is not located in the first area.
[0211] It is understandable that in the above application scenario, the target person's intent to steal the target item is only definitively determined when the target item changes from being located in the first area to being located outside the first area, and the initial state information indicates that the target person has the intent to steal the target item. This further improves the accuracy of identifying theft intent.
[0212] It should be noted that, in addition to the contents described above, this embodiment may also include... Figure 1 The corresponding technical features described in the corresponding embodiments, thereby achieving Figure 1 For details on the technical effectiveness of the method for identifying theft intent shown, please refer to [link / reference needed]. Figure 1 The relevant descriptions are presented concisely and will not be elaborated upon here.
[0213] The theft intent identification method provided in this application generates theft discrimination information based on the associated video frame sequence of the target image when the overlap between the first detection box and the second detection box is greater than or equal to a preset threshold. This further improves the efficiency and accuracy of theft intent identification.
[0214] The embodiments of this application are described below by way of example. However, it should be noted that the embodiments of this application may have the features described below, but the following description does not constitute a limitation on the protection scope of the embodiments of this application.
[0215] Before introducing this solution, the concept behind its design will be explained as follows:
[0216] The "edge computing + smart terminal" model (hereinafter referred to as "edge + terminal") refers to utilizing the computing resources at the network edge and the local processing capabilities of smart devices to improve data processing speed, reduce network bandwidth requirements, enhance system real-time performance and reliability, and improve user experience. Edge computing is a network architecture that migrates data processing from the data center to the network edge, closer to the data source, to reduce latency, lower bandwidth consumption, improve reliability, and enhance privacy protection. A smart terminal refers to a device with certain computing, storage, and network connectivity capabilities, capable of performing local data processing, providing personalized services, and supporting offline operation.
[0217] Object detection algorithms are used in the field of computer vision to identify and locate specific objects (such as people, vehicles, and objects) in images. These algorithms are able to determine the location of the object (usually by drawing a rectangle or a more complex shape) and may include classifying the object.
[0218] Intent determination refers to enabling computer systems to understand and recognize the intent or purpose of a user's language or behavior.
[0219] LLMs are artificial intelligence models designed to understand and generate human language. They are trained on massive amounts of text data and can perform a wide range of tasks, including text summarization, translation, sentiment analysis, and more. LLMs are characterized by their massive scale, typically containing billions of parameters.
[0220] MLLM builds upon LLM's ability to understand language by combining it with the ability to understand other modal information, enabling the understanding and generation of content involving multiple data types. Here, "modality" refers to different types of data input, such as text, images, audio, and video. Through training on massive amounts of data, large multimodal models can learn the complementarity and correlation between different modalities.
[0221] CNN (Convolutional Neural Network) is a deep learning architecture widely used in visual computing tasks such as image and video recognition, image classification, object detection, face recognition, and medical image analysis. A CNN is a feedforward neural network with a convolutional structure, capable of extracting complex features from images and used for various pattern recognition tasks.
[0222] OVOD is an object detection technique that allows models to detect and identify new object categories not seen during the training phase. Traditional object detection methods typically require all object categories to be detected to be included in the training data, while OVOD technology can achieve generalized object detection.
[0223] This solution enhances security monitoring and crime prediction capabilities in home environments by understanding video surveillance data, thereby protecting users' private property. The following devices are used with this method:
[0224] 1. Smart Cameras: Installed inside or around the home, these cameras monitor and record video in real time. They typically feature high resolution and a wide field of view to cover the main activity areas of the home. In this solution, smart cameras can be used to capture target images and videos.
[0225] 2. Edge computing device: Processes video data (including target images and videos) acquired from smart cameras. This device has certain computing power, enabling real-time customized target property detection and human detection. It also includes a microphone and some audio-visual equipment to repel threats to property security and play welcome messages to family members or those on a whitelist. In this solution, the edge computing device can serve as the executing entity or part of the aforementioned method for identifying theft intent.
[0226] 3. Home Intelligent Central Control System: Serving as the computing and intelligence center, this system is equipped with a high-performance computing chip, capable of building multiple video streams and processing video stream behavior in real time. In this solution, the home intelligent central control system can serve as the second device in the aforementioned method for identifying theft intent.
[0227] 4. Mobile devices (such as smartphones and tablets): Users can customize settings for the property they need to protect in their home environment via mobile devices. Users can also provide feedback on the behavioral judgments learned by the system, allowing the system to learn and iterate based on user-defined settings. Furthermore, when the system detects an act that could harm property, the smart device can send a notification or directly trigger an alarm. In this solution, mobile devices can serve as the preset terminal in the aforementioned method for identifying theft intent.
[0228] This solution can be applied to the following scenarios:
[0229] Home security monitoring: Smart cameras continuously monitor the surrounding environment of the home, and edge computing devices identify theft intentions based on the user-defined warning area (i.e., the preset area mentioned above) and whitelist (i.e., the set of personnel information mentioned above). For example, when a stranger (i.e., the target person mentioned above) is detected approaching and exhibiting abnormal behavior (such as stealing packages or prying open car locks), the system will quickly predict the intent and immediately issue a deterrent alarm and notify the user.
[0230] Smart Home Integration: By combining video information captured by home security cameras, edge computing devices process the video stream in real time, responding promptly to simple situations. For example, when a family member enters a "hotspot" (i.e., the aforementioned preset area), the edge computing device can identify them in real time and play a "Welcome Home" message. For complex situations, the central control system uses larger-scale models and more complex logic for centralized and coordinated calculations. For example, when the edge computing device identifies a stranger entering a hotspot and engaging in actions suspected of theft, it transmits the relevant video frames (i.e., the aforementioned associated video frame sequence) to the central control system for further judgment to improve prediction accuracy.
[0231] Specifically, this solution may include the following steps:
[0232] I. User Settings:
[0233] 1. Users can set the hot zone (i.e. the preset area mentioned above) around their home or vehicle in the APP (Application).
[0234] 2. Users enter the appearance and facial features of their family members into the system, and add the characteristics of trusted relatives and friends as well as delivery personnel into a whitelist, thereby obtaining a collection of personnel information.
[0235] 3. Users can add individuals requiring vigilance or with existing criminal records to a blacklist. An alarm can be triggered when a blacklisted individual is detected in a video frame.
[0236] 4. Users can set personalized voice settings to play different voices for strangers, family members, and deliverymen, and can also record their own messages.
[0237] II. Algorithm Logic:
[0238] See Figure 3 , Figure 3 This is a flowchart illustrating another method for identifying theft intent provided in an embodiment of this application. Figure 3 As shown, the method includes:
[0239] 1. Function trigger: The edge computing device detects packages in real time. When a package is detected, the package guard function is activated.
[0240] 2. In the first stage, the edge device performs pedestrian detection in real time and determines whether a pedestrian has entered the hot zone R. If someone enters the hot zone, different greetings are given based on different ID results (family member, deliveryman, stranger, unidentified).
[0241] 3. In the second stage, when someone enters the hot zone, the intention of the pedestrian is judged. If it is a stranger with the intention to steal, the person is driven away; if it is a family member, the family member is reminded to take the package home; if it is a courier, the courier is reminded to put the package in the designated location; if the identity cannot be identified, the person is informed that they are visiting and to wait while they are on their way.
[0242] Intent determination refers to the ability to judge whether there is a theft intent in the current frame of the image by combining historical information of the video. The specific implementation logic can be divided into two schemes based on the device's computing power requirements: a lightweight scheme (single-frame image overlay logic judgment) and a more computationally demanding scheme (multi-frame image end-to-end implementation). Considering the lightweight nature and practical operability of the intent determination system, this scheme mainly focuses on single-frame image overlay logic judgment, with the specific steps as follows:
[0243] a. The user specifies the hot zone R0 (i.e. the preset area mentioned above) and records the current state of the package (i.e. the target object mentioned above) as Z0 (i.e. the first state information mentioned above);
[0244] b. Perform parcel detection on the real-time video stream. If a parcel bbox_parcel (i.e., the first detection box mentioned above) is detected, it is recorded as P0.
[0245] c. Simultaneously, human detection is performed in the real-time video stream, and the detected human bounding box bbox_person (i.e., the second detection box mentioned above) is denoted as P1;
[0246] d. Perform IOU judgment on the overlapping area of P0 and P1, denoted as Q (i.e., the degree of overlap mentioned above). If both the person and the package are in the hot zone R, and Q is greater than Q0 (i.e., the preset threshold mentioned above), then perform key action judgment on P1. If P1 hits the key action, the key action for the package theft behavior is defined as bending over (i.e., the theft behavior mentioned above), then it is considered that there is an intent to steal. In the whole system, based on the frame where P1 is located (i.e., the target image mentioned above), based on historical and future video frames (i.e., the associated video frame sequence mentioned above), key frame sampling will be performed, and continuous frames (i.e., the associated video frame sequence mentioned above) will be sent to the multimodal large model for continuous behavior sequence judgment.
[0247] e. Finally, combining the package state Z1 (i.e., the second state information mentioned above) and the results of the multimodal large model (i.e., the initial theft detection information mentioned above), a final judgment (i.e., the final theft detection information mentioned above) is given. If the state of the package changes (for example, the state represented by the first state information is different from the state represented by the second state information), and the large model gives a judgment of theft behavior, then it is considered that package theft has occurred.
[0248] Taking a package as an example, this method can be extended to guarding any valuables (i.e., the aforementioned target items). The specific steps are as follows:
[0249] a. The user designates a hot zone RA0 (i.e., the aforementioned preset area) and records the current status of the assets that need to be protected as ZA0 (i.e., the aforementioned first status information);
[0250] b. Perform arbitrary object detection on the real-time video stream. If an object bbox_A (i.e. the first detection box mentioned above) is detected, it is denoted as PA0.
[0251] c. Simultaneously, human detection is performed in the real-time video stream, and the detected human bounding box bbox_person (i.e., the second detection box mentioned above) is denoted as PA1;
[0252] d. Perform IOU judgment on the overlapping area of PA0 and PA1, denoted as QA (i.e., the degree of overlap mentioned above). If it is within the hot zone RA0 and QA is greater than QA0 (i.e., the preset threshold mentioned above), then perform key action judgment on PA1. If PA1 hits the key action, it is considered to have theft intention. In the whole system, based on the frame where PA1 is located (i.e., the target image), based on historical and future video frames (i.e., the associated video frame sequence), key frame sampling will be performed, and continuous frames (i.e., the associated video frame sequence) will be sent to the multimodal large model for continuous behavior sequence judgment.
[0253] e. Finally, combining the arbitrary financial state ZA1 (i.e., the second state information mentioned above) and the results of the multimodal large model (i.e., the initial theft detection information mentioned above), a final judgment (i.e., the final theft detection information mentioned above) is given. If the state of any property changes (for example, the state represented by the first state information is different from the state represented by the second state information), and the large model gives a judgment of theft, then theft is considered to have occurred.
[0254] 4. If the intent judgment in the second stage is Yes (that is, the final theft judgment information above indicates that the target person has the intent to steal the target item), then the third stage is entered. In this stage, a multimodal large model is used to analyze the person's behavior, trajectory and interaction with the object, and push event messages to the user, that is, send prompt information to the preset terminal.
[0255] 5. The same logic applies to vehicle guarding and guarding any property.
[0256] III. User Interaction
[0257] 1. In the first phase, the settings are fully customized for users with different identities. When talking to family members, the voice message is "Welcome home". When talking to strangers, deliverymen and people whose identities are not identified, the voice message is "How can I help you?"
[0258] 2. In the second stage, based on the intent judgment results and combined with identity information, customized voice broadcasts are made to fully match the identity of different intents. The voice broadcast is "Please pick up the package in time" to family members, "Put down my package, I'm leaving right away" to strangers, "Thank you for delivery" to deliverymen, and "Wait a moment, I'm leaving right away" to people whose identities have not been identified.
[0259] In the third stage, the results of the multimodal large model (i.e., the initial theft detection information), the original state of the package (i.e., the first state information), and the state of the package after the time expires (i.e., the second state information) are combined to provide customized push messages. For example, for strangers, if the package exists initially but disappears after the time expires, a warning message "The package has been stolen" is pushed.
[0260] It should be noted that, in addition to the contents described above, this embodiment may also include the technical features described in the above embodiments, thereby achieving the technical effect of the above-described method for identifying theft intent. For details, please refer to the above description. For the sake of brevity, it will not be elaborated here.
[0261] The theft intent identification method provided in this application embodiment achieves personalized settings for each user: users can customize settings, such as setting hotspot ranges around their doorstep or vehicle in the app. Furthermore, the system supports learning property definitions and guarded areas based on user behavior habits. The system can respond in a customized manner according to user settings; for example, if a stranger enters the warning range and attempts to steal a package or pry open a car lock, the system will issue a voice warning, automatically sound an alarm, and notify the user. This improves user experience and system usability. In addition, this solution allows for real-time processing and rapid response: employing an "edge + terminal" distributed computing approach maximizes the efficiency of computing resources. A small video classification model is used for initial screening; when the small model determines abnormal behavior in the current frame, a large multimodal model is triggered for fine-grained logical judgment. This method not only saves computing resources but also ensures response speed, improving efficiency and accuracy. Real-time analysis and processing of video streams can be performed, and timely notifications can be sent to users. Furthermore, this solution combines multimodal large-scale data: based on over 40 billion internet graphs and texts, the MLLM model is pre-trained and fine-tuned using data from the security vertical industry. By combining extensive internet knowledge with the inference capabilities of LLM, it can accurately understand multiple behavioral trajectories and precisely push user messages. Based on the identity of the person approaching the camera, a welcome message is accurately pushed. When intent judgment is triggered, MLLM model inference is performed based on the frames before and after the intent judgment, enabling precise analysis of user behavior and personalized push notifications. A hybrid "edge + terminal" multi-model system integrates a multimodal large model and a video classification small model into an end-to-end system. It processes video streams in real time, performing multi-step judgments and understanding based on video content, achieving highly accurate user message pushes while maintaining extreme performance. For the unique needs of home security, a theft intent prediction solution based on an AI (Artificial Intelligence) "edge + terminal" architecture has been launched. This solution not only deeply integrates high-precision sensors (such as cameras, depth sensors, and LiDAR), smart cameras, and deep learning algorithms, but also achieves accurate identification of abnormal behavior through continuous learning of the surrounding environment. It enables "personalized security protection," meaning that algorithms are optimized and interfaces are customized based on the specific environment of each family, truly achieving "one policy per family" and providing a unique property protection experience tailored to each household. Adopting an "edge + terminal" architecture, edge computing is equipped with real-time, efficient target detection algorithms and intent judgment, enabling rapid response to interactions between people and their surroundings. Customized voice prompts are used to guide or discourage potential behaviors. The terminal is intelligent, equipped with a high-performance computing chip and deploys MLLM (Multimodal Large Model) with inference capabilities to perform inference on real-time video streams. Based on a large amount of pre-trained data, MLLM exhibits strong generalization ability in behavioral understanding.By combining a small CNN model with a large MLLM multimodal model, the system's performance can be maximized, and the accuracy of behavior understanding is significantly improved compared to the small CNN model, with both generalization ability and accuracy being greatly enhanced. "Personalized Family Planning" combines OVOD, Open Vocabulary Object Detection, and MLLM to perform structured analysis of family member behavior, and integrates user settings to protect customized assets for each family member.
[0262] Figure 4 This is a schematic diagram of a device for identifying theft intent provided in an embodiment of this application. Specifically, it includes:
[0263] The first acquisition unit 401 is used to acquire a target image, wherein the target image includes an object image and a person image, the object image represents a target object, and the person image represents a target person;
[0264] The first determining unit 402 is used to determine a first detection box and a second detection box in the target image, wherein the first detection box is the detection box of the object image and the second detection box is the detection box of the person image.
[0265] The second determining unit 403 is used to determine the degree of overlap between the first detection box and the second detection box;
[0266] The generation unit 404 is used to generate theft identification information based on the degree of overlap, wherein theft identification information indicates whether the target person has the intention to steal the target item.
[0267] In one possible implementation, generating theft detection information based on the degree of overlap includes:
[0268] Determine whether the degree of overlap is greater than or equal to a preset threshold;
[0269] If the degree of overlap is greater than or equal to the preset threshold, determine whether the behavior of the target person represented by the personnel image is theft, so as to obtain a first determination result;
[0270] Based on the first determination result, theft detection information is generated.
[0271] In one possible implementation, the target image is a video frame from a target video; and the generation of theft detection information based on the degree of overlap includes:
[0272] Determine whether the degree of overlap is greater than or equal to a preset threshold;
[0273] If the degree of overlap is greater than or equal to the preset threshold, extract the associated video frame sequence of the target image from the target video;
[0274] Theft detection information is generated based on the associated video frame sequence.
[0275] In one possible implementation, the apparatus is applied to a first device, wherein the data processing amount of the target image is less than the data processing amount of the video frames in the associated video frame sequence; and
[0276] The process of generating theft detection information based on the associated video frame sequence includes:
[0277] The associated video frame sequence is sent to a second device; wherein the second device is used to generate theft detection information based on the associated video frame sequence; the computing power of the second device is greater than that of the first device;
[0278] The theft identification information returned by the second device is received to generate the theft identification information.
[0279] In one possible implementation, the associated video frame sequence includes: a preceding video frame and a following video frame of the target image; the preceding video frame is a video frame in the target video that precedes the target image, and the following video frame is a video frame in the target video that follows the target image; and
[0280] The process of generating theft detection information based on the associated video frame sequence includes:
[0281] Determine the first state information of the target item in the previous video frame;
[0282] Determine the second state information of the target item in the subsequent video frame;
[0283] Based on the first state information, the second state information, and the associated video frame sequence, theft detection information is generated.
[0284] In one possible implementation, the first state information indicates whether the target item represented by the item image in the previous video frame is located in the first region, and the second state information indicates whether the target item represented by the item image in the subsequent video frame is located in the first region; and
[0285] The process of generating theft detection information based on the first state information, the second state information, and the associated video frame sequence includes:
[0286] If the first state information indicates that the target item represented by the item image in the previous video frame is located in the first area, initial theft detection information is generated based on the associated video frame sequence.
[0287] If the initial theft identification information indicates that the target person has the intention to steal the target item, and the second state information indicates that the target item represented by the item image in the subsequent video frame is not located in the first area, then final theft identification information indicating that the target person has the intention to steal the target item is generated.
[0288] In one possible implementation, generating theft detection information based on the degree of overlap includes:
[0289] Determine whether both the target person and the target item are located in a preset area to obtain a second determination result;
[0290] Based on the second determination result and the degree of overlap, theft detection information is generated.
[0291] In one possible implementation, before determining the degree of overlap between the first detection frame and the second detection frame, the device further includes:
[0292] The second acquisition unit (not shown in the figure) is used to acquire a pre-entered set of personnel information;
[0293] The third determining unit (not shown in the figure) is used to determine whether the personnel information set includes target personnel information representing the target personnel; and
[0294] Determining the degree of overlap between the first detection box and the second detection box includes:
[0295] If the target personnel information is not included in the personnel information set, the degree of overlap between the first detection box and the second detection box is determined.
[0296] In one possible implementation, after generating the theft detection information, the device further includes at least one of the following:
[0297] A control unit (not shown in the figure) is used to control the expulsion device to perform an expulsion operation when the theft identification information indicates that the target person has the intent to steal;
[0298] The sending unit (not shown in the figure) is used to send a prompt message to a preset terminal when the theft identification information indicates that the target person has the intention to steal.
[0299] The theft intent identification device provided in this embodiment can be as follows: Figure 4 The theft intent identification device shown can execute all the steps of the theft intent identification methods described above, thereby achieving the technical effects of the theft intent identification methods described above. For details, please refer to the relevant descriptions above. For the sake of brevity, it will not be elaborated here.
[0300] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 5 The illustrated electronic device 500 includes at least one processor 501, a memory 502, at least one network interface 504, and other user interfaces 503. The various components in the electronic device 500 are coupled together via a bus system 505. It is understood that the bus system 505 is used to implement communication between these components. In addition to a data bus, the bus system 505 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 5 The general designated all buses as Bus System 505.
[0301] The user interface 503 may include a display, keyboard, or clicking device (e.g., mouse, trackball, touchpad, or touchscreen).
[0302] It is understood that the memory 502 in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDRSDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The memory 502 described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0303] In some implementations, memory 502 stores elements, executable units or data structures, or subsets thereof, or extended sets thereof: operating system 5021 and application program 5022.
[0304] The operating system 5021 includes various system programs, such as the framework layer, core library layer, and driver layer, used to implement various basic business functions and handle hardware-based tasks. The application program 5022 includes various applications, such as a media player and a browser, used to implement various application functions. Programs implementing the methods of this application embodiment can be included in application program 5022.
[0305] In this embodiment, by calling the program or instructions stored in memory 502, specifically the program or instructions stored in application program 5022, processor 501 executes the method steps provided in each method embodiment, including, for example:
[0306] Acquire a target image, wherein the target image includes an object image and a person image, the object image representing a target object and the person image representing a target person;
[0307] A first detection box and a second detection box are determined in the target image, wherein the first detection box is the detection box of the object image and the second detection box is the detection box of the person image;
[0308] Determine the degree of overlap between the first detection box and the second detection box;
[0309] Based on the degree of overlap, theft detection information is generated, wherein theft detection information indicates whether the target person has the intention to steal the target item.
[0310] The methods disclosed in the embodiments of this application can be applied to or implemented by processor 501. Processor 501 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 501 or by instructions in the form of software. The processor 501 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software units in the decoding processor. The software units may be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 502. Processor 501 reads the information in memory 502 and, in conjunction with its hardware, completes the steps of the above method.
[0311] It is understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described above in this application, or combinations thereof.
[0312] For software implementation, the techniques described herein can be implemented by units that perform the functions described above. The software code can be stored in memory and executed by a processor. The memory can be implemented within the processor or external to the processor.
[0313] The electronic device provided in this embodiment may be as follows: Figure 5 The electronic device shown can execute all the steps of the above-described methods for identifying theft intent, thereby achieving the technical effects of the above-described methods for identifying theft intent. For details, please refer to the above descriptions. For the sake of brevity, further details are omitted here.
[0314] This application also provides a storage medium (computer-readable storage medium). This storage medium stores one or more programs. The storage medium may include volatile memory, such as random access memory; it may also include non-volatile memory, such as read-only memory, flash memory, hard disk, or solid-state drive; and it may also include combinations of the above types of memory.
[0315] When one or more programs in the storage medium can be executed by one or more processors, the above-mentioned method for identifying theft intent executed on the electronic device side can be implemented.
[0316] The processor described above is used to execute a theft intent identification program stored in memory to implement the following steps of the theft intent identification method executed on the electronic device side:
[0317] Acquire a target image, wherein the target image includes an object image and a person image, the object image representing a target object and the person image representing a target person;
[0318] A first detection box and a second detection box are determined in the target image, wherein the first detection box is the detection box of the object image and the second detection box is the detection box of the person image;
[0319] Determine the degree of overlap between the first detection box and the second detection box;
[0320] Based on the degree of overlap, theft detection information is generated, wherein theft detection information indicates whether the target person has the intention to steal the target item.
[0321] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0322] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0323] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also include the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.
[0324] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for identifying intent to steal, characterized in that, The method includes: Acquire a target image, wherein the target image includes an object image and a person image, the object image representing a target object and the person image representing a target person; A first detection box and a second detection box are determined in the target image, wherein the first detection box is the detection box of the object image and the second detection box is the detection box of the person image; Determine the degree of overlap between the first detection box and the second detection box; Based on the degree of overlap, theft detection information is generated, wherein theft detection information indicates whether the target person has the intention to steal the target item.
2. The method according to claim 1, characterized in that, The process of generating theft detection information based on the degree of overlap includes: Determine whether the degree of overlap is greater than or equal to a preset threshold; If the degree of overlap is greater than or equal to the preset threshold, determine whether the behavior of the target person represented by the personnel image is theft, so as to obtain a first determination result; Based on the first determination result, theft detection information is generated.
3. The method according to claim 1, characterized in that, The target image is a video frame from the target video; as well as The process of generating theft detection information based on the degree of overlap includes: Determine whether the degree of overlap is greater than or equal to a preset threshold; If the degree of overlap is greater than or equal to the preset threshold, extract the associated video frame sequence of the target image from the target video; Theft detection information is generated based on the associated video frame sequence.
4. The method according to claim 3, characterized in that, The method is applied to a first device, where the amount of data processing for the target image is less than the amount of data processing for the video frames in the associated video frame sequence. as well as The process of generating theft detection information based on the associated video frame sequence includes: The associated video frame sequence is sent to a second device; wherein the second device is used to generate theft detection information based on the associated video frame sequence; the computing power of the second device is greater than that of the first device; The theft identification information returned by the second device is received to generate the theft identification information.
5. The method according to claim 3, characterized in that, The associated video frame sequence includes: a preceding video frame and a following video frame of the target image; the preceding video frame is the video frame in the target video that precedes the target image, and the following video frame is the video frame in the target video that follows the target image; and The process of generating theft detection information based on the associated video frame sequence includes: Determine the first state information of the target item in the previous video frame; Determine the second state information of the target item in the subsequent video frame; Based on the first state information, the second state information, and the associated video frame sequence, theft detection information is generated.
6. The method according to claim 5, characterized in that, The first status information indicates whether the target item represented by the item image in the previous video frame is located in the first region, and the second status information indicates whether the target item represented by the item image in the subsequent video frame is located in the first region. as well as The process of generating theft detection information based on the first state information, the second state information, and the associated video frame sequence includes: If the first state information indicates that the target item represented by the item image in the previous video frame is located in the first area, initial theft detection information is generated based on the associated video frame sequence. If the initial theft identification information indicates that the target person has the intention to steal the target item, and the second state information indicates that the target item represented by the item image in the subsequent video frame is not located in the first area, then final theft identification information indicating that the target person has the intention to steal the target item is generated.
7. The method according to claim 1, characterized in that, The process of generating theft detection information based on the degree of overlap includes: Determine whether both the target person and the target item are located in a preset area to obtain a second determination result; Based on the second determination result and the degree of overlap, theft detection information is generated.
8. The method according to any one of claims 1-7, characterized in that, Before determining the degree of overlap between the first detection box and the second detection box, the method further includes: Retrieve the pre-entered set of personnel information; Determine whether the personnel information set includes target personnel information representing the target personnel; and Determining the degree of overlap between the first detection box and the second detection box includes: If the target personnel information is not included in the personnel information set, the degree of overlap between the first detection box and the second detection box is determined.
9. The method according to any one of claims 1-7, characterized in that, After generating the theft detection information, the method further includes at least one of the following: If the theft detection information indicates that the target person has the intent to steal, the expulsion device is controlled to perform an expulsion operation. If the theft detection information indicates that the target person has the intent to steal, a prompt message is sent to a preset terminal.
10. A device for identifying intent to steal, characterized in that, The device includes: The first acquisition unit is used to acquire a target image, wherein the target image includes an object image and a person image, the object image representing a target object and the person image representing a target person; The first determining unit is used to determine a first detection box and a second detection box in the target image, wherein... The first detection box is the detection box for the image of the item, and the second detection box is the detection box for the image of the person; The second determining unit is used to determine the degree of overlap between the first detection box and the second detection box; The generation unit is used to generate theft identification information based on the degree of overlap, wherein theft identification information indicates whether the target person has the intention to steal the target item.