Security and protection method and device and electronic equipment
By acquiring and recognizing security videos and combining them with security preference information, the system dynamically determines and executes matching security response operations, thus solving the problem of insufficient intelligence in existing security systems and achieving more efficient security responses.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANKER INNOVATIONS TECH CO LTD
- Filing Date
- 2024-10-31
- Publication Date
- 2026-05-01
AI Technical Summary
Existing security systems lack intelligent monitoring and have limited response methods, failing to meet the growing security needs of modern families.
By acquiring and identifying security videos, and combining them with pre-determined security preference information, the system dynamically determines and executes security response actions that match the videos, utilizing smart devices for targeted responses.
This improves the targeting and effectiveness of security systems, reduces false alarms and missed alarms, and ensures a rapid and accurate response to various emergencies.
Smart Images

Figure CN121963366A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of security, and more particularly to a security method, device, and electronic device. Background Technology
[0002] In related technologies, some security systems mainly rely on simple monitoring and alarm devices, lacking intelligent monitoring and failing to meet the growing security needs of modern families. Furthermore, their response methods are limited, typically relying on fixed alarm methods and fixed security equipment to provide simple alerts to users.
[0003] Clearly, improving the targeting and effectiveness of security measures is a technical issue that deserves attention. Summary of the Invention
[0004] In view of this, in order to solve some or all of the above-mentioned technical problems, embodiments of this application provide a security method, device and electronic device.
[0005] In a first aspect, embodiments of this application provide a security method, the method comprising:
[0006] Acquire the first security video and pre-determined security preference information;
[0007] The first security video is identified to obtain the video identification result of the first security video;
[0008] Based on the video recognition results and the security preference information, a security response operation matching the first security video is determined to obtain the first security response operation;
[0009] Identify the security device used to perform the first security response operation;
[0010] Control the security device to execute the first security response operation.
[0011] Secondly, embodiments of this application provide a security device, the device comprising:
[0012] The first acquisition unit is used to acquire the first security video and the pre-determined security preference information;
[0013] The first identification unit is used to identify the first security video to obtain the video identification result of the first security video;
[0014] The first determining unit is used to determine a security response operation that matches the first security video based on the video recognition result and the security preference information, so as to obtain the first security response operation;
[0015] The second determining unit is used to determine the security device used to perform the first security response operation;
[0016] The first control unit is used to control the security device to perform the first security response operation.
[0017] Thirdly, embodiments of this application provide an electronic device, including:
[0018] Memory, used to store computer programs;
[0019] A processor is configured to execute a computer program stored in the memory, wherein, when the computer program is executed, it implements the method of any embodiment of the security method of the first aspect of this application.
[0020] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, it implements the method of any embodiment of the security method of the first aspect described above.
[0021] The security method provided in this application embodiment can acquire a first security video and pre-determined security preference information. Then, the first security video is identified to obtain a video recognition result. Subsequently, based on the video recognition result and the security preference information, a security response operation matching the first security video is determined to obtain a first security response operation. Then, a security device for executing the first security response operation is determined. Finally, the security device is controlled to execute the first security response operation. In this way, based on the specific event represented by the video recognition result of the security video and the pre-determined security preference information, a security response operation matching the security video can be dynamically determined and executed. Each response operation is determined for the specific current situation, thereby ensuring that appropriate measures can be taken for different types of events. Furthermore, the security device for executing the first security response operation can be dynamically determined according to different situations, thus enabling faster and more accurate responses to various emergencies, reducing false alarms and missed alarms, and improving the overall effectiveness of the security system. Attached Figure Description
[0022] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0023] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0025] Figure 1 A flowchart illustrating a security method provided in an embodiment of this application;
[0026] Figure 2 A flowchart illustrating another security method provided in an embodiment of this application;
[0027] Figure 3A A three-dimensional panoramic view of a house provided as an embodiment of the security method of this application;
[0028] Figure 3B A schematic diagram of a knowledge graph for a security method provided in an embodiment of this application;
[0029] Figure 3C This is a schematic diagram illustrating an application scenario of a security method provided in an embodiment of this application;
[0030] Figure 3D An architecture diagram of a security method provided in an embodiment of this application;
[0031] Figure 3E A schematic diagram of an environmental file for a security method provided in an embodiment of this application;
[0032] Figure 3F A schematic diagram of a personal file and family relationship map provided for a security method according to an embodiment of this application;
[0033] Figure 3G A schematic diagram illustrating a multi-modal interaction entry point for a security method provided in this application embodiment;
[0034] Figure 3H A schematic diagram of a visitor reception module for a security method provided in an embodiment of this application;
[0035] Figure 3I A schematic diagram of a family guard module in a security method provided in an embodiment of this application;
[0036] Figure 3J A schematic diagram of a property guarding module of a security method provided in an embodiment of this application;
[0037] Figure 3K A schematic diagram illustrating an artificial intelligence generation service for a security method provided in an embodiment of this application;
[0038] Figure 3LA schematic diagram illustrating a user-generated service for a security method provided in this application embodiment;
[0039] Figure 4A A schematic diagram of the knowledge graph in a dynamic response method for security video provided in an embodiment of this application;
[0040] Figure 4B A schematic diagram illustrating the generation method of an instruction to perform a security response operation in a dynamic response method for security video provided in an embodiment of this application;
[0041] Figure 4C A schematic diagram of user settings in a dynamic response method for security video provided in an embodiment of this application;
[0042] Figure 4D A schematic diagram of event recognition in a dynamic response method for security video provided in an embodiment of this application;
[0043] Figure 4E A schematic diagram illustrating behavior prediction in a dynamic response method for security videos provided in this application embodiment;
[0044] Figure 4F A schematic diagram illustrating the decision execution in a dynamic response method for security video provided in an embodiment of this application;
[0045] Figure 4G A schematic diagram illustrating information sharing in a dynamic response method for security video provided in an embodiment of this application;
[0046] Figure 5A This is a schematic diagram illustrating an application scenario of an object behavior recognition method provided in an embodiment of this application;
[0047] Figure 5B A schematic diagram of a captured image provided in an embodiment of this application;
[0048] Figure 5C A schematic diagram illustrating another captured image provided in an embodiment of this application;
[0049] Figure 5D This is a schematic diagram of the structure of an object behavior recognition system provided in an embodiment of this application;
[0050] Figure 5E A flowchart illustrating an embodiment of an object behavior recognition method provided in this application;
[0051] Figure 6A A flowchart illustrating a device linkage monitoring operation method provided in an embodiment of this application;
[0052] Figure 6BA flowchart illustrating the construction of a three-dimensional model in a device linkage monitoring operation method provided in this application embodiment;
[0053] Figure 6C A schematic diagram of the intersection area between the first monitoring area and the second monitoring area in a device linkage monitoring operation method provided in an embodiment of this application;
[0054] Figure 7 A flowchart illustrating a method for identifying theft intent provided in an embodiment of this application;
[0055] Figure 8 A flowchart illustrating a video frame extraction method provided in an embodiment of this application;
[0056] Figure 9 This is a schematic diagram illustrating an application scenario of an information push method provided in an embodiment of this application;
[0057] Figure 10A This is a schematic diagram illustrating an application scenario of a linkage control strategy generation method provided in an embodiment of this application.
[0058] Figure 10B A flowchart illustrating the framework of a linkage control strategy generation method provided in this application embodiment;
[0059] Figure 11A A flowchart illustrating a control method for a security device provided in an embodiment of this application;
[0060] Figures 11B-11D This is a schematic diagram illustrating an application scenario of stranger tracking and removal in a control method for a security device provided in an embodiment of this application.
[0061] Figures 11E-11G This is a schematic diagram illustrating an application scenario of tracking and removing suspicious vehicles in a security equipment control method provided in this application embodiment.
[0062] Figure 12 This is a schematic diagram of the structure of a security device provided in an embodiment of this application;
[0063] Figure 13 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0064] Various exemplary embodiments of this application will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of this application.
[0065] Those skilled in the art will understand that the terms "first" and "second" in the embodiments of this application are only used to distinguish different steps, devices or modules, and do not represent any specific technical meaning, nor do they indicate the logical order between them.
[0066] It should also be understood that in this embodiment, "multiple" can refer to two or more, and "at least one" can refer to one, two or more.
[0067] It should also be understood that any component, data or structure mentioned in the embodiments of this application can generally be understood as one or more unless explicitly defined or given contrary guidance in the context.
[0068] Furthermore, the term "and / or" in this application is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this application generally indicates that the preceding and following related objects have an "or" relationship.
[0069] It should also be understood that the description of the various embodiments in this application emphasizes the differences between the various embodiments, and the similarities or similarities can be referred to each other. For the sake of brevity, they will not be described in detail.
[0070] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the scope of this application and its application or use.
[0071] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0072] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0073] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. To facilitate understanding of the embodiments of this application, the application will be described in detail below with reference to the accompanying drawings and embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0074] Furthermore, it should be noted that the users described in this application can be distinguished by user identifiers. For example, a user identifier can be a login account. In this scenario, if different people log in using the same account, they can be considered the same user; if the same person logs in using different accounts, they can be considered different users. As another example, when a device is not logged in, a user identifier can be assigned based on the device's identifier. In this scenario, if different people operate using devices with the same device identifier, they can be considered the same user; if the same person operates using devices with different device identifiers, they can be considered different users.
[0075] To address the technical problem of how to improve the targeting and effectiveness of security in existing technologies, this application provides a dynamic response method and apparatus for security videos, which can improve the targeting and effectiveness of security.
[0076] Figure 1 This is a flowchart illustrating a security method provided in an embodiment of this application. Embodiments of this disclosure can be applied to one or more electronic devices such as security equipment, smartphones, laptops, desktop computers, portable computers, and servers. Furthermore, the executing entity of embodiments of this disclosure can be hardware or software. When the executing entity is hardware, it can be one or more of the aforementioned electronic devices. For example, a single electronic device can execute embodiments of this disclosure, or multiple electronic devices can cooperate with each other to execute embodiments of this disclosure. When the executing entity is software, embodiments of this disclosure can be implemented as multiple software programs or software modules, or as a single software program or software module. No specific limitations are imposed here.
[0077] like Figure 1 As shown, the method specifically includes:
[0078] Step 101: Obtain the first security video and the pre-determined security preference information.
[0079] In this embodiment, security video can be surveillance video used to ensure security and prevent risks. In practice, security video can be captured by a camera.
[0080] The first security video can be any security video.
[0081] Security preference information can represent the security preferences and needs of users and other entities. For example, security preference information may include at least one of the following: events, people, actions that users or other entities are concerned about, the manner in which they perform security response actions, the frequency of performing security response actions, and the time at which they perform security response actions.
[0082] Step 102: Identify the first security video to obtain the video identification result of the first security video.
[0083] In this embodiment, the video recognition result can be the result obtained by recognizing the first security video. As an example, the video result can indicate whether the current behavior of the target object belongs to the target behavior type (e.g., theft, dangerous behavior). In addition, the video result can indicate the person recognition result of the target object, etc.
[0084] In this embodiment, Support Vector Machine (SVM), multimodal models, etc., can be used to identify the first security video, thereby obtaining the video recognition result of the first security video.
[0085] As an example, video recognition results can represent events such as theft, fire, and falls.
[0086] Step 103: Based on the video recognition result and the security preference information, determine the security response operation that matches the first security video to obtain the first security response operation.
[0087] In this embodiment, the security response operation can be a security video response operation.
[0088] The first security response operation can be a security response operation that matches the first security video, determined based on the video recognition result and the security preference information.
[0089] As an example, a pre-determined first correspondence table can be used to determine the security response operation that corresponds to the video recognition result and the security preference information. This security response operation is then identified as the security response operation that matches the first security video, i.e., the first security response operation. The aforementioned first correspondence table can represent the correspondence between the video recognition result, the security preference information, and the security response operation.
[0090] As another example, the video recognition result and the security preference information can also be input into a pre-trained first model to obtain a security response operation, and this security response operation can be identified as the security response operation that matches the first security video, i.e., the first security response operation. The first model can represent the correspondence between the video recognition result, the security preference information, and the security response operation. The first model can be a convolutional neural network or a large language model trained using machine learning algorithms based on training samples containing the video recognition result, security preference information, and the security response operation.
[0091] Step 104: Determine the security device to perform the first security response operation.
[0092] In this embodiment, after determining the first security response operation, the security device used to perform the first security response operation can be further determined.
[0093] In some cases, security equipment can be smart devices. Examples include cameras, alarms, and access control systems.
[0094] In practice, the security equipment used to perform the first security response operation can be determined based on the correspondence table that represents the correspondence between the first security response operation and the security equipment.
[0095] Step 105: Control the security device to execute the first security response operation.
[0096] In this embodiment, after determining the security response operation that matches the first security video, the security device can be further controlled to execute the first security response operation.
[0097] In some optional implementations of this embodiment, after identifying the first security video to obtain the video recognition result of the first security video, the following steps may also be performed:
[0098] The first step is to obtain first feedback information regarding the video recognition result, wherein the first feedback information indicates an adjustment to the recognition strategy of the second security video, and the recognition strategy includes at least one of the following: recognition efficiency and recognition method.
[0099] The second security video is either the first security video or a security video acquired after the first security video.
[0100] The identification strategy is a strategy for identifying the second security video to obtain a video identification result of the second security video. For example, the identification strategy may be: identifying the second security video with efficiency A to obtain a video identification result of the second security video, identifying the second security video with efficiency B to obtain a video identification result of the second security video, identifying the second security video with identification method A to obtain a video identification result of the second security video, identifying the second security video with identification method B to obtain a video identification result of the second security video, etc.
[0101] As an example, the initial feedback information can be represented by text, voice, or other means. In some cases, the initial feedback information can be determined by an object such as a user.
[0102] Here, if the first feedback message indicates that "the efficiency of recognizing security videos is too slow," then the efficiency of recognizing security videos can be improved, thereby recognizing subsequent security videos more efficiently. Alternatively, the security videos that have already been recognized can be re-recognized (such as the first security video obtained in step 101 above). In this case, the recognition strategy can be expressed as "recognizing the second security video with a higher efficiency than the current recognition of the security video, so as to obtain the video recognition result of the second security video."
[0103] Specifically, before obtaining the first feedback information, if an event of "person approaching an item - bending down to pick up the item - taking the item" is identified in the security video, and the video identification result is determined to be "theft", then after obtaining the first feedback information indicating that "the efficiency of identifying security videos is too slow", if an event of "person approaching an item - bending down to pick up the item" is identified in the security video, the video identification result can be determined to be "theft", in order to improve the efficiency of identifying security videos.
[0104] In addition, if the first feedback message indicates that "the efficiency of identifying security videos is too slow", then the identification method of security videos can be changed, for example, changing the current identification method A of security videos to another identification method B.
[0105] The second step is to determine the identification strategy to be adjusted according to the first feedback information.
[0106] The third step is to identify the second security video according to the identification strategy adjusted according to the first feedback information, so as to obtain the video identification result of the second security video.
[0107] Here, the security video obtained afterward can be identified according to the identification strategy adjusted according to the first feedback information, so as to obtain the video identification result of the subsequent security video; or the security video that has been identified (such as the first security video obtained in step 101) can be identified, so as to obtain the video identification result of the security video that has been identified.
[0108] The fourth step is to determine the security response operation that matches the second security video based on the video recognition results and security preference information of the second security video, so as to obtain the second security response operation.
[0109] The second security response operation can represent a security response operation that matches the security video, determined based on the video recognition results and security preference information of the security video.
[0110] Fifth step, execute the second security response operation.
[0111] Here, the execution methods for steps four and five can refer to steps 103 and 104 above. For the sake of brevity, they will not be repeated here.
[0112] It is understandable that, among the above-mentioned optional implementation methods, the security video recognition strategy can be dynamically adjusted based on the feedback information of the video recognition results. This can meet the different recognition needs of different users or the same user at different times, thereby improving user satisfaction with the security effect.
[0113] In some optional implementations of this embodiment, after determining the security response operation matching the first security video based on the video recognition result and the security preference information to obtain the first security response operation, the following steps may be further performed:
[0114] The first step is to obtain second feedback information for the first security response operation, wherein the second feedback information indicates the determination strategy for adjusting the second security response operation, and the determination strategy includes at least one of the following: determining efficiency and determining method.
[0115] The second security video is either the first security video or a security video acquired after the first security video.
[0116] The determination strategy is as follows: based on video recognition results and security preference information, determine the security response operation that matches the second security video to obtain the strategy for the third security response operation. For example, the determination strategy could be: using efficiency A to determine the security response operation that matches the second security video based on video recognition results and security preference information to obtain the third security response operation; using efficiency B to determine the security response operation that matches the second security video based on video recognition results and security preference information to obtain the third security response operation; using determination method A to determine the security response operation that matches the second security video based on video recognition results and security preference information to obtain the third security response operation; using determination method B to determine the security response operation that matches the second security video based on video recognition results and security preference information to obtain the third security response operation, etc.
[0117] As an example, the second feedback information can be represented in the form of text, voice, or other means. In some cases, the second feedback information can be determined by an object such as a user.
[0118] Here, if the second feedback indicates that "the efficiency of determining the security response operation is too slow," then the efficiency of determining the security response operation can be improved, thereby determining subsequent security response operations more efficiently. Alternatively, the security response operation that has already been determined can be re-determined (e.g., the security response operation determined in step 103 above). Specifically, if the security response operation is determined serially before obtaining the second feedback, then after obtaining the second feedback indicating that "the efficiency of determining the security response operation is too slow," the security response operation can be determined in parallel. In this case, the determination strategy can be expressed as "obtaining the third security response operation by determining the security response operation matching the first security video with a higher efficiency than the current efficiency based on video recognition results and security preference information."
[0119] Furthermore, if the second feedback message indicates that "the efficiency of determining the security response operation is too slow," then the method of determining the security response operation can be changed, for example, changing the current method A of determining the security response operation to another method B.
[0120] The second step is to determine the adjustment strategy indicated by the second feedback information.
[0121] The third step involves determining the security response operation that matches the second security video based on the video recognition results and security preference information, according to the adjustment strategy indicated by the second feedback information, so as to obtain the third security response operation.
[0122] Here, the subsequent security response operation can be determined according to the determination strategy indicated by the second feedback information; or the security response operation that has been determined (such as the security response operation determined in step 103) can be re-determined.
[0123] The third security response operation can refer to a security response operation that matches the security video, determined based on the video recognition results and security preference information, according to the determined strategy adjusted according to the second feedback information.
[0124] The fourth step is to execute the third security response operation.
[0125] It is understandable that, among the above-mentioned optional implementation methods, the determination strategy of security response operation can be dynamically adjusted through the feedback information of security response operation. In this way, the different needs of different users or the same user at different times can be met, thereby improving the user's satisfaction with the security effect.
[0126] In some optional implementations of this embodiment, after performing the first security response operation, the following steps may also be performed:
[0127] The first step is to obtain third feedback information for the first security response operation, wherein the third feedback information indicates the adjustment of the execution strategy of the security response operation, and the execution strategy includes at least one of the following: execution efficiency and execution method.
[0128] The execution strategy refers to the strategy for performing the fourth security response operation, which is a security response operation performed after the first security response operation. For example, the execution strategy could be: perform the fourth security response operation with efficiency A; perform the fourth security response operation with efficiency B, etc.
[0129] As an example, third-party feedback can be represented using text, voice, or other methods. In some cases, third-party feedback can be determined by objects such as users.
[0130] Here, if the third feedback indicates that "the execution efficiency of the security response operation is too slow," then the execution efficiency of the security response operation can be improved, thereby executing subsequent security response operations more efficiently. Specifically, if the security response operation is executed 10 seconds after it is determined before receiving the third feedback, then after receiving the third feedback indicating "the execution efficiency of the security response operation is too slow," it can be executed 8 seconds after it is determined. In this case, the execution strategy can be expressed as "execute the fourth security response operation with a higher efficiency than the current execution of the first security response operation."
[0131] Furthermore, if the third feedback indicates that "the execution efficiency of the security response operation is too slow", then the execution method of the security response operation can be changed, for example, changing the current execution method A of the security response operation to another execution method B.
[0132] The second step is to determine the execution strategy indicated by the third feedback information.
[0133] The third step is to execute the fourth security response operation according to the execution strategy adjusted according to the third feedback information.
[0134] The fourth security response operation is a security response operation performed after the first security response operation.
[0135] Here, the security response operations determined afterward can be executed according to the execution strategy adjusted according to the third feedback information.
[0136] It is understandable that, among the above-mentioned optional implementation methods, the execution strategy of the security response operation can be dynamically adjusted through the feedback information of the security response operation. In this way, the different execution needs of different users or the same user at different times can be met, thereby improving the user's satisfaction with the security effect.
[0137] In some optional implementations of this embodiment, the security response operation matching the first security video can be determined based on the video recognition result and the security preference information in the following manner:
[0138] The first step is to determine the response probability of the first security video based on the video recognition results.
[0139] As an example, a predetermined fourth correspondence table can be used to determine the response probability that corresponds to the video recognition result, thereby determining the response probability as the response probability of the first security video. The aforementioned fourth correspondence table can represent the correspondence between the video recognition result and the response probability.
[0140] As another example, the video recognition result can be input into a pre-trained fourth model to obtain a response probability, and this response probability can be determined as the response probability of the first security video. The fourth model can represent the correspondence between the recognition result and the response probability. This fourth model can be a convolutional neural network or a large language model trained using machine learning algorithms based on training samples containing the recognition result and response probability.
[0141] The second step is to determine whether the response probability is greater than or equal to a preset threshold.
[0142] Third, if the response probability is greater than or equal to the preset threshold, determine the security response operation that matches the first security video based on the security preference information.
[0143] It is understood that in the above-mentioned optional implementation methods, a security response operation matching the first security video is determined based on the security preference information only when the response probability is greater than or equal to the preset threshold, and then the security response operation is executed. This can avoid the security response operation being executed too frequently.
[0144] In some application scenarios of the above-mentioned optional implementation methods, after the first security response operation is performed, the preset threshold is adjusted in the following manner:
[0145] The first step is to obtain the fourth feedback information in response to the first security response operation.
[0146] The fourth feedback information indicates that the preset threshold should be adjusted.
[0147] As an example, the fourth feedback information can be represented in the form of text, voice, or other means. In some cases, the fourth feedback information can be determined by an object such as a user.
[0148] The second step is to adjust the preset threshold according to the adjustment method indicated by the fourth feedback information to obtain the adjusted preset threshold.
[0149] Based on this, the following steps can also be performed:
[0150] The third step is to determine the security response operation that matches the second security video based on the security preference information when the response probability is greater than or equal to the adjusted preset threshold, so as to obtain the fifth security response operation.
[0151] The second security video is either the first security video or a security video acquired after the first security video.
[0152] The fourth step is to execute the fifth security response operation.
[0153] It is understandable that in the above application scenarios, the preset threshold can be dynamically adjusted through the feedback information of the security response operation. In this way, the frequency requirements of different users or the same user at different times for performing security response operations can be met, thereby improving the user's satisfaction with the security effect.
[0154] In some optional implementations of this embodiment, the video recognition result includes a first person in the first security video, and the security preference information indicates a preset relationship with the first person.
[0155] The preset relationship can be a kinship relationship, such as father and son, mother and son, father and daughter, mother and daughter, etc. The first person can be a person in the first security video.
[0156] Based on this, the security response operation matching the first security video can be determined using the following method: based on the video recognition result and the security preference information.
[0157] The first step is to identify a second person who has the preset relationship with the first person in the first security video.
[0158] The second person can be someone who has the preset relationship with the first person in the first security video.
[0159] The second step is to determine the security response action that matches the first security video: send a security alert message to the second person's terminal.
[0160] Here, the second person's terminal can be a terminal logged into with the second person's account, or a terminal bound to the second person's personnel information.
[0161] Security alert messages can be used to provide security notifications. Their content can be set by users or other entities, or generated using preset strategies.
[0162] The first security response operation can be defined as sending a security alert message to the second person's terminal. Furthermore, the first security response operation can be a security response operation determined based on the video recognition results and the security preference information, and matched with the first security video.
[0163] It is understandable that, among the above-mentioned optional implementation methods, when a specific person is identified in the security video, a security alert message can be sent in a timely manner to the terminal of a person with a preset relationship to that person, so that the corresponding person can be informed of the specific person's movements in a timely manner.
[0164] In some application scenarios of the above-mentioned optional implementation methods, the second person who has the preset relationship with the first person in the first security video can be determined in the following way:
[0165] The first step is to determine whether the pre-constructed knowledge graph includes a first node representing the first person.
[0166] In this knowledge graph, nodes represent people, and edges represent relationships between people.
[0167] The first node represents the node that represents the first person.
[0168] Knowledge graphs can include nodes and edges. Edges can represent the relationships between the nodes they connect.
[0169] Here, the nodes in the knowledge graph representing individuals can be pre-entered. For example, personnel information can be collected and entered into the knowledge graph. The relationships represented by the edges in the knowledge graph can be determined through objects such as users.
[0170] For example, see Figure 4A , Figure 4A This is a schematic diagram of the knowledge graph in a dynamic response method for security videos provided in an embodiment of this application.
[0171] The second step, if the knowledge graph includes the first node, is to determine the edge representing the preset relationship from the edges in the knowledge graph that are connected to the first node.
[0172] The third step is to determine the second node that represents the edge connection of the preset relationship.
[0173] The second node is the other node besides the first node among the two nodes connected by the edge representing the preset relationship.
[0174] The fourth step is to identify the person represented by the second node as a second person who has the preset relationship with the first person in the first security video.
[0175] It is understandable that, in the above application scenarios, a knowledge graph can be used to more accurately identify a second person who has the preset relationship with the first person in the first security video.
[0176] To address the technical problem of inaccurate behavior recognition, misjudgment, or inability to determine behavior in home security, which affects user experience, this application can also determine whether the preset target behavior type is met by recognizing the moving distance and the endpoint of the target object in the image. This determines whether the behavior of the target object belongs to the target behavior type. It does not rely on external devices, nor does it require obtaining the object's movement trajectory outside the preset area. Since the preset area is the area traversed by the behavior corresponding to the target behavior type, it prevents the situation where the object changes its movement trajectory and the target object's behavior cannot be determined. This achieves fast and accurate identification of whether the object's behavior belongs to the target behavior type.
[0177] See Figure 5A This is a schematic diagram illustrating an application scenario of an object behavior recognition method provided in an embodiment of this application. Figure 5A As shown, the application scenario 10 may include: object 11, preset area 12, and camera module 13.
[0178] The aforementioned object 11 refers to the object that enters the preset area 12 and is to be identified. It can be a pre-set target object. For example, when the application scenario 10 is a family scenario, the target object can be a family member. Or, when the application scenario 10 is a work scenario, the target object can be a company employee. This application embodiment does not limit this.
[0179] The aforementioned preset area 12 refers to the preset area that the object passes through when its behavior belongs to the target behavior type. The target behavior type can be the object's behavior of going home, leaving home, going to work, or delivering a package. Accordingly, the aforementioned preset area can be an area that the object passes through when entering its home, an area that the object passes through when entering its company, an area that a courier passes through when delivering a package, etc. This application embodiment does not limit this.
[0180] Taking the behavior of an object returning home as an example, if an object wants to return home, it can pass through an area in front of the door, or if the user's home has a courtyard, the above-mentioned preset area can be the courtyard area. When the user returns home, it can go from the entrance of the courtyard to the entrance of the house, thus returning home.
[0181] The aforementioned camera module 13 can be a camera or other device with shooting function, and this application embodiment does not limit it in this regard. The aforementioned camera module 13 can be used to shoot a preset area, and it can be installed within the preset area or outside the preset area, and this application embodiment does not limit it in this regard.
[0182] The camera module 13 can capture images of the entire preset area 12 or only a portion of the preset area 12. In other words, the entire or the main area of the preset area 12 is within the field of view of the camera module 13. The main area can be a portion of the preset area 12 that the target object will pass through when it performs the target behavior.
[0183] In one embodiment, the execution subject of this application embodiment may be the camera module 13 or the base station corresponding to the camera module 13, and this application embodiment does not limit it in this way.
[0184] In this embodiment of the application, the executing entity can acquire images captured by the camera module 13 in a preset area and identify the images to determine whether the movement behavior of the target object in the preset area belongs to the target behavior type using the object behavior recognition method provided in this application.
[0185] In some optional implementations of this embodiment, the first security video includes an image of the target object within a preset area.
[0186] Based on this, the first security video can be identified in the following way to obtain the video recognition result of the first security video:
[0187] The first step is to determine the moving distance and the endpoint of the target object within the preset area based on the image.
[0188] The aforementioned target objects refer to pre-defined objects whose behavior needs to be identified, such as pre-defined family members.
[0189] The aforementioned preset area refers to the area that the target object will pass through when the behavior it exhibits is a target behavior type. For example, when the target object is going home, the entrance area or the courtyard area of the family area it will pass through; or when the target object is delivering a package, the area in front of the delivery locker it will pass through.
[0190] The aforementioned movement distance refers to the straight-line distance between the starting point and the ending point of the target object's movement within the preset area.
[0191] The aforementioned starting point position is the position where the target object begins to move when it enters the preset area, and the aforementioned ending point position is the position where the target object ends its movement or disappears within the preset area.
[0192] When the image includes a region object (such as a house door or a parcel locker), the aforementioned endpoint position is the position where the target object ends its movement within the preset region, or the position where the target object reaches the specified location. However, when the image does not include a region object within the preset region, for example, when the camera module is located on a house door, it cannot capture the house door, and therefore the camera module cannot capture the target object ending its movement. Thus, the aforementioned endpoint position is the position where the target object disappears within the preset region.
[0193] In one embodiment, there is at least one shooting module (e.g. Figure 5A The shooting module 13 shown is used to capture the aforementioned preset area. Based on this, the execution subject of this application embodiment can use the shooting module to detect the preset area in real time, and identify the object within the preset area if an object is detected. Optionally, if the object within the preset area is identified as a target object, an image of the target object within the preset area is acquired.
[0194] As an optional implementation, the execution entity of this application embodiment can detect the presence of a target object in a preset area in the following way: First, when an object is detected in the image captured by the camera module, the object features of the object in the image can be extracted. The aforementioned object features refer to features that can be used to identify the target object, which may include, but are not limited to, physiological features, appearance features, and body posture features. The aforementioned physiological features may be features such as the object's face or iris; the aforementioned appearance features may be features such as corresponding clothing; and the aforementioned body posture features may include the object's gait or posture. For example, when the object is a courier, the execution entity of this application embodiment can further determine that the object is the target object, i.e., the courier, by identifying the object's clothing or courier number.
[0195] Then, based on the aforementioned object characteristics, it can be determined whether a target object matching those characteristics is stored in the preset database. Optionally, if a target object is matched, it can be determined that the target object exists in the aforementioned preset area.
[0196] In one embodiment, when a target object is identified within a preset area, an image of the target object within the preset area can be acquired, and the moving distance and the endpoint of the moving object within the preset area can be determined based on the image.
[0197] The aforementioned images may include multiple frames of images involved in the process of the target object moving from the start to the end of its movement (or disappearing from the shooting frame of the preset area) within the preset area. Based on this, the execution subject of this application embodiment can determine the moving distance and the position of the moving end point of the target object within the preset area from the aforementioned multiple frames of images.
[0198] As for how the execution subject of this application embodiment specifically determines the moving distance and the moving endpoint of the target object within the preset area, please refer to the following explanation, which will not be detailed here.
[0199] The second step is to determine whether the moving distance is greater than or equal to a preset first distance threshold, and whether the moving endpoint is located within a preset endpoint area.
[0200] Third, if it is determined that the moving distance is greater than or equal to the first distance threshold and the moving endpoint is located within the endpoint area, the video recognition result of the first security video indicates that the current behavior of the target object belongs to the target behavior type.
[0201] The following provides a unified explanation of steps two and three:
[0202] The aforementioned first distance threshold refers to the minimum distance that the target behavior type needs to move within the preset area.
[0203] The aforementioned endpoint area refers to the area where the target object corresponding to the target behavior type is located in the preset area when it stops moving or disappears from the shooting frame of the preset area.
[0204] The aforementioned target behavior types refer to the types corresponding to the preset behaviors, which may include, but are not limited to: returning home behavior, leaving home behavior, going to work behavior, and express delivery behavior.
[0205] In this embodiment of the application, in order to more accurately determine whether the current behavior of the target object belongs to the target behavior type, the executing entity of this embodiment can identify the behavior of the target object in the preset area from two aspects: First, it can determine whether the moving distance of the target object in the preset area is greater than or equal to a preset first distance threshold; second, it can determine whether the endpoint position of the target object in the preset area is located within a preset endpoint area. It is understood that this embodiment of the application does not restrict the timing order of the above-mentioned determination of whether the moving distance is greater than or equal to the first distance threshold and whether the endpoint position is located within the endpoint area.
[0206] Regarding the movement distance, as described above, the movement distance is the straight-line distance between the starting point and the ending point of the target object within the preset area. Therefore, when the movement distance is greater than or equal to the first distance threshold, it can exclude the situation where the target object returns to the target area (e.g., home) due to temporary activities, wandering, lingering, or playing.
[0207] Furthermore, the aforementioned first distance threshold can be determined by: determining the minimum distance between a preset starting point position and a preset ending point area, and using the minimum distance as the aforementioned first distance threshold. The preset starting point position can be determined by responding to the starting point position of the target behavior type, that is, the starting point position corresponding to the target behavior type can be preset by the user.
[0208] Regarding the destination location, since the target object moves in different directions, its final destination location will also be different. Therefore, the user's behavior type can be determined based on the destination location of the target object. For example, when a user goes home, its destination location will be near the door of the house, while when a user leaves home, its destination location will be in an area far away from the door of the house.
[0209] Based on this, the endpoint area corresponding to the target behavior type can be preset. When the target object's movement endpoint is located in the preset endpoint area, it can be determined that the target object's current behavior belongs to the target behavior type.
[0210] As for how the target object's movement endpoint is determined to be within the preset endpoint area, please refer to the explanation below, which will not be detailed here.
[0211] In one embodiment, the aforementioned target behavior type may include the behavior of a target entering a target area, where the target area is the area where the target object's behavior is intended (e.g., the target object's house, or the area where a parcel locker is located). Based on this, after determining that the target object's current behavior belongs to the target behavior type, behavior logs can be generated from videos related to the target behavior type within a preset time period.
[0212] Furthermore, as an exemplary implementation, in order to more efficiently and accurately determine whether the current behavior of the target object belongs to the target behavior type, the execution entity of this application embodiment can use the center of the aforementioned endpoint region as the center of a circle and the aforementioned first distance threshold as the radius to determine a first circle. Based on this, when it is determined that the target object moves from a point on or outside the circumference of the first circle within the preset area to the aforementioned endpoint region, it can be determined that the current behavior of the target object belongs to the target behavior type.
[0213] Furthermore, in order to ensure the accuracy of the object identified by the executing entity in this application embodiment as the target object, the executing entity in this application embodiment can re-identify the target object after the target object enters the endpoint area.
[0214] As an exemplary implementation, a target camera module for physiological feature recognition may be present within the aforementioned endpoint area. For example, when the endpoint area includes a door, the target camera module may be a doorbell camera. Based on this, when it is determined that an object has entered the endpoint area, the physiological features of the object (such as facial features, iris features, etc.) can be recognized through the aforementioned target camera module.
[0215] Subsequently, based on the aforementioned physiological characteristics, it can be determined whether a target object matching the aforementioned physiological characteristics is stored in the preset database. Optionally, if a target object is found to be matched, it can be determined that a target object exists within a preset area. Following the above method, since the target camera module can detect the physiological characteristics of the object at close range, the physiological characteristics it acquires are more accurate. Therefore, the target camera module can be used as an auxiliary device for further identification of the target object to ensure the accuracy of identification.
[0216] Furthermore, after determining that the current behavior of the target object belongs to the target behavior type, the executing entity in this application embodiment can send the image and / or prompt information corresponding to the target behavior type of the target object to an external terminal. The prompt information may include at least one of the following: no other preset target objects were detected within a preset time period, or a non-target object entered the preset area. The prompt information may also be a notification of the identified target behavior type, such as a child returning home or a courier from a certain express delivery company delivering a package.
[0217] The aforementioned external terminal can be a terminal capable of communicating with a camera or base station, such as a smartphone. Optionally, the external terminal can have a corresponding application installed or can receive emails to monitor the behavior of the target object in real time.
[0218] For example, when the target is a child at home and the target behavior type corresponds to the behavior of returning home, when the camera or base station detects that the child has returned home, a notification message can be generated and sent to the application installed on the external terminal in the form of a pop-up, or the notification message can be sent to the external terminal via email.
[0219] As an exemplary implementation, the user can pre-set a deadline for each target object to perform the corresponding behavior of the target behavior type each day (e.g., setting the latest time for each target object to return home). Based on this, when the execution subject of this application embodiment determines that the behavior of the target object belongs to the target behavior type, it can record the target object in a preset user table.
[0220] Then, the deadline for each preset target object to enter the preset area can be obtained. When each of the aforementioned deadlines is reached, the user table is searched to see if there is a target object corresponding to that deadline.
[0221] Optionally, if there is no target object corresponding to the deadline, a prompt message can be sent to a preset terminal device. This prompt message is used to indicate that the deadline has arrived and the target user has not yet performed the action corresponding to the target behavior type.
[0222] As an exemplary implementation, if it is determined that the current behavior of the target object belongs to the target behavior type, it can be determined whether the target object is a first preset object. The first preset object is an object whose age is less than a preset age value, such as a child in the family.
[0223] Optionally, if the target object is determined to be a first preset object, it is further determined whether other objects exist in the image of the preset area. If other objects are found to exist, preset alarm information is sent to a preset terminal. These other objects refer to objects not preset and that can be identified as strangers. In this case, it can be identified that a stranger is following a child home. Therefore, to ensure the child's safety, an alarm message can be sent to the terminal corresponding to the target object to alert the child that they are being followed and require special attention, thus ensuring the child's safety.
[0224] It is understood that in the above optional implementation, by acquiring an image of the target object within a preset area, determining the target object's movement distance and endpoint position within the preset area based on the image, determining whether the movement distance is greater than or equal to a preset first distance threshold, and determining whether the endpoint position is located within a preset endpoint area, and if the movement distance is greater than or equal to the first distance threshold and the endpoint position is located within the endpoint area, the current behavior of the target object is determined to belong to the target behavior type. This technical solution, by acquiring a movement image of the target object within a preset area and identifying whether the movement distance and endpoint position of the target object in the image meet the preset conditions of the target behavior type, determines whether the behavior of the target object belongs to the target behavior type. It does not rely on external devices, nor does it require acquiring the object's movement trajectory outside the preset area. Since the preset area is the area traversed by the behavior corresponding to the target behavior type, it prevents the situation where the object changes its movement trajectory, making it impossible to determine the target object's behavior, thus achieving fast and accurate identification of whether the object's behavior belongs to the target behavior type.
[0225] In some application scenarios of the above-mentioned optional implementation methods, the following method can be used to determine the movement distance and movement endpoint position of the target object within the preset area based on the image:
[0226] Step 1: Determine the starting point and ending point of the target object's movement within the preset area.
[0227] Step 2: Determine the first straight-line distance between the starting point and the ending point of the movement, and define the first straight-line distance as the movement distance of the target object within the preset area.
[0228] The following provides a unified explanation of steps one and two:
[0229] The aforementioned starting point position refers to the position where the target object begins to move within the preset area. Correspondingly, the aforementioned ending point position refers to the position where the target object ends to move within the preset area or disappears from the captured image of the preset area.
[0230] The aforementioned first straight-line distance refers to the straight-line distance between the starting point and the ending point of the movement, which is also the shortest distance between the two points.
[0231] In this embodiment of the application, the starting point and ending point of the target object's movement within the preset area can be determined by recognizing the image of the target object within the preset area, and the first straight-line distance between the starting point and ending point is determined as the movement distance of the target object within the preset area.
[0232] As an exemplary implementation, the executing entity of this application embodiment can obtain the shooting depth of the camera at the starting point of the movement and determine the starting shooting distance between the camera and the starting point of the movement based on the shooting depth of the camera. Similarly, the shooting depth of the camera at the ending point of the movement can be obtained and the ending shooting distance between the camera and the ending point of the movement can be determined based on the shooting depth of the camera. Based on this, the first straight-line distance can be determined based on the above-mentioned starting shooting distance and ending shooting distance.
[0233] As another exemplary implementation, the camera described above may be equipped with a distance sensor (including but not limited to: optical distance sensor, infrared distance sensor, ultrasonic distance sensor, etc.). Based on this, the distance sensor can be used to determine a first distance between the camera and the starting point of the movement, and a second distance between the camera and the ending point of the movement. Then, the difference between the first distance and the second distance can be calculated to obtain a first straight-line distance.
[0234] As another exemplary implementation, the executing entity of this application embodiment can obtain the starting coordinates of the above-mentioned starting point position and the ending coordinates of the above-mentioned ending point position, and then directly calculate the starting coordinates and ending coordinates to obtain the above-mentioned first straight-line distance.
[0235] As an optional implementation, the executing entity of this application embodiment can obtain the movement trajectory of the target object within a preset area, and determine the starting point and ending point of the target object's movement on the aforementioned movement trajectory, so as to determine the starting point as the starting point position of the target object's movement within the preset area, and the ending point as the ending point position of the target object's movement within the preset area.
[0236] As an exemplary implementation, an image of the target object within a preset area (which may be a video of the target object moving within the preset area) can be input into a pre-trained motion trajectory extraction model so that the motion trajectory extraction model can extract the motion trajectory of the target object within the target area from the image.
[0237] Furthermore, the aforementioned motion trajectory extraction model may include: N spatiotemporal graph convolutional layers and at least one classifier. The spatiotemporal graph convolutional layers may include graph convolution and spatiotemporal convolution. Based on this, when using the pre-trained motion trajectory extraction model to extract the motion trajectory of a target object in a preset region, the N spatiotemporal graph convolutional layers can be used to extract the spatial features and temporal features of the key points of the target object in the image across different frames. Specifically, the spatial features are extracted through graph convolution, and the temporal features are extracted through spatiotemporal convolution. When extracting spatial features, the graph convolution employs a unified processing method for the key points of the target object's limbs.
[0238] The aforementioned unified processing refers to the removal of edge weights in the model. That is, when determining the object model of the target object, existing technology requires assigning edge weights to the edges of the models corresponding to the limbs of the target object, thus generating an object model that better matches the target object. However, in this application, since the focus of the motion trajectory extraction model is to extract the motion trajectory of the target object, fuzzy processing can be used when generating the object model of the target object. That is, the edge weights of the edges corresponding to the limbs are removed, generating a simpler object model (e.g., a stick figure) corresponding to the target object. This simplifies the model, reduces processing time, and allows for faster determination of the target object's motion trajectory within the preset area.
[0239] Then, at least one classifier can be used to classify the above spatial features and the above temporal features to obtain the movement trajectory of the object model in the target blank image. The above object model is the shape model corresponding to the target object, and the above target blank image is a blank image without background corresponding to the image.
[0240] Furthermore, the aforementioned motion trajectory extraction model can be trained as follows: First, acquire multiple sample images containing the object's movement process, and the standard motion trajectory of the object corresponding to each sample image. Then, input each sample image into the initial motion trajectory extraction model to obtain the predicted motion trajectory of each sample image output by the initial motion trajectory extraction model. Next, for each sample image, determine the corresponding loss value based on the standard motion trajectory and the predicted motion trajectory of that sample image.
[0241] Then, based on the loss value corresponding to each sample image, it can be determined whether the preset convergence condition is met. Optionally, if the convergence condition is met, a pre-trained motion trajectory extraction model is obtained. Optionally, if the convergence condition is not met, the initial motion trajectory extraction model is further trained using sample images.
[0242] The convergence condition can be that the loss value of each sample image is less than a first preset threshold, or that the number of sample images with a loss value less than the first preset threshold is greater than a second preset threshold. This application embodiment does not limit this.
[0243] Furthermore, the aforementioned initial motion trajectory extraction model may include: N initial spatiotemporal graph convolutional layers and at least one initial classifier. The initial spatiotemporal graph convolutional layers may include initial graph convolution and initial spatiotemporal convolution. Based on this, when inputting each sample image into the initial motion trajectory extraction model to obtain the predicted motion trajectory of each sample image output by the initial motion trajectory extraction model, each sample image can be sequentially input into the N initial spatiotemporal graph convolutional layers. This allows the initial graph convolution in each initial spatiotemporal graph convolutional layer to extract the initial spatial features of the key points of the object in the sample image, and the initial spatiotemporal convolution to extract the initial temporal features of the key points of the object in different frames. Specifically, when extracting the initial spatial features of the key points of the object, the initial graph convolution employs uniform processing for the key points of the object's limbs.
[0244] Then, the initial spatial features and initial temporal features can be input into at least one initial classifier to obtain the predicted movement trajectory of the initial object model output by at least one initial classifier in the sample blank image. The initial object model is the shape model corresponding to the object, and the sample blank image is the blank image without background corresponding to the sample image.
[0245] It should be noted that the images and sample images input to the motion trajectory extraction model mentioned above, as well as the output blank images and sample blank images, are all multi-frame images. The movement process of the target object's shape model in the blank image can be clearly observed through the video segment composed of multiple frames.
[0246] It is understandable that in the above application scenario, by determining the starting and ending positions of the target object's movement within a preset area, a first straight-line distance between the starting and ending positions is determined, and this first straight-line distance is defined as the movement distance of the target object within the preset area. This technical solution, by determining the starting and ending positions of the target object's movement within the preset area, thereby determining the movement distance of the target object within the preset area, and since this movement distance is the straight-line distance between the two, it can directly reflect the movement changes of the target object within the preset area. This allows for a more accurate identification of whether the target object's movement behavior within the preset area belongs to the target behavior type, achieving accurate determination of the target object's movement distance within the preset area, and thus more accurately determining whether the target object's current behavior belongs to the target behavior type.
[0247] In some application scenarios of the above-mentioned optional implementation methods, the following method can be used to determine whether the moving endpoint location is located within the preset endpoint area:
[0248] The first step is to determine the distance between the moving endpoint location and the preset endpoint location, wherein the preset endpoint location is determined by responding to the endpoint location of the target behavior type.
[0249] The second step is to determine that the moving endpoint location is located within a preset endpoint area if the distance to the endpoint is less than or equal to the second distance threshold; otherwise, it is determined that the moving endpoint location is located outside the preset endpoint area.
[0250] The following provides a unified explanation of the solutions described in the above application scenarios:
[0251] The aforementioned preset endpoint position refers to the endpoint position that the target object, set by the user for the corresponding behavior type, will traverse. Specifically, when the image captured by the camera module includes an area object (such as a door), the preset endpoint position can be the position where the user ends their movement; when the image captured by the camera module does not include an area object, the preset endpoint position can be the position where the user disappears from the image of the preset area.
[0252] In this embodiment, the user can pre-set an endpoint location corresponding to the target behavior type. Furthermore, the execution entity in this embodiment can determine an endpoint region within a preset area based on the preset endpoint location. Therefore, the execution entity can compare the determined endpoint location of the target object with the preset endpoint location to determine whether the target object has reached the endpoint region of the preset area.
[0253] As an optional implementation, the distance between the aforementioned moving endpoint position and the preset endpoint position can be determined, that is, the straight-line distance between the two.
[0254] Then, it can be determined whether the distance to the endpoint is less than or equal to the second distance threshold. Optionally, if it is determined that the distance to the endpoint is less than or equal to the second distance threshold, it can be determined that the moving endpoint is located within a preset endpoint area. Conversely, if it is determined that the distance to the endpoint is greater than the second distance threshold, it can be determined that the moving endpoint is located outside the preset endpoint area.
[0255] The aforementioned second distance threshold can be determined based on at least one historical movement endpoint location of at least one target object within a historical time period.
[0256] As an exemplary implementation, historical mobile endpoint locations can be obtained, and the historical distance value between the historical mobile endpoint locations and the aforementioned preset endpoint locations can be determined. The preset endpoint locations can be determined by setting the endpoint location through a response to a target behavior type, that is, the preset endpoint locations can be set by the user.
[0257] Then, the second distance threshold can be determined based on the historical distance values mentioned above.
[0258] Optionally, when there is only one historical distance value, the historical distance value can be directly determined as the second distance threshold; when there are multiple historical distance values, the average or maximum value of the multiple historical distance values can be determined as the second distance threshold.
[0259] Based on the above description, it can be deduced that the aforementioned endpoint region can be a circle or semicircle with the preset endpoint position as the center and the second distance threshold as the radius.
[0260] Based on this, when determining whether the endpoint of the target object's movement is located in the endpoint area of the preset area, the execution subject of this application embodiment can first determine a second distance threshold, and then determine a target circle or a target semicircle with the preset endpoint as the center and the second distance threshold as the radius.
[0261] Then, it can be determined whether the above-mentioned moving endpoint position is located within the above-mentioned target circle or target semicircle. If so, it can be determined that the moving endpoint position of the target object is located within the above-mentioned endpoint area; if not, it can be determined that the moving endpoint position of the target object is located outside the above-mentioned endpoint area.
[0262] As another optional implementation, the execution entity of this application embodiment can determine the coordinates of the aforementioned endpoint region based on the coordinates of the preset endpoint position, such as the aforementioned determined target circle or target semicircle.
[0263] Subsequently, when determining whether the destination location is within the preset destination area, the location coordinates of the destination location can be obtained, and it can be determined whether the location coordinates are within the coordinates of the destination area. If so, it can be determined that the destination location is within the preset destination area.
[0264] It is understandable that the technical solution in the above application scenario determines whether the distance between the target object's moving endpoint and the preset endpoint is less than or equal to a second distance threshold. If so, the target object is determined to be within the preset endpoint area; otherwise, it is determined to be outside the preset endpoint area. This technical solution, by comparing the distance between the target object's moving endpoint within the preset area and the preset endpoint with the second distance threshold, can simply and accurately determine whether the target object's moving endpoint is within the preset endpoint area, thereby enabling a simple and accurate determination of whether the target object's current behavior belongs to the target behavior type.
[0265] In some optional implementations of this embodiment, the first security video includes an image of the target object within a preset area.
[0266] Based on this, the first security video can be identified in the following way to obtain the video recognition result of the first security video:
[0267] The first step is to determine the starting point and ending point of the target object's movement within the preset area based on the image.
[0268] The aforementioned target objects refer to pre-defined objects whose behavior needs to be identified, such as pre-defined family members.
[0269] The aforementioned preset area refers to the area that the target object will pass through when the behavior it exhibits is a target behavior type. For example, when the target object is going home, the entrance area or the courtyard area of the family area it will pass through; or when the target object is delivering a package, the area in front of the delivery locker it will pass through.
[0270] The aforementioned starting point position is the position where the target object begins to move when it enters the preset area, and the aforementioned ending point position is the position where the target object ends its movement or disappears within the preset area.
[0271] Specifically, when the image includes a pre-defined area (such as a house door or a parcel locker), the aforementioned endpoint position is the location where the target object ends its movement within the pre-defined area, or it can be the location where the target object reaches a designated position. However, when the image does not include a pre-defined area, such as when the camera module is positioned on a house door and cannot capture the door itself, the camera module cannot capture the target object ending its movement. Therefore, the aforementioned endpoint position is the location where the target object disappears within the pre-defined area. Furthermore, the endpoint position can also be a designated area, such as a pre-defined area defined by the user within the pre-defined area.
[0272] In one embodiment, there is at least one shooting module (e.g. Figure 5A The shooting module 13 shown is used to capture the aforementioned preset area. Based on this, the execution subject of this application embodiment can use the shooting module to detect the preset area in real time, and identify the object within the preset area if an object is detected. Optionally, if the object within the preset area is identified as a target object, an image of the target object within the preset area is acquired.
[0273] As an optional implementation, the execution entity of this application embodiment can detect the presence of a target object in a preset area in the following way: First, when an object is detected in the image captured by the camera module, the object features of the object in the image can be extracted. The aforementioned object features refer to features that can be used to identify the target object, which may include, but are not limited to, physiological features, appearance features, and body posture features. The aforementioned physiological features may be features such as the object's face or iris; the aforementioned appearance features may be features such as corresponding clothing; and the aforementioned body posture features may include the object's gait or posture. For example, when the object is a courier, the execution entity of this application embodiment can further determine that the object is the target object, i.e., the courier, by identifying the object's clothing or courier number.
[0274] Then, based on the aforementioned object characteristics, it can be determined whether a target object matching those characteristics is stored in the preset database. Optionally, if a target object is matched, it can be determined that the target object exists in the aforementioned preset area.
[0275] In one embodiment, when a target object is identified within a preset area, an image of the target object within the preset area can be acquired, and the starting point and ending point of the target object's movement within the preset area can be determined based on the image.
[0276] The aforementioned images may include multiple frames of images involved in the process of the target object moving from the start to the end of its movement (or disappearing from the captured image of the preset area) within the preset area. Based on this, the execution subject of this application embodiment can determine the starting point position and the ending point position of the target object's movement within the preset area from the aforementioned multiple frames of images.
[0277] The second step is to determine that, if the starting point of the movement is located within a preset starting area and the ending point of the movement is located within a preset ending area, the video recognition result of the first security video indicates that the current behavior of the target object belongs to the target behavior type.
[0278] The aforementioned starting area refers to the area to which the starting point of the behavior corresponding to the pre-set target behavior type belongs.
[0279] The aforementioned endpoint area refers to the area where the movement ends or the target object disappears in the pre-defined target behavior type.
[0280] In this embodiment of the application, the executing entity can determine whether the target object has moved from the starting area to the ending area, thereby determining whether the current behavior of the target object belongs to the target behavior type.
[0281] As an optional implementation, it can be determined whether the aforementioned starting point position is located in the starting area, and whether the ending point position is located in the ending area. Optionally, if it is determined that the starting point position is located in the starting area and the ending point position is located in the ending area, it is determined that the current behavior of the target object belongs to the target behavior type.
[0282] As an exemplary implementation, the starting distance between the starting point position and the preset starting point position can be determined, as can the ending distance between the ending point position and the preset starting point position. The preset starting point position is determined by responding to the starting point position of the target behavior type, and the preset ending point position is determined by responding to the ending point position of the target behavior type. The specific method for determining the preset starting point position and the preset ending point position will be explained below and will not be detailed here.
[0283] Subsequently, if it is determined that the starting point distance is less than the third distance threshold and the ending point distance is less than the fourth distance threshold, it can be determined that the starting point position is located within the preset starting point area and the ending point position is located within the preset ending point area.
[0284] The third and fourth distance thresholds can be preset distance thresholds, or they can be determined by the executing entity in this application embodiment based on multiple historical starting points and multiple historical ending points generated when the object performs a behavior corresponding to the target behavior type in a historical time period. For example, the maximum distance value between multiple historical starting points can be determined as the third distance threshold, and the maximum distance value between multiple historical ending points can be determined as the fourth distance threshold.
[0285] Furthermore, the executing entity in this application embodiment can determine the movement distance of the target object within a preset area based on the starting point and ending point of the movement. Then, it can determine whether the movement distance is greater than or equal to a preset first distance threshold, and whether the ending point is located within a preset ending area. Optionally, if it is determined that the movement distance is greater than or equal to the first distance threshold and the ending point is located within the ending area, it is determined that the target object's current behavior belongs to a target behavior type.
[0286] As for how the movement distance is determined and how the destination location is determined to be within the destination area, please refer to the description above, which will not be elaborated here.
[0287] It is understandable that in the above optional implementation methods, by acquiring an image of the target object within a preset area, and determining the starting and ending positions of the target object's movement within the preset area based on the image, and if the starting position is located in the starting area and the ending position is located in the ending area, the current behavior of the target object is determined to belong to the target behavior type. This technical solution, by acquiring an image of the target object's movement within a preset area and identifying whether the starting and ending positions of the target object's movement in the image are located in the starting and ending areas respectively, determines whether the behavior of the target object belongs to the target behavior type. It does not rely on external devices, nor does it require acquiring the object's movement trajectory outside the preset area. Since the preset area is the area traversed by the behavior corresponding to the target behavior type, it prevents the situation where the object changes its movement trajectory, making it impossible to determine the target object's behavior, thus achieving fast and accurate identification of whether the object's behavior belongs to the target behavior type.
[0288] Alternatively, object behavior recognition can also be performed using the following methods:
[0289] The first step is to acquire an image of the target object within a preset area, and then determine the movement information of the target object within the preset area based on the image.
[0290] The aforementioned target objects refer to pre-defined objects whose behavior needs to be identified, such as pre-defined family members.
[0291] The aforementioned preset area refers to the area that the target object will pass through when the behavior it exhibits is a target behavior type. For example, when the target object is going home, the entrance area or the courtyard area of the family area it will pass through; or when the target object is delivering a package, the area in front of the delivery locker it will pass through.
[0292] The aforementioned movement information refers to the movement information of the target object within the preset area, which may include, but is not limited to, movement distance, movement endpoint position, and movement destination position. The aforementioned movement distance refers to the straight-line distance between the starting position and the ending position of the target object's movement within the preset area. The aforementioned starting position is the position where the target object begins its movement upon entering the preset area, and the aforementioned ending position is the position where the target object ends its movement or disappears within the preset area.
[0293] When the image includes a region object (such as a house door or a parcel locker), the aforementioned endpoint position is the position where the target object ends its movement within the preset region, or the position where the target object reaches the specified location. However, when the image does not include a region object within the preset region, for example, when the camera module is located on a house door, it cannot capture the house door, and therefore the camera module cannot capture the target object ending its movement. Thus, the aforementioned endpoint position is the position where the target object disappears within the preset region.
[0294] In one embodiment, there is at least one shooting module (e.g. Figure 5A The shooting module 13 shown is used to capture the aforementioned preset area. Based on this, the execution subject of this application embodiment can use the shooting module to detect the preset area in real time, and identify the object within the preset area if an object is detected. Optionally, if the object within the preset area is identified as a target object, an image of the target object within the preset area is acquired.
[0295] As an optional implementation, the execution entity of this application embodiment can detect the presence of a target object in a preset area in the following way: First, when an object is detected in the image captured by the camera module, the object features of the object in the image can be extracted. The aforementioned object features refer to features that can be used to identify the target object, which may include, but are not limited to, physiological features, appearance features, and body posture features. The aforementioned physiological features may be features such as the object's face or iris; the aforementioned appearance features may be features such as corresponding clothing; and the aforementioned body posture features may include the object's gait or posture. For example, when the object is a courier, the execution entity of this application embodiment can further determine that the object is the target object, i.e., the courier, by identifying the object's clothing or courier number.
[0296] Then, based on the aforementioned object characteristics, it can be determined whether a target object matching those characteristics is stored in the preset database. Optionally, if a target object is matched, it can be determined that the target object exists in the aforementioned preset area.
[0297] In one embodiment, when a target object is identified within a preset area, an image of the target object within the preset area can be acquired, and the movement information of the target object within the preset area can be determined based on the image.
[0298] The aforementioned images may include multiple frames of images involved in the process of the target object moving from the start to the end of its movement (or disappearing from the captured image of the preset area) within the preset area. Based on this, the execution subject of this application embodiment can determine the movement information of the target object within the preset area from the aforementioned multiple frames of images.
[0299] The second step is to determine, based on the aforementioned movement information, whether the target object's current behavior belongs to the target behavior type.
[0300] In one embodiment, the executing entity of this application embodiment can determine whether the current behavior of the target object belongs to the target behavior type based on the movement information.
[0301] As an optional implementation, the aforementioned movement information may include the movement distance and destination location of the target object within a preset area. Based on this, the execution entity of this application embodiment can determine whether the current behavior of the target object belongs to the target behavior type based on the aforementioned movement distance and destination location.
[0302] As for how the current behavior of the target object is determined to belong to the target behavior type based on the movement distance and the location of the movement endpoint, please refer to the description above, which will not be repeated here.
[0303] As another optional implementation, the aforementioned movement information may include the starting point position and ending point position of the target object within a preset area. Based on this, the execution entity in this application embodiment can determine whether the current behavior of the target object belongs to the target behavior type based on the starting point position and the ending point position.
[0304] As for how the current behavior of the target object is determined to belong to the target behavior type based on the starting position and ending position of the movement, please refer to the description above, which will not be repeated here.
[0305] As can be understood, in the above solution, an image of the target object within a preset area is acquired, and the movement information of the target object within the preset area is determined based on this image. Based on this movement information, it is then determined whether the target object's current behavior belongs to the target behavior type. This technical solution, by acquiring an image of the target object's movement within a preset area and identifying whether the movement information of the target object in the image meets the conditions of a preset target behavior type, determines whether the behavior of the target object belongs to the target behavior type. It does not rely on external devices, nor does it require acquiring the object's movement trajectory outside the preset area. Since the preset area is the area traversed by the behavior corresponding to the target behavior type, it prevents situations where the object changes its movement trajectory, making it impossible to determine the target object's behavior. This achieves rapid and accurate identification of whether an object's behavior belongs to the target behavior type.
[0306] In some optional implementations of this embodiment, the first security video includes footage captured of a preset area.
[0307] Based on this, the first security video can be obtained in the following way: acquire the captured image of the preset area and display the captured image.
[0308] The aforementioned preset area refers to the preset area that an object passes through when it performs the behavior corresponding to the target behavior type.
[0309] The aforementioned captured footage refers to the footage captured by the camera module in the preset area.
[0310] In one embodiment, the execution subject of this application embodiment may be a preset terminal (which may be a terminal device or an application installed on the terminal device; this application embodiment does not limit this). The preset terminal may be a terminal connected to a camera module or a base station. The terminal can acquire the shooting image of the preset area by the camera module in real time and display the shooting image through a visual interface so that the user can make settings based on the shooting image.
[0311] In one embodiment, since there may be multiple shooting modules used to shoot the preset area in actual applications, the executing entity of this application embodiment can obtain a list of shooting modules used to shoot the preset area when acquiring the shooting image of the preset area, and display the list of shooting modules.
[0312] Users can select a target camera module from the camera module list. Based on this, the execution entity of this application embodiment can respond to the camera module selection operation of the above-mentioned camera module list, determine the target camera module from the above-mentioned camera module list, and obtain the shooting image of the preset area through the target camera module.
[0313] Based on this, the video recognition result of the first security video can be obtained in the following way:
[0314] The first step is to identify the captured image to determine the preset start point and preset end point positions of the target behavior type set for the captured image.
[0315] The aforementioned preset starting point position refers to the starting point when the user-defined target behavior type begins to move within the preset area.
[0316] The aforementioned preset endpoint position refers to the endpoint when the user-defined target behavior type ends its movement within the preset area or disappears from the image within the preset area.
[0317] Specifically, when the image captured by the camera module includes an area object (such as a door), the preset endpoint position can be the position where the user ends their movement; when the image captured by the camera module does not include an area object, the preset endpoint position can be the position where the user disappears from the image of the preset area.
[0318] The aforementioned target behavior type refers to the type corresponding to the preset behavior, which may include behaviors such as going home, going to work, and delivering packages.
[0319] In this embodiment of the application, the executing entity can acquire and output the captured image of a preset area. Based on this, the user can set a preset start position and a preset end position corresponding to the target behavior type for the captured image.
[0320] Based on this, the executing entity of this application embodiment can determine the preset starting point position and preset ending point position of the target behavior type set for the above-mentioned captured image.
[0321] As an optional implementation, the user can trigger (double-click, single-click, or long-press, etc.) the display screen of the execution subject in this application embodiment. The execution subject in this application embodiment can respond to the user's triggering operation by outputting a preset icon in the captured image. The preset icon can be the user image of the user who logged into the execution subject in this application embodiment, or it can be a default icon, such as an arrow, a dot, a circle, etc. This application embodiment does not impose any restrictions on this.
[0322] Based on this, the execution subject of this application embodiment can identify the preset start position and preset end position in response to the setting of preset icons in the captured image.
[0323] As an exemplary implementation, the setting of the preset icon may include a click operation. Furthermore, each time a user clicks the shooting screen, the execution entity of this application embodiment may, in response to the user's click operation, set a preset icon at the position corresponding to the click operation. Based on this, a user may perform at least two click operations on the shooting screen, and the execution entity of this application embodiment may, in response to at least two click operations on the shooting screen, set a preset icon at the position corresponding to each click operation.
[0324] Subsequently, the executing entity of this application embodiment can obtain the positions of at least two preset icons clicked by the user, and determine the preset start position and preset end position based on the positions of the preset icons.
[0325] Furthermore, optionally, the executing entity in this application embodiment can determine the preset starting position and the preset ending position according to the order in which the user clicks, for example, the position clicked first is determined as the preset starting position, and the position clicked later is determined as the preset ending position.
[0326] Optionally, the execution entity in this embodiment may output a selection box at each location, indicating whether it should be a preset start point or a preset end point. The user can use this selection box to determine whether to use that location as the preset start point or the preset end point. Based on this, the execution entity in this embodiment can determine the final preset start point and preset end point according to the user's selection.
[0327] As another exemplary implementation, setting a preset icon may include a drag operation. Further, after the user triggers the shooting screen and the preset icon is displayed on the shooting screen, the user can drag the icon to draw a movement trajectory corresponding to the target behavior type on the shooting screen. The execution entity of this application embodiment can, in response to the drag operation of the preset icon in the shooting screen, determine the drag trajectory corresponding to the drag operation and identify the starting and ending points of the drag trajectory.
[0328] Then, the above-mentioned drag starting point can be determined as the preset starting point position of the target behavior type set for the shooting screen, and the drag ending point can be determined as the preset ending point position of the target behavior type set for the shooting screen.
[0329] For example, assuming the target behavior type corresponds to the target object's behavior of going home, and the camera module connected to the execution subject in this application embodiment is a camera mounted high by the user, then the captured image obtained by the execution subject in this application embodiment can be as follows: Figure 5B As shown, see Figure 5BThis is a schematic diagram of a captured image provided in an embodiment of this application. Assuming the camera module connected to the execution subject of this application is a camera installed on a door, the captured image obtained by the execution subject of this application can be as follows: Figure 5C As shown, see Figure 5C This is a schematic diagram of another captured image provided in an embodiment of this application. Figure 5B and Figure 5C As shown, Figure 5B and Figure 5C The difference lies in whether the camera module can capture the location of the door.
[0330] As mentioned above Figure 5B or Figure 5C As shown, when the executing entity of this application embodiment detects the user's trigger operation, it can output a preset icon in the shooting screen. The user can drag the preset icon to draw the movement trajectory corresponding to the target behavior type in the shooting screen, that is... Figure 5C The image shows the movement trajectory from point A to point B. Based on this, the starting point A of the identified movement trajectory can be determined as the preset starting position, and the ending point B can be determined as the preset ending position.
[0331] As another optional implementation, the user can record a video, which may include the movement process of any target object within a preset area, corresponding to a target behavior type, such as the behavior of going home. The user can then send the recorded video to the execution entity of this embodiment.
[0332] Based on this, the executing entity of this application embodiment can acquire the input captured video, which may include a preset object performing a target behavior type within a preset area. Then, the captured video can be analyzed to determine the appearance and disappearance points of the preset object in the video, and the appearance point is determined as a preset starting point position for the target behavior type set for the captured image, and the disappearance point is determined as a preset ending point position for the target behavior type set for the captured image.
[0333] Furthermore, to improve the accuracy of the aforementioned preset start and end positions, the executing entity in this embodiment can mark the appearance and disappearance points in the captured video on the captured screen when identifying them, and output a prompt indicating whether to use the appearance point as the preset start position and the disappearance point as the preset end position. When the user clicks the confirmation button in the prompt, the executing entity in this embodiment can determine the appearance point as the preset start position and the disappearance point as the preset end position.
[0334] The second step involves determining a distance threshold corresponding to the target behavior type based on the preset starting point position and the preset ending point position. This threshold is used to identify whether the object behavior belongs to the target behavior type based on the distance threshold and the preset ending point position, so as to obtain the video recognition result of the first security video.
[0335] The aforementioned distance threshold refers to the shortest movement distance of the target behavior type within the preset area. In other words, if the object's current behavior belongs to the target behavior type, then its movement distance within the preset area can be greater than or equal to this distance threshold.
[0336] In one embodiment, the executing entity of this application embodiment can determine the distance threshold based on the determined preset starting position and the preset ending position, and determine whether the behavior of the target object belongs to the target behavior type based on the distance threshold and the preset ending position as described above.
[0337] As an optional implementation, the straight-line distance between the preset starting position and the preset ending position can be determined as the distance threshold.
[0338] Furthermore, after determining the distance threshold, the distance threshold and the preset endpoint location can be sent to the camera module or the base station of the camera module, so that the camera module or the base station can determine whether the behavior of the target object belongs to the target behavior type based on the distance threshold and the preset endpoint location.
[0339] Furthermore, the endpoint area corresponding to the target behavior type can be determined based on the preset endpoint location, and the distance threshold and endpoint area can be sent to the camera module or the base station of the camera module, so that the camera module or the base station can determine whether the behavior of the target object belongs to the target behavior type based on the distance threshold and endpoint area.
[0340] As an exemplary implementation, when the shooting module is located outside the area object, the area object may exist in the shooting screen (for example, when the target behavior type corresponds to the behavior of going home, the area object refers to the door used to indicate that the user is going home; when the target behavior type refers to the behavior of express delivery, the area object refers to the delivery locker). In this case, the preset endpoint position can be located on the area object. When the shooting module is located on the area object, there is no area object in the shooting screen, so the preset endpoint position is not located on the area object.
[0341] Based on this, when determining the endpoint region of the target behavior type according to the preset endpoint location, it can be determined first whether the preset endpoint location is located on the preset region object.
[0342] Optionally, if the preset starting point is located within a region object, the endpoint region corresponding to the target behavior type can be determined with that region object as the center. For example, a circle with the region object as the center and a first preset distance threshold as the radius.
[0343] Conversely, if the preset starting point is not located within a region object, the endpoint region corresponding to the target behavior type can be determined using the preset ending point as the center. For example, a circle with the preset ending point as the center and a second preset distance threshold as the radius. The first and second preset distance thresholds can be the same or different distance thresholds.
[0344] Furthermore, in practical applications, there may be multiple target objects. Therefore, the execution entity in this embodiment can output a list of objects, and the user can select the target object for which behavior recognition is required. Then, the user-selected target object, along with its object characteristics, can be sent to the camera module or base station, so that the camera module or base station can identify whether the behavior of the user-selected target object belongs to the target behavior type when performing object behavior recognition.
[0345] As an example implementation, a list of objects can be obtained and displayed. The user can then select the target object from this list for behavior recognition.
[0346] Based on this, the execution entity of this application embodiment can respond to an object selection operation on the object list, determine the target object from the object list, and identify whether the behavior of the target object belongs to the target behavior type.
[0347] It is understandable that in the above optional implementation, by acquiring and displaying a captured image of a preset area, a preset start position and a preset end position for the target behavior type set for the captured image are determined. Based on the preset start position and the preset end position, a distance threshold corresponding to the target behavior type is determined, which is used to identify whether the object behavior belongs to the target behavior type based on the distance threshold and the preset end position. This technical solution, by displaying a captured image of a preset area on a preset terminal, allows the user to set a preset start position and a preset end position for the target object to perform the behavior corresponding to the target behavior type on the captured image. The distance threshold corresponding to the target behavior type is determined based on the preset start position and the preset end position, enabling the shooting module or base station to identify the target object's behavior based on the distance threshold and the preset end position. This provides an "immersive" experience and a sense of animation, bringing the user into a familiar homecoming scenario, achieving fast and accurate homecoming trajectory drawing, thereby improving the efficiency and accuracy of object behavior identification.
[0348] In some optional implementations of this embodiment, the first security video includes footage captured of a preset area.
[0349] Based on this, the first security video can be obtained in the following way: acquire the captured image of the preset area and display the captured image.
[0350] Furthermore, the first security video can be identified in the following way to obtain the video recognition result of the first security video:
[0351] The first step is to identify the captured footage to determine the movement trajectory of the target behavior type set for the captured footage.
[0352] The second step is to determine the distance threshold and endpoint region corresponding to the target behavior type based on the movement trajectory, and to identify whether the object behavior belongs to the target behavior type based on the distance threshold and the endpoint region, so as to obtain the video recognition result of the first security video.
[0353] The following provides a unified explanation of steps one and two:
[0354] The aforementioned movement trajectory refers to the trajectory of a user moving within a preset area when the user performs a behavior corresponding to the target behavior type.
[0355] The aforementioned target behavior type refers to the type corresponding to the preset behavior, which may include behaviors such as going home, going to work, and delivering packages.
[0356] The aforementioned distance threshold refers to the shortest movement distance of the target behavior type within the preset area. In other words, if the object's current behavior belongs to the target behavior type, then its movement distance within the preset area can be greater than or equal to this distance threshold.
[0357] The aforementioned endpoint area refers to the area where the user's movement stops or disappears within a preset area when the user performs the action corresponding to the target behavior type.
[0358] In this embodiment of the application, the executing entity can acquire and output the captured image of a preset area. Based on this, the user can set the movement trajectory corresponding to the target behavior type for the captured image.
[0359] Subsequently, the distance threshold and endpoint area corresponding to the target behavior type can be determined based on the movement trajectory, and the distance threshold and endpoint area can be sent to the shooting module or base station so that the shooting module or base station can determine whether the object behavior belongs to the target behavior type based on the distance threshold or endpoint area.
[0360] As an optional implementation, the user can trigger (double-click, single-click, or long-press, etc.) the display screen of the execution subject in this application embodiment. The execution subject in this application embodiment can respond to the user's triggering operation by outputting a preset icon in the captured image. The preset icon can be the user image of the user who logged into the execution subject in this application embodiment, or it can be a default icon, such as an arrow, a dot, a circle, etc. This application embodiment does not impose any restrictions on this.
[0361] Based on this, the execution subject of this application embodiment can identify the movement trajectory of the target behavior type in response to the setting of preset icons in the captured image.
[0362] As an exemplary implementation, setting a preset icon may include a drag operation. Further, after the user triggers the shooting screen and the preset icon is displayed, the user can drag the icon to draw a movement trajectory corresponding to the target behavior type on the shooting screen. The execution entity of this application embodiment can determine the drag trajectory corresponding to the drag operation in response to the aforementioned drag operation of the preset icon in the shooting screen.
[0363] Subsequently, the executing entity of this application embodiment can determine the preset starting position and the preset ending position based on the movement trajectory, and determine the distance threshold and the ending area based on the preset starting position and the preset ending position.
[0364] As an exemplary implementation, the starting point and ending point of the dragging trajectory can be identified, and the starting point is determined as a preset starting point position for the target behavior type set for the shooting screen, and the ending point is determined as a preset ending point position for the target behavior type set for the shooting screen.
[0365] For example, assuming the target behavior type corresponds to the target object's behavior of going home, and the camera module connected to the execution subject in this application embodiment is a camera mounted high by the user, then the captured image obtained by the execution subject in this application embodiment can be as follows: Figure 5B As shown, see Figure 5B This is a schematic diagram of a captured image provided in an embodiment of this application. Assuming the camera module connected to the execution subject of this application is a camera installed on a door, the captured image obtained by the execution subject of this application can be as follows: Figure 5C As shown, see Figure 5C This is a schematic diagram of another captured image provided in an embodiment of this application. Figure 5B and Figure 5C The difference lies in whether the camera module can capture the location of the door.
[0366] As mentioned above Figure 5B or Figure 5CAs shown, when the executing entity of this application embodiment detects the user's trigger operation, it can output a preset icon in the shooting screen. The user can drag the preset icon to draw the movement trajectory corresponding to the target behavior type in the shooting screen, that is... Figure 5C The image shows the movement trajectory from point A to point B. Based on this, the starting point A of the identified movement trajectory can be determined as the preset starting position, and the ending point B can be determined as the preset ending position.
[0367] In one embodiment, the executing entity of this application embodiment can determine the distance threshold and the endpoint region based on the determined preset starting point position and the preset ending point position, and determine whether the behavior of the target object belongs to the target behavior type in the context based on the distance threshold and the endpoint region.
[0368] As an optional implementation, the straight-line distance between the preset starting position and the preset ending position can be determined as the distance threshold.
[0369] Furthermore, when the shooting module is located outside the area object, the area object may exist in the shooting screen (for example, when the target behavior type corresponds to the behavior of going home, the area object refers to the door used to indicate that the user is going home; when the target behavior type corresponds to the behavior of express delivery, the area object refers to the delivery locker). In this case, the preset endpoint position can be located on the area object. When the shooting module is located on the area object, there is no area object in the shooting screen, so the preset endpoint position is not located on the area object.
[0370] Based on this, when determining the endpoint region of the target behavior type according to the preset endpoint location, it can be determined first whether the preset endpoint location is located on the preset region object.
[0371] Optionally, if the preset endpoint location is determined to be within a region object, the endpoint region corresponding to the target behavior type can be determined with that region object as the center. For example, a circle with the region object as the center and a first preset distance threshold as the radius.
[0372] Conversely, if the preset endpoint location is determined not to be located within a region object, the endpoint region corresponding to the target behavior type can be determined using the preset endpoint location as the center. For example, a circle with the preset endpoint location as the center and a second preset distance threshold as the radius. The first and second preset distance thresholds can be the same or different distance thresholds.
[0373] As another optional implementation, a starting point can be preset as the center, and an area with a first distance threshold as the radius can be defined as the ending area. The distance value from the preset starting point to any edge position of the ending area can be determined as the distance threshold. The aforementioned first distance threshold can be a distance value preset by the user.
[0374] It is understood that in the above optional implementation, by acquiring and displaying a captured image of a preset area, the movement trajectory of the target behavior type set for the captured image is determined. Based on the movement trajectory, a distance threshold and endpoint area corresponding to the target behavior type are determined, which are used to identify whether the object's behavior belongs to the target behavior type. This technical solution, by displaying a captured image of a preset area on a preset terminal, allows the user to set a movement trajectory for the target object when it performs a behavior corresponding to the target behavior type. Based on the movement trajectory, the distance threshold and endpoint area corresponding to the target behavior type are determined. This enables the shooting module or base station to identify the target object's behavior based on the distance threshold and endpoint area, achieving fast and accurate homecoming trajectory drawing, thereby improving the efficiency and accuracy of object behavior identification.
[0375] In some optional implementations of this embodiment, the first security video includes footage captured of a preset area.
[0376] Based on this, the first security video can be obtained in the following way: acquire the captured image of the preset area and display the captured image.
[0377] Furthermore, the first security video can be identified in the following way to obtain the video recognition result of the first security video:
[0378] The first step is to identify the captured image to determine the preset start point and preset end point positions of the target behavior type set for the captured image.
[0379] The second step involves determining a starting area and an ending area based on the preset starting point and the preset ending point, respectively. These areas are used to identify whether the object's behavior belongs to the target behavior type, thereby obtaining the video recognition result of the first security video.
[0380] The following provides a unified explanation of steps one and two:
[0381] The aforementioned target behavior type refers to the type corresponding to the preset behavior. The preset behavior can be home behavior, leaving home behavior, going to work behavior, delivering express delivery behavior, etc.
[0382] The aforementioned preset starting point position refers to the starting point when the user-defined target behavior type begins to move within the preset area.
[0383] The aforementioned preset endpoint position refers to the endpoint when the user-defined target behavior type ends its movement within the preset area or disappears from the image within the preset area.
[0384] The aforementioned starting area refers to the area to which the starting point of the behavior corresponding to the pre-set target behavior type belongs.
[0385] The aforementioned endpoint area refers to the area where the movement ends or the target object disappears in the pre-defined target behavior type.
[0386] In this embodiment of the application, the executing entity can acquire and output the captured image of a preset area. Based on this, the user can set the starting point position and ending point position of the target behavior type for the captured image.
[0387] Based on this, the execution entity of this application embodiment can determine the preset start position and preset end position of the target behavior type set for the shooting screen, and determine the start area and end area respectively according to the preset start position and preset end position, and send the start area and end area to the shooting module or base station, so that the shooting module or base station can determine whether the object behavior belongs to the target behavior type according to the start area or end area.
[0388] As an optional implementation, the execution entity of this application embodiment can determine the preset starting point position and the preset ending point position based on the movement trajectory of the target behavior type set for the captured image.
[0389] As an exemplary implementation, a user can trigger (double-click, single-click, or long-press, etc.) the display screen of the execution subject of this application embodiment. The execution subject of this application embodiment can respond to the user's triggering operation by outputting a preset icon in the captured image. The preset icon may be a user image of the user who has logged into the execution subject of this application embodiment, or it may be a default icon, such as an arrow, a dot, a circle, etc. This application embodiment does not impose any restrictions on this.
[0390] Based on this, the execution subject of this application embodiment can identify the movement trajectory of the target behavior type in response to the setting of preset icons in the captured image.
[0391] As an exemplary implementation, setting a preset icon may include a drag operation. Further, after the user triggers the shooting screen and the preset icon is displayed on the shooting screen, the user can drag the icon to draw a movement trajectory corresponding to the target behavior type on the shooting screen. The execution entity of this application embodiment can, in response to the drag operation of the preset icon on the shooting screen, determine the drag trajectory corresponding to the drag operation and define the drag trajectory as the movement trajectory of the target behavior type set for the shooting screen.
[0392] Then, the starting point and ending point of the movement trajectory can be identified, and the starting point is determined as the preset starting point position of the target behavior type set for the shooting screen, and the ending point is determined as the preset ending point position of the target behavior type set for the shooting screen.
[0393] In one embodiment, the executing entity of this application embodiment may determine a third distance threshold and a fourth distance threshold, and determine a starting region based on a preset starting position and the third distance threshold, and determine an ending region based on a preset ending position and the fourth distance threshold.
[0394] The third and fourth distance thresholds can be preset distance thresholds, or they can be determined by the executing entity in this application embodiment based on multiple historical starting points and multiple historical ending points generated when the object performs a behavior corresponding to the target behavior type in a historical time period. For example, the maximum distance value between multiple historical starting points can be determined as the third distance threshold, and the maximum distance value between multiple historical ending points can be determined as the fourth distance threshold.
[0395] As an optional implementation, the starting area can be defined as the region centered on the starting position and with a third distance threshold as the radius, and the ending area can be defined as the region centered on the ending position and with a fourth distance threshold as the radius.
[0396] It is understood that in the above optional implementation, by acquiring and displaying a captured image of a preset area, a preset start position and a preset end position for the target behavior type set for the captured image are determined. Based on the preset start position and preset end position, a start area and an end area are determined respectively, which are used to determine whether the object's behavior belongs to the target behavior type. This technical solution, by displaying a captured image of a preset area on a preset terminal, allows the user to set a preset start position and a preset end position for the target object to perform the behavior corresponding to the target behavior type on the captured image. Based on the preset start position and preset end position, the start area and end area of the behavior corresponding to the target behavior type are determined respectively. This allows the shooting module or base station to identify the target object's behavior based on the aforementioned start area and end area, achieving fast and accurate homecoming trajectory drawing, thereby improving the efficiency and accuracy of object behavior identification.
[0397] In some optional implementations of this embodiment, the first security video includes a target image, which includes an object image and a person image, wherein the object image represents a target object and the person image represents a target person.
[0398] The target image can be any image that includes both object and person images. As an example, the target image can be a video frame extracted from a video that includes both object and person images. As yet another example, the target image can also be a video frame extracted from a video that includes both object and person images, where the object image represents the target object and the person image represents the target person, both located within a predetermined area.
[0399] The aforementioned preset area can be set by a user or other object, or it can be an area determined by the aforementioned executing entity or other electronic device that meets preset conditions. For example, the preset condition could be that the area includes a preset item. In this case, the preset item can refer to the same item as the target item.
[0400] The preset area can be a fixed area or an area whose position changes. For example, if the preset condition is that the area includes a preset item, and if the preset item (such as a robot vacuum cleaner, a pet, etc.) can move, then the preset area can be an area whose position changes.
[0401] The target item can be an item represented by an image, and the target person can be a person represented by an image.
[0402] Based on this, the first security video can be identified in the following way to obtain the video recognition result of the first security video:
[0403] The first step is to identify the first security video to determine the first detection box and the second detection box in the target image, wherein the first detection box is the detection box of the object image and the second detection box is the detection box of the person image.
[0404] Here, an object detection algorithm can be used to detect objects in the target image, thereby determining the first and second detection boxes in the target image.
[0405] Among them, object detection algorithms are algorithms in the field of computer vision used to identify and locate specific targets (such as target objects or people) in images (including the aforementioned target images). These algorithms are able to determine the location of the target (usually by drawing a rectangle or more complex shapes) and may include classifying the target.
[0406] Here, OVOD (Open Vocabulary Object Detection) can be used for object detection. OVOD allows the model to detect and recognize new object categories that have not been seen during the training phase, thereby achieving generalized object detection.
[0407] The second step is to determine the degree of overlap between the first detection box and the second detection box.
[0408] Here, the degree of overlap can represent the proportion or level of the overlapping portion of the first and second detection boxes within the whole. As an example, the degree of overlap can be represented by at least one of the following: the number of pixels overlapping between the image regions corresponding to the first and second detection boxes, the proportion of overlapping areas between the image regions corresponding to the first and second detection boxes, the intersection-union ratio (IUGR) of the image regions corresponding to the first and second detection boxes, etc.
[0409] The third step is to generate theft detection information based on the degree of overlap, wherein theft detection information indicates whether the target person has the intention to steal the target item.
[0410] Here, various methods can be used to generate theft detection information based on the degree of overlap.
[0411] As an example, when the degree of overlap is greater than or equal to a preset threshold, theft identification information indicating that the target person has the intent to steal the target item can be generated; when the degree of overlap is less than the preset threshold, theft identification information indicating that the target person does not have the intent to steal the target item can be generated. The preset threshold can be set by a user or other object, or it can be determined by statistically analyzing the correspondence between the degree of overlap and theft identification information.
[0412] As another example, when the degree of overlap is greater than or equal to a preset threshold, and both the target person and the target item are located in a preset area, theft identification information indicating that the target person has the intent to steal the target item can be generated; when the degree of overlap is less than the preset threshold, or when at least one of the target person and the target item is not located in the preset area, theft identification information indicating that the target person does not have the intent to steal the target item can be generated.
[0413] The fourth step is to determine the video recognition result of the first security video based on the theft detection information.
[0414] Here, the theft identification information can be determined as the video recognition result of the first security video to determine whether the first security video indicates the existence of a theft. Alternatively, the video recognition result of the first security video can be determined by determining whether the target of the theft indicated by the theft identification information has preset characteristics (e.g., whether it is wearing a courier's uniform, whether it is a family member, whether it is a stranger, etc.).
[0415] It is understood that, in some of the above-mentioned optional implementation methods, the degree of overlap between the detection bounding boxes of the object image and the detection bounding boxes of the person image in a single image can be used to determine whether the target person has the intention to steal the target item, thereby improving the efficiency and accuracy of theft intent identification.
[0416] In some application scenarios of the above-mentioned optional implementation methods, theft detection information can be generated based on the degree of overlap in the following manner:
[0417] The first step is to determine whether the degree of overlap is greater than or equal to a preset threshold.
[0418] The aforementioned preset threshold can be set by users or other objects, or it can be determined by statistically analyzing the correspondence between the degree of overlap and theft identification information.
[0419] The second step is to determine whether the behavior of the target person represented by the personnel image is theft if the degree of overlap is greater than or equal to the preset threshold, so as to obtain the first determination result.
[0420] The first determination result can indicate whether the behavior of the target person represented by the personnel image is theft.
[0421] Theft can be one or more acts that constitute theft. For example, theft could be bending down, reaching out and glancing sideways, or reaching out and glancing sideways while walking or running quickly.
[0422] The third step is to generate theft detection information based on the first determination result.
[0423] Here, various methods can be used to generate theft detection information based on the first determination result.
[0424] As an example, if the first determination result indicates that the behavior of the target person represented by the personnel image is theft, theft discrimination information indicating that the target person has the intent to steal the target item can be generated; if the first determination result indicates that the behavior of the target person represented by the personnel image is not theft, theft discrimination information indicating that the target person does not have the intent to steal the target item can be generated.
[0425] In addition, other methods can be used to generate theft identification information based on the first determination result. Please refer to the following description for details, which will not be elaborated here.
[0426] It is understandable that in the above application scenarios, theft detection information can be generated by determining whether the behavior of the target person represented by the personnel image constitutes theft. This can improve the accuracy of theft intent recognition.
[0427] In some application scenarios of the above-mentioned optional implementation methods, theft detection information can be generated based on the degree of overlap in the following manner:
[0428] The first step is to determine whether the degree of overlap is greater than or equal to a preset threshold.
[0429] The aforementioned preset threshold can be set by users or other objects, or it can be determined by statistically analyzing the correspondence between the degree of overlap and theft identification information.
[0430] The second step is to extract the associated video frame sequence of the target image from the first security video if the degree of overlap is greater than or equal to the preset threshold.
[0431] Here, the associated video frame sequence can be composed of video frames in the first security video that are associated with the target image.
[0432] As an example, the associated video frame sequence may include: the target image, N video frames of the target image preceding it, and M video frames of the target image following it.
[0433] As yet another example, the associated video frame sequence may also include: N video frames of the target image preceding the target image, and M video frames of the target image following the target image.
[0434] In the above examples, N and M are both positive integers, and the values of N and M can be equal or unequal. The preceding video frame is the video frame in the first security video that precedes the target image, and the following video frame is the video frame in the first security video that follows the target image.
[0435] As another example, the associated video frame sequence may also include: video frames in the first security video that include images of the target item, and / or video frames in the first security video that include images of the target person.
[0436] The third step is to generate theft detection information based on the associated video frame sequence.
[0437] Here, various methods can be used to generate theft detection information based on the associated video frame sequence.
[0438] As an example, the aforementioned associated video frame sequence can be input into a pre-trained Large Language Model (LLM) to generate theft detection information. The LLM can represent the correspondence between the prompt words, the associated video frame sequence, and theft detection information.
[0439] Among them, large-scale language models are natural language processing models based on deep learning technology, which have a very high number of parameters and powerful language understanding and generation capabilities.
[0440] As an example, the aforementioned large language model could be MLLM (Multimodal Large Language Models). MLLM builds upon LLM's ability to understand language by incorporating the ability to understand other modalities, enabling the understanding and generation of content involving multiple data types. Here, "modality" refers to different types of data input, such as text, images, audio, and video. Through training on massive amounts of data, multimodal large models can learn the complementarity and correlation between different modalities. For example, MLLM's input data could include the aforementioned associated video frame sequences and cue words, while its output data could be theft detection information.
[0441] In addition, other methods can be used to generate theft detection information based on the associated video frame sequence. Please refer to the description below for details, which will not be elaborated here.
[0442] It is understandable that in the above application scenario, when the overlap between the first and second detection boxes is greater than or equal to a preset threshold, theft discrimination information is generated based on the associated video frame sequence of the target image. This can further improve the efficiency and accuracy of theft intent recognition.
[0443] In some of the above application scenarios, the above solution is applied to the first device, where the data processing amount of the target image is less than the data processing amount of the video frames in the associated video frame sequence.
[0444] Based on this, theft detection information can be generated using the associated video frame sequence in the following manner:
[0445] The first step is to send the associated video frame sequence to the second device.
[0446] The second device is used to generate theft detection information based on the associated video frame sequence. The computing power of the second device is greater than that of the first device.
[0447] As an example, the first device described above could be an edge computing device. The first device can process video data (such as the target image mentioned above) acquired from a smart camera. This first device has a certain computing power, enabling real-time customized target property detection and human detection. Furthermore, the device includes a microphone and some audio-visual equipment, capable of repelling threats to property security and playing welcome messages to family members or those on a whitelist.
[0448] The second device mentioned above can be a home intelligent central control system (server). The second device can serve as a computing power center and intelligent center, equipped with a high-performance computing chip, capable of building multiple video streams, and capable of real-time processing of video stream behavior.
[0449] Here, the second device can input the associated video frame sequence into a pre-trained large-scale language model to generate theft detection information. Alternatively, the second device can also generate theft detection information based on the associated video frame sequence, the state information of the target item in the preceding video frame, and the state information of the target item in the subsequent video frame.
[0450] The second step is to receive the theft detection information returned by the second device in order to generate the theft detection information.
[0451] It is understandable that in the above situation, a first device with lower computing power can process video frames with smaller data processing volume, while a second device with higher computing power can process multiple video frames with larger data processing volume. Thus, the efficiency of identifying theft intent can be improved by the cooperation of the two.
[0452] In some of the above application scenarios, the associated video frame sequence includes: a preceding video frame and a following video frame of the target image. The preceding video frame is the video frame in the first security video that precedes the target image. The following video frame is the video frame in the first security video that follows the target image.
[0453] Based on this, theft detection information can be generated using the associated video frame sequence in the following manner:
[0454] The first step is to determine the first state information of the target item in the previous video frame.
[0455] The first state information can indicate the state of the target item in the preceding video frame. For example, the first state information can indicate the position of the target item in the preceding video frame, or it can indicate whether the target item represented by the image of the item in the preceding video frame is located in a preset area.
[0456] The second step is to determine the second state information of the target item in the subsequent video frame.
[0457] The second state information can indicate the state of the target item in the subsequent video frame. For example, the second state information can indicate the position of the target item in the subsequent video frame, or it can indicate whether the target item represented by the image in the subsequent video frame is located in a preset area.
[0458] The third step is to generate theft detection information based on the first state information, the second state information, and the associated video frame sequence.
[0459] Here, theft detection information can be generated in various ways based on the first state information, the second state information, and the associated video frame sequence.
[0460] As an example, if the first state information indicates the position of the target item in the preceding video frame, and the second state information indicates the position of the target item in the following video frame, then if the distance between the position indicated by the first state information and the position indicated by the second state information is greater than or equal to a preset distance threshold, then theft identification information can be further generated based on the associated video frame sequence; if the distance between the position indicated by the first state information and the position indicated by the second state information is less than the preset distance threshold, then theft identification information indicating that the target person does not have the intent to steal the target item can be generated.
[0461] In addition, other methods can be used to generate theft detection information based on the first state information, the second state information, and the associated video frame sequence. Please refer to the description below for details, which will not be elaborated upon here.
[0462] It is understood that, in the above situation, theft discrimination information can be generated based on the first state information, the second state information, and the associated video frame sequence, thereby further improving the accuracy of theft intent recognition.
[0463] In some of the examples described above, the first state information indicates whether the target item represented by the item image in the previous video frame is located in the first region, and the second state information indicates whether the target item represented by the item image in the subsequent video frame is located in the first region.
[0464] The first region can represent the aforementioned preset region, or the first region can be a region with a preset area and shape, and whose position is variable.
[0465] Based on this, theft detection information can be generated using the following method, based on the first state information, the second state information, and the associated video frame sequence:
[0466] The first step is to generate initial theft detection information based on the associated video frame sequence, when the first state information indicates that the target item represented by the item image in the previous video frame is located in the first area.
[0467] Here, various methods can be used to generate initial theft detection information based on the associated video frame sequence.
[0468] As an example, the aforementioned associated video frame sequence can be input into a pre-trained large-scale language model to generate initial theft detection information. This large-scale language model can represent the correspondence between the prompt words, the associated video frame sequence, and the initial theft detection information.
[0469] As another example, initial theft detection information can also be generated based on the associated video frame sequence and whether the target person and the target item are both located in a preset area.
[0470] The initial state information indicates whether the target person has the intention to steal the target item.
[0471] The second step involves generating final theft identification information indicating that the target person has the intent to steal the target item when the initial theft identification information indicates that the target person has the intent to steal the target item, and the second state information indicates that the target item represented by the item image in the later video frame is not located in the first area.
[0472] As can be understood, in the above example, the theft intent of the target person is only finally determined when the target item changes from being located in the first area to being located outside the first area, and the initial state information indicates that the target person has the intent to steal the target item. This further improves the accuracy of theft intent identification.
[0473] In some of the above-mentioned optional implementation methods, theft detection information can be generated based on the degree of overlap in the following manner:
[0474] The first step is to determine whether the target personnel and the target items are both located in a preset area, in order to obtain a second determination result.
[0475] The second determination result mentioned above can indicate whether the target person and the target item are both located in the preset area.
[0476] The aforementioned preset area can be set by a user or other object, or it can be an area determined by the aforementioned executing entity or other electronic device that meets preset conditions. For example, the aforementioned preset condition could be that the area includes a preset item. In this case, the aforementioned preset item can refer to the same item as the aforementioned target item.
[0477] The preset area can be a fixed area or an area whose position changes. For example, if the preset condition is that the area includes a preset item, and if the preset item (such as a robot vacuum cleaner, a pet, etc.) can move, then the preset area can be an area whose position changes.
[0478] The second step involves generating theft detection information based on the second determination result and the degree of overlap.
[0479] Here, theft detection information can be generated in various ways based on the second determination result and the degree of overlap.
[0480] As an example, if the second determination result indicates that both the target person and the target item are located in a preset area, and the degree of overlap is greater than or equal to the preset threshold, theft discrimination information indicating that the target person has the intention to steal the target item can be generated; if the second determination result indicates that at least one of the target person and the target item is not located in the preset area, or the degree of overlap is less than the preset threshold, theft discrimination information indicating that the target person does not have the intention to steal the target item can be generated.
[0481] As another example, if the second determination result indicates that both the target person and the target item are located in a preset area, and the degree of overlap is greater than or equal to the preset threshold, it can be further determined whether the behavior of the target person represented by the personnel image is theft, thus obtaining a first determination result. The first determination result indicates whether the behavior of the target person represented by the personnel image is theft. Subsequently, based on the first determination result, theft discrimination information is generated.
[0482] In addition, other methods can be used to generate theft detection information based on the second determination result and the degree of overlap. Please refer to the description below for details, which will not be elaborated here.
[0483] It is understandable that, in the above situation, theft detection information can be generated based on both the second determination result and the degree of overlap. This can further improve the accuracy of identifying theft intent.
[0484] In some of the above-mentioned optional implementations, the following steps may also be performed before determining the degree of overlap between the first detection box and the second detection box:
[0485] The first step is to obtain a pre-entered set of personnel information.
[0486] The personnel information in the above-mentioned personnel information set can represent a family member and their relatives and friends.
[0487] In practice, personnel information can be entered by collecting images of relevant personnel.
[0488] The second step is to determine whether the personnel information set includes target personnel information representing the target personnel.
[0489] Based on this, the degree of overlap between the first detection box and the second detection box can be determined even if the target personnel information is not included in the personnel information set.
[0490] Optionally, if the target person's information is included in the personnel information set, it is not necessary to determine the degree of overlap between the first and second detection boxes. Furthermore, theft discrimination information indicating that the target person does not have the intent to steal the target item can be generated.
[0491] It is understandable that, in the above situation, if the target person's information is not included in the personnel information set, the degree of overlap between the detection boxes of the item image and the personnel image in the image can be used to determine whether the target person has the intention to steal the target item. This can improve the accuracy of theft intent recognition. Furthermore, in scenarios where it is determined that the target person has the intention to steal the target item and an alarm is required, the disturbance caused to users and other parties by frequent alarm prompts can be reduced.
[0492] In some of the above-mentioned optional implementation methods, after generating theft identification information, if the theft identification information indicates that the target person has the intent to steal, the expulsion device can be further controlled to perform an expulsion operation, and / or a prompt message can be sent to a preset terminal.
[0493] The expulsion device can be an audio output device, a mobile robot, etc.
[0494] When the ejection device is an audio output device, the ejection operation can output an alarm prompt audio. This audio can be set by a user or other entity.
[0495] When the expulsion device is a mobile robot, the expulsion operation can be carried out by moving towards the location of the target person.
[0496] The default terminal can be a device pre-associated with the aforementioned execution entity. As an example, the default terminal can be a device logged in with an administrator account.
[0497] It is understandable that in the above situation, by controlling the expulsion device to perform the expulsion operation, the probability of the target item being stolen can be reduced, and by sending a prompt message to the preset terminal, the user of the preset terminal can be promptly informed that the target item may be or is about to be stolen.
[0498] In some optional implementations of this embodiment, the first security video is a frame-sampling result of the target video.
[0499] Based on this, the frame extraction result is generated in the following manner:
[0500] The first step is to obtain an image description data set and a target video, wherein the image description data in the image description data set is used to describe the content of the target image, and the target video consists of an image sequence.
[0501] Here, the image description data set may include at least one image description data. In some cases, the cardinality of the image description data set may be less than or equal to a preset value, such as 30. The cardinality of the image description data set can be the number of image description data included in the image description data set.
[0502] Image description data in an image description dataset can be determined in a variety of ways.
[0503] As an example, the image description data in the image description dataset can be event description content input by users or other objects, or it can be the feature vector of event description content input by users or other objects.
[0504] The event description can be text, audio, or image used to describe the event in the image. For example, the event description can be used to describe at least one of the following events in the image: an elderly person falls, a child goes out, a stranger breaks in, a pet damages property, or a window is opened.
[0505] As yet another example, the image description data in the image description dataset can also be image feature data of one or more frames of images.
[0506] A single image description in the aforementioned image description dataset can be used to describe one or more events in an image. For example, a single image description can be used to describe two events: a child going out and a stranger intruding.
[0507] The aforementioned image description data set can be fixed or updated according to a preset strategy.
[0508] The second step is to calculate the similarity between the images in the image sequence and each image description data in the image description data set, and obtain the target similarity corresponding to the image.
[0509] Here, the target similarity corresponding to the image can be: the result of a weighted summation of the similarities between the images in the calculated image sequence and each image description data in the image description data set. Alternatively, the target similarity corresponding to the image can also be: the similarity between the image and each image description data in the image description data set that meets the first similarity condition. The first similarity condition is used to determine the target similarity corresponding to each image in the image sequence.
[0510] Here, for each frame in the image sequence, the similarity between the image and each image description data in the image description data set can be calculated to obtain the target similarity corresponding to the image.
[0511] The target similarity corresponding to the image is: the similarity between the image and each image description data in the image description data set that meets the first similarity condition.
[0512] The first similarity condition can be a pre-defined similarity condition.
[0513] As an example, the first similarity condition could be: the highest numerical similarity among the similarities between the image and each image description data in the image description data set. In this case, the target similarity for the image is: the highest numerical similarity among the similarities between the image and each image description data in the image description data set.
[0514] As another example, the first similarity condition can also be: the similarity between the image and each image description data in the image description data set whose numerical value is greater than or equal to the first similarity threshold. In this case, the target similarity corresponding to the image is: the similarity between the image and each image description data in the image description data set whose numerical value is greater than or equal to the first similarity threshold.
[0515] As another example, the first similarity condition could also be: the third (e.g., 1, 3, 5, etc.) similarity scores among the image descriptions in the image description data set, arranged in descending order of numerical value. In this case, the target similarity score for the image is the third (e.g., 1, 3, 5, etc.) similarity score among the image descriptions in the image description data set, selected in descending order of numerical value.
[0516] For example, if the image sequence includes two images: image 1 and image 2, and the image description data set includes three image description data: image description data 1, image description data 2, and image description data 3, the first similarity condition is: the image with the highest similarity value among the image description data in the image description data set. That is, the third similarity condition is 1.
[0517] Therefore, we can calculate the similarity between image 1 and image description data 1, the similarity between image 1 and image description data 2, and the similarity between image 1 and image description data 3, thus obtaining three similarity scores for image 1: similarity 1, similarity 2, and similarity 3. The similarity score with the highest value among similarity 1, similarity 2, and similarity 3 is determined as the target similarity score for image 1.
[0518] Similarly, the similarity between image 2 and image description data 1, the similarity between image 2 and image description data 2, and the similarity between image 2 and image description data 3 can be calculated to obtain three similarity scores for image 2: similarity 4, similarity 5, and similarity 6. The similarity score with the highest value among similarity scores 4, 5, and 6 is determined as the target similarity score for image 2.
[0519] In some cases, the target similarity corresponding to the above image includes the maximum similarity between the image and each image description data in the image description data set.
[0520] It can be understood that through the second step described above, each frame in the image sequence can obtain at least one corresponding target similarity. Therefore, each target similarity can correspond to one image frame.
[0521] The third step is to select a first number of target similarities from the calculated target similarities.
[0522] Here, a first number of target similarities that meet the second similarity condition can be selected from the calculated target similarities.
[0523] The second similarity is used to select at least a portion of the target similarities from the calculated target similarities.
[0524] The first quantity can be any quantity. The first quantity is only used to represent the number of selected target similarities. For example, the first quantity can be a preset quantity or a non-preset quantity. When the first quantity is a non-preset quantity, step 103 can be implemented as: selecting multiple target similarities from the calculated target similarities. In this case, the number of target similarities among the selected multiple target similarities is the first quantity.
[0525] The second similarity condition can be a pre-set similarity condition that is different from the first similarity condition mentioned above.
[0526] As an example, the second similarity condition could be: the highest similarity among the calculated target similarities. In this case, the first quantity could be represented as 1.
[0527] As another example, the second similarity condition could also be: among the calculated target similarities, those with values greater than or equal to the second similarity threshold. In this case, the first quantity could be the number of similarities with values greater than or equal to the second similarity threshold.
[0528] As another example, the second similarity condition could also be: the first number (e.g., 1, 3, 5, etc.) of similarities that appear at the top of the calculated target similarities. In this case, the first number can be a preset positive integer.
[0529] Here, the first number of target similarities that meet the second similarity condition include the target similarity with the largest value among the calculated target similarities.
[0530] The fourth step is to determine the first image set corresponding to the first number of target similarities, wherein the images in the first image set correspond one-to-one with the target similarities in the first number of target similarities.
[0531] Since each target similarity corresponds to one image frame, the first image set corresponding to the first number of target similarities can be determined, wherein the images in the first image set correspond one-to-one with the target similarities in the first number of target similarities.
[0532] Here, the number of images in the first image set is the first number.
[0533] Fifth step: Based on the first image set, determine the frame extraction result of the target video.
[0534] The frame extraction result can be at least one frame from the target video.
[0535] As an example, the first set of images can be determined as the frame extraction result of the target video.
[0536] In addition, other methods can be used to determine the frame extraction results of the target video based on the first image set. Please refer to the following description for details, which will not be elaborated here.
[0537] It is understandable that, in the above optional implementation methods, since image description data can be used to describe the target image content that users and other objects are concerned about, such as behaviors and events appearing in images that users and other objects are more concerned about, the matching degree between the frame extraction result and the target image content can be improved by first determining the target similarity corresponding to each frame image, and then selecting a first number of frame images from the images included in the target video based on the target similarity corresponding to each frame image, and determining the frame extraction result of the target video accordingly.
[0538] In some application scenarios of the above-mentioned optional implementation methods, the frame extraction result of the target video can be determined based on the first image set in the following manner:
[0539] The first step is to display the first set of images.
[0540] The device executing the above video frame extraction method can be a terminal, such as a smartphone, tablet, or computer.
[0541] Here, after determining the first set of images, they can be displayed.
[0542] As an example, each frame of the target video can be displayed, wherein the images in the first set of images can be highlighted in the target video.
[0543] The second step is to determine whether an adjustment operation is detected for the images in the first image set; wherein the adjustment operation is used to: adjust the images in the first image set to obtain a second image set.
[0544] The second image set can be a set of images obtained by adjusting the images in the first image set. The number of images in the second image set can be a second number. The second number can be a preset number or a number determined through adjustment operations. The second number can be greater than, less than, or equal to the first number.
[0545] The second image set can represent the image obtained after performing adjustment operations on the images in the first image set.
[0546] As an example, in the case where the adjustment operation means replacing image A in a first number of frames with image B in the target video, the first number can be equal to the second number.
[0547] As yet another example, in the case where the adjustment operation indicates that image A is deleted from a first number of frames, the first number can be greater than the second number.
[0548] As another example, in the case where the adjustment operation indicates that image C is added to a first number of frames, the first number can be less than the second number.
[0549] Third, upon detecting the adjustment operation, the second image set is determined as the frame extraction result of the target video.
[0550] It is understandable that in the above application scenarios, adjusting the operation to adjust the final frame extraction result can further improve the matching degree between the final determined frame extraction result and the target image content that users and other objects are concerned about.
[0551] In some of the above application scenarios, upon detecting the adjustment operation, the image description data set can be updated in the following manner:
[0552] The first step is to determine the image features of each image in the second image set to obtain an image feature set, wherein the image features in the image feature set correspond one-to-one with the images in the second image set.
[0553] In the second image set, each frame can correspond to an image feature, thus obtaining the image feature set.
[0554] Image features may include at least one of the following: color features, texture features, corner features, etc.
[0555] The second step is to determine the image description data based on the image feature set.
[0556] As an example, the above set of image features can be identified as image description data.
[0557] As another example, when image features are represented by vectors, the mean of the second number of vectors representing the image feature set can be used as the image description data.
[0558] The third step is to update the image description data set based on the determined image description data.
[0559] As an example, the determined image description data can be added to the image description data set to update the image description data set.
[0560] In addition, other methods can be used to perform the third step mentioned above, which will not be elaborated here.
[0561] Understandably, in the above scenario, the image description data set can be automatically iteratively updated based on the adjustment operations performed by users and other objects at different times, so that the image description data set is more in line with the current preferences of users and other objects. In turn, the matching degree between the frame extraction results at different times and the target image content that users and other objects are interested in at different times can be dynamically maintained or even improved.
[0562] In some of the examples above, the image description data set can be updated based on the determined image description data in the following manner:
[0563] The first step is to determine the cardinality of the image description data set before the update.
[0564] The cardinality of the image description dataset can represent the number of image description data in the image description dataset.
[0565] The second step is to determine whether the base number is less than a preset value.
[0566] The preset value can be a predetermined integer. For example, the preset value can be 10, 20, 30, 40, etc.
[0567] Third, if the base number is less than the preset value, add the determined image description data to the image description data set to obtain the updated image description data set; if the base number is greater than or equal to the preset value, replace any image description data included in the image description data set before the update with the determined image description data to obtain the updated image description data set.
[0568] It is understandable that, in the above example, when the cardinality of the image description data set is small, the determined image description data can be added to the image description data set to increase the cardinality of the updated image description data set; when the cardinality of the image description data set is large, the image description data set can be updated by replacement. Thus, it can be ensured that the cardinality of the image description data set is within a certain range, avoiding resource waste caused by an excessive cardinality of the image description data set.
[0569] In some of the solutions in the above examples, the image description data included in the set of image description data before the update can be replaced with the determined image description data in the following ways:
[0570] The first step is to determine the image description data with the earliest addition time from the image description data set before the update.
[0571] The addition time refers to the time when the image description data is added to the image description data set.
[0572] During the process of adding descriptive data to the image description dataset, the time of addition can be recorded. Therefore, the image description data with the earliest addition time can be determined from the image description dataset before the update.
[0573] The second step is to use the determined image description data to replace the image description data with the earliest added time in the image description data set before the update.
[0574] It is understandable that in the above scheme, by using the determined image description data to replace the image description data with the earliest added time in the image description data set before the update, the updated image description data set can more accurately reflect the preferences of users and other objects for the frame extraction results at the current time.
[0575] In some application scenarios of the above-mentioned optional implementation methods, the image description data set includes at least one image description data.
[0576] Based on this, the image description data can be determined in the following way:
[0577] The first step is to obtain the event description content input by the object, wherein the event description content is used to describe one or more events.
[0578] The event description can be entered by users or other entities.
[0579] The second step is to determine the characteristic data of the event description content.
[0580] Among them, the feature data of the above event description content can represent semantic features, lexical features, etc.
[0581] The third step is to determine the feature data as image description data.
[0582] It is understandable that in the above application scenario, since the image description data set includes feature data of the event description content input by the object, users and other objects can determine the events they are interested in by inputting the event description content. Thus, the subsequent frame extraction results can be used to determine whether the events that users and other objects are interested in have occurred.
[0583] The target video can be any video. As an example, the target video can be a video captured by a camera.
[0584] In some application scenarios of the above optional implementation methods, the similarity between the images in the image sequence and each image description data in the image description data set can be calculated in the following way to obtain the target similarity corresponding to the image:
[0585] The first step is to calculate the similarity between the images in the image sequence and each image description data in the image description data set, thereby obtaining the similarity set corresponding to the image.
[0586] The second step is to determine the target similarity of the image as the image with the highest similarity value in the similarity set corresponding to the images in the image sequence.
[0587] For example, if the image sequence includes two frames of images: image 1 and image 2, the image description data set includes three image description data: image description data 1, image description data 2, and image description data 3.
[0588] Therefore, we can calculate the similarity between image 1 and image description data 1, the similarity between image 1 and image description data 2, and the similarity between image 1 and image description data 3, thus obtaining three similarity scores for image 1: similarity 1, similarity 2, and similarity 3. In this case, the similarity set corresponding to image 1 includes three similarity scores, namely similarity 1, similarity 2, and similarity 3. Then, the similarity score with the highest value among similarity scores 1, 2, and 3 is determined as the target similarity score for image 1.
[0589] Similarly, the similarity between image 2 and image description data 1, the similarity between image 2 and image description data 2, and the similarity between image 2 and image description data 3 can be calculated, thus obtaining three similarity scores for image 2: similarity 4, similarity 5, and similarity 6. In this case, the similarity set corresponding to image 2 includes three similarity scores, namely similarity 4, similarity 5, and similarity 6. Then, the similarity score with the highest value among similarity scores 4, 5, and 6 is determined as the target similarity score for image 2.
[0590] It is understood that in the above application scenario, the target similarity corresponding to each frame in the image sequence is the maximum similarity among the calculated similarities. Therefore, the target similarity corresponding to each image can more accurately reflect the degree of matching between the image and the image description data. Thus, when the image description data is used to describe events of interest to users or other objects, the above-mentioned optional implementation method can determine the frame extraction results that are of greater interest to users or other objects.
[0591] In addition, various methods can be used to determine the similarity between an image and its image description data.
[0592] As an example, we can first use convolutional neural networks (CNNs), autoencoders, generative adversarial networks (GANs), etc., to extract feature vectors from the image and feature vectors from the image description data. Then, we can determine the similarity between the image and the image description data by using the Euclidean distance, cosine similarity, or Pearson correlation coefficient between the feature vectors of the image and the image description data.
[0593] In some application scenarios of the above optional implementation methods, after generating the frame extraction result, the following steps can also be performed:
[0594] The first step is to determine the push information to be pushed to the preset terminal based on the frame extraction results.
[0595] Here, the first step described above can be performed in several ways.
[0596] The preset terminal can be a pre-defined terminal. For example, the preset terminal can be a terminal that performs the above adjustment operations.
[0597] In some of the above application scenarios, the following method can be used to determine the push information to be pushed to the preset terminal based on the frame extraction results:
[0598] The first step is to generate video segments of the target video based on the frame extraction results.
[0599] For example, the extracted frames can be used as video segments of the target video. As another example, the multiple frames represented by the extracted frames can be processed with background music, dubbing, filters, etc., to obtain video segments of the target video.
[0600] The second step is to determine the video information of the video clip as push information to be pushed to a preset terminal.
[0601] The aforementioned video information may refer to at least one of the following: the video clip's cover, title, or summary.
[0602] It is understandable that, in the above situation, after the video clip is generated, the video information of the video clip can be pushed to the preset terminal in a timely manner so that the user of the preset terminal can understand the video information of the video clip in a timely manner.
[0603] In some of the above application scenarios, the following method can be used to determine the push information to be pushed to the preset terminal based on the frame extraction results:
[0604] The first step is to perform behavior recognition on the extracted frame results to obtain the recognition results.
[0605] The recognition result refers to the result obtained by performing behavior recognition on the frame extraction result. For example, the recognition result can represent the behavior that occurs in the image represented by the frame extraction result, such as a stranger loitering or an elderly person falling down.
[0606] The second step is to generate push information for pushing to preset terminals based on the recognition results.
[0607] Here, the identification results can be used as push notifications to preset terminals. Alternatively, the hazard level corresponding to the identification results can also be used as push notifications to preset terminals.
[0608] It is understood that in the above situation, behavior recognition can be performed on the frame extraction results, and then information can be pushed to the preset terminal based on the recognition results. In this way, the user of the preset terminal can know in a timely manner whether the behavior they are interested in has occurred.
[0609] Optionally, after obtaining the above identification results, one or more devices that match the above identification results, as well as the control method of the one or more matching devices, can be determined, and then the one or more devices can be controlled according to the control method.
[0610] For example, if the above identification result indicates that a stranger is loitering, then the device that matches the above identification result may include a camera, and the control method may indicate that the stranger is being tracked and filmed.
[0611] The second step is to push the push data to the preset terminal.
[0612] It is understood that in the above application scenario, the frame extraction results of the target video can be obtained. These frame extraction results are determined using the video frame extraction method described in the first aspect. Then, based on the frame extraction results, push information for pushing to a preset terminal is determined, and the push data is then pushed to the preset terminal. Therefore, by determining the push information using the frame extraction results determined by the above video frame extraction method and pushing this push information to the preset terminal, the timeliness of pushing push information matching the target image content to the preset terminal can be improved.
[0613] In some optional implementations of this embodiment, the security preference information is in natural language.
[0614] The natural language is used to determine the video to be pushed.
[0615] Here, natural language can be a language that evolves naturally with culture. For example, it can be languages such as Chinese and English.
[0616] In some cases, natural language can be collected via a terminal. When the executing entity in this embodiment is a server, after the terminal collects the natural language, it can send it to the executing entity in this embodiment. The terminal can be communicatively connected to the executing entity in this embodiment. Alternatively, when the executing entity in this embodiment is a terminal, natural language can be directly collected by the executing entity.
[0617] The aforementioned terminal can be either hardware or software. For example, a terminal can be an electronic device such as a mobile phone or computer, or it can be an application running on such an electronic device.
[0618] The natural language mentioned above can be represented in the form of text, audio, etc. As an example, natural language can be audio or text such as "Remind me when the child comes home starting tomorrow," or audio or text such as "Remind me when the cat wakes up tomorrow."
[0619] The aforementioned camera can be used to monitor a preset area, thereby generating video of the preset area, i.e., the first security video. The preset area can be the monitoring area of the camera.
[0620] In practice, the executing entity of this embodiment may first acquire natural language and then acquire the first security video; or it may first acquire the first security video and then acquire natural language; or it may acquire both natural language and the first security video simultaneously. This embodiment does not limit the order in which natural language and the first security video are acquired.
[0621] Based on this, the following steps can also be performed:
[0622] The first step is to determine the feature data of the natural language to obtain the first feature data.
[0623] Here, the first feature data can be feature data of the natural language. As an example, the first feature data can be semantic features of the natural language.
[0624] The second step involves determining a first video based on the first security video, and determining whether the first video matches the natural language based on at least two types of feature data from the first video and the first feature data.
[0625] Here, the first video can be any video directly captured by the camera, or it can be a video generated after processing the video captured by the camera.
[0626] The at least two feature data of the first video can be data features obtained by extracting features from the first video using at least two different data feature extraction methods.
[0627] In practice, various methods can be used to determine whether the first video matches the natural language, based on at least two feature data of the first video and the first feature data.
[0628] As an example, a first video and natural language can be input into a pre-trained artificial intelligence model, which extracts at least two types of feature data from the first video and a first feature data from the natural language, thereby obtaining discrimination information indicating whether the first video matches the natural language.
[0629] The aforementioned AI model can be trained in a supervised manner based on training samples, including sample videos, sample natural language, and sample discrimination information. The sample discrimination information indicates whether the sample video matches the sample natural language.
[0630] In addition, other methods can be used to determine whether the first video matches the natural language, as described below, and will not be elaborated here.
[0631] Third, if the first video matches the natural language, the first video is used as the video to be pushed, and video information of the video to be pushed is pushed, wherein the video information represents the information of the video to be pushed.
[0632] Here, in the case where the executing entity in this embodiment is a terminal, the executing entity can push the video information of the video to be pushed to users or other objects by displaying the video information of the video to be pushed. For example, the executing entity can be a smartphone, computer, etc.
[0633] In embodiments of this disclosure where the executing entity is a server, the executing entity can send the video information of the video to be pushed to the user's terminal (e.g., smartphone, computer) to push the video information of the video to be pushed to the terminal. After receiving the video information of the video to be pushed, the terminal can display the video information of the video to be pushed.
[0634] The video to be pushed can be used to push to the aforementioned terminals or users.
[0635] Video information can be a notification message or address of the video to be pushed, or it can be the video itself.
[0636] In practice, if the first video matches the natural language, the first video can be pushed as the video to be pushed. Alternatively, the video information (e.g., a prompt message) of the video to be pushed can be pushed first. If a preset operation (e.g., a video playback operation) is detected through the video information, the first video can be pushed as the video to be pushed.
[0637] It is understandable that, among the above optional implementation methods, by pre-setting natural language, automatic reminders can be achieved when an event represented by that natural language is triggered in the first security video. This can push video information triggered by the corresponding event more accurately and / or in a timely manner, thereby improving the accuracy and / or timeliness of users and other objects in judging whether the events they are concerned about have occurred.
[0638] In some application scenarios of the above-mentioned optional implementation methods, the following method can be used to determine whether the first video matches the first feature data based on at least two feature data of the first video:
[0639] The first step is to determine at least two types of feature data from the image feature data, text feature data, and audio feature data of the first video to obtain the second feature data.
[0640] Here, the second feature data can be at least two of the image feature data, text feature data, and audio feature data of the first video.
[0641] In some cases, the second feature data may include image feature data, text feature data, and audio feature data of the first video.
[0642] The second step is to determine whether the first video matches the natural language based on the first feature data and the second feature data.
[0643] Here, the similarity between the first feature data and the second feature data can be calculated. Based on the relationship between this similarity and a preset similarity threshold, it can be determined whether the first video matches the natural language. For example, if the similarity is greater than or equal to the preset similarity threshold, then it can be determined that the first video matches the natural language. If the similarity is less than the preset similarity threshold, then it can be determined that the first video does not match the natural language.
[0644] It is understandable that in the above application scenario, by using multimodal fusion feature data from at least two dimensions of the natural language feature data and the image feature data, text feature data and audio feature data of the first video, it can be determined whether the first video matches the natural language. This can improve the accuracy of determining whether the first video matches the natural language.
[0645] In some application scenarios of the above optional implementation methods, the first video can be generated in the following manner:
[0646] The first step is to extract event frames from the first security video.
[0647] An event frame can be one or more video frames that represent an event.
[0648] There are several ways to extract event frames from the primary security video, which will be described later and will not be elaborated here.
[0649] The second step is to identify the extracted multi-frame event that represents the same event as the first video.
[0650] Wherein, the similarity between multiple event frames representing the same event is greater than or equal to a preset first threshold, and the similarity between multiple event frames representing different events is less than the preset first threshold.
[0651] Here, the acquired video may include video frames from multiple different events. Therefore, the multiple event frames corresponding to each event can be defined as a first video. That is, each event can correspond to one first video. For example, the acquired video may include video frames representing event 1 as follows: video frame 5, video frame 6, and video frame 7, and also includes video frames representing event 2 as follows: video frame 15, video frame 16, video frame 17, and video frame 19. Therefore, the video composed of video frames 5, 6, and 7 can be defined as one first video, and the video composed of video frames 15, 16, 17, and 19 can be defined as another first video.
[0652] It is understandable that in the above application scenario, the first video can be extracted from the video continuously captured by the camera, on an event-by-event basis. In this way, once it is determined that the video captured by the camera has triggered the event represented by the natural language, the video information of the event video can be pushed. This can further improve the timeliness of users receiving video information triggered by the events they are interested in.
[0653] In some examples of the above application scenarios, event frames can be extracted from the first security video in the following way:
[0654] An event extraction model is used to extract event frames from the first security video.
[0655] Based on this, the event extraction model can be trained in the following way:
[0656] The first step is to obtain a training sample set.
[0657] The training samples in the training sample set include videos (i.e., sample videos), event times (sample event times), and event labels (i.e., sample event labels).
[0658] Event time can represent the start and end times of an event; alternatively, it can also represent the position of the event frame in the video.
[0659] Event labels can be event names, such as "electric shock" or "near a safe".
[0660] The third step involves using a machine learning algorithm, taking the videos included in the training samples of the training sample set as input data, and the event time and event labels as the expected output data, to train an event extraction model.
[0661] It is understandable that, in the above example, the event extraction model can be trained using video, event time, and event tags, and then the event extraction model can be used to extract event frames from the video. In this way, the accuracy of event frame extraction can be improved by using event tags to assist in the extraction of event frames.
[0662] In some examples of the above application scenarios, the following steps can also be performed:
[0663] The first step is to determine the playback speed of the non-target video segments in the acquired video as the first speed.
[0664] The second step is to determine the playback speed of the target video segment in the acquired video as the second speed.
[0665] The target video segment is composed of event frames; in other words, the target video segment is also the event video.
[0666] The second speed is less than the first speed.
[0667] Understandably, in the example above, the target video segment can be played at a slower speed than non-target video segments, which can enhance the user's immersive experience and reduce the cost for users to record their lives.
[0668] In some examples of the above application scenarios, the following steps can also be performed: generating descriptive text for the first video.
[0669] The descriptive text can be used to describe the content of the first video.
[0670] In practice, artificial intelligence models (such as large language models) can be used to generate descriptive text for the first video.
[0671] As an example, a first video consisting of one or more frames is input into an AI model, and the number of words in the output text is set. The AI model can then output a descriptive text for the first video. For instance, if a first video of a child and his mother returning home is input into the AI model, and the model is set to output a description of no more than 30 words, including information such as time, location, people and their attributes, and behavior, then the AI model can output, "At 11:30 today, there was a little boy wearing blue clothes returning home with his mother at the door."
[0672] In some cases, the descriptive text is used for terminal display. For example, if the executing entity is a server, it can send the descriptive text to the terminal so that the terminal displays the descriptive text. If the executing entity is a terminal, it can directly display the descriptive text.
[0673] As an example, see Figure 9 ,exist Figure 9 In the middle, the terminal displayed the descriptive text "08:30 AM, Robert and Lisa, Robert and Lisa brought the skateboard home" and "07:10 AM, courier, the courier wearing a blue hat delivered the package to the house and then left immediately".
[0674] As can be understood, in the example above, the user of the terminal can obtain the content of the first video through the descriptive text without playing the video, thus allowing the user to obtain the content of the first video more quickly.
[0675] In some application scenarios among the above optional implementation methods, the following steps can also be performed:
[0676] The first step is to identify at least one of the text and music that matches the first video, thereby obtaining the matching information for the first video.
[0677] The matching information may include at least one of the text and music that match the first video.
[0678] The second step is to fuse the matching information with the first video to obtain the second video.
[0679] The second video can be the result of fusing the matching information with the first video. For example, the second video can be a video obtained by adding matching text to the first video, or a video obtained by adding matching music to the first video.
[0680] In practice, one can first establish the relationships between the text and / or music corresponding to each of the various types of first videos. Therefore, by identifying the text and / or music associated with the first video, the matching information for that first video can be determined.
[0681] The third step is to perform the target operation on the second video if a target operation is detected.
[0682] The target operation includes at least one of the following: sharing, downloading, storing, and sending.
[0683] It is understandable that in the above application scenarios, the matching information of the first video can be automatically determined, and then the second video can be generated by merging the two, so as to more quickly perform operations such as sharing, downloading, storing and sending the second video.
[0684] In some application scenarios among the above optional implementation methods, the following steps can also be performed:
[0685] The first step is to determine one or more target video frames from the first video.
[0686] Wherein, the similarity between the target video frame and the preceding video frame is less than or equal to a preset second threshold, and the similarity between the target video frame and the following video frame is less than or equal to the preset second threshold. The preceding video frame is the video frame before the target video frame in the first video. The following video frame is the video frame after the target video frame in the first video.
[0687] The preset second threshold may be equal to or different from the preset first threshold. In some cases, the preset second threshold may be greater than the first similarity threshold, thereby allowing for more accurate identification of highlight video frames.
[0688] In some cases, machine learning models can be used to determine the target video frame from the first video.
[0689] The machine learning model described above can use unsupervised contrastive learning to measure the inter-frame similarity in the image encoder (e.g., image similarity and audio over time), defining video frames with significant differences between frames as target video frames, i.e., highlight video frames.
[0690] The second step is to identify the target video frame as the highlight video frame in the first video.
[0691] In some cases, a highlight video frame can be one or more consecutive video frames in the first video. Based on this, for video frames in the first video preceding the highlight video frame, the number of video frames between them and the highlight video frame is positively correlated with their playback speed; that is, the more video frames between them, the faster their playback speed. Conversely, for video frames in the first video following the highlight video frame, the number of video frames between them and the highlight video frame is negatively correlated with their playback speed; that is, the more video frames between them, the slower their playback speed.
[0692] It is understandable that in the above application scenario, the highlight video frames can be determined from the first video.
[0693] In some of the application scenarios of the above-mentioned optional implementation methods, the method is applied to the first device.
[0694] Here, "first device" can refer to either a terminal or a server. As an example, "first device" can be a camera.
[0695] Based on this, the video information of the video to be pushed can be sent in the following manner:
[0696] The first step is to obtain the location information of the second device.
[0697] The second device can be a different device from the first device. For example, the second device can represent a different terminal or server than the first device. For instance, if the first device is a camera, the second device could be a smartphone, computer, etc.
[0698] The location information mentioned above can indicate the location of the second device.
[0699] The second step is to determine the video information of the video to be pushed based on the location information.
[0700] Here, after establishing a pre-established correspondence between location information and video information, video information that corresponds to the location information obtained in the first step can be determined and used as the video information of the video to be pushed.
[0701] As an example, location information 1 can correspond to video information 1 of the video to be pushed, and location information 2 can correspond to video information 2 of the video to be pushed.
[0702] The third step is to push the video information to the second device.
[0703] It is understandable that in the above application scenarios, different video information of the video to be pushed can be sent to the second device even when the device is located in different locations.
[0704] Optionally, the following steps may also be performed: generating descriptive text for the first video.
[0705] Here, the steps described above in this application scenario can be implemented by referring to the application scenario described above. Please refer to the description above for details, which will not be repeated here.
[0706] In some cases, the descriptive text is used to determine whether the first video is a search result of a video search request sent by the terminal. For example, when the executing entity is a server, it can receive the video search request sent by the terminal and determine whether the first video is a search result of the video search request based on the descriptive text. As another example, when the executing entity is a terminal, it can obtain a video search request input by a user or other object and determine whether the first video is a search result of the video search request based on the descriptive text.
[0707] Among them, the video search request is used to perform video searches.
[0708] In practice, the similarity between the descriptive text and the first video can be calculated to determine whether the first video is a search result of the video search request.
[0709] It is understandable that the above solution can achieve faster video search by using the descriptive text of the first video.
[0710] In some examples of the above application scenarios, the video information of the video to be pushed can be determined based on the location information in the following way:
[0711] First, determine the location of the camera to obtain the target location.
[0712] The target location can refer to the location of the camera.
[0713] Next, the distance between the location indicated by the location information and the target location is determined to obtain the target distance.
[0714] The target distance can represent the distance between the location indicated by the location information and the target location.
[0715] Then, determine whether the target distance is greater than or equal to a preset distance threshold.
[0716] Subsequently, if the target distance is greater than or equal to the preset distance threshold, the video information of the video to be pushed is determined to be first information. If the target distance is less than the preset distance threshold, the video information of the video to be pushed is determined to represent second information.
[0717] The first information indicates a request to control the camera to monitor a target area. The second information indicates the location of the target area.
[0718] The target area is the region where the video to be pushed was generated by monitoring.
[0719] In the case where the second device represents a smartphone, if the target distance is greater than or equal to the preset distance threshold, then it can be assumed that the user using the second device is not at home; if the target distance is less than the preset distance threshold, then it can be assumed that the user using the second device is at home.
[0720] It is understandable that in the above example, when the camera is far from the second device (e.g., not at home), the user can be asked to control the camera to capture video of the event-triggered area so that the user can monitor it remotely. When the camera is close to the second device (e.g., at home), the location of the event-triggered area can be communicated to the user through the second device so that the user can arrive at the event-triggered area immediately.
[0721] To address the technical problem that the process of manually generating linkage control strategies by users in existing technologies is lengthy and prone to errors, thus affecting user experience, this application provides a linkage control strategy generation method and apparatus. This method automatically generates linkage control strategies for multiple target security devices that meet the target scenarios and configuration requirements included in the configuration requirements, by combining the device capability set and behavior rules of the security devices when receiving configuration requirement information input by the user. This achieves automatic generation of linkage control strategies that meet the user's linkage needs, saving time and costs while generating linkage control strategies efficiently and accurately, thereby improving user experience.
[0722] See Figure 10A This is a schematic diagram illustrating an application scenario provided by an embodiment of this application. For example... Figure 10A As shown, the application scenario may include: control terminal 10, base station 11, security equipment 12, security equipment 13, security equipment 14, and security equipment 15.
[0723] The control terminal 10 may include an input module, which may be a voice input module or a text input module. Users can speak configuration requirements to the control terminal 10 through the input module, or input text content corresponding to the configuration requirements to the control terminal 10 through the input module. This application embodiment does not limit this.
[0724] The aforementioned control terminal 10 can be a hardware device or software that supports network connectivity to provide various network services. When the control terminal is hardware, it can be a variety of electronic devices with a display screen, including but not limited to smartphones, tablets, laptops, desktop computers, etc. When the control terminal is software, it can be installed in the electronic devices listed above. Figure 10A The following explanation uses only the control terminal 10 as a computer as an example.
[0725] The input module can be a voice acquisition device connected to the control terminal 10, such as a microphone, or it can be the display screen of the control terminal, which has a text input box so that the user can input text information from the text input box.
[0726] The aforementioned base station 11 can be a type of home control device in the home security category, serving as the central hub for the product in the home, managing connected devices such as cameras and sensors. It can be used to learn behavioral information within a preset area (e.g., a user's home area) to derive daily behavioral rules for security devices.
[0727] Furthermore, the aforementioned base station 11 may include a controller, which may include a receiving module, which may be a data receiving device connected to the controller in the base station 11. That is, the user can upload configuration requirement information through the control terminal 10, and the control terminal can send the configuration requirement information to the receiving module. Therefore, the receiving module can obtain the configuration requirement information input by the user.
[0728] The aforementioned security devices 12, 13, 14, and 15 are home security devices such as cameras and projectors. They can also be smart home devices installed in the home, such as smart refrigerators, smart air conditioners, and smart TVs. This application embodiment does not limit this. The figures illustrate security devices 12 and 13 as cameras, and security devices 14 and 15 as smart TVs. These security devices can be used in different scenarios or in the same scenario; for example, security device 12 can be a camera in the living room, and security device 13 can be a camera in the yard. This application embodiment does not limit this.
[0729] In one embodiment, the execution subject of this application embodiment may be a control terminal 10, which can obtain data (such as the daily behavior rules of security equipment) by interacting with the base station 11, thereby realizing the linkage control strategy generation method provided in this application.
[0730] In another embodiment, the execution subject of this application embodiment may be the controller in the base station 11. The controller can obtain the data sent by the control terminal 10 (such as the configuration requirement information input by the user in the control terminal 10) through the receiving module, thereby realizing the linkage control strategy generation method provided in this application.
[0731] In one embodiment, when a user wants to control security devices in a home area, they can input the corresponding configuration requirements into the control terminal 10 via the input module using voice commands.
[0732] In another embodiment, when a user wants to control security devices in a home area, they can input the corresponding configuration requirements into the control terminal 10 via the input module.
[0733] Based on this, when the execution subject of this application embodiment is the control terminal 10, after receiving the above configuration requirement information, the control terminal 10 can obtain the daily behavior rules of the user's home area from the base station 11, and based on this, use the linkage control strategy generation method provided in this application embodiment to determine the linkage control strategy corresponding to the user's needs, and perform linkage control on all or part of the security devices in the above security devices 12, 13, 14 and 15 according to the linkage control strategy.
[0734] Optionally, when the executing entity in this embodiment is the controller within the base station 11, the control terminal 10, after receiving the configuration requirement information, can send the configuration requirement information to the controller of the base station 11. Then, the controller can use its own receiving module to receive the configuration requirement information sent by the control terminal 10, and use the linkage control strategy generation method provided in this embodiment to determine the linkage control strategy corresponding to the user's requirements, and perform linkage control on all or part of the security devices 12, 13, 14, and 15 according to the linkage control strategy.
[0735] In some optional implementations of this embodiment, the security device for performing the first security response operation can be determined in the following manner:
[0736] The first step is to receive configuration requirements information.
[0737] The aforementioned configuration requirement information is used to characterize the user's need for the linkage of multiple security devices, and can also be understood as the effect that multiple security devices want to achieve when linked. The aforementioned configuration requirement information can be voice information input by the user, for example, if the user wants the security devices to intelligently monitor the security of their backyard, they can say the following voice: "Help me protect the security of my backyard on weekdays," or it can be text information input by the user. This application embodiment does not limit this.
[0738] In one embodiment, the executing entity of this application embodiment can receive voice or text information input by the user in real time, and determine the collected voice or text information as configuration requirement information.
[0739] As an exemplary implementation, the input module provided in this application embodiment may include a linkage scene setting page, which contains linkage buttons. Based on this, when a user wants to perform linkage control on security devices, they can enter the linkage scene setting page, long-press the voice input button, and then speak the corresponding configuration requirements to the voice input module of this application embodiment, such as "Help me protect the security of my backyard on weekdays," "When a stranger lingers at my door, help me warn and drive them away," or "I want my house to be cleaned when I come home every day," and other different configuration requirements.
[0740] As another exemplary implementation, the input module of this application embodiment may have a linkage scene setting page, which contains linkage buttons. Based on this, when a user wants to perform linkage control on security devices, they can enter the linkage scene setting page and input the corresponding configuration requirement text in the input box of the linkage scene setting page, such as inputting configuration requirement text like "protect the security of the backyard on weekdays".
[0741] The second step is to parse the configuration requirement information using a preset model. The configuration requirement information includes the target scenario and configuration requirements.
[0742] The aforementioned preset model is a pre-trained information parsing model that can parse configuration requirement information. It can parse the input text or voice information to obtain the scene and configuration requirements corresponding to the input text or voice information.
[0743] The aforementioned target scenario refers to the scenario where a user wants to control multiple security devices in a coordinated manner, that is, the scenario corresponding to the configuration requirement information. It can be any sub-scenario of the user's daily work, study, or residence, such as the living room, front yard, or bedroom in the user's home. For example, if the user's configuration requirement information is to monitor the activities of a baby in the nursery and living room, then the target scenario could be the nursery and living room.
[0744] The above configuration requirements refer to the needs of security devices when users want to achieve coordinated control of multiple security devices, such as cooling, shooting, or lighting requirements.
[0745] In one embodiment, after receiving configuration requirement information, the execution subject of this application embodiment can parse the received configuration requirement information through a pre-trained preset model in order to more accurately understand the user's linkage control intention, thereby obtaining the target scenario and configuration requirements included in the configuration requirement information.
[0746] In order to more accurately determine the target scenario and configuration requirements of linkage control from the configuration requirement information, the preset model provided in this application embodiment can be a preset model obtained by further training on the basis of an existing model. For example, the preset model can be a large language model, which can be used to parse the input speech text. Such a large language model also exists in the prior art. Therefore, in this application embodiment, the above-mentioned existing large language model can be further trained to obtain the preset model in this application.
[0747] Based on this, when training the above-mentioned preset model, the execution subject of this application embodiment can obtain a set of configuration requirement information samples. The set of configuration requirement information samples may include multiple configuration requirement information samples, and each configuration requirement information sample may correspond to a standard scenario and standard configuration requirements.
[0748] Furthermore, in order to enable the preset model to more comprehensively recognize speech or text in different scenarios, the above-mentioned configuration requirement information sample set may include configuration requirement information samples for different scenarios. Each scenario may correspond to speech configuration requirement information samples with different accents and languages but the same meaning, or each scenario may correspond to text configuration requirement information samples with different expressions but the same text meaning.
[0749] Based on this, the executing entity of this application embodiment can use the aforementioned configuration requirement information sample set to train a preset initial model, thereby obtaining the predicted scenario and predicted configuration requirement corresponding to each configuration requirement information sample output by the initial model. The aforementioned initial model can be a model that already exists in the prior art, or it can be a reconstructed model; this application embodiment does not impose any restrictions on this.
[0750] Then, for each configuration requirement information sample, a first loss value can be determined between the predicted scenario and the corresponding standard scenario, and a second loss value can be determined between the predicted configuration requirement and the standard configuration requirement.
[0751] Then, based on the first and second loss values corresponding to each of the above configuration requirement information samples, it can be determined whether the preset stop training condition is met.
[0752] Optionally, if the above-mentioned conditions for stopping training are met, a pre-set model that has completed training can be obtained.
[0753] Conversely, if it is determined that the above-mentioned conditions for stopping training are not met, the initial model can continue to be trained using the above-mentioned configuration requirement information sample set until the trained model meets the above-mentioned conditions for stopping training, thereby obtaining the preset model that has been trained.
[0754] As an exemplary implementation, the execution entity of this application embodiment can determine whether the first loss value and the second loss value corresponding to each configuration requirement information sample are both less than a preset loss value threshold. Optionally, if the first loss value and the second loss value corresponding to each configuration requirement information sample are both less than the preset loss value threshold, or if the first loss value and the second loss value of configuration requirement information samples greater than a preset number threshold are both less than the preset loss value threshold, then it can be determined that the preset stop training condition is currently met.
[0755] The third step is to determine the target security device related to the configuration requirements from the set of device capabilities of multiple security devices, so as to obtain the security device used to perform the first security response operation.
[0756] Here, the behavioral rules of the target security devices in the target scenario can also be determined. Based on these behavioral rules and the device capability set of each target security device, a linkage control strategy is determined to control the target security devices to achieve the aforementioned configuration requirements in the target scenario according to the aforementioned linkage control strategy.
[0757] The above set of equipment capabilities corresponds to the functions of security equipment, that is, what the security equipment can be used for and what role it plays.
[0758] The aforementioned behavioral rules refer to the behavioral rules that target security devices typically follow in the aforementioned target scenarios. For example, in daily use, the air conditioner in a bedroom typically operates at what time and in what working mode.
[0759] In this embodiment of the application, in order to generate a linkage control strategy that more accurately corresponds to the user's configuration requirements, the execution entity of this embodiment of the application may determine the device capability set of each security device in a preset scenario, determine the target security device related to the above configuration requirements, and the behavior rules of the above target security device in the target scenario, and on this basis determine the linkage control strategy according to the above behavior rules and the device capability set of each target security device.
[0760] In one embodiment, the executing entity of this application embodiment can determine the device identifier of each security device among a plurality of security devices, and determine the device capability set corresponding to each security device based on the device identifier.
[0761] As an exemplary implementation, the executing entity of this application embodiment may pre-store the object model of each security device under a preset scenario, thereby determining the device capability set of each security device based on this, and determining the target security device related to the configuration requirements from multiple device capability sets. The preset scenario may at least include the target scenario.
[0762] The aforementioned object model refers to the digital representation of the physical entity (such as a camera, sensor, etc.) corresponding to each type of security equipment. It describes what the entity is (attribute), what it can do (service), and what information it can provide (event) from three dimensions: attributes, services, and events. Each object model can correspond to a type of security equipment, and the object model of this type of security equipment can include different models of security equipment within the same category.
[0763] Based on this, the executing entity of this application embodiment can determine the target object model corresponding to each of the multiple security devices in a preset scenario from a preset object model set. The aforementioned object model set may include multiple object models, and each object model may correspond to a type of security device.
[0764] Based on this, the device identifier of each security device can be determined, and the device capability set corresponding to each security device can be determined from the target object model according to the device identifier of each security device.
[0765] As a feasible implementation, in a preset scenario containing the target scenario, for each security device installed, the execution subject of this application embodiment can interact with the security device to obtain the device identifier of the security device and store the correspondence between the security device and the corresponding device identifier in the database.
[0766] Based on this, the executing entity of this application embodiment can directly obtain the device identifier of each security device from the database.
[0767] Then, the installation location of each security device can be determined.
[0768] As an exemplary implementation, in a preset scenario including the target scenario, each time a user installs a security device, the installation location of that security device can be stored in a preset database via a display interface. Based on this, the execution subject of this application embodiment can directly obtain the installation location of each security device from the aforementioned database.
[0769] The aforementioned display interface may include a 3D scene diagram of a preset scenario. Users can click on the installation location of a new security device in the 3D scene diagram and enter information such as the type of security device to store the installation location of the security device in the database.
[0770] As another exemplary implementation, the executing entity of this application embodiment can call a preset image acquisition device to acquire scene images corresponding to multiple security devices, and identify the scene images to determine the installation location of each security device.
[0771] Finally, based on the aforementioned equipment capability set and installation location, it can be determined whether the security equipment is the target security equipment related to the above configuration requirements.
[0772] As an example implementation, the target functions included in the user's configuration requirements can be determined first, and based on the device capability set of each security device, it can be determined whether the corresponding security device can achieve the above target functions.
[0773] Optionally, if it is determined that the security device can achieve the above-mentioned target functions, it means that the security device meets the functional requirements in the user's configuration requirements. In this case, in order to further determine whether the security device is a target security device that is completely related to the configuration requirements, the executing entity of this application embodiment can determine whether the installation location of the security device is located within the target scenario of the above-mentioned configuration requirements.
[0774] Optionally, when it is determined that the security device is located in the target scenario of the configuration requirements, it indicates that the security device meets the scenario requirements of the configuration requirements. Therefore, the security device can be determined as the target security device related to the configuration requirements.
[0775] For example, suppose a user's configuration requirement is "turn on the living room lights and set the living room temperature to 27 degrees Celsius after the child comes home." Then the target scenario for this configuration requirement could be the living room, and the configuration requirement would be: turn on the lights and set the temperature to 27 degrees Celsius after the child comes home.
[0776] Continuing the assumption, the security devices in a user's home may include: an air conditioner, a television, a refrigerator, a washing machine, smart lighting (living room), and smart lighting (bedroom), with the air conditioner, television, and refrigerator located in the living room. In this scenario, analyzing the above configuration requirements reveals that the target functions are lighting and cooling. Therefore, based on the device function sets of each security device, it can be determined that the air conditioner and all smart lights satisfy the aforementioned target functions. Thus, the air conditioner and smart lights can be identified as the initial target security devices that meet the functional requirements of the configuration.
[0777] Furthermore, in order to determine whether the aforementioned initial target security devices meet the scenario requirements of the aforementioned configuration requirements, the installation locations of the aforementioned air conditioner and all smart lights can be determined separately, and the initial target security devices installed in the living room are determined as the target security devices related to the configuration requirements, namely the air conditioner and the smart lights installed in the living room.
[0778] In one embodiment, in order to determine a linkage control strategy that better conforms to the daily behavior rules of users and devices, the executing entity of this application embodiment can further determine the behavior rules of the target security device in the target scenario, and then determine the linkage control strategy based on the behavior rules and the device capability set of each target security device, so as to control the target security device to achieve the above configuration requirements in the target scenario according to the linkage control strategy.
[0779] As an exemplary implementation, when determining the behavior rules of the target security device in a target scenario, the executing entity of this application embodiment can obtain the daily behavior rules of the target security device in a preset scenario within a preset historical time period. These daily behavior rules can be obtained from a base station, which can acquire the daily behavior rules in the preset scenario through local self-learning. The preset scenario may at least include the target scenario; for example, if the target scenario is a living room, the preset scenario may be the user's home, which may include bedrooms and a living room, etc.
[0780] Then, based on the daily behavior rules of the preset scenario, the behavior rules of the target security device in the target scenario can be determined.
[0781] As an exemplary implementation, the executing entity of this application embodiment can directly use the daily behavior rules of each target security device in a preset historical time period and in a preset scenario as the behavior rules of the target security device in the target scenario.
[0782] In order to obtain daily behavior rules under preset scenarios in real time, the execution subject in this application embodiment can perform AI self-learning through a base station to learn the daily behavior rules under the preset scenarios. This method of obtaining daily behavior rules through local self-learning via a base station protects user privacy because the data does not need to leave the device for learning and improvement. Furthermore, since it does not require network transmission, it reduces network latency and improves response speed.
[0783] Based on this, when the execution subject of this application embodiment obtains the daily behavior rules under the preset scenario, it can obtain the daily behavior rules under the preset scenario from the base station. The base station can obtain the daily behavior rules under the preset scenario through local self-learning.
[0784] The aforementioned base station can obtain the daily behavior rules under the preset scenario in the following ways: First, when a behavior event is detected under the preset scenario, it can be determined whether the execution object of the behavior event is a security device under the preset scenario.
[0785] Optionally, if the execution object of the behavior event is determined to be a security device in the aforementioned preset scenario, the device behavior information corresponding to the behavior event can be obtained. The aforementioned behavior event refers to an event in which the security device executes a corresponding function after being triggered in the preset scenario, thereby causing a change in the device state of the security device, such as a change in the monitoring direction of the camera or a change in the operating state of the air conditioner.
[0786] Then, the device behavior information can be stored in a preset database, and if no behavioral event is detected, the device behavior information stored in the database can be self-learned to obtain the daily behavior rules of the security device under the preset scenario.
[0787] As for how the executing entity of this application determines the linkage control strategy based on the behavior rules and the device capability set of each target security device, it can be explained in the following description through the process shown in Figure 3, which will not be detailed here.
[0788] Furthermore, to ensure that the determined linkage control strategy better meets the user's linkage control needs, the executing entity in this embodiment, after determining the linkage control strategy based on the behavior rules and device capability set of each target security device, can output the linkage control strategy through a visual interface. Based on this, the user can modify the linkage control strategy. Upon receiving a user's modification operation regarding the linkage control strategy, the executing entity in this embodiment can obtain the modified target linkage control strategy and update the linkage control strategy accordingly.
[0789] Based on this, the execution entity of this application embodiment can perform linkage control on the security equipment in the preset scenario according to the updated target linkage control strategy described above.
[0790] For example, let's take the user's interaction intent as tracking and removing strangers as an example:
[0791] The linkage control strategy can be generated through the following steps:
[0792] 1. The user enters their own ID into the main controller, starts the calibration process, and slowly circles the house once.
[0793] 2. The main controller determines the relative positional relationship between cameras by identifying the time sequence of users and the pan-tilt rotation angle of each camera.
[0794] 3. When a user opens the APP (Application), starts the voice input function, and submits a configuration request, such as "I hope that when strangers linger around my house on weekdays, I can receive push notifications and help drive them away, and then provide me with complete video recording data."
[0795] 4. After a period of processing, the APP will automatically generate the following message on the APP interface: "When the sensors and cameras in the yard (front yard / back yard) detect a non-familiar person, a message alarm will be pushed, all recording devices in the yard (front yard / back yard) will be activated to record, and the person will be tagged. At the same time, an alarm will be triggered. The linkage is effective 24 / 7 from Monday to Friday."
[0796] 5. Users can manually adjust the settings to include the effective time on weekdays from 9:00 AM to 6:00 PM, then click save to create a complete set of rules for tracking and removing strangers.
[0797] Furthermore, the following linkage effects can be achieved by following the above linkage control strategy:
[0798] 1. If a stranger intrudes into the perimeter of a house, any camera that detects the stranger will trigger a stranger tracking and removal linkage, and a message will be pushed to the user.
[0799] 2. The main controller collects images returned by each camera and uses AI image algorithms to determine the area the stranger is about to enter based on the direction and speed of movement in the image. It then wakes up the cameras in the designated area to start recording, tracking, and sound and light alarms.
[0800] 3. Once a stranger leaves the perimeter of the home, the camera will stop recording and trigger an alarm.
[0801] 4. The main controller collects recordings from multiple cameras; uses AI face detection algorithms to extract clear facial portraits from multiple recordings; and draws the intrusion trajectory of the stranger based on the order in which the stranger appears in the images of each camera.
[0802] 5. The main controller combines video recordings, facial images, intrusion trajectories, first appearance time, first appearance location, departure time, departure location, and other multimodal key information to form event cards for users to view.
[0803] It is understood that in the above-mentioned optional implementation methods, configuration requirement information is received and parsed using a preset model. This configuration requirement information includes a target scenario and configuration requirements. Target security devices related to the configuration requirements are determined from the device capability set of multiple security devices, and the behavior rules of these target security devices in the target scenario are also determined. Based on these behavior rules and the device capability set of each target security device, a linkage control strategy is determined to control the target security devices to fulfill the configuration requirements in the target scenario according to the linkage control strategy. This technical solution, by automatically generating linkage control strategies for multiple target security devices that conform to the target scenario and configuration requirements included in the configuration requirement information when receiving configuration requirement information input by the user, combined with the device capability set and behavior rules of the security devices, achieves automatic generation of linkage control strategies that meet the user's linkage requirements. This saves time and costs while efficiently and accurately generating linkage control strategies, improving the user experience.
[0804] In some application scenarios of the above-mentioned optional implementation methods, the security device can be controlled to perform the first security response operation in the following manner:
[0805] The first step is to determine the behavior rules of the target security device in the target scenario.
[0806] The second step is to determine a linkage control strategy based on the behavior rules and the device capability set of each target security device, so as to control the target security device to perform the first security response operation according to the linkage control strategy in the target scenario, so as to achieve the configuration requirements.
[0807] Please refer to the above description for the specific implementation methods disclosed above; they will not be repeated here.
[0808] In some of the above application scenarios, the following approach can be used to determine the linkage control strategy based on the behavioral rules and the device capability set of each target security device:
[0809] The first step is to analyze the configuration requirements and determine the triggering conditions for each target security device.
[0810] The triggering conditions mentioned above refer to the triggering conditions of the target security device in the target scene. For example, if the user's configuration requirement is to monitor the activities of the baby in the nursery and living room, then the target scene can be the nursery and living room, the target security device is the camera in the nursery and living room, and the triggering condition for the camera to capture the image is to identify the baby's activity.
[0811] In one embodiment, after inputting configuration requirements, the user can write their own linkage control requirements into the configuration requirements information. These linkage control requirements encompass the target scene, the target security devices within the target scene, and the triggering conditions for each target security device. Based on this, the execution entity of this application embodiment can parse the above configuration requirements to determine the triggering conditions for each target security device.
[0812] The second step is to determine the initial linkage control strategy based on the triggering conditions and the device capability set of each target security device.
[0813] The aforementioned initial linkage control strategy is a linkage control strategy between target security devices. It is a preliminary linkage rule set for the target security devices. For example, if the target security devices are cameras in the nursery and living room, then the initial linkage control strategy can be that when a camera with recognition function identifies that the baby is moving, the camera with specific camera function will take a picture of the baby.
[0814] In one embodiment, when determining the initial linkage control strategy between multiple target security devices, the executing entity of this application embodiment can determine the initial linkage strategy between multiple target security devices based on the triggering conditions and device capability set corresponding to each target security device.
[0815] As an exemplary implementation, the functions that the target security device can achieve can be determined based on the device capability set of the target security device, and the triggering conditions for triggering each target security device in the target scenario can be determined based on the triggering conditions.
[0816] Based on this, the target functions that each target security device needs to achieve in the target scenario can be determined according to the triggering conditions and the functions that it can achieve, and an initial linkage control strategy between one or more target security devices can be generated.
[0817] For example, suppose the configuration requirement is "to cool and illuminate the living room upon detecting that any family member has returned home." The preset scenarios corresponding to this configuration requirement could include two air conditioners, a television, multiple smart lights, a washing machine, and a refrigerator located in the living room. Based on this configuration requirement, the target security devices can be determined to be the two air conditioners and the smart lights located in the living room.
[0818] Assuming that both air conditioners (Air Conditioner 1 and Air Conditioner 2) can perform cooling functions based on their device capabilities, and the smart light can perform illumination functions, then from the above description, we can determine that Air Conditioner 1 will perform the cooling function, and the smart light will perform the illumination function. Furthermore, based on the configuration requirements, the trigger condition for both Air Conditioner 1 and the smart light is that a family member returns home. Therefore, the initial linkage control strategy can be determined as follows: when a family member returns home, Air Conditioner 1 and the smart light in the living room can be turned on.
[0819] The third step is to determine the linkage control strategy based on the initial linkage control strategy and the behavior rules.
[0820] After determining the initial linkage control strategy between target security devices, a further linkage control strategy can be determined based on the initial linkage control strategy and the behavior rules of each target security device.
[0821] As an exemplary implementation, the execution entity of this application embodiment can combine the initial linkage control strategy and behavior rules to determine at least one preset dimension of linkage control conditions. The preset dimensions include, but are not limited to, time dimension, device dimension, user dimension, trigger action, and execution action.
[0822] Then, the linkage control conditions can be input into a preset linkage control strategy model to obtain the linkage control strategy output by the linkage control strategy model. The aforementioned linkage control strategy model can be a pre-stored linkage rule format specification model, through which linkage control strategies conforming to preset format conditions can be generated.
[0823] For example, suppose the received configuration requirement is "monitor the security of the backyard during the user's working hours and issue a warning after detecting a stranger." Further suppose the target security devices corresponding to the above configuration requirement include: a human body sensor in the backyard, camera A, camera B, and a speaker. The initial linkage control strategy determined according to the above steps is: when the human body sensor in the backyard is triggered or the camera triggers detection, camera A in the backyard is linked to record video, while camera B performs patrol and flashing. When the target is determined to be a stranger, the speaker is linked to issue a warning.
[0824] Continuing to assume the behavior rules of each target security device are as follows: the human body sensor in the backyard detects and identifies users from Monday to Friday, specifically at 9:00 AM, 6:00 PM, and 9:00 PM each day; camera A records video upon receiving a signal from the human body sensor; camera B performs patrol and flashing actions upon receiving a signal from the human body sensor; and the speaker issues a warning upon receiving a stranger signal from the human body sensor.
[0825] Based on the initial linkage control strategy and the behavior rules of each target security device, the following linkage control conditions can be obtained, including but not limited to: Time dimension: Monday to Friday, 9:00 AM to 6:00 PM or 9:00 PM daily; Device dimension: Human body sensor, Camera A, Camera B, and speaker in the backyard; User dimension: Each family member who needs to go to and from get off work; Trigger action: The human body sensor detects an object and determines that the object is a stranger, then the speaker is triggered; Execution action: The human body sensor is used to detect objects and determine whether the detected object is a stranger, Camera A is used to record video, Camera B is used for patrol and flashing light processing, and the speaker is used to issue a warning.
[0826] Then, the above-mentioned linkage control conditions can be input into the preset linkage control strategy model, so that the model generates a linkage control strategy that conforms to the preset format conditions. The resulting linkage control strategy can be as follows: when the human body sensor in the backyard is triggered and detects an object, the backyard camera A is linked to record video, while camera B performs patrol and flashing. The human body sensor is used to identify whether the object is a stranger. If so, the horn will sound a warning. The effective time is from Monday to Friday, from 9:00 am to 6:00 pm or 9:00 pm every day, with a weekly cycle.
[0827] Furthermore, since user configuration requirements are generally set based on themselves and their family members, each configuration requirement can correspond to at least one target user, which can be a family member or a stranger. Based on this, the executing entity in this embodiment can further adjust the obtained linkage control strategy according to the target user's user behavior rules, so that the linkage control strategy better conforms to the user's daily behavior patterns. Here, the aforementioned target user refers to the user involved in the configuration requirement information, who can be a family member with historical behavior rules or a stranger without historical behavior rules; this embodiment does not impose any restrictions on this.
[0828] In one embodiment, when the target user is a pre-recorded family member, meaning the target user has historical behavioral rules in the target scenario, the user behavior rules of the target user in the target scenario can be obtained. The linkage control strategy is then adjusted based on these user behavior rules.
[0829] As an exemplary implementation, when determining the user behavior rules of a target user in a target scenario, the executing entity of this application embodiment may first obtain the daily behavior rules of the target user in a preset scenario within a preset historical time period from the base station. The base station can obtain the daily behavior rules in the preset scenario through local self-learning, and the preset scenario may at least include the target scenario. For example, if the target scenario is a living room, the preset scenario may be the user's home, which may include bedrooms and a living room, etc.
[0830] Then, based on the daily behavior rules of the target users in the preset scenarios, the user behavior rules of the target users in the target scenarios can be determined.
[0831] In order to obtain daily behavior rules under preset scenarios in real time, the execution subject in this application embodiment can perform AI self-learning through a base station to learn the daily behavior rules under the preset scenarios. This method of obtaining daily behavior rules through local self-learning via a base station protects user privacy because the data does not need to leave the device for learning and improvement. Furthermore, since it does not require network transmission, it reduces network latency and improves response speed.
[0832] Based on this, when the execution subject of this application embodiment obtains the daily behavior rules under the preset scenario, it can obtain the daily behavior rules under the preset scenario from the base station corresponding to the preset scenario. The base station can obtain the daily behavior rules of each user under the preset scenario through local self-learning.
[0833] The aforementioned base station can obtain the daily behavior rules of each user in a preset scenario through the following methods: First, when a behavioral event is detected in the preset scenario, it can be determined whether the execution object of the behavioral event is a user in the preset scenario. Optionally, if it is determined that the execution object of the behavioral event is a user in the preset scenario, the user behavior information corresponding to the behavioral event can be obtained. The aforementioned behavioral event refers to the event corresponding to the behavior when the user's state changes in the preset scenario, such as the user entering the camera's shooting range.
[0834] The user behavior information can then be stored in a preset database. If no behavioral event is detected, the user behavior information stored in the database can be self-learned to obtain the user's daily behavior rules under the preset scenario.
[0835] Furthermore, in one embodiment, after determining the linkage control strategy between target security devices, the linkage control strategy can be adjusted according to the user behavior rules.
[0836] As an exemplary implementation, the execution entity of this application embodiment can combine the linkage control strategy and user behavior rules to redetermine at least one preset dimension of linkage control conditions. The preset dimensions include, but are not limited to, time dimension, device dimension, user dimension, triggering action, and execution action.
[0837] Then, the linkage control conditions can be input into a preset linkage control strategy model to obtain the linkage control strategy output by the linkage control strategy model. The aforementioned linkage control strategy model can be a pre-stored linkage rule format specification model, through which linkage control strategies conforming to preset format conditions can be generated.
[0838] For example, suppose the received configuration requirement is "monitor the security of the backyard during the user's working hours and issue a warning after detecting a stranger." Further suppose the target security devices corresponding to the above configuration requirement include: a human body sensor in the backyard, camera A, camera B, and a speaker. The initial linkage control strategy determined according to the above steps is: when the human body sensor in the backyard is triggered or a camera that has moved into the backyard triggers detection, camera A in the backyard is linked to record video, while camera B performs patrol and flashing actions. When the target is determined to be a stranger, the speaker is linked to issue a warning.
[0839] Continuing to assume the behavior rules of each target security device are as follows: the human body sensor in the backyard detects and identifies users from Monday to Friday, specifically at 9:00 AM, 6:00 PM, and 9:00 PM each day; camera A records video upon receiving a signal from the human body sensor; camera B performs patrol and flashing actions upon receiving a signal from the human body sensor; and the speaker issues a warning upon receiving a stranger signal from the human body sensor.
[0840] Based on the initial linkage control strategy and the behavior rules of each target security device, the following linkage control conditions can be obtained, including but not limited to: Time dimension: Monday to Friday, 9:00 AM to 6:00 PM or 9:00 PM daily; Device dimension: Human body sensor, Camera A, Camera B, and speaker in the backyard; User dimension: Each family member who needs to go to and from get off work; Trigger action: The human body sensor detects an object and determines that the object is a stranger, then the speaker is triggered; Execution action: The human body sensor is used to detect objects and determine whether the detected object is a stranger, Camera A is used to record video, Camera B is used for patrol and flashing light processing, and the speaker is used to issue a warning.
[0841] Then, the above-mentioned linkage control conditions can be input into the preset linkage control strategy model, so that the model generates a linkage control strategy that conforms to the preset format conditions. The resulting linkage control strategy can be as follows: when the human body sensor in the backyard is triggered and detects an object, the backyard camera A is linked to record video, while camera B performs patrol and flashing. The human body sensor is used to identify whether the object is a stranger. If so, the horn will sound a warning. The effective time is from Monday to Friday, from 9:00 am to 6:00 pm or 9:00 pm every day, with a weekly cycle.
[0842] Continuing to assume that the behavioral rules of the target users are: working from Monday to Friday, starting work at 9:00 am and finishing get off work at 6:00 pm, and taking a walk in the backyard at 9:00 pm.
[0843] Based on this behavioral rule, it can be determined that 9:00 PM is the target user's walking time, not their working hours. Therefore, the above-mentioned linkage control strategy can be adjusted as follows: when the human body sensor in the backyard is triggered and detects an object, it will trigger the backyard camera A to record video, while camera B will patrol and flash its lights. The human body sensor will also identify whether the object is a stranger. If so, it will trigger the horn to issue a warning. This will be effective from Monday to Friday, from 9:00 AM to 6:00 PM, on a weekly cycle.
[0844] Understandably, in the above scenario, by parsing the configuration requirements, the triggering conditions for each target security device are determined. Based on the triggering conditions and device capability set of each target security device, an initial linkage strategy is determined. Then, based on the initial linkage strategy and behavioral rules, a linkage control strategy is determined. This disclosed technology, by first determining the initial linkage control strategy corresponding to the target security device's functions, and then generating a corresponding linkage control strategy based on the target security device's behavioral rules and the initial linkage control strategy, can accurately and quickly automatically generate linkage control strategies that meet user configuration requirements. This saves time and costs while efficiently and accurately generating linkage control strategies, thus improving the user experience.
[0845] In some application scenarios of the above-mentioned optional implementation methods, the security device used to perform the first security response operation can be determined in the following manner:
[0846] The first step is to obtain the set of target images corresponding to the set of security devices.
[0847] A security device set may include at least two security devices. Security devices may include at least one of the following: a camera, an alarm, a sensor, etc.
[0848] For example, a single security device in a security equipment set can be a camera, an alarm, or a sensor. As another example, a single security device in a security equipment set can include both a camera and an alarm.
[0849] The target image set corresponding to the security equipment set can be a set of images obtained by capturing images of the security equipment in the security equipment set (hereinafter referred to as Method 1). For example, images of each security equipment in the security equipment set can be captured by one or more cameras to obtain the target image set corresponding to the security equipment set.
[0850] Alternatively, the target image set corresponding to the security equipment set can also be a set of images acquired by cameras in the security equipment set (hereinafter referred to as Method Two). For example, each security device in the security equipment set may include a camera. Thus, by acquiring images of the area within the shooting range of the cameras included in the security equipment set, the target image set corresponding to the security equipment set can be obtained.
[0851] The second step is to extract the calibration objects from the target image set.
[0852] The calibrator can be a pre-defined physical entity or an image of multiple physical entities with known positional relationships. For example, the calibrator can be an image of a pre-defined tree at the user's residence, or it can be an image of the user's front door and a pre-defined tree. Alternatively, the calibrator can be an image of the user, an object held by the user, etc., captured by the security equipment set as the user walks around the security equipment set.
[0853] The third step is to establish the association between the security devices based on the calibrated objects.
[0854] The association relationship can include at least one of the following: positional relationship (e.g., whether they are adjacent, relative, etc.), whether they work in pairs, or whether they share a field of view. Specifically, if two security devices include camera A and camera B, and if camera A is used to capture telephoto images of object A, and camera B is used to capture wide-angle images of object A, then the association relationship between the two security devices indicates that they work in pairs. The association relationship can include the association relationships between all or some of the security devices in the set of security devices.
[0855] As an example, when the target image set is obtained using the method described in Method 1 above, the relative positions of each security device in the security device set can be determined based on the target image set, and the determined relative positions can be used as the association relationship.
[0856] As another example, when the target image set is obtained by using the method described in Method 2 above, the positional relationship (e.g., whether they are adjacent, relative positions, etc.), whether they work in pairs, and whether they have a shared field of view can be determined based on the target image set, thereby obtaining the association relationship.
[0857] Fourth step: If a target object is detected in the security area monitored by the set of security devices, one or more linked security devices are determined from the set of security devices based on the association relationship and the target information of the target object, so as to obtain the security device for performing the first security response operation.
[0858] The security area can be the security zone of the security devices in the aforementioned set of security equipment. For example, the security area can be the shooting range of the cameras included in the aforementioned set of security equipment.
[0859] The target object can be an object located within the aforementioned security area. As an example, the target object can be an image of an object in motion within the aforementioned security area, either historically or currently. Furthermore, the target object can include images of at least one of the following: people, vehicles, animals, etc.
[0860] Target information may include at least one of the following: pose information, focal length information, shape feature information, and movement speed of the target object.
[0861] As an example, if the target information of the target object indicates that the moving speed of the target object is less than or equal to a preset speed value, then the security area of a single security device A where the target object is located can be determined first. Then, from the set of security devices, the security device that has a pairing relationship with security device A (which is a kind of association relationship, such as a telephoto camera and a wide-angle camera can be paired) can be determined, and the security device can be used as a linked security device.
[0862] As another example, if the target information of the target object indicates that the moving speed of the target object is greater than the above-mentioned preset speed value, then the security area of the single security device A where the target object is located can be determined first. Then, from the set of security devices, the security device that has a shared field of view with security device A (which is a kind of association relationship, such as two adjacent cameras can share the field of view) is determined, and the security device is used as the linked security device.
[0863] As another example, it can be determined whether the objects indicated by the two images are the same object based on pose information and / or shape feature information. If they are the same object, it can be determined that the two security devices that captured the two images have a shared field of view (which is a type of association, such as two adjacent cameras sharing a field of view). If the target object is located in the security area monitored by one of the security devices, the other security device can be used as a linked security device.
[0864] As another example, it can be determined whether the objects indicated by the two images are the same object based on pose information and / or shape feature information. If they are the same object, and the focal length information of the two images is different, it can be determined that the two security devices that acquired the two images have a pairing relationship (a type of association, such as a telephoto camera and a wide-angle camera can work together). If the target object is located in the security area monitored by one of the security devices, the other security device can be used as a linked security device.
[0865] It is understandable that, when the same security device A has a primary relationship with security device B (i.e., sharing a field of view) and a secondary relationship with security device C (i.e., a paired working relationship), the system can determine whether security device B or security device C is the linked security device based on the movement speed of objects in the security area monitored by the security device.
[0866] Based on this, the security equipment can be controlled to perform the first security response operation in the following way: control one or more of the linked security equipment to perform monitoring operations on the target object in order to perform the first security response operation.
[0867] Each security device in the aforementioned security equipment set can be associated with one or more monitoring operations. Based on this association, the monitoring operation to be executed by the linked security device can then be determined.
[0868] Monitoring operations may include at least one of the following: image acquisition operations and alert operations. For example, if the security device is a camera, then the security device can perform image acquisition operations. If the security device is an alarm, then the security device can perform alert operations.
[0869] Here, each security device in the aforementioned set of security devices can be associated with one or more monitoring operations. Based on this association, the monitoring operation to be executed by the linked security device can then be determined.
[0870] Furthermore, linkage rules can be predetermined, and based on these rules, the monitoring operations to be performed by the linked security devices can be determined. For example, these linkage rules can be determined based on user needs and at least one of the aforementioned relationships. The linkage rules are used to instruct the linked security devices to perform monitoring operations. For example, when the front yard camera detects a delivery person at the front yard gate, the front door camera (the linked security device) can be activated to detect the package. In this context, all operations performed by the devices that take action throughout the event can be considered monitoring operations. For instance, the front door camera detects a face and identifies the delivery person; the yard camera records footage after the delivery person enters the yard; and the doorbell camera detects the package and broadcasts a signal at the door. All of these can be considered monitoring operations.
[0871] Furthermore, after identifying the linked security device and the monitoring operation to be performed by the linked security device, the linked security device can be controlled to perform the monitoring operation.
[0872] It is understood that, among the above optional implementation methods, the association relationship of the security devices in the security device set can be automatically determined based on the target image set corresponding to the security device set, and the linked security devices can be controlled to perform monitoring operations. In this way, the automation level of users using security devices for monitoring can be improved, and the complexity of users using security devices for monitoring can be reduced.
[0873] In some of the above application scenarios, the security device includes a camera.
[0874] Based on this, the target image set corresponding to the security device set can be obtained in the following way: obtain the images collected by the security devices in the security device set to obtain the target image set, wherein the security devices and the target images correspond one-to-one.
[0875] A security device set may include at least two security devices. Security devices may include at least one of the following: a camera, an alarm, a sensor, etc.
[0876] For example, a single security device in a security equipment set can be a camera, an alarm, or a sensor. As another example, a single security device in a security equipment set can include both a camera and an alarm.
[0877] The target image set corresponding to the security equipment set can be a collection of images obtained by capturing images of the security equipment in the security equipment set. For example, the target image set corresponding to the security equipment set can be obtained by capturing images of each security equipment in the security equipment set through one or more cameras.
[0878] Alternatively, the target image set corresponding to the security device set can also be a set of images captured by cameras within the security device set. For example, each security device in the security device set can include a camera. Thus, by capturing images of the area within the shooting range of the cameras included in the security device set, the target image set corresponding to the security device set can be obtained.
[0879] Specifically, if the set of security devices includes security device A, security device B, and security device C, and security device A has acquired image A, security device B has acquired image B, and security device C has acquired image C, then the target image set may include image A, image B, and image C.
[0880] In this embodiment, the correspondence between security equipment and target images is as follows: the target image is acquired by security equipment that has a corresponding relationship with the target image.
[0881] Furthermore, the association between the security devices can be established based on the calibrated object in the following manner:
[0882] The first step is to determine the acquisition time of the target image in the target image set.
[0883] When acquiring target images, the acquisition time can be recorded.
[0884] The second step is to establish the association relationship between the security devices in the security device set based on the determined collection time and the calibration object.
[0885] After obtaining the acquisition time and the calibration object, the association relationship of the security devices in the security device set can be determined based on the determined acquisition time and the calibration object.
[0886] For example, if a target object appears simultaneously in the target images of both camera A and camera B, and at another moment, the target object appears simultaneously in the target images of both camera A and camera C, it means that at the first location where the target object appears, camera A and camera B are cameras with a shared field of view (i.e., the aforementioned relationship), and at the second location where the target object appears, camera A and camera C are cameras with a shared field of view (i.e., the aforementioned relationship).
[0887] Based on this relationship, comprehensive tracking can be achieved. For example, when a target is about to leave the field of view of camera A and it is determined from its target information that it is about to enter the monitoring area of camera B, camera B takes over recording, while camera C remains in sleep mode to save power; conversely, when the target is about to leave the field of view of camera A and it is determined that it is about to enter the monitoring area of camera C, camera C is awakened to record, and camera B goes into sleep mode.
[0888] For example, if camera A and camera B simultaneously capture images of a vehicle (the target object mentioned above), camera A can capture a clear image of the license plate, while camera B can only capture the vehicle's exterior. In this case, cameras A and B are paired telephoto and wide-angle cameras, with camera A using a telephoto lens and camera B using a wide-angle lens. They can work together to track the vehicle, with the telephoto lens recognizing the license plate and the wide-angle lens recording the vehicle's exterior features.
[0889] It is understandable that, in the above situation, the correlation between the security devices in the security device set is determined by the determined collection time and calibration object, which can improve the accuracy of determining the correlation between security devices.
[0890] In some of the examples above, the association between the security devices in the security device set can be determined based on the determined acquisition time and the calibration object in the following manner:
[0891] First, based on the determined acquisition time and the calibration object, it is determined whether the target image set contains a second target image subset.
[0892] The second subset of target images includes telephoto and wide-angle images of the same subject.
[0893] Subsequently, if the target image set contains the second target image subset, the association relationship of the security devices in the second security device subset is determined to represent the second relationship.
[0894] In this context, the security devices in the second security device subset acquire the telephoto image or the wide-angle image, and the second relationship represents a pair of telephoto cameras and a wide-angle camera working together.
[0895] As an example, if camera A and camera B simultaneously capture images of a vehicle target (i.e., the aforementioned target object), camera A can capture a clear image of the license plate, while camera B can only capture the vehicle's exterior. In this case, cameras A and B are paired telephoto and wide-angle cameras, with camera A using a telephoto lens and camera B using a wide-angle lens. They can work together to track the vehicle, with the telephoto lens recognizing the license plate and the wide-angle lens recording the vehicle's exterior features.
[0896] It is understood that, in the above example, the security equipment (camera) can be automatically determined to be paired and working based on the determined acquisition time and the calibration object.
[0897] In some of the disclosures in the above examples, one or more linked security devices are determined from the set of security devices based on the association and the target information of the target object:
[0898] If the association represents the first relationship, and the target object is located within the security area monitored by the first security device subset, then from the first security device subset, the security device corresponding to the security area to be monitored by the target object is determined to be a linked security device.
[0899] Among them, the linked security equipment can be security equipment used to monitor the security area corresponding to the target object to be entered.
[0900] Based on this, one or more of the linked security devices can be controlled to perform monitoring operations on the target object in the following manner: after waking up the linked security device, control the linked security device to acquire images of the target object in order to perform the monitoring operation.
[0901] It is understood that, in the above disclosure, when the first subset of security devices includes security device A and security device B, and security device A and security device B have the aforementioned first relationship (i.e., shared field of view), if the current target object is located within the security area of security device A, then security device B can be in a dormant state. As the target object moves, security device B (i.e., the aforementioned linked security device) can be awakened, thereby controlling security device B to acquire images of the target object to execute the monitoring operation. This allows for more intelligent security monitoring.
[0902] In some of the examples above, the association between the security devices in the security device set can be determined based on the determined acquisition time and the calibration object in the following manner:
[0903] First, based on the determined acquisition time and the calibration object, it is determined whether the target image set contains a first target image subset.
[0904] Wherein, the acquisition time of each target image in the first target image subset is the same, and at least a portion of the markers in each target image in the first target image subset indicate the same subject.
[0905] Subsequently, if the target image set contains the first target image subset, the association relationship of the security devices in the first security device subset is determined to represent the first relationship.
[0906] Wherein, the security devices in the first security device subset acquire the target images in the first target image subset, and the first relationship represents a shared field of view.
[0907] As an example, if the target object appears in the target images of both camera A and camera B, it means that camera A and camera B are cameras with a shared field of view (i.e., the aforementioned relationship).
[0908] It is understood that, in the above example, it is possible to automatically determine whether the security device (camera) shares the field of view based on the determined acquisition time and the calibration object.
[0909] In some of the disclosures in the above examples, one or more linked security devices can be determined from the security device set based on the association relationship and the target information of the target object:
[0910] When the association relationship represents the second relationship, the security devices in the second security device subset are respectively identified as linked security devices.
[0911] Based on this, one or more of the linked security devices can be controlled to perform monitoring operations on the target object in the following manner:
[0912] The first step is to control the telephoto camera to capture a telephoto image of the target object.
[0913] The second step is to control the wide-angle camera to capture wide-angle images of the target object.
[0914] It is understood that, in the above disclosure, when the second subset of security devices includes security device C and security device D, and security device C and security device D have the aforementioned second relationship (i.e., a telephoto camera and a wide-angle camera used in pairs), if the current target object is located within the security area of security device C, then security device D can be identified as a linked security device, and security device D can be controlled to acquire images of the target object to perform the monitoring operation. Thus, both the overall image of the target object and its detailed features can be obtained, enabling more intelligent security monitoring and improving the security of the installation.
[0915] In some application scenarios of the above optional implementation methods, the target information includes at least the following information: the object type of the target object, and the security device that detected the target object.
[0916] Based on this, one or more of the linked security devices can be controlled to perform monitoring operations on the target object in the following manner:
[0917] The first step is to identify the events triggered by the target object based on the target information.
[0918] The events triggered by the target object may include at least one of the following: indoor intrusion, visits from relatives and friends, etc.
[0919] The second step is to control the linked security equipment to perform the monitoring operation for monitoring the event.
[0920] Here, each event can be pre-associated with one or more monitoring operations. The monitoring operations pre-associated with an event are the monitoring operations used to monitor that event. Therefore, the linked security equipment can be controlled to execute the monitoring operations used to monitor the event.
[0921] It is understandable that in the above application scenarios, the linked security equipment can be controlled to perform corresponding monitoring operations based on the events triggered by the target object. In this way, different monitoring operations can be performed based on different events, thereby improving the targeting and security of security.
[0922] In some application scenarios of the above-mentioned optional implementation methods, one or more linked security devices can be determined from the security device set based on the association relationship and the target information of the target object in the following manner:
[0923] The first step is to determine whether the target object is located in the security area monitored by the target camera.
[0924] The target camera can be any camera in the aforementioned set of security devices. The target camera is any one of the security devices in the set.
[0925] The second step is to determine the linked security devices associated with the target camera from the security device set based on the association relationship when the target object is located in the security area monitored by the target camera.
[0926] Here, since the association relationship can indicate the security devices in the security device set associated with the target camera, the security device associated with the target camera can be determined from the security device set based on the association relationship, and that security device can be designated as a linked security device.
[0927] Furthermore, one or more of the linked security devices can be controlled to perform monitoring operations on the target object in the following manner: if the linked security device includes an audio output device, the audio output device is controlled to output a prompt audio corresponding to the target object in order to perform the monitoring operation.
[0928] The aforementioned audio prompt can be a preset audio corresponding to the target object. For example, the audio prompt could be the audio stating "You have entered the security area." Furthermore, the audio prompt can also be determined based on the object type captured by the target camera. For example, if the image captured by the target camera includes an image of a courier (i.e., an object type), then the audio prompt could be the audio stating "Please place the package at location XX."
[0929] It is understood that in the above application scenarios, when the linked security equipment includes an audio output device (such as a walkie-talkie), the corresponding prompt audio for the target object can be automatically output via the audio output device to execute the monitoring operation. This enables more intelligent monitoring.
[0930] In some application scenarios of the above-mentioned optional implementation methods, when the monitoring operation is used for image acquisition, the following steps can also be performed:
[0931] The first step is to acquire the images obtained by the linked security equipment performing the monitoring operation, thus obtaining the monitoring images.
[0932] The surveillance video can be an image obtained by the security equipment performing the surveillance operation.
[0933] The second step is to extract at least one of the following feature information of the target object from the surveillance image: face image, license plate number, movement trajectory, first appearance time, first appearance location, departure time, and departure location.
[0934] The third step is to send the feature information to a preset user terminal so that the user terminal displays event information determined based on the feature information; or, to determine event information based on the feature information and send the event information to a preset user terminal so that the user terminal displays the event information; wherein the event information represents an event triggered by the target object.
[0935] The event information may include at least one of the following: indoor intrusion, visit by relatives or friends, etc.
[0936] It is understandable that in the above application scenarios, the aforementioned feature information can be automatically extracted from surveillance images for secondary processing, thereby improving the efficiency for users to obtain information about security status.
[0937] In some optional implementations of this embodiment, the security device performing the first security response operation includes a first security device and a second security device.
[0938] Security equipment can be a collective term for various devices and systems that ensure the safety of people, property, and the environment inside and around a building. For example, security equipment may include, but is not limited to: surveillance cameras, intrusion detectors, and alarm systems.
[0939] The first security device and the second security device can be two different security devices installed on the building. As an example, the first security device and the second security device can be two pan-tilt cameras installed on the building.
[0940] Based on this, before controlling the security device to execute the first security response operation, a three-dimensional model of the building can also be obtained, wherein the first security device and the second security device are installed on the building.
[0941] Buildings can be places where people live, work, study, entertain, and engage in various activities. As examples, buildings can include, but are not limited to: residential buildings, such as apartments and villas; commercial buildings, such as shopping malls and office buildings; public buildings, such as libraries, museums, and stadiums; and industrial buildings, such as factories and warehouses.
[0942] The 3D model can be a digital 3D representation of the aforementioned building. The acquired 3D model may consist only of the building's 3D model, or it may include the building's 3D model, the first security device's 3D model, and the second security device's 3D model.
[0943] As an example, a 3D model can be constructed in the following way:
[0944] First, guide the user to take photos (e.g., videos) around the aforementioned building (e.g., a house), as well as the installation points of the first and second security devices, to capture images of the building's interior and surroundings from all angles.
[0945] Subsequently, 3D reconstruction technologies such as SLAM (Simultaneous Localization and Mapping) were used to construct 3D models of the buildings.
[0946] Specifically, the process begins with feature extraction: features are extracted from the captured image data, such as corner points and edges. Next, data association and matching are performed: the features of the data to be associated and matched are compared with features in the previously constructed 3D model to determine the relative position and orientation of the image acquisition device in space. Then, pose estimation is performed: based on the results of data association and matching, the position and orientation (including position coordinates and rotation angle) of the image acquisition device at each moment are estimated. Finally, the 3D model is constructed: newly acquired data is integrated into the existing 3D model, gradually refining and updating the model.
[0947] Furthermore, the security device can be controlled to perform the first security response operation in the following manner:
[0948] The first step is to determine the first mapping pose of the first security device in the three-dimensional model and the second mapping pose of the second security device in the three-dimensional model, wherein the first mapping pose represents the pose of the first security device mapped to the three-dimensional model and the second mapping pose represents the pose of the second security device mapped to the three-dimensional model.
[0949] The first mapping pose of the first security device in the three-dimensional model can be determined by the relative pose between the first security device and the building; the second mapping pose of the second security device in the three-dimensional model can be determined by the relative pose between the second security device and the building.
[0950] Furthermore, when the acquired 3D models include 3D models of buildings, 3D models of the first security device, and 3D models of the second security device, a coordinate system containing the acquired 3D models can be constructed first. Then, the pose of the 3D model of the first security device in the aforementioned coordinate system is determined to obtain the first mapped position; the pose of the 3D model of the second security device in the aforementioned coordinate system is determined to obtain the second mapped position.
[0951] Alternatively, other methods can be used to perform the first step described above, which will not be elaborated here.
[0952] The second step is to determine whether there is an intersection between the first monitoring area of the first security device and the second monitoring area of the second security device, based on the first mapping pose and the second mapping pose, so as to obtain the determination result.
[0953] The first monitoring area can represent the monitoring area obtained after mapping the actual monitoring area of the first security device to the three-dimensional model; or it can represent the actual monitoring area of the first security device.
[0954] The second monitoring area can represent the monitoring area obtained after mapping the actual monitoring area of the second security device onto the three-dimensional model; or it can represent the actual monitoring area of the second security device.
[0955] In the case where the first monitoring area can represent the monitoring area obtained after mapping the actual monitoring area of the first security device to the three-dimensional model, the second monitoring area represents the monitoring area obtained after mapping the actual monitoring area of the second security device to the three-dimensional model; in the case where the first monitoring area represents the actual monitoring area of the first security device, the second monitoring area represents the actual monitoring area of the second security device.
[0956] Here, various methods can be used to determine whether there is an overlap between the first monitoring area of the first security device and the second monitoring area of the second security device, based on the first mapping pose and the second mapping pose.
[0957] As an example, the first and second mapped poses can be input into a pre-trained decision model to determine whether there is an intersection between the first and second monitoring regions.
[0958] The aforementioned determination model can represent the correspondence between the first mapped pose, the second mapped pose, and the discrimination information. The discrimination information can indicate whether there is an intersection between the first monitoring area and the second monitoring area.
[0959] The aforementioned discrimination model can be a convolutional neural network or other model trained using training samples that include a first mapping pose, a second mapping pose, and discrimination information. Alternatively, it can be a formula or table representing the correspondence between the first mapping pose, the second mapping pose, and the discrimination information.
[0960] In addition, other methods can be used to determine whether there is an overlap between the first monitoring area of the first security device and the second monitoring area of the second security device, based on the first and second mapped poses. Please refer to the description below for details, which will not be elaborated here.
[0961] The third step is to determine, based on the determination results, whether the first security device and the second security device are linked devices.
[0962] Linked devices can refer to devices that can cooperate with each other to perform monitoring operations.
[0963] Here, various methods can be used to determine whether the first security device and the second security device are linked devices based on the first mapping pose and the second mapping pose.
[0964] In some application scenarios of the above-mentioned optional implementations, the first mapped pose includes a first position and a first orientation of the first security device mapped onto the 3D model. The second mapped pose includes a second position and a second orientation of the second security device mapped onto the 3D model.
[0965] The first position can represent the location of the first security device mapped onto the three-dimensional model. For example, the first position can be represented by coordinates.
[0966] The first attitude can represent the attitude of the first security device mapped onto the 3D model. For example, the first position may include, but is not limited to, the pitch angle, yaw angle, etc. of the first security device mapped onto the 3D model.
[0967] The second position can represent the location of the second security device mapped onto the three-dimensional model. For example, the second position can be represented using coordinates.
[0968] The second attitude can represent the attitude of the second security device mapped onto the 3D model. For example, the second position may include, but is not limited to, the pitch angle, yaw angle, etc., of the second security device mapped onto the 3D model.
[0969] Based on this, the method can be adopted to determine whether the first security device and the second security device are linked devices based on the first mapped pose and the second mapped pose:
[0970] The first step is to determine the first monitoring area in the three-dimensional model to which the first security device is mapped, based on the first position, the first posture, and the first monitoring parameters of the first security device.
[0971] The first monitoring parameter may include, but is not limit...
Claims
1. A security method, characterized in that, The method includes: Acquire the first security video and pre-determined security preference information; The first security video is identified to obtain the video identification result of the first security video; Based on the video recognition results and the security preference information, a security response operation matching the first security video is determined to obtain the first security response operation; Identify the security device used to perform the first security response operation; Control the security device to execute the first security response operation.
2. The method according to claim 1, characterized in that, The video recognition result is the event recognition result of the first security video, and the event recognition result represents the event characterized by the first security video; as well as The step of determining the security response operation matching the first security video based on the video recognition result and the security preference information includes: Based on the event identification results, the urgency level of the event represented by the first security video is determined; Based on the urgency level and the security preference information, a security response operation matching the first security video is determined.
3. The method according to claim 1, characterized in that, After identifying the first security video to obtain the video identification result of the first security video, the method further includes: Obtain first feedback information regarding the video recognition result; wherein, the first feedback information indicates an adjustment to the recognition strategy of the second security video; the second security video is: the first security video, or a security video obtained after the first security video; the recognition strategy includes at least one of the following: recognition efficiency, recognition method; Determine the identification strategy to be adjusted as indicated by the first feedback information; The second security video is identified according to the identification strategy adjusted according to the first feedback information, so as to obtain the video identification result of the second security video; Based on the video recognition results and security preference information of the second security video, a security response operation matching the second security video is determined to obtain the second security response operation; Execute the second security response operation.
4. The method according to claim 1, characterized in that, After performing the first security response operation, the method further includes: Obtain third feedback information for the first security response operation, wherein the third feedback information indicates an adjustment to the execution strategy of the security response operation, and the execution strategy includes at least one of the following: execution efficiency and execution method; The third feedback information indicates the adjusted execution strategy; The fourth security response operation is executed according to the execution strategy adjusted according to the third feedback information, wherein the fourth security response operation is a security response operation executed after the first security response operation.
5. The method according to claim 1, characterized in that, The step of determining the security response operation matching the first security video based on the video recognition result and the security preference information includes: Based on the video recognition results, the response probability of the first security video is determined; Determine whether the response probability is greater than or equal to a preset threshold; If the response probability is greater than or equal to the preset threshold, a security response operation matching the first security video is determined based on the security preference information.
6. The method according to claim 5, characterized in that, After the first security response operation is performed, the preset threshold is adjusted in the following manner: Obtain fourth feedback information in response to the first security response operation, wherein the fourth feedback information indicates adjustment of the preset threshold; Adjust the preset threshold according to the adjustment method indicated by the fourth feedback information to obtain the adjusted preset threshold; and The method further includes: If the response probability is greater than or equal to the adjusted preset threshold, a security response operation matching the second security video is determined based on security preference information to obtain a fifth security response operation, wherein the second security video is: the first security video, or a security video obtained after the first security video; Perform the fifth security response operation.
7. The method according to any one of claims 1-6, characterized in that, The first security video includes a target image, which includes an object image and a person image, wherein the object image represents a target object and the person image represents a target person; as well as The step of identifying the first security video to obtain the video identification result of the first security video includes: The first security video is identified to determine a first detection box and a second detection box in the target image, wherein the first detection box is the detection box of the object image and the second detection box is the detection box of the person image; Determine the degree of overlap between the first detection box and the second detection box; Based on the degree of overlap, theft detection information is generated, wherein theft detection information indicates whether the target person has the intention to steal the target item; Based on the theft detection information, the video recognition result of the first security video is determined.
8. The method according to claim 7, characterized in that, The process of generating theft detection information based on the degree of overlap includes: Determine whether the degree of overlap is greater than or equal to a preset threshold; If the degree of overlap is greater than or equal to the preset threshold, determine whether the behavior of the target person represented by the personnel image is theft, so as to obtain a first determination result; Based on the first determination result, theft detection information is generated.
9. The method according to any one of claims 1-6, characterized in that, The first security video is a frame-sampling result of the target video; and The frame extraction results are generated in the following manner: Obtain an image description data set and a target video, wherein the image description data in the image description data set is used to describe the content of the target image, and the target video consists of an image sequence; Calculate the similarity between the images in the image sequence and each image description data in the image description data set to obtain the target similarity corresponding to the image; From the calculated target similarities, select a first number of target similarities; A first image set corresponding to the first number of target similarities is determined, wherein the images in the first image set correspond one-to-one with the target similarities in the first number of target similarities; Based on the first image set, the frame extraction result of the target video is determined.
10. The method according to claim 9, characterized in that, The step of determining the frame extraction result of the target video based on the first image set includes: Display the first image set; Determine whether an adjustment operation is detected for an image in the first image set; wherein the adjustment operation is used to: adjust the image in the first image set to obtain a second image set; Upon detecting the adjustment operation, the second set of images is determined as the frame extraction result of the target video.
11. The method according to claim 10, characterized in that, Upon detecting the adjustment operation, the image description data set is updated as follows: The image features of each image in the second image set are determined to obtain an image feature set, wherein the image features in the image feature set correspond one-to-one with the images in the second image set; Based on the image feature set, determine the image description data; The image description data set is updated based on the determined image description data.
12. The method according to claim 11, characterized in that, The step of updating the image description data set based on the determined image description data includes: Determine the cardinality of the image description data set before the update; Determine whether the base number is less than a preset value; If the base number is less than the preset value, the determined image description data is added to the image description data set to obtain the updated image description data set. If the base number is greater than or equal to the preset value, the image description data included in the image description data set before the update is replaced with the determined image description data to obtain the updated image description data set.
13. The method according to any one of claims 1-6, characterized in that, The security devices that perform the first security response operation include a first security device and a second security device; as well as Before controlling the security device to perform the first security response operation, the method further includes: Obtain a 3D model of the building, wherein the building is equipped with the first security device and the second security device; and The control of the security device to execute the first security response operation includes: A first mapping pose of the first security device in the three-dimensional model and a second mapping pose of the second security device in the three-dimensional model are determined, wherein the first mapping pose represents the pose of the first security device mapped to the three-dimensional model and the second mapping pose represents the pose of the second security device mapped to the three-dimensional model. Based on the first mapped pose and the second mapped pose, determine whether there is an intersection between the first monitoring area of the first security device and the second monitoring area of the second security device, so as to obtain the determination result; Based on the determination result, it is determined whether the first security device and the second security device are linked devices; When the first security device and the second security device are part of the linked devices, determine the motion information of the target object detected by the first security device; Based on the motion information, the second security device is controlled to monitor the target object.
14. The method according to claim 13, characterized in that, The first mapping pose includes the first position and first posture of the first security device mapped to the 3D model, and the second mapping pose includes the second position and second posture of the second security device mapped to the 3D model; as well as The step of determining whether there is an intersection between the first monitoring area of the first security device and the second monitoring area of the second security device based on the first mapped pose and the second mapped pose includes: Based on the first position, the first posture, and the first monitoring parameters of the first security device, the first monitoring area of the first security device is determined to be mapped to the three-dimensional model; Based on the second position, the second posture, and the second monitoring parameters of the second security device, the second monitoring area of the second security device is determined to be mapped to the three-dimensional model; Determine whether there is an overlap between the first monitoring area and the second monitoring area; and The step of determining whether the first security device and the second security device are linked devices based on the determination result includes: If the determination result indicates that the first monitoring area and the second monitoring area have an intersection area, then the first security device and the second security device are determined to be linked devices. If the determination result indicates that there is no intersection between the first monitoring area and the second monitoring area, then the first security device and the second security device are determined not to be part of the linkage device.
15. The method according to claim 14, characterized in that, Determining whether the first monitoring area and the second monitoring area have an overlapping area includes: Determine the first projection of the first monitoring area onto a preset plane; Determine the second projection of the second monitoring area onto the preset plane; Determine whether there is an overlapping area between the first projection and the second projection; In the case where the first projection and the second projection have an overlapping area, it is determined that the first monitoring area and the second monitoring area have an intersection area; If there is no overlapping area between the first projection and the second projection, it is determined that there is no intersection area between the first monitoring area and the second monitoring area. The preset plane is parallel to the ground plane and mapped to the plane in the three-dimensional model.
16. The method according to claim 13, characterized in that, Determining the first mapped pose of the first security device in the 3D model and the second mapped pose of the second security device in the 3D model includes: Display the three-dimensional model; The detection includes a first operation and a second operation for the displayed 3D model, wherein the first operation is used to determine the mapping pose of the first security device in the 3D model, and the second operation is used to determine the mapping pose of the second security device in the 3D model. Upon detecting the first operation, the mapping pose indicated by the first operation is determined as the first mapping pose of the first security device in the three-dimensional model; Upon detecting the second operation, the mapping pose indicated by the second operation is determined as the second mapping pose of the second security device in the three-dimensional model.
17. The method according to claim 16, characterized in that, The three-dimensional model includes: a first initial pose of the first security device and a second initial pose of the second security device; the first operation is used to adjust the first initial pose in the three-dimensional model, and the second operation is used to adjust the second initial pose in the three-dimensional model; and Determining the mapped pose of the first operation instruction as the first mapped pose of the first security device in the three-dimensional model includes: Adjust the first initial pose according to the adjustment method indicated by the first operation instruction; The adjusted first initial pose is determined as the first mapped pose of the first security device in the three-dimensional model; and The step of determining the mapped pose of the second operation instruction as the second mapped pose of the second security device in the three-dimensional model includes: Adjust the second initial pose according to the adjustment method indicated by the second operation instruction; The adjusted second initial pose is determined as the second mapped pose of the second security device in the three-dimensional model.
18. The method according to claim 13, characterized in that, The second security device is a PTZ camera, and the motion information includes the motion trajectory of the security object; as well as The step of controlling the second security device to monitor the target object based on the motion information includes: Based on the motion trajectory included in the motion information, the initial position of the target object in the monitoring area of the second security device is determined; The field of view of the second security device is controlled to move to the initial position so that the second security device can monitor the target object.
19. A security device, characterized in that, The device includes: The first acquisition unit is used to acquire the first security video and the pre-determined security preference information; The first identification unit is used to identify the first security video to obtain the video identification result of the first security video; The first determining unit is used to determine a security response operation that matches the first security video based on the video recognition result and the security preference information, so as to obtain the first security response operation; The second determining unit is used to determine the security device used to perform the first security response operation; The first control unit is used to control the security device to perform the first security response operation.
20. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor is configured to execute a computer program stored in the memory, wherein when the computer program is executed, it implements the security method according to any one of claims 1-18.