Video recording control method, video recording equipment and storage medium

By introducing a video recording control method in a low-power camera, using AI model to analyze and control the recording process, the problems of limited detection distance of PIR sensors and insensitive to static object detection are solved, and the integrity and power consumption of the recording process are achieved.

CN120075558APending Publication Date: 2025-05-30SHENZHEN BAICHUAN SECURITY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510080352.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing low-power cameras trigger wake-up via PIR sensors, resulting in limited detection distance and insensitive to radial motion and rest state detection, which may cause the monitoring target to stop recording in advance before it has been drawn, resulting in the missing monitoring data.

Method used

The video recording control method is adopted to obtain the trigger signal through the preset infrared sensor, and in response to the start of the recording program, obtain the audio and video information of the target environment, and determine the status of the dynamic target based on the preset artificial intelligence model. The recording device is controlled to continue or stop recording to ensure the completeness and accuracy of the recording process.

Benefits of technology

Through the analysis of AI model and image recognition, the start and stop of the recording process are accurately controlled, the probability of the recording process ending prematurely, the loss of monitoring data is avoided, and the power consumption of the recording device is reduced and the life of the device is extended.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075558A_ABST
    Figure CN120075558A_ABST
Patent Text Reader

Abstract

The invention discloses a video recording control method, video recording equipment, computer equipment and a computer readable storage medium. The method comprises the following steps: acquiring a trigger signal generated by a preset infrared sensor; in response to the trigger signal, controlling a preset recording device to start a recording program; controlling a preset recording device to obtain first target media information of the target environment; determining target artificial intelligence model type information according to the first target media information; and controlling the preset recording device to continuously record the second target media information of the target environment. According to the video recording control method, analysis can be carried out based on the preset AI model under the condition that the recording equipment is triggered to start to acquire the environment video and audio information based on infrared sensors such as PIR, so that supplementary verification is carried out on triggering of the PIR for the recording process; therefore, the situation that in the prior art, monitoring data are lost due to the fact that recording is stopped in advance in the scheme that the recording process is triggered and maintained only by means of PIR can be avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of video surveillance, and particularly to a video recording control method, a video recording device, a computer device, and a computer-readable storage medium. Background Art

[0002] A low-power camera refers to a surveillance camera that significantly reduces the device's energy consumption while ensuring the surveillance effect by adopting technical means such as low-power chips and optimized software algorithms; this type of camera is different from the continuous surveillance recording of a power camera. It is in a dormant state most of the time, and the built-in pyroelectric infrared sensor (PIR) continuously monitors the surrounding environment. When a moving object such as a person or an animal is detected, it will be awakened and enter the working state to start recording or perform other operations. This working mode can significantly reduce the energy consumption of the device and extend the battery life.

[0003] In the current related technical solutions, most low-power cameras are triggered and awakened by PIR. The PIR sensor is a device that uses the pyroelectric effect to detect the infrared radiation emitted by the human body or other objects. It usually consists of a pyroelectric material, an optical filter, and a signal processing circuit. When a human body or other heat-emitting object moves within the detection range of the sensor, it will cause a change in the intensity of the infrared radiation received by the sensor, thereby generating an electrical signal output. However, limited by the electronic characteristics of the PIR itself, its detection distance is limited, and its detection of radial movement and static state is also limited, which may lead to the situation where the monitoring target stops the recording process in advance before leaving the frame, resulting in missing monitoring data. Summary of the Invention

[0004] This application provides a video recording control method, a video recording device, a computer device, and a computer-readable storage medium.

[0005] The video recording control method involved in the embodiments of this application includes the following steps:

[0006] Obtain a trigger signal generated by a preset infrared sensor;

[0007] In response to the trigger signal, control a preset recording device to start a recording program;

[0008] Control the preset recording device to obtain first target media information of a target environment;

[0009] Based on the first target media information and a preset artificial intelligence model, determine target artificial intelligence model type information;

[0010] When the target artificial intelligence model type information meets the first preset condition, control the preset recording device to continuously record the second target media information of the target environment until the second target media information meets the second preset condition.

[0011] Wherein the second preset condition is configured to indicate that based on the preset artificial intelligence model, it is determined that the state of the dynamic target in the second target media information in the picture meets the condition for stopping recording.

[0012] In this way, the video recording control method in the embodiments of the present application can, when triggering the recording device to start and obtain environmental video and audio information based on an infrared sensor such as a PIR, perform element analysis based on a pre-set AI model, so as to accurately control and supplement the verification of the triggering and continuation of the PIR for the recording process, thereby reducing the probability of premature termination of the recording process. At the same time, the AI model can also be used to perform image recognition on the obtained video and audio information, determine the conditions of the actions of the dynamic targets in the picture, and then use the determination results to end the recording process, so as to avoid the situation of missing monitoring data caused by premature termination of recording in the existing solutions that only rely on PIR to trigger and maintain the recording process. In addition, the video recording control method proposed in the embodiments of the present application can not only ensure the integrity of the recording process, but also can reduce the power consumption of the recording device as much as possible by accurately controlling the start and stop of the recording process, and extend the service life of the recording device.

[0013] In some embodiments, the controlling the preset recording device to start the recording program in response to the trigger signal includes:

[0014] In response to the trigger signal generated by the preset infrared sensor due to detecting an energy change within the target range, control the preset recording device to enter the system bootloader.

[0015] In some embodiments, the controlling the preset recording device to obtain the first target media information of the target environment includes:

[0016] In response to the completion of the execution of the system bootloader, control the preset recording device to obtain the original media information of the target environment;

[0017] Perform encoding processing on the original media information and store it in a preset cache space to generate the first target media information.

[0018] In some embodiments, the controlling the preset recording device to continuously record the second target media information of the target environment until the target artificial intelligence model type information meets the second preset condition when the target artificial intelligence model type information meets the first preset condition includes;

[0019] When the target artificial intelligence model type information matches the information recognition type preset by the user, cache the feature information corresponding to the dynamic target in the first target media information;

[0020] Control the preset recording device to continuously record the second target media information of the target environment;

[0021] Within the time range when it is determined according to the feature information that all the dynamic targets have left the screen and lasted for a first preset duration, if there are no secondary dynamic targets newly entering the screen, control the preset recording device to stop recording; and / or

[0022] When it is determined according to the feature information that all the dynamic targets remain stationary and last for a second duration, control the preset recording device to stop recording.

[0023] In some embodiments, the step of, within the time range when it is determined according to the feature information that all the dynamic targets have left the screen and lasted for a first preset duration, if there are no secondary dynamic targets newly entering the screen, controlling the preset recording device to stop recording includes:

[0024] Obtain the current feature information of the dynamic target;

[0025] According to the cached feature information and the current feature information, determine the similarity between the dynamic targets in the current moment's screen and the dynamic targets in the first target media information;

[0026] When the similarity meets the second preset condition, determine that the dynamic target has left the screen;

[0027] Within the time range when all the dynamic targets have left the screen and lasted for a first preset duration, if there are no secondary dynamic targets newly entering the screen, control the preset recording device to stop recording.

[0028] In some embodiments, the step of, when it is determined according to the feature information that all the dynamic targets remain stationary and last for a second duration, controlling the preset recording device to stop recording includes:

[0029] Obtain the current feature information of the dynamic target;

[0030] According to the cached feature information and the current feature information, determine the position offset of the dynamic target;

[0031] When the position offset meets the third preset condition, determine that the dynamic target remains stationary;

[0032] In the case where the dynamic objects all remain stationary and continue for a second duration, control the preset recording device to stop recording.

[0033] In some embodiments, in the case where the target artificial intelligence model type information meets the first preset condition, control the preset recording device to continuously record the second target media information of the target environment until the target artificial intelligence model type information meets the second preset condition, including;

[0034] In the case where the target artificial intelligence model type information does not match the information recognition type preset by the user, control the preset recording device to enter the sleep state.

[0035] In some embodiments, the method further includes:

[0036] In the time range where all the dynamic objects leave the screen and continue for a first preset duration, if there are secondary dynamic objects newly entering the screen, re - perform numbering processing on the dynamic objects in the first target media information; and / or

[0037] In the second duration range after the dynamic objects all become stationary, if there are dynamic objects that start to move among the dynamic objects, re - perform numbering processing on the dynamic objects in the first target media information.

[0038] In some embodiments, in the case where the target artificial intelligence model type information meets the information recognition type preset by the user, when performing numbering processing on the dynamic objects in the first target media information, it further includes:

[0039] In the case where the target artificial intelligence model type information matches the information recognition type preset by the user, filter out the dynamic objects that remain stationary within the preset image frame range according to the first target information;

[0040] Cache the feature information corresponding to the remaining dynamic objects.

[0041] The video recording device in the embodiments of the present application includes the above - mentioned preset infrared sensor and preset recording device.

[0042] The computer device in the embodiments of the present application includes a memory and a processor. When the computer program stored in the memory is executed by the processor, the above - mentioned video recording control method is implemented.

[0043] The computer - readable storage medium in the embodiments of the present application stores a computer program. When the computer program is executed by one or more processors, the above - mentioned video recording control method is implemented.

[0044] Additional aspects and advantages of embodiments of the present application will be given in part in the following description, become apparent in part from the following description, or be learned by practice of the embodiments of the present application. Description of the Drawings

[0045] The above-mentioned and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of embodiments in conjunction with the drawings, where:

[0046] Figure 1 is a schematic diagram of event timing for the process of triggering recording by a low-power camera in the current related art;

[0047] Figure 2 is a schematic diagram of event timing for the process of triggering recording by a low-power camera in the current related art;

[0048] Figure 3 is a schematic flowchart of a video recording control method in an embodiment of the present application;

[0049] Figure 4 is a schematic execution flowchart of a video recording control method in an embodiment of the present application;

[0050] Figure 5 is a schematic diagram of event timing for a video recording control method in an embodiment of the present application. Detailed Embodiments

[0051] The following details the embodiments of the present application. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals always denote the same or similar elements or elements with the same or similar functions. The embodiments described below by referring to the drawings are exemplary and are only used to explain the embodiments of the present application and should not be construed as a limitation to the embodiments of the present application.

[0052] In the current related art, for a solution of performing video recording and alarm for a target environment, in order to ensure the monitoring effect while reducing the energy consumption of video recording devices such as cameras as much as possible, the generally used video recording device is a low-power camera. A so-called low-power camera refers to a monitoring camera that significantly reduces the energy consumption of the device while ensuring the monitoring effect by adopting technical means such as using low-power chips and optimizing software algorithms; this type of camera is different from the continuous monitoring and recording of a power camera. It is in a dormant state most of the time, and the built-in pyroelectric infrared sensor (i.e., PIR sensor) continuously monitors the surrounding environment. When a moving object such as a person or an animal is detected, it will be awakened and enter the working state to start recording or perform other operations. This working mode can significantly reduce the energy consumption of the device and extend the battery life;

[0053] In current related solutions, most low-power cameras are triggered to wake up the device by PIR, and further determine whether the recording process needs to continue based on the detection signal of PIR. The PIR sensor is a device that uses the pyroelectric effect to detect infrared radiation emitted by the human body or other objects. It usually consists of a pyroelectric material, an optical filter, and a signal processing circuit. When a human body or other heat-emitting object moves within the detection range of the sensor, it will cause a change in the intensity of the infrared radiation received by the sensor, thereby generating an electrical signal output. The normal body temperature of a human is between 36-37°C, and the released infrared radiation is approximately between 9-10μm. The detection wavelength range of PIR is generally in the range of 5-14μm;

[0054] That is, within the detection range of PIR, an energy change with a wavelength of 5-14μm will cause a current change in the PIR sensor, thereby giving a trigger signal. It should be noted here that PIR detects energy changes, not infrared radiation of a fixed wavelength. If an object of a fixed wavelength remains stationary within the detection range, PIR cannot detect the energy change, and thus cannot generate a current change or give a trigger signal.

[0055] The PIR selection of conventional low-power cameras is basically an adjustable-sensitivity digital PIR, which adopts the trigger output high-level mode. That is, when the pyroelectric infrared signal received by the built-in PIR probe exceeds the pre-configured sensitivity threshold, a counting pulse will be generated inside, and at the same time, its REL (Relay out) pin will output a continuous high-level signal to update the PIR status to triggered;

[0056] The electrical characteristics of different PIR sensors will be different. When the signal changes sign (or does not need to change sign) and exceeds the set threshold again, it will also be used as a second pulse, outputting the conditions for triggering or alarming. It can be seen from this that the number of pulses and the duration of the REL high level are configurable;

[0057] For example, please refer to Figure 1 , set the single-trigger condition to 1 pulse number, and the REL duration (i.e., the O-C stage) is set to 4s; for the single-trigger scenario, before the O stage, the PIR sensor of the low-power camera will be in the powered-on state, and other modules in the camera will be in the powered-off state;

[0058] Triggered by PIR, during the O-A stage (about 0.5s), the camera quickly starts in response to the PIR trigger and enters the RTOS system bootloader, wakes up the encoding chip from the sleep state, enables audio collection and image capture, and stores the collected encoded data in the buffer; from stage A, the built-in Real Time Operating System (hereinafter referred to as the RTOS system) of the recording device starts up. During the A-B stage (about 0.5s - 1s), the system starts to parse AI type information in real time. The system will determine whether to start recording after judging according to the AI type selected by the user and the triggered AI information. If it is within the range, recording will start normally from node B; generally speaking, considering power saving, the single-trigger alarm recording duration of the low-power camera is often 8 - 15s, and then the recording ends and it enters the sleep state again. For the convenience of subsequent narration, the length of a single trigger segment is exemplarily set to 8s, that is, after a single trigger, if it is judged that the above recording start conditions are met, a single event recording segment with a duration of 8s from stage B to D is stored. Considering the precise control requirements for starting and stopping the recording process during the above recording process, the above RTOS system is exemplarily selected as a hard real-time system so that the camera can complete the above actions within a preset time range.

[0059] In addition to the above single-trigger scenario, multiple-trigger and continuous-trigger scenarios may occur during the operation of the low-power camera. That is, after the first target initially triggers the PIR, if there are large movements after entering the detection range or the second target enters the detection range during the single-trigger recording period, a continuous trigger phenomenon will occur; please refer to Figure 2 , at node O, the PIR is initially triggered to generate pulse count ①, and the REL duration is set to 4s, that is, the O-C stage. During this period, if it is triggered again, pulse count ② should be generated, and a new REL 4s window is reset starting from node E corresponding to count ②, and the PIR status is updated again.

[0060] In the actual design of the low-power camera, for various considerations such as reducing power consumption, avoiding false triggers and repeated alarms, and optimizing signal processing and algorithm design, in the case of continuous PIR triggers, there will actually be a REL window between two consecutive counting pulses. Within 4s after the front PIR trigger, a high-level signal is continuously output from the REL pin to maintain the effective trigger signal. That is, within the REL duration (O-C stage), the PIR will not update its status again.

[0061] At the end of the REL duration set by the pulse count ①, that is, at node C, the PIR returns to the normal power-on waiting-to-be-triggered state. At this time, it will re-determine whether there is a new PIR signal. From this, it can be known that after a single trigger, if there is a new PIR signal pulse count during the C-D stage, a new round of REL output will be triggered, and on the original 8s recording segment of the single trigger, it will start timing 8s backward for recording again, so as to achieve the purpose of continuous alarm recording. And because the camera is always in the working state throughout the C-D stage, it can continue recording without going through the additional power-on wake-up stage from O to B. If during the C-D stage, that is, after the end of the first round of REL duration, the PIR is not triggered a second time to wake up a new round of REL, the recording will stop at stage D, and the B-D segment will be recorded, with each segment lasting 8s.

[0062] In such a continuous PIR trigger scenario, the above scheme has the following problems: First, affected by the electronic characteristics of the PIR itself, the detection distance of the PIR sensor is limited. The detection distance of most PIR sensors is generally between several meters and a dozen meters. For scenarios such as large parking lots and squares that require long-distance detection, the PIR sensor cannot meet the detection requirements. Second, based on the detection principle of the PIR for energy changes, the PIR sensor is very sensitive to objects moving laterally, but very insensitive to objects moving radially or stationary objects within the detection range. When the object has almost no position change laterally, the PIR sensor cannot accurately generate a trigger signal, which may lead to the situation where the target has not left the frame but the recording process ends prematurely. Third, based on the above continuous trigger mechanism, after a single target triggers the PIR, the camera wakes up normally and starts recording. After maintaining a REL window, the PRI will update the status again (that is, within the 6s after the end of the REL window during a single trigger, which is the C-D stage in Figure 2 ). And during this period, as long as the PIR is triggered twice / multiple times, assuming that the recording lasts for t0 (2s < t0 < 8s) after a single trigger, a segment of t0 + 8s will be recorded on this basis, thereby extending the alarm recording duration. That is, under the above mechanism, if you want to extend the recording segment, you must detect a heat source change again and trigger the PIR during the C-D stage to extend the original 8s recording of the single trigger. Such conditions are relatively harsh for the actual monitoring situation. In some cases, the actual point of heat source change is not within the above time range, which may lead to the recording process ending prematurely due to the inability to extend, resulting in information omission.

[0063] Based on the above possible problems, please refer to Figure 3 , the embodiments of the present application provide a video recording control method, and the above method includes:

[0064] S301: Obtain a trigger signal generated by a preset infrared sensor;

[0065] S302: In response to the trigger signal, control the preset recording device to start the recording program;

[0066] S303: Control the preset recording device to obtain the first target media information of the target environment;

[0067] S304: Based on the first target media information and the preset artificial intelligence model, determine the target artificial intelligence model type information;

[0068] S305: When the target artificial intelligence model type information meets the first preset condition, control the preset recording device to continuously record the second target media information of the target environment until the second target media information meets the second preset condition,

[0069] wherein the second preset condition is configured to indicate that based on the preset artificial intelligence model, it is determined that the state of the dynamic target in the second target media information in the picture meets the condition for stopping recording.

[0070] Specifically, in the above-mentioned embodiment, the preset infrared sensor and the preset recording device are both components of video recording devices such as low-power cameras. The first target media information can be, with reference to the above example, the audio-visual data collected when the recording device in the low-power camera performs audio-visual acquisition after the RTOS system bootloader is executed in response to the trigger signal generated by the PIR. The second target media information can be, with reference to the above example, the audio-visual data recorded after the recording is triggered between nodes B-D.

[0071] To be able to more specifically illustrate the actual execution process of the video recording control method in the above-mentioned embodiment, please further refer to Figure 4 , Figure 4 which shows the specific execution process of the video recording control method in the present application. The video recording control method in the embodiment of the present application can be executed according to the following steps:

[0072] S401: In response to the PIR trigger signal, control the recording device to enter the RTOS system bootloader.

[0073] Specifically, in the embodiment of the present application, the above-mentioned video recording control method is implemented by a video recording device such as a low-power camera. Exemplarily, the above-mentioned low-power camera includes a PIR sensor and a recording device for obtaining audio-visual information. It should be noted in advance that in the initial state, the above-mentioned low-power camera and the computer device controlling the camera are in a dormant state as a whole, but the PIR sensor on the low-power camera is in a working state and continuously detects the target area.

[0074] Specifically, when an object enters the target area, the PIR sensor detects the energy change in the target area, and then generates a PIR trigger signal. The above trigger signal is transmitted to the recording device on the low-power camera and the computer device controlling the low-power camera based on the communication connection relationship. The recording device responds to the PIR trigger signal and quickly starts to enter the RTOS system bootloader, waking up the encoding chip from the sleep state. At the same time, the computer device responds to the trigger signal to prepare for data processing and data caching. The specific operation timing can refer to Figure 5 Triggered by the PIR, in the O-A stage (about 0.5 s), the camera responds to the PIR trigger and quickly starts to enter the RTOS system bootloader, waking up the encoding chip from the sleep state.

[0075] S402: In response to the completion of the execution of the RTOS system bootloader, control the recording device to obtain the video and audio information of the target environment.

[0076] Specifically, in the embodiment of the present application, on the basis of the above example, after the recording device starts the RTOS system bootloader and completes, the recording device executes to obtain the video information and the synchronized audio information of the target area. The main function of the above video information and audio information is to provide basic data materials for subsequent artificial intelligence model recognition to determine whether to start or not start the monitoring video recording for the target area.

[0077] Specifically, please refer to Figure 5 , Figure 5 Schematically shows the main timing of the above video recording control method starting from the generation of the PIR trigger signal. After the PIR trigger signal is sent, node A represents the completion of the RTOS system bootloader. The above recording device starts to obtain the video and audio information of the target environment from node A. The running period of obtaining the video and audio information corresponds to Figure 5 the A-C segment in Figure 5 for subsequent identification and processing by the artificial intelligence model. During the time period corresponding to the A-C segment shown in

[0078] S403: Perform encoding processing on the video and audio information and store it in a preset cache space.

[0079] Specifically, in the embodiment of the present application, in order to determine whether to start or not start the monitoring video recording process for the target area, exemplarily, while the above recording device starts to acquire the video and audio information of the target environment, the low-power camera can encode and process the acquired video and audio information based on the above encoding chip, and store the encoded information in the cache memory set by itself or the cache storage space set on the computer that controls its operation, so that the artificial intelligence model can conveniently read and identify it. Regarding the above encoding method, any encoding method in the current related technologies can be adopted. The main factor determining the encoding method lies in the video file encoding methods supported by the artificial intelligence model, including but not limited to H.264, MPEG, AVC, H.265, etc.

[0080] S404: Based on the preset AI model and the video and audio information, determine the AI type information

[0081] Specifically, in the embodiment of the present application, in order to be able to identify the above video and audio information in a timely manner, exemplarily, during the process that the recording device continuously acquires video information and audio information of the target area and transfers them to the cache storage, the artificial intelligence model (hereinafter abbreviated as AI model) set on the control computer of the low-power camera reads and identifies the encoded video and audio data stored in the above cache storage in real time. For example, please refer to again Figure 5 , within the time range from node A to node B, while the recording device acquires video and audio information, encodes it, and transfers it, the AI model synchronously identifies and processes the encoded video and audio information. The main objects of identification are the objects entering the target area or the objects moving in the target area. After the identification and processing, for each identified object, the AI model will save the identified feature information as the corresponding AI type information. The main function of the above AI type information is to determine whether to start or not start the monitoring video recording for the target area in the subsequent process. The feature information included in the AI type information can include data such as the semantic information, size, and distance from the recording device of the object.

[0082] Optionally, the AI model selects a pre-trained ReID model based on the improved architecture of MobileNetV3 as the backbone network. The parameters and computational complexity of the above model are relatively low, and it can cope with different lighting conditions, pedestrian postures, and occlusion situations while ensuring a certain recognition accuracy, and is suitable for running on the low-power camera in the above embodiment.

[0083] S405: Judge whether the AI type information conforms to the AI type pre-selected by the user. If so, enter step S406; if not, enter step S408.

[0084] Specifically, in the implementation manner of the present application, in order to finally determine whether to start the surveillance video recording for the target area, illustratively, please refer to Figure 5 , within the range from node A to node B, while the AI ​​model identifies the above-mentioned audio and video information to determine the corresponding AI type information, the AI ​​model will further compare and match each AI type information according to the AI ​​type information identification standard pre-set by the user at the AI ​​model. For example, illustratively, the AI ​​type information identification standard pre-selected by the user in the AI ​​model is to identify people, vehicles, and objects within a preset range of distance from the recording device. Then, when the AI ​​model identifies the AI ​​type information, it compares the AI ​​type information based on the above-mentioned identification standard range pre-selected by the user. If there is an object that meets the above-mentioned identification standard range in the AI ​​type information, it can be determined to start the surveillance video recording for the target area. Conversely, if there is no object that meets the above-mentioned identification standard range in the AI ​​type information, it can be determined not to start the surveillance video recording for the target area.

[0085] Optionally, the above-mentioned AI type information can be expressed in the form of a feature vector corresponding to the object, and the feature vector includes various types of label information of the corresponding object, such as label information of people, cars, animals, etc. When comparing and matching each AI type information based on the AI ​​type information recognition standard, the comparison and matching of each AI type information can be achieved by comparing and matching the characters of a specific field in the feature vector. In the comparison and matching process, exemplarily, the extracted feature vectors can be stored in a cache queue, thereby providing a data comparison basis for identifying each dynamic target using an AI model.

[0086] S406: Cache feature information of dynamic targets in the video and audio information, and filter out all static targets in the picture within a preset image frame range.

[0087] Specifically, in the implementation manner of the present application, when it is determined to start the surveillance video recording for the target area, exemplarily, in order to identify the dynamic targets that can trigger and maintain the recording process, and at the same time to reduce the interference of objects that are stationary in the target area to the video recording process, before the recording process is officially started, the feature information of all dynamic objects that meet the identification standard range in the above example is first acquired. And the above feature information is cached in a preset cache queue. When each dynamic object is detected later, it can be quickly located based on the feature information of each dynamic object based on the above cache queue. Secondly, the stationary objects that appear in the image frames within a certain time before the start of the recording process are filtered out to prevent these stationary objects from erroneously extending the time of the actual recording process. The above image frame range is generally within the first 2s of the recording process, that is, corresponding to the time of the first 2s of the recording process.Figure 5 The image frame of segment B-C. Among them, the feature information can be expressed as the feature vector in the above embodiments, and mainly can show information such as the position and motion state of the dynamic target in the picture.

[0088] S407: Control the recording device to start continuously recording the current video and audio information of the target environment.

[0089] Specifically, in the embodiment of the present application, immediately following step S406, after the feature information acquisition for each dynamic object and the filtering for static objects are completed, all the preparatory work before the recording process starts has been completed. In this case, the recording device officially starts recording for the target area. It should be noted that according to the example in the above embodiment and referring to Figure 5 , during the process from node A to node B, the process of the recording device acquiring video and audio information for the target area has been ongoing. Therefore, the above recording process can be directly based on the above process of acquiring video and audio information. However, the difference from the above example is that for the purpose of facilitating the recognition by the AI model, the video and audio information obtained in the above example is stored in the cache space after being encoded. But in this embodiment, the video and audio information recorded during the recording process will be directly saved in the permanent storage space of the low-power camera or the permanent storage space of the control computer of the low-power camera in the form of video archives.

[0090] S408: Do not start the recording process and control the low-power camera to enter the sleep state again as a whole.

[0091] Specifically, in the embodiment of the present application, when it is determined not to start the monitoring video recording for the target area, for example, when it has been confirmed not to start the recording process, there is no need to continue the ongoing process of acquiring video and audio information for the target area. Therefore, in this case, the low-power camera enters the sleep state again as a whole, and keeps the PIR sensor performing real-time detection on the target area to wait for the next PIR trigger.

[0092] S409: Start timing when all the dynamic targets with cached feature information leave the picture, and continue for the first preset duration.

[0093] Specifically, in the embodiment of the present application, during the continuous execution of the recording process, in order to minimize the overall energy consumption of the low-power camera while ensuring the integrity of the recorded content, for example, the AI model will perform real-time monitoring on the motion state of the above-mentioned numbered dynamic targets in the frame information obtained by the recording device, and will also perform real-time tracking on the dynamic targets that appear in the frame to facilitate determining whether the recording process needs to be further extended. For example, during the recording process, if the AI model detects that all the dynamic targets with cached feature information in the frame have left the frame, then in order to determine whether the recording process needs to be further extended in this case, the AI model starts timing from the moment when all the above-mentioned dynamic targets leave the frame, and the timing length is the first preset duration. The above-mentioned first preset duration can be set to 2s exemplarily, or can be adjusted to other duration values according to actual detection requirements. The main purpose of setting the above-mentioned first preset duration is to provide a time window for judging whether the recording process needs to be extended, so as to reduce the judgment error of the short-term disappearance of dynamic targets. Please refer to Figure 5 Suppose that at node M, the AI model determines that all the dynamic targets with cached feature information have left the frame, then start a 2s timing from node M.

[0094] Regarding the method of determining whether all dynamic targets have left the frame, in some examples, starting from Figure 5 node C in, for each new video frame, the above-mentioned AI model continuously extracts feature vectors for each unfiltered dynamic target, and compares these feature vectors with the feature vectors pre-cached in the queue to calculate the similarity. Next, the above-mentioned similarity is used as a criterion to determine whether all the above-mentioned dynamic targets have left the frame.

[0095] Specifically, for each dynamic target, the similarity data corresponding to the dynamic target is calculated by comparing the current feature vector of the dynamic target with the feature vectors in the cache queue. When the above-mentioned similarity data is greater than the set threshold, that is, when the similarity does not meet the second preset condition, it can be determined that the corresponding dynamic target still remains within the frame range. On the contrary, when the above-mentioned similarity data is less than or equal to the set threshold, that is, when the similarity meets the second preset condition, it can be determined that the corresponding dynamic target has left the frame range. Then, start timing at the moment when it is detected that all the unfiltered dynamic targets have left the frame range, so as to determine when the recording process stops.

[0096] S410: Determine whether there are new dynamic targets entering the frame within the first preset duration. If so, return to step S406. If not, end the recording process.

[0097] Specifically, based on the above embodiments, within the timing range of the first preset duration, the AI model will continuously determine the captured images during the recording process. If a new dynamic target enters the frame within the first preset duration, it indicates that the recording process should not end at this time, and the movement of the above new dynamic target within the target area still needs to be monitored through the recording process. Then, in order to further maintain the recording process and avoid errors during the maintenance of the recording process, return to step S406 at this time. The AI model re-obtains the feature information of each dynamic target for the image with the new dynamic target and caches it in the cache queue, and performs the process of filtering out static targets to ensure that there will be no error delay in the continuous recording process.

[0098] On the contrary, if no new dynamic target is detected entering the frame within the first preset duration, then for the purpose of reducing the overall energy consumption of the low-power camera, the recording process can be declared stopped, and the video and audio information generated during the recording process is stored in the permanent storage space of the low-power camera or the permanent storage space of the control computer of the low-power camera in the form of a video file to form a video archive. Based on the above example, please refer to Figure 5 , at node M, the AI model determines that all dynamic targets with cached feature information have left the frame, then start timing for 2 seconds from node M to node N. If no new dynamic target is detected entering the frame within these 2 seconds, then the recording process stops at node N, and the camera stores the video and audio information obtained during the recording process in the form of a video file.

[0099] S411: Start timing when all dynamic targets with cached feature information remain stationary, and continue for the second preset duration.

[0100] Specifically, in the embodiment of the present application, during the continuous execution of the recording process, in order to minimize the overall energy consumption of the low-power camera while ensuring the integrity of the recorded content, for example, the AI model will perform real-time monitoring on the motion state of the above-mentioned numbered dynamic targets in the picture information obtained by the recording device, and will also perform real-time tracking on the dynamic targets that appear in the picture, so as to determine whether the recording process needs to be further extended. For example, during the recording process, the AI model detects that all the numbered dynamic targets in the picture have become stationary relative to the target area. Based on the above embodiment and the current related technical solutions, only moving objects in the picture will continue the entire recording process, and stationary objects cannot continue the entire recording process. Therefore, in the above scenario, the AI model starts timing from the moment when it detects that all the numbered dynamic targets have become stationary relative to the target area, and the timing length is the second preset duration. The above second preset duration can be set to 2s exemplarily, or can be adjusted to other duration values according to actual detection requirements. The main purpose of setting the above second preset duration is to provide a time window for judging whether the recording process needs to be extended, so as to reduce the judgment error of the situation where the dynamic target disappears briefly. Please refer to Figure 5 , it is assumed that at node M, the AI model determines that all the dynamic targets with cached feature information have left the picture, then start a 2s timing from node M.

[0101] Regarding how to determine whether all dynamic targets remain stationary, in some examples, starting from Figure 5 node C in, for each new video frame, the above AI model continuously extracts feature vectors for each dynamic target that has not been filtered out, and uses the feature vectors of adjacent frames as basic data for comparison calculations, so as to obtain the position offset of the corresponding dynamic target on the picture over time. Next, use the above position offset as a criterion to determine whether the above dynamic target remains stationary.

[0102] Specifically, for each dynamic target, the position offset corresponding to the dynamic target is obtained by comparing and calculating the feature vectors of the dynamic target in adjacent frames. If the above position offset is less than or equal to the preset threshold, that is, when the position offset meets the third preset condition, it can be determined that the corresponding dynamic target remains stationary. Correspondingly, if the above position offset is greater than the preset threshold, that is, when the position offset does not meet the third preset condition, it can be determined that the corresponding dynamic target is in a moving state. In addition, exemplarily, the moving speed and moving direction of each dynamic target can be further analyzed according to the magnitude and direction of the above position offset.

[0103] S412: Determine whether each dynamic target starts to move again within the second preset duration. If so, return to step S406; if not, end the recording process.

[0104] Specifically, based on the above embodiment, within the timing range of the second preset duration, the AI model will continuously determine the images obtained during the recording process. If a dynamic target that has been numbered and marked in the image starts to move again within the second preset duration, it indicates that the recording process should not end at this time. The continuous movement of the above dynamic target within the target area still needs to be monitored through the recording process. Then, in order to further maintain the recording process and avoid errors during the maintenance of the recording process, return to step S406 at this time. The AI model re-obtains the feature information of each dynamic target from the image where the dynamic target starts to move again from the stationary state and caches it in the cache queue, and performs the process of filtering out stationary targets, which can ensure that the continuous recording process will not have an error delay.

[0105] On the contrary, if all the dynamic targets that have been numbered and marked remain stationary relative to the target area within the second preset duration, then for the purpose of reducing the overall energy consumption of the low-power camera, the recording process can be declared to stop. The video and audio information generated during the recording process is stored in the permanent storage space of the low-power camera or the permanent storage space of the control computer of the low-power camera in the form of a video file to form a video archive. Based on the above example, please refer to Figure 5 , at node M, the AI model determines that all the dynamic targets with cached feature information have left the image. Then, start a 2s timer from node M to node N. If all the above dynamic targets remain stationary within this 2s time range, then the recording process stops at node N, and the camera stores the video and audio information obtained during the recording process in the form of a video file.

[0106] It can be understood that the situations in steps S409, S410, S411, and S412 may overlap. That is, when the timing process starts for each dynamic target to meet the corresponding conditions, it is possible that new dynamic targets enter the image, or it is possible that some stationary dynamic targets start to move again. At this time, the S409 and S410 branches and the S411 and S412 branches can be executed respectively according to the above examples, and there is no conflict between the two groups of branches.

[0107] Thus, in the video recording control method according to the embodiments of the present application, when the recording device is triggered to start acquiring environmental video and audio information based on an infrared sensor such as a PIR, element analysis can be performed based on a pre-set AI model, so as to supplement and verify the trigger of the PIR for the recording process, reduce the probability of premature termination of the recording process, and at the same time, the above AI model can be used for image recognition to determine the conditions of the dynamic target actions in the picture, and then use the determination result to end the recording process, so as to reduce the power consumption of the recording device as much as possible while ensuring the integrity of the recording process and extend the service life of the recording device.

[0108] An embodiment of the present application also provides a video recording device, including the above-mentioned preset infrared sensor and preset recording device. The above video recording device can be served by the low-power camera in the above embodiment, the preset infrared sensor corresponds to the PIR sensor in the above embodiment, and the preset recording device corresponds to the recording device in the above embodiment.

[0109] An embodiment of the present application also provides a computer device, which can serve as the control computer of the above video recording device. The above computer device includes a memory and a processor, and a computer program is stored in the memory. When the computer program is executed by the processor, the method steps of the above video recording control method can be implemented.

[0110] An embodiment of the present application also provides a computer-readable storage medium, storing a computer program, which when executed by one or more processors, implements the method steps of the above video recording control method.

[0111] In the description of this specification, the descriptions with reference to terms such as "certain embodiments", "in an example", "exemplarily", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiments or examples. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.

[0112] Any process or method description, whether in a flowchart or otherwise described herein, can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present application includes additional implementations where functions may be executed not in the order shown or discussed, including substantially concurrently or in reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application pertain.

[0113] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A video recording control method, characterized in that: The method comprises: Acquire a trigger signal generated by a preset infrared sensor; In response to the trigger signal, controlling a preset recording device to start a recording program; Controlling the preset recording device to obtain first target media information of a target environment; According to the first target media information, based on a preset artificial intelligence model, determine target artificial intelligence model type information; When the target artificial intelligence model type information satisfies the first preset condition, the preset recording device is controlled to continuously record the second target media information of the target environment until the second target media information satisfies the second preset condition. The second preset condition is configured to indicate that, based on the preset artificial intelligence model, it is determined that the state of the dynamic target in the second target media information in the picture meets the condition for stopping recording.

2. The method according to claim 1, characterized in that In response to the trigger signal, controlling a preset recording device to start a recording program includes: In response to a trigger signal generated by the preset infrared sensor due to detecting energy changes within a target range, the preset recording device is controlled to enter a system boot loader.

3. The method according to claim 2, characterized in that The controlling the preset recording device to obtain first target media information of the target environment includes: In response to the system boot loader being executed, controlling the preset recording device to obtain original media information of the target environment; The original media information is encoded and stored in a preset cache space to generate the first target media information.

4. The method according to claim 1, characterized in that: The method of controlling the preset recording device to continuously record the second target media information of the target environment when the target artificial intelligence model type information satisfies the first preset condition until the target artificial intelligence model type information satisfies the second preset condition includes: When the target artificial intelligence model type information matches the information recognition type preset by the user, cache the feature information corresponding to the dynamic target in the first target media information; Controlling the preset recording device to continuously record the second target media information of the target environment; Within the time range in which it is determined according to the characteristic information that all the dynamic targets have left the screen and lasted for a first preset time period, if there is no secondary dynamic target newly entering the screen, controlling the preset recording device to stop recording; and / or When it is determined according to the characteristic information that the dynamic targets remain stationary for a second period of time, the preset recording device is controlled to stop recording.

5. The method according to claim 4, characterized in that The step of controlling the preset recording device to stop recording if there is no secondary dynamic target newly entering the screen within a time range in which the dynamic targets are determined to have all left the screen and last for a first preset time period according to the characteristic information, comprises: Acquire current feature information of the dynamic target; Determine, based on the cached feature information and the current feature information, the similarity between the dynamic target in the current picture and the dynamic target in the first target media information; When the similarity satisfies a second preset condition, determining that the dynamic target leaves the screen; Within the time range when all the dynamic objects leave the screen and last for a first preset time period, if there is no secondary dynamic object newly entering the screen, the preset recording device is controlled to stop recording.

6. The method according to claim 4, characterized in that When it is determined according to the characteristic information that the dynamic targets all remain stationary for a second period of time, controlling the preset recording device to stop recording includes: Acquire current feature information of the dynamic target; Determining a position offset of the dynamic target according to the cached feature information and the current feature information; When the position offset satisfies a third preset condition, determining that the dynamic target remains stationary; When the dynamic targets all remain still for a second period of time, the preset recording device is controlled to stop recording.

7. The method according to claim 4, characterized in that The method of controlling the preset recording device to continuously record the second target media information of the target environment when the target artificial intelligence model type information satisfies the first preset condition until the target artificial intelligence model type information satisfies the second preset condition includes: When the target artificial intelligence model type information does not match the information identification type preset by the user, the preset recording device is controlled to enter a sleep state.

8. The method according to claim 4, characterized in that The method further comprises: Within the time range when all the dynamic targets leave the screen and last for a first preset time, if there is a secondary dynamic target that newly enters the screen, re-numbering the dynamic targets in the first target media information; and / or Within the second time range after the dynamic objects are all stationary, if there is a moving dynamic object among the dynamic objects, the dynamic objects in the first target media information are renumbered.

9. The method according to any one of claims 4 to 8, characterized in that: When the target artificial intelligence model type information satisfies the information identification type preset by the user, performing numbering processing on the dynamic target in the first target media information also includes: When the target artificial intelligence model type information matches the information recognition type preset by the user, the dynamic target that remains stationary within the preset image frame range is filtered out according to the first target information; The feature information corresponding to the remaining dynamic targets is cached.

10. A video recording device, characterized in that: The video recording device includes a preset infrared sensor and a preset recording device as described in any one of claims 1 to 9.

11. A computer device, characterized in that: The device comprises a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the video recording control method according to any one of claims 1 to 9 is implemented.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by one or more processors, the video recording control method according to any one of claims 1 to 9 is implemented.