Monitoring system control method and device and storage-computation integrated system

By employing an in-memory computing system and neural network technology in the monitoring system, images are segmented and processed according to the shooting device's perspective, solving the high power consumption problem of the monitoring system during long-term display and achieving low-power, high-precision image recognition and display control.

CN119678192BActive Publication Date: 2026-08-25BOE TECHNOLOGY GROUP CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380009259.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-31
Publication Date
2026-08-25
Estimated Expiration
2043-05-31

AI Technical Summary

Technical Problem

Existing monitoring systems consume a lot of power when displaying surveillance footage for extended periods, which affects system lifespan. Achieving low-power operation is a technical problem that urgently needs to be solved.

Method used

By employing an in-memory computing system combined with neural network technology, the image to be detected is divided into multiple target sub-images according to the perspective of the shooting device. Each sub-image is processed using a pre-trained target neural network model. The recognition results are combined to determine whether there is a target object in the image to be detected, thereby controlling the display state of the terminal and reducing system power consumption.

Benefits of technology

By using precise image recognition and reasonable sub-image segmentation, the recognition accuracy of the monitoring system has been improved, the system power consumption has been reduced, long-term standby time has been achieved, and maintenance costs have been reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119678192B_ABST
    Figure CN119678192B_ABST
Patent Text Reader

Abstract

The present disclosure provides a kind of monitoring system control method, device, computer equipment and storage medium, belong to image recognition and terminal monitoring technical field, wherein the control method of monitoring system includes obtaining image to be detected;According to the shooting angle of the shooting device for shooting image to be detected, image to be detected is divided into multiple target sub-images;For each target sub-image in multiple target sub-images, a pre-trained target neural network model is used to process target sub-image, and the recognition result of target sub-image is obtained;Based on the recognition result corresponding to each target sub-image, the detection result of whether target object exists in image to be detected is obtained, and the detection result is sent to terminal, to make terminal at least based on detection result determination display state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure belongs to the field of image recognition and terminal monitoring technology, specifically relating to a control method, device and storage-computing integrated system for a monitoring system. Background Technology

[0002] With the rapid development of information technology, modern electronic devices are also rapidly evolving towards intelligence, lightweight design, and portability. For smart terminals, displaying monitoring images on the screen for extended periods results in high system power consumption, impacting system lifespan. Therefore, achieving low-power operation is a pressing technical challenge in the monitoring field. Summary of the Invention

[0003] This disclosure aims to at least solve one of the technical problems existing in the prior art, and to provide a control method, device and storage-computing integrated system for a monitoring system.

[0004] Firstly, the technical solution adopted to solve the technical problem of this disclosure is a control method for a monitoring system, including:

[0005] Acquire the image to be detected;

[0006] Based on the shooting angle of the shooting device that captured the image to be detected, the image to be detected is divided into multiple target sub-images;

[0007] For each of the plurality of target sub-images, a pre-trained target neural network model is used to process the target sub-image to obtain the recognition result of the target sub-image;

[0008] Based on the recognition results corresponding to each of the target sub-images, a detection result is obtained as to whether a target object exists in the image to be detected, and the detection result is sent to the terminal so that the terminal can determine the display state based at least on the detection result.

[0009] In some embodiments, dividing the image to be detected into multiple target sub-images according to the shooting angle of the shooting device includes:

[0010] When the shooting angle is within a preset angle range, the image to be detected is divided into a first sub-region, a second sub-region, and a third sub-region arranged sequentially along a first direction, and the first sub-region and the third sub-region partially overlap with the second sub-region; wherein, the width of the first sub-region in the first direction is equal to the width of the third sub-region in the first direction, and the width of the second sub-region in the first direction is greater than the width of the first sub-region in the first direction; the first direction is the height direction of the image to be detected;

[0011] The portion of the image to be detected located in the first sub-region is divided into multiple first sub-images arranged side by side along the second direction, at least some of the first sub-images having equal width in the second direction; the second direction is the width direction of the image to be detected;

[0012] The portion of the image to be detected located in the second sub-region is divided into multiple second sub-images arranged side by side along the second direction, and at least some of the second sub-images have equal widths in the second direction;

[0013] The portion of the image to be detected located in the third sub-region is divided into multiple third sub-images arranged side by side along the second direction, at least some of the third sub-images having equal width in the second direction; the target sub-image includes a first sub-image, a second sub-image, and a third sub-image;

[0014] The width of the third sub-image in the second direction is greater than the width of the second sub-image in the second direction, and the width of the second sub-image in the second direction is greater than the width of the first sub-image in the second direction.

[0015] In some embodiments, each of the first sub-images at least partially overlaps in the second direction; each of the second sub-images at least partially overlaps in the second direction; and each of the third sub-images does not overlap in the second direction.

[0016] In some embodiments, the width ratio of the second sub-region to the width of the first sub-region in the first direction is 2:1;

[0017] The width ratio of the third sub-image to the second sub-image in the second direction is 3:2, and the width ratio of the second sub-image to the first sub-image in the second direction is 4:3.

[0018] In some embodiments, among a plurality of first sub-images arranged side by side along the second direction, except for the first and last first sub-images arranged side by side along the second direction, the widths of the remaining first sub-images in the second direction are all the same; the widths of the first and last first sub-images in the second direction are the same and smaller than the widths of the remaining first sub-images in the second direction; the ratio of the overlap width of two adjacent first sub-images in the second direction to the width of the remaining first sub-images in the second direction is 1:10.

[0019] Among the multiple second sub-images arranged side by side along the second direction, except for the first and last second sub-images arranged side by side along the second direction, the widths of the remaining second sub-images in the second direction are all the same; the widths of the first and last second sub-images in the second direction are the same and smaller than the widths of the remaining second sub-images in the second direction; the ratio of the overlap width of two adjacent second sub-images in the second direction to the width of the remaining second sub-images in the second direction is 1:10.

[0020] Among the multiple third sub-images arranged side by side along the second direction, except for the first and last third sub-images arranged side by side along the second direction, the widths of the remaining third sub-images are all the same in the second direction; the widths of the first and last third sub-images are the same in the second direction and are smaller than the widths of the remaining third sub-images in the second direction; the ratio of the overlap width of two adjacent third sub-images in the second direction to the width of the remaining third sub-images in the second direction is 1:10.

[0021] The first sub-image and the second sub-image that are adjacent in the first direction have an overlap width in the first direction that is 1:10 to the width of the second sub-image in the first direction; the third sub-image and the second sub-image that are adjacent in the first direction have an overlap width in the first direction that is 1:10 to the width of the second sub-image in the first direction.

[0022] In some embodiments, the target neural network model is trained by the following steps:

[0023] Obtain a first training dataset and a second training dataset; the second training dataset is obtained by filtering the first training dataset.

[0024] The teacher machine learning model to be trained is trained based on the first training dataset to obtain the preliminarily trained teacher machine learning model.

[0025] The pre-trained teacher machine learning model is trained based on the second training dataset to obtain the trained teacher machine learning model.

[0026] Based on the second training dataset and the trained teacher machine learning model, the knowledge distillation training method is used to train the student machine learning model to be trained, and the trained student machine learning model is used as the target neural network model.

[0027] In some embodiments, the second training dataset includes multiple training images labeled with sample labels;

[0028] The step of training the student machine learning model to be trained using a knowledge distillation method based on the second training dataset and the trained teacher machine learning model to obtain the trained student machine learning model as the target neural network model includes:

[0029] The training image is input into the trained teacher machine learning model to obtain the first output result of the trained teacher machine learning model;

[0030] The training image is input into the student machine learning model to be trained, and the second output result of the student machine learning model to be trained is obtained.

[0031] Based on the first output result and the second output result, determine the first loss function;

[0032] Based on the second output result and the sample labels of the training images, a second loss function is determined;

[0033] Based on the first loss function and the second loss function, the weighted loss function is obtained;

[0034] The parameters of the student machine learning model to be trained are adjusted according to the weighted loss function until the weighted loss function converges, and the trained student machine learning model is obtained as the target neural network model.

[0035] In some embodiments, the first training dataset is determined by the following steps:

[0036] Obtain the original dataset; the original dataset includes multiple initial sample images;

[0037] The target object is identified in the initial sample image. If the target object exists, a first reference box containing the target object is determined.

[0038] Based on the position information of the first reference frame, update the position of the first reference frame to obtain the second reference frame;

[0039] Based on the position information of the second reference frame and the position information of the first reference frame, the overlap between the second reference frame and the first reference frame is determined;

[0040] Based on the comparison result between the overlap and the first preset threshold, the sample images located in the second reference box of the initial sample image are labeled with the sample label, and the labeled sample images are used as training images in the first training dataset.

[0041] In some embodiments, updating the position of the first reference frame based on the position information of the first reference frame to obtain the second reference frame includes:

[0042] Based on the position information of the first reference frame, move a specific coordinate point in the first reference frame to obtain a third reference frame;

[0043] Based on the position information of the third reference frame, the width and height of the third reference frame are adjusted according to a preset scaling factor, with the center point of the third reference frame as the center, to obtain the fourth reference frame;

[0044] Based on the position information of the fourth reference frame, a new center point is determined according to the first preset center point range; and based on the new center point and the preset clipping range, the width and height of the fourth reference frame are adjusted to obtain the second reference frame.

[0045] In some embodiments, labeling the sample images of the portion of the initial sample image located within the second reference frame with the sample label based on the comparison result between the overlap and a first preset threshold includes:

[0046] When the overlap is greater than or equal to the first preset threshold, a first preset range is generated according to a first preset probability, a second preset range is generated according to a second preset probability, a third preset range is generated according to a third preset probability, and a fourth preset range is generated according to a fourth preset probability; the first preset probability is greater than the second preset probability, the second preset probability is greater than the third preset probability, and the third preset probability is greater than or equal to the fourth preset probability; the sum of the first preset probability, the second preset probability, the third preset probability, and the fourth preset probability is 1;

[0047] When the overlap is within the first preset range, the portion of the sample images located in the second reference box in the initial sample image is determined and labeled as positive sample labels, so that the ratio between the number of training images with the overlap within the first preset range and the total number of training images with the positive sample labels in the first training dataset is the first preset probability.

[0048] When the overlap is within the second preset range, the portion of the sample images located in the second reference box in the initial sample image is determined and labeled as positive sample labels, so that the ratio between the number of training images with the overlap within the second preset range and the total number of training images with the positive sample labels in the first training dataset is the second preset probability.

[0049] When the overlap is within the third preset range, the sample images located in the second reference box in the initial sample image are identified and labeled as positive sample labels, so that the ratio between the number of training images with the overlap within the third preset range and the total number of training images with the positive sample labels in the first training dataset is the third preset probability.

[0050] When the overlap is within the fourth preset range, the sample images located in the second reference box in the initial sample image are identified and labeled as positive sample labels, so that the ratio between the number of training images with the overlap within the fourth preset range and the total number of training images with the positive sample labels in the first training dataset is the fourth preset probability.

[0051] The overlap within the first preset range is less than the overlap within the second preset range, the overlap within the second preset range is less than the overlap within the third preset range, and the overlap within the third preset range is less than the overlap within the fourth preset range.

[0052] In some embodiments, the step of obtaining the first training dataset, if it is determined that the target object does not exist when identifying the target object in the initial sample image, further includes:

[0053] The initial sample images are labeled with negative sample labels, and the initial sample images labeled with negative sample labels are used as training images in the first training dataset.

[0054] In some embodiments, after identifying the target object in the initial sample image to determine its existence and determining a first reference bounding box containing the target object, the method further includes:

[0055] Based on the size of the initial sample image, a target center point is determined in the initial sample image according to the second preset center point range, and a fifth reference frame is determined with the target center point as the center.

[0056] Based on the position information of the fifth reference frame and the position information of the first reference frame, the overlap between the fifth reference frame and any first reference frame in the initial sample image is determined;

[0057] If the overlap between the fifth reference box and any first reference box in the initial sample image is less than or equal to the second preset threshold, then the portion of the initial sample image located in the fifth reference box is labeled as a negative sample, and the portion of the sample image labeled with the negative sample is used as the training image.

[0058] In some embodiments, acquiring the image to be detected includes:

[0059] Acquire multiple consecutive video frames captured by the camera device;

[0060] Based on the multiple consecutive video frames, the frame difference method is used to determine whether there is a moving target in the shooting scene of the shooting device. If the moving target exists, the video frames captured by the shooting device are used as images to be detected.

[0061] Secondly, this disclosure also provides a control method for a monitoring system, the monitoring system including a storage-computing integrated system and a terminal; the storage-computing integrated system includes a processing module, a storage-computing integrated module, and an external storage module; the control method for the monitoring system includes:

[0062] The processing module acquires the image to be detected; according to the shooting angle of the shooting device that captured the image to be detected, the image to be detected is divided into multiple target sub-images and stored in the external storage module;

[0063] The in-memory computing module reads the target sub-image; uses a pre-trained target neural network model to process the target sub-image, obtains the recognition result of the target sub-image, and stores the recognition result in the external storage module;

[0064] The processing module obtains the detection result of whether a target object exists in the image to be detected based on the recognition result corresponding to each of the target sub-images, and sends the detection result to the terminal;

[0065] The terminal determines the display status based on the detection results.

[0066] In some embodiments, the terminal includes a display module and a main control module;

[0067] The terminal determines the display status based on the detection result, including:

[0068] The terminal determines the display status based on the detection result, including:

[0069] When the main control module determines that the detection result indicates that the target object exists in the image to be detected, it sends a wake-up request to the display module and resets the sleep timer to respond to the terminal request to receive the detection result sent by the processing module; when it determines that the detection result indicates that the target object does not exist in the image to be detected, and the current system time is greater than the time since the last wake-up of the display module is greater than a preset time, it sends a sleep request to the display module and controls itself to enter a sleep state.

[0070] The display module responds to the wake-up request and determines that the display state is to normally display the captured image from the shooting device; it also responds to the sleep request and determines that the display state is to sleep.

[0071] Thirdly, this disclosure also provides an in-memory computing system, comprising a processing module, an in-memory computing module, and an external storage module;

[0072] The processing module acquires the image to be detected; according to the shooting angle of the shooting device that captured the image to be detected, the image to be detected is divided into multiple target sub-images and stored in the external storage module;

[0073] The in-memory computing module reads the target sub-image; uses a pre-trained target neural network model to process the target sub-image, obtains the recognition result of the target sub-image, and stores the recognition result in the external storage module;

[0074] The processing module obtains the detection result of whether a target object exists in the image to be detected based on the recognition result corresponding to each of the target sub-images, and sends the detection result to the terminal so that the terminal can determine the display state based on the detection result.

[0075] Fourthly, this disclosure also provides a control device for a monitoring system, which includes an in-memory computing system and a terminal; the in-memory computing system includes a processing module, an in-memory computing module, and an external storage module;

[0076] The processing module is configured to acquire an image to be detected; divide the image to be detected into multiple target sub-images according to the shooting angle of the shooting device that captured the image to be detected, and store them in the external storage module;

[0077] The in-memory computing module is configured to read the target sub-image; process the target sub-image using a pre-trained target neural network model to obtain the recognition result of the target sub-image; and store the recognition result in the external storage module.

[0078] The processing module is configured to obtain a detection result of whether a target object exists in the image to be detected based on the recognition result corresponding to each of the target sub-images, and send the detection result to the terminal;

[0079] The terminal is configured to determine the display status based on the detection result.

[0080] Fifthly, embodiments of this disclosure also provide a computer non-transient readable storage medium, wherein a computer program is stored on the computer non-transient readable storage medium, and the computer program, when executed by a processor, performs the steps of the control method of the monitoring system as described in any one of the embodiments of the first aspect above, and / or the steps of the control method of the monitoring system as described in any one of the embodiments of the second aspect above. Attached Figure Description

[0081] Figure 1 A flowchart of a control method for a monitoring system provided in this embodiment of the present disclosure;

[0082] Figure 2 This is a schematic diagram illustrating the effect of reasonably dividing the image to be detected according to an embodiment of the present disclosure;

[0083] Figure 3 A schematic diagram illustrating an exemplary training process for a target neural network model provided in this disclosure embodiment;

[0084] Figure 4 A schematic diagram of an exemplary network architecture for knowledge distillation provided in this disclosure embodiment;

[0085] Figure 5 A schematic diagram of an exemplary student model provided for embodiments of this disclosure;

[0086] Figure 6a An exemplary terminal workflow diagram provided for embodiments of this disclosure;

[0087] Figure 6b A flowchart illustrating an exemplary in-memory computing system provided for embodiments of this disclosure;

[0088] Figure 7 A schematic diagram of an in-memory computing system provided in an embodiment of this disclosure;

[0089] Figure 8 A schematic diagram of the control device for the monitoring system provided in this embodiment of the disclosure. Detailed Implementation

[0090] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0091] Unless otherwise defined, the technical or scientific terms used in this disclosure shall have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms “first,” “second,” and similar terms used in this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms “an,” “a,” or “the,” and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms “comprising,” “including,” or “including,” and similar terms mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects.

[0092] In this disclosure, "multiple or several" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0093] In related technologies, with the in-depth research and popularization of artificial intelligence algorithms represented by deep learning neural networks, intelligent electronic devices and related application scenarios are ubiquitous, such as facial recognition, voice recognition, smart homes, security monitoring, and autonomous driving. However, limited by the classic von Neumann computing architecture, data storage and processing are separated, with computing and storage functions handled by the central processing unit (CPU) and memory respectively. The performance gap between the two forms a "memory wall," resulting in significant energy consumption due to frequent data transfers between the computing core and memory. However, in-memory computing designs can now solve this problem. In-memory computing refers to combining memory and computing more tightly than traditional computer architectures, thereby reducing the overhead of memory access and solving the "memory wall" problem. The three key elements of artificial intelligence are computing power, data, and algorithms, and in-memory computing features high computing power, low power consumption, and low latency. Therefore, in-memory computing will play a significant role in the future of artificial intelligence. Currently, common in-memory computing methods include flash memory and static random-access memory (SRAM). Typically, many national or industry standards associated with smart terminal products have certain standby power consumption requirements. To design a smart wake-up function on such products and support 24-hour standby wake-up through voice, personnel detection, and other methods, it is necessary to ensure that the entire detection system has low standby power consumption in order to meet the corresponding energy efficiency standards.

[0094] In view of this, the present disclosure provides a control method for a monitoring system. Specifically, the method involves: acquiring an image to be detected; dividing the image to be detected into multiple target sub-images according to the shooting angle of the shooting device that captured the image; processing each of the multiple target sub-images using a pre-trained target neural network model to obtain a recognition result for the target sub-image; obtaining a detection result for whether a target object exists in the image to be detected based on the recognition result corresponding to each target sub-image; and sending the detection result to a terminal so that the terminal determines the display state based at least on the detection result.

[0095] This disclosure employs a neural network to detect images captured by a camera, feeding the detection results back to the terminal. The terminal then determines the current display state based on these results. For example, if no target object is detected for an extended period, the system can enter a sleep state, saving power. If, after entering sleep mode, the detection result indicates the presence of a target object, the system can automatically restart and resume normal display. Furthermore, in the intelligent wake-up application scenario, this disclosure, while saving system power, combines neural network image recognition technology. Based on the shooting angle, the image to be detected captured by the camera is rationally divided into multiple target sub-images. Each target sub-image is identified separately, and the recognition results are integrated to obtain the detection result, improving the accuracy of target object recognition in the image and thus more precisely controlling the terminal's wake-up timing. Additionally, utilizing an integrated storage and computing module—a system architecture that combines storage and computing—supports complex data operations in the neural network, greatly avoiding energy loss caused by data transfer and reducing system power consumption.

[0096] To facilitate understanding, a control method for a monitoring system provided in this disclosure will first be described in detail. The execution entity of this control method can be a memory-based computing system with certain computing capabilities. This system includes a processing module, a memory-based computing module, and an external storage module. The processing module is, for example, the main processor (CPU) of the memory-based computing system. The memory-based computing module can be, for example, a memory-based computing core integrating Flash or SRAM. The external storage module can be, for example, an external storage module (DRAM), offering faster response speeds and reduced power consumption of the monitoring system. The memory-based computing system can be part of the monitoring system. Furthermore, this disclosure designs a lightweight target neural network model, migrates it to the memory-based computing module, and uses a low-power memory-based computing approach to assist the target neural network in inference, completing the wake-up service of the smart terminal, reducing power consumption and increasing response speed.

[0097] Figure 1 A flowchart of a control method for a monitoring system provided in this disclosure embodiment is shown below. Figure 1 As shown, steps S11 to S14 are included, wherein:

[0098] S11. Obtain the image to be detected.

[0099] In this step, the image to be detected is an image captured by a camera device, prepared for detecting the presence of a target object. Step S11 can, for example, be the execution process of a processing module in a memory computing system. In one case, the image to be detected is an image captured in real-time by the camera device; in another case, the image to be detected can be an image selected by the processing module from the images captured in real-time by the camera device that has a high probability of containing a target object. The selection method can employ techniques such as frame difference analysis, and this disclosure does not impose specific limitations.

[0100] The shooting device can be a camera or other device with image capture capabilities. The shooting device captures images of the shooting scene in real time, where target objects may be present, such as people moving around; however, target objects may not be present in the shooting scene in real time.

[0101] Afterwards, the acquired image to be detected can be stored in an external storage module for later use in the image segmentation process.

[0102] S12. Based on the shooting angle of the shooting device that captured the image to be detected, divide the image to be detected into multiple target sub-images.

[0103] Step S12 can be, for example, the execution process of a processing module in a storage-computing system. The processing module reads the image to be detected from the external storage module, and divides the image to be detected into multiple target sub-images according to the shooting angle of the shooting device that captured the image. The divided multiple target sub-images are then stored in the external storage module for later use in the target neural network model for processing.

[0104] Since the terminal (including the display screen) in the intelligent wake-up application scenario determines the display state based on the detection results, if the confidence level of the detection results is low, the terminal may mistakenly enter a sleep state, thus preventing the user from accurately monitoring the scene being captured. Therefore, while saving system power consumption, it is necessary to improve the accuracy of the detection results. Based on this, in this embodiment, the processing module divides the image to be detected into multiple target sub-images according to the shooting angle of the shooting device. By refining the image to be detected and processing each refined target sub-image separately using a pre-trained target neural network model, the image recognition accuracy can be improved.

[0105] It should be noted that the higher the level of detail in the image to be detected, that is, the more target sub-images a single image can be divided into, the more accurate the detection result will be after processing by the target neural network model and combining the recognition results of each target sub-image. However, a higher level of detail in the image to be detected requires a large amount of subsequent model inference, which increases the system's power consumption. Based on this, the embodiments of this disclosure reasonably divide the image to be detected according to the shooting angle of the imaging device, which can both ensure image recognition accuracy and reasonably divide the image to be detected to avoid a large number of image detections, thereby reducing system power consumption.

[0106] For example, multiple shooting angles can be pre-set for the shooting device, such as normal angle, close-up (angle smaller than normal shooting angle), medium angle (angle larger than normal shooting angle), and distant angle (angle larger than medium angle), etc. Different shooting angles correspond to different image segmentation methods. For instance, the image to be detected corresponding to a distant angle has a higher level of detail than the image to be detected corresponding to a medium angle, the image to be detected corresponding to a medium angle has a higher level of detail than the image to be detected corresponding to a normal angle, and the image to be detected corresponding to a normal angle has a higher level of detail than the image to be detected corresponding to a close-up.

[0107] S13. For each target sub-image in multiple target sub-images, use a pre-trained target neural network model to process the target sub-image and obtain the recognition result of the target sub-image.

[0108] Step S13 can be, for example, the execution process of an in-memory computing module in an in-memory computing system. The in-memory computing module integrates a target neural network model. The in-memory computing module reads the target sub-image from the external storage module. For any target sub-image, it processes the target sub-image using the pre-trained target neural network model to obtain the recognition result of the target sub-image, and stores the recognition result of the target sub-image in the external storage module.

[0109] Specifically, the target sub-image is input into the target neural network model, and features are extracted and analyzed using multiple network layers to determine the presence of a target object. During feature extraction and analysis of the target sub-image, the operations and outputs of different network layers are processed using a memory-in-memory (Flash or SRAM) architecture. For example, for a convolutional layer in a convolutional neural network, the intermediate results of the internal convolution downsampling process are cached during the convolution operation until the convolutional output of a target sub-image is output. This output can then be used as the input to another convolutional layer. This process is repeated sequentially across convolutional layers until the target neural network model outputs the recognition result of the target sub-image.

[0110] This in-memory computing approach caches the intermediate results of internal convolutions, making them easier for the next convolutional layer to process. Compared to traditional systems with memory walls, this avoids energy loss caused by data transfer, improves computational efficiency, and reduces system power consumption, making it a promising candidate for portable devices.

[0111] The target neural network model can employ a lightweight classification model, and its identification result may indicate the presence or absence of a target object. Here, the target object can be, for example, a person or an object.

[0112] S14. Based on the recognition results corresponding to each target sub-image, obtain the detection result of whether there is a target object in the image to be detected, and send the detection result to the terminal so that the terminal can determine the display state based at least on the detection result.

[0113] Step S14 can be, for example, the execution process of a processing module in a storage-based computing system. The processing module obtains the recognition results of each target sub-image of the image to be detected from the external storage module, and based on the recognition results corresponding to each target sub-image, obtains the detection result of whether a target object exists in the image to be detected, and sends the detection result to the terminal.

[0114] For example, the processing module can obtain the detection result "a target object exists" based on the recognition results corresponding to each target sub-image, provided that any recognition result indicates the existence of a target object.

[0115] For example, the processing module can obtain the detection result "a target object exists" based on the recognition results corresponding to each target sub-image, when the proportion of the results indicating the presence of a target object in the total recognition results exceeds a set value.

[0116] The processing module sends the detection results to the terminal, and the main control module in the terminal can determine the display status based at least on the detection results; the display screen in the terminal can determine whether to display a normal image or enter sleep mode based on the currently determined display status.

[0117] The control method of the monitoring system disclosed herein adopts an in-memory computing system to realize data processing, avoiding energy loss caused by data transfer, improving computing efficiency, and reducing system power consumption. At the same time, based on the shooting angle, the image to be detected captured by the shooting device is reasonably divided into multiple target sub-images. Combined with a lightweight target neural network model, the system can accurately identify whether there is a target object in the image to be detected, thereby determining a more accurate display state. When the display screen enters sleep mode, it saves system power consumption and can maintain standby for a long time when the terminal is not actively woken up, thereby solving the problem of high power consumption and reducing a lot of maintenance costs.

[0118] In some embodiments, if the entire image to be detected is input into the model for classification and recognition, the target objects will not be concentrated, resulting in inaccurate recognition results. Therefore, in order to improve the model's recognition accuracy, for step S12, the image to be detected can be reasonably divided according to the shooting angle of the shooting device. Specifically:

[0119] When the shooting angle is within a preset angle range, the image to be detected is divided into a first sub-region 01, a second sub-region 02, and a third sub-region 03 arranged side-by-side along a first direction Y, with the first sub-region 01 and the third sub-region 03 partially overlapping with the second sub-region 02; wherein the width of the first sub-region 01 in the first direction Y is equal to the width of the third sub-region 03 in the first direction Y, and the width of the second sub-region 02 in the first direction Y is greater than the width of the first sub-region 01 in the first direction Y; the first direction Y is the height direction of the image to be detected; the portion of the image to be detected located in the first sub-region 01 is divided into multiple first sub-images 011 arranged side-by-side along a second direction X, at least some of the first sub-images 011 having equal width in the second direction X; the second direction X represents the width direction of the image to be detected; the portion of the image to be detected located in the second sub-region 02 is divided into multiple second sub-images 012 arranged side by side along the second direction X, at least some of the second sub-images 012 having equal widths in the second direction X; the portion of the image to be detected located in the third sub-region 03 is divided into multiple third sub-images 013 arranged side by side along the second direction X, at least some of the third sub-images 013 having equal widths in the second direction X; the target sub-image includes a first sub-image 011, a second sub-image 012, and a third sub-image 013; the width of the third sub-image 013 in the second direction X is greater than the width of the second sub-image 012 in the second direction X, and the width of the second sub-image 012 in the second direction X is greater than the width of the first sub-image 011 in the second direction X.

[0120] In this embodiment, the preset viewing angle range can be set according to the actual situation. Figure 2 This is a schematic diagram illustrating the effect of reasonably segmenting the image to be detected according to the embodiments of this disclosure, such as... Figure 2As shown, for example, the preset viewing angle range is a range within the normal viewing angle. The image to be detected is divided into three regions, specifically a first sub-region 01, a second sub-region 02, and a third sub-region 03 arranged side by side along the first direction Y, located at the top, middle, and bottom positions of the image to be detected when placed face-up. The second sub-region 02 is located between the first sub-region 01 and the third sub-region 03, and the deployment scene positions corresponding to the first sub-region 01, the second sub-region 02, and the third sub-region 03 are sequentially moved away from the camera. For example, the width of the first sub-region 01 in the first direction Y is equal to the width of the third sub-region 03 in the first direction Y, and the width of the second sub-region 02 in the first direction Y is greater than the width of the first sub-region 01 in the first direction Y. The widths of the first sub-region 01, the second sub-region 02, and the third sub-region 03 in the second direction X are equal. The second direction X is the width direction of the image to be detected.

[0121] Continue as Figure 2 As shown, in practical application scenarios, the closer a person is to the camera, the larger the area occupied by their body in the frame; the farther away a person is from the camera, the smaller the area occupied by their body in the frame. Therefore, the number of multiple first sub-images 011 located in the first sub-region 01 is greater than the number of multiple second sub-images 012 located in the second sub-region 02, and the number of multiple second sub-images 012 located in the second sub-region 02 is greater than the number of multiple third sub-images 013 located in the third sub-region 03. For example, the width of the third sub-image 013 in the second direction X is greater than the width of the second sub-image 012 in the second direction X, and the width of the second sub-image 012 in the second direction X is greater than the width of the first sub-image 011 in the second direction X. At least some of the first sub-images 011 have equal widths in the second direction X. For example, among a plurality of first sub-images 011 arranged side by side along the second direction X, except for the first and last first sub-images 011 arranged side by side along the second direction X, the remaining first sub-images 011 have the same width in the second direction X; the first and last first sub-images 011 have the same width in the second direction X. At least some of the second sub-images 012 have equal widths in the second direction X. For example, among a plurality of second sub-images 012 arranged side by side along the second direction X, except for the first and last second sub-images 012 arranged side by side along the second direction X, the remaining second sub-images 012 have the same width in the second direction X; the first and last second sub-images 012 have the same width in the second direction X. At least some of the third sub-images 013 have equal widths in the second direction X. For example, each third sub-image 013 has the same width in the second direction X.

[0122] Furthermore, for locations relatively far from the camera in the deployment scenario (e.g., the location corresponding to the first sub-image 011 and the second sub-image 012), each first sub-image 011 overlaps at least partially in the second direction X; each second sub-image 012 overlaps at least partially in the second direction X. This further addresses the problem of inaccurate identification at the stitching points of adjacent target sub-images and improves the recognition accuracy of target sub-images. For locations relatively close to the camera in the deployment scenario (e.g., the location corresponding to the third sub-image 013), to simplify the image segmentation method, each third sub-image 013 can be set to have no overlap in the second direction X. Since target objects closer to the camera occupy a larger proportion of the image than target objects farther from the camera, the lack of overlap between adjacent third sub-images 013 in the second direction X will not affect the recognition accuracy of the third sub-image 013, or the impact is small and negligible.

[0123] Furthermore, the width ratio of the second sub-region 02 to the first sub-region 01 in the first direction Y is 2:1; the width ratio of the third sub-image 013 to the second sub-image 012 in the second direction X is 3:2, and the width ratio of the second sub-image 012 to the first sub-image 011 in the second direction X is 4:3.

[0124] Of course, depending on the different application scenarios and the size of the image captured by the camera, different sub-regions can be divided, and different width ratios can be set. This disclosure does not impose specific limitations on this.

[0125] Furthermore, among the plurality of first sub-images 011 arranged side-by-side along the second direction X, except for the first and last first sub-images 011 arranged side-by-side along the second direction X, the widths of the remaining first sub-images 011 in the second direction X are all the same; the widths of the first and last first sub-images 011 in the second direction X are the same and smaller than the widths of the remaining first sub-images 011 in the second direction X; the ratio of the overlap width of two adjacent first sub-images 011 in the second direction X to the width of the remaining first sub-images 011 in the second direction X is 1:10. Among the plurality of second sub-images 012 arranged side-by-side along the second direction X, except for the first and last second sub-images 012 arranged side-by-side along the second direction X, the widths of the remaining second sub-images 012 in the second direction X are all the same; the widths of the first and last second sub-images 012 in the second direction X are the same and smaller than the widths of the remaining second sub-images 012 in the second direction X; the overlap width of two adjacent second sub-images 012 in the second direction X is 1:10. The ratio of the overlap width on X to the width of the remaining second sub-image 012 on the second direction X is 1:10; among the multiple third sub-images 013 arranged side by side along the second direction X, except for the first and last third sub-images 013 arranged side by side along the second direction X, the widths of the remaining third sub-images 013 on the second direction X are all the same; the widths of the first and last third sub-images 013 on the second direction X are the same and smaller than the widths of the remaining third sub-images on the second direction X; the ratio of the overlap width of two adjacent third sub-images 013 on the second direction X to the width of the remaining third sub-images 013 on the second direction X is 1:10; the ratio of the overlap width of the first sub-image 011 and the second sub-image 012 adjacent on the first direction Y to the width of the second sub-image 012 on the first direction Y is 1:10; the ratio of the overlap width of the third sub-image 013 and the second sub-image 012 adjacent on the first direction Y to the width of the second sub-image 012 on the first direction Y is 1:10.

[0126] In conjunction with the above embodiments, under normal viewing conditions, the image to be detected is divided into multiple target sub-images using the above division method, specifically including multiple first sub-images 011, multiple second sub-images 012, and multiple third sub-images 013. This refines the image to be detected and improves the accuracy of subsequent detection results while keeping the power consumption of the control system within a reasonable range.

[0127] In some embodiments, for step S11, a frame difference method is used to detect whether a moving target exists in the shooting scene of the shooting device. This moving target can be considered a moving object or person. Specifically, the processing module acquires multiple consecutive video frames captured by the shooting device; based on these multiple consecutive video frames, the frame difference method is used to determine whether a moving target exists in the shooting scene of the shooting device, and the determination result is stored in an external storage module. For example, if a moving target exists in the shooting scene, the position of the target's image will be different in different video frames. The frame difference algorithm performs pixel value difference operations on two or three consecutive video frames in time, subtracting the pixel values ​​corresponding to different frames, and determining the absolute value of the grayscale value. When the absolute value exceeds a certain threshold, it can be determined that a moving target exists. If a moving target exists, the video frames captured by the shooting device are used as the images to be detected.

[0128] In this embodiment, before using the model for detection, the frame difference method is used to initially screen video frames to find video frames that are likely to contain the target object, which simplifies the subsequent model processing and improves detection efficiency.

[0129] In some embodiments, Figure 3 This is a schematic diagram illustrating an exemplary training process for a target neural network model provided in an embodiment of the present disclosure, such as... Figure 3 As shown, the processing module includes the training process that executes steps S31 to S34, wherein:

[0130] S31. Obtain the first training dataset and the second training dataset.

[0131] Here, the first training dataset can be a labeled public dataset and / or collected video data. Alternatively, the first training dataset can also be a dataset further enriched with the sample types based on a labeled public dataset and / or collected video data, making the resulting training images more diverse. Methods for enriching the sample types can be found in steps S311 to S315 below; detailed processes are not described here. The first training dataset includes multiple training images, each with its corresponding sample label.

[0132] It should be noted that in order to improve the model training accuracy, a large amount of sample data is required. However, it is not easy to produce samples. Therefore, in order to obtain a large number of samples in the early stage, this embodiment obtains a first training dataset with noise as the training set for the teacher's machine learning model.

[0133] Here, the second training dataset can be obtained by filtering the first training dataset; that is, it is a carefully selected, noise-free dataset for accurate classification. The second training dataset includes a portion of the training images from the first training dataset. Specifically, the second training dataset includes multiple training images, each with its corresponding sample label.

[0134] In some embodiments, the first training dataset includes noisy training images. For example, positive samples are pre-defined as images containing human head features, and negative samples are images without human head features (e.g., images with no human features at all, and images containing partial human limb features (excluding head features)). Training images in the first training dataset that contain the feature of a target object are labeled as positive samples; those that do not contain the feature of a target object are labeled as negative samples. For example, a training image in the first training dataset contains human features, so this training image is labeled as a positive sample image; however, this training image does not contain head features, therefore this training image is noisy.

[0135] One approach is to manually select accurately classified training images from a noisy first training dataset as samples for the second training dataset. For example, training images containing the human's shoulders and above, including head features, can be used as positive samples in the second training dataset, while training images containing other limb features (but no head features) can be used as negative samples. Alternatively, traditional image recognition techniques can be employed to detect whether training images in the first training dataset contain features of the human's shoulders and above, including the head. If such features are found, the image is selected as a positive sample in the second training dataset; otherwise, it is selected as a negative sample.

[0136] S32. Train the teacher machine learning model to be trained based on the first training dataset to obtain the preliminarily trained teacher machine learning model.

[0137] For example, a training image A from the first training dataset is input into the teacher machine learning model to be trained, and an output result A corresponding to the training image A is obtained. The output result A indicates whether a target object exists in the training image A. If the sample label of the training image A indicates the presence of a target object, a third loss function is determined based on the output result A and the sample label of the training image A. The above process is repeated to iterate through the training images in the first training dataset and determine the third loss function in order to continuously train the teacher machine learning model to be trained, and finally obtain a preliminarily trained teacher machine learning model.

[0138] S33. Train the pre-trained teacher machine learning model based on the second training dataset to obtain the trained teacher machine learning model.

[0139] For example, a training image B from the second training dataset is input into the initially trained teacher machine learning model to obtain an output result B corresponding to the training image B. This output result B indicates whether a target object exists in the training image B. If the sample label of the training image B indicates the presence of a target object, a fourth loss function is determined based on the second output result and the sample label of the training image B. The above process is repeated, iterating through the training images in the second training dataset to determine the fourth loss function, thereby continuously training the initially trained teacher machine learning model, and finally obtaining the trained teacher machine learning model.

[0140] S34. Based on the second training dataset and the trained teacher machine learning model, the knowledge distillation training method is used to train the student machine learning model to be trained, and the trained student machine learning model is used as the target neural network model.

[0141] Here, the trained teacher machine learning model is the teacher model in the knowledge distillation process; the student machine learning model to be trained is the student model in the knowledge distillation process.

[0142] Figure 4 This is a schematic diagram of an exemplary network architecture for knowledge distillation provided in this disclosure, such as... Figure 4 As shown, there are teacher and student models. The teacher model consists of m network layers, and the student model consists of n network layers. Here, the teacher and student models can be homogeneous or heterogeneous networks.

[0143] The process of training the student model also involves using the hard labels in the second training dataset (i.e., the sample labels carried by the training images) to train the teacher model, and then using the soft labels obtained in the teacher model (i.e., the output of the softmax layer in the teacher model) in combination with the hard labels in the second training dataset to determine the loss function of the student model.

[0144] Continue as Figure 4 As shown, the output of the m-th layer of the teacher model is processed by a softmax layer to obtain the first output result; the output of the n-th layer of the student model is processed by a softmax layer to obtain the second output result. A loss function is constructed based on the first output result, the second output result, and the sample labels.

[0145] This embodiment employs a knowledge distillation method to transform a larger teacher model into a smaller student model while retaining performance close to that of the teacher model, thereby addressing the issue of insufficient hardware for deploying the target neural network model at edge segments.

[0146] In some embodiments, due to the hardware limitations of in-memory computing chips in the smart wake-up application scenario of this disclosure, a lightweight target neural network model needs to be designed to complete the target detection task. The target neural network model architecture is also the network architecture of the student model.

[0147] Figure 5 A schematic diagram of an exemplary student model provided in this disclosure is shown below. Figure 5 As shown, it includes five convolutional units (blocks) and one fully connected layer (fc); each convolutional unit (block) includes two convolutional layers. For example, in the two convolutional layers of the first convolutional unit (block1), the second convolutional unit (block2), and the third convolutional unit (block3), the first convolutional layer (i.e., the convolutional layer that processes the data first) uses a 3×3 kernel for convolution, with a stride of 2 and a padding value of 1; the second convolutional layer (i.e., the convolutional layer that receives the output data from the first convolutional layer) also uses a 3×3 kernel for convolution, with a stride of 1 and a padding value of 1. For the two convolutional layers in the fourth convolutional unit block4 and the fifth convolutional unit block5, the first convolutional layer (that is, the convolutional layer that processes the data first) uses a 5×5 convolutional kernel, a stride of 2, and a padding value of 1; the second convolutional layer (that is, the convolutional layer that receives the output data of the first convolutional layer) uses a 3×3 convolutional kernel, a stride of 1, and a padding value of 1.

[0148] The aforementioned target neural network model meets the lightweight requirement. When migrated to an in-memory computing chip, it offers faster response speed and lower power consumption during its application.

[0149] In some embodiments, the second training dataset includes multiple training images labeled with sample labels.

[0150] like Figure 4 As shown, for step S34, the specific process of training the target neural network model includes the following steps S341 to S346, wherein:

[0151] S341. Input the training image into the trained teacher machine learning model to obtain the first output result of the trained teacher machine learning model.

[0152] For example, the training image B in the second training dataset is input into the teacher machine learning model to be trained, and the first output result corresponding to the training image B is obtained. The first output result is the result output by the teacher machine learning model to be trained, indicating whether the training image B contains a target object.

[0153] S342. Input the training image into the student machine learning model to be trained, and obtain the second output result of the student machine learning model to be trained.

[0154] Continuing the previous example, the training image B in the second training dataset is input into the student machine learning model to be trained, and the second output result corresponding to the training image B is obtained. The second output result is the result output by the student machine learning model to be trained, indicating whether the training image B contains a target object.

[0155] S343. Determine the first loss function based on the first output result and the second output result.

[0156] Following steps S341 and S342, the first loss function can be determined based on the mean square error between the first and second output results, or based on the first and second output results, using the KLD loss function algorithm (i.e., KLDloss).

[0157] S344. Determine the second loss value based on the second output result and the sample labels of the training images.

[0158] The second loss function can be determined based on the overlap between the second output (i.e., the predicted label of training image B predicted by the student machine learning model to be trained) and the sample labels of the training image (i.e., the true hard labels).

[0159] S345. Based on the first loss function and the second loss function, the weighted loss function is obtained.

[0160] The first loss function and the second loss function are weighted according to preset weights to obtain a weighted loss function. Here, the preset weights can be determined based on experience, and this embodiment does not impose specific limitations.

[0161] Obtain the proportion coefficients of the first loss function and the second loss function. In this embodiment, since the weighted loss function is composed of the first loss function and the second loss function, determining one proportion coefficient (α) allows us to determine the other proportion coefficient (i.e., 1-α). Then, the first loss function L1 and the second loss function L2 are weighted and summed according to the proportion coefficients to determine the weighted loss function. The weighted loss function L, based on the above description, can be specifically expressed as: L=α×L1+(1-α)×L2.

[0162] S346. Adjust the parameters of the student machine learning model to be trained according to the weighted loss value until the weighted loss value converges, and obtain the trained student machine learning model as the target neural network model.

[0163] In some embodiments, the first training dataset can be a dataset further enriched with more sample types based on the original dataset. The specific processing module includes the step of generating diverse training images, as detailed in steps S311 to S317 below:

[0164] S311. Obtain the original dataset.

[0165] The original dataset includes multiple initial sample images.

[0166] For example, the original dataset might be a labeled public dataset and / or collected video data. Here, "labeling" refers to labeling the first bounding boxes of target objects present in the initial sample images.

[0167] The initial sample images in the original dataset may or may not contain target objects. For those initial sample images that do not contain target objects (meaning they have no features of target objects at all), there is no first reference box, or the position information of the first reference box is empty.

[0168] S312. Identify the target object in the initial sample image. If the target object exists, proceed to step S313. If the target object does not exist, proceed to step S317.

[0169] For example, object detection networks such as RetinaFace or YOLOv5 neural networks can be used to identify target objects in initial sample images.

[0170] Here, there is no initial sample image of the target object, that is, an initial sample image that has no features of the target object at all. This initial sample image is labeled with a negative sample, and the initial sample image labeled with the negative sample is used as a negative sample (training image) in the first training dataset.

[0171] It should be noted that for an initial sample image containing a target object, the following steps S313 to S316 are included.

[0172] S313. Determine the first reference frame containing the target object.

[0173] If a target object is determined to exist in the initial sample image, a first reference box is marked around the region containing the target object, so that the target object is enclosed by the first reference box. The position information of the first reference box indicates the position of the target object in the initial sample image.

[0174] S314. Update the position of the first reference box according to the position information of the first reference box to obtain the second reference box.

[0175] Specific update methods may include at least one of the following: randomly moving a specific coordinate point, randomly scaling, or randomly cropping.

[0176] Example 1: For randomly moving detection points, the specific process includes: based on the position information of the first reference box (i.e., the position coordinates in the initial sample image), the coordinates of a specific coordinate point pre-set in the first reference box are randomly moved to obtain a new contour coordinate point, forming a second reference box.

[0177] Example 2: For random scaling, the specific process includes: based on the position information of the first reference box (i.e., the position coordinates in the initial sample image), keeping the center point coordinates of the first reference box unchanged, randomly scaling the width and height of the first reference box to obtain new contour coordinate points, forming a second reference box.

[0178] Example 3: For random cropping, the specific process includes: based on the position information of the first reference frame (i.e., the position coordinates in the initial sample image), a new center point is determined according to the first preset center point range; while keeping the new center point coordinates unchanged, the width and height of the first reference frame are randomly adjusted according to the preset cropping range to obtain new contour coordinate points, forming a second reference frame.

[0179] In other embodiments, the position of the first reference frame is updated, and the specific update method can be combined with the above examples. For example, executing Examples 1 and 2 sequentially, specifically: based on the position information of the first reference frame, the coordinates of a pre-set specific coordinate point in the first reference frame are randomly moved to obtain a new contour coordinate point, forming a transition frame; based on the position information of the transition frame, while keeping the center point coordinates of the transition frame unchanged, the width and height of the transition frame are randomly scaled to obtain a new contour coordinate point, forming a second reference frame. As another example, executing Examples 1 and 3 sequentially, specifically: based on the position information of the first reference frame, the coordinates of a pre-set specific coordinate point in the first reference frame are randomly moved to obtain a new contour coordinate point, forming a transition frame; based on the position information of the transition frame, a new center point is determined according to a preset center point range; and while keeping the new center point coordinates unchanged, the width and height of the transition frame are randomly adjusted according to a preset clipping range to obtain a new contour coordinate point, forming a second reference frame. For example, if we execute Examples 2 and 3 in sequence, specifically: based on the position information of the first reference frame, while keeping the center point coordinates of the first reference frame unchanged, we randomly scale the width and height of the first reference frame to obtain new outline coordinate points and form a transition frame; based on the position information of the transition frame, we determine a new center point according to a preset center point range; and while keeping the new center point coordinates unchanged, we randomly adjust the width and height of the transition frame according to a preset clipping range to obtain new outline coordinate points and form a second reference frame.

[0180] In other embodiments, the position of the first reference frame is updated. Specifically, the update method can be a combination of all the above examples (Example 1, Example 2, and Example 3), including the following steps S314-1 to S314-3, wherein:

[0181] S314-1. Based on the position information of the first reference frame, move a specific coordinate point in the first reference frame to obtain the third reference frame.

[0182] Here, the specific coordinate points can be vertices of the pre-defined outer contour of the first reference frame, such as vertex 1 at the top left, vertex 2 at the top right, vertex 3 at the bottom left, and vertex 4 at the bottom right. In this embodiment, vertex 1 at the top left and vertex 4 at the bottom right can be selected as specific coordinate points for diversified updates of the reference frame.

[0183] For example, the coordinates of vertex 1 in the top left corner are (lt... x ,lt y The coordinates of vertex 4 in the lower right corner are (rb) x ,rb y Let the coordinates of vertex 1 be (lt). x ,lt y The random shift value is δ x δ yThe coordinates of vertex 4 (rb) x ,rb y The random shift value σ x , σ y The values ​​can be positive or negative. After random movement, the coordinates of the new vertex 1 are (lt1) x ,lt1 y ), that is:

[0184] lt1 x =lt x +δ x

[0185] lt1 y =lt y +δ y

[0186] The coordinates of the new vertex 4 are (rb1) x ,rb1 y ), that is:

[0187] rb1 x =rb x +σ x

[0188] rb1 y =rb y +σ y

[0189] S314-2. Based on the position information of the third reference frame, take the center point of the third reference frame as the center, and adjust the width and height of the third reference frame according to the preset scaling factor to obtain the fourth reference frame.

[0190] Continuing the previous example, the coordinates of the center point of the third reference frame are (c x ,c y The third reference frame has a width of w and a height of h; the preset scaling factors include a width scaling factor s. x and height scaling factor s y Adjust the width and height of the third reference frame according to the preset scaling factor to obtain the coordinates (lt2) of vertex 1 at the top left corner of the fourth reference frame. x ,lt2 y The coordinates of vertex 4 at the bottom right corner of the fourth reference frame (rb2) x ,rb2 y ),in:

[0191] c x =(lt1) x +rb1 x ) / 2, c y =(lt1) y +rb1y ) / 2

[0192] w = (rb1) x -lt1 x ), h=(rb1 y -lt1 y )

[0193] lt2 x =c x -s x ×w / 2

[0194] lt2 y =c y -s y ×h / 2

[0195] rb2 x ==c x +s x ×w / 2

[0196] rb2 y =c y +s y ×h / 2

[0197] S314-3. Based on the position information of the fourth reference frame, determine a new center point according to the first preset center point range; and adjust the width and height of the fourth reference frame according to the new center point and the preset cutting range to obtain the second reference frame.

[0198] Here, the width of the fourth reference frame is w′=(rb2) x -lt2 x ), height h' = (rb2) y -lt2 y The first preset center point range is also the randomly generated center point (cnew). x cnew y The range in which cnew is located, where x ∈[0.45×w', 0.55×w'], cnew y ∈[0.45×h′,0.55×h']. The preset clipping range is also the range of the width wnew and height hnew of the randomly generated fourth reference frame, where wnew∈[w'×0.9,w′],hnew∈[h'×0.9,h'].

[0199] According to the new center point (cnew) x cnew y ) and preset clipping range, adjust the width and height of the fourth reference frame to obtain a new width wnew' and height hnew', and then obtain the new outline coordinate points, the upper left corner coordinates (lt3)x ,lt3 y ), bottom right corner coordinates (rb3) x ,rb3 y ).

[0200] The above update method can obtain diverse training images, which can improve the training accuracy of subsequent models.

[0201] S315. Determine the overlap between the second reference frame and the first reference frame based on the position information of the second reference frame and the position information of the first reference frame.

[0202] The Intersection over Union (IOU) characterizes the degree of overlap between the first and second bounding boxes. A larger IOU value indicates a larger overlap area between the second and first bounding boxes, and a greater proportion of the target object within the second bounding box; conversely, a smaller IOU value indicates a smaller overlap area.

[0203] S316. Based on the comparison result between the overlap and the first preset threshold, label the sample images of the initial sample images located in the second reference box, and use the labeled sample images as training images in the first training dataset.

[0204] Here, the first preset threshold can be set empirically. When the overlap is greater than or equal to the first preset threshold, the corresponding portion of the sample images is determined to be positive samples and labeled as positive samples. It should be noted that although the overlap is greater than or equal to the first preset threshold, the overlap only reflects the presence of a portion of the target object in the second reference frame. The features of this portion of the target object may be features that can clearly identify the user, such as the head, or features that cannot clearly identify the user, such as limbs.

[0205] Generating diverse training images can enhance the richness of the first training dataset, thereby improving the model training accuracy. Diversity refers to the variety of features of the segmented target objects in the training images. For example, there may be a large number of training images containing head features, some containing only limb features, some containing torso features, and some containing full-body features of the target object, etc. In some embodiments, following step S315, each initial sample image corresponds to its own overlap (or each training image corresponds to its own overlap).

[0206] In specific implementation, when the overlap is greater than or equal to the first preset threshold, a first preset range can be generated according to a first preset probability, a second preset range can be generated according to a second preset probability, a third preset range can be generated according to a third preset probability, and a fourth preset range can be generated according to a fourth preset probability; the first preset probability is greater than the second preset probability, the second preset probability is greater than the third preset probability, and the third preset probability is greater than or equal to the fourth preset probability; the sum of the first preset probability, the second preset probability, the third preset probability, and the fourth preset probability is 1.

[0207] When the overlap is within a first preset range, the sample images located in the second reference box in the initial sample image are identified and labeled as positive sample labels, so that the ratio between the number of training images with overlap within the first preset range and the total number of training images with positive sample labels in the first training dataset is the first preset probability.

[0208] When the overlap is within a second preset range, the portion of the sample images located in the second reference box in the initial sample image are identified and labeled as positive sample labels, so that the ratio between the number of training images with overlap within the second preset range and the total number of training images with positive sample labels in the first training dataset is the second preset probability.

[0209] When the overlap is within a third preset range, the sample images located in the second reference box in the initial sample image are identified and labeled as positive sample labels, so that the ratio between the number of training images with overlap within the third preset range and the total number of training images with positive sample labels in the first training dataset is the third preset probability.

[0210] When the overlap is within the fourth preset range, the sample images located in the second reference box in the initial sample image are identified and labeled as positive sample labels, so that the ratio between the number of training images with the overlap within the fourth preset range and the total number of training images with positive sample labels in the first training dataset is the fourth preset probability.

[0211] Among them, the overlap within the first preset range is less than the overlap within the second preset range, the overlap within the second preset range is less than the overlap within the third preset range, and the overlap within the third preset range is less than the overlap within the fourth preset range.

[0212] For example, the first preset threshold is 0.6; if the overlap is greater than or equal to 0.6, then the first preset range of 0.6 to 0.7 (excluding 0.7) is selected with a 50% probability (first preset probability), the range of 0.7 to 0.8 (excluding 0.8) is selected with a 30% probability (second preset probability), the range of 0.8 to 0.9 (excluding 0.9) is selected with a 10% probability (third preset probability), and the range of 0.9 to 1.0 is selected with a 10% probability (fourth preset probability). When the overlap is within a first preset range, the sample images located within the second reference box are labeled as positive samples, so that the training images with the overlap within the first preset range account for 50% of the total number of training images with positive samples; when the overlap is within a second preset range, the sample images located within the second reference box are labeled as positive samples, so that the training images with the overlap within the second preset range account for 30% of the total number of training images with positive samples; when the overlap is within a third preset range, the sample images located within the second reference box are labeled as positive samples, so that the training images with the overlap within the third preset range account for 10% of the total number of training images with positive samples; when the overlap is within a fourth preset range, the sample images located within the second reference box are labeled as positive samples, so that the training images with the overlap within the fourth preset range account for 10% of the total number of training images with positive samples.

[0213] In this embodiment, the proportion of positive samples in the first training dataset is preset. A first preset range of 0.6–0.7 is selected with a 50% probability, meaning that positive samples with an IOU in the range [0.6, 0.7) account for 50% of the total positive sample data. A second preset range of 0.7–0.8 is selected with a 30% probability, meaning that positive samples with an IOU in the range [0.7, 0.8) account for 30% of the total positive sample data. A third preset range of 0.8–0.9 is selected with a 10% probability, meaning that positive samples with an IOU in the range [0.8, 0.9) account for 10% of the total positive sample data. A fourth preset range of 0.9–1.0 is selected with a 10% probability, meaning that positive samples with an IOU in the range [0.9, 1.0] account for 10% of the total positive sample data. Since the target neural network model in this embodiment uses a small model, and the target image to be detected is segmented into multiple target sub-images, the training image will also be a segmented image from the entire image to improve the model training accuracy. In the process of generating the first training dataset, as many segmented images of the target object as possible (i.e., partial sample images) are generated as training images.

[0214] Of course, the above division of the proportion of positive samples in the first training dataset is only one embodiment. Other proportions can be set as long as the condition of rich training samples is met. This embodiment is not limited.

[0215] S317. Label the initial sample images with negative sample labels, and use the initial sample images with negative sample labels as training images in the first training dataset.

[0216] This step involves labeling the initial sample images that do not contain the target object with negative sample labels.

[0217] In some embodiments, if the overlap between the first and second reference frames is less than a third preset threshold, the portion of the initial sample image located within the second reference frame is labeled with a negative sample label. Here, the second preset threshold can be set empirically, and is less than the first preset threshold. For example, the third preset threshold is 0.1. Therefore, the portion of the sample image labeled with a negative sample label does not necessarily mean that there are no features of the target object; it simply means that the feature is very small and can be ignored, so it can be used as a negative sample.

[0218] Of course, in addition to the above-mentioned methods of creating negative samples, a portion of the sample image that does not contain the target object can be extracted from the initial sample image as a negative sample; and / or, a portion of the sample image that contains only a small portion of the target object features can be extracted from the initial sample image as a negative sample, etc. The embodiments disclosed herein do not impose specific limitations on this.

[0219] In some embodiments, after S312 is executed, another method for generating negative samples is further included. Specifically: based on the size of the initial sample image, a target center point is determined in the initial sample image according to the second preset center point range, and a fifth reference box is determined with the target center point as the center; based on the position information of the fifth reference box and the position information of the first reference box, the overlap between the fifth reference box and any first reference box in the initial sample image is determined; if the overlap between the fifth reference box and any first reference box in the initial sample image is less than or equal to the second preset threshold, then the sample images of the initial sample image located in the fifth reference box are labeled as negative samples, and the sample images labeled as negative samples are used as training images in the first training dataset.

[0220] Here, the second preset threshold can be selected from a threshold in the range of 0 to 0.1; by randomly setting the target center point on the initial sample image, a fifth reference box is defined; since the part of the sample image corresponding to the first reference box contains the target object, if the overlap between the fifth reference box and any first reference box is less than or equal to the second preset threshold, it means that the overlap between the fifth reference box and any first reference box is very small, or even 0. Then, the part of the sample image corresponding to the fifth reference box can be considered to not contain the target object. Therefore, the part of the sample image at the position corresponding to the fifth reference box can be used as a negative sample training image.

[0221] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0222] It should be noted that, in addition to detecting the target object and waking up the device for operation as mentioned in the embodiments of this disclosure, portable smart terminal application services in smart homes and wearable devices can also adopt this energy-efficient in-memory computing approach, reducing significant costs. Moreover, with the development of in-memory computing technology, it can support larger deep learning models and more extensive and universal services. Furthermore, since different in-memory computing platforms support different functions and operators, this disclosure can design and build networks according to actual hardware platforms and is not limited to the target neural network model proposed in this disclosure.

[0223] In addition, this disclosure also provides another control method for a monitoring system. The execution entity of this control method can be the monitoring system itself, which includes a memory-to-computing system and a terminal. The memory-to-computing system includes a processing module, a memory-to-computing module, and an external storage module. The processing module is, for example, the main processor CPU of the memory-to-computing system; the memory-to-computing module is, for example, a memory-to-computing core integrating Flash or SRAM; and the external storage module is, for example, an external storage module DRAM, which offers faster response speeds and reduces power consumption of the monitoring system. The control method for the monitoring system includes:

[0224] The processing module acquires the image to be detected; based on the shooting angle of the shooting device that captured the image, the image to be detected is divided into multiple target sub-images and stored in the external storage module. It should be noted that the specific execution process of the image division module in this embodiment can be found in steps S11 and S12 of the control method of the monitoring system described above, and the repeated parts will not be described again.

[0225] For each target sub-image among multiple target sub-images, the in-memory computing module uses a pre-trained target neural network model to process the target sub-image, obtain the recognition result of the target sub-image, and store the recognition result in an external storage module. It should be noted that the specific execution process of the in-memory computing module in this embodiment can be found in step S13 of the control method of the aforementioned monitoring system; repeated parts will not be described again. The in-memory computing module integrates a target neural network model.

[0226] Based on the recognition results corresponding to each target sub-image, the processing module obtains the detection result of whether a target object exists in the image to be detected, and sends the detection result to the terminal. It should be noted that the specific execution process of the processing module in this embodiment can be referred to step S13 in the control method of the above-mentioned monitoring system, and the repeated parts will not be described again.

[0227] The terminal determines the display status based on the detection results.

[0228] This disclosure provides a control method for a monitoring system. It employs a neural network to detect images captured by a camera and feeds the detection results back to the terminal. The terminal determines its current display state based on these results. For example, if no target object is detected for an extended period, the system can enter a sleep state to save power. If, after entering sleep mode, the detection result indicates the presence of a target object, the system can automatically restart and resume normal display. Furthermore, in the intelligent wake-up application scenario, this disclosure, while saving power, combines neural network image recognition technology. Based on the shooting angle, the image to be detected captured by the camera is rationally divided into multiple target sub-images. Each target sub-image is identified separately, and the recognition results are integrated to obtain the detection result, improving the accuracy of target object recognition in the image and thus more precisely controlling the terminal's wake-up timing. Additionally, utilizing an integrated storage and computing module, a system architecture that combines storage and computing, supports complex data operations in the neural network, greatly avoiding energy loss caused by data transfer and reducing system power consumption.

[0229] In some embodiments, the terminal includes a display module and a main control module; the terminal determines the display state based on the detection result, including: when the main control module determines that the detection result indicates that a target object exists in the image to be detected, it sends a wake-up request to the display module and resets the sleep timer to respond to the terminal request to receive the detection result sent by the processing module; when the main control module determines that the detection result indicates that no target object exists in the image to be detected, and the current system time is greater than the time since the last wake-up of the display module is greater than a preset duration, it sends a sleep request to the display module and controls itself to enter a sleep state; the display module responds to the wake-up request and determines that the display state is the normal display of the captured image from the shooting device; responds to the sleep request and determines that the display state is sleep.

[0230] Here, the preset duration can be set based on experience. It should be noted that setting a reasonable preset duration ensures that the terminal enters sleep mode appropriately in actual application scenarios. If the preset duration is too short, it will lead to frequent switching between sleep and wake-up states, which is not conducive to reducing power consumption. If the preset duration is too long, it will not minimize the terminal's power consumption.

[0231] Here, the display module can be, for example, a display screen, and the main control module can be, for example, a system-on-a-chip (SOC). The main control module and the image segmentation module can communicate using protocols such as UART.

[0232] Figure 6a An exemplary terminal workflow diagram provided for embodiments of this disclosure, such as... Figure 6aAs shown, the main control module performs the following steps: S61, Start system initialization; S62, Send a start command to the image segmentation module and receive a response confirmation from the image segmentation module, then execute S63 sequentially; S63, Reset the sleep timer; S64, Poll the serial port to see if the target object is detected and send a response confirmation to the image segmentation module; If yes, return to step S63; If no, execute S65; S65, Determine if the preset time has been exceeded; If yes, execute S66; If no, return to S64; S66, The main control SOC enters sleep mode and sends a sleep request to the display module, then executes S67 sequentially; S67, Poll the serial port to see if the target object is detected and send a response confirmation to the image segmentation module; If yes, execute S68; If no, repeat S67; S68, Wake up the main control SOC, send a wake-up request to the display module, and return to step S63.

[0233] Figure 6b An exemplary in-memory computing system workflow diagram provided for embodiments of this disclosure, such as... Figure 6b As shown, the in-memory computing system proceeds as follows: S601, system initialization begins; S602, the processing module polls the serial port to identify the start command sent by the main control module and sends a response confirmation to the main control module, then executes S603 sequentially; S603, the processing module performs motion target detection, that is, uses the frame difference method to determine whether there is a moving target in the shooting scene of the shooting device; if yes, execute step 604, that is, start the in-memory computing module to perform target object recognition function, the detailed process is described in steps S12 to S14 above for determining the detection result; if no, repeat step 603. S604, detect whether there is a target object and send the detection result to the terminal.

[0234] In some embodiments, the monitoring system also includes a camera, and the image segmentation module can acquire data from the low-power camera via protocols such as SPI.

[0235] The following example illustrates the control process of the monitoring system implemented by the in-memory computing system. For detailed process, please refer to the description of the specific embodiment of the control method of the monitoring system with the in-memory computing system as the execution subject. Repeated parts will not be repeated.

[0236] In some embodiments, the processing module divides the image to be detected into multiple target sub-images according to the shooting angle of the shooting device, including:

[0237] When the shooting angle is within the preset angle range, the image to be detected is divided into a first sub-region, a second sub-region, and a third sub-region set sequentially along a first direction, and the first sub-region and the third sub-region partially overlap with the second sub-region; wherein, the width of the first sub-region in the first direction is equal to the width of the third sub-region in the first direction, and the width of the second sub-region in the first direction is greater than the width of the first sub-region in the first direction; the first direction is the height direction of the image to be detected;

[0238] The portion of the image to be detected located in the first sub-region is divided into multiple first sub-images arranged side by side along the second direction, at least some of the first sub-images having equal width in the second direction; the second direction is the width direction of the image to be detected.

[0239] The portion of the image to be detected located in the second sub-region is divided into multiple second sub-images arranged side by side along the second direction, and at least some of the second sub-images have equal widths in the second direction;

[0240] The portion of the image to be detected located in the third sub-region is divided into multiple third sub-images arranged side by side along the second direction, at least some of the third sub-images having equal width in the second direction; the target sub-image includes a first sub-image, a second sub-image, and a third sub-image;

[0241] The width of the third sub-image in the second direction is greater than the width of the second sub-image in the second direction, and the width of the second sub-image in the second direction is greater than the width of the first sub-image in the second direction.

[0242] In some embodiments, the first sub-images overlap at least partially in the second direction; the second sub-images overlap at least partially in the second direction; and the third sub-images do not overlap in the second direction.

[0243] In some embodiments, the width ratio of the second sub-region to the first sub-region in the first direction is 2:1;

[0244] The width ratio of the third sub-image to the second sub-image in the second direction is 3:2, and the width ratio of the second sub-image to the first sub-image in the second direction is 4:3.

[0245] In some embodiments, among a plurality of first sub-images arranged side by side along the second direction, except for the first and last first sub-images arranged side by side along the second direction, the widths of the remaining first sub-images in the second direction are all the same; the widths of the first and last first sub-images in the second direction are the same and smaller than the widths of the remaining first sub-images in the second direction; the ratio of the overlap width of two adjacent first sub-images in the second direction to the width of the remaining first sub-images in the second direction is 1:10.

[0246] Among the multiple second sub-images arranged side by side along the second direction, except for the first and last second sub-images arranged side by side along the second direction, the widths of the remaining second sub-images are all the same in the second direction; the widths of the first and last second sub-images are the same in the second direction and are smaller than the widths of the remaining second sub-images in the second direction; the ratio of the overlap width of two adjacent second sub-images in the second direction to the width of the remaining second sub-images in the second direction is 1:10.

[0247] Among the multiple third sub-images arranged side by side along the second direction, except for the first and last third sub-images arranged side by side along the second direction, the widths of the remaining third sub-images are all the same in the second direction; the widths of the first and last third sub-images are the same in the second direction and are smaller than the widths of the remaining third sub-images in the second direction; the ratio of the overlap width of two adjacent third sub-images in the second direction to the width of the remaining third sub-images in the second direction is 1:10.

[0248] The first and second sub-images adjacent in the first direction have an overlap width in the first direction that is 1:10 to the width of the second sub-image in the first direction; the third and second sub-images adjacent in the first direction have an overlap width in the first direction that is 1:10 to the width of the second sub-image in the first direction.

[0249] In some embodiments, the target neural network model is trained by the steps performed by the following processing modules:

[0250] Obtain the first training dataset and the second training dataset; the second training dataset is obtained by filtering the first training dataset.

[0251] The teacher machine learning model to be trained is trained based on the first training dataset to obtain the preliminarily trained teacher machine learning model.

[0252] The teacher machine learning model is initially trained based on the second training dataset to obtain the fully trained teacher machine learning model.

[0253] Based on the second training dataset and the trained teacher machine learning model, the knowledge distillation training method is used to train the student machine learning model to be trained, and the trained student machine learning model is used as the target neural network model.

[0254] In some embodiments, the second training dataset includes multiple training images labeled with sample labels;

[0255] The processing module, based on the second training dataset and the trained teacher machine learning model, uses a knowledge distillation training method to train the student machine learning model to be trained, obtaining the trained student machine learning model as the target neural network model, including:

[0256] Input the training images into the trained teacher machine learning model to obtain the first output result of the trained teacher machine learning model;

[0257] The training images are input into the student machine learning model to be trained, and the second output result of the student machine learning model to be trained is obtained.

[0258] Based on the first and second output results, determine the first loss function;

[0259] The second loss function is determined based on the second output and the sample labels of the training images;

[0260] Based on the first loss function and the second loss function, the weighted loss function is obtained;

[0261] The parameters of the student machine learning model to be trained are adjusted according to the weighted loss function until the weighted loss function converges, and the trained student machine learning model is obtained as the target neural network model.

[0262] In some embodiments, the first training dataset is determined by steps performed by the following processing module:

[0263] Obtain the original dataset; the original dataset includes multiple initial sample images;

[0264] The initial sample image is used to identify the target object. If the target object exists, the first reference box containing the target object is determined.

[0265] Based on the position information of the first reference box, update the position of the first reference box to obtain the second reference box;

[0266] Based on the position information of the second reference frame and the position information of the first reference frame, determine the overlap between the second reference frame and the first reference frame;

[0267] Based on the comparison between the overlap and the first preset threshold, sample labels are assigned to the sample images of the initial sample images that are located in the second reference box, and the sample images with labeled sample labels are used as training images in the first training dataset.

[0268] In some embodiments, the processing module updates the position of the first reference frame based on the position information of the first reference frame to obtain the second reference frame, including:

[0269] Based on the position information of the first reference frame, move a specific coordinate point in the first reference frame to obtain the third reference frame;

[0270] Based on the position information of the third reference frame, the width and height of the third reference frame are adjusted according to the preset scaling factor, with the center point of the third reference frame as the center, to obtain the fourth reference frame;

[0271] Based on the position information of the fourth reference frame, a new center point is determined according to the first preset center point range; and based on the new center point and the preset clipping range, the width and height of the fourth reference frame are adjusted to obtain the second reference frame.

[0272] In some embodiments, the processing module labels the portion of the initial sample image located within the second reference frame with sample labels based on a comparison result between the overlap and a first preset threshold, including:

[0273] When the overlap is greater than or equal to the first preset threshold, a first preset range is generated according to the first preset probability, a second preset range is generated according to the second preset probability, a third preset range is generated according to the third preset probability, and a fourth preset range is generated according to the fourth preset probability; the first preset probability is greater than the second preset probability, the second preset probability is greater than the third preset probability, and the third preset probability is greater than or equal to the fourth preset probability; the sum of the first preset probability, the second preset probability, the third preset probability, and the fourth preset probability is 1.

[0274] When the overlap is within a first preset range, the sample images located in the second reference box in the initial sample image are identified and labeled as positive sample labels, so that the ratio between the number of training images with the overlap within the first preset range and the total number of training images with positive sample labels in the first training dataset is the first preset probability.

[0275] When the overlap is within a second preset range, the sample images located in the second reference box in the initial sample image are identified and labeled as positive sample labels, so that the ratio between the number of training images with overlap within the second preset range and the total number of training images with positive sample labels in the first training dataset is the second preset probability.

[0276] When the overlap is within the third preset range, the sample images located in the second reference box in the initial sample image are identified and labeled as positive sample labels, so that the ratio between the number of training images with the overlap within the third preset range and the total number of training images with positive sample labels in the first training dataset is the third preset probability.

[0277] When the overlap is within the fourth preset range, the sample images located in the second reference box in the initial sample image are identified and labeled as positive sample labels, so that the ratio between the number of training images with the overlap within the fourth preset range and the total number of training images with positive sample labels in the first training dataset is the fourth preset probability.

[0278] The overlap within the first preset range is less than the overlap within the second preset range, the overlap within the second preset range is less than the overlap within the third preset range, and the overlap within the third preset range is less than the overlap within the fourth preset range.

[0279] In some embodiments, the step of the processing module obtaining the first training dataset, if it is determined that no target object exists when identifying the target object in the initial sample image, further includes:

[0280] Label the initial sample images with negative sample labels, and use the initial sample images with negative sample labels as training images in the first training dataset.

[0281] In some embodiments, after the processing module identifies the presence of a target object in the initial sample image and determines a first reference bounding box containing the target object, it further includes:

[0282] Based on the size of the initial sample image, a target center point is determined in the initial sample image according to the second preset center point range, and a fifth reference frame is determined with the target center point as the center.

[0283] Based on the position information of the fifth reference box and the position information of the first reference box, determine the overlap between the fifth reference box and any first reference box in the initial sample image;

[0284] If the overlap between the fifth reference box and any first reference box in the initial sample image is less than or equal to the second preset threshold, then the sample images of the initial sample image located in the fifth reference box are labeled as negative samples, and the sample images labeled as negative samples are used as training images in the first training dataset.

[0285] In some embodiments, the processing module acquires the image to be detected, including:

[0286] Acquire multiple consecutive video frames captured by the camera device;

[0287] Based on multiple consecutive video frames, the frame difference method is used to determine whether there is a moving target in the shooting scene of the shooting device. If there is a moving target, the video frame captured by the shooting device is used as the image to be detected.

[0288] In addition, this disclosure also provides a storage-computing integrated system corresponding to the control method of the monitoring system of the first aspect. Since the principle of solving the problem by the storage-computing integrated system in this disclosure is similar to the control method of the monitoring system of the first aspect of this disclosure, the implementation of the storage-computing integrated system can refer to the implementation of the method corresponding to the first aspect, and the repeated parts will not be described again.

[0289] Figure 7 A schematic diagram of the in-memory computing system provided in the embodiments of this disclosure, as shown below. Figure 7 As shown, the in-memory computing system includes a processing module 71, an in-memory computing module 72, and an external storage module 73. The processing module is, for example, the main processor (CPU) of the in-memory computing system. The in-memory computing module can be, for example, a computing core integrating Flash or SRAM. The external storage module can be, for example, an external storage module (DRAM), offering faster response speeds and reduced power consumption in the monitoring system. The in-memory computing system can be part of the monitoring system. Furthermore, this disclosure designs a lightweight target neural network model, which is migrated to the in-memory computing module, employing a low-power in-memory computing approach to assist the target neural network in inference, completing the wake-up service of the smart terminal, reducing power consumption, and improving response speed.

[0290] The processing module 71 is configured to acquire the image to be detected; divide the image to be detected into multiple target sub-images according to the shooting angle of the shooting device that captures the image to be detected, and store them in the external storage module 73;

[0291] The in-memory computing module 72 is configured to process each target sub-image in multiple target sub-images using a pre-trained target neural network model, obtain the recognition result of the target sub-image, and store the recognition result in the external storage module 73.

[0292] The processing module 71 obtains the detection result of whether there is a target object in the image to be detected based on the recognition result corresponding to each target sub-image, and sends the detection result to the terminal so that the terminal can determine the display status based on the detection result.

[0293] In some embodiments, the processing module 71 is specifically configured to, when the shooting angle is within a preset viewing angle range, divide the image to be detected into a first sub-region, a second sub-region, and a third sub-region arranged sequentially along a first direction, wherein the first sub-region and the third sub-region partially overlap with the second sub-region; wherein the width of the first sub-region in the first direction is equal to the width of the third sub-region in the first direction, and the width of the second sub-region in the first direction is greater than the width of the first sub-region in the first direction; the first direction is the height direction of the image to be detected; and divide the portion of the image to be detected located in the first sub-region into multiple first sub-images arranged side by side along a second direction, wherein at least some of the first sub-images are located in the first sub-region. The widths in the second direction are equal; the second direction is the width direction of the image to be detected; the portion of the image to be detected located in the second sub-region is divided into multiple second sub-images arranged side by side along the second direction, at least some of the second sub-images having equal widths in the second direction; the portion of the image to be detected located in the third sub-region is divided into multiple third sub-images arranged side by side along the second direction, at least some of the third sub-images having equal widths in the second direction; the target sub-image includes a first sub-image, a second sub-image, and a third sub-image; the width of the third sub-image in the second direction is greater than the width of the second sub-image in the second direction, and the width of the second sub-image in the second direction is greater than the width of the first sub-image in the second direction.

[0294] In some embodiments, the first sub-images overlap at least partially in the second direction; the second sub-images overlap at least partially in the second direction; and the third sub-images do not overlap in the second direction.

[0295] In some embodiments, the width ratio of the second sub-region to the first sub-region in the first direction is 2:1; the width ratio of the third sub-image to the second sub-image in the second direction is 3:2; and the width ratio of the second sub-image to the first sub-image in the second direction is 4:3.

[0296] In some embodiments, among a plurality of first sub-images arranged side-by-side along a second direction, except for the first and last first sub-images arranged side-by-side along the second direction, the remaining first sub-images have the same width in the second direction; the first and last first sub-images have the same width in the second direction, and are smaller than the widths of the remaining first sub-images in the second direction; the ratio of the overlap width of two adjacent first sub-images in the second direction to the width of the remaining first sub-images in the second direction is 1:10; among a plurality of second sub-images arranged side-by-side along a second direction, except for the first and last second sub-images arranged side-by-side along the second direction, the remaining second sub-images have the same width in the second direction; the first and last second sub-images have the same width in the second direction, and are smaller than the widths of the remaining second sub-images in the second direction; the ratio of the overlap width of two adjacent second sub-images in the second direction to the width of the remaining second sub-images in the second direction is 1:10. The ratio of the overlap width of the first sub-image to the width of the remaining second sub-images in the second direction is 1:10; among the multiple third sub-images arranged side by side along the second direction, except for the first and last third sub-images arranged side by side along the second direction, the widths of the remaining third sub-images in the second direction are all the same; the widths of the first and last third sub-images in the second direction are the same and smaller than the widths of the remaining third sub-images in the second direction; the ratio of the overlap width of two adjacent third sub-images in the second direction to the width of the remaining third sub-images in the second direction is 1:10; the ratio of the overlap width of the first and second adjacent sub-images in the first direction to the width of the second sub-image in the first direction is 1:10; the ratio of the overlap width of the third and second adjacent sub-images in the first direction to the width of the second sub-image in the first direction is 1:10.

[0297] In some embodiments, the processing module 71 is further configured to train a target neural network model, specifically: acquiring a first training dataset and a second training dataset; the second training dataset is obtained after filtering based on the first training dataset; training a teacher machine learning model to be trained according to the first training dataset to obtain a pre-trained teacher machine learning model; training the pre-trained teacher machine learning model according to the second training dataset to obtain a trained teacher machine learning model; and training a student machine learning model to be trained using a knowledge distillation training method according to the second training dataset and the trained teacher machine learning model to obtain a trained student machine learning model as the target neural network model.

[0298] In some embodiments, the second training dataset includes multiple training images labeled with sample labels; the processing module 71 is specifically configured to input the training images into a trained teacher machine learning model to obtain a first output result of the trained teacher machine learning model; input the training images into a student machine learning model to be trained to obtain a second output result of the student machine learning model to be trained; determine a first loss function based on the first and second output results; determine a second loss function based on the second output result and the sample labels of the training images; obtain a weighted loss function based on the first and second loss functions; adjust the parameters of the student machine learning model to be trained according to the weighted loss function until the weighted loss function converges, thereby obtaining a trained student machine learning model as the target neural network model.

[0299] In some embodiments, the processing module 71 is further configured to determine a first training dataset, specifically, to obtain an original dataset; the original dataset includes multiple initial sample images; to identify target objects in the initial sample images, and if a target object exists, to determine a first reference box containing the target object; to update the position of the first reference box according to the position information of the first reference box, and to obtain a second reference box; to determine the overlap between the second reference box and the first reference box according to the position information of the second reference box and the position information of the first reference box; and to label the sample images of the initial sample images located in the second reference box according to the comparison result between the overlap and a first preset threshold, and to use the labeled sample images as training images in the first training dataset.

[0300] In some embodiments, in conjunction with the above embodiments, the processing module 71 is specifically configured to: move a specific coordinate point in the first reference frame according to the position information of the first reference frame to obtain a third reference frame; adjust the width and height of the third reference frame according to the position information of the third reference frame, with the center point of the third reference frame as the center, and according to a preset scaling factor, to obtain a fourth reference frame; determine a new center point according to the position information of the fourth reference frame, according to a first preset center point range; and adjust the width and height of the fourth reference frame according to the new center point and a preset clipping range to obtain a second reference frame.

[0301] In some embodiments, in conjunction with the above embodiments, the processing module 71 is specifically configured to, when the overlap is greater than or equal to a first preset threshold, generate a first preset range according to a first preset probability, generate a second preset range according to a second preset probability, generate a third preset range according to a third preset probability, and generate a fourth preset range according to a fourth preset probability; the first preset probability is greater than the second preset probability, the second preset probability is greater than the third preset probability, and the third preset probability is greater than or equal to the fourth preset probability; the sum of the first preset probability, the second preset probability, the third preset probability, and the fourth preset probability is 1; when the overlap is within the first preset range, determine and label a portion of the sample images located in the second reference box in the initial sample image as positive sample labels, so that the ratio between the number of training images with overlap within the first preset range and the total number of training images with positive sample labels in the first training dataset is the first preset probability; when the overlap is within the second preset range, determine and label a portion of the sample images located in the second reference box in the initial sample image as positive samples. Labels are defined such that the ratio between the number of training images with overlap within a second preset range and the total number of training images with positive sample labels in the first training dataset is a second preset probability; when the overlap is within a third preset range, a portion of the initial sample images located within the second reference box are identified and labeled as positive sample labels, so that the ratio between the number of training images with overlap within the third preset range and the total number of training images with positive sample labels in the first training dataset is a third preset probability; when the overlap is within a fourth preset range, a portion of the initial sample images located within the second reference box are identified and labeled as positive sample labels, so that the ratio between the number of training images with overlap within the fourth preset range and the total number of training images with positive sample labels in the first training dataset is a fourth preset probability; the overlap within the first preset range is less than the overlap within the second preset range, the overlap within the second preset range is less than the overlap within the third preset range, and the overlap within the third preset range is less than the overlap within the fourth preset range.

[0302] In some embodiments, the processing module 71 is further configured to, when it is determined that no target object exists in the initial sample image after identifying the target object, label the initial sample image with a negative sample label and use the initial sample image with the negative sample label as a training image in the first training dataset.

[0303] In some embodiments, after identifying the target object in the initial sample image and determining the existence of the target object, and determining the first reference box containing the target object, the processing module 71 is further configured to determine a target center point in the initial sample image according to the size of the initial sample image and according to the second preset center point range, and determine a fifth reference box with the target center point as the center; determine the overlap between the fifth reference box and any first reference box in the initial sample image according to the position information of the fifth reference box and the position information of the first reference box; if the overlap between the fifth reference box and any first reference box in the initial sample image is less than or equal to the second preset threshold, then the sample images of the initial sample image located in the fifth reference box are marked as negative sample labels, and the sample images marked with negative sample labels are used as training images in the first training dataset.

[0304] In some embodiments, the processing module 71 is specifically configured to acquire multiple consecutive video frames captured by the shooting device; based on the multiple consecutive video frames, use the frame difference method to determine whether there is a moving target in the shooting scene of the shooting device; if there is a moving target, then use the video frames captured by the shooting device as images to be detected.

[0305] In addition, this disclosure also provides a control device for a monitoring system corresponding to the control method of the monitoring system in the second aspect. Since the principle of the control device for the monitoring system in this disclosure is similar to the control method of the monitoring system in the second aspect of this disclosure, the implementation of the control device for the monitoring system can refer to the implementation of the method corresponding to the second aspect, and the repeated parts will not be described again.

[0306] Figure 8 A schematic diagram of the control device for the monitoring system provided in this embodiment of the present disclosure is shown below. Figure 8 As shown, the in-memory computing system 81 and terminal 82 are included. The in-memory computing system includes a processing module 811, an in-memory computing module 812, and an external storage module 813. The processing module is, for example, the main processor CPU of the in-memory computing system. The in-memory computing module can be, for example, an in-memory computing core integrating Flash or SRAM. The external storage module can be, for example, an external storage module DRAM, offering faster response speed and reduced power consumption of the monitoring system. The in-memory computing system can be part of the monitoring system. Furthermore, this disclosure designs a lightweight target neural network model, migrates it to the in-memory computing module, and uses a low-power in-memory computing approach to assist the target neural network in inference, completing the wake-up service of the smart terminal, reducing power consumption and increasing response speed.

[0307] The processing module 811 is configured to acquire the image to be detected; according to the shooting angle of the shooting device that captures the image to be detected, the image to be detected is divided into multiple target sub-images and stored in the external storage module 813.

[0308] The in-memory computing module 812 is configured to read the target sub-image; process the target sub-image using a pre-trained target neural network model to obtain the recognition result of the target sub-image, and store the recognition result in the external storage module 813.

[0309] The processing module 811 is configured to obtain the detection result of whether there is a target object in the image to be detected based on the recognition result corresponding to each target sub-image, and send the detection result to the terminal 82.

[0310] Terminal 82 is configured to determine the display status based on the detection results.

[0311] In some embodiments, the terminal 82 includes a display module 821 and a main control module 822; the main control module 822 is configured to send a wake-up request to the display module and reset the sleep timer when the detection result indicates that a target object exists in the image to be detected, in response to the terminal request to receive the detection result sent by the processing module; when the detection result indicates that no target object exists in the image to be detected, and the current system time is greater than the time since the last wake-up of the display module is greater than a preset duration, the main control module 822 sends a sleep request to the display module and controls itself to enter a sleep state; the display module 821 is configured to respond to the wake-up request and determine that the display state is the normal display of the shooting screen of the shooting device; and respond to the sleep request and determine that the display state is sleep.

[0312] In some embodiments, the control device of the monitoring system further includes a shooting device 83, and the image segmentation module 811 can acquire data from the low-power shooting device 83 through protocols such as SPI.

[0313] According to embodiments of this disclosure, a computer non-transient readable storage medium is also provided. This computer non-transient readable storage medium stores a computer program, wherein, when executed by a processor, the program implements the steps in the control method of any of the monitoring systems described in the above embodiments.

[0314] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a machine-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), it performs the functions defined above in the system of this disclosure.

[0315] It should be noted that the computer-readable non-transient readable medium disclosed herein may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable storage media may be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any non-transient readable computer storage medium other than a computer-readable storage medium, which can transmit, propagate, or transfer a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the non-transient readable computer storage medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0316] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two adjacent blocks may actually represent substantially parallel execution, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0317] It is understood that the above embodiments are merely exemplary embodiments used to illustrate the principles of this disclosure, and this disclosure is not limited thereto. For those skilled in the art, various modifications and improvements can be made without departing from the spirit and substance of this disclosure, and these modifications and improvements are also considered to be within the scope of protection of this disclosure.

Claims

1. A control method for a monitoring system, wherein, include: Acquire the image to be detected; Based on the shooting angle of the shooting device that captured the image to be detected, the image to be detected is divided into multiple target sub-images; For each of the plurality of target sub-images, a pre-trained target neural network model is used to process the target sub-image to obtain the recognition result of the target sub-image; Based on the recognition results corresponding to each of the target sub-images, a detection result is obtained regarding whether a target object exists in the image to be detected, and the detection result is sent to the terminal so that the terminal determines the display state based at least on the detection result; wherein, dividing the image to be detected into multiple target sub-images according to the shooting angle of the shooting device includes: When the shooting angle is within a preset angle range, the image to be detected is divided into a first sub-region, a second sub-region, and a third sub-region arranged sequentially along a first direction, and the first sub-region and the third sub-region partially overlap with the second sub-region; wherein, the width of the first sub-region in the first direction is equal to the width of the third sub-region in the first direction, and the width of the second sub-region in the first direction is greater than the width of the first sub-region in the first direction; the first direction is the height direction of the image to be detected; The portion of the image to be detected located in the first sub-region is divided into multiple first sub-images arranged side by side along a second direction, at least some of the first sub-images having equal width in the second direction; the second direction is the width direction of the image to be detected; The portion of the image to be detected located in the second sub-region is divided into multiple second sub-images arranged side by side along the second direction, and at least some of the second sub-images have equal widths in the second direction; The portion of the image to be detected located in the third sub-region is divided into multiple third sub-images arranged side by side along the second direction, at least some of the third sub-images having equal width in the second direction; the target sub-image includes a first sub-image, a second sub-image, and a third sub-image; The width of the third sub-image in the second direction is greater than the width of the second sub-image in the second direction, and the width of the second sub-image in the second direction is greater than the width of the first sub-image in the second direction.

2. The control method for the monitoring system according to claim 1, wherein, Each of the first sub-images overlaps at least partially in the second direction; each of the second sub-images overlaps at least partially in the second direction; Each of the third sub-images does not overlap in the second direction.

3. The control method for the monitoring system according to claim 2, wherein, The width ratio of the second sub-region to the first sub-region in the first direction is 2:1; The width ratio of the third sub-image to the second sub-image in the second direction is 3:2, and the width ratio of the second sub-image to the first sub-image in the second direction is 4:

3.

4. The control method for the monitoring system according to claim 2, wherein, Among the multiple first sub-images arranged side by side along the second direction, except for the first and last first sub-images arranged side by side along the second direction, the widths of the remaining first sub-images in the second direction are all the same; the widths of the first and last first sub-images in the second direction are the same and smaller than the widths of the remaining first sub-images in the second direction; the ratio of the overlap width of two adjacent first sub-images in the second direction to the width of the remaining first sub-images in the second direction is 1:

10. Among the multiple second sub-images arranged side by side along the second direction, except for the first and last second sub-images arranged side by side along the second direction, the widths of the remaining second sub-images in the second direction are all the same; the widths of the first and last second sub-images in the second direction are the same and smaller than the widths of the remaining second sub-images in the second direction; the ratio of the overlap width of two adjacent second sub-images in the second direction to the width of the remaining second sub-images in the second direction is 1:

10. Among the multiple third sub-images arranged side by side along the second direction, except for the first and last third sub-images arranged side by side along the second direction, the widths of the remaining third sub-images are all the same in the second direction; the widths of the first and last third sub-images are the same in the second direction and are smaller than the widths of the remaining third sub-images in the second direction; the ratio of the overlap width of two adjacent third sub-images in the second direction to the width of the remaining third sub-images in the second direction is 1:

10. The first sub-image and the second sub-image that are adjacent in the first direction have an overlap width in the first direction that is 1:10 to the width of the second sub-image in the first direction; the third sub-image and the second sub-image that are adjacent in the first direction have an overlap width in the first direction that is 1:10 to the width of the second sub-image in the first direction.

5. The control method for the monitoring system according to claim 1, wherein, The target neural network model is trained using the following steps: Obtain a first training dataset and a second training dataset; the second training dataset is obtained by filtering the first training dataset. The teacher machine learning model to be trained is trained based on the first training dataset to obtain the preliminarily trained teacher machine learning model. The pre-trained teacher machine learning model is trained based on the second training dataset to obtain the trained teacher machine learning model. Based on the second training dataset and the trained teacher machine learning model, the knowledge distillation training method is used to train the student machine learning model to be trained, and the trained student machine learning model is used as the target neural network model.

6. The control method for the monitoring system according to claim 5, wherein, The second training dataset includes multiple training images with labeled sample tags; The step of training the student machine learning model to be trained using a knowledge distillation method based on the second training dataset and the trained teacher machine learning model to obtain the trained student machine learning model as the target neural network model includes: The training image is input into the trained teacher machine learning model to obtain the first output result of the trained teacher machine learning model; The training image is input into the student machine learning model to be trained, and the second output result of the student machine learning model to be trained is obtained. Based on the first output result and the second output result, determine the first loss function; Based on the second output result and the sample labels of the training images, a second loss function is determined; Based on the first loss function and the second loss function, the weighted loss function is obtained; The parameters of the student machine learning model to be trained are adjusted according to the weighted loss function until the weighted loss function converges, and the trained student machine learning model is obtained as the target neural network model.

7. The control method for the monitoring system according to claim 6, wherein, The first training dataset is determined by the following steps: Obtain the original dataset; the original dataset includes multiple initial sample images; The target object is identified in the initial sample image. If the target object exists, a first reference box containing the target object is determined. Based on the position information of the first reference frame, update the position of the first reference frame to obtain the second reference frame; Based on the position information of the second reference frame and the position information of the first reference frame, the overlap between the second reference frame and the first reference frame is determined; Based on the comparison result between the overlap and the first preset threshold, the sample images located in the second reference box of the initial sample image are labeled with the sample label, and the labeled sample images are used as training images in the first training dataset.

8. The control method for the monitoring system according to claim 7, wherein, The step of updating the position of the first reference frame based on the position information of the first reference frame to obtain the second reference frame includes: Based on the position information of the first reference frame, move a specific coordinate point in the first reference frame to obtain a third reference frame; Based on the position information of the third reference frame, the width and height of the third reference frame are adjusted according to a preset scaling factor, with the center point of the third reference frame as the center, to obtain the fourth reference frame; Based on the position information of the fourth reference frame, a new center point is determined according to the first preset center point range; and based on the new center point and the preset clipping range, the width and height of the fourth reference frame are adjusted to obtain the second reference frame.

9. The control method for the monitoring system according to claim 7, wherein, The step of labeling the sample images located within the second reference frame of the initial sample image with the sample label based on the comparison result between the overlap degree and the first preset threshold includes: When the overlap is greater than or equal to the first preset threshold, a first preset range is generated according to a first preset probability, a second preset range is generated according to a second preset probability, a third preset range is generated according to a third preset probability, and a fourth preset range is generated according to a fourth preset probability; the first preset probability is greater than the second preset probability, the second preset probability is greater than the third preset probability, and the third preset probability is greater than or equal to the fourth preset probability; the sum of the first preset probability, the second preset probability, the third preset probability, and the fourth preset probability is 1; When the overlap is within the first preset range, the portion of the sample images located in the second reference box in the initial sample image is determined and labeled as positive sample labels, so that the ratio between the number of training images with the overlap within the first preset range and the total number of training images with the positive sample labels in the first training dataset is the first preset probability. When the overlap is within the second preset range, the portion of the sample images located in the second reference box in the initial sample image is determined and labeled as positive sample labels, so that the ratio between the number of training images with the overlap within the second preset range and the total number of training images with the positive sample labels in the first training dataset is the second preset probability. When the overlap is within the third preset range, the sample images located in the second reference box in the initial sample image are identified and labeled as positive sample labels, so that the ratio between the number of training images with the overlap within the third preset range and the total number of training images with the positive sample labels in the first training dataset is the third preset probability. When the overlap is within the fourth preset range, the sample images located in the second reference box in the initial sample image are identified and labeled as positive sample labels, so that the ratio between the number of training images with the overlap within the fourth preset range and the total number of training images with the positive sample labels in the first training dataset is the fourth preset probability. The overlap within the first preset range is less than the overlap within the second preset range, the overlap within the second preset range is less than the overlap within the third preset range, and the overlap within the third preset range is less than the overlap within the fourth preset range.

10. The control method for the monitoring system according to claim 7, wherein, In the step of obtaining the first training dataset, if it is determined that the target object does not exist when identifying the target object in the initial sample image, the method further includes: The initial sample images are labeled with negative sample labels, and the initial sample images labeled with negative sample labels are used as training images in the first training dataset.

11. The control method for the monitoring system according to claim 7, wherein, After identifying the target object in the initial sample image, determining the presence of the target object, and determining a first reference bounding box containing the target object, the method further includes: Based on the size of the initial sample image, a target center point is determined in the initial sample image according to the second preset center point range, and a fifth reference frame is determined with the target center point as the center. Based on the position information of the fifth reference frame and the position information of the first reference frame, the overlap between the fifth reference frame and any first reference frame in the initial sample image is determined; If the overlap between the fifth reference box and any first reference box in the initial sample image is less than or equal to the second preset threshold, then the portion of the initial sample image located in the fifth reference box is labeled as a negative sample, and the portion of the sample image labeled with the negative sample is used as the training image in the first training dataset.

12. The control method for the monitoring system according to claim 1, wherein, The acquisition of the image to be detected includes: Acquire multiple consecutive video frames captured by the camera device; Based on the multiple consecutive video frames, the frame difference method is used to determine whether there is a moving target in the shooting scene of the shooting device. If the moving target exists, the video frames captured by the shooting device are used as images to be detected.

13. A control method for a monitoring system, wherein, The monitoring system includes a storage-and-computing integrated system and a terminal; the storage-and-computing integrated system includes a processing module, a storage-and-computing integrated module, and an external storage module; the control method of the monitoring system includes: The processing module acquires the image to be detected; according to the shooting angle of the shooting device that captured the image to be detected, the image to be detected is divided into multiple target sub-images and stored in the external storage module; The in-memory computing module reads the target sub-image; uses a pre-trained target neural network model to process the target sub-image, obtains the recognition result of the target sub-image, and stores the recognition result in the external storage module; The processing module obtains the detection result of whether a target object exists in the image to be detected based on the recognition result corresponding to each of the target sub-images, and sends the detection result to the terminal; The terminal determines the display status based on the detection results; The step of dividing the image to be detected into multiple target sub-images according to the shooting angle of the shooting device includes: When the shooting angle is within a preset angle range, the image to be detected is divided into a first sub-region, a second sub-region, and a third sub-region arranged sequentially along a first direction, and the first sub-region and the third sub-region partially overlap with the second sub-region; wherein, the width of the first sub-region in the first direction is equal to the width of the third sub-region in the first direction, and the width of the second sub-region in the first direction is greater than the width of the first sub-region in the first direction; the first direction is the height direction of the image to be detected; The portion of the image to be detected located in the first sub-region is divided into multiple first sub-images arranged side by side along a second direction, at least some of the first sub-images having equal width in the second direction; the second direction is the width direction of the image to be detected; The portion of the image to be detected located in the second sub-region is divided into multiple second sub-images arranged side by side along the second direction, and at least some of the second sub-images have equal widths in the second direction; The portion of the image to be detected located in the third sub-region is divided into multiple third sub-images arranged side by side along the second direction, at least some of the third sub-images having equal width in the second direction; the target sub-image includes a first sub-image, a second sub-image, and a third sub-image; The width of the third sub-image in the second direction is greater than the width of the second sub-image in the second direction, and the width of the second sub-image in the second direction is greater than the width of the first sub-image in the second direction.

14. The control method for the monitoring system according to claim 13, wherein, The terminal includes a display module and a main control module; The terminal determines the display status based on the detection result, including: When the main control module determines that the detection result indicates that the target object exists in the image to be detected, it sends a wake-up request to the display module and resets the sleep timer to respond to the terminal request to receive the detection result sent by the processing module; when it determines that the detection result indicates that the target object does not exist in the image to be detected, and the current system time is greater than the time since the last wake-up of the display module is greater than a preset time, it sends a sleep request to the display module and controls itself to enter a sleep state. The display module responds to the wake-up request and determines that the display state is to normally display the captured image of the shooting device; responds to the sleep request and determines that the display state is to sleep.

15. An in-memory computing system, comprising a processing module, an in-memory computing module, and an external storage module; The processing module is configured to acquire an image to be detected; divide the image to be detected into multiple target sub-images according to the shooting angle of the shooting device that captured the image to be detected, and store them in the external storage module; The in-memory computing module is configured to read the target sub-image; The target sub-image is processed using a pre-trained target neural network model to obtain the recognition result of the target sub-image, and the recognition result is stored in the external storage module. The processing module is configured to obtain a detection result of whether a target object exists in the image to be detected based on the recognition result corresponding to each of the target sub-images, and send the detection result to the terminal so that the terminal determines the display state based on the detection result; The step of dividing the image to be detected into multiple target sub-images according to the shooting angle of the shooting device includes: When the shooting angle is within a preset angle range, the image to be detected is divided into a first sub-region, a second sub-region, and a third sub-region arranged sequentially along a first direction, and the first sub-region and the third sub-region partially overlap with the second sub-region; wherein, the width of the first sub-region in the first direction is equal to the width of the third sub-region in the first direction, and the width of the second sub-region in the first direction is greater than the width of the first sub-region in the first direction; the first direction is the height direction of the image to be detected; The portion of the image to be detected located in the first sub-region is divided into multiple first sub-images arranged side by side along a second direction, at least some of the first sub-images having equal width in the second direction; the second direction is the width direction of the image to be detected; The portion of the image to be detected located in the second sub-region is divided into multiple second sub-images arranged side by side along the second direction, and at least some of the second sub-images have equal widths in the second direction; The portion of the image to be detected located in the third sub-region is divided into multiple third sub-images arranged side by side along the second direction, at least some of the third sub-images having equal width in the second direction; the target sub-image includes a first sub-image, a second sub-image, and a third sub-image; The width of the third sub-image in the second direction is greater than the width of the second sub-image in the second direction, and the width of the second sub-image in the second direction is greater than the width of the first sub-image in the second direction.

16. A control device for a monitoring system, comprising an in-memory computing system and a terminal; the in-memory computing system comprising a processing module, an in-memory computing module, and an external storage module; The processing module is configured to acquire an image to be detected; divide the image to be detected into multiple target sub-images according to the shooting angle of the shooting device that captured the image to be detected, and store them in the external storage module; The in-memory computing module is configured to read the target sub-image; The target sub-image is processed using a pre-trained target neural network model to obtain the recognition result of the target sub-image, and the recognition result is stored in the external storage module. The processing module is configured to obtain a detection result of whether a target object exists in the image to be detected based on the recognition result corresponding to each of the target sub-images, and send the detection result to the terminal; The terminal is configured to determine the display status based on the detection result; The step of dividing the image to be detected into multiple target sub-images according to the shooting angle of the shooting device includes: When the shooting angle is within a preset angle range, the image to be detected is divided into a first sub-region, a second sub-region, and a third sub-region arranged sequentially along a first direction, and the first sub-region and the third sub-region partially overlap with the second sub-region; wherein, the width of the first sub-region in the first direction is equal to the width of the third sub-region in the first direction, and the width of the second sub-region in the first direction is greater than the width of the first sub-region in the first direction; the first direction is the height direction of the image to be detected; The portion of the image to be detected located in the first sub-region is divided into multiple first sub-images arranged side by side along a second direction, at least some of the first sub-images having equal width in the second direction; the second direction is the width direction of the image to be detected; The portion of the image to be detected located in the second sub-region is divided into multiple second sub-images arranged side by side along the second direction, and at least some of the second sub-images have equal widths in the second direction; The portion of the image to be detected located in the third sub-region is divided into multiple third sub-images arranged side by side along the second direction, at least some of the third sub-images having equal width in the second direction; the target sub-image includes a first sub-image, a second sub-image, and a third sub-image; The width of the third sub-image in the second direction is greater than the width of the second sub-image in the second direction, and the width of the second sub-image in the second direction is greater than the width of the first sub-image in the second direction.

17. A computer-defined non-transient readable storage medium, wherein, The computer non-transient readable storage medium stores a computer program that, when executed by a processor, performs the steps of the control method of the monitoring system as described in any one of claims 1 to 12, and / or the steps of the control method of the monitoring system as described in any one of claims 13 to 14.

Citation Information

Patent Citations

  • Target detection method and apparatus

    CN105740792A

  • Smartphone wake-on-approach method

    CN106303065A