Personnel state detection method and device and nonvolatile storage medium
By filtering images with image similarity in video stream data with a lower image similarity than the preset value and using the improved neural network model for object detection, the problem of waste of computing resources and slow data processing speed caused by large parameters of the deep learning object detection algorithm is solved, and the effect of improving data processing speed is achieved.
Patent Information
- Application Number
- CN202510045520.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-09
AI Technical Summary
The large amount of parameters of deep learning object detection algorithm model leads to waste of computing resources and slow data processing speed.
By acquiring video stream data, we filtered images with images that have a similarity of image less than a preset value, and analyzed these images using a target neural network model that deleted the high-dimensional convolutional layer to generate target images marked with labels.
The amount of data to be processed is reduced, the amount of model parameters is reduced, computing resources is saved, and data processing speed is improved.
Smart Images

Figure CN119964081A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular, to a method and device for detecting a person's status, and a non-volatile storage medium. Background Art
[0002] With the development of computer technology, especially the breakthrough progress of deep learning technology, artificial intelligence (AI) technology is increasingly widely used in all walks of life. Among them, deep learning target detection algorithm, as an important branch of computer vision, plays a vital role in the field of video image processing. In recent years, with the improvement of service quality in business halls and the increase in demand for optimizing customer experience, the use of deep learning target detection algorithms to monitor, evaluate and warn the behavioral norms of personnel in business halls has become a new trend. In related technologies, the deep learning target detection algorithm contains a high-dimensional convolution process. The model that executes the deep learning target detection algorithm needs to perform high-dimensional convolution processing when processing input data. Therefore, the requirements for the model are high, the model parameters are large, there is a waste of computing resources, and the data processing speed is slow.
[0003] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention
[0004] The embodiments of the present application provide a method and device for detecting the status of a person, and a non-volatile storage medium, so as to at least solve the technical problem that a large amount of computing resources is consumed and the data processing speed is slow when applying a deep learning target detection algorithm due to the large number of parameters of the model that executes the deep learning target detection algorithm.
[0005] According to one aspect of an embodiment of the present application, a method for detecting the status of a person is provided, comprising: obtaining video stream data of an area to be detected, wherein the video stream data is used to record dynamic changes in the area to be detected; determining from the video stream data a plurality of images to be processed whose image similarity is less than a preset value, wherein each image to be processed is used to record a frame of the video stream data; using a target detection model to analyze the image to be processed to obtain a processing result, wherein the processing result is a target image marked with a label, the label is used to indicate the status of a target object located in the area to be detected, and the target detection model is obtained by training a target neural network model using historical video data of the area to be detected as training data, and the target neural network model is a neural network model with a high-dimensional convolutional layer deleted.
[0006] Optionally, a plurality of images to be processed whose image similarities are less than a preset value are generated based on the video data, including: intercepting a plurality of initial images in the video stream data at a preset frequency; for each initial image, determining target pixels of a target area in the initial image, wherein the target area is an area where the initial image overlaps with a template image, and the template image is a preset image that only includes a position to be detected in the area to be detected; determining the image similarity between every two initial images based on the plurality of target pixels; and determining the image to be processed based on the image similarity.
[0007] Optionally, determining the image similarity of every two initial images based on multiple target pixels includes: determining a first coordinate of a target pixel in one of the initial images in the initial image pair to be compared, and a second coordinate of the target pixel in the other initial image; determining the image similarity based on the multiple first coordinates, the multiple second coordinates and a transformation relationship between the two initial images in the initial image pair to be compared, wherein the transformation relationship includes: scaling, rotation, and translation; determining the target image based on the image similarity includes: determining multiple target initial images corresponding to multiple image similarities greater than or equal to a preset value; and deleting the target initial image from all the initial images to obtain multiple target images.
[0008] Optionally, after determining the image similarity of every two initial images based on multiple target pixels, the method also includes: determining a group of initial images corresponding to the image similarity, and determining a target device corresponding to a group of initial images, wherein the target device is a device that generates video stream data belonging to a group of initial images, and the number of target devices is multiple; determining the image similarity as the weight of the target device, wherein the weight is related to the calling order of the target device, and the calling order of the target device is the order of obtaining video stream data from the target device when detecting the status of people in the area to be detected.
[0009] Optionally, a target detection model is used to process and analyze the image to be processed to obtain a processing result, including: normalizing and activating the image to be processed in the first neural network layer of the target detection model to obtain a first feature map output by the first neural network layer; convolving the first feature map with a one-dimensional convolution kernel in the second neural network layer of the target detection model to obtain a convolution result, and normalizing and activating the convolution result to obtain a second feature map output by the second neural network layer; fusing the first feature map and the second feature map into a target feature map to be marked; marking the target feature map to generate a processing result.
[0010] Optionally, the first neural network layer of the target detection model performs normalization and activation processing on the image to be processed to obtain a first feature map output by the first neural network layer, including: extracting a feature vector of each image to be processed to obtain multiple feature vectors, wherein the feature vector is a vector generated based on feature information of the image to be processed, and the feature information includes at least: the edge of the image, the texture of the image, and the color of the image; determining the mean of the multiple feature vectors, and converting each feature vector into a feature value belonging to a preset value range based on the mean; processing each pixel in the image to be processed based on the feature value to obtain a first feature map.
[0011] Optionally, the target feature map is marked and a processing result is generated, including: determining multiple prediction probabilities corresponding to each grid in the target feature map, wherein multiple grids constitute the target feature map, and the multiple prediction probabilities are probabilities that the target object is in multiple states, and the states include: idle, present, away, and in service; for each grid, determining the target prediction probability with the largest value among the multiple prediction probabilities; recording the state corresponding to the target prediction probability as a label, and marking the grid with the label.
[0012] Optionally, after obtaining the processing result, it also includes: determining whether there is a target tag in the target image, wherein the target tag includes: a tag whose recorded status is idle, and a tag whose recorded status is away; in the case that the target tag exists in the target image, sending an alarm message, wherein the alarm message is used to indicate that there is a target object in an abnormal state in the area to be detected.
[0013] According to another aspect of an embodiment of the present application, a device for detecting the status of a person is also provided, including: an acquisition module, used to acquire video stream data of an area to be detected, wherein the video stream data is used to record dynamic changes in the area to be detected; a determination module, used to determine from the video stream data a plurality of images to be processed whose image similarity is less than a preset value, wherein each image to be processed is used to record a frame of the video stream data; a processing module, used to use a target detection model to process and analyze the image to be processed to obtain a processing result, wherein the processing result is a target image marked with a label, the label is used to indicate the status of a target object located in the area to be detected, and the target detection model is obtained by training a target neural network model using historical video data of the area to be detected as training data, and the target neural network model is a neural network model with a high-dimensional convolutional layer deleted.
[0014] According to another aspect of an embodiment of the present application, a non-volatile storage medium is further provided, in which a computer program is stored, wherein the above-mentioned method for detecting personnel status is executed by running the computer program on a device where the non-volatile storage medium is located.
[0015] According to another aspect of an embodiment of the present application, there is also provided an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the above-mentioned method for detecting the status of a person through the computer program.
[0016] According to another aspect of an embodiment of the present application, a computer program product is also provided, including computer instructions, which implement the steps of the above-mentioned method for detecting personnel status when the computer instructions are executed by a processor.
[0017] In an embodiment of the present application, video stream data of an area to be detected is obtained, wherein the video stream data is used to record dynamic changes in the area to be detected; multiple images to be processed whose image similarity is less than a preset value are determined from the video stream data, wherein each image to be processed is used to record one frame of the video stream data; a target detection model is used to analyze the image to be processed to obtain a processing result, wherein the processing result is a target image marked with a label, and the label is used to indicate the state of the target object located in the area to be detected. The target detection model is obtained by training a target neural network model using historical video data of the area to be detected as training data, and the target neural network model is a neural network model with a high-dimensional convolutional layer deleted. By screening the acquired video data, the purpose of reducing the amount of data to be processed is achieved, and an improved neural network model is used to perform target detection on the video data to be processed, and the states of multiple persons appearing in the video data to be processed are determined, thereby achieving the purpose of reducing model parameters and saving computing resources, thereby achieving the technical effect of improving data processing speed, and further solving the technical problem of consuming a lot of computing resources and slow data processing speed when applying a deep learning target detection algorithm due to the large number of parameters of the model executing the deep learning target detection algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0019] Figure 1 is a hardware structure block diagram of a computer terminal for implementing a method for detecting a personnel status according to an embodiment of the present application;
[0020] Figure 2 is a flowchart of a method for detecting a person's status according to an embodiment of the present application;
[0021] Figure 3 is a model structure comparison diagram of a neural network model before improvement and a neural network model after improvement according to an embodiment of the present application;
[0022] Figure 4is a structural diagram of a device for detecting a person's status according to an embodiment of the present application;
[0023] Figure 5 This is a workflow diagram of a device for detecting a person's status according to an embodiment of the present application. DETAILED DESCRIPTION
[0024] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this application.
[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0026] In order to better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application are explained as follows:
[0027] Object detection algorithm: an algorithm used to detect and locate objects in videos or images.
[0028] In the deep learning target detection algorithm in the related art, there is a high-dimensional convolution process. Therefore, a large model with a large number of model parameters is usually used when executing the deep learning target detection algorithm. Therefore, the large number of model parameters leads to the waste of computing resources and the long data processing time. In order to solve this problem, the embodiments of the present application provide a relevant solution, which is described in detail below.
[0029] According to an embodiment of the present application, a method embodiment for detecting the status of a person is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0030] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 FIG. 1 shows a hardware structure block diagram of a computer terminal for implementing a method for detecting a person's status. Figure 1 As shown, the computer terminal 10 may include one or more (102a, 102b, ..., 102n are used to illustrate) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.
[0031] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuits". The data processing circuits may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the computer terminal 10. As described in the embodiments of the present application, the data processing circuit acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0032] The memory 104 can be used to store software programs and modules of application software, such as program instructions / data storage devices corresponding to the detection method of personnel status in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, realizing the above-mentioned detection method of personnel status. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0033] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0034] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 .
[0035] The embodiment of the present application provides a method for detecting the status of a person that can be operated in the above operating environment. Figure 2 is a flowchart of the steps of the method for detecting the status of a person provided in an embodiment of the present application, such as Figure 2 As shown, the method comprises the following steps:
[0036] Step S202, obtaining video stream data of the area to be detected, wherein the video stream data is used to record dynamic changes of the area to be detected.
[0037] The method provided in the embodiment of the present application can be used to process different video / image data, and can be particularly used to process video data in areas with large personnel flow, for example, it can be used to process recorded videos of service halls and shopping malls. In step S202, video stream data recording the area to be detected in a preset time period (for example, one hour, one day) is obtained, and the video stream data records the dynamic changes in the area to be detected, for example, records the dynamic changes of each person in the area to be detected in the preset time period, and records the dynamic changes of each object in the area to be detected in the preset time period.
[0038] Step S204: determining a plurality of to-be-processed images whose image similarity is less than a preset value from the video stream data, wherein each to-be-processed image is used to record a frame of the video stream data.
[0039] In areas with a large flow of people, there are usually multiple camera devices (i.e., target devices) for recording dynamic changes in the area to be detected in the form of video, and the coverage of these monitoring devices may overlap. In step S202, the video stream data obtained from these monitoring devices may contain data recording the same information. In step S204, in order to reduce the amount of data to be processed, multiple frame segments with a similarity less than a preset value are determined from the multiple frame segments contained in the video stream image, and they are used as objects for analysis by the target detection algorithm (i.e., images to be processed), wherein each frame segment is presented in the form of a static image, and the similarity between the above frame segments is also called image similarity; the above preset value is a preset image similarity threshold. When the similarity between two frame segments is greater than or equal to the image similarity threshold (i.e., the preset value), it is determined that the information recorded in the two frame segments is the same.
[0040] According to an optional embodiment of the present application, a plurality of images to be processed whose image similarities are less than a preset value are generated based on video data, including: intercepting a plurality of initial images in the video stream data at a preset frequency; for each initial image, determining target pixels of a target area in the initial image, wherein the target area is an area where the initial image overlaps with a template image, and the template image is a preset image that only includes a position to be detected in the area to be detected; determining the image similarity between every two initial images based on the plurality of target pixels; and determining the image to be processed based on the image similarity.
[0041] As mentioned in step S204, different camera devices (i.e., target devices) may record dynamic changes in the same position of the detection area, and there are video stream data recording the same information. In order to reduce the amount of data to be processed, the video stream data is screened by horizontal comparison, so that the data obtained after the screening are all recorded in the dynamic changes of different positions of the detection area. Specifically, in this embodiment, first, multiple initial images are extracted from the video data acquired in step S202. For example, different frames of the video stream data can be intercepted by intercepting the video stream data by frequency, and each frame of the image intercepted is the initial image; when intercepting by frequency, 1 second, 5 seconds, or 8 seconds can be set as the interception frequency (i.e., the preset frequency). The setting of the interception frequency is positively correlated with the flow of people in the detection area. The larger the flow of people in the detection area in the preset time period, the higher the interception frequency (the smaller the value). For example, if the flow of people in the detection area in the preset time period is 50 people, the interception frequency is set to 5 seconds; if the flow of people in the detection area in the preset time period is 100 people, the interception frequency is set to 1 second. In the embodiment of the present application, if the structures of the areas recorded in the two images are the same, it is considered that the two images record the same information. Therefore, in the present embodiment, when determining the similarity of the two initial images, it is essentially to determine whether the markers recorded in the two images are the same. Therefore, after obtaining the initial images, it is necessary to detect the markers in each initial image. In the present embodiment, the markers contained in each initial image are determined by matching each initial image with a template image. The markers are objects that will not change dynamically in the area to be detected, and the markers are composed of multiple (target) pixels in the image; wherein the template image is a pre-set image that only records the markers. An image of the position of the marker in the area to be detected (i.e., the position to be detected) is obtained. Objects that do not contain the marker in the template image are processed by masking. Therefore, when matching, the edges of the initial image and the template image are aligned and then fitted. If there is an unmasked part (i.e., the overlapping area) in the fitted image, it is determined that the initial image contains the marker, and the multiple (target) pixels that constitute the marker can be determined in each initial image; next, the image similarity of the two initial images can be determined based on the multiple (target) pixels that constitute the marker; finally, the image similarity is used to screen out the objects (i.e., the images to be processed) to be analyzed by the target detection algorithm.
[0042] Optionally, determining the image similarity of every two initial images based on multiple target pixels includes: determining a first coordinate of a target pixel in one of the initial images in the initial image pair to be compared, and a second coordinate of the target pixel in the other initial image; determining the image similarity based on the multiple first coordinates, the multiple second coordinates and a transformation relationship between the two initial images in the initial image pair to be compared, wherein the transformation relationship includes: scaling, rotation, and translation; determining the target image based on the image similarity includes: determining multiple target initial images corresponding to multiple image similarities greater than or equal to a preset value; and deleting the target initial image from all the initial images to obtain multiple target images.
[0043] In this embodiment, when determining the image similarity based on multiple (target) pixels constituting the marker, the two initial images to be compared constitute an initial image pair; for each initial image in the initial image pair, the position (i.e., the first coordinate and the second coordinate) of the marker constituent pixel (i.e., the target pixel) contained therein is determined, and the image similarity of the initial image pair is determined based on the target pixel at each coordinate, for example, according to the formula R(x,y)=∑{x , ,y ,}(T(x , ,y , )-l(x+x , ,y+y , )) 2 Determine the image similarity R(x,y) of the initial image pair, where (x,y) represents the relative displacement of the two initial images in the initial image pair, and the relative displacement can be determined by the first coordinate and the second coordinate, for example, it can be the difference between the first coordinate and the second coordinate. , ,y , ) is the first coordinate or the second coordinate, indicating the position of the target pixel in one of the initial images of the initial image pair; T(x , ,y , ) indicates that the coordinate is (x , ,y , )’s target pixel value, l(x+x , ,y+y , ) represents the coordinates (x+x , ,y+y , ), x is a value generated based on information related to a transformation of a target pixel's horizontal coordinate from a horizontal coordinate in a first coordinate to a horizontal coordinate in a second coordinate, and y is a value generated based on information related to a transformation of a target pixel's vertical coordinate from a vertical coordinate in a first coordinate to a vertical coordinate in a second coordinate; a target pixel may undergo one or more of translation, rotation, and scaling when it is transformed from a first coordinate to a second coordinate.
[0044] According to another optional embodiment of the present application, after determining the image similarity of every two initial images based on multiple target pixels, the method also includes: determining a group of initial images corresponding to the image similarity, and determining a target device corresponding to a group of initial images, wherein the target device is a device that generates video stream data belonging to a group of initial images, and the number of target devices is multiple; determining the image similarity as the weight of the target device, wherein the weight is related to the calling order of the target device, and the calling order of the target device is the order of obtaining video stream data from the target device when detecting the status of people in the area to be detected.
[0045] In this embodiment, when acquiring the video stream data of the area to be detected in step S202, the existence of image data with image similarity greater than a preset value in the acquired video stream data can be avoided by acquiring only the video stream data from a selected few camera devices (i.e., target devices), so as to further improve the speed of data processing. Specifically, the weight of each camera device (i.e., target device) can be reversely determined through multiple rounds of image data screening processes and results included in the training process of the target detection model, and the camera devices (i.e., target devices) that are selected can be determined according to the size of the weight and the preset number of target devices. For example, in this embodiment, during the training process of the target detection model, after the image similarities of multiple initial image pairs to be compared are determined using the method in the previous embodiment, the obtained image similarities are determined as the weights of the camera devices (i.e., target devices) that generate the initial image pairs (i.e., a group of initial images) corresponding to the image similarities, wherein, in order to determine the camera devices that generate the initial image pairs (i.e., a group of initial images) corresponding to the image similarities, the video stream data to which each initial image in the initial image pair belongs can be determined first, and then the camera devices corresponding to the video stream data can be determined; in the above manner, after the training process is completed, the corresponding weights are determined for each camera device in the area to be detected, and then in the subsequent use process, the order of weights from large to small can be used as the order of calling the camera devices, and the above order of calling the camera devices is the order of obtaining the video stream data from the camera devices. If the selected camera devices (i.e., target devices) are determined according to the weights and the preset number of target devices, the calling order of the target devices is first determined according to the weights, and video stream data is obtained starting from the target device that is first in the calling order, and video stream data is no longer obtained until the video stream data of the nth (the number of target devices is n) camera device is obtained.
[0046] Step S206, using the target detection model to analyze the image to be processed to obtain a processing result, wherein the processing result is a target image marked with a label, the label is used to indicate the state of the target object located in the area to be detected, and the target detection model is obtained by training the target neural network model using the historical video data of the area to be detected as training data, and the target neural network model is a neural network model with the high-dimensional convolutional layer deleted.
[0047] After the object to be analyzed by the target detection algorithm is screened out in step S204 (i.e., the image to be processed), in step S206, the target detection model is used to analyze the image to be processed, and a labeled image (i.e., the processing result) output by the target detection model is obtained. The target detection model uses the target detection algorithm to identify the state of each person contained in the image to be processed, generates a label containing information about the state of the person, and marks the label at the position of the person in the image; the above-mentioned state of the person can be divided into two categories, one category indicates that the person is in the corresponding preset position in the area to be detected, and the other category indicates that the person is not in the corresponding preset position. The target detection model is an improved neural network model. In this embodiment, the neural network model is improved by deleting the high-dimensional convolution layer. By using the historical video data that records the dynamic changes of the area to be detected before the current moment to train the above-mentioned improved neural network model, a target detection model with fewer model parameters can be obtained.
[0048] According to an optional embodiment of the present application, a target detection model is used to process and analyze the image to be processed to obtain a processing result, including: normalizing and activating the image to be processed in the first neural network layer of the target detection model to obtain a first feature map output by the first neural network layer; convolving the first feature map with a one-dimensional convolution kernel in the second neural network layer of the target detection model to obtain a convolution result, and normalizing and activating the convolution result to obtain a second feature map output by the second neural network layer; fusing the first feature map and the second feature map into a target feature map to be marked; marking the target feature map to generate a processing result.
[0049] Figure 3 This is a comparison diagram of the model structure of the neural network model before and after improvement, such as Figure 3As shown, the improved neural network model reduces one high-dimensional convolution layer (i.e., 3*3 convolution layer) compared with the neural network model before improvement. The target detection model is trained by the improved neural network model. Then, the network structure of the target detection model is the same as that of the improved neural network model, both of which include a neural network layer for normalization and activation processing (i.e., the first neural network layer), and a neural network layer for low-dimensional convolution and normalization and activation processing after convolution (i.e., the second neural network layer). In this embodiment, when the target detection model processes the image to be processed, it first performs batch normalization and activation processing on the image to be processed in the first neural network layer to obtain a feature map (i.e., the first feature map). Furthermore, in the second neural network layer, the (first) feature map output by the first neural network layer is subjected to low-dimensional convolution and normalization and activation processing after convolution, and the image to be processed (i.e., the second feature map) is output after the image processing is completed. Finally, the target detection model predicts and marks the state of each target object in the target feature map after the image processing is completed (i.e., the second feature map), and obtains a feature map marked with a label (i.e., the processing result). Among them, the activation function (SiLU) can be used during the activation process, and the activation process can be expressed as the formula: f(x) = x*sigmoid(x), wherein sigmoid is an activation function, x is the first feature map to be processed, and * represents multiplication. Still as Figure 3 As shown in the figure, the low-dimensional convolution layer uses a 1-dimensional (i.e. 1*1) convolution kernel. The process of low-dimensional convolution can be expressed as the formula: param =(ksize*ksize*inchannel+bias)*out_channel, where conv param represents the convolution result, ksize represents the size of the convolution kernel. In this embodiment, ksize is 1, inchannel represents the number of channels of the first input feature map (for example, the color channel contains three primary colors, and its channel number is 3), bias represents the number of bias items, and each channel of the output image has one bias item. Out_channel represents the number of channels of the output feature map.
[0050] Optionally, the first neural network layer of the target detection model performs normalization and activation processing on the image to be processed to obtain a first feature map output by the first neural network layer, including: extracting a feature vector of each image to be processed to obtain multiple feature vectors, wherein the feature vector is a vector generated based on feature information of the image to be processed, and the feature information includes at least: the edge of the image, the texture of the image, and the color of the image; determining the mean of the multiple feature vectors, and converting each feature vector into a feature value belonging to a preset value range based on the mean; processing each pixel in the image to be processed based on the feature value to obtain a first feature map.
[0051] In this embodiment, the processing of the image to be processed in the first neural network layer (i.e., batch normalization (Batch Normalization, BN) + activation function (SiLU) layer) can be expressed as the following formula, where x i represents the feature vector of each image to be processed. Each image to be processed generates a feature vector, which is generated based on the feature information of the image to be processed (image edge, image texture, image color, etc.). m represents the total number of images to be processed. represents the average value of multiple eigenvectors, Represents the variance of m eigenvectors, ∈ is a preset positive number with a very small value to prevent the denominator from being 0 during the calculation process and ensure the accuracy of the result. γ and β are the preset model parameters of the BN layer. Represents the eigenvalue of the i-th image to be processed, y i Represents the vector form corresponding to the (first) feature map of the i-th image to be processed.
[0052]
[0053] Optionally, the target feature map is marked and a processing result is generated, including: determining multiple prediction probabilities corresponding to each grid in the target feature map, wherein multiple grids constitute the target feature map, and the multiple prediction probabilities are probabilities that the target object is in multiple states, and the states include: idle, present, away, and in service; for each grid, determining the target prediction probability with the largest value among the multiple prediction probabilities; recording the state corresponding to the target prediction probability as a label, and marking the grid with the label.
[0054] In this embodiment, each target object in the feature map is marked, wherein if the state of a person in the area to be detected is detected, the target object is a person; if the state of an object is detected, the target object is an object. When marking, the target feature map is divided into multiple grids. If the grid contains the target object, the state of the target object is determined according to the image of the target object, wherein the target detection model learns the following four states of the target object during the training process: idle, present, absent, and in service; then when marking the target feature map, the target detection model determines the probability (i.e., predicted probability) of the target object in each state respectively, and further, records the state corresponding to the probability value with the largest value (i.e., the target predicted probability) as a label, and uses the generated label to mark the grid; otherwise, if the grid does not contain the target object, it will not be marked with a label.
[0055] In step S206, the target detection model may be loaded into the memory, for example, the raw data of the target detection model may be loaded from the non-volatile memory into the volatile memory, so that the processor runs the target detection model. The raw data of the target detection model refers to unprocessed data, and generally includes parameters and structural data of the target detection model. The structural data may be a calculation relationship based on the parameters, such as a forward propagation calculation relationship between intermediate layers and neurons. Specifically, the structural data may include code related to the structure of the target detection model, such as code for performing related calculations between intermediate layers and neurons.
[0056] In one embodiment, an area for loading the target detection model can be divided in the memory, which can include a structure data storage area and a parameter storage area. The structure data storage area is used to store structure-related codes, and the parameters referenced by it can point to the address of specific parameters in the parameter storage area through a pointer. During the training process of the target detection model, it may be necessary to update the parameters frequently, and the parameter values in the parameter storage area can be updated.
[0057] In the embodiment of the present application, the loss function (life) applied during the training process can be set to Among them, k is the target existence coefficient, which indicates the probability that the target object to be detected appears in a certain image to be processed; f is the target disappearance coefficient, which indicates the probability that the target object to be detected does not appear in a certain image to be processed; s and j are the preset sensitivity coefficients, which are used to balance the influence of the target existence coefficient and the target disappearance coefficient on the training results and improve the training effect of the model; e is the natural number base.
[0058] According to some optional embodiments of the present application, after obtaining the processing results, it also includes: determining whether there is a target tag in the target image, wherein the target tag includes: a tag whose recorded status is idle, and a tag whose recorded status is away; in the case where there is a target tag in the target image, sending an alarm message, wherein the alarm message is used to indicate that there is a target object in an abnormal state in the area to be detected.
[0059] After detecting the status of a person, the method provided in the embodiment of the present application can further identify whether there is a target object in an abnormal state in the area to be detected. For example, in some embodiments, after obtaining an image marked with a label (i.e., a processing result), each label in the image marked with a label is detected. If there is a label with preset information recorded (i.e., a target label), an alarm message is generated to inform the user of the target detection model that there is a target object in an abnormal state in the area to be detected, wherein the above-mentioned preset information can be "absent" and "idle".
[0060] Through the above steps, it is possible to filter the data to be processed by comparing the similarity of image data, thereby reducing the amount of data to be processed; use the improved neural network model (i.e., the target detection model) to process the filtered data to be processed, wherein the improved neural network model has fewer model parameters, thereby achieving the purpose of reducing model parameters and realizing the technical effect of reducing the computing resources consumed in the data processing process. By improving the model and reducing the amount of data to be processed, the technical effect of increasing the data processing speed is also achieved.
[0061] Figure 4 is a structural diagram of a device for detecting a person's status according to an embodiment of the present application, such as Figure 4 As shown, the detection device for the status of a person includes: an acquisition module 40, which is used to acquire video stream data of the area to be detected, wherein the video stream data is used to record the dynamic changes of the area to be detected; a determination module 42, which is used to determine from the video stream data a plurality of images to be processed whose image similarity is less than a preset value, wherein each image to be processed is used to record a frame of the video stream data; a processing module 44, which is used to use a target detection model to process and analyze the image to be processed to obtain a processing result, wherein the processing result is a target image marked with a label, and the label is used to indicate the status of the target object located in the area to be detected, and the target detection model is obtained by training a target neural network model using historical video data of the area to be detected as training data, and the target neural network model is a neural network model with a high-dimensional convolutional layer deleted.
[0062] Figure 5 It is the working flow chart of the detection device of personnel status, such as Figure 5 As shown, the detection device starts to work, and the acquisition module 40 acquires the video stream data (i.e., video images) used to record the dynamic changes of the area to be detected. The acquisition module 40 transmits the video stream data to the determination module 42. The determination module 42 executes the process of monitoring screen screening and regional adaptive screening, and screens multiple images to be processed whose image similarity is less than a preset value in the video stream data. Further, the determination module 42 outputs the screened images to be processed to the processing module 44. The processing module 44 uses the target detection model to execute the deep learning target detection algorithm to determine the state of the target object in the image to be processed (idle, present, absent, in service), and outputs an image marked with a label for recording the state of the target object (i.e., the processing result). In addition. Figure 5 As shown, the detection device can also perform an early warning function based on the processing results. Figure 5As shown, since the target detection model is obtained by training the improved neural network model, that is, the target detection model needs to be trained. If it is in the training process of the target detection model, after the acquisition module 40 obtains the video stream data, it will create a data set (including training data, verification data, and test data), and use the data set to train the improved neural network model.
[0063] It should be noted that Figure 4 The preferred implementation of the illustrated embodiment can be found in Figure 2 The relevant description of the illustrated embodiment will not be repeated here.
[0064] An embodiment of the present application further provides a non-volatile storage medium, in which a computer program is stored, wherein the above personnel status detection method is executed by running the computer program on a device where the non-volatile storage medium is located.
[0065] The above-mentioned non-volatile storage medium is used to store programs that perform the following functions: obtaining video stream data of the area to be detected, wherein the video stream data is used to record the dynamic changes of the area to be detected; determining multiple images to be processed whose image similarity is less than a preset value from the video stream data, wherein each image to be processed is used to record a frame of the video stream data; using a target detection model to analyze the image to be processed to obtain a processing result, wherein the processing result is a target image marked with a label, and the label is used to indicate the state of the target object located in the area to be detected. The target detection model is obtained by training a target neural network model using historical video data of the area to be detected as training data, and the target neural network model is a neural network model with a high-dimensional convolutional layer deleted.
[0066] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the above personnel status detection method through the computer program.
[0067] The processor in the above-mentioned electronic device is used to run a program that performs the following functions: obtaining video stream data of the area to be detected, wherein the video stream data is used to record the dynamic changes of the area to be detected; determining multiple images to be processed whose image similarity is less than a preset value from the video stream data, wherein each image to be processed is used to record a frame of the video stream data; using a target detection model to analyze the image to be processed to obtain a processing result, wherein the processing result is a target image marked with a label, and the label is used to indicate the state of the target object located in the area to be detected. The target detection model is obtained by training a target neural network model using historical video data of the area to be detected as training data, and the target neural network model is a neural network model with a high-dimensional convolutional layer deleted.
[0068] The embodiment of the present application also provides a computer program product, including computer instructions, which implement the steps of the above personnel status detection method when executed by a processor.
[0069] It should be noted that the various modules in the above-mentioned personnel status detection device can be program modules (for example, a set of program instructions that implement a certain specific function) or hardware modules. For the latter, it can be expressed in the following forms, but is not limited to this: the expression form of each of the above-mentioned modules is a processor, or the functions of each of the above-mentioned modules are implemented by a processor.
[0070] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0071] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0072] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0073] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0074] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0075] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the relevant technology or all or part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk, etc. Various media that can store program codes.
[0076] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for detecting a person's status, characterized in that: include: Acquire video stream data of the area to be detected, wherein the video stream data is used to record dynamic changes of the area to be detected; Determine from the video stream data a plurality of images to be processed whose image similarity is less than a preset value, wherein each of the images to be processed is used to record a frame of the video stream data; The target detection model is used to analyze the image to be processed to obtain a processing result, wherein the processing result is a target image marked with a label, and the label is used to indicate the state of the target object located in the area to be detected. The target detection model is obtained by training a target neural network model using historical video data of the area to be detected as training data, and the target neural network model is a neural network model with a high-dimensional convolutional layer deleted.
2. The method according to claim 1, characterized in that Generating a plurality of to-be-processed images whose image similarities are less than a preset value according to the video data includes: intercepting and obtaining a plurality of initial images from the video stream data at a preset frequency; For each of the initial images, determining a target pixel of a target area in the initial image, wherein the target area is an area where the initial image overlaps with a template image, and the template image is a preset image that only includes a position to be detected in the area to be detected; Determining the image similarity between each two of the initial images according to the plurality of target pixels; The image to be processed is determined according to the image similarity.
3. The method according to claim 2, characterized in that Determining the image similarity of each two initial images according to the plurality of target pixels, comprising: determining a first coordinate of the target pixel in one of the initial images in the initial image pair to be compared, and a second coordinate of the target pixel in the other initial image; determining the image similarity according to the plurality of the first coordinates, the plurality of the second coordinates, and a conversion relationship between the two initial images in the initial image pair to be compared, wherein the conversion relationship comprises: scaling, rotation, and translation; Determining the target image according to the image similarity includes: determining a plurality of target initial images corresponding to a plurality of image similarities greater than or equal to the preset value; and deleting the target initial image from all the initial images to obtain a plurality of the target images.
4. The method according to claim 2, characterized in that: After determining the image similarity between each two of the initial images according to the plurality of target pixels, the method further comprises: Determine a group of initial images corresponding to the image similarity, and determine a target device corresponding to the group of initial images, wherein the target device is a device that generates video stream data to which the group of initial images belongs, and the number of the target devices is multiple; The image similarity is determined as the weight of the target device, wherein the weight is related to the calling order of the target device, and the calling order of the target device is the order of obtaining the video stream data from the target device when detecting the status of people in the area to be detected.
5. The method according to claim 1, characterized in that The target detection model is used to process and analyze the image to be processed to obtain a processing result, including: Normalizing and activating the image to be processed in the first neural network layer of the target detection model to obtain a first feature map output by the first neural network layer; Using a one-dimensional convolution kernel to perform convolution processing on the first feature map in the second neural network layer of the target detection model to obtain a convolution result, and performing the normalization and activation processing on the convolution result to obtain a second feature map output by the second neural network layer; Merging the first feature map and the second feature map into a target feature map to be marked; The target feature map is marked to generate the processing result.
6. The method according to claim 5, characterized in that Normalizing and activating the image to be processed in the first neural network layer of the target detection model to obtain a first feature map output by the first neural network layer, including: Extracting a feature vector of each of the images to be processed to obtain a plurality of feature vectors, wherein the feature vector is a vector generated according to feature information of the images to be processed, and the feature information at least includes: an edge of the image, a texture of the image, and a color of the image; Determine a mean value of the plurality of feature vectors, and convert each of the feature vectors into a feature value within a preset value range according to the mean value; Each pixel in the image to be processed is processed according to the characteristic value to obtain the first characteristic map.
7. The method according to claim 5, characterized in that Marking the target feature map to generate the processing result includes: Determine a plurality of prediction probabilities corresponding to each grid in the target feature map, wherein the plurality of grids constitute the target feature map, and the plurality of prediction probabilities are probabilities that the target object is in a plurality of states, wherein the states include: idle, present, absent, and in service; For each of the grids, determining a target prediction probability with the largest value among the plurality of prediction probabilities; The state corresponding to the target prediction probability is recorded as the label, and the grid is marked with the label.
8. The method according to claim 7, characterized in that After getting the processing results, it also includes: Determine whether a target tag exists in the target image, wherein the target tag includes: a tag whose recorded status is the idle tag and a tag whose recorded status is the away tag; In the case where the target tag exists in the target image, an alarm message is sent, wherein the alarm message is used to indicate that the target object in an abnormal state exists in the area to be detected.
9. A device for detecting a person's status, characterized in that: include: An acquisition module, used for acquiring video stream data of the area to be detected, wherein the video stream data is used for recording dynamic changes of the area to be detected; A determination module, used to determine a plurality of to-be-processed images whose image similarity is less than a preset value from the video stream data, wherein each of the to-be-processed images is used to record a frame of the video stream data; A processing module is used to use a target detection model to process and analyze the image to be processed to obtain a processing result, wherein the processing result is a target image marked with a label, and the label is used to indicate the state of the target object located in the area to be detected. The target detection model is obtained by training a target neural network model using historical video data of the area to be detected as training data, and the target neural network model is a neural network model with a high-dimensional convolutional layer deleted.
10. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a computer program, wherein the method for detecting the status of a person as claimed in any one of claims 1 to 8 is executed by running the computer program on the device where the non-volatile storage medium is located.
11. An electronic device comprising a memory and a processor, characterized in that: The memory stores a computer program, and the processor is configured to execute the method for detecting a person's status according to any one of claims 1 to 8 through the computer program.
12. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the method for detecting the status of a person described in any one of claims 1 to 8 are implemented.