Intelligent store checking system and method for IPC
Real-time video streams are obtained through IPC and image preprocessing and feature extraction are performed. Combined with UDP protocol transmission and image recognition, the cost and reliability problems of the existing IPC store viewing methods are solved, and a low-cost and high-reliability intelligent store viewing system is realized.
Patent Information
- Application Number
- CN202510328616.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-11
AI Technical Summary
The existing IPC store viewing methods cannot meet the needs of low-cost, high-reliability real-time monitoring and customer flow statistics, especially in terms of labor costs and 24-hour monitoring and review.
Real-time video stream is obtained through IPC, and face and hand recognition is used to use image preprocessing, feature extraction, image recognition and marking modules to track hand movements and judge abnormalities. UDP protocol transmits video stream data to achieve a low-cost and highly reliable store viewing system.
It realizes low-cost and high-reliability store-viewing needs, improves the reliability of video streaming data transmission and the accuracy of image feature extraction, and provides accurate customer flow statistics and abnormal judgment capabilities.
Smart Images

Figure CN120298945A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of integrated circuits and electronic communications, and particularly relates to an intelligent store-watching system and method for IPC. Background Art
[0002] With the update and development of computer technology, artificial intelligence technology has been applied in more and more industries, and AI algorithms have penetrated into every corner of life. Applications related to artificial intelligence analysis have quickly entered people's lives, and the use of IPC (Internet Protocol Camera) with intelligent store-watching algorithms has become the mainstream method for various offline stores to monitor their stores. The existing methods for store clerks to monitor stores can no longer meet the needs of users, as customers have requirements in aspects such as labor costs, real-time monitoring and viewing, 24-hour monitoring and playback, statistics of customer flow entering the store, and backtracking of intrusion alarm information in sensitive areas.
[0003] Therefore, how to develop an intelligent store-watching system for IPC to meet the requirements of low-cost and high-reliability store-watching is a technical problem that urgently needs to be solved at present. Summary of the Invention
[0004] The purpose of the present invention is to provide an intelligent store-watching system and method for IPC to meet the requirements of low-cost and high-reliability store-watching.
[0005] To solve the above technical problems, the technical solutions adopted by the present invention are as follows:
[0006] In the first aspect, an intelligent store-watching method for IPC is provided, including the following steps:
[0007] S1: Obtain a real-time video stream through the IPC, transmit the real-time video stream data to a remote management terminal, and the remote management terminal obtains real-time image frame data based on the real-time video stream;
[0008] S2: The image preprocessing module preprocesses the real-time image frame data;
[0009] S3: The feature extraction module extracts image features from the preprocessed real-time image frame data;
[0010] S4: The image recognition module recognizes the faces and hands in the image frame based on the extracted image features;
[0011] S5: Add a unique label to the same face in the image frame through the labeling module, perform face tracking based on the unique label, track hand movements based on the face tracking data, and determine whether there is an abnormality according to the tracked hand movements. If so, execute step S6; if not, continue tracking;
[0012] S6: Determine the abnormal category and the severity of the abnormality based on the tracked hand movements, transmit the abnormal category to the warning module, and the warning module informs the management personnel of the abnormal category and the severity of the abnormality.
[0013] Preferably, the specific process of transmitting the real-time video stream data to the remote management terminal in step S1 is as follows:
[0014] S11: Create a socket based on the UDP protocol and encode and compress the video stream data;
[0015] S12: Package the encoded and compressed video stream data into an RTP packet. The RTP packet includes a fixed header, an extension header, and a video stream payload;
[0016] S13: Send the RTP packet through the socket based on the UDP protocol based on the specified target IP address and port;
[0017] S14: The remote management terminal receives the RTP packet through the socket based on the UDP protocol. After receiving the RTP packet, it unpacks the packet and extracts the video stream payload data.
[0018] Preferably, the specific process of the image preprocessing module preprocessing the real-time image frame data in step S2 is as follows:
[0019] S21: For each pixel point in each frame of the image, obtain the pixel values of the pixel points in the specified-sized neighborhood around it, take the average value of the pixel values of the pixel points in the neighborhood, and use the average value as the pixel value of the pixel point;
[0020] S23: Obtain the two-dimensional data matrix f(x, y) of each image frame, obtain the coefficient matrix [A] of the discrete cosine transform for the two-dimensional data matrix f(x, y), and obtain its corresponding transpose matrix [A] based on the coefficient matrix [A] of the discrete cosine transform; T ;
[0021] S24: Calculate the discrete cosine transform value according to the formula [A]f(x, y)[A]; T Perform a frequency domain transformation on the image frame according to the calculated discrete cosine transform value;
[0022] S25: Perform a correction process on the image frame, map the R, G, B values of each pixel point of the image from the range of 0 - 255 to a non-linear space of 0 - 1, perform correction through a preset power function, calculate the new pixel value, and map the new pixel value back to the range of 0 - 255;
[0023] S26: Convert the corrected image frame into a grayscale image.
[0024] Preferably, the specific process of extracting image features from the preprocessed real-time image frame data in step S3 is as follows:
[0025] S31: Initially calibrate the area to be subjected to image feature extraction and define the size of the pixel neighborhood;
[0026] S32: Take each pixel in the area of image feature extraction as the central pixel, obtain its corresponding neighborhood, calculate the gray value of the pixels in the neighborhood, and compare the gray value of the pixels in the neighborhood with the gray value of the central pixel in the neighborhood;
[0027] S33: If the gray value of the pixels in the neighborhood is greater than or equal to the gray value of the central pixel in the neighborhood, set the binary bit at the corresponding position of the central pixel to 1; if the gray value of the pixels in the neighborhood is less than the gray value of the central pixel in the neighborhood, set the binary bit at the corresponding position of the central pixel to 0, to obtain the binary coding of the area of image feature extraction;
[0028] S34: Calculate the frequency distribution of the binary coding of each pixel in the area of image feature extraction in this area, and obtain the texture features in the area of image feature extraction.
[0029] Preferably, after obtaining the texture features in the area of image feature extraction in step S34, extract the edge features of the image. The specific process is as follows:
[0030] S35: Define two convolution templates of specified sizes, the x-axis convolution template and the y-axis convolution template, which are respectively used to detect the edges in the x-axis direction and the y-axis direction of the area of image feature extraction;
[0031] S36: Perform pixel scanning in the area of image feature extraction, and perform convolution operations on each pixel with the x-axis convolution template and the y-axis convolution template respectively. The result of the convolution operation of the pixel with the x-axis convolution template is the edge strength in the x-axis direction of the area of image feature extraction, and the result of the convolution operation of the pixel with the y-axis convolution template is the edge strength in the y-axis direction of the area of image feature extraction;
[0032] S37: Calculate the gradient amplitude of each pixel point according to the results of the convolution operations of each pixel with the x-axis convolution template and the y-axis convolution template respectively, set a gradient amplitude threshold, and mark the pixel points with a gradient amplitude greater than the gradient amplitude threshold as edge pixel points.
[0033] Preferably, the image recognition module is a neural network model. The specific process of the image recognition module for face and hand recognition is as follows:
[0034] S41: Input image samples with specified numbers of labeled facial and hand features into the image recognition module for model training, and verify the trained image recognition module to determine whether the model performance meets the preset requirements. If so, execute step S42; if not, retrain the model until the model performance meets the preset requirements.
[0035] S42: Input the image frame data after image feature extraction in step S3 into the image recognition module. The image recognition module recognizes the face and hand based on the extracted image features and outputs the face recognition result and the hand recognition result.
[0036] S43: Statistic the passenger flow data within a specified time period according to the face recognition result, and obtain the action information according to the hand recognition result.
[0037] In a second aspect, a smart store - watching system for IPC is provided to implement the described smart store - watching method for IPC, including an IPC, a data transmission module, an image pre - processing module, a feature extraction module, an image recognition module, a marking module, an anomaly judgment module, and a warning module. The IPC is connected to the data transmission module, the data transmission module is connected to the image pre - processing module, the image pre - processing module is connected to the feature extraction module, the feature extraction module is connected to the image recognition module, the image recognition module is connected to the marking module, the marking module is connected to the anomaly judgment module, and the anomaly judgment module is connected to the warning module;
[0038] The IPC is used to obtain a real - time video stream;
[0039] The data transmission module is used to transmit the real - time video stream data to a remote management terminal;
[0040] The image pre - processing module is used to pre - process the real - time image frame data;
[0041] The feature extraction module is used to extract image features from the pre - processed real - time image frame data;
[0042] The image recognition module is used to recognize the face and hand in the image frame based on the extracted image features;
[0043] The marking module is used to add a unique mark to the same face in the image frame and perform face tracking based on the unique mark;
[0044] The anomaly judgment module is used to judge whether there is an anomaly according to the tracked hand action, judge the anomaly category and the anomaly severity according to the tracked hand action, and transmit the anomaly category to the warning module;
[0045] The warning module is used to inform the management personnel of the abnormal category and the severity of the abnormality.
[0046] The beneficial effects of the present invention include:
[0047] The intelligent store monitoring system and method for IPC provided by the present invention obtain a real-time video stream through an IPC and transmit it to a remote management terminal. After the remote management terminal obtains the real-time image frame data, preprocessing is performed; the feature extraction module extracts image features from the preprocessed real-time image frame data; the image recognition module recognizes the faces and hands in the image frame based on the extracted image features; the same faces in the image frame are added with unique marks through the marking module, face tracking is performed based on the unique marks, the hand movements are tracked based on the face tracking data, it is judged whether there are abnormalities in the tracked hand movements, the abnormal category and the severity of the abnormality are judged according to the tracked hand movements, the abnormal category is transmitted to the warning module, and the process of the warning module informing the management personnel of the abnormal category and the severity of the abnormality realizes the low-cost and highly reliable store monitoring requirements.
[0048] First, a socket based on the UDP protocol is created, and the video stream data is encoded, compressed and encapsulated into an RTP packet; the RTP packet is sent through the socket based on the UDP protocol based on the specified target IP address and port; the remote management terminal receives the RTP packet through the socket based on the UDP protocol, and after de-encapsulating the RTP packet, the video stream payload data is extracted, realizing the secure transmission of the video stream data of the terminal IPC and improving the reliability of the video stream data transmission.
[0049] Secondly, by obtaining the pixel values of each pixel point in a specified-sized neighborhood around it, taking the average of the pixel values of each pixel point in the neighborhood as the pixel value of the pixel point, obtaining the two-dimensional data matrix of each image frame, obtaining the coefficient matrix of the discrete cosine transform for the two-dimensional data matrix, and obtaining the corresponding transpose matrix based on the coefficient matrix of the discrete cosine transform; calculating the discrete cosine transform value and performing frequency domain transformation; performing correction calculation on the image frame to obtain new pixel values, mapping the new pixel values back to the range of 0-255 from the non-linear space; the preprocessing process of converting the corrected image frame into a grayscale image provides a data basis for subsequent image feature extraction and recognition, and can improve the accuracy of the image feature extraction and recognition results.
[0050] Again, by preliminarily calibrating the area where image feature extraction is to be performed, taking each pixel within the area of image feature extraction as the central pixel, calculating the gray values of the pixels within the neighborhood, comparing the gray values of the pixels within the neighborhood with the gray value of the central pixel within the neighborhood, setting the corresponding binary bits according to the comparison result, and obtaining the binary encoding of the area of image feature extraction; calculating the frequency distribution of the binary encoding of each pixel in the area of image feature extraction within this area, and obtaining the texture features within this area of image feature extraction, the accurate extraction of the texture features within this area is realized.
[0051] Again, by defining two convolution templates of specified sizes, performing pixel scanning in the area of image feature extraction, performing convolution operations on each pixel with the two convolution templates respectively, calculating the gradient amplitude of each pixel point according to the convolution operation result of each pixel, setting a gradient amplitude threshold, and marking the pixel points with a gradient amplitude greater than the gradient amplitude threshold as edge pixel points, the accurate extraction of edge features is realized.
[0052] Finally, inputting a specified number of image samples with labeled face and hand features into the image recognition module for model training, inputting the image frame data after image feature extraction into the image recognition module to recognize the face and hand, counting the passenger flow data within a specified time period according to the face recognition result, and obtaining the action information according to the hand recognition result, realizing the accurate recognition of the face and hand, providing an accurate data basis for subsequent passenger flow statistics and anomaly judgment, and facilitating the realization of intelligent store monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 It is a schematic flowchart of the intelligent store monitoring method for IPC according to the present invention.
[0054] Figure 2 It is a schematic flowchart of the image feature extraction according to the present invention.
[0055] Figure 3 It is a schematic architecture diagram of the intelligent store monitoring system for IPC according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] The following further describes the present invention in detail with reference to the attached Figures 1 to 3 :
[0057] Embodiment 1
[0058] Referring to the attached Figure 1 As shown, an intelligent store monitoring method for IPC includes the following steps:
[0059] S1: Obtain a real-time video stream through IPC, transmit the real-time video stream data to a remote management terminal, and the remote management terminal obtains real-time image frame data based on the real-time video stream;
[0060] S2: The image preprocessing module preprocesses the real-time image frame data;
[0061] S3: The feature extraction module extracts image features from the preprocessed real-time image frame data;
[0062] S4: The image recognition module recognizes the faces and hands in the image frame based on the extracted image features;
[0063] S5: The marking module adds a unique mark to the same face in the image frame, performs face tracking based on the unique mark, tracks hand movements based on the face tracking data, and determines whether there is an abnormality according to the tracked hand movements. If so, execute step S6; if not, continue tracking;
[0064] S6: Determine the abnormality category and severity level according to the tracked hand movements, transmit the abnormality category to the warning module, and the warning module informs the management personnel of the abnormality category and severity level.
[0065] Since in the prior art, store watching requires management personnel to watch the store on-site. However, during some periods, the customer flow is large, and during some periods, the customer flow is small. If personnel are required to watch the store on-site all the time, the labor cost is very high and the efficiency is very low. Therefore, the present invention obtains a real-time video stream through an IPC and transmits it to a remote management terminal. After the remote management terminal obtains the real-time image frame data, it performs preprocessing; the feature extraction module extracts image features from the preprocessed real-time image frame data; the image recognition module recognizes the faces and hands in the image frame based on the extracted image features; the marking module adds a unique mark to the same face in the image frame, performs face tracking based on the unique mark, tracks hand movements based on the face tracking data, determines whether there is an abnormality in the tracked hand movements, determines the abnormality category and severity level according to the tracked hand movements, transmits the abnormality category to the warning module, and the warning module informs the management personnel of the abnormality category and severity level, thus realizing the low-cost and highly reliable store watching requirements.
[0066] Embodiment 2
[0067] Based on Embodiment 1, the specific process of transmitting the real-time video stream data in step S1 to the remote management terminal is as follows:
[0068] S11: Create a socket based on the UDP protocol and encode and compress the video stream data;
[0069] S12: Package the encoded and compressed video stream data into an RTP packet. The RTP packet includes a fixed header, an extension header, and a video stream payload;
[0070] S13: Send the RTP packet through a socket based on the UDP protocol based on the specified target IP address and port;
[0071] S14: After the remote management terminal receives the RTP packet through a socket based on the UDP protocol, it performs decapsulation on the received RTP packet and extracts the video stream payload data.
[0072] The above process creates a socket based on the UDP protocol, encodes and compresses the video stream data and encapsulates it into an RTP packet; sends the RTP packet through a socket based on the UDP protocol based on the specified target IP address and port; the remote management terminal receives the RTP packet through a socket based on the UDP protocol and performs decapsulation on the received RTP packet to extract the video stream payload data, realizing the secure transmission of the video stream data of the terminal IPC and improving the reliability of the video stream data transmission.
[0073] In this embodiment, the specific process of the image preprocessing module in step S2 for preprocessing the real-time image frame data is as follows:
[0074] S21: For each pixel point in each frame of the image, obtain the pixel values of each pixel point in the specified-sized neighborhood around it, take the average of the pixel values of each pixel point in the neighborhood, and use the average as the pixel value of the pixel point;
[0075] S23: Obtain the two-dimensional data matrix f(x, y) of each image frame, obtain the coefficient matrix [A] of the discrete cosine transform for the two-dimensional data matrix f(x, y), and obtain its corresponding transposed matrix [A] based on the coefficient matrix [A] of the discrete cosine transform T ;
[0076] S24: Calculate the discrete cosine transform value according to the formula [A]f(x, y)[A] T Perform frequency domain transformation on the image frame according to the calculated discrete cosine transform value;
[0077] S25: Perform correction processing on the image frame, map the R, G, B values of each pixel point of the image from the range of 0 - 255 to the non-linear space of 0 - 1, perform correction through a preset power function, calculate the new pixel value, and map the new pixel value back to the range of 0 - 255;
[0078] S26: Convert the corrected image frame into a grayscale image.
[0079] By obtaining the pixel values of each pixel point within a specified-size neighborhood around it, taking the average of the pixel values of each pixel point in the neighborhood as the pixel value of the pixel point, obtaining the two-dimensional data matrix of each image frame, obtaining the coefficient matrix of the discrete cosine transform for the two-dimensional data matrix, and obtaining its corresponding transpose matrix based on the coefficient matrix of the discrete cosine transform; calculating the discrete cosine transform value and performing frequency-domain transformation; performing correction calculation on the image frame to obtain new pixel values, and mapping the new pixel values back to the range of 0 - 255 from the non-linear space; the preprocessing process of converting the corrected image frame into a grayscale image provides a data basis for subsequent image feature extraction and recognition, and can improve the accuracy of image feature extraction and recognition results.
[0080] Embodiment 3
[0081] Based on Embodiment 1 or Embodiment 2, the specific process of image feature extraction from the preprocessed real-time image frame data in step S3 is as follows:
[0082] S31: Initially calibrate the region to be subjected to image feature extraction, and define the size of the pixel point neighborhood;
[0083] S32: For each pixel in the region to be subjected to image feature extraction as the central pixel, obtain its corresponding neighborhood, calculate the grayscale values of the pixels in the neighborhood, and compare the grayscale values of the pixels in the neighborhood with the grayscale value of the central pixel in the neighborhood;
[0084] S33: If the grayscale value of the pixel in the neighborhood is greater than or equal to the grayscale value of the central pixel in the neighborhood, set the binary bit at the corresponding position of the central pixel to 1. If the grayscale value of the pixel in the neighborhood is less than the grayscale value of the central pixel in the neighborhood, set the binary bit at the corresponding position of the central pixel to 0, to obtain the binary coding of the region to be subjected to image feature extraction;
[0085] S34: Calculate the frequency distribution of the binary coding of each pixel in the region to be subjected to image feature extraction in this region, and obtain the texture features within this region to be subjected to image feature extraction.
[0086] The above process initially calibrates the region to be subjected to image feature extraction, takes each pixel in the region to be subjected to image feature extraction as the central pixel, calculates the grayscale values of the pixels in the neighborhood, compares the grayscale values of the pixels in the neighborhood with the grayscale value of the central pixel in the neighborhood, sets the corresponding binary bits according to the comparison results, to obtain the binary coding of the region to be subjected to image feature extraction; calculates the frequency distribution of the binary coding of each pixel in the region to be subjected to image feature extraction in this region, and obtains the texture features within this region to be subjected to image feature extraction, realizing the accurate extraction of the texture features within this region.
[0087] After obtaining the texture features within the region for image feature extraction in step S34, the edge features of the image are extracted. The specific process is as follows:
[0088] S35: Define two convolution templates of specified sizes, an x-axis convolution template and a y-axis convolution template, which are respectively used to detect the edges of the region for image feature extraction in the x-axis direction and the y-axis direction;
[0089] S36: Perform pixel scanning in the region for image feature extraction, and perform convolution operations on each pixel with the x-axis convolution template and the y-axis convolution template respectively. The result of the convolution operation of the pixel with the x-axis convolution template is the edge intensity of the region for image feature extraction in the x-axis direction, and the result of the convolution operation of the pixel with the y-axis convolution template is the edge intensity of the region for image feature extraction in the y-axis direction;
[0090] S37: Calculate the gradient magnitude of each pixel point according to the results of the convolution operations of each pixel with the x-axis convolution template and the y-axis convolution template respectively. Set a gradient magnitude threshold, and mark the pixel points with a gradient magnitude greater than the gradient magnitude threshold as edge pixel points.
[0091] By defining two convolution templates of specified sizes, performing pixel scanning in the region for image feature extraction, performing convolution operations on each pixel with the two convolution templates respectively, calculating the gradient magnitude of each pixel point according to the convolution operation results of each pixel, setting a gradient magnitude threshold, and marking the pixel points with a gradient magnitude greater than the gradient magnitude threshold as edge pixel points, the accurate extraction of edge features is achieved.
[0092] Embodiment 4
[0093] Based on Embodiment 1 or Embodiment 2 or Embodiment 3, the image recognition module is a neural network model. The specific process of the image recognition module for face and hand recognition is as follows:
[0094] S41: Input a specified number of image samples with labeled face and hand features into the image recognition module for model training, and verify the trained image recognition module to determine whether the model performance meets the preset requirements. If so, execute step S42; if not, re-perform model training until the model performance meets the preset requirements;
[0095] S42: Input the image frame data after image feature extraction in step S3 into the image recognition module. The image recognition module recognizes the face and hand based on the extracted image features and outputs the face recognition result and the hand recognition result;
[0096] S43: Statistically calculate the passenger flow data within a specified time period according to the face recognition result, and obtain the action information according to the hand recognition result.
[0097] The above process trains the model by inputting a specified number of image samples with labeled face and hand features into the image recognition module. The image frame data after image feature extraction is input into the image recognition module to recognize faces and hands. The passenger flow data within a specified time period is counted based on the face recognition results, and the action information is obtained based on the hand recognition results, achieving accurate recognition of faces and hands, providing an accurate data basis for subsequent passenger flow statistics and anomaly judgment, and facilitating the realization of intelligent store monitoring.
[0098] In another implementation manner of this embodiment, the passenger flow data statistics function for entering and leaving the store is realized by obtaining the real-time video stream of the IPC and performing intelligent recognition and analysis on each frame of image data. For example: when a consumer enters the store from outside the store and the algorithm recognizes this event, the upper-layer application of the IPC increments the count of the number of people entering by one. Correspondingly, when a consumer exits the store from inside the store, the upper-layer application of the IPC increments the count of the number of people leaving by one. The statistical time interval can be configured to report to the cloud platform once every 1 minute, or the reporting interval can be customized. The store manager can view the daily passenger flow data through the mobile app.
[0099] The area intrusion module mainly realizes the function of monitoring key areas of concern by obtaining the real-time video stream of the IPC and performing intelligent recognition and analysis on each frame of image data. For example: when there are certain areas that the store manager is inconvenient to supervise, the area information can be sent to the IPC through the platform at this time, and then the area monitoring algorithm is initialized. When someone enters this area, the IPC pushes the captured picture and alarm information to the remote management terminal, and the store manager can trace this abnormal situation through the mobile app.
[0100] An intelligent store monitoring system for IPC, used to implement the described intelligent store monitoring method for IPC, includes an IPC, a data transmission module, an image preprocessing module, a feature extraction module, an image recognition module, a marking module, an anomaly judgment module, and an early warning module. The IPC is connected to the data transmission module, the data transmission module is connected to the image preprocessing module, the image preprocessing module is connected to the feature extraction module, the feature extraction module is connected to the image recognition module, the image recognition module is connected to the marking module, the marking module is connected to the anomaly judgment module, and the anomaly judgment module is connected to the early warning module. Offline store operators use this system for intelligent monitoring and management of the store, improving the security and reliability of the store, and meeting the user's requirements in aspects such as labor cost, real-time monitoring and viewing, 24-hour monitoring and playback, in-store passenger flow statistics, and backtracking of intrusion alarm information in sensitive areas.
[0101] The IPC is used to obtain a real-time video stream. The data transmission module is used to transmit the real-time video stream data to a remote management terminal. The image preprocessing module is used to preprocess the real-time image frame data. The feature extraction module is used to extract image features from the preprocessed real-time image frame data. The image recognition module is used to recognize faces and hands in the image frame based on the extracted image features. The marking module is used to add a unique mark to the same face in the image frame and perform face tracking based on the unique mark. The anomaly judgment module is used to judge whether there is an anomaly according to the tracked hand movements, judge the anomaly category and anomaly severity according to the tracked hand movements, and transmit the anomaly category to the warning module. The warning module is used to inform the management personnel of the anomaly category and anomaly severity.
[0102] In another implementation manner of this embodiment, the IPC device can implement passenger flow statistics and automatic switching of the regional intrusion algorithm in different time periods. The time periods can all be configured in the mobile app. The passenger flow statistics time period can be configured from 8:00 am to 10:00 pm, and the regional intrusion time period can be configured from 10:00 pm to 8:00 am the next day. During the business hours, the passenger flow data information is mainly concerned. Outside the business hours, the monitoring alarm messages of sensitive areas are mainly concerned. After the time periods are configured, the platform can send an algorithm switching command in the corresponding time period, and the IPC switches to run the corresponding AI algorithm after receiving the command. In addition, the alarm audio content of both algorithms can be customized. For example, when entering the store, "Welcome" can be broadcast, and when leaving the store, "Welcome next time" can be broadcast; when the regional intrusion alarm is triggered, "You have entered the warning area" can be broadcast. The operators of offline stores use this system for intelligent store monitoring management, which improves the safety and reliability of the store, reduces the labor cost, and also enhances the shopping experience of consumers.
[0103] In summary, for the intelligent store viewing system and method provided by the present invention, the IPC obtains a real-time video stream and transmits it to a remote management terminal. After the remote management terminal obtains the real-time image frame data, it performs preprocessing. The feature extraction module extracts image features from the preprocessed real-time image frame data. The image recognition module recognizes faces and hands in the image frame based on the extracted image features. A unique mark is added to the same face in the image frame through the marking module, and face tracking is performed based on the unique mark. The hand movements are tracked based on the face tracking data, and it is judged whether there is an anomaly in the tracked hand movements. The anomaly category and anomaly severity are judged according to the tracked hand movements, and the anomaly category is transmitted to the warning module. The warning module informs the management personnel of the anomaly category and anomaly severity, thus realizing the low-cost and highly reliable store viewing requirements.
[0104] By creating a socket based on the UDP protocol, encoding, compressing, and encapsulating video stream data into RTP packets; sending the RTP packets through the UDP - based socket based on the specified destination IP address and port; the remote management terminal receives the RTP packets through the UDP - based socket, unpacks them after receiving, and extracts the video stream payload data, thus realizing the secure transmission of the video stream data of the terminal IPC and improving the reliability of the video stream data transmission. By obtaining the pixel values of each pixel point within a specified - sized neighborhood around it, taking the average of the pixel values of each pixel point in the neighborhood as the pixel value of the pixel point, obtaining the two - dimensional data matrix of each image frame, obtaining the coefficient matrix of the discrete cosine transform for the two - dimensional data matrix, and obtaining its corresponding transpose matrix based on the coefficient matrix of the discrete cosine transform; calculating the discrete cosine transform value and performing frequency - domain transformation; performing correction calculation on the image frame to obtain new pixel values, mapping the new pixel values back to the range of 0 - 255 from the non - linear space; the pre - processing process of converting the corrected image frame into a grayscale image provides a data basis for subsequent image feature extraction and recognition and can improve the accuracy of image feature extraction and recognition results.
[0105] By preliminarily calibrating the region to be subjected to image feature extraction, taking each pixel in the region of image feature extraction as the central pixel, calculating the grayscale values of the pixels in the neighborhood, comparing the grayscale values of the pixels in the neighborhood with the grayscale value of the central pixel in the neighborhood, and setting the corresponding binary bits according to the comparison result, the binary encoding of the region of image feature extraction is obtained; calculating the frequency distribution of the binary encoding of each pixel in the region of image feature extraction in this region, obtaining the texture features within this region of image feature extraction, and realizing the accurate extraction of the texture features within this region. By defining two convolution templates of specified sizes, performing pixel scanning in the region of image feature extraction, performing convolution operations on each pixel with the two convolution templates respectively, calculating the gradient amplitude of each pixel point according to the convolution operation results of each pixel, setting a gradient amplitude threshold, and marking the pixel points with gradient amplitudes greater than the gradient amplitude threshold as edge pixel points, the accurate extraction of edge features is realized. Inputting a specified number of image samples with labeled face and hand features into the image recognition module for model training, inputting the image frame data after image feature extraction into the image recognition module to recognize faces and hands, counting the passenger flow data within a specified time period according to the face recognition result, and obtaining action information according to the hand recognition result, realizing the accurate recognition of faces and hands, providing an accurate data basis for subsequent passenger flow statistics and anomaly judgment, and facilitating the realization of intelligent store monitoring.
Claims
1. An intelligent store monitoring method for IPC, characterized in that, Including the following steps: S1: Obtain a real-time video stream through IPC, transmit the real-time video stream data to a remote management terminal, and the remote management terminal obtains real-time image frame data based on the real-time video stream; S2: The image preprocessing module preprocesses the real-time image frame data; S3: The feature extraction module extracts image features from the preprocessed real-time image frame data; S4: The image recognition module recognizes the faces and hands in the image frame based on the extracted image features; S5: Add a unique mark to the same face in the image frame through the marking module, perform face tracking based on the unique mark, track hand movements based on face tracking data, and determine whether there is an abnormality according to the tracked hand movements. If so, execute step S6; if not, continue tracking; S6: Determine the abnormality category and severity level according to the tracked hand movements, transmit the abnormality category to the warning module, and the warning module informs the management personnel of the abnormality category and severity level.
2. The intelligent store monitoring method for IPC according to claim 1, wherein, The specific process of transmitting the real-time video stream data in step S1 to the remote management terminal is as follows: S11: Create a socket based on the UDP protocol and encode and compress the video stream data; S12: Package the encoded and compressed video stream data into an RTP packet. The RTP packet includes a fixed header, an extension header, and a video stream payload; S13: Send the RTP packet through the socket based on the UDP protocol based on the specified destination IP address and port; S14: The remote management terminal receives the RTP packet through the socket based on the UDP protocol, unpacks it after receiving the RTP packet, and extracts the video stream payload data.
3. The intelligent store monitoring method for IPC according to claim 1, characterized in that The specific process of the image preprocessing module in step S2 preprocessing the real-time image frame data is as follows: S21: For each pixel point in each frame of the image, obtain the pixel values of each pixel point in the specified-size neighborhood around it, take the average of the pixel values of each pixel point in the neighborhood, and use the average value as the pixel value of the pixel point; S23: Obtain the two-dimensional data matrix f(x, y) of each image frame, obtain the coefficient matrix [A] of the discrete cosine transform for the two-dimensional data matrix f(x, y), and obtain its corresponding transposed matrix [A] based on the coefficient matrix [A] of the discrete cosine transform T ; S24: Calculate the discrete cosine transform value according to the formula [A]f(x, y)[A] T Perform a frequency domain transformation on the image frame based on the calculated discrete cosine transform value; S25: Perform correction processing on the image frame, map the R, G, B values of each pixel point in the image from the range of 0-255 to a non-linear space of 0-1, perform correction through a preset power function, calculate the new pixel value, and map the new pixel value back to the range of 0-255; S26: Convert the corrected image frame into a grayscale image.
4. The intelligent store monitoring method for IPC according to claim 1, wherein, The specific process of extracting image features from the preprocessed real-time image frame data in step S3 is as follows: S31: Initially calibrate the area to be subjected to image feature extraction and define the size of the pixel point neighborhood; S32: Take each pixel in the area for image feature extraction as the central pixel, obtain its corresponding neighborhood, calculate the grayscale value of the pixels in the neighborhood, and compare the grayscale value of the pixels in the neighborhood with the grayscale value of the central pixel in the neighborhood; S33: If the grayscale value of the pixels in the neighborhood is greater than or equal to the grayscale value of the central pixel in the neighborhood, set the binary bit corresponding to the central pixel to 1. If the grayscale value of the pixels in the neighborhood is less than the grayscale value of the central pixel in the neighborhood, set the binary bit corresponding to the central pixel to 0, to obtain the binary encoding of the region for image feature extraction; S34: Calculate the frequency distribution of the binary encoding of each pixel in the region for image feature extraction in this region, and obtain the texture feature within this region for image feature extraction.
5. The intelligent store monitoring method for IPC according to claim 4, characterized in that, After obtaining the texture feature within the region for image feature extraction in step S34, extract the edge feature of the image. The specific process is as follows: S35: Define two convolution templates of specified sizes, an x-axis convolution template and a y-axis convolution template, which are respectively used to detect the edges of the region for image feature extraction in the x-axis direction and the y-axis direction; S36: Perform pixel scanning in the region for image feature extraction, and perform convolution operations on each pixel with the x-axis convolution template and the y-axis convolution template respectively. The result of the convolution operation of the pixel with the x-axis convolution template is the edge intensity of the region for image feature extraction in the x-axis direction, and the result of the convolution operation of the pixel with the y-axis convolution template is the edge intensity of the region for image feature extraction in the y-axis direction; S37: Calculate the gradient magnitude of each pixel point according to the results of the convolution operations of each pixel with the x-axis convolution template and the y-axis convolution template respectively. Set a gradient magnitude threshold, and mark the pixel points with a gradient magnitude greater than the gradient magnitude threshold as edge pixel points.
6. The intelligent store-watching method for IPC according to claim 1, characterized in that, The image recognition module is a neural network model. The specific process of the image recognition module for face and hand recognition is as follows: S41: Input a specified number of image samples with labeled face and hand features into the image recognition module for model training, and verify the trained image recognition module to determine whether the model performance meets the preset requirements. If so, execute step S42. If not, re-perform model training until the model performance meets the preset requirements; S42: Input the image frame data after image feature extraction in step S3 into the image recognition module. The image recognition module recognizes the face and hand based on the extracted image features, and outputs the face recognition result and the hand recognition result; S43: Count the passenger flow data within a specified time period according to the face recognition result, and obtain the action information according to the hand recognition result.
7. An intelligent store monitoring system for IPC, which is used to implement an intelligent store monitoring method for IPC described in any one of claims 1-6, characterized in that, It includes an IPC, a data transmission module, an image preprocessing module, a feature extraction module, an image recognition module, a marking module, an anomaly judgment module, and a warning module. The IPC is connected to the data transmission module, the data transmission module is connected to the image preprocessing module, the image preprocessing module is connected to the feature extraction module, the feature extraction module is connected to the image recognition module, the image recognition module is connected to the marking module, the marking module is connected to the anomaly judgment module, and the anomaly judgment module is connected to the warning module; The IPC is used to obtain the real-time video stream; The data transmission module is used to transmit the real-time video stream data to the remote management terminal; The image preprocessing module is used to preprocess the real-time image frame data; The feature extraction module is used to extract image features from the preprocessed real-time image frame data; The image recognition module is used to recognize the faces and hands in the image frame based on the extracted image features; The marking module is used to add a unique mark to the same face in the image frame and perform face tracking based on the unique mark; The anomaly judgment module is used to judge whether there is an anomaly according to the tracked hand movements, judge the anomaly category and anomaly severity according to the tracked hand movements, and transmit the anomaly category to the warning module; The warning module is used to inform the management personnel of the anomaly category and anomaly severity.