A method and device for generating a heat map of a hall and store scene with multi-angle fusion

Through multi-angle shooting and three-dimensional model reconstruction technology, a multi-angle-fusion thermal map of the store scene is generated, which solves the problem that a single camera cannot deeply analyze the human posture and movements, and realizes accurate analysis of the passenger flow density area and customer preference identification, improving operational efficiency and customer experience.

CN119229385BActive Publication Date: 2025-06-24CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411774225.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-05
Publication Date
2025-06-24
Estimated Expiration
2044-12-05

AI Technical Summary

Technical Problem

In the prior art, the heat map generated by a single camera cannot accurately capture important areas and cannot deeply analyze human postures and movements, resulting in the inability to understand customer behavior and demand information, affecting the effectiveness of marketing activities.

Method used

Multi-angle shooting is used to collect video data, identify humanoid information and key points of the human body through the Yolov8 algorithm, reconstruct the three-dimensional scene model, map the humanoid information to the three-dimensional model, generate a heat map, and analyze the passenger flow density area.

Benefits of technology

It realizes accurate analysis of the traffic density area in the store, and can accurately determine customers' shopping preferences and preference areas, improving operational efficiency, customer experience and customer satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119229385B_ABST
    Figure CN119229385B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of passenger flow data processing. A method and device for generating a heat map of a hall store scene with multi-angle fusion are provided. The method includes: collecting multi-angle video data through multi-angle shooting, and extracting an image time series at the same moment; using the Yolov8 algorithm to process the extracted image time series to identify human form information; reconstructing a three-dimensional scene model, including using the SIFT algorithm to extract image features and perform feature matching on the image time series for object surface reconstruction and texture mapping to obtain a reconstructed three-dimensional model of the hall store scene; identifying human form information in a specified historical time period, forming human form data by using the reconstructed three-dimensional model of the hall store scene, and generating a heat map of the hall store scene; and obtaining a passenger flow density area corresponding to the data to be analyzed according to the generated heat map of the hall store scene. The present invention determines the passenger flow density area in the hall store and determines customers' shopping preferences and habits.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of passenger flow data analysis, and provides a method and device for generating a heat map of a hall store scene with multi-angle fusion. Background Art

[0002] With the continuous development of information technology, as an intuitive data visualization tool, the heat map has been widely used in various hall stores such as retail, catering, and commercial complexes. By performing heat map analysis on the passenger flow data in the hall store, it can help merchants better understand customers' behavior habits, consumption preferences, and the rationality of the store layout, so as to optimize the allocation of in-store resources, enhance the customer experience, and achieve sales growth.

[0003] However, the existing technologies often rely on a single camera to generate a heat map. Due to the viewing angle limitation of the camera, some important areas cannot be accurately captured. This limitation may miss important information and even lead to incomplete data in the heat map. In addition, the heat map generated by a single camera usually cannot conduct in-depth analysis of human postures and actions. The lack of analysis of human postures makes it impossible to understand customers' behavior and demand information, thereby affecting the effect of formulating personalized marketing activity plans. And accurate posture recognition and action analysis are crucial for understanding customers' interest points, behavior patterns, and shopping habits. Without this data, it is impossible to effectively adjust marketing strategies according to the needs of different customer groups, resulting in poor marketing activity effects and unable to maximize the improvement of customer experience and sales performance.

[0004] Therefore, it is necessary to provide a method and device for generating a heat map of a hall store scene with multi-angle fusion to solve the above problems. Summary of the Invention

[0005] The present invention provides a method and device for generating a heat map of a hall store scene with multi-angle fusion to solve the technical problems in the prior art that the heat map generated by a single camera usually cannot conduct in-depth analysis of human postures and actions, cannot understand customers' behavior and demand information, cannot achieve accurate posture recognition and action analysis, cannot quickly identify customers' interest points, and thus cannot implement personalized marketing activity plans for users and has a poor user experience. The technical problems to be solved by the present invention are achieved through the following technical solutions.

[0006] A method for generating a heat map of a hall store scene with multi - angle fusion is proposed in the first aspect of the present invention. The method includes: collecting multi - angle video data through multi - angle shooting to extract an image time series at the same moment; using the Yolov8 algorithm to process the extracted image time series to identify human form information, where the human form information includes human position information and human body key point information; reconstructing a scene three - dimensional model, specifically including the following steps: using multi - angle template images to recalibrate the camera, and mapping the scene two - dimensional images collected by the calibrated camera into three - dimensional space to reconstruct the three - dimensional world of the hall store scene; using the SIFT algorithm to extract image features and perform feature matching on the image time series to perform object surface reconstruction and texture mapping to obtain the reconstructed three - dimensional model of the hall store scene; using the Yolov8 algorithm to identify the human form information in a specified historical time period, forming human form data using the reconstructed three - dimensional model of the hall store scene to generate a heat map of the hall store scene, and representing the activity density of people in colors for analyzing the passenger flow density area; analyzing the data to be analyzed according to the generated heat map of the hall store scene to obtain the passenger flow density area corresponding to the data to be analyzed.

[0007] A hall store scene heat map generation device is proposed in the second aspect of the present invention. It executes the hall store scene heat map generation method described in the first aspect of the present invention. The hall store scene heat map generation device includes: a collection and processing module that collects multi - angle video data through multi - angle shooting to extract an image time series at the same moment; an identification and processing module that uses the Yolov8 algorithm to process the extracted image time series to identify human form information, where the human form information includes human position information and human body key point information; a reconstruction module for reconstructing a scene three - dimensional model, specifically including the following steps: using multi - angle template images to recalibrate the camera, and mapping the scene two - dimensional images collected by the calibrated camera into three - dimensional space to reconstruct the three - dimensional world of the hall store scene; using the SIFT algorithm to extract image features and perform feature matching on the image time series to perform object surface reconstruction and texture mapping to obtain the reconstructed three - dimensional model of the hall store scene; a generation module that uses the Yolov8 algorithm to identify the human form information in a specified historical time period, forms human form data using the reconstructed three - dimensional model of the hall store scene to generate a heat map of the hall store scene, and represents the activity density of people in colors for analyzing the passenger flow density area; an analysis module that analyzes the data to be analyzed according to the generated heat map of the hall store scene to obtain the passenger flow density area corresponding to the data to be analyzed.

[0008] A third aspect of the present invention provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method for generating a heat map of the hall store scene according to the first aspect of the present invention.

[0009] A fourth aspect of the present invention provides a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, the method for generating a heat map of the hall store scene according to the first aspect of the present invention is implemented.

[0010] The embodiments of the present invention include the following advantages:

[0011] Compared with the prior art, the present invention captures real-time images of the hall store based on multiple cameras, uses the Yolov8 algorithm to identify human form information, and simultaneously detects human key points. The method reconstructs a three-dimensional model of the scene using the images captured by multiple cameras, maps the human form information to the three-dimensional model of the scene, generates a heat map to determine the passenger flow density area in the hall store, and analyzes the obtained human body postures based on the human form data formed by mapping the human form information to the three-dimensional model of the scene, so as to accurately determine the shopping preferences, preferred areas, etc. of customers, and further accurately determine the placement area and placement height of commodities, which can effectively improve the operation efficiency, customer experience and customer satisfaction of the hall store. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a flowchart of the steps of the method for generating a heat map of the hall store scene with multi-angle fusion according to the present invention;

[0013] Figure 2 is a flowchart of the steps from another angle of the method for generating a heat map of the hall store scene with multi-angle fusion according to the present invention;

[0014] Figure 3 is a schematic diagram of an example of reconstructing a three-dimensional model of the scene in the method for generating a heat map of the hall store scene with multi-angle fusion according to the present invention;

[0015] Figure 4 is a block diagram of the structure of the system for generating a heat map of the hall store scene with multi-angle fusion according to the present invention;

[0016] Figure 5 is a schematic diagram of the structure of an embodiment of an electronic device according to the present invention;

[0017] Figure 6 is a schematic diagram of the structure of an embodiment of a computer-readable medium according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0019] In view of the above problems, the present invention proposes a method for generating a heat map of a hall and store scene with multi-angle fusion. This method is based on multiple cameras to capture real-time images of the hall and store, uses the Yolov8 algorithm to identify human figures, and at the same time detects human key points. It reconstructs a three-dimensional model of the scene using the images captured by multiple cameras, maps the human figure information to the three-dimensional model of the scene, generates a heat map to determine the passenger flow density area in the hall and store, and can accurately determine the shopping preferences, preferred areas, etc. of customers based on the human pose obtained by analyzing the human figure data formed by mapping the human figure information to the three-dimensional model of the scene, effectively improving the operation efficiency, customer experience and customer satisfaction of the hall and store.

[0020] Embodiment 1

[0021] The following will refer to Figure 1 、 Figure 2 、 Figure 3 to describe the content of the present invention in detail.

[0022] Figure 1 is the step flow chart of the method for generating a heat map of a hall and store scene with multi-angle fusion of the present invention. Figure 2 is the step flow chart from another angle of the method for generating a heat map of a hall and store scene with multi-angle fusion of the present invention.

[0023] See Figure 1 and Figure 2 , in step S101, multi-angle video data is collected by multi-angle shooting to extract an image time series at the same moment.

[0024] In a specific implementation, multiple identical high-definition cameras are installed at different angles in store A to form a comprehensive monitoring network. For details, see Figure 3 . Through multi-angle shooting, a set of image sequences at the same moment is extracted from the captured video, denoted as , where is the number of cameras, .

[0025] It should be noted that the above is only described as an optional example and should not be construed as a limitation to the present invention.

[0026] Next, in step S102, the Yolov8 algorithm is used to process the extracted image time series to identify human figure information, and the human figure information includes the position information of a person and the information of human key points.

[0027] Specifically, the Yolov8 algorithm is used to process the image time series at the same moment to identify people, the number of people, human key points, and the position information of each person within the detection frame. The human key point information includes the head, arms, and legs. In a specific embodiment, a framework is selected to build a human form recognition model. Specifically, the Yolov8 framework is used to build the human form recognition model for object detection and pose estimation.

[0028] Specifically, image data containing humans is collected, and image bounding box labels (such as x and y representing width and height respectively) and pose key points (i.e., human key points) and the coordinate labels of each pose key point are annotated to establish a training dataset. The pose key points are, for example, shoulders, elbows, knees, head, arms, legs, etc.

[0029] Preferably, the image data of the human body is an image sequence at the same moment t.

[0030] Pre - processing such as normalization and resizing is performed on the above - mentioned image data.

[0031] For the training process, it also includes setting hyperparameters, specifically setting hyperparameters such as learning rate, batch size, and number of training epochs.

[0032] Furthermore, the object detection loss and pose estimation loss are defined, and cross - entropy loss and mean squared error are adopted.

[0033] Multiple rounds of iterative training are carried out. By continuously updating the network weights to reduce the object detection loss and pose estimation loss, for example, when the two loss values are within the range of 0.001 - 0.006, the model training is stopped, that is, the model training is completed, and a trained human form recognition model is obtained.

[0034] Specifically, the image sequence to be processed at the same moment

[0035] is pre - processed, and the pre - processed image sequence to be processed is input into the human form recognition model, and the coordinate position of the image detection frame, such as the center position is output. is the total number of people recognized, i represents the i - th camera, i is a positive integer, specifically i = 1, 2, 3,..., N, , is the number of cameras, At the same time, for each recognized person, the human key points are accurately located, denoted as . Specifically, it represents the -th person in the image collected by the i - th camera, and the -th The coordinate information of key points. The key point coordinates of the human body are usually returned in the form of an array. The detection of these key points can help understand the customer's posture and movement, providing basic data for subsequent analysis.

[0036] It should be noted that the above is only for illustrative purposes as an optional example and should not be construed as a limitation to the present invention.

[0037] Next, in step S103, reconstruct the three-dimensional model of the scene, which specifically includes the following steps: Use template images from multiple angles to recalibrate the camera, and map the two-dimensional scene images captured by the calibrated camera into the three-dimensional space to reconstruct the three-dimensional world of the hall store scene; Use the SIFT algorithm to extract image features and perform feature matching on the image time series to perform object surface reconstruction and texture mapping to obtain the reconstructed three-dimensional model of the hall store scene.

[0038] Specifically, reconstructing the three-dimensional model of the scene specifically includes the following steps:

[0039] Step S201, use template images from multiple angles to recalibrate the camera, and map the two-dimensional scene images captured by the calibrated camera into the three-dimensional space to reconstruct the three-dimensional world of the hall store scene;

[0040] For the calibration of each camera, fix a black and white grid calibration board on a plane, that is, the template plane, and take several template images from different angles. Detect the feature points in the template images. Assume that the template plane is on the plane of the world coordinate system The relationship between the three-dimensional points in space projected onto the template image is:

[0041] ;

[0042] Among them, represents projecting the corresponding point on the template plane image into the three-dimensional point in space; (u, v) is the point on the template plane, is the internal parameter matrix of the camera, and are the rotation matrix and translation variable of the camera coordinate system relative to the world coordinate system respectively, is the homogeneous coordinate of the point on the template plane projected onto the corresponding point on the image plane. u represents the abscissa of the point on the template plane, v represents the ordinate of the point on the template plane, and l represents the weight of the homogeneous coordinate, specifically the number 1; is the homogeneous coordinate of the point on the template plane to solve for , and .

[0043] It should be noted that in camera calibration, the 1 in the homogeneous coordinate (u, v, 1) represents the weight of the homogeneous coordinate, which is used to convert the planar point (u, v) into the homogeneous coordinate form (u, v, 1) for various geometric transformations and projection operations.

[0044] For the camera calibration process, calibration board: Select an object with known geometric characteristics as the calibration board (i.e., the template plane), such as a checkerboard. The checkerboard is composed of black and white squares, and the actual size of each square is known, for example, 10mm x 10mm. The camera is used to capture images of the calibration board.

[0045] For example, use the OpenCV library for image processing and calibration calculations.

[0046] Specifically, collect calibration data by taking multiple pictures of the checkerboard from different angles and distances. Ensure that all parts of the checkerboard appear in different images, which can provide sufficient data to accurately estimate the camera parameters. Then, detect corner points using the findChessboardCorners function in OpenCV to automatically detect the corner points of the checkerboard in each image. The positions of these corner points in the image will be used to calculate the internal and external parameters of the camera. Further, calculate the camera parameters using the collected object points and image points, and call the cv2.calibrateCamera function to calculate the internal parameters of the camera (focal length, principal point position, distortion coefficients) and external parameters (rotation matrix and translation vector).

[0047] Use the calculated internal parameters (focal length, principal point position, distortion coefficients) and external parameters (rotation matrix and translation vector) of the camera to map the points in the two-dimensional image back to the three-dimensional world.

[0048] Step S202: Use the SIFT algorithm to perform image feature extraction and feature matching on the image time series for object surface reconstruction and texture mapping to obtain a reconstructed three-dimensional model of the hall and store scene.

[0049] Use the SIFT (Scale-Invariant Feature Transform) algorithm to extract sparse feature points of display items such as shelves and cabinets in the hall from the image time series to obtain a set of feature points.

[0050] For example, by finding key point pairs with similar feature point descriptors in two captured hall images, the corresponding feature points in the two images can be matched for three-dimensional reconstruction.

[0051] Calculate the descriptors of each feature point, and use nearest neighbor matching to obtain the point cloud corresponding to the display items such as shelves and cabinets in the hall to reconstruct the object surface of the key items.

[0052] Specifically, calculate the descriptor of each feature point , and use nearest neighbor matching to obtain the matching pairs . Among them, the descriptor is a vector representing the local area around the feature point.

[0053] For each feature point in Image A , calculate its descriptor and the Euclidean distance between the descriptor

[0054] of the feature point in Image B to find the feature point with the minimum distance as the nearest neighbor:

[0055] ;

[0056] Among them, represents the Euclidean distance between the descriptor of each feature point in Image A and the descriptor

[0057] of the feature point in Image B ; k is a positive integer, specifically 1, 2,..., 128.

[0058] For each feature point , find the nearest neighbor and the second nearest neighbor .

[0059] Calculate the distance ratio r:

[0060] ;

[0061] If <a predetermined value (e.g., represented by threshold), then and are considered matching points.

[0062] After completing the feature point matching, use the PMVS (Patch-Match Stereo) algorithm to obtain a denser point cloud corresponding to the display items such as shelves and cabinets in the hall, reconstruct the object surface based on these point clouds, and perform texture mapping, that is, restore the three-dimensional scene and objects.

[0063] Specifically, the PMVS algorithm is a multi-view graph matching algorithm.

[0064] Taking a small supermarket as an example, the process of reconstructing the surface of an object from point clouds will be described.

[0065] Specifically, obtain the camera parameters inside the small supermarket. Set up a calibration board in the small supermarket and take multiple photos at different angles to obtain the internal parameters and distortion parameters of the camera. By identifying and matching the feature points of key objects (or target objects) such as shelves, freezers, and cash desks in the supermarket, feature point matching is performed.

[0066] Next, generate sparse three-dimensional point clouds from the extracted feature points. By identifying the layout feature points of the shelves, such as the arrangement, spacing, and height of the shelves, it is used to reconstruct the entrance and exit positions of the supermarket, as well as possible customer passageways. The shelf layout feature points specifically include the promotional signs at the top.

[0067] Use algorithms such as PMVS to generate more detailed dense point clouds. Specifically, generate dense point clouds of various commodities on the shelves, including dense point clouds of details of commodities such as bottled beverages, canned foods, and daily necessities. In addition, it also includes refining the shape and size of the refrigerator and the arrangement of the commodities inside. In addition, it also includes capturing the shape of the cash desk and the equipment placed on it (such as cash registers, shopping baskets, etc.). Convert the generated dense point clouds into a three-dimensional mesh model for object surface reconstruction to generate a mesh model of the shelf, refine the models of various commodities, and ensure that their shapes and sizes are truly reflected.

[0068] For object surface reconstruction, it specifically includes reconstructing the material of the supermarket floor (such as tiles or carpets), as well as the color and decoration of the walls. It also includes promotional posters and indication signs on the walls, adding to the brand atmosphere of the supermarket.

[0069] Preferably, add realistic textures to the three-dimensional model. For example, apply the extracted textures of the walls and floors to the corresponding three-dimensional models to create a realistic shopping environment and make it more realistic.

[0070] Specifically, apply the package and label images of the extracted commodities to the corresponding models on the shelves to ensure the accuracy of the information. Finally, use a rendering engine to generate the final effect, showing the lighting effects of the supermarket, the colors of the commodities, and the overall shopping atmosphere.

[0071] Through the above process, perform texture mapping on the point clouds of display items such as shelves and cabinets in the store to restore the hall store scene and various objects, so as to obtain a three-dimensional model of the reconstructed hall store scene.

[0072] Use the relationship between the three-dimensional points in the space projected onto the template image to map the two-dimensional images captured by the camera into the three-dimensional space, that is, the three-dimensional model of the hall store scene.

[0073] It should be noted that the above is only for illustrative purposes as an optional example and should not be construed as a limitation to the present invention.

[0074] Next, in step S104, the Yolov8 algorithm is used to identify humanoid information in a specified historical time period, and the reconstructed three-dimensional model of the hall and store scene is utilized to form humanoid data, so as to generate a heat map of the hall and store scene, and the activity density of people is represented by colors for analyzing the passenger flow density area.

[0075] Specifically, according to the internal parameter matrix of the calibrated camera, the points representing humanoid information in the camera coordinate system are converted into corresponding mapped points in the world coordinate system, so as to convert the humanoid information into humanoid data in the three-dimensional model of the hall and store scene. The humanoid data includes the bounding box coordinates of people and human body pose data.

[0076] Internal parameter matrix of the camera , where 、 are the coordinates of the focal length, are the coordinates of the principal point. Use to convert the coordinates of the j-th customer in the image captured by the i-th camera into the normalized coordinates in the camera coordinate system: . Calculate the depth , is the actual height of the customer, i.e., a person, and the projected height in the camera coordinate system is , z c is the calculated depth of the person in the three-dimensional world, and f y represents the focal length of the camera on the y-axis.

[0077] Combine the external parameter matrix , and convert the points in the camera coordinate system in a certain image to be processed into corresponding points in the world coordinate system , where respectively represent the coordinate information of the abscissa X, ordinate Y, and third coordinate Z of the corresponding points converted into the world coordinate system; respectively represent the coordinate information of the points in the camera in the abscissa X, ordinate Y, and third coordinate Z, represents the result obtained after the points in the camera are rotated, and R(Xc, Yc, Zc)+T represents the result obtained after the points in the camera are rotated and translated.

[0078] Use the reconstructed 3D model of the hall and store scenario to form humanoid data. Specifically, map all humanoid information into the 3D space through the above method to obtain humanoid data, and generate a heat map of the hall and store scenario. Specifically, visualize the humanoid data within the specified historical time period to generate a heat map of the hall and store scenario. The specified historical time period includes the time of the generated store heat map, specifically from the start time to the end time.

[0079] For example, from October 1, 2023 to October 2, 2023.

[0080] In a specific embodiment, perform visualization processing on the humanoid data obtained within the specified historical time period from the start time to the cut-off time to generate a heat map of the hall and store scenario.

[0081] Furthermore, represent the activity density of people in colors for analyzing the passenger flow density area. Represent the activity density of customers in the hall and store scenario through gradient colors and divide them into different density areas. Specifically, adopt color mapping technology to represent the activity density of customers through gradient colors (such as from yellow to red). Yellow represents the low-density area, while red represents the high-density area. This visualization method enables the hall and store managers to clearly identify the areas where customers concentrate their activities. Specifically, refer to Figure 3 the different passenger flow density areas represented by circles a, b, c, d, and e in it. Among them, circle a uses red and circle e uses yellow.

[0082] In addition, use the reconstructed 3D model of the hall and store scenario to form humanoid data. Specifically, map all humanoid information into the 3D space through the above method to obtain humanoid data, and analyze and determine the posture changes of people based on the heat map of the hall and store scenario, which can further understand the behavior patterns of customers, identify popular products or preferred areas, and provide accurate data basis for optimizing the operation and sales layout.

[0083] Preferably, perform comparative analysis on the human body postures according to the image time series and the identified humanoid information to determine the behavior habits and consumption preferences of each user.

[0084] It should be noted that the above is only for illustrative purposes as an optional example and should not be construed as a limitation to the present invention.

[0085] Next, in step S105, analyze the data to be analyzed according to the generated heat map of the hall and store scenario to obtain the passenger flow density area corresponding to the data to be analyzed.

[0086] According to the generated heat map of the hall and store scenarios, corresponding passenger flow density regions are determined, and circular regions a, b, c, d, and e are used to represent different passenger flow density regions respectively. Specifically, circular regions a to e represent passenger flow density regions with gradually decreasing passenger flow density. For details, see Figure 3 .

[0087] Specifically, the data to be analyzed are all human form data obtained from the start time to the end time period, specifically including the position of the coordinate frame of the person and the data of the key points of the person's posture.

[0088] First, based on the human form data in the image sequence, specifically the human form data within an accumulated time period from the start time to the end time, the overlapping effect of the human form data can be obtained. The places with more people have darker colors, which are represented by red, that is, a red area is formed. The overlapping effect of the human form data (i.e., the red area) is defined as the high passenger flow density area, and the places with fewer people have lighter colors, which are represented by yellow, that is, a yellow area is formed. The yellow area is defined as the low passenger flow density area. Thus, the hall and store merchants can obtain the areas that customers like and can accurately determine the areas for placing goods.

[0089] Secondly, based on the human body posture data in the image sequence, through the accumulation of time, it is possible to obtain which goods the customers are interested in. Therefore, the hall and store merchants can accurately determine the placement height of the goods.

[0090] In an alternative embodiment, according to the recognized people, the number of people, the key points of the human body, and the position information of each person within the detection frame, customer behavior data are determined, specifically including the staying time of the customer in front of a certain shelf, the action of leaning forward to approach the goods, the action of stretching out both hands, or other actions indicating interest in the goods.

[0091] Further, according to the determined customer behavior data, the display position of the goods is adjusted or the inventory of the goods is increased.

[0092] Specifically, a depth network algorithm is used to construct an item prediction model for predicting the display position of the item, the display height, and the quantity of the item (i.e., the quantity of each commodity).

[0093] Based on the above determined customer behavior data, specifically including the staying time of the customer in front of a certain shelf, the action of leaning forward to approach the goods, the action of stretching out both hands, or other actions indicating interest in the goods, the customer behavior data marked with the display position of the item, the display height, and the quantity of the item are used to establish a training data set to train the item prediction model. When the accuracy of the model reaches more than 90%, the model training is stopped, and a trained item prediction model is obtained.

[0094] Further, input the customer behavior data of the customer to be predicted (the customer behavior data determined based on step S105) into the item prediction model to obtain the product display position corresponding to the customer to be predicted or increase the inventory of products.

[0095] In addition, for the managers of the hall store, first, the layout can be optimized according to the analysis results, and the product display, shelf placement, rest area setting, etc. in the hall store can be optimized and adjusted. Secondly, according to the division of the passenger flow density area, the number of salesclerks and working hours can be reasonably allocated, and customized marketing plans can be made according to the consumption habits of customers in different areas to attract customers to consume.

[0096] It should be noted that the above is only described as an optional example and should not be construed as a limitation to the present invention.

[0097] Compared with the prior art, the present invention captures real-time images of the hall store based on multiple cameras, uses the Yolov8 algorithm to identify human form information, and simultaneously detects human body key points. The three-dimensional model of the scene is reconstructed using the images captured by multiple cameras, and the human form information is mapped onto the three-dimensional model of the scene to generate a heat map to determine the passenger flow density area in the hall store. According to the human form data formed by mapping the human form information onto the three-dimensional model of the scene and the analyzed human body postures, the shopping preferences, preferred areas, etc. of customers can be accurately determined, and thus the placement area and placement height of products can be accurately determined, which can effectively improve the operation efficiency, customer experience, and customer satisfaction of the hall store.

[0098] Embodiment 2

[0099] The following is an embodiment of the system of the present invention, which can be used to execute the method embodiment of the present invention. For the details not disclosed in the device embodiment of the present invention, please refer to the method embodiment of the present invention.

[0100] Figure 4 It is a schematic structural diagram of an example of the hall store scene heat map generation device according to the present invention. The hall store scene heat map generation device will be described below with reference to Figure 4 , and the hall store scene heat map generation device will be described. The hall store scene heat map generation device is used to execute the hall store scene heat map generation method described in the first aspect of the present invention.

[0101] As Figure 4 shown, the hall store scene heat map generation device 300 includes a collection and processing module 310, an identification and processing module 320, a reconstruction module 330, a generation module 340, and an analysis module 350.

[0102] In a specific embodiment, the acquisition processing module 310 acquires multi-angle video data by shooting from multiple angles to extract the image time series at the same time. The recognition processing module 320 uses the Yolov8 algorithm to process the extracted image time series and identify the human figure information, which includes the position information of the person and the key point information of the human body. The reconstruction module 330 is used to reconstruct the scene three-dimensional model, which specifically includes the following steps: using the template image of multiple angles to recalibrate the camera, and use the scene two-dimensional image acquired by the calibrated camera to map to the three-dimensional space to reconstruct the three-dimensional world of the hall and store scene; using the SIFT algorithm to extract image features and match features of the image time series to reconstruct the object surface and texture mapping to obtain the reconstructed hall and store scene three-dimensional model. The generation module 340 uses the Yolov8 algorithm to identify the human figure information of the specified historical time period, and uses the reconstructed hall and store scene three-dimensional model to form human figure data to generate a hall and store scene heat map, and the activity density of the characters is represented by color for analyzing the passenger flow density area. The analysis module 350 analyzes the data to be analyzed based on the generated store scene heat map to obtain the passenger flow density area corresponding to the data to be analyzed.

[0103] According to an optional implementation, the SIFT algorithm is used to extract sparse feature points corresponding to the displayed items in the store from the image time series to obtain a feature point set. The descriptor of each feature point is calculated, and the nearest neighbor matching is used to obtain the point cloud corresponding to the displayed items to reconstruct the object surface of the key items; after completing the feature point matching, the PMVS algorithm is used to obtain a denser point cloud corresponding to the displayed items, and the object surface is reconstructed based on these denser point clouds, and texture mapping is performed; the displayed items include shelves, refrigerators, and cash registers.

[0104] According to an optional implementation, texture mapping is performed on the point clouds of the following key objects to restore the store scene and each object to obtain a reconstructed three-dimensional model of the store scene: shelves, cabinets, and cash registers in the store;

[0105] The surface reconstruction of objects includes the reconstruction of the material of the store floor, the color and decoration of the walls, the promotional posters and signs on the walls.

[0106] According to an optional implementation manner, the reconstructed three-dimensional model of the hall and store scene is used to form human shape data to generate a heat map of the hall and store scene.

[0107] According to the calibrated camera's internal parameter matrix, the points representing the human figure information in the camera coordinate system are converted into corresponding mapping points in the world coordinate system, so as to convert the human figure information into human figure data in the three-dimensional model of the hall and store scene. The human figure data in the specified historical time period is visualized to generate a hall and store scene heat map.

[0108] According to an alternative embodiment, the activity density of customers in the store scenario is represented by a gradient color and divided into different passenger flow density regions, and the number of store clerks and working hours are allocated according to different passenger flow density regions.

[0109] According to an alternative embodiment, based on the image time series and the recognized human form information, a comparative analysis of human postures is performed to determine the behavior habits and consumption preferences of each user.

[0110] Specifically, the Yolov8 algorithm is used to process the image time series at the same moment to identify the number of people, the key points of the human body, and the position information of each person within the detection frame. The key point information of the human body includes the head, arms, and legs.

[0111] Based on the recognized number of people, the key points of the human body, and the position information of each person within the detection frame, customer behavior data is determined, specifically including the stay time of customers in front of a certain shelf, the action of leaning forward to approach the product, the action of stretching out both hands, or other actions indicating interest in the product.

[0112] Furthermore, based on the determined customer behavior data, the product display position is adjusted or the product inventory is increased.

[0113] According to an alternative embodiment, a black and white checkerboard calibration plate is fixed on a plane, i.e., the template plane, and several template images are taken from different angles. The feature points in the template images are detected. Assuming that the template plane is on the plane of the world coordinate system the relationship between the three-dimensional points in space projected onto the template image is:

[0114] ;

[0115] where represents projecting the corresponding point on the template plane image onto the three-dimensional point in space; (u, v) is the point on the template plane; is the internal parameter matrix of the camera, and are the rotation matrix and translation variable of the camera coordinate system relative to the world coordinate system respectively, is the homogeneous coordinate of the point on the template plane projected onto the corresponding point on the image plane. u represents the abscissa of the point on the template plane, v represents the ordinate of the point on the template plane, and l represents the weight of the homogeneous coordinate, specifically the number 1; is the homogeneous coordinate of the point on the template plane to solve for , and ;

[0116] The two-dimensional image captured by the camera is mapped into a three-dimensional space, that is, the three-dimensional model of the hall store scene, by using the relationship between the three-dimensional points in the space projected onto the template image.

[0117] It should be noted that since Figure 4 the method for generating the heat map of the hall store scene executed by the heat map generation device of the hall store scene is substantially the same as Figure 1 the method for generating the heat map of the hall store scene in the example of

[0118] Compared with the prior art, the present invention captures real-time images of the hall store based on multiple cameras, uses the Yolov8 algorithm to identify human figures, and simultaneously detects human key points. The method reconstructs the three-dimensional model of the scene by using the images captured by multiple cameras, maps the human figure information to the three-dimensional model of the scene, generates a heat map to determine the passenger flow density area in the hall store, and analyzes the human body postures obtained from the human figure data formed by mapping the human figure information to the three-dimensional model of the scene, so as to accurately determine the shopping preferences, preferred areas, etc. of customers, effectively improving the operation efficiency, customer experience and customer satisfaction of the hall store.

[0119] Figure 5 FIG. is a schematic structural diagram of an embodiment of an electronic device according to the present invention.

[0120] As Figure 5 shown, the electronic device is presented in the form of a general-purpose computing device. The processor can be one or multiple and work cooperatively. The present invention does not exclude distributed processing, that is, the processors can be dispersed in different physical devices. The electronic device of the present invention is not limited to a single entity, but can also be the sum of multiple physical devices.

[0121] The memory stores computer-executable programs, usually machine-readable codes. The computer-readable programs can be executed by the processor so that the electronic device can execute the method of the present invention or at least some of the steps in the method.

[0122] The memory includes volatile memory, such as a random access storage unit (RAM) and / or a cache storage unit, and can also be non-volatile memory, such as a read-only storage unit (ROM).

[0123] Optionally, in this embodiment, the electronic device further includes an I / O interface for data exchange between the electronic device and external devices. The I / O interface can represent one or more of several bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in multiple bus structures.

[0124] It should be understood that Figure 5The electronic device shown is merely an example of the present invention. The electronic device of the present invention may also include elements or components not shown in the above examples. For example, some electronic devices also include a display unit such as a display screen, and some electronic devices also include human-computer interaction elements, such as buttons, keyboards, etc. As long as the electronic device can execute the computer-readable program in the memory to implement at least part of the steps of the method of the present invention, it can be considered as the electronic device covered by the present invention.

[0125] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, as Figure 6 shown, the technical solution according to the embodiment of the present invention can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several commands to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the above method according to the embodiment of the present invention.

[0126] The software product may adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0127] The computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program used by or in conjunction with a command execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted by any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.

[0128] The program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0129] The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by a device, the computer-readable medium implements the data interaction method of the present disclosure.

[0130] Those skilled in the art can understand that the above-mentioned modules can be distributed in the device according to the description of the embodiments, or can be correspondingly changed and distributed in one or more devices that are uniquely different from this embodiment. The modules of the above embodiments can be combined into one module, or can be further split into multiple sub-modules.

[0131] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solution according to the embodiments of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several commands to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present invention.

[0132] It should be noted that the above detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0133] In the foregoing detailed description, reference has been made to the accompanying drawings, which form a part hereof. In the drawings, like numerals typically identify like components, unless the context indicates otherwise. The illustrated embodiments described in the detailed description, the drawings, and the claims are not meant to be limiting. Other embodiments may be used and other changes may be made without departing from the spirit or scope of the subject matter presented herein.

[0134] The foregoing are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for generating a heat map of a store scene with multi-angle fusion, characterized in that: The method comprises: Collect multi-angle video data through multi-angle shooting to extract the image time series at the same moment; The Yolov8 algorithm is used to establish a human figure recognition model, and the extracted image time series is processed to identify human figure information, which includes the position information of the person and the key point information of the human body; the Yolov8 algorithm is used to process the image time series at the same time t Processing is performed to identify people, the number of people, key points of the human body, and the position information of each person in the detection frame, wherein the key point information of the human body includes the head, arms, and legs; target detection loss and posture estimation loss are defined, and cross entropy loss and mean square error are used to perform multiple rounds of iterative training, and the target detection loss and posture estimation loss are reduced by continuously updating the network weights. When the two loss values ​​are within the range of 0.001 to 0.006, the model training is stopped, i.e., the model training is completed, and a trained human figure recognition model is obtained; customer behavior data is determined based on the identified people, the number of people, key points of the human body, and the position information of each person in the detection frame, and the specific Including the time a customer stays in front of a shelf, the action of leaning towards the product, the action of stretching out both hands, or other actions that indicate interest in the product; further adjusting the product display position or increasing the product inventory based on the determined customer behavior data; using a deep network algorithm to build an item prediction model for predicting the item display position, display height, and number of items; using customer behavior data labeled with the item display position, display height, and number of items to establish a training data set, train the item prediction model, input the customer behavior data of the customer to be predicted into the item prediction model, and obtain the product display position or increase the product inventory corresponding to the customer to be predicted; Reconstructing a three-dimensional model of a scene specifically includes the following steps: using a template image from multiple angles to recalibrate the camera, and using the two-dimensional image of the scene collected by the calibrated camera to map it into a three-dimensional space to reconstruct the three-dimensional world of the hall and store scene; using the SIFT algorithm to perform image feature extraction and feature matching on the image time series to perform object surface reconstruction and texture mapping to obtain a reconstructed three-dimensional model of the hall and store scene; wherein, using the SIFT algorithm to extract sparse feature points corresponding to the displayed items in the hall and store from the image time series to obtain a feature point set; calculating the descriptor of each feature point, using the nearest neighbor matching to obtain the point cloud corresponding to the displayed item to reconstruct the object surface of the key item; after completing the feature point matching, using the PMVS algorithm to obtain a denser point cloud corresponding to the displayed item, reconstructing the object surface based on these denser point clouds, and performing texture mapping; the displayed items include shelves, refrigerators, and cash registers; Use the Yolov8 algorithm to identify human figure information in a specified historical time period, and use the reconstructed three-dimensional model of the hall and store scene to form human figure data to generate a hall and store scene heat map, and represent the activity density of the characters with colors to analyze the passenger flow density area; According to the generated heat map of the hall and store scene, the data to be analyzed is analyzed to obtain the passenger flow density area corresponding to the data to be analyzed.

2. The method for generating a heat map of a store scene according to claim 1 is characterized in that: Further including: Texture mapping is performed on the point clouds of the following key objects to restore the store scene and each object to obtain a reconstructed 3D model of the store scene: shelves, cabinets, and cash registers in the store; The surface reconstruction of objects includes the reconstruction of the material of the store floor, the color and decoration of the walls, the promotional posters and signs on the walls.

3. The method for generating a heat map of a store scene according to claim 1, characterized in that: The method of using the reconstructed three-dimensional model of the hall and store scene to form human shape data to generate a hall and store scene heat map includes: According to the calibrated internal parameter matrix of the camera, the points representing the human figure information in the camera coordinate system are converted into corresponding mapping points in the world coordinate system, so as to convert the human figure information into human figure data in the three-dimensional model of the store scene; The humanoid data within a specified historical time period is visualized to generate a heat map of the hall and store scene.

4. The method for generating a heat map of a store scene according to claim 1, characterized in that: Further including: The activity density of customers in the hall store scene is expressed by gradient colors and divided into different customer flow density areas. The number of store clerks and working hours are configured according to different customer flow density areas.

5. The method for generating a heat map of a store scene according to claim 1, characterized in that: Further including: Based on the image time series and the recognized human figure information, a comparative analysis of human body postures is performed to determine the behavioral habits and consumption preferences of each user.

6. The method for generating a heat map of a store scene according to claim 1, characterized in that: Further including: Fix a black and white grid calibration plate on a plane, i.e., the template plane, take several template images from different angles, and detect the feature points in the template image. Assuming that the template plane is on the plane of the world coordinate system Z=0, the relationship between the three-dimensional point in space projected onto the template image is: s[u,v,1] T =K[r1,r2,r3,t] [X,Y,0,1] T =K[r1,r2,t] [X,Y,1] T ; Among them, s[u,v,1] T Indicates the projection of the corresponding point on the template plane image to a three-dimensional point in space; (u, v) is a point on the template plane; K[r1, r2, r3, t] is the intrinsic parameter matrix of the camera, [r1, r2, r3] and t are the rotation matrix and translation variable of the camera coordinate system relative to the world coordinate system, [u, v, 1] T [X, Y, 1] is the homogeneous coordinate of the corresponding point on the template plane projected onto the image plane. u represents the horizontal coordinate of the point on the template plane, v represents the vertical coordinate of the point on the template plane, and l represents the weight of the homogeneous coordinate, which is specifically 1; T is the homogeneous coordinates of the point on the template plane to solve for K, [r1, r2, r3] and t; The relationship between the three-dimensional points in the space and the template image is used to map the two-dimensional image captured by the camera into the three-dimensional space, that is, the three-dimensional model of the store scene.

7. A device for generating a heat map of a store scene, characterized in that: The method for generating a heat map of a hall and store scene according to any one of claims 1 to 6 is executed, and the heat map generating device for the hall and store scene comprises: The acquisition and processing module acquires video data from multiple angles by shooting from multiple angles to extract the image time series at the same moment; The recognition processing module uses the Yolov8 algorithm to process the extracted image time series and recognize the human figure information, which includes the position information of the person and the key point information of the human body; The reconstruction module is used to reconstruct the three-dimensional model of the scene, which specifically includes the following steps: using the template images from multiple angles to recalibrate the camera, and using the two-dimensional images of the scene collected by the calibrated camera to map into the three-dimensional space to reconstruct the three-dimensional world of the hall and store scene; using the SIFT algorithm to extract image features and match features on the image time series to perform object surface reconstruction and texture mapping to obtain the reconstructed three-dimensional model of the hall and store scene; The generation module uses the Yolov8 algorithm to identify human figure information in a specified historical time period, and uses the reconstructed three-dimensional model of the hall and store scene to form human figure data to generate a hall and store scene heat map, and represents the activity density of the characters by color, which is used to analyze the passenger flow density area; The analysis module analyzes the data to be analyzed based on the generated store scene heat map to obtain the customer flow density area corresponding to the data to be analyzed.

8. The device for generating a heat map of a store scene according to claim 7, characterized in that: include: Using SIFT algorithm to extract sparse feature points related to the displayed items in the store from the image time series to obtain a feature point set; Calculate the descriptor of each feature point, use nearest neighbor matching to obtain the point cloud corresponding to the displayed items to reconstruct the object surface of the key items; After completing the feature point matching, the PMVS algorithm is used to obtain a denser point cloud corresponding to the displayed items, and the object surface is reconstructed based on these denser point clouds, and texture mapping is performed; the displayed items include shelves, refrigerators, and cash registers.

Citation Information

Patent Citations

  • Three-dimensional model reconstruction method based on smart phone

    CN110533774A

  • Building space passenger flow simulation system and construction method

    CN117973069A

  • Method for identifying stealing behavior in unmanned convenience store

    CN118865497A