Customer behavior tracking method, customer behavior tracking system and behavior analyzing unit
Patent Information
- Authority / Receiving Office
- TW · TW
- Patent Type
- Patents
- Current Assignee / Owner
- IND TECH RES INST
- Filing Date
- 2024-11-28
- Publication Date
- 2026-08-01
AI Technical Summary
Existing retail monitoring systems fail to accurately identify key customer shopping behaviors, such as missed or accidental item scanning, leading to losses for merchants and lack of understanding of customer preferences and shopping intentions.
A customer behavior tracking system utilizing multi-target tracking elements, AIoT components, and Vision Language Models to capture key video clips of customer actions, identify products, and provide real-time notifications for accurate checkout or purchase behaviors.
Accurately identifies checkout or purchase behaviors, reducing errors and providing timely prompts to enhance customer purchasing motivation.
Smart Images

Figure TWG2TB001903695_001 
Figure TWG2TB001903695_002 
Figure TWG2TB001903695_003
Abstract
Description
Technical Field
[0001] This disclosure relates to a customer behavior tracking method, a customer behavior tracking system, and a behavior analysis unit. Prior Technology
[0002] With the digital transformation of the retail environment and the increasing maturity of mall monitoring technology, many businesses are beginning to rely on cameras and sensors to monitor and analyze customer traffic and behavior. However, most existing systems focus on real-time recording and monitoring, failing to accurately identify key actions in customer shopping behavior. For example, in self-checkout environments, abnormal behaviors such as missed or accidental item scanning are easily overlooked. Furthermore, these systems cannot build an understanding of customer purchasing behavior in physical retail settings, nor can they reveal customer preferences and potential shopping intentions.
[0003] Therefore, how to avoid missing or accidentally scanning products to reduce losses for merchants, and how to proactively and promptly provide relevant prompts to customers to increase their purchasing motivation, is one of the goals of current research and development in the industry. Summary of the Invention
[0004] This disclosure relates to a customer behavior tracking method, a customer behavior tracking system, and a behavior analysis unit, which captures key video clips based on key actions to accurately identify a user's checkout or purchase behavior through these key video clips.
[0005] According to one aspect of this disclosure, a customer behavior tracking method is proposed. The customer behavior tracking method includes the following steps: A key user entering an identification area is identified by a multi-target tracking element connected to several image capturing units. At least one Artificial Intelligence of Things (AIoT) element is used to identify at least one key action of the key user on a single product, including a picking-up action. Based on the key action, a key video clip is captured. The time point of the picking-up action is the starting point of the key video clip. Based on the key video clip, a product item of the single product is identified using a Vision Language Model (VLM). Based on the product item of the single product, a checkout action or a purchase action of the key user on the single product is identified.
[0006] According to another aspect of this disclosure, a customer behavior tracking system is proposed. The customer behavior tracking system includes a behavior recognition module. The behavior recognition module includes several image capture units and a behavior analysis unit. The behavior analysis unit includes a multi-target tracking element, at least one artificial intelligence (AI) IoT element, a video capture element, and a visual language model. The multi-target tracking element is connected to the image capture units. The multi-target tracking element is used to identify a key user entering a recognition area. The AI IoT element is used to identify at least one key action of the key user towards a single product. The key action includes a picking action. The video capture element is used to capture a key video clip based on the key action, with the time point of the picking action being the starting point of the key video clip. The visual language model is used to identify a product item of a single product based on the key video clip. The product item of a single product is used to identify a checkout behavior or a selection behavior of a key user towards a single product.
[0007] According to another aspect of this disclosure, a behavior analysis unit is proposed. The behavior analysis unit includes a multi-target tracking element, at least one artificial intelligence (AI) IoT element, a video capture element, and a visual language model. The multi-target tracking element is connected to a plurality of video capture units and is used to track a key user entering a recognition area. The AI IoT element is used to identify at least one key action of the key user towards a single product. The at least one key action includes a picking action. The video capture element is used to capture a key video clip based on the at least one key action. The time point of the picking action is the starting point of the key video clip. The visual language model is used to identify a product item of the single product based on the key video clip. The product item of the single product is used to identify a checkout behavior of the key user towards the single product.
[0008] To provide a better understanding of the above and other aspects of this disclosure, specific embodiments are described below in conjunction with the accompanying drawings: Simple Explanation of the Diagram
[0009] Figure 1 illustrates a missed brushing behavior according to an embodiment of this disclosure. Figure 2 illustrates a false brushing behavior according to an embodiment of this disclosure. Figure 3 illustrates a block diagram of a customer behavior tracking system according to an embodiment of this disclosure. Figure 4 illustrates a flowchart of a customer behavior tracking method applied to a merchandise checkout area according to an embodiment of this disclosure. Figure 5 illustrates one example of step S110 in Figure 4. Figure 6 illustrates an example of step S120 in Figure 4. Figures 7A-7B illustrate an example of step S130 in Figure 4. Figures 8A-8B illustrate another example of step S130 in Figure 4. Figures 9A-9B illustrate another example of step S130 in Figure 4. Figure 10 illustrates an example of steps S140-S150 in Figure 4. Figure 11 illustrates a detailed flowchart of step S160 according to an embodiment of the present disclosure. Figure 12 illustrates a block diagram of a smart self-checkout system according to an embodiment of this disclosure. Figure 13 illustrates a flowchart of a customer behavior tracking method applied to a merchandise selection area according to an embodiment of this disclosure. Figure 14 illustrates the steps in Figure 13. Implementation
[0010] The technical terms used in this specification are based on common terminology in the field. Where this specification provides explanations or definitions for certain terms, the interpretation of those terms shall be based on the explanations or definitions provided in this specification. Each of the embodiments disclosed herein has one or more technical features. Where feasible, those skilled in the art may selectively implement some or all of the technical features in any embodiment, or selectively combine some or all of the technical features in these embodiments.
[0011] Please refer to Figure 1, which illustrates a missed scan behavior according to an embodiment of this disclosure. In one embodiment, in the identification area RG1 in front of the self-checkout system, a key user US1 is checking out an item PD2. However, the barcode BC of item PD2 may not be correctly mapped to an item scanning module 300, resulting in a missed scan. In this embodiment, the checkout screen 500 will immediately display "Missed Scan" to provide an interactive reminder and request the key user US1 to rescan item PD2.
[0012] Please refer to Figure 2, which illustrates a case of accidental scanning according to one embodiment of this disclosure. In another embodiment, when the key user US1 is checking out item PD3, they may mistakenly scan the barcode BC of item PD4, resulting in accidental scanning. In this embodiment, the checkout screen 500 will immediately display "Accidental Scan" to provide an interactive reminder and request the key user US1 to rescan item PD3.
[0013] In the aforementioned examples of missed or erroneous scans, the product scanning module 300 alone cannot identify them. Instead, the technology disclosed herein must be combined with image recognition technology and a behavioral database for cross-comparison to improve the abnormal behavior recognition rate.
[0014] Please refer to Figure 3, which illustrates a block diagram of a customer behavior tracking system 1000 according to an embodiment of this disclosure. The customer behavior tracking system 1000 includes a behavior recognition module 100 and a behavior processing module 200. The behavior recognition module 100 includes several image capturing units 110 and a behavior analysis unit 120.
[0015] The behavior analysis unit 120 includes a multi-target tracking element 122, at least one Artificial Intelligence of Things (AIoT) element 123, a video capture element 124, a Vision Language Model (VLM) 125, and a behavior database 126. The behavior processing module 200 includes a real-time interaction unit 210, a notification unit 220, and a recording unit 230.
[0016] Multi-target tracking element 122 is used for tracking people. Artificial intelligence (AI) IoT element 123 is used for behavior detection. Video capture element 124 is used for capturing video. Visual language model 125 is used for inference processes. Real-time interaction unit 210 is used for generating interactive messages. Notification unit 220 is used for information notification processes. The multi-target tracking element 122, AI IoT element 123, video capture element 124, visual language model 125, real-time interaction unit 210, and / or notification unit 220 are, for example, a circuit board, a storage device for stored code, or a chip. The chip is, for example, a central processing unit (CPU), or other programmable general-purpose or special-purpose microcontroller (MCU), microprocessor, digital signal processor (DSP), programmable controller, application-specific integrated circuit (ASIC), graphics processing unit (GPU), image signal processor (ISP), image processing unit (IPU), arithmetic logic unit (ALU), complex programmable logic device (CPLD), field programmable gate array (FPGA), or other similar components or combinations thereof.
[0017] The behavior database 126 and recording unit 230 are used to record and store data, such as any type of fixed or removable random access memory (RAM), read-only memory (ROM), flash memory, hard disk drive (HDD), solid-state drive (SSD), or similar components or combinations thereof, for storing multiple modules or various applications executable by a processor. The behavior database 126 records various consumer behaviors or abnormal behaviors. During the operation of the customer behavior tracking system 1000, the behavior database 126 can be updated based on feedback from staff regarding the identification results or additional augmented information provided by staff, so that the customer behavior tracking system 1000 can obtain more accurate identification results.
[0018] In the checkout area, if key user US1 makes a mistake or misses a transaction, the customer behavior tracking system 1000 can immediately notify key user US1 to re-checkout to avoid errors. The following flowchart illustrates in detail how each of these components operates.
[0019] Please refer to Figure 4, which illustrates a flowchart of a customer behavior tracking method applied to a merchandise checkout area according to an embodiment of this disclosure. The customer behavior tracking method in Figure 4 includes steps S110 to S170. The customer behavior tracking system 1000 and customer behavior tracking method disclosed herein can be applied to both the merchandise checkout area and the merchandise selection area. The following explanation will focus on the merchandise checkout area as an example.
[0020] Please refer to Figure 5, which illustrates an example of step S110 in Figure 4. In step S110, as shown in Figures 3 and 5, at least one image VD is captured by these image capturing units 110. In this step, when multiple customers move around in the mall, the multi-target tracking element 122 of the image capturing unit 110 tracks these customers using multi-target multi-camera (MTMC) technology and displays multiple person frames BX on the image VD.
[0021] Next, please refer to Figure 6, which illustrates an example of step S120 in Figure 4. In step S120, as shown in Figures 3 and 6, key users US1 entering the identification area RG1 are bound by the multi-target tracking element 122 connected to these image capturing units 110. In this step, when the multi-target tracking element 122 detects that a customer has entered the identification area RG1, this customer is bound as key user US1. In one embodiment, the multi-target tracking element 122 can bind the user with the highest overlap with the identification area RG1 as key user US1, or bind the user who has been in the identification area RG1 for the longest time as key user US1. Once a customer leaves the identification area RG1, they are not bound as key user US1.
[0022] Next, please refer to Figures 7A-7B, which illustrate an example of step S130 in Figure 4. In step S130, as shown in Figures 3 and 7A-7B, at least one key action performed by the key user US1 on a single product PD1 is identified through the AI IoT component 123 (for example, key action AT1 is a picking action, and key action AT2 is a putting action). The AI IoT component 123 is, for example, the image recognition component 123a illustrated in Figures 7A and 7B.
[0023] In this step, when the image recognition element 123a detects that the product PD1 is picked up by the key user US1, it is determined to be a key action AT1; when the image recognition element 123a detects that the product PD1 is put down, it is determined to be a key action AT2.
[0024] Please refer to Figures 8A-8B for another example of step S130 in Figure 4. The AI IoT component 123 is, for example, the weight sensing component 123b illustrated in Figures 8A and 8B. When the weight sensing component 123b senses a decrease in weight, it is determined to be a critical action AT1; when the weight sensing component 123b senses a return to normal weight, it is determined to be a critical action AT2.
[0025] Please refer to Figures 9A and 9B, which illustrate another example of step S130 in Figure 4. The AI IoT component 123 is, for example, the infrared sensing component 123c illustrated in Figures 9A and 9B. When the infrared sensing component 123c senses an object passing through infrared lines L1 and L2 in a certain order, it is determined to be a critical action AT1. When the infrared sensing component 123c senses an object passing through infrared lines L1 and L2 in the reverse order, it is determined to be a critical action AT2.
[0026] In addition to the image recognition element 123a, weight sensing element 123b, and infrared sensing element 123c mentioned above, the artificial intelligence Internet of Things element 123 can also be a light-sensitive element.
[0027] Next, please refer to Figure 10, which illustrates an example of steps S140-S150 in Figure 4. In step S140, as shown in Figures 3 and 10, the video recording element 124 records a key video segment SG1 based on at least one key action. The time point T1 of the key action AT1 of picking up the item is the starting point of the key video segment SG1, and the time point T2 of the key action AT2 of putting down the item is the ending point of the key video segment SG1. By setting the starting and ending points, the picking process of the product PD1 can be accurately recorded, so that the subsequent identification process can be more accurate.
[0028] Then, in step S150, as shown in Figures 3 and 10, the visual language model 125 identifies item IT1, one of the product PD1, based on the key video clip SG1. For example, the visual language model 125 can identify information about item IT1 such as the name, quantity, and color of product PD1. Since the key video clip SG1 is captured specifically for product PD1, the visual language model 125 is not affected by too many other products and can thus accurately identify product PD1.
[0029] Next, in step S160, as shown in Figure 3, based on item IT1 of item PD1, the checkout behavior BV1 of key user US1 for item PD1 is identified. This step, for example, involves visual language model 125 using behavior database 126 to identify key user US1's checkout behavior BV1 for item PD1. Behavior database 126 can be updated based on feedback from staff regarding the identification results or additional augmentation information provided by staff.
[0030] Please refer to Figure 11, which illustrates a detailed flowchart of step S160 according to an embodiment of this disclosure. In one embodiment, step S160 includes steps S162 to S163.
[0031] In step S162, the visual language model 125 compares the product item IT1 with multimodal data. The multimodal data is, for example, a product scan record RD1.
[0032] Then, in step S163, the visual language model 125 is queried to confirm whether the checkout behavior is an abnormal behavior. As shown in Figure 1, when product item IT1 does not exist in product scan record RD1, the visual language model 125 will use the behavior database 126 to identify the abnormal behavior of key user US1 missing scan of product PD1.
[0033] As shown in Figure 2, when the product item IT1 is different from the product scan record RD1, the visual language model 125 will use the behavior database 126 to identify the abnormal behavior of the key user US1 accidentally scanning the product PD1. The behavior database 126 can be updated based on feedback from staff regarding the recognition results or additional augmented information provided by staff.
[0034] Next, in step S170, as shown in Figure 3, the notification unit 220 sends a notification signal MS1 based on the checkout behavior BV1 of the key user US1 on product PD1. In the example of the product checkout area, the real-time interaction unit 210 generates the content of the notification signal MS1 based on the checkout behavior BV1 of the key user US1 on product PD1. The notification signal MS1 is displayed, for example, on the checkout screen 500 in Figures 1 and 2, to remind the key user US1 of any abnormalities such as accidental or missed swipes on product PD1. Alternatively, the notification signal MS1 can also directly notify the operator so that staff can go to assist in handling the situation.
[0035] Based on the aforementioned customer behavior tracking method, key video clips SG1 are captured using key actions AT1 and AT2 to accurately identify the checkout behavior BV1 of key user US1. In the event of accidental or missed transactions, timely alerts and interventions can be provided.
[0036] Please refer to Figure 12, which shows a block diagram of a smart self-checkout system 2000 according to an embodiment of this disclosure. The customer behavior tracking system 1000 in the above embodiment can be applied to a smart self-checkout system 2000. The smart self-checkout system 2000 includes the above-described product scanning module 300, the above-described behavior recognition module 100, and the above-described behavior processing module 200.
[0037] The components of the behavior recognition module 100 and the behavior processing module 200 are described in the same way as those in the above embodiments, and will not be repeated here.
[0038] In this embodiment, during the self-checkout process, the intelligent self-checkout system 2000 can capture key video clips SG1 based on key actions AT1 and AT2, so as to correctly identify the checkout behavior of the key user US1 through the key video clips SG1. In case of accidental or missed transactions, timely reminders and handling can be provided.
[0039] In another embodiment, the customer behavior tracking system 1000 described above can also be applied to the merchandise selection area to accurately identify the purchasing behavior BV2 of key user US2 (shown in Figure 14). The following flowchart illustrates in detail how the various components of the customer behavior tracking system 1000 operate in the merchandise selection area.
[0040] Please refer to Figures 13 and 14 simultaneously. Figure 13 illustrates a flowchart of a customer behavior tracking method applied to a shopping area according to an embodiment of this disclosure. Figure 14 illustrates the steps of Figure 13. The customer behavior tracking method in Figure 13 includes steps S210 to S270.
[0041] In step S210, as shown in Figures 3 and 13, at least one image VD is captured by the image capturing unit 110. In this step, when multiple customers move around in the mall, the multi-target tracking element 122 of the image capturing unit 110 tracks these customers using multi-target tracking technology (MTMC).
[0042] Next, in step S220, as shown in Figures 3 and 13, key users US2 entering the identification area RG1 are bound by the multi-target tracking element 122 connected to these image capturing units 110. In this step, when the multi-target tracking element 122 detects that a customer has entered the identification area RG2, this customer is bound as key user US2. In one embodiment, the multi-target tracking element 122 can bind the user with the highest degree of overlap with the identification area RG2 as key user US2, or bind the user who has been in the identification area RG2 for the longest time as key user US2. Once a customer leaves the identification area RG2, they are no longer bound as key user US2.
[0043] Then, in step S230, as shown in Figures 3 and 13, at least one key action performed by the key user US2 on the product PD5 is identified through the AI IoT component 123 (for example, key action AT1 is a picking action, key action AT2 is a putting action, and key action AT3 is a leaving action, such as the key user US2 leaving the identification area RG2). The AI IoT component 123 is, for example, the weight sensing component 123b in Figure 13. When the weight sensing component 123b senses a decrease in weight, it is determined to be key action AT1; when the weight sensing component 123b senses a return to normal weight, it is determined to be key action AT2.
[0044] The AI IoT component 123 is, for example, the image recognition component 123a in Figure 13. When the image recognition component 123a detects that the product PD5 is picked up by the key user US2, it is determined to be a key action AT1; when the image recognition component 123a detects that the product PD5 is put down, it is determined to be a key action AT2.
[0045] Next, in step S240, as shown in Figures 3 and 10, the video recording element 124 records a key video segment SG1 based on at least one key action. The time point T1 of the key action AT1 of picking up the item is the starting point of the key video segment SG2, and the time point T2 of the key action AT2 of putting down the item or the key action AT3 of leaving the item is the ending point of the key video segment SG2. By setting the starting and ending points, the picking process of the product PD5 can be accurately recorded, so that the subsequent identification process can be more accurate.
[0046] Then, in step S250, as shown in Figures 3 and 10, the visual language model 125 identifies item IT2 of item PD5 based on the key video clip SG2.
[0047] Next, in step S260, as shown in Figure 3, based on item IT2 of product PD5, the purchasing behavior BV2 of key user US2 regarding product PD5 is identified. This step, for example, involves visual language model 125 using behavioral database 126 to identify the purchasing behavior BV2 of key user US2 regarding product PD5. Purchasing behavior BV2 includes actions such as directly taking items, repeatedly comparing prices, and checking prices. Behavioral database 126 can be updated based on feedback from staff regarding the identification results or additional augmented information provided by staff.
[0048] Next, in step S270, as shown in Figure 3, the notification unit 220 sends a notification signal MS2 based on the key user US2's purchase behavior BV2 regarding product PD5. In the example of the product selection area, the real-time interaction unit 210 generates the content of the notification signal MS2 based on the key user US2's purchase behavior BV2 regarding product PD5. The notification signal MS2 may be, for example, promotional information or recommendation information displayed on the advertising screen 600 in Figure 13, or a text message displayed on the key user US2's mobile phone to increase purchase motivation.
[0049] According to the above embodiment, through the coordinated operation of AI IoT components 123 (such as ceiling cameras and shelf weight sensors), key actions AT1, AT2, and AT3, such as the action of holding goods and changes in weight on the machine, are analyzed to capture the key video clip SG2 required by the visual language model 125. The visual language model 125 can accurately identify the purchasing behavior BV2 of the key user US2 through the short key video clip SG2. Using the information from the purchasing behavior BV2, the key user US2 can receive real-time notifications of promotional activities and related products during the shopping process, which can not only provide the key user US2 with a reference for purchasing, but also effectively increase the key user US2's purchase motivation.
[0050] In summary, although this disclosure has been presented above with examples, it is not intended to limit the scope of this disclosure. Those skilled in the art to which this disclosure pertains can make various modifications and refinements without departing from the spirit and scope of this disclosure. Therefore, the scope of protection of this disclosure shall be determined by the appended claims.
[0051] 100: Behavior Recognition Module 110: Image Capture Unit 120: Behavioral Analysis Unit 122: Multi-target tracking element 123: Artificial Intelligence Internet of Things Components 123a: Image recognition element 123b: Weight sensing element 123c: Infrared sensing element 124: Video capture element 125: Visual Language Model 126: Behavioral Database 200: Behavior Processing Module 210: Real-time Interactive Unit 220: Notification Unit 230: Recording Unit 300: Product Scanning Module 500: Checkout screen 600: Advertising Screen 1000: Customer Behavior Tracking System 2000: Intelligent Self-Checkout System AT1, AT2, AT3: Key Actions BX: Character Frame BC: Barcode BV1: Checkout Action BV2: Purchasing Behavior IT1, IT2: Product Items MS1, MS2: Notification signals PD1, PD2, PD3, PD4, PD5: Commodity RG1, RG2: Identification areas SG1, SG2: Key Video Clips T1, T2: Time points US1, US2: Key Users VD: Image S110, S120, S130, S140, S150, S160, S162, S163, S170, S210, S220, S230, S240, S250, S260, S270: Steps
Claims
1. A customer behavior tracking method, comprising: By connecting a multi-target tracking element to one of a plurality of image capturing units, a key user entering a recognition area is identified; by using at least one Artificial Intelligence of Things (AIoT) element, at least one key action of the key user on a single product is identified, the key action including a picking-up action and a putting-down action; based on the at least one key action, a key video clip is captured, the time point of the picking-up action being the start point of the key video clip, and the time point of the putting-down action being the end point of the key video clip; based on the key video clip, a visual language model (VLM) is used to identify a product item of the single product; and based on the product item of the single product, the key user's checkout behavior or purchase behavior on the single product is identified; wherein the step of identifying the key user's checkout behavior or purchase behavior on the single product based on the product item of the single product includes: The visual language model compares the product item with multimodal data, including a product scan record, a customer behavior video, or product weight change data; and by querying the visual language model, confirms whether the checkout behavior is an abnormal behavior.
2. The customer behavior tracking method as described in claim 1, wherein the at least one artificial intelligence Internet of Things element is an image recognition element, a weight sensing element, an infrared sensing element, or a light sensing element.
3. The customer behavior tracking method as described in claim 1, wherein the at least one key action further includes a departure action in which the key user leaves the identification area, and the time point of the departure action is the end point of the key video segment.
4. The customer behavior tracking method as described in claim 1 further includes: A notification signal is sent based on the key user's checkout or purchase behavior for a single product.
5. A customer behavior tracking system, comprising: A behavior recognition module includes: a plurality of image capturing units; a behavior analysis unit including: a multi-target tracking element connected to the image capturing units, the multi-target tracking element being used to identify a key user entering a recognition area; at least one Artificial Intelligence of Things (AIoT) element for identifying at least one key action of the key user on a single product, the at least one key action including a picking-up action and a putting-down action; a video recording element for recording a key video clip based on the at least one key action, the time point of the picking-up action being the start point of the key video clip, the time point of the putting-down action being the end point of the key video clip; and a Vision Language Model (VLM) for identifying a product item of the single product based on the key video clip, the product item of the single product being used to identify a checkout behavior or a purchase behavior of the key user on the single product; The visual language model is further used to compare the product item with multimodal data, which includes a product scan record, a customer behavior video, or product weight change data; and the visual language model responds in a questioning manner to whether the checkout behavior is an abnormal behavior.
6. The customer behavior tracking system as described in claim 5, wherein the at least one artificial intelligence Internet of Things element is an image recognition element, a weight sensing element, an infrared sensing element, or a light sensing element.
7. The customer behavior tracking system as described in claim 5, wherein the at least one key action further includes a departure action in which the key user leaves the identification area, and the time point of the departure action is the end point of the key video segment.
8. The customer behavior tracking system as described in claim 5 further includes: An action processing module includes: a notification unit for issuing a notification signal based on the key user's checkout action or purchase action for a single product.
9. The customer behavior tracking system as described in claim 5 further includes: A behavioral database is used by the visual language model to identify the checkout or purchase behavior of the key user for a single product. The behavioral database is updated based on feedback or augmentation information.
10. A behavior analysis unit, comprising: A multi-target tracking element connected to a plurality of image capturing units, the multi-target tracking element being used to track a key user entering a recognition area; The system includes at least one Artificial Intelligence of Things (AIoT) component for identifying at least one key action performed by the key user on a single product, the key action including a picking-up action and a putting-down action; a video recording component for recording a key video clip based on the at least one key action, the time point of the picking-up action being the start point of the key video clip, and the time point of the putting-down action being the end point of the key video clip; and a Vision Language Model (VLM) for identifying a product item of the single product based on the key video clip, the product item of the single product being used to identify a checkout behavior of the key user on the single product; wherein the VLM is further used to compare the product item with multimodal data, the multimodal data including a product scan record, a customer behavior video, or product weight change data; and the VLM responds in a questioning manner to whether the checkout behavior is an abnormal behavior.
11. The behavior analysis unit as described in claim 10, wherein the at least one AI Internet of Things element is an image recognition element, a weight sensing element, an infrared sensing element, or a light sensing element.