Method for managing offline market using artificial intelligence agent, and artificial intelligence agent using same
An AI agent with deep learning models and LLMs analyzes customer behaviors in offline markets, addressing the limitations of conventional methods by providing accurate marketing strategies and incident detection.
Patent Information
- Application Number
- PCT/KR2023/022005
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-03
AI Technical Summary
Conventional methods for managing offline markets fail to accurately reflect customer tendencies and behaviors, leading to ineffective marketing strategies and difficulty in identifying incidents due to reliance on statistical data and real-time video monitoring.
An artificial intelligence agent utilizing image analysis deep learning models to detect and analyze customer behaviors, combined with a Large Language Model (LLM) to generate marketing strategies based on real-time data and historical sales information.
Enables accurate identification of customer behaviors and market status, allowing for effective marketing strategies and efficient incident detection, reducing the need for manual video review.
Smart Images

Figure KR2023022005_03072025_PF_FP_ABST
Abstract
Description
A method for managing an offline market using an artificial intelligence agent and an artificial intelligence agent using the same.
[0001] The present invention relates to a method for managing an offline market using an artificial intelligence agent and an artificial intelligence agent using the same.
[0002] Typically, offline markets use POS data to manage sales and inventory, manage customers through memberships, and analyze sales data using statistical data. This data, in turn, allows for the development of marketing strategies to generate profits by leveraging product-specific trends and consumption patterns.
[0003] Additionally, in order to check the current status within the offline market, the manager monitors the footage captured by the camera in real time to check for accidents or incidents that occur within the offline market.
[0004] However, these conventional methods have the problem of not reflecting the tendencies or characteristics of customers located in actual offline markets because they only use statistical data.
[0005] In addition, since the manager must monitor the video in real time to check the current status within the offline market, it is difficult to identify specific situations, and there is a problem of not being able to accurately recognize situations that occurred during the time when the video was not monitored, and there is also a problem of having to check the recorded videos one by one to identify specific situations that were not monitored.
[0006] Therefore, the applicant proposes a method to establish an accurate marketing strategy based on product-specific trends and consumption patterns, and to accurately confirm the current status of the offline market.
[0007] The purpose of the present invention is to solve all of the problems of the above-mentioned prior art.
[0008] In addition, another purpose of the present invention is to enable the establishment of a marketing strategy by reflecting the behavior and characteristics of customers within an offline market.
[0009] In addition, another purpose of the present invention is to enable accurate information on the current status of an offline market using vision sensor data.
[0010] In addition, the present invention allows the image analysis deep learning model to analyze the image, and to detect object detection information that detects at least one of a person and an object in the image, face and head detection information that detects at least one of a face and a head of a person in the image, posture and behavior analysis information that extracts key points of a person detected in the image and analyzes the person's posture and behavior, facial expression analysis information that estimates the expression of a face detected in the image, facial feature detection information that extracts feature points of a face detected in the image, attribute analysis information that estimates at least one of a gender and an age of a person detected in the image, clothing analysis information that analyzes the clothing of a person detected in the image, safety equipment non-wearing detection information that tracks whether a person detected in the image is wearing safety equipment, interest measurement information that analyzes where a person detected in the image is interested in based on the location and head direction of a person detected in the image, crowd count measurement information that estimates the number of people located in a specific area in the image, abnormal behavior judgment information that estimates abnormal behavior of a person detected in the image, and abnormal behavior information that estimates the current situation of a person detected in the image. Another purpose is to output the task results that include at least some of the situational judgment information.
[0011] In addition, the present invention, with reference to the task results, includes: person identification information distinguishing each person located within the specific space; entry / exit counting information counting the number of people who passed through a specific area within the specific space; person attribute information for each person located within the specific space; dangerous area intrusion information detecting a person who has intruded into a dangerous area within the specific space; waiting time information measuring a waiting time for using a specific service within the specific space; interest information of people located within the specific space; parking lot information within the specific space; movement tracking information tracking the movement of at least one of a person and an object within the specific space; heat map information estimating the density of people detected within the specific space; group recognition information analyzing a group of people detected within the specific space; employee recognition information analyzing whether a person located within the specific space is an employee; lost item detection information detecting a newly detected or missing object within the specific space; abnormal behavior information detecting an abnormal behavior of a person located within the specific space; crime information detecting a crime that occurred within the specific space; and information about a crime detected within a video recorded for a preset period of time. Another purpose is to obtain the image analysis status information including at least one of image summary information that organizes only the necessary parts according to preset conditions and product inventory information that detects the inventory of products placed within the specific space.
[0012] A representative configuration of the present invention to achieve the above purpose is as follows.
[0013] According to one embodiment of the present invention, there is provided a method for managing an offline market using an artificial intelligence agent, comprising: (a) when at least one video stream is transmitted from at least one camera filming a specific space of at least one offline market, the artificial intelligence agent inputs at least one video belonging to the video stream into at least one video analysis deep learning model, causes the video analysis deep learning model to perform each task on the video through video analysis, and outputs task results for each task, acquires video analysis status information of the offline market by referring to the task results, vectorizes the video analysis status information and stores it in a vector database, and vectorizes text analysis status information including at least a portion of POS (Point Of Sales) data and POG (planogram) data corresponding to the offline market and stores it in the vector database; And (b) when a natural language command related to the management of the offline market is input, the artificial intelligence agent inputs the natural language command into an LLM (Large Language Model) to cause the LLM to generate a first text command to an n-th text command for performing the natural language command, where n is an integer greater than or equal to 1, and, according to a k-th text command, where k is an integer greater than or equal to 1 and less than or equal to n, obtains at least some of the image analysis state information and the text analysis state information corresponding to the k-th text command from the vector database as k-th state information, thereby obtaining first state information to n-th state information corresponding to the first text command to the n-th text command, and generates response information corresponding to the natural language command with reference to the first state information to the n-th state information.
[0014] In the above embodiment, in the step (a), the artificial intelligence agent causes the image analysis deep learning model to analyze the image, and detects object detection information that detects at least one of a person and an object in the image, face and head detection information that detects at least one of a person's face and head in the image, posture and behavior analysis information that extracts key points of a person detected in the image and analyzes the person's posture and behavior, facial expression analysis information that estimates the expression of a face detected in the image, facial feature detection information that extracts feature points of a face detected in the image, attribute analysis information that estimates at least one of a gender and an age of a person detected in the image, clothing analysis information that analyzes the clothing of a person detected in the image, safety equipment non-wearing detection information that tracks whether a person detected in the image is wearing safety equipment, interest measurement information that analyzes where a person detected in the image is interested in based on the location and head direction of a person detected in the image, crowd count measurement information that estimates the number of people located in a specific area in the image, abnormal behavior judgment information that estimates abnormal behavior of a person detected in the image, and The task results may be output that include at least a portion of the abnormal situation judgment information that estimates the current situation of the person detected in the above image.
[0015] In the above embodiment, in the step (a), the artificial intelligence agent, with reference to the task results, identifies each person located within the specific space, counts entry / exit counting information counting the number of people who passed through a specific area within the specific space, identifies each person located within the specific space, detects a person who has intruded into a dangerous area within the specific space, measures waiting time information for using a specific service within the specific space, identifies interest information of people located within the specific space, identifies parking lot information within the specific space, identifies movement paths of at least one person and object within the specific space, identifies heat map information estimating the density of people detected within the specific space, identifies group recognition information analyzing a group of people detected within the specific space, identifies employee recognition information analyzing whether a person located within the specific space is an employee, identifies lost items detecting newly detected or missing items within the specific space, identifies abnormal behavior information detecting abnormal behavior of a person located within the specific space, identifies crimes detected occurring within the specific space Information, video summary information that organizes only the necessary parts according to preset conditions within a video recorded for a preset period of time, and product inventory information that detects the inventory of products placed within the specific space can be obtained.
[0016] In the above embodiment, in the step (b), the artificial intelligence agent may cause the LLM to obtain at least a portion of the image analysis status information and at least a portion of the POS data as the first status information to the n-th status information, and may generate at least one of business status information and business plan information related to the business of the offline market as the response information by referring to at least a portion of the image analysis status information and at least a portion of the POS data.
[0017] In the above embodiment, the artificial intelligence agent may cause the LLM to generate at least one of the business status information and the business plan information as the response information by referencing a plurality of pieces of image analysis status information and a plurality of pieces of POS data corresponding to two or more offline markets operated in different locations.
[0018] In the above embodiment, in the step (b), the artificial intelligence agent may cause the LLM to obtain at least some of the image analysis status information, at least some of the POS data, and at least some of the POG data as the first status information to the n-th status information, and generate POG update information for updating the POG as the response information by referring to at least some of the image analysis status information, at least some of the POS data, and at least some of the POG data.
[0019] In the above embodiment, the artificial intelligence agent may cause the LLM to analyze a purchasing trend for the specific product at the specific location by referring to at least some of the following information: movement path information for moving to the specific location within the specific space, exposure viewing angle information for the specific location within the specific space, movement information of people moving to the specific location within the specific space, density of people at the specific location, interest in the specific location, time information for which people stay within a preset radius from the specific location, identifiers of products displayed by floor of a shelf at the specific location, identifiers of other products displayed near the displayed products, information on picking up the specific product displayed at the specific location, sales information of the specific product according to the POS data, and attribute information of a person who purchased the specific product, and generate the POG update information that determines whether the specific product displayed at the specific location has changed in response to the analyzed purchasing trend.
[0020] In the above embodiment, in the step (b), the artificial intelligence agent may cause the LLM to obtain at least some of the image analysis status information as the first status information to the n-th status information, and generate market status information related to the status of the offline market as the response information by referring to at least some of the image analysis status information.
[0021] In the above embodiment, the LLM may be generated by fine-tuning a pre-trained model using large-scale language data using target learning data related to management of the offline market.
[0022] In the above embodiment, after the step (a), the artificial intelligence agent may further include a step of adding at least one of the following to the image displayed on the administrator terminal and displaying it: entry / exit counting information counting the number of people who have passed through a specific area within the specific space; dangerous area intrusion information detecting a person who has intruded into a dangerous area located within the specific space; waiting time information measuring a waiting time for using a specific service within the specific space; parking lot information within the specific space; heat map information estimating the density of people detected within the specific space; abnormal behavior information detecting abnormal behavior of people located within the specific space; crime information detecting a crime that has occurred within the specific space; and product inventory information detecting the inventory of products placed within the specific space;
[0023] In the above embodiment, after step (a), if status search request information of the offline market according to a specific condition is obtained, the artificial intelligence agent may further include a step of causing the LLM to obtain at least one specific image analysis information corresponding to the specific condition from the vector database, and to output status information of the offline market according to the specific condition as the response information by referring to the obtained specific image analysis information.
[0024] According to another embodiment of the present invention, an artificial intelligence agent for managing an offline market comprises: a memory storing instructions for managing the offline market; and a processor performing operations for managing the offline market according to the instructions stored in the memory;, wherein the processor comprises: (I) when at least one video stream is transmitted from at least one camera filming a specific space of at least one offline market, inputting at least one video belonging to the video stream into at least one video analysis deep learning model to cause the video analysis deep learning model to perform each task on the video through video analysis and output task results for each task, obtaining video analysis status information of the offline market with reference to the task results, vectorizing the video analysis status information and storing it in a vector database, and vectorizing text analysis status information including at least a part of POS (Point Of Sales) data and POG (planogram) data corresponding to the offline market and storing it in the vector database, and (II) when a natural language command related to the management of the offline market is input, inputting the natural language command into an LLM (Large Language Model) to cause the LLM to generate a first text command to an n-th text command for performing the natural language command, wherein n is an integer greater than or equal to 1, and the k-th text command - An artificial intelligence agent is provided that performs a process of acquiring first state information to n-th state information corresponding to the first text command to the n-th text command by acquiring at least some of the image analysis state information and the text analysis state information corresponding to the k-th text command from the vector database as k-th state information, and generating response information corresponding to the natural language command by referring to the first state information to the n-th state information.;
[0025] In the other embodiment, the processor, in the process (I), causes the image analysis deep learning model to analyze the image, and detect object detection information that detects at least one of a person and an object in the image, face and head detection information that detects at least one of a face and a head of a person in the image, posture and behavior analysis information that extracts key points of a person detected in the image and analyzes the person's posture and behavior, facial expression analysis information that estimates an expression of a face detected in the image, facial feature detection information that extracts feature points of a face detected in the image, attribute analysis information that estimates at least one of a gender and an age of a person detected in the image, clothing analysis information that analyzes the clothing of a person detected in the image, safety equipment non-wearing detection information that tracks whether a person detected in the image is wearing safety equipment, interest measurement information that analyzes where a person detected in the image is interested in based on a location and a head direction of a person detected in the image, crowd count measurement information that estimates the number of people located in a specific area in the image, abnormal behavior judgment information that estimates an abnormal behavior of a person detected in the image, and The task results may be output that include at least a portion of the abnormal situation judgment information that estimates the current situation of the person detected in the above image.
[0026] In the above other embodiment, in the process (I), with reference to the task results, person identification information distinguishing each person located in the specific space, entry / exit counting information counting the number of people who passed through a specific area in the specific space, person attribute information for each person located in the specific space, dangerous area intrusion information detecting a person who has intruded into a dangerous area located in the specific space, waiting time information measuring a waiting time for using a specific service in the specific space, interest information of people located in the specific space, parking lot information in the specific space, movement tracking information tracking the movement of at least one of a person and an object in the specific space, heat map information estimating the density of people detected in the specific space, group recognition information analyzing a group of people detected in the specific space, employee recognition information analyzing whether a person located in the specific space is an employee, lost item detection information detecting a newly detected or missing object in the specific space, abnormal behavior information detecting an abnormal behavior of a person located in the specific space, crime information detecting a crime that occurred in the specific space, and a preset time It is possible to obtain the video analysis status information including at least one of video summary information that organizes only the necessary parts according to preset conditions within the video recorded during the recording, and product inventory information that detects the inventory of products placed within the specific space.
[0027] In the other embodiment, the processor, in the process (II), causes the LLM to obtain at least a portion of the image analysis status information and at least a portion of the POS data as the first status information to the n-th status information, and generate at least one of business status information and business plan information related to the business of the offline market as the response information by referring to at least a portion of the image analysis status information and at least a portion of the POS data.
[0028] In another embodiment, the processor may cause the LLM to generate at least one of the business status information and the business plan information as the response information by referencing a plurality of pieces of image analysis status information and a plurality of pieces of POS data corresponding to two or more offline markets operated at different locations.
[0029] In the other embodiment, the processor, in the process (II), may cause the LLM to obtain at least some of the image analysis state information, at least some of the POS data, and at least some of the POG data as the first state information to the n-th state information, and generate POG update information for updating the POG as the response information by referring to at least some of the image analysis state information, at least some of the POS data, and at least some of the POG data.
[0030] In another embodiment, the processor may cause the LLM to analyze a purchasing trend for the specific product at the specific location by referring to at least some of, with respect to a specific location in the POG, movement path information for moving to the specific location within the specific space, exposure viewing angle information for the specific location within the specific space, movement information of people moving to the specific location within the specific space, density of people at the specific location, interest in the specific location, time information for which people stay within a preset radius from the specific location, identifiers of products displayed by floor of a shelf at the specific location, identifiers of other products displayed near the displayed products, information on picking up the specific product displayed at the specific location, sales information of the specific product according to the POS data, and attribute information of a person who purchased the specific product, and generate the POG update information that determines whether the specific product displayed at the specific location has changed in response to the analyzed purchasing trend.
[0031] In the other embodiment, the processor, in the process (II), may cause the LLM to obtain at least some of the image analysis status information as the first status information to the n-th status information, and generate market status information related to the status of the offline market as the response information by referring to at least some of the image analysis status information.
[0032] In another embodiment, the LLM may be generated by fine-tuning a pre-trained model using large-scale language data using target learning data related to management of the offline market.
[0033] In another embodiment, the processor may further perform a process of adding, to the image displayed on the administrator terminal, at least one of the following: entry / exit counting information counting the number of people who have passed through a specific area within the specific space; dangerous area intrusion information detecting a person who has intruded into a dangerous area within the specific space; waiting time information measuring a waiting time for using a specific service within the specific space; parking lot information within the specific space; heat map information estimating the density of people detected within the specific space; abnormal behavior information detecting abnormal behavior of people located within the specific space; crime information detecting a crime that has occurred within the specific space; and product inventory information detecting the inventory of products placed within the specific space.
[0034] In the above other embodiment, after the process (I), if status search request information of the offline market according to a specific condition is obtained, the processor may further perform a process of causing the LLM to obtain at least one specific image analysis information corresponding to the specific condition from the vector database, and outputting status information of the offline market according to the specific condition as the response information by referring to the obtained specific image analysis information.
[0035] In addition, a computer-readable recording medium for recording a computer program for executing the method of the present invention is further provided.
[0036] According to the present invention, it is also possible to establish a marketing strategy by reflecting the behavior and characteristics of customers within an offline market.
[0037] According to the present invention, it is possible to accurately determine the current status of an offline market using vision sensor data.
[0038] According to the present invention, the image analysis deep learning model analyzes the image and detects object detection information that detects at least one of a person and an object in the image, face and head detection information that detects at least one of a face and a head of a person in the image, posture and behavior analysis information that extracts key points of a person detected in the image and analyzes the person's posture and behavior, facial expression analysis information that estimates the expression of a face detected in the image, facial feature detection information that extracts feature points of a face detected in the image, attribute analysis information that estimates at least one of a gender and an age of a person detected in the image, clothing analysis information that analyzes the clothing of a person detected in the image, safety equipment non-wearing detection information that tracks whether a person detected in the image is wearing safety equipment, interest measurement information that analyzes where a person detected in the image is interested based on the location and head direction of a person detected in the image, crowd count measurement information that estimates the number of people located in a specific area in the image, abnormal behavior judgment information that estimates abnormal behavior of a person detected in the image, and abnormal behavior information that estimates the current situation of a person detected in the image. The task results may be output that include at least some of the situational judgment information.
[0039] According to the present invention, with reference to the task results, person identification information distinguishing each person located in the specific space, entry / exit counting information counting the number of people who passed through a specific area in the specific space, person attribute information for each person located in the specific space, dangerous area intrusion information detecting a person who has intruded into a dangerous area located in the specific space, waiting time information measuring a waiting time for using a specific service in the specific space, interest information of people located in the specific space, parking lot information in the specific space, movement tracking information tracking the movement of at least one of a person and an object in the specific space, heat map information estimating the density of people detected in the specific space, group recognition information analyzing a group of people detected in the specific space, employee recognition information analyzing whether a person located in the specific space is an employee, lost item detection information detecting a newly detected or missing object in the specific space, abnormal behavior information detecting an abnormal behavior of a person located in the specific space, crime information detecting a crime that occurred in the specific space, and information about a crime detected in a video recorded for a preset period of time. It is possible to obtain the image analysis status information including at least one of image summary information that organizes only the necessary parts according to preset conditions and product inventory information that detects the inventory of products placed within the specific space.
[0040] The drawings attached below for use in explaining embodiments of the present invention are only some of the embodiments of the present invention, and a person having ordinary knowledge in the technical field to which the present invention pertains (hereinafter “ordinary skilled in the art”) can obtain other drawings based on these drawings without performing an inventive work.
[0041] FIG. 1 schematically illustrates an artificial intelligence agent that manages an offline market according to one embodiment of the present invention.
[0042] FIG. 2 schematically illustrates a system for managing an offline market using an artificial intelligence agent according to one embodiment of the present invention.
[0043] FIG. 3 schematically illustrates a method for managing an offline market using an artificial intelligence agent according to another embodiment of the present invention.
[0044] FIG. 4 illustrates image analysis status information of an offline market as an example in a method for managing an offline market using an artificial intelligence agent according to another embodiment of the present invention.
[0045] The following detailed description of the present invention refers to the accompanying drawings, which illustrate specific embodiments in which the present invention may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the present invention. It should be understood that the various embodiments of the present invention, while different from each other, are not necessarily mutually exclusive. For example, specific shapes, structures, and characteristics described herein may be modified and implemented from one embodiment to another without departing from the spirit and scope of the present invention. Furthermore, it should be understood that the positions or arrangements of individual components within each embodiment may be modified without departing from the spirit and scope of the present invention. Accordingly, the following detailed description is not intended to be limiting, and the scope of the present invention is to include the scope of the claims and all equivalents thereof. Like reference numerals in the drawings represent the same or similar elements throughout the several aspects.
[0046] Hereinafter, various preferred embodiments of the present invention will be described in detail with reference to the attached drawings so that a person having ordinary skill in the art to which the present invention pertains can easily practice the present invention.
[0047] FIG. 1 schematically illustrates an artificial intelligence agent for managing an offline market according to one embodiment of the present invention. Referring to FIG. 1, the artificial intelligence agent (100) may include a memory (101) in which instructions for managing an offline market are stored, and a processor (102) for performing an operation for managing an offline market according to the instructions stored in the memory (101).
[0048] Specifically, the artificial intelligence agent (100) may typically achieve desired system performance using, but is not limited to, a combination of computing devices (e.g., devices that may include computer processors, memory, storage, input devices and output devices, and other components of conventional computing devices; electronic communication devices such as routers, switches, etc.; electronic information storage systems such as network attached storage (NAS) and storage area networks (SAN)) and computer software (i.e., instructions that cause the computing device to function in a particular manner).
[0049] Additionally, the processor of the computing device may include hardware components such as a Micro Processing Unit (MPU) or a Central Processing Unit (CPU), cache memory, and a data bus. Furthermore, the computing device may further include software components such as an operating system and applications that perform specific purposes.
[0050] However, this does not exclude the case where the computing device includes an integrated processor in which the medium, processor, and memory for implementing the present invention are integrated.
[0051] Meanwhile, referring to FIG. 2, when at least one video stream is transmitted from at least one camera (10_1, 10_2, …, 10_m) that films a specific space of at least one offline market, the artificial intelligence agent (100) performs each task in the video through at least one video analysis deep learning model (110) to obtain task results for each task, obtains video analysis status information of the offline market by referring to the task results through the post process module (120), vectorizes the video analysis status information through the vector generation module (130) and stores it in a vector database (140), and vectorizes text analysis status information including at least a part of POS (Point Of Sales) data (20) and POG (planogram) data (30) corresponding to the offline market and stores it in the vector database (140). And, when a natural language command (1) related to the management of an offline market is input, the artificial intelligence agent (100) inputs the natural language command (1) into an LLM (Large Language Model) (150) so that the LLM generates a first text command to an n-th text command for performing the natural language command, and obtains at least some of the image analysis state information and the text analysis state information corresponding to the k-th text command from the vector database (140) as k-th state information according to the k-th text command, thereby obtaining the first state information to the n-th state information corresponding to the first text command to the n-th text command (ultimately by changing k from 1 to n), and generates response information (2) corresponding to the natural language command (1) by referring to the first state information to the n-th state information. In the above, n may be an integer greater than or equal to 1, and k may be an integer greater than or equal to 1 and less than or equal to n.
[0052] A method for managing an offline market using an artificial intelligence agent according to one embodiment of the present invention configured as described above is described below with reference to FIG. 3.
[0053] First, the artificial intelligence agent (100) can obtain each task result (S10) by analyzing at least one image acquired from at least one camera filming a specific space of at least one offline market through at least one image analysis deep learning model.
[0054] That is, when at least one video stream is transmitted from at least one camera that films a specific space of at least one offline market, the artificial intelligence agent (100) can input at least one video belonging to the video stream into at least one video analysis deep learning model, and cause the video analysis deep learning model to perform each task in the video through video analysis and output task results for each task. At this time, each camera has a respective camera view (i.e., a viewing frustum), and each camera can be installed so that the entire specific space of the offline market can be filmed through each camera view, and each video analysis deep learning model can be composed of different video analysis deep learning models for performing each task.
[0055] For example, an artificial intelligence agent (100) may have an image analysis deep learning model analyze an image and generate object detection information that detects at least one person or object in the image. The object may include a product, an object located within a specific space in an offline market, or a vehicle located in a parking lot within the specific space.
[0056] Additionally, the artificial intelligence agent (100) can cause the image analysis deep learning model to generate face and head detection information by detecting at least one human face and head in the image. At this time, there is no limit to the number of detectable faces and heads, and faces and heads can be detected in environments where cameras are installed at various angles.
[0057] In addition, the artificial intelligence agent (100) can cause the image analysis deep learning model to extract key points of a person detected in the image, i.e., the human skeleton, and generate posture and behavior analysis information by analyzing the person's posture and behavior. That is, the image analysis deep learning model can predict the person's posture by referring to the key points of the person detected in the image, and can identify the actions taken by the person in the image, i.e., falling, sitting or standing, picking up an object, etc. based on the predicted posture, and can be used to analyze risk detection, interest measurement, risk detection, etc. based on human behavior analysis.
[0058] Additionally, the artificial intelligence agent (100) can cause a video analysis deep learning model to generate facial expression analysis information that estimates facial expressions detected in the video. The detected expressions can be analyzed to predict the person's mood, such as neutral, happy, surprised, angry, or sad. Furthermore, the analysis of facial expressions of visitors to offline markets can be used to analyze interest, marketing, and advertising effectiveness.
[0059] In addition, the artificial intelligence agent (100) can cause the image analysis deep learning model to generate facial feature point detection information by extracting facial feature points detected in the image, i.e., eyebrows, eyes, nose, mouth, facial outline, etc. In this case, the image analysis deep learning model can accurately predict the location of the nose / mouth / facial outline even for parts covered by a mask, and can be used to analyze information such as the location the face is looking at and emotions by identifying key points of the face.
[0060] Additionally, the artificial intelligence agent (100) can cause the image analysis deep learning model to generate attribute analysis information that estimates at least one of the gender and age of a person detected in the image. In this case, the image analysis deep learning model predicts attributes based on the characteristics of the entire body as well as the face of the detected person, so it can maintain accuracy even in situations where small objects or masks are worn. The attribute analysis information can be used to analyze the characteristics of visitors to an offline market.
[0061] In addition, the artificial intelligence agent (100) can cause the image analysis deep learning model to generate clothing analysis information by analyzing the clothing of a person detected in the image. At this time, the image analysis deep learning model can estimate the clothing of the person detected in the image by distinguishing between the upper and lower garments, and can extract information on the shape and color of the upper garment (e.g., short sleeves / long sleeves and color distinction) and the shape and color of the lower garment (e.g., short pants / long pants / skirt and color distinction), and the clothing analysis information can be used to analyze the characteristics of the detected person or track the movement line.
[0062] Additionally, the artificial intelligence agent (100) can have a video analysis deep learning model generate safety equipment non-wearing detection information by tracking whether a person detected in the video is wearing safety equipment. In this case, the information tracking real-time safety equipment wearing can be used for risk detection and prevention at work sites within offline markets, and can be used to track workers' movements in real time.
[0063] In addition, the artificial intelligence agent (100) can cause the image analysis deep learning model to generate interest measurement information that analyzes where the detected person is interested based on the location and head direction of the person detected in the image. At this time, the image analysis deep learning model can adjust the time the person stays in front of the area of interest and the time the person looks at the area of interest so that the interest can be applied differently depending on the situation, and the interest measurement information can be used to analyze the interest of visitors to an offline market.
[0064] Additionally, the artificial intelligence agent (100) can cause the image analysis deep learning model to generate crowd counting information that estimates the number of people located in a specific area within the image. In this case, the image analysis deep learning model can roughly estimate the number of people gathered in a specific area within the image, and since it estimates the approximate number of people without having to count each person individually, it can perform the task with a small amount of computation.
[0065] In addition, the artificial intelligence agent (100) can cause the video analysis deep learning model to generate abnormal behavior judgment information that estimates the abnormal behavior of a person detected in the video. At this time, the video analysis deep learning model detects the person in the video and detects abnormal behavior and events occurring, and can detect problematic behaviors such as a specific person persistently following others as if stalking them or stealing items.
[0066] Additionally, the artificial intelligence agent (100) can cause a video analysis deep learning model to generate abnormal situation judgment information by estimating the current situation of a person detected in the video. At this time, the video analysis deep learning model can identify the behavior and characteristics of the person detected in the video and predict the situation in which the person is located, such as falling, being caught, or remaining motionless for a long period of time.
[0067] However, the task results of the present invention are not limited to the tasks exemplified above, and various task results can be generated to identify behavioral patterns and characteristics of people, objects, etc. in offline markets.
[0068] Next, the artificial intelligence agent (100) can obtain image analysis status information of an offline market by referring to the task results, and vectorize the image analysis situation information and store it in a vector database (S20).
[0069] For example, referring to FIG. 4, the image analysis status information of an offline market obtained by an artificial agent (100) by referring to task results is described as follows.
[0070] The artificial intelligence agent (100) can obtain person identification information (○01) that distinguishes each person located within a specific space by referring to the task results. At this time, the person identification information may be arbitrary information for distinguishing each person detected in the video, and unique identification information may be assigned to each person entering the offline market, and may be assigned until the person exits the offline market and used to track the person moving within the offline market. In addition, the person identification information may be assigned in an anonymized state to protect the personal information of the person detected in the video and may be assigned to distinguish the anonymized people.
[0071] In addition, the artificial intelligence agent (100) can obtain entry / exit counting information (○02) that counts the number of people who passed through a specific area within a specific space by referring to the task results.
[0072] At this time, the artificial intelligence agent (100) can count the number of people passing through the area of interest based on the people detected in the video and their movement paths, and can count the number of people passing through based on lines or count the number of people entering and exiting based on regions. In other words, the artificial intelligence agent (100) can count people passing through a designated line or entering or exiting a designated area in the video captured over a preset period of time.
[0073] In addition, the artificial intelligence agent (100) can obtain human attribute information (○03) for each person located within a specific space by referring to the task results. At this time, the artificial intelligence agent (100) can obtain human attribute information for each person located within a specific space by matching the attribute information of the people detected in the image to each person located within the specific space.
[0074] In addition, the artificial intelligence agent (100) can obtain information on a dangerous area intrusion (○04) by detecting a person who has intruded into a dangerous area located within a specific space by referring to the task results. At this time, the artificial intelligence agent (100) can obtain information on a dangerous area intrusion by referring to information on whether a person detected in the video is located in an area set as a dangerous area within a specific space or has moved to a dangerous area within a specific space through movement tracking.
[0075] In addition, the artificial intelligence agent (100) can obtain waiting time information (○05) by measuring the waiting time for using a specific service within a specific space by referencing task results. At this time, the artificial intelligence agent (100) can obtain waiting time information by measuring the time spent in a specific area providing a specific service through movement tracking. For example, the artificial intelligence agent (100) can obtain waiting time information by measuring the time of moving to the checkout counter of an offline market and the time of leaving the counter after completing the checkout.
[0076] In addition, the artificial intelligence agent (100) can obtain interest information (○06) of people located within a specific space by referring to task results. For example, the artificial intelligence agent (100) can obtain interest information by referring to the interest measurement information of a person detected in an image, such as the time for which the person looks at a product or billboard located in the direction of the person's gaze for a preset period of time or more, or through a behavioral pattern such as looking at a specific product or picking up a specific product to check information about a specific product.
[0077] In addition, the artificial intelligence agent (100) can obtain parking lot information (○07) within a specific space by referring to task results. That is, it can count the number of vehicles detected in the image entering and exiting the parking area, and obtain information on currently available parking lots by referring to the number of parking lots in the parking area and the number of currently counted vehicles.
[0078] In addition, the artificial intelligence agent (100) can obtain heat map information (○08) that estimates the density of people detected in a specific space by referring to the task results. That is, the artificial intelligence agent (100) visualizes areas with high population density by time zone within the area of interest, and obtains heat map information that makes it easy to understand the density of the entire space by normalizing the entire data based on the area with the highest density, thereby enabling the use of spatial layout information of an offline market by understanding popular spaces and relatively less popular areas within the entire space. For example, the artificial intelligence agent (100) can obtain heat map information by accumulating the number of people located within a certain radius on the floor plan of an offline market for a certain period of time and generating a histogram.
[0079] In addition, the artificial intelligence agent (100) can obtain movement tracking information (○09) that tracks the movement of at least one person and object within a specific space by referring to the task results. At this time, the artificial intelligence agent (100) converts the location of a person detected in each image into the location of a specific space in an offline market and tracks it, and even when moving out of the field of view of the current camera and into the field of view of another camera, the agent can obtain movement tracking information that tracks the movement of a person within the offline market by recognizing and tracking the same person by referring to facial feature detection information, attribute analysis information, clothing analysis information, and safety equipment non-wearing detection information detected in the image.
[0080] In addition, the artificial intelligence agent (100) can obtain crime information (○10) that detects a crime that occurred within the specific space by referring to the task results. For example, the artificial intelligence agent (100) can obtain crime information, such as theft, that occurred within the specific space by referring to abnormal behavior judgment information detected in the video.
[0081] In addition, the artificial intelligence agent (100) can obtain abnormal behavior information (○11) that detects abnormal behavior of a person located within a specific space by referring to task results. For example, the artificial intelligence agent (100) can obtain abnormal behavior information related to a dangerous situation, such as an accident, that occurred within a specific space by referring to abnormal situation judgment information detected in an image.
[0082] In addition, the artificial intelligence agent (100) can obtain product inventory information (○12) by detecting the inventory of products placed within a specific space by referring to the task results. For example, the artificial intelligence agent (100) can obtain product inventory information by checking whether the displayed products detected in the image are continuously detected to check whether there is product inventory or by counting the number of displayed products.
[0083] In addition, the artificial intelligence agent (100) can also obtain group recognition information by analyzing the group of detected people by referencing the task results. At this time, the artificial intelligence agent (100) can analyze the group of detected people by tracking the movement path and analyzing the characteristics of the people, thereby creating a group of the group. Even if the grouped group temporarily disperses, the recognized group can be continuously tracked based on the characteristic of regrouping and moving together.
[0084] In addition, the artificial intelligence agent (100) can obtain employee recognition information by analyzing whether a person located within a specific space is an employee by referencing task results. At this time, the employee recognition information can be recognized by detecting characteristics such as employee uniforms through an image analysis deep learning model. However, if the employee is not wearing a uniform and cannot be distinguished from a visitor, the artificial intelligence agent (100) can recognize the detected person as an employee if the detected person stays in an area accessible only to employees for a certain period of time or a certain number of times based on the movement path of the detected person.
[0085] Additionally, the AI agent (100) can obtain lost item detection information by referencing task results to detect newly detected or missing objects within a specific space. That is, the AI agent (100) can detect missing or newly appeared objects by comparing them with previous video frames, or detect lost items that have been dropped or left behind in a specific location for an extended period of time.
[0086] In addition, the artificial intelligence agent (100) can obtain video summary information that organizes only the necessary parts according to preset conditions within the video recorded for a preset period of time by referring to the task results. That is, the artificial intelligence agent (100) can analyze the video stored for a long period of time and summarize and organize meaningful and interesting parts. For example, it can summarize only the parts in which people enter and exit within the video recorded for a preset period of time, or summarize the parts in which only specific people with specific characteristics such as gender and age enter and exit. This reduces the burden of the manager having to monitor the video for a long period of time and allows for quick video retrieval to induce accurate monitoring results.
[0087] However, the image analysis status information obtained in the present invention is not limited thereto, and may include various information related to management of offline markets.
[0088] Referring again to FIG. 3, the artificial intelligence agent (100) can vectorize text analysis status information including at least some of the POS (Point Of Sales) data and POG (planogram) data and store it in a vector database (S30).
[0089] That is, the artificial intelligence agent (100) can vectorize text analysis status information including at least a portion of sales information obtained in real time from the POS system and POG data, which is a product layout map within a specific space of an offline market, and store the vectorized information in a vector database.
[0090] Meanwhile, in the above, it has been described that image analysis status information and text analysis status information are sequentially acquired, but this is for convenience of explanation. In the present invention, image analysis status information and text analysis status information can be acquired in real time, vectorized, and stored in a vector database.
[0091] In addition, vectorization of image analysis status information is to convert each image analysis status information into a vector representing it. For example, human identification information can be vectorized by using coordinate information within a specific space where the person is located and assigned human identification information. Human attribute information can be vectorized by adding human attribute information, such as age and gender, to location information and identification information. Interest information can be vectorized by adding gaze direction information to human identification information and location information. Path tracking information can be vectorized by using human identification information, previous location information, and current location information.
[0092] Next, when a natural language command is acquired, the artificial intelligence agent (100) can generate (S40) a first to nth text command for performing the natural language command through an LLM (Large Language Model). At this time, n can be an integer greater than or equal to 1.
[0093] That is, when a natural language command related to management of an offline market is input, the artificial intelligence agent (100) can input the natural language command into the LLM and cause the LLM to generate a first text command to an nth text command for performing the natural language command.
[0094] For example, the artificial intelligence agent (100) can cause the LLM to generate various text commands to obtain information for performing the natural language command, such as generating text commands such as “yesterday’s sales information,” “today’s sales information,” “yesterday’s number of visitors,” and “today’s number of visitors” to obtain information to use in determining today’s sales status in response to the natural language command “What is the sales status today?”
[0095] In this case, the LLM may be generated by fine-tuning a pre-trained model using large-scale language data, using target learning data related to offline market management. For example, the target learning data may consist of various information related to offline market operation, such as basic knowledge and manuals required for offline market operation, sales calculation methods, product cost concepts, distribution, information on the format and meaning of sales data provided with time information, schemas for video analysis status information, and various examples of combining and analyzing these data.
[0096] Next, the artificial intelligence agent (100) can obtain first state information to n-th state information according to the first text command to n-th text command from the vector database through LLM (S50), and generate response information according to the first state information to n-th state information through LLM (S60).
[0097] That is, the artificial intelligence agent (100) may obtain at least some of the image analysis state information and text analysis state information corresponding to the k-th text command from the vector database as the k-th state information according to the k-th text command (ultimately by changing k from 1 to n), thereby obtaining the first state information to the n-th state information corresponding to the first text command to the n-th text command, and may generate response information corresponding to the natural language command by referring to the first state information to the n-th state information. At this time, k may be an integer greater than or equal to 1 and less than or equal to n.
[0098] At this time, the artificial intelligence agent may cause the LLM to obtain at least a portion of the image analysis status information and at least a portion of the POS data as the first to nth state information, and may generate at least one of the business status information and the business plan information related to the offline market as response information by referring to at least a portion of the image analysis status information and at least a portion of the POS data.
[0099] For example, in response to the question “What are the sales figures today?”, LLM can obtain “yesterday’s sales information”, “today’s sales information”, “number of visitors yesterday”, and “number of visitors today”, and then generate response information such as “the number of visitors today has increased compared to yesterday, but today’s sales have decreased compared to yesterday.”
[0100] And, in response to the natural language command, “What is the reason for the decrease in sales today?”, LLM can obtain and analyze “interest in a specific product”, “sales for a specific product”, etc., and then generate response information such as “There were many visitors who showed interest in a specific product, but few visitors actually purchased a specific product.”
[0101] In addition, in response to the natural language command, “What are some ways to increase sales for a specific product?”, LLM can obtain and analyze “detailed information about a specific product,” “attribute information and group information of visitors who showed interest in a specific product,” and “attribute information and group information of visitors who purchased a specific product,” and then generate response information such as “A specific product is a product with a large capacity. Everyone showed interest in a specific product, but visitors who actually purchased the specific product were families, and individual visitors who showed interest did not purchase it due to the product’s capacity. Therefore, it is determined that it would be better to configure the specific product with a small capacity.”
[0102] Meanwhile, while one offline market has been described as an example above, the artificial intelligence agent (100) can cause the LLM to generate at least one of business status information and business plan information as response information by referencing multiple pieces of image analysis status information and multiple POS data corresponding to two or more offline markets operated in different locations.
[0103] For example, the artificial intelligence agent (100) can cause the LLM to generate response information so that a product with high sales in one offline market can be sold in another offline market if the product is not sold in another offline market.
[0104] In addition, the artificial intelligence agent may cause the LLM to obtain at least some of the image analysis state information, at least some of the POS data, and at least some of the POG data as the first state information to the n-th state information, and generate POG update information for updating the POG as response information by referring to at least some of the image analysis state information, at least some of the POS data, and at least some of the POG data.
[0105] That is, the artificial intelligence agent (100) may cause the LLM to analyze a purchase trend for a specific product at a specific location by referring to at least some of the following information: movement path information for moving to a specific location within a specific space, exposure viewing angle information for a specific location within a specific space, movement information for people moving to a specific location within a specific space, density of people at a specific location, interest in a specific location, time information for people staying within a preset radius from a specific location, identifiers of products displayed by floor of a shelf at the specific location, identifiers of other products displayed near the displayed products, information on picking up a specific product displayed at a specific location, sales information for a specific product according to POS data, and attribute information for a person who purchased a specific product, in relation to a specific location in a POG, and generate POG update information that determines whether a specific product displayed at a specific location has changed.
[0106] For example, the artificial intelligence agent (100) can cause the LLM to generate POG update information to display products with a high purchase rate among family visitors in a specific area of an offline market, when family visitors spend a lot of time in a specific area of an offline market but have low purchase rates for specific products displayed in the specific area.
[0107] As another example, the artificial intelligence agent (100) may cause the LLM to generate POG update information to update the POG so that the movement path to the specific location in the offline market becomes easier, when there is a high level of interest and purchase rate for a product displayed at a specific location with a good viewing angle in the entire space of the offline market, but the movement path to reach the specific location is not easy, and many visitors give up on purchasing a specific product while moving to the specific location and move to another location.
[0108] In addition, the artificial intelligence agent (100) can cause the LLM to obtain at least some of the image analysis status information as the first status information to the n-th status information, and generate market status information related to the status of the offline market as response information by referring to at least some of the image analysis status information.
[0109] For example, the artificial intelligence agent (100) can cause the LLM to analyze the attribute information, interest information, and waiting time information of visitors located in a specific area with high traffic density, and generate response information by identifying the cause of the high traffic density in the specific area. In other words, the agent can analyze whether the high traffic density in a specific area is due to waiting time or the high traffic density is due to the purchase of a specific product, and generate response information.
[0110] Meanwhile, the artificial intelligence agent (100) monitors the video analysis status information, and may add at least one of the following to the video displayed on the administrator terminal: entry / exit counting information that counts the number of people who have passed through a specific area within a specific space; dangerous area intrusion information that detects a person who has intruded into a dangerous area within a specific space; waiting time information that measures the waiting time to use a specific service within a specific space; parking lot information within a specific space; heat map information that estimates the density of people detected within a specific space; abnormal behavior information that detects abnormal behavior of people located within a specific space; crime information that detects a crime that has occurred within a specific space; and product inventory information that detects the inventory of products placed within a specific space. Through this, the administrator can easily visually check information related to the offline market.
[0111] In addition, when request information for searching the status of an offline market according to specific conditions is obtained, the artificial intelligence agent (100) can cause the LLM to obtain at least one specific image analysis information corresponding to the specific conditions from a vector database, and output the status information of the offline market according to the specific conditions as response information by referring to the obtained specific image analysis information.
[0112] For example, in response to a search request for the movement path of a specific visitor, the artificial intelligence agent (100) may cause the LLM to track the movement path of the specific visitor and output, as response information, movement path tracking information from the time the specific visitor entered the offline market to the time he or she left the market.
[0113] As another example, in response to a search request for the current location of a specific visitor at a specific location at a specific time, the artificial intelligence agent (100) may cause the LLM to identify a specific visitor at a specific location at a specific time, and then track the movement of the specific visitor to output response information indicating the current location of the specific visitor.
[0114] As another example, in response to a search request for the top five products that had high interest from visitors during a specific time period, the artificial intelligence agent (100) can cause the LLM to obtain the interests of visitors during a specific time period, analyze the information, and output a list of five products with high interest as response information.
[0115] The embodiments of the present invention described above may be implemented in the form of program commands that can be executed through various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program commands, data files, data structures, etc., either singly or in combination. The program commands recorded on the computer-readable recording medium may be those specially designed and configured for the present invention or may be those known and available to those skilled in the art of computer software. Examples of the computer-readable recording medium include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program commands, such as ROMs, RAMs, and flash memories. Examples of the program commands include not only machine language codes generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc. The hardware devices may be configured to operate as one or more software modules to perform processing according to the present invention, and vice versa.
[0116] Although the present invention has been described above with specific details such as specific components and limited examples and drawings, these are provided only to help a more general understanding of the present invention, and the present invention is not limited to the above examples, and those with ordinary knowledge in the technical field to which the present invention pertains can make various modifications and variations from this description.
[0117] Therefore, the idea of the present invention should not be limited to the embodiments described above, and all things that are modified equally or equivalently to the following claims as well as the claims are considered to fall within the scope of the idea of the present invention.
Claims
1. A method for managing an offline market using an artificial intelligence agent, (a) when at least one video stream is transmitted from at least one camera filming a specific space of at least one offline market, an artificial intelligence agent inputs at least one video belonging to the video stream into at least one video analysis deep learning model, causes the video analysis deep learning model to perform each task in the video through video analysis and output task results for each task, acquires video analysis status information of the offline market by referring to the task results, vectorizes the video analysis status information and stores it in a vector database, and vectorizes text analysis status information including at least a part of POS (Point Of Sales) data and POG (planogram) data corresponding to the offline market and stores it in the vector database; and (b) a step of causing the artificial intelligence agent, when a natural language command related to the management of the offline market is input, to cause the artificial intelligence agent to input the natural language command into an LLM (Large Language Model) to cause the LLM to generate a first text command to an n-th text command for performing the natural language command, wherein n is an integer greater than or equal to 1, and, according to the k-th text command, wherein k is an integer greater than or equal to 1 and less than or equal to n, to obtain at least some of the image analysis state information and the text analysis state information corresponding to the k-th text command from the vector database as k-th state information, thereby obtaining first state information to n-th state information corresponding to the first text command to the n-th text command, and generating response information corresponding to the natural language command with reference to the first state information to the n-th state information; A method including:
2. In paragraph 1, In step (a) above, The artificial intelligence agent analyzes the image using the image analysis deep learning model to detect object detection information that detects at least one of a person and an object in the image, face and head detection information that detects at least one of a person's face and head in the image, posture and behavior analysis information that extracts key points of a person detected in the image and analyzes the person's posture and behavior, facial expression analysis information that estimates an expression of a face detected in the image, facial feature detection information that extracts feature points of a face detected in the image, attribute analysis information that estimates at least one of a gender and an age of a person detected in the image, clothing analysis information that analyzes the clothing of a person detected in the image, safety equipment non-wearing detection information that tracks whether a person detected in the image is wearing safety equipment, interest measurement information that analyzes where a person detected in the image is interested in based on the location and head direction of a person detected in the image, crowd count measurement information that estimates the number of people located in a specific area in the image, abnormal behavior judgment information that estimates abnormal behavior of a person detected in the image, and current situation information that estimates a situation of a person detected in the image. A method for outputting task results including at least a portion of the above situation judgment information.
3. In paragraph 1, In step (a) above, The artificial intelligence agent, with reference to the task results, identifies each person located in the specific space, counts the number of people passing through a specific area in the specific space, counts the number of people, counts the number of people passing through, counts the number of people ... A method for obtaining image analysis status information including at least one of image summary information that organizes only the necessary parts according to preset conditions and product inventory information that detects the inventory of products placed within the specific space.
4. In paragraph 1, In step (b) above, A method in which the artificial intelligence agent causes the LLM to obtain at least a portion of the image analysis status information and at least a portion of the POS data as the first state information to the nth state information, and generates at least one of the business status information and the business plan information related to the business of the offline market as the response information by referring to at least a portion of the image analysis status information and at least a portion of the POS data.
5. In paragraph 4, A method in which the artificial intelligence agent causes the LLM to generate at least one of the business status information and the business plan information as the response information by referencing a plurality of pieces of image analysis status information and a plurality of POS data corresponding to two or more offline markets operated in different locations.
6. In paragraph 1, In step (b) above, A method in which the artificial intelligence agent causes the LLM to obtain at least some of the image analysis state information, at least some of the POS data, and at least some of the POG data as the first state information to the n-th state information, and generate POG update information for updating the POG by referring to at least some of the image analysis state information, at least some of the POS data, and at least some of the POG data as the response information.
7. In paragraph 6, The artificial intelligence agent causes the LLM to analyze a purchase trend for the specific product at the specific location by referring to at least some of the following information: movement path information for moving to the specific location within the specific space, exposure viewing angle information for the specific location within the specific space, movement line information of people moving to the specific location within the specific space, density of people at the specific location, interest in the specific location, time information for which people stay within a preset radius from the specific location, identifiers of products displayed by floor of a shelf at the specific location, identifiers of other products displayed near the displayed products, information on picking up the specific product displayed at the specific location, sales information of the specific product according to the POS data, and attribute information of people who purchased the specific product, in relation to a specific location in the POG, and generates the POG update information that determines whether the specific product displayed at the specific location has changed.
8. In paragraph 1, In step (b) above, A method in which the artificial intelligence agent causes the LLM to obtain at least some of the image analysis status information as the first state information to the nth state information, and generate market status information related to the status of the offline market as the response information by referring to at least some of the image analysis status information.
9. In paragraph 1, The above LLM is a method of creating a pre-trained model using large-scale language data by fine-tuning it using target learning data related to the management of the offline market.
10. In paragraph 1, After step (a) above, A step for causing the artificial intelligence agent to monitor the image analysis status information, and display at least one of the following information: entry / exit counting information counting the number of people who have passed through a specific area within the specific space, dangerous area intrusion information detecting a person who has intruded into a dangerous area located within the specific space, waiting time information measuring a waiting time for using a specific service within the specific space, parking lot information within the specific space, heat map information estimating the density of people detected within the specific space, abnormal behavior information detecting abnormal behavior of people located within the specific space, crime information detecting a crime that has occurred within the specific space, and product inventory information detecting the inventory of products placed within the specific space; How to include more.
11. In paragraph 1, After step (a) above, When status search request information of the offline market according to a specific condition is obtained, the step of causing the artificial intelligence agent to cause the LLM to obtain at least one specific image analysis information corresponding to the specific condition from the vector database, and to output status information of the offline market according to the specific condition as the response information by referring to the obtained specific image analysis information; How to include more.
12. For artificial intelligence agents that manage offline markets, Memory storing instructions for managing the offline market; and A processor that performs operations for managing the offline market according to the instructions stored in the memory; Including, The processor comprises: (I) a process for inputting at least one video belonging to the video streaming into at least one video analysis deep learning model when at least one video streaming is transmitted from at least one camera filming a specific space of at least one offline market, causing the video analysis deep learning model to perform each task on the video through video analysis and output task results for each task, obtaining video analysis status information of the offline market by referring to the task results, vectorizing the video analysis status information and storing it in a vector database, and vectorizing text analysis status information including at least a part of POS (Point Of Sales) data and POG (planogram) data corresponding to the offline market and storing it in the vector database, and (II) a process for inputting the natural language command related to the management of the offline market into an LLM (Large Language Model) to cause the LLM to generate first to n-th text commands for performing the natural language command, wherein n is an integer greater than or equal to 1, and a k-th text command, wherein k is greater than or equal to 1. An artificial intelligence agent that performs a process of acquiring first state information to n-th state information corresponding to the first text command to the n-th text command by acquiring at least some of the image analysis state information and the text analysis state information corresponding to the k-th text command from the vector database as k-th state information according to which n is an integer or less, and generating response information corresponding to the natural language command by referring to the first state information to the n-th state information.
13. In paragraph 12, The processor, in the process (I), causes the image analysis deep learning model to analyze the image, and detects object detection information that detects at least one of a person and an object in the image, face and head detection information that detects at least one of a face and a head of a person in the image, posture and behavior analysis information that extracts key points of a person detected in the image and analyzes the person's posture and behavior, facial expression analysis information that estimates an expression of a face detected in the image, facial feature detection information that extracts feature points of a face detected in the image, attribute analysis information that estimates at least one of a gender and an age of a person detected in the image, clothing analysis information that analyzes the clothing of a person detected in the image, safety equipment non-wearing detection information that tracks whether a person detected in the image is wearing safety equipment, interest measurement information that analyzes where a person detected in the image is interested in based on a location and a head direction of a person detected in the image, crowd count measurement information that estimates the number of people located in a specific area in the image, abnormal behavior judgment information that estimates an abnormal behavior of a person detected in the image, and An artificial intelligence agent that outputs task results including at least a portion of the abnormal situation judgment information that estimates the current situation of the detected person.
14. In paragraph 12, The processor, in the (I) process, refers to the task results, and identifies each person located in the specific space, access counting information counting the number of people who passed through a specific area in the specific space, person attribute information for each person located in the specific space, dangerous area intrusion information detecting a person who has intruded into a dangerous area located in the specific space, waiting time information measuring a waiting time for using a specific service in the specific space, interest information of people located in the specific space, parking lot information in the specific space, path tracking information tracking the path of at least one of a person and an object in the specific space, heat map information estimating the density of people detected in the specific space, group recognition information analyzing a group of people detected in the specific space, employee recognition information analyzing whether a person located in the specific space is an employee, lost item detection information detecting a newly detected or missing object in the specific space, abnormal behavior information detecting an abnormal behavior of a person located in the specific space, crime information detecting a crime that occurred in the specific space, and a preset time. An artificial intelligence agent that obtains video analysis status information including at least one of video summary information that organizes only the necessary parts according to preset conditions from a video recorded during the recording, and product inventory information that detects the inventory of products placed within the specific space.
15. In paragraph 12, The processor, in the (II) process, causes the LLM to obtain at least a portion of the image analysis status information and at least a portion of the POS data as the first status information to the nth status information, and generate at least one of the business status information and the business plan information related to the business of the offline market as the response information by referring to at least a portion of the image analysis status information and at least a portion of the POS data.
16. In paragraph 15, The above processor is an artificial intelligence agent that causes the LLM to generate at least one of the business status information and the business plan information as the response information by referencing a plurality of pieces of image analysis status information and a plurality of POS data corresponding to two or more offline markets operated in different locations.
17. In paragraph 12, The processor, in the (II) process, causes the LLM to obtain at least some of the image analysis state information, at least some of the POS data, and at least some of the POG data as the first state information to the nth state information, and generates POG update information for updating the POG as the response information by referring to at least some of the image analysis state information, at least some of the POS data, and at least some of the POG data.
18. In paragraph 17, The processor is an artificial intelligence agent that causes the LLM to analyze a purchase trend for the specific product at the specific location by referencing at least some of, with respect to a specific location in the POG, movement path information for moving to the specific location within the specific space, exposure viewing angle information for the specific location within the specific space, movement line information of people moving to the specific location within the specific space, density of people at the specific location, interest in the specific location, time information of people staying within a preset radius from the specific location, identifiers of products displayed by floor of a shelf at the specific location, identifiers of other products displayed near the displayed products, information on picking up the specific product displayed at the specific location, sales information of the specific product according to the POS data, and attribute information of people who purchased the specific product, and generates the POG update information that determines whether the specific product displayed at the specific location has changed in response to the analyzed purchase trend.
19. In paragraph 12, The above processor is an artificial intelligence agent that causes the LLM, in the (II) process, to obtain at least some of the image analysis status information as the first status information to the n-th status information, and to generate market status information related to the status of the offline market as the response information by referring to at least some of the image analysis status information.
20. In paragraph 12, The above LLM is an artificial intelligence agent created by fine-tuning a pre-trained model using large-scale language data and target learning data related to management of the offline market.
21. In paragraph 12, The above processor, after the (I) process, monitors the image analysis status information, and further performs a process of adding and displaying at least one of the following information to the image displayed on the administrator terminal: entry / exit counting information counting the number of people who have passed through a specific area within the specific space; dangerous area intrusion information detecting a person who has intruded into a dangerous area located within the specific space; waiting time information measuring a waiting time for using a specific service within the specific space; parking lot information within the specific space; heat map information estimating the density of people detected within the specific space; abnormal behavior information detecting abnormal behavior of people located within the specific space; crime information detecting a crime that has occurred within the specific space; and product inventory information detecting the inventory of products placed within the specific space.
22. In paragraph 12, The above processor is an artificial intelligence agent that, after the (I) process, if status search request information of the offline market according to a specific condition is obtained, causes the LLM to obtain at least one specific image analysis information corresponding to the specific condition from the vector database, and further performs a process of outputting status information of the offline market according to the specific condition as the response information by referencing the obtained specific image analysis information.
Citation Information
Patent Citations
Method and system for automatically measuring retail store display compliance
JP2008537226A
Method for producing instant grain noodle capable of cooking by cold water
KR1020200131969A
Raid system with fault resilient storage devices
KR1020220008214A
Method, program, and apparatus for monitoring behaviors based on artificial intelligence
KR102599020B1
KR20230025714A
Cited By
Systems and methods for end-to-end modular intelligence
US12572877B2