Simulation data generation device, simulation data generation method, and simulation data generation system
The simulation data generation system using a large-scale language model addresses inefficiencies in metadata generation by providing interactive and efficient data creation, resulting in cost-effective and realistic simulation data for imaging system applications.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SONY SEMICON SOLUTIONS CORP
- Filing Date
- 2025-09-30
- Publication Date
- 2026-05-07
AI Technical Summary
Existing methods for generating metadata for application development in imaging systems are costly, time-consuming, and often produce unnatural results due to the need for physical setups and three-dimensional model programming, which are inefficient and impractical.
A simulation data generation system utilizing a large-scale language model to generate and modify time-series simulation data, allowing for interactive input and visualization of metadata through prompts, with the aid of a user database to enhance naturalness.
Enables efficient and natural simulation data generation, reducing costs and time while ensuring accurate and realistic metadata for application development.
Smart Images

Figure JP2025034645_07052026_PF_FP_ABST
Abstract
Description
SIMULATION DATA GENERATION DEVICE, SIMULATION DATA GENERATION METHOD, AND SIMULATION DATA GENERATION SYSTEMCROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of Japanese Priority Patent Application JP 2024-190552 filed on October 30, 2024, the entire contents of which are incorporated herein by reference.
[0002] The present technology relates to a simulation data generation device, a simulation data generation method, and a simulation data generation system related to generation of simulation data simulating metadata output from an imaging device.
[0003] Systems that analyze an event that has occurred in a space to be sensed using a monitoring camera or the like are known. Some of these systems use metadata output from an imaging device as a monitoring camera (for example, see PTL 1 below).
[0004] JP 2010-273125 ASummary
[0005] For development of an application used in such a system, appropriate metadata is necessary. In the metadata for development, for example, an imaging device is actually arranged in a specific space, and an actor performs a predetermined motion within an angle of view of the imaging device, whereby the metadata is output from the imaging device. Development of the application is advanced so that it can be appropriately estimated that a movement to be detected has been performed on the basis of the metadata. However, the method of actually placing the imaging device in the specific space and capturing an image has a problem that it takes too much time and cost. Furthermore, an acting performance of the actor may be different from a motion of a normal person who is not conscious of the imaging device, and there is a high possibility that unnatural metadata is output.
[0006] Furthermore, it is also conceivable to arrange a person of a three-dimensional model in a virtual space, program the three-dimensional model to perform a predetermined motion, further arrange a virtual imaging device in the virtual space, generate an image estimated to be obtained within the angle of view, and output simulation data simulating metadata. This method has an advantage that the cost for employing an actor and the time cost in capturing an image can be reduced, but has a problem that it takes a lot of time to create and program the three-dimensional model.
[0007] The present technology has been made in view of such problems, and it is desirable to generate natural simulation data that is data simulating metadata.
[0008] A simulation data generation device comprising: processing circuitry configured to input a first prompt to a large-scale language model, wherein the first prompt includes instructions for generation of time-series simulation data simulating metadata of a captured image output from an imaging device, wherein the time-series data enables confirmation of motion of a subject in the captured image through visualization, obtain the time-series simulation data as answer information from the large-scale language model, input a second prompt to the large-scale language model instructing modification of the answer information, obtain modified time-series simulation data as answer information from the large-scale language model, and output the modified time-series simulation data, the output being a visualized presentation to a user.
[0009] A simulation data generation method executed by an information processing apparatus, the method comprising inputting a first prompt to a large-scale language model, wherein the first prompt includes instructions for generation of time-series simulation data simulating metadata of a captured image output from an imaging device, wherein the time-series data enables confirmation of motion of a subject in the captured image through visualization; obtaining the time-series simulation data as answer information from the large-scale language model; inputting a second prompt to the large-scale language model instructing modification of the answer information; obtaining modified time-series simulation data as answer information from the large-scale language model; and outputting the modified time-series simulation data, the output being a visualized presentation to a user.
[0010] A non-transitory computer-readable medium storing instructions that, when executed by processing circuitry, cause the processing circuitry to perform a method, the method comprising inputting a first prompt to a large-scale language model, wherein the first prompt includes instructions for generation of time-series simulation data simulating metadata of a captured image output from an imaging device, wherein the time-series data enables confirmation of motion of a subject in the captured image through visualization; obtaining the time-series simulation data as answer information from the large-scale language model; inputting a second prompt to the large-scale language model instructing modification of the answer information; obtaining modified time-series simulation data as answer information from the large-scale language model; and outputting the modified time-series simulation data, the output being a visualized presentation to a user.
[0011] Fig. 1 is a block diagram illustrating a schematic configuration of a development system including a simulation data generation system according to the present embodiment.Fig. 2 is a block diagram illustrating a hardware configuration of a computer device such as an information processing apparatus.Fig. 3 is a block diagram illustrating a functional configuration of a simulation data generation device.Fig. 4 is a diagram for describing a rough flow of processing executed by a developer terminal and the simulation data generation device.Fig. 5 is a diagram illustrating an example of an interaction screen, and is a diagram of an initial screen to be presented to a developer.Fig. 6 is a diagram illustrating an example of the interaction screen, and is a diagram of the screen in a state where a first prompt is input by the developer.Fig. 7 is a diagram for describing a specific flow of auxiliary information setting processing.Fig. 8 is a diagram illustrating an example of visualized simulation data.Fig. 9 is a diagram illustrating another example of the visualized simulation data.Fig. 10 is a diagram illustrating still another example of the visualized simulation data.Fig. 11 is a diagram illustrating an example of the interaction screen, and is a diagram of the screen in a state where answer information obtained from a large-scale language model M by the developer is presented to the developer.Fig. 12 is a diagram illustrating an example of the interaction screen, and is a diagram of the screen in a state where a second prompt is input by the developer.Fig. 13 is a schematic explanatory diagram illustrating an example of exchange between the simulation data generation device and the large-scale language model.Fig. 14 is a diagram illustrating a state in which corrected simulation data is visualized.Fig. 15 is a diagram illustrating an example of the interaction screen, and is a diagram of the screen in a state where the first prompt and a third prompt are input by the developer.Fig. 16 is a diagram illustrating an example of a display mode of a bounding box related to a subject that has performed a specific action.
[0012] Hereinafter, embodiments of a system according to the present technology will be described in the following order with reference to the accompanying drawings. <1. Configuration of Development System> <2. Hardware Configuration of Each Device> <3. Function of Simulation Data Generation Device> <4. Flow of Processing> <5. Modification> <6. Variations> <6-1. Specific Action> <6-2. Example of Interactive Prompt> <6-3. Effective Use Example of User Database> <6-4. Setting of Information Not Output as Answer Information> <6-5. Simulation Data Suitable for Application Development> <6-6. Simulation Data Including Unsteady but Natural Changes> <6-7. Varied Simulation Data> <6-8. Simulation Data Generated over Long Period of Time> <7. Implementation of Functions> <8. Summary> <9. Present Technology>
[0013] <1. Configuration of Development System> A configuration of a development system S1 of the present technology will be described with reference to Fig. 1.
[0014] The development system S1 is a system used by a developer to develop an application. The development system S1 includes a simulation data generation system S2 that is a system for generating simulation data MD used for developing an application, and a developer terminal 1 used by the developer.
[0015] The simulation data generation system S2 includes a simulation data generation device 2, and a first server device 3 and a second server device 4 that are server devices used in a case where the simulation data generation device 2 generates the simulation data MD.
[0016] The developer terminal 1, the simulation data generation device 2, the first server device 3, and the second server device 4 included in the development system S1 are mutually connected in a communicable manner by a communication network NW.
[0017] The simulation data MD generated by the simulation data generation system S2 will be described. Examples of the application to be developed include an application that analyzes data output from an imaging device such as a monitoring camera and issues an alert or notification according to an analysis result, and an application that visualizes the analysis result.
[0018] Some of such applications use not only image data output from the imaging device but also metadata related to the image data as an analysis target. Development of these applications needs the metadata output from the imaging device.
[0019] However, it may be difficult to prepare appropriate metadata output from the imaging device in a development stage of the application. For example, the imaging device is installed in a space SP to be sensed of the imaging device, and a state in which a person gives a performance such as shopping within an angle of view of the imaging device is captured. Thereby, the image data and the metadata can be actually output from the imaging device.
[0020] However, in this method, there are various costs such as a financial cost and a time and effort for actually preparing the space SP and the imaging device, a time and effort for securing the person giving a performance in front of the imaging device, and a time and effort for actually capturing images for ten hours in a case where data for ten hours is desired. Furthermore, in a case where the space SP to be sensed is the space SP in a store, building of the store itself may not be completed.
[0021] Furthermore, in a case where an actor who performs a predetermined motion in front of the imaging device is employed, the financial cost also occurs.
[0022] Moreover, there is also a problem that a deviation occurs between a motion of a person in a case where the person performs a predetermined motion in front of the imaging device and a motion performed by a general person who is not conscious of the imaging device in order to achieve his / her purpose, and natural and appropriate metadata may not be able to be obtained.
[0023] Furthermore, as another method, it is also conceivable to arrange a virtual imaging device and a three-dimensional model of a person in a virtual space, program a motion of the three-dimensional model, generate image data obtained at the angle of view of the imaging device using a technology such as ray tracing, and generate metadata including an analysis result of the image data and the like.
[0024] However, in this method, first, it is difficult to cause the three-dimensional model of the person to perform a natural motion by programming, and the motion of the three-dimensional model is also limited only to a motion that can be conceived by the person.
[0025] Furthermore, even if a video is created using generative artificial intelligence (AI), quality of a still image has been improved, but there are still many unnatural points regarding a video for expressing a motion, and it is not possible to generate the metadata obtained by capturing a natural motion of a person.
[0026] The simulation data generation system S2 is a system in view of these problems. The simulation data generation system S2 generates the simulation data MD directly simulating the metadata without generating the image data by using a large-scale language model (large language model (LLM)) M stored in the first server device 3.
[0027] As is well known, the large-scale language model M is an artificial intelligence (AI) model that generates response information according to content of an input prompt and is capable of executing various natural language processing tasks. For example, in a case where a question sentence is input as a prompt, the large-scale language model M generates answer information to the question by performing data search processing according to the question content, sentence generation processing according to a search result, and the like. Furthermore, in addition to such a task of generating answer information to the question, the large-scale language model M can execute various tasks of generating the response information according to the content of an input prompt, such as a task of creating a computer program that realizes processing specified by the prompt, and a task of generating an image or music that satisfies a condition specified by the prompt, for example.
[0028] Examples of the large-scale language model M that can be used by the simulation data generation device 2 in the present embodiment include the following.
[0029] -Generative pre-trained transformer (GPT) -Bidirectional encoder representations from transformers (BERT) -Text-to-text transfer transformer (T5) -XLNet -Enhanced representation through knowledge integration (ERNIE) -Efficiently learning an encoder that classifies token replacements accurately (ELECTRA) -BigScience large open-science open-access multilingual language model (BLOOM) -Mistral -Open pre-trained transformer (OPT) -Topher -Language model for dialogue applications (LaMDA) -Turing natural language generation (Turing-NLG) -Pathways language mode (PaLM) -Language large models meta AI (Llama)
[0030] Note that a new large-scale language model M appears every day, and the large-scale language model M available in the present embodiment is not limited to these models.
[0031] Furthermore, the simulation data generation system S2 generates the simulation data MD based on a more natural motion of a person by using a user database that is a database on a user side provided in the second server device 4. The user database may be a relational database (RDB) or another database. Note that, in the following description, a vector database VD is given as an example of the user database.
[0032] The vector database VD is stored with a vector attached to each piece of information. In other words, the vector database VD is a database that stores each piece of information in a vector format. Note that each piece of information stored in the vector database VD is, for example, a text itself.
[0033] The vector stored in the vector database VD is information for expressing a meaning or a concept of target information, and has several hundred to several thousand dimensions, for example. That is, the vector database VD can also be said to be a database expressing a similarity of the meanings or the concepts of the pieces of information.
[0034] The simulation data generation device 2 presents an input screen, a presentation screen, and the like to the developer terminal 1. The developer who uses the developer terminal 1 can perform interactive exchange with the large-scale language model M included in the first server device 3 via the simulation data generation device 2 by viewing the screen or performing an input operation on the screen.
[0035] The simulation data generation device 2 inputs the developer's input as a prompt to the large-scale language model M to obtain the response information. Furthermore, the simulation data generation device 2 uses the vector database VD included in the second server device 4 in order to make the response information more appropriate.
[0036] The simulation data generation device 2 visualizes time-series data of the simulation data MD obtained from the large-scale language model M as the response information and presents the time-series data to the developer terminal 1. Specific examples thereof will be described below.
[0037] <2. Hardware Configuration of Each Device> Fig. 2 illustrates a hardware configuration example of the developer terminal 1, the simulation data generation device 2, the first server device 3, and the second server device 4. Note that, here, in a case where the developer terminal 1, the simulation data generation device 2, the first server device 3, and the second server device 4 are not distinguished, they are referred to as computer devices Com.
[0038] The computer device Com includes a processing circuit 51, a read only memory (ROM) 52, and a random access memory (RAM) 53.
[0039] The processing circuit 51, the ROM 52, and the RAM 53 can perform data communication with each other via a bus 54.
[0040] An input / output interface (I / F) 55 is further connected to the bus 54.
[0041] An input device 56, an output device 57, a storage unit 58, a communication interface 59, and a drive 60 are connected to the input / output interface 55.
[0042] The processing circuit 51 is, for example, a circuit that performs various operations, such as a central processing unit (CPU) or a graphics processing unit (GPU). Note that the processing circuit 51 may include a plurality of circuits such as both the CPU and the GPU.
[0043] The processing circuit 51 includes an arithmetic circuit that performs an arithmetic operation, a control circuit that controls the arithmetic circuit, a storage circuit such as a register and a cache, and an internal bus used as a data transmission path.
[0044] The processing circuit 51 executes various types of processing according to a program stored in the ROM 52 or a program loaded from the storage unit 58 to the RAM 53. The RAM 53 appropriately stores data and the like necessary for the processing circuit 51 to execute various types of processing.
[0045] The input device 56 is assumed to be, for example, a pointing device 56a such as a mouse, a keyboard 56b, a camera 56c, a microphone 56d, or various operators and operation devices such as a key, a dial, a touch panel, a touch pad, and a remote controller. The input device 56 detects a user's operation and transmits a signal corresponding to the input operation to the processing circuit 51. The processing circuit 51 interprets the signal and executes corresponding processing.
[0046] The output device 57 is, for example, a speaker 57a, a display device 57b including a liquid crystal display (LCD), an organic electro-luminescence (EL) panel, or the like.
[0047] The display device 57b is used for displaying various types of information, and includes, for example, a device provided in a housing of the computer device Com, a separate device connected to the computer device Com, or the like.
[0048] The display device 57b executes display of an image for various types of image processing, a video to be processed, and the like on a display screen on the basis of an instruction from the processing circuit 51. In addition, the display device 57b displays various operation menus, icons, messages, and the like, that is, performs display as a graphical user interface (GUI), on the basis of an instruction from the processing circuit 51.
[0049] For example, an interaction screen G1, a confirmation screen G2, and the like to be described below are displayed on the display device 57b of the developer terminal 1.
[0050] The storage unit 58 includes a hard disc drive (HDD), a solid-state memory, and the like. Note that the large-scale language model M is stored in the storage unit 58 of the first server device 3 illustrated in Fig. 1. Furthermore, the storage unit 58 of the second server device 4 functions as the vector database VD in a case where various types of information are stored together with vectors. The large-scale language model M and the vector database VD are indicated by the broken lines in Fig. 2 because it is sufficient that some of the computer devices Com include the large-scale language model M and the vector database VD. Note that each of the computer devices Com does not need to include all the other components illustrated in Fig. 2.
[0051] The communication interface 59 is a transmission unit or a reception unit related to wired or wireless communication performed with another computer device Com. The communication interface 59 includes an interface circuit for performing communication using various communication standards such as a universal serial bus (USB), Bluetooth (registered trademark), and Wi-Fi (registered trademark). Furthermore, the communication interface 59 may be configured as a circuit included in a network interface card (NIC), a modem, a Wi-Fi adapter, or the like.
[0052] The drive 60 is a device to which a removable recording medium 61 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory is appropriately mounted.
[0053] The drive 60 can read a data file such as a program used for each processing from the mounted removable recording medium 61. The read data file is stored in the storage unit 58, an image included in the data file is displayed on the display device 57b, or a sound is output from the speaker 57a. Furthermore, a computer program or the like read from the removable recording medium 61 by the drive 60 is installed in the storage unit 58 as necessary.
[0054] In the computer device Com having the hardware configuration as described above, for example, software for the processing of the present embodiment can be installed via network communication by the communication interface 59 or the removable recording medium 61. Alternatively, the software may be stored in advance in the ROM 52, the storage unit 58, or the like. In the computer device Com, the processing circuit 51 performs processing operation on the basis of various programs, thereby executing information processing and communication processing necessary as the developer terminal 1, the simulation data generation device 2, the first server device 3, or the second server device 4.
[0055] Note that the computer device Com such as the developer terminal 1 or the simulation data generation device 2 is not limited to a single device as illustrated in Fig. 3, and may be configured by systematizing a plurality of computer devices. The plurality of computer devices may be systematized by a local area network (LAN) or the like, or may be arranged in remote places by a virtual private network (VPN) or the like using the Internet or the like. The plurality of computer devices may include a computer device as a server group (cloud) that can be used by a cloud computing service.
[0056] <3. Function of Simulation Data Generation Device> Each function implemented by the processing circuit 51 of the simulation data generation device 2 executing a program will be described with reference to Fig. 3.
[0057] The processing circuit 51 functions as a prompt acquisition unit F1, an answer information acquisition unit F2, a presentation processing unit F3, an information provision processing unit F4, and an auxiliary information acquisition unit F5 by executing a predetermined program.
[0058] The prompt acquisition unit F1 performs processing of acquiring a prompt to be input to the large-scale language model M.
[0059] Specifically, the prompt acquisition unit F1 acquires, as a first prompt P1, a sentence that is input by the developer and requests generation of the simulation data MD. Furthermore, the prompt acquisition unit F1 acquires, as a second prompt P2, a sentence that is input by the developer and requests correction or revision of the generated simulation data MD.
[0060] Moreover, the prompt acquisition unit F1 acquires, as a third prompt P3, a prompt for making the simulation data MD obtained as answer information of the large-scale language model M more natural and for instructing setting of information that is not output as the answer information. In the following description, the information set by the input of the third prompt P3 is referred to as an “internal parameter”.
[0061] The third prompt P3 may be acquired by acquiring a sentence input by the developer, or may be acquired by acquiring an instruction sentence generated by the simulation data generation device 2.
[0062] Here, an example of each prompt is specifically described. Note that, here, an example will be given in which the space SP where an imaging device is disposed is the space SP to be sensed and is the space SP of a selling space of a retail store, a subject to be detected is a person such as a customer to the store or a clerk, and the simulation data MD of the metadata output from the imaging device is information including coordinates of a bounding box BB related to the detected person.
[0063] The first prompt P1 is, for example, a sentence such as “Create metadata output from the imaging device that is arranged in the retail store and detects a person. The metadata includes coordinate information of the bounding box of the detected person.”
[0064] The second prompt P2 is a sentence such as “The number of people seems to be small. Please re-create the metadata with the maximum of five people in the store.”
[0065] The third prompt P3 is a sentence giving an instruction to set information other than the coordinates of the bounding box BB output as metadata. The third prompt P3 is, for example, a sentence giving an instruction to set information regarding the subject to be detected, set a personal space of the subject, or set information regarding the space SP to be sensed. The information set here can be said to be information regarding the captured image captured by the imaging device.
[0066] Note that the setting of information that is not output as the answer information performed in response to the input of the third prompt P3 is not necessarily performed. That is, the developer may acquire the desired simulation data MD by inputting only the first prompt P1 and the second prompt P2.
[0067] Note that the “personal space” is an area that a person does not want other people to enter, and can be set for each person. Furthermore, the personal space tends to be wider for men than for women and wider for adults than for children. Furthermore, the personal space also depends on a relationship between persons, and tends to be narrow between men or between women, and wide between opposite genders. Moreover, in a case where two persons are a married couple or a parent and a child, the personal space may be narrower between the two persons than between strangers.
[0068] Furthermore, there is an inter-vehicle distance as a concept similar to the personal space. The personal space is set to make the position of the bounding box BB more natural in a case where the subject is a person. The inter-vehicle distance is set to make the position of the bounding box BB more natural in a case where the subject is a vehicle.
[0069] If the bounding box BB of the vehicle is automatically generated, there is a high possibility that data of a state in which the vehicles are extremely close to each other is generated. However, by setting the “inter-vehicle distance” as an internal parameter, the positional relationship between the vehicles maintains an appropriate inter-vehicle distance, and it is possible to increase a possibility of generating appropriate simulation data MD. Furthermore, since the inter-vehicle distance of the vehicle increases in proportion to a speed of the vehicle, it is also effective to further set information of “speed” as an internal parameter. Yet another similar example is a transport vehicle in a warehouse. Considering a video of a monitoring camera of the warehouse, there is a high possibility that not only a person but also a transport vehicle (autonomous mobile robot (AMR)) such as a forklift is present in the warehouse. At this time, a distance to surroundings is set in the transport vehicle for ensuring safety. Therefore, for the transport vehicle, a “safe distance” may be set as an internal parameter. Meanwhile, since the transport vehicles in the warehouse may travel in a coupled manner, there may be a case where the inter-vehicle distance proportional to the speed of the vehicle does not need to be considered.
[0070] As described above, a sense of distance between the subjects varies depending on types and situations of the subjects, and it is possible to generate natural simulation data MD by setting the personal space, the inter-vehicle distance, or the like as the internal parameter in consideration of the types and situations of the subjects. Note that the sense of distance between the subjects is a sense of distance in three dimensions, and is different from the distance on the image of the bounding box BB projected in two dimensions. For example, the bounding boxes BB for the subjects arranged in an optical axis direction of the imaging device are not uncomfortable even if they overlap each other. Furthermore, in a case where persons pass each other in a store passage, the bounding boxes BB may overlap each other.
[0071] As described above, even if there is a three-dimensional gap between humans, the distance between the bounding boxes BB on the projected image varies. A human has an ability to grasp a three-dimensional space from a two-dimensional image, and can imagine what kind of three-dimensional space has been projected to be a visualized two-dimensional image. Therefore, the human can understand how the overlapping of the bounding boxes BB has three-dimensionally occurred in a positional relationship, and can determine the presence or absence of discomfort.
[0072] Meanwhile, since the large-scale language model M has also learned large-scale data, and thus has empirically learned “the distance between subjects in a depth direction of the image corresponds to a depth direction of the imaging device and may overlap”, “the distance between the subjects in a horizontal direction of the image corresponds to the distance between the subjects and it is natural that there is a gap”, and further, “the bounding box BB becomes larger in a front side in the image”, although the large-scale language model M may not have acquired the ability to grasp a three-dimensional space like a human. Therefore, it can be said that the large-scale language model M has learned the natural arrangement of the bounding boxes BB. Therefore, in a case where a human instructs the setting of the personal space, the large-scale language model M can generate the simulation data MD that reflects the intention.
[0073] The third prompt P3 is a sentence such as “Please set a purpose of visit for each customer, and assume that each customer is moving in the store so as to achieve the purpose.”
[0074] The information regarding the subject to be detected is, for example, attribute information such as gender, age, height, weight, dominant arm, and stride of the subject, information indicating that the subject is a wheelchair user or a customer who visits the store on a bicycle, information of a direction that a body is facing, and the like.
[0075] By setting such information regarding the subject, the large-scale language model M can bring a change in the coordinates of the bounding box BB based on a motion of the customer closer to a natural change. That is, it is possible to bring the simulation data MD output from the large-scale language model M closer to more natural one.
[0076] As an example, by setting the direction that the body is facing to the subject, a changing direction of the bounding box BB becomes natural. Furthermore, in a case where the subject changes a traveling direction, it is considered that the subject performs a movement of changing the direction of the body, and it is possible to generate more natural simulation data MD such that the position of the bounding box BB does not change much while the subject is performing the movement of changing the direction.
[0077] The information of the personal space of the subject is information that can change depending on the gender, age, body type, personality, race, or the like of the subject.
[0078] By setting the personal space, the movement of the coordinates of the bounding box BB in a case where the customer moves along a natural route can be reflected in the simulation data MD.
[0079] The information regarding the space SP to be sensed is information such as size and location conditions of the space SP, or weather, time information, season information, and the like. Furthermore, the location conditions are conditions such as information of other stores existing in the surroundings or a density of buildings. Alternatively, the information regarding the space SP may be information of a type of the space SP. Information for identifying whether or not the space SP is the space SP in a theme park, the space SP in a stadium, or the space SP in a store is an example of the information regarding the space SP to be sensed. For example, the motion of a person walking on a road and the motion of a person walking on a road in a theme park may be different, and by inputting these pieces of information to the large-scale language model M, it is possible to reflect the motion of a person according to the situation in the simulation data MD.
[0080] The third prompt P3 is a sentence giving an instruction to set these various types of information, and is, for example, a sentence such as “Please generate metadata after setting attribute information for each customer.”
[0081] The answer information acquisition unit F2 performs processing of acquiring the output of the large-scale language model M as the answer information.
[0082] The presentation processing unit F3 performs processing of presenting various screens to be described below to the developer. On the screen presented to the developer, a sentence is input by the developer, or the answer information of the large-scale language model M acquired by the answer information acquisition unit F2 is presented.
[0083] The information provision processing unit F4 performs processing of providing information to assist the generation of the simulation data MD to the large-scale language model M. Here, the information is described as “auxiliary information SD”.
[0084] The auxiliary information SD is, for example, information regarding a data format of the simulation data MD such as a JSON format.
[0085] The auxiliary information SD may be, for example, information regarding the imaging device arranged in the space SP. Specifically, the auxiliary information SD may include model number information and manufacturer information of the imaging device, pixel number information and angle of view information, and the like. Furthermore, sensor type information capable of identifying a type of an image sensor included in the imaging device, such as a color sensor such as a monochrome sensor, an RGB sensor, or a CMYK sensor, a distance measuring sensor that outputs a distance image, a temperature sensor that outputs a heat map image, an EVS sensor that outputs an image including event data, can be regarded as an example of the auxiliary information SD.
[0086] The auxiliary information SD may be information of the number and positions of the imaging devices arranged in the space SP to be sensed.
[0087] The auxiliary information SD may be information obtained using the vector database VD. For example, the simulation data generation device 2 searches the vector database VD using the first prompt P1 input by the developer to obtain a search result. The search result obtained by the vector database VD can be said to be the auxiliary information SD.
[0088] By inputting these pieces of auxiliary information SD, the information provision processing unit F4 brings the simulation data MD as the answer information obtained from the large-scale language model M closer to more natural one.
[0089] The auxiliary information acquisition unit F5 performs processing of acquiring the auxiliary information SD described above. The acquisition of the auxiliary information SD may be performed using, for example, the vector database VD of the second server device 4, or may be performed in response to an input of the developer in the developer terminal 1.
[0090] The simulation data generation device 2 acquires the sentence input by the developer as the above-described various prompts and inputs the sentence to the large-scale language model M, thereby obtaining the simulation data MD desired by the developer. That is, the developer brings the generated simulation data MD closer to desired one while performing interactive exchange with the large-scale language model M. Therefore, the developer does not need to consider a perfect first prompt P1 for generating the desired simulation data MD at a time. Furthermore, by inputting the second prompt P2 for correcting the generated simulation data MD many times, it is possible to efficiently and effectively obtain the desired simulation data MD.
[0091] Furthermore, it is possible to efficiently obtain the natural simulation data MD by appropriately inputting the third prompt P3 and the auxiliary information SD in the process.
[0092] An example of obtaining desired simulation data MD by accumulating interactive exchanges will be described.
[0093] Note that some preconditions are set for the description. It is assumed that the developer desires the simulation data MD as metadata that simulates a travel route of a customer of a convenience store.
[0094] Furthermore, the developer desires the simulation data MD for verification of an algorithm implemented in the application to be developed.
[0095] Furthermore, the desired simulation data MD is the simulation data MD for a monitoring camera that captures an image of an inside of the convenience store during daytime of August. However, it is assumed that the vector database VD as the user database included in the second server device 4 stores video data of the monitoring camera that captures the inside of the convenience store in the morning of February, and does not store similar video data for the daytime of August.
[0096] First, the developer performs the following input together with the first prompt P1 such as “We will create metadata indicating a travel route of a person in the convenience store by referring to a video obtained by capturing the inside of the convenience store in February of this year. Please generate the metadata in the data format specified below as output data including coordinates of a human bounding box.”
[0097] { "descriptions": { { "category": "person", "bbox": { 40, 30, 100, 200}, "score" : 0.80 }, { "category": "person", "bbox": { 34, 28, 90, 165}, "score" : 0.60 } } }
[0098] The developer browses presentation data obtained by visualizing the simulation data MD generated by the large-scale language model M by the above-described input sentence.
[0099] Furthermore, the developer grasps correction points on the basis of the presented visualized data. Here, as an example, it is assumed that the developer notices that a large number of bounding boxes forms a line in a passage in the convenience store and human movement is stopped (congested).
[0100] Then, it is assumed that the developer feels that the video specified in the previous first prompt P1 is for the morning of February, and thus unnatural data is obtained as the daytime simulation data MD in August.
[0101] In this case, the developer can further input the second prompt P2 such as “The number of people seems to be too large. Please set the data with the maximum of five people in the convenience store.”
[0102] Therefore, the large-scale language model M generates new simulation data MD obtained by correcting the simulation data MD described above. The generated new simulation data MD is visualized and presented to the developer.
[0103] It is assumed that the developer who has confirmed the data visualizing the simulation data MD further feels unnaturalness in “The line of people in the store has disappeared, but people are concentrating on a specific place in the convenience store and picking up an item.” Then, it is assumed that the developer has come to think that this unnaturalness is caused by the fact that the video specified in the previous first prompt P1 is one in the morning of February, and there are many customers who buy sandwiches, rice balls, and the like during commuting hours.
[0104] In response, the developer may input the second prompt P2 such as “Consider that customers at convenience stores often buy cold drinks or ice cream.”
[0105] Moreover, the developer can appropriately input the following second prompt P2 each time the visualized new simulation data MD is confirmed.
[0106] “This is a summer scene. Please keep a little distance from each other because it is hot.” “Please increase a percentage of children a little more in consideration of the fact that the children do not need to go to school during the summer vacation.” “Please create variation in the walking speed of persons by adding a person who walks a little faster.” “It is unnatural that all the people stop in front of the shelf for the same length of time. Please create various patterns so that some people walk fast and take an item immediately, and some people walk slowly and take time to choose an item.”
[0107] By sequentially inputting such a second prompt P2 by the developer, the simulation data MD can be gradually corrected to match the intention of the developer.
[0108] Furthermore, for the visualized simulation data MD, after the coordinate change of the appearing person (bounding box BB) is corrected to a desired one, the developer may further input the second prompt P2 such as “Please change the format of the output data from the current JSON format to the CSV format specified below.”
[0109] Furthermore, the developer performs the following input together with the second prompt P2.
[0110] { object, id, x1, y1, x2, y2 }
[0111] Here, “id” indicates a number uniquely assigned to a human as the subject located within the angle of view, or the like. Furthermore, “x1, y1, x2, y2” indicates the coordinate information of two diagonal points of the bounding box BB for the human to be detected.
[0112] As described above, reasons why it is possible to obtain the desired simulation data MD by accumulating interactive exchanges are as follows.
[0113] It is difficult for the human as the developer to completely specify all conditions from zero. If the developer knows in advance what kind of correction points or unnaturalness is included in the simulation data MD generated by the large-scale language model M, he / she can give a condition to the first prompt P1 that is the first instruction so as not to be such simulation data MD. However, there is no way for the developer to know in advance what the simulation data MD generated by the large-scale language model M is, which is difficult.
[0114] Furthermore, even if the simulation data MD is incomplete, the human as the developer can feel a sense of discomfort by checking the generated and visualized simulation data MD for the time being.
[0115] Moreover, it is difficult for the human as the developer to quantitatively describe a change amount of coordinate data included in the simulation data MD in order to eliminate the sense of discomfort. Meanwhile, it is relatively easy for the human as the developer to qualitatively express a method of correcting the coordinate data included in the simulation data MD in a natural language in order to eliminate the sense of discomfort.
[0116] Then, the human as the developer can compare the presented data to determine which is closer to the desired simulation data MD. However, it is difficult for the human as the developer to explain a determination basis deriving a comparison result of the presented data.
[0117] Moreover, for the simulation data MD generated by the large-scale language model M, there is no need for the developer to instruct again that there is no discomfort. For example, for the simulation data MD generated by the large-scale language model M, if the number of people in the store is appropriate from the beginning, the developer does not need to intentionally give an instruction regarding the number of people as a prompt from the beginning to the end.
[0118] Due to such characteristics of the human or the large-scale language model M, the developer can easily acquire the simulation data MD for a natural scene as compared with a case where the developer performs simulation from zero and needs to intentionally and completely specify all the conditions. In other words, the developer can acquire the natural simulation data MD by appropriately inputting a prompt using a natural language that humans are good at, in other words, by simply pointing out the sense of discomfort about the simulation data MD as a prompt.
[0119] Note that, in the example illustrated immediately before, data as a template of metadata is first extracted from the user database prepared by the developer and provided to the large-scale language model M, and the simulation data MD is improved by interactively repeating the input of the prompt. The present embodiment is not limited thereto, and a template of metadata acquired by the large-scale language model M without referring to another database may be used.
[0120] Moreover, in the input of the second prompt P2, a correction policy is indicated in natural language. However, in a case where there is data to be reference in the user database or trained data of the large-scale language model M, a similar effect may be obtained by inputting the second prompt P2 such as “Please correct the metadata with reference to the video captured at the xxx convenience store in month x of year xxx.”
[0121] <4. Flow of Processing> An outline of a flow of processing executed in the developer terminal 1 and the simulation data generation device 2 will be described with reference to Fig. 4. Note that each processing illustrated in Fig. 4 is processing executed by the processing circuit 51 of each device. Here, the “processing circuit 51 of the developer terminal 1” as a processing execution entity will be simply described as the “developer terminal 1”. The same similarly applies to the simulation data generation device 2.
[0122] In response to the developer performing an operation of activating an application for generating the simulation data MD for the developer terminal 1, the developer terminal 1 activates a simulation data generation application AP in step S101. In this activation processing, for example, processing of accessing a cloud application provided by the simulation data generation device 2 and requesting the interaction screen G1 is performed.
[0123] The simulation data generation device 2 performs processing of presenting the interaction screen G1 in step S201 in response to a request of the interaction screen G1.
[0124] In response to the request, the developer terminal 1 performs processing of displaying the interaction screen G1 in step S102.
[0125] Fig. 5 illustrates an example of the interaction screen G1. The interaction screen G1 includes a title bar 21, an interaction field 22, and an input field 23.
[0126] In the title bar 21, a name or the like of a service that is used by the developer and is provided by the simulation data generation device 2 is displayed. In the present example, since the service generates and provides the simulation data MD that simulates the metadata output from the imaging device, the title “simulation data generation application” is displayed in the title bar 21.
[0127] Although specifically described below, an input sentence input by the developer and a response sentence output from the simulation data generation device 2 according to the input sentence are displayed in chronological order in the interaction field 22. Note that not only sentences but also buttons, options, and the like may be displayed in the interaction field 22.
[0128] Note that, in the interaction field 22 of the interaction screen G1 illustrated in Fig. 5, an explanatory sentence of data generated by using the application and a sentence prompting an input of a condition of the simulation data MD to be generated are displayed as initial display.
[0129] The input field 23 is provided for the developer to input a sentence. In the input field 23, a sentence being input is displayed. Note that, in the illustrated interaction screen G1, a button or the like for transmitting an input sentence is omitted, but the button may be arranged in or near the input field 23.
[0130] In a case where the developer inputs a sentence instructing generation of the simulation data MD in the input field 23, the developer terminal 1 performs processing of accepting the input sentence as the first prompt P1 in step S103 of Fig. 4.
[0131] Fig. 6 illustrates an example of the interaction screen G1 at the time of executing the processing of step S103. As illustrated, a sentence “I want data output from the monitoring camera of the retail store.”, which is the sentence input by the developer, is displayed in the interaction field 22 of the interaction screen G1. This sentence can be treated as the first prompt P1 input to the large-scale language model M. Note that, in a case where the sentence displayed in the interaction field 22 of the interaction screen G1 is not appropriate as the first prompt P1 for the large-scale language model M, the simulation data generation device 2 may appropriately perform correction, addition, or the like.
[0132] The first prompt P1 input by the developer is acquired by the simulation data generation device 2 in step S202 of Fig. 4.
[0133] Note that the developer terminal 1 may transmit the sentence input by the developer to the simulation data generation device 2 without determining whether or not the sentence is the first prompt P1, and the simulation data generation device 2 may determine whether or not the sentence corresponds to the first prompt P1, that is, whether or not the sentence is a sentence for requesting the large-scale language model M to generate the answer information.
[0134] Subsequently, the simulation data generation device 2 performs information setting processing in step S203. The information set here is information set by inputting the above-described third prompt P3 to the large-scale language model M. That is, the processing in step S203 can also be said to be processing of inputting the third prompt P3 to the large-scale language model M.
[0135] Therefore, before the generation of the simulation data MD, the information regarding the captured image, specifically, the information regarding the subject to be detected, the information regarding the personal space of the subject, and the information regarding the space SP to be sensed are set.
[0136] Subsequently, in step S204, the simulation data generation device 2 executes processing of setting the auxiliary information SD.
[0137] Here, specific content of the processing of setting the auxiliary information SD is illustrated in Fig. 7. The processing of setting the auxiliary information SD is implemented by, for example, cooperation of the simulation data generation device 2 and the second server device 4.
[0138] In step S221, the simulation data generation device 2 transmits the first prompt P1 to the second server device 4. Note that this transmission processing can be rephrased as processing of transmitting a search query in the vector database VD included in the second server device 4.
[0139] In step S401, the second server device 4 receives the first prompt P1.
[0140] In step S402, the second server device 4 performs processing of adding a vector to the sentence as the first prompt P1.
[0141] Note that the processing of adding the vector may be performed on the entire sentence as the first prompt P1, or may be performed for each of a plurality of keywords extracted from the first prompt P1.
[0142] In step S403, the second server device 4 executes search processing using the vector database VD. With this processing, the second server device 4 searches for the search query based on the first prompt P1, and obtains a search result.
[0143] In step S404, the second server device 4 transmits the search result extracted from the vector database VD to the simulation data generation device 2.
[0144] In step S222, the simulation data generation device 2 receives the search result from the second server device 4. The information received by the simulation data generation device 2 from the second server device 4 is the auxiliary information SD.
[0145] The description returns to Fig. 4. In step S205, the simulation data generation device 2 obtains the answer information from the large-scale language model M. The answer information from the large-scale language model M is the above-described simulation data MD. Note that the acquisition of the answer information from the large-scale language model M is implemented by inputting the third prompt P3 and the auxiliary information SD to the large-scale language model M as necessary in addition to the first prompt P1.
[0146] In step S206, the simulation data generation device 2 performs visualization processing. The simulation data MD is data simulating the metadata, and may be difficult for the developer to understand as it is. The processing of step S206 is processing of processing the simulation data MD output from the large-scale language model M so as to be easily viewed. Note that the visualization processing may be implemented by instructing the large-scale language model M to perform processing such as graphing processing for easy viewing.
[0147] The simulation data generation device 2 obtains data such as a graph, a video, and a table as a result of the visualization processing.
[0148] An example is illustrated in Fig. 8. Note that Fig. 8 is an example of a video, and is a diagram illustrating an extracted image of an nth frame and an extracted image of an (n + 1)th frame in the video.
[0149] The image of each frame includes an outer frame 31 imitating the angle of view of the imaging device and a plurality of bounding boxes BB arranged inside the outer frame 31.
[0150] Note that, in the image of each frame, to facilitate understanding, the image may be expressed such that positions of furniture and obstacles, walls, windows, doors, and the like arranged within the angle of view of the imaging device can be recognized. For example, in the example illustrated in Fig. 9, a video in a state where not only the bounding boxes BB but also furniture 32 arranged within the angle of view of the imaging device is displayed is presented to the developer.
[0151] Another example of data obtained by the visualization processing is illustrated in Fig. 10. Fig. 10 is a graph illustrating a counting result of the number of subjects appearing within the angle of view, and is an example of the confirmation screen G2 for the developer to confirm the result of the visualization processing.
[0152] On the confirmation screen G2, as illustrated, the number of people (that may be an average number of people) in the angle of view for each predetermined time such as ten minutes is represented by a graph by the visualization processing. By confirming the graph, the developer can confirm a recognition result of the imaging device, and thus, an outline of the simulation data MD.
[0153] The description returns to Fig. 4. In step S207, the simulation data generation device 2 performs presentation processing. In this processing, the generation of the simulation data MD is notified to the developer on the interaction screen G1, and a means for downloading the generated various data is presented to the developer.
[0154] Fig. 11 illustrates an example of the interaction screen G1 presented to the developer by the presentation processing. The interaction screen G1 presented to the developer by the processing in step S207 displays an outline of various types of information set by the third prompt P3 and not output as the simulation data MD. Specifically, in generating the simulation data MD, it is notified that the size of the store, the number of cameras, and the attribute information for each customer to the store are set.
[0155] Furthermore, the imaging device actually installed in the space SP recognizes the subject captured in the image for each imaging frame, but not all the recognition results are necessarily correct.
[0156] For example, processing of detecting a person and setting the bounding box BB will be considered. Even if two persons are detected up to a frame immediately before a certain frame, the two persons overlap in the frame so as to be aligned in an optical axis direction of the imaging device, and it is conceivable that only one person can be detected. In this case, in an actual imaging device, one bounding box BB is set as a detection result. However, in the simulation, since it is known that there are two persons, two bounding boxes BB are set to overlap each other.
[0157] In the present example, in consideration of occurrence of such an erroneous recognition in the imaging device with a certain probability, the developer is notified that the large-scale language model M has been instructed to set the setting so that an incorrect recognition result is obtained with a certain probability. Therefore, the simulation data MD in which two persons are erroneously recognized as one person with a certain probability is obtained. That is, the simulation data MD appropriately simulating the metadata output from the actual imaging device can be obtained.
[0158] Note that various other erroneous detections are conceivable. For example, there are a case where a frame in which a part of the subject is not able to be detected due to backlight occurs, a case where the subject repeatedly comes out of the angle of view and enters the angle of view in a peripheral edge portion of the angle of view of the imaging device is not able to be detected, a case where occlusion occurs in a case where a passenger car overtakes a bus in detection of vehicles and the vehicles are not able to be detected, and a case where a passenger car and a bus are integrated and detected as one vehicle.
[0159] The simulation data MD appropriately simulating such a case where such erroneous detection occurs can be obtained.
[0160] Moreover, there is a case where a normally unassumed action is performed within the angle of view of the imaging device. For example, in the case of a retail store, a shoplifting action or the like is conducted. In addition, for example, a behavior of a vehicle entering the store due to an erroneous operation of an accelerator and a brake, an action of drowning in a river, and the like are conceivable. These actions can also be rephrased as actions that can be caused by an abnormal subject.
[0161] In the following description, these normally unassumed actions are referred to as “specific actions”. Note that “unassumed” means an action that is not assumed by a subject who performs an action or an object such as a salesclerk or a manager who receives the action, and does not mean that the action is not expected as a monitoring purpose. Specifically, it is not assumed for a normal person that a person performs a shoplifting action instead of shopping, but it is an action that can be assumed as a monitoring purpose of the imaging device.
[0162] The developer is notified that generation of such a “specific action” without an instruction of the developer, and generation of the simulation data MD related to the metadata to be output from the imaging device at the generation of the specific action have been instructed. These instructions are implemented by inputting a predetermined prompt to the large-scale language model M in step S205. Specifically, this is implemented by inputting a sentence such as “Please also generate simulation data in a case where there is a shoplifting person.”, a sentence such as “Please generate simulation data in consideration of a case where a vehicle enters the store.”, or a sentence such as “Please generate simulation data for a case where there is a person who has drowned in a river.” to the large-scale language model M together with the first prompt P1 and the third prompt P3 as a prompt.
[0163] As described above, in the presentation processing in step S207, not only the simulation data MD as the answer information obtained from the large-scale language model M is presented to the developer, but also various conditions, settings, and the like used for generating the simulation data MD may be appropriately notified.
[0164] Note that, in the example illustrated in Fig. 11, presentation of the simulation data MD is performed by presenting a button Btn1 linked to download the simulation data MD, instead of by directly presenting the simulation data MD. Furthermore, presentation of visualized data is also performed by presenting a button Btn2 linked to download a video file.
[0165] The description returns to Fig. 4. The developer terminal 1 appropriately performs processing of displaying the simulation data MD in step S104 by the developer's operation.
[0166] For example, in response to the developer operating the button Btn2, video data obtained by visualizing the simulation data MD is downloaded, and the downloaded video data is played back and displayed.
[0167] The developer can check the content of the simulation data MD and appropriately give a correction instruction of the simulation data MD by playing back and displaying the downloaded video data. In the case of giving a correction instruction, the developer inputs a sentence for instructing correction in the input field 23 of the interaction screen G1.
[0168] In a case where the developer inputs a sentence instructing correction of the simulation data MD in the input field 23, the developer terminal 1 performs processing of accepting the input sentence as the second prompt P2 in step S105 of Fig. 4. Note that determination as to whether or not the input sentence is the second prompt P2 may be performed by the developer terminal 1 or by the simulation data generation device 2.
[0169] Various sentences are conceivable as the second prompt P2. For example, as illustrated in Fig. 12, there are countless examples including a sentence such as “The number of people seems to be small. Please re-create the metadata with the maximum of five people in the store.”, a sentence such as “Please diversify your personal space.”, a sentence such as “Please consider the motion of a customer in a case of viewing an item advertisement posted on the wall and purchasing the item.”, a sentence such as “Please set the number of customers who come to the store in consideration of time zone.”, and a sentence such as “Consider a customer to the store who is just coming to cool without any purpose of item purchase in the summer.”
[0170] In response to the developer terminal 1 accepting the input of the second prompt P2 and transmitting the second prompt P2 to the simulation data generation device 2 in step S105 of Fig. 4, the simulation data generation device 2 performs processing of acquiring the second prompt P2 in step S208.
[0171] Subsequently, in step S209, the simulation data generation device 2 obtains the answer information from the large-scale language model M. The acquisition of the answer information from the large-scale language model M is implemented by inputting the second prompt P2 to the large-scale language model M. Furthermore, additional auxiliary information SD or the like may be input to the large-scale language model M as necessary, or a prompt for instructing additional setting may be input.
[0172] Note that the processing in step S209 is similar to the processing in step S205 described above, and is processing of inputting a prompt to the large-scale language model M to obtain the answer information.
[0173] After step S209, each processing of steps S206 and S207 is performed in the simulation data generation device 2. Furthermore, the developer terminal 1 performs both pieces of processing of step S104 and step S105 accordingly.
[0174] That is, in the developer terminal 1 and the simulation data generation device 2, each processing of steps S205, S206, S207, S104, S105, and S208 may be repeated a plurality of times.
[0175] Next, the prompt stored in the first server device 3 and input from the simulation data generation device 2 to the large-scale language model M and the answer information output from the large-scale language model M to the simulation data generation device 2 will be described with reference to Fig. 13. Note that, in Fig. 13, they are described in chronological order from the top.
[0176] First, the first prompt P1, the third prompt P3, and the auxiliary information SD are input to the large-scale language model M.
[0177] From the large-scale language model M, the simulation data MD is output as the answer information.
[0178] Moreover, in a case where the second prompt P2 is input to the large-scale language model M, the corrected simulation data MD is output from the large-scale language model M as the answer information.
[0179] The input of the second prompt P2 and the output of the corrected simulation data MD can be executed a plurality of times.
[0180] Fig. 14 illustrates an example of the confirmation screen G2 for confirming information generated as a result of performing visualization for the corrected simulation data MD.
[0181] Note that the example of the confirmation screen G2 illustrated in Fig. 14 is a screen example for confirming a result of visualizing a graph generated as a result of improving the example illustrated in Fig. 10 in which the simulation data MD before correction is graphed. Furthermore, the correction in the present example is performed in response to an input of the sentence “Please set the number of customers who come to the store in consideration of time zone.” to the large-scale language model M as the second prompt P2.
[0182] As illustrated in each drawing, in the graph illustrated in Fig. 10, the detection result of the number of people has no deviation regardless of the time zone, and the mode of increase or decrease is uniform, whereas in the graph illustrated in Fig. 14, the maximum value of the detection result of the number of people varies depending on the time zone, and the deviation occurs in the number of people. That is, the corrected simulation data MD is obtained by taking into account a change in the degree of crowding of the store depending on the time zone, and is more natural simulation data MD.
[0183] <5. Modification> An example in which the third prompt P3 is input by the developer by inputting the sentence indicating the instruction content for the large-scale language model M has been described. Specifically, the developer inputs a sentence such as “Please generate simulation data after setting attribute information for each customer.”
[0184] The example is not limited thereto, and the developer may simply input an evaluation for the simulation data MD generated by the large-scale language model M. The evaluation input by the developer is appropriately used for generating a prompt by the simulation data generation device 2. For example, in response to the developer inputting “Variations in the number of customers to the store depending on the time zone are small.”, the simulation data generation device 2 may generate a prompt “Please generate the simulation data MD in consideration of variations of the customers to the store depending on the time zone.” as the third prompt P3 and input the third prompt P3 to the large-scale language model M in step S209 in Fig. 4.
[0185] In the example illustrated in Fig. 6, the sentence input by the developer is only the sentence corresponding to the first prompt P1. Alternatively, as illustrated in Fig. 15, the developer may input a sentence specifying the auxiliary information SD or a sentence corresponding to the third prompt P3 together with the sentence corresponding to the first prompt P1.
[0186] Note that, in the example illustrated in Fig. 15, the developer specifies that the detection target is the “customer to the store”. In this way, by specifying a class to be detected, it is possible to appropriately generate the simulation data MD desired by the developer. Note that class information for specifying the class to be detected is information indicating categories of the subject, and distinguishes, for example, “person”, “automobile”, “airplane”, “ship”, “truck”, “bird”, “cat”, “dog”, “deer”, “frog”, “horse”, and the like. That is, the class information for specifying the class of the subject to be detected may be the above-described auxiliary information SD.
[0187] In Figs. 8 and 9, it has been described that the developer is shown the image in which the plurality of bounding boxes BB is arranged inside the outer frame 31 imitating the angle of view of the imaging device. For the bounding box BB related to the above-described normally unassumed action, a presentation mode to the developer may be changed.
[0188] For example, Fig. 16 is a diagram illustrating a certain frame in a case where the simulation data MD that simulates the metadata output from the imaging device that captures an image of tourists visiting a river is presented in a video. That is, a state in which the tourists are detected by the number of bounding boxes BB is simulated. In Fig. 16, a bounding box BB1 related to a detection result of a drowning person is displayed with the thicker line than the other bounding boxes BB.
[0189] In addition to the above example, display of changing a line type of the bounding box BB for a person who has performed the normally unassumed specific action, or display of changing a line color may be performed.
[0190] Furthermore, in the display in the video, the frame before drowning may be displayed in the same mode as the other bounding boxes BB, and the frame at the moment of drowning may be displayed in a different mode from the other bounding boxes BB.
[0191] According to these modes, the developer can recognize the bounding box BB for the person performing the specific action, and it becomes easy to determine whether or not the movement of the bounding box BB is natural.
[0192] Note that such a change in the line type or line color of the bounding box BB is a measure for improving efficiency of work of interactively correcting the simulation data MD by the large-scale language model M and the developer, and the metadata as the simulation data MD to be finally output may be a normal bounding box. That is, it is also possible to add a flag indicating that the action is the specific action to the metadata, but this is not indispensable.
[0193] <6. Variations> <6-1. Specific Action> Note that the normally unassumed specific action may be various actions other than those described above. For example, the specific action can be included in an unconscious action or a natural motion. Specifically, various motions are conceivable, such as “a motion in which a standing person swings the body”, “a motion of standing upright but standing with the center of gravity on the right foot”, “a motion of alternately repeating a state in which the center of gravity is placed on the right foot and a state in which the center of gravity is placed on the left foot”, “a motion of looking at the upper left in a case of thinking something or looking at the lower right in a case of trying to recall something while sitting and writing", and “a motion of re-sitting on a chair”.
[0194] Furthermore, the specific action may also be included in an unexpected but usually seen motion. Specifically, various motions are conceivable, such as “a person stumbles”, “a motion of stopping because clothes are caught by an obstacle”, “a motion of coming across an acquaintance and starting talking”, “a motion in which a vehicle suddenly stops in front of a camera”, “a motion in which a vehicle enters an intersection immediately after a traffic light changes from yellow to red”, and “a motion in which a vehicle turns left at a place where a left turn is prohibited”.
[0195] Moreover, rare movements may also include the specific action. Specifically, in a case where the subject is a “person”, “getting involved in an accident”, “quarreling on the road”, “shoulders of the subjects collide”, and the like are conceivable. Furthermore, in a case where the subject is a “vehicle”, “colliding with another vehicle or an obstacle”, “hitting a person”, “getting road rage”, and the like are conceivable. Furthermore, in a case where the subject is a “dog”, “there is no owner around”, “biting a person”, and the like are conceivable. Note that these are merely examples.
[0196] <6-2. Example of Interactive Prompt> An input example of various prompts will be described.
[0197] The first prompt P1 is, for example, a sentence for providing a file formed in a sample data format and then specifying the file. Specifically, the first prompt P1 such as “aaa1234.json is metadata in which the coordinates of the bounding box BB surrounding a person in the convenience store are output. Please create metadata with a time length of twenty minutes in the same format as this data format.” is input.
[0198] It has been described that the second prompt P2 can be performed a plurality of times. Here, an example in which the input of the second prompt P2 is performed three times is illustrated. The first second prompt P2 is a sentence “The number of people seems to be small. Please re-create the metadata with the maximum of five people in the store.”
[0199] The second second prompt P2 is a sentence “Human motion is monotonous. Please create metadata having various people, such as a person who stops and picks up an item from a shelf, and a person who leaves immediately because there is no desired item on the shelf.”
[0200] The third second prompt P2 is a sentence “Please create metadata assuming that one of five people is a child and two of them are women.”
[0201] By repeatedly inputting the second prompt P2 in this manner, the generated simulation data MD is gradually improved closer to desired data. Furthermore, the generated simulation data MD may be visualized as described above each time the simulation data MD is corrected. Thereby, the developer can find a sense of discomfort or an improvement point with respect to the presented simulation data MD, and can input the second prompt P2 for correction. Therefore, it is possible to gradually improve the generated simulation data MD to be more natural.
[0202] Note that it is difficult to input the second prompt P2 particularly instructing an implicit assumption in such an interactive exchange. For example, the kind of second prompt P2 indicating an implicit assumption such as “Do not allow humans to float in the air.” exists indefinitely, and it is difficult to input all of them. However, the large-scale language model M is naturally trained with an implicit assumption such as a human not floating in the air in the course of training, and the second prompt P2 instructing such correction is often unnecessary. That is, the developer can obtain the natural simulation data MD with a simple instruction using a natural language.
[0203] Note that some examples of variations of the second prompt P2 will be further given.
[0204] -"The position of the camera capturing the inside of the convenience store seems to be low. Please capture data of the camera set at a higher position." -"The format of the data has improved. Please increase the data length to thirty minutes instead of twenty minutes." -"Please set the output interval of the metadata to five seconds instead of ten seconds. Please be noted, at the setting, the travel amount on the screen is reduced because the travel speed of a human does not change." -"In the convenience store, there are shelves in the vertical direction. Please re-create the travel route in consideration of the fact that a human can travel laterally only on the innermost or near passage."
[0205] These second prompts P2 are input to further perform correction in a case where there is a sense of discomfort in the presented simulation data MD.
[0206] Furthermore, the second prompt P2 for generating the new simulation data MD is also conceivable, not for solving the sense of discomfort. That is, the second prompts P2 exemplified below are examples of the second prompts P2 for changing the conditions and generating the new simulation data MD.
[0207] -"Please change the data to the data of the monitoring camera that captures the lobby of the office building instead of the data in the convenience store." -"Please double the size of the convenience store only in the depth direction."
[0208] <6-3. Effective Use Example of User Database> The user database included in the second server device 4 will be described. The vector database VD as the user database described above has a function to search for similar data similar for the first prompt P1 input by the developer.
[0209] For example, in response to the first prompt P1 such as “Please create data of travel trajectories of a human in the convenience store.”, the vector database VD is searched for past cases, and the search result is input together with the first prompt P1 to the large-scale language model M.
[0210] Thereby, the large-scale language model M can present the simulation data MD with reference to past cases.
[0211] The user database included in the second server device 4 is not limited to the vector database VD, and may be a DB in another format such as an RDB. That is, the user database included in the second server device 4 can be regarded as a database in which the auxiliary information SD (hint information) that can be provided to the large-scale language model M is accumulated.
[0212] For example, the developer can search for the data format of the metadata created in the past and provide the data format to the large-scale language model M as the auxiliary information SD. The metadata output from the imaging device is data in a JSON format or a CSV format. For example, the data format of the JSON format is complicated, and it is difficult to express the JSON format in a natural language. That is, it is difficult for the developer to describe and provide the JSON format in an appropriate sentence.
[0213] In such a case, the developer can help generation of the MD by the large-scale language model M by acquiring the past metadata in the JSON format and providing the metadata as the auxiliary information SD to the large-scale language model M. Since the metadata in the JSON format is not generally disclosed, the large-scale language model M does not have knowledge thereof. However, if the developer can acquire the metadata in the JSON format acquired in the past by using the user database included in the second server device 4, the developer can provide the acquired metadata in the JSON format together with the first prompt P1 to the large-scale language model M. Thereby, the large-scale language model M can present the simulation data MD created in the appropriate JSON format to the developer.
[0214] Note that the auxiliary information SD may be input to the large-scale language model M in response to the second prompt P2, instead of inputting the metadata in the JSON format as the auxiliary information SD together with the first prompt P1 to the large-scale language model M. That is, the user database has a function to search for data similar to the second prompt P2 or the third prompt P3 input by the developer.
[0215] For example, by inputting a sentence “Please generate the metadata in the same JSON format as the previous AAA project's person detection.” as the second prompt P2 after acquiring the simulation data MD other than the JSON format from the large-scale language model M by inputting the first prompt P1, the second server device 4 that has received the input sentence can acquire the past metadata in the JSON format using the user database included in the second server device 4 and input the past metadata together with the second prompt P2 to the large-scale language model M.
[0216] Similarly, in a case where the developer inputs “Please generate the metadata in the CSV format used for the person detection in the past BBB project.” as the second prompt P2, the second server device 4 can acquire the past metadata in the CSV format using the user database included in the second server device 4 and input the past metadata together with the second prompt P2 to the large-scale language model M.
[0217] Furthermore, the developer can further enrich the simulation data MD based on various formats.
[0218] Specifically, by inputting the second prompt P2 such as “Please add posture information about each bounding box BB to the JSON format used for person detection in the past AAA project. Please add the posture information to the JSON format as a posture variable such as {"posture":"standing"} or {"posture":"sitting"}.”, it is possible to include more information to enrich the simulation data MD.
[0219] Furthermore, the developer can change the conditions of the past project to generate the desired simulation data MD. These changes and corrections can proceed in an interactive manner by repeating input of various prompts and acquisition of the simulation data MD.
[0220] Note that it is also possible to obtain the metadata as the simulation data MD not as a text file in the JSON format but as data serialized and encoded in a Base64 format.
[0221] For example, the developer can obtain the simulation data MD in a desired format by inputting a prompt “Please serialize and encode the simulation data MD in the same method as the method used in the CCC project.”
[0222] <6-4. Setting of Information Not Output as Answer Information> Variations of the third prompt P3 will be described.
[0223] As described above, the third prompt P3 is a prompt for instructing setting of information that is not output as the answer information from the large-scale language model M, and is a prompt for obtaining more natural simulation data MD.
[0224] That is, the internal parameters are information used for generating the natural simulation data MD and are not included in the simulation data MD output from the large-scale language model M.
[0225] Some examples of the internal parameters are given.
[0226] A first example of the internal parameters is a parameter regarding mass or direction. For example, an internal parameter for making vehicle travel trajectories natural is considered. Since trucks and buses have a large mass, energy necessary for acceleration and deceleration is larger than that of passenger cars. Therefore, for example, sudden stop by sudden braking is less likely to occur, and acceleration and deceleration are smaller than those of a passenger car.
[0227] Note that, in a case where the assumed metadata is the bounding box BB, the difference between trucks and buses can be ignored, and thus it is only necessary to generate the simulation data MD by dividing the vehicles as the subjects into two classes of passenger cars and buses.
[0228] That is, the developer inputs the third prompt P3 such as "Please set the vehicles to be detected to two classes of passenger cars and buses”.
[0229] Furthermore, the second prompt P2 may be input in addition to the third prompt P3. For example, the developer inputs a sentence “Compared to passenger cars, buses have twice the mass. The acceleration is approximately 1 / 2.” as the second prompt P2.
[0230] Therefore, the large-scale language model M can generate the simulation data MD including both a heavy vehicle and a small vehicle and having natural movement.
[0231] Furthermore, by further inputting the second prompt P2 such as “There are directions for buses and passenger cars. The probability of moving forward is 90%, and the probability of turning right or left is 10%. There is no need to consider retraction.”, the developer can bring the corrected simulation data MD closer to more natural data.
[0232] A second example of the internal parameters is a parameter regarding gender such as female or male, or an attribute regarding age such as child or adult.
[0233] These internal parameters are useful in creating a travel trajectory for person detection.
[0234] In general, there is a higher possibility that a male has a larger bounding box BB than a female, and an adult has a larger bounding box BB than a child.
[0235] Therefore, the developer can acquire the natural simulation data MD in which an adult and a child mix by inputting the third prompt P3 such as “The persons to be detected should be in two classes of adults and children.” Furthermore, it is also possible to input the third prompt P3 according to a place.
[0236] For example, the developer can cause the large-scale language model M to generate the appropriate simulation data MD according to the place by inputting the third prompt P3 such as “The place where the camera is installed is a vacation facility for children, so please increase the number of children.”
[0237] Moreover, the developer may input, as the third prompt P3, information such as “having a heavy load”, “having a large load”, or “being injured and can only walk slowly” for a certain person.
[0238] Furthermore, by inputting the third prompt P3 for setting a “weight of a load inside” for a heavy load or container, it is possible to generate the simulation data MD in consideration of ease of traveling, ease of movement, shaking caused by vibration, or the like of the subject to be detected.
[0239] In this manner, the third prompt P3 can be used in various ways such as expressing a difference between the subjects to be detected.
[0240] A third example of the internal parameters is for an internal parameter for the personal space described above. This internal parameter is for making a positional relationship among a plurality of persons natural.
[0241] The developer can obtain the simulation data MD in consideration of the personal space by inputting the third prompt P3 such as “Please set the personal space for each detected person.”
[0242] Furthermore, the developer may more specifically input the third prompt P3 such as “Please set a personal space distance for each bounding box and generate data in which adjacent bounding boxes move while keeping their personal spaces.” Note that, once a prompt for requesting the setting of the personal space is input, it can be regarded as the second prompt P2 that is a prompt for correcting the setting of the personal space.
[0243] Furthermore, the developer may input a sentence such as “When picking up an item in front of the shelf, please correct the data so that the personal space distance becomes small.”, as the second prompt P2.
[0244] By inputting these prompts, it is possible to obtain the natural simulation data MD in consideration of an appropriate personal space size according to the situation.
[0245] A fourth example of the internal parameters is an internal parameter for human habits.
[0246] For example, there are individual habits in the way of exercise such as sports. Taking running as an example, there are a person who starts from the right foot and a person who starts from the left foot. Furthermore, in a case where a break is interposed in the running, there is a person who stops and faces downward to recover physical strength, and there is a person who recovers physical strength while slowly walking.
[0247] Furthermore, as for a person who sits on a chair, there is a person who sits with his / her legs crossed or there is a person who shakes his / her knees like jiggling his / her knee up and down (jigging or foot shaking).
[0248] Such a difference in motion between individuals is not randomly generated but personal. That is, in a case where a certain person sits on a chair with his / her legs crossed, and in a case where the person walks around and sits again on the chair, there is a high possibility that the person performs a similar motion.
[0249] Therefore, the developer inputs setting of these personal habits for each person as the third prompt P3, so that the large-scale language model M can generate the simulation data MD simulating the movement of the bounding box BB appropriately reflecting the personality.
[0250] <6-5. Simulation Data Suitable for Application Development> As a case of using the simulation data MD generated by the large-scale language model M for the development of the application, the simulation data MD may be used for verification of an algorithm of the application.
[0251] Here, an example of a prompt for generating the simulation data MD appropriate for verification of the algorithm will be described.
[0252] The large-scale language model M having been trained with knowledge of a software test technique can generate the simulation data MD suitable for a test. Specifically, the large-scale language model M has been trained with a concept of a boundary value test and an equivalence division test.
[0253] For such a large-scale language model M, the developer inputs a sentence such as “Please create metadata for test using the software testing technique. First, please create the metadata so that the boundary value test and the equivalence division test can be performed for the coordinates of the four corners of the bounding box.”, as the first prompt P1.
[0254] This allows the developer to obtain the simulation data MD with the software tests in mind.
[0255] Moreover, the developer can further obtain the simulation data MD according to an intention by inputting the second prompt P2 such as “Please set boundaries to 0 and 10 on an X-axis and a Y-axis, respectively.” or the second prompt P2 such as “Please make two bounding boxes.”
[0256] That is, the second prompt P2 can cause the large-scale language model M to generate the simulation data MD according to the intention of the developer by dividing a plurality of instructions into each instruction instead of including the plurality of instructions at a time.
[0257] Another example of the prompt for generating the simulation data MD appropriate for verification of the algorithm will be described.
[0258] In the verification of the algorithm, it is possible to further improve perfection of the application by using the simulation data MD for test that is out of the developer's idea.
[0259] To obtain the simulation data MD that the developer may not be able to conceive by himself / herself, for example, the developer inputs the first prompt P1 such as “Please remove physical constraints of the coordinates of the four corners of the bounding box and create test data that is not normal.”
[0260] Thereby, the developer can acquire the simulation data MD that is not normally assumed, such as the simulation data MD in which the right and left and the up and down are reversed for the coordinates of the four corners of the bounding box BB, the simulation data MD in which the width and height of the bounding box BB rapidly change, and the simulation data MD in which some of the coordinates of the four corners of the bounding box BB are missing.
[0261] Moreover, the desired simulation data MD may be obtained by inputting a pre-preparation prompt before the first prompt P1, instead of generating the simulation data MD by inputting the first prompt P1 once.
[0262] For example, the developer first inputs the pre-preparation prompt such as “Please consider ten test cases that are not normal for the metadata including the coordinates of the bounding box.”
[0263] Subsequently, the developer inputs the first prompt P1 such as “For each of the ten cases, please create the metadata that actually includes the coordinate information of the bounding box as the test data that is not normal.”
[0264] By using such a method, it is possible to prevent generation of the simulation data MD created by a person with his / her own knowledge and biased or occurrence of omission or leakage in the case, and it is possible to create the test data with high completeness.
[0265] <6-6. Simulation Data Including Unsteady but Natural Changes> The large-scale language model M is trained using various videos available via a communication network such as the Internet. Then, these videos naturally include videos of monitoring cameras installed at various places.
[0266] Note that it is highly likely that most of the videos of the monitoring cameras are not suitable for development of an analysis algorithm of the metadata in the application. That is, the videos of the monitoring cameras tend to be videos with little change in the angle of view, such as a video in which a passerby travels right to left and passes through.
[0267] Furthermore, since a capturing position (camera position) of the monitoring camera and a surrounding environment are different for each space to be monitored, each video is regarded as a different new training material, and a large number of videos with little change are used for training. Furthermore, since the data desired by the developer is the simulation data MD itself as the metadata and is not a video, a large amount of video with little change is not indispensable.
[0268] Note that these videos are not completely unnecessary, and are very useful as statistical data indicating a tendency of human travel. For example, the videos by the monitoring cameras installed in station premises are videos having a difference in density and a difference in walking directions of users in the morning, daytime, and evening, a difference in density of users on a rainy day and a sunny day, a difference in distance between humans, a difference in walking speed, and the like. Use of the large-scale language model M constructed by training such a large amount of statistical video data is more likely to obtain the natural and appropriate simulation data MD than use of the human imagining to set statistical parameters.
[0269] That is, the developer can generate the simulation data MD incorporating a natural motion of the user based on the past videos used for training by the large-scale language model M by inputting, in stages, the first prompt P1 such as “Please generate the metadata including the coordinates of the bounding box for the user captured at a ticket gate of the station in the morning.” and the second prompt P2 such as “Please generate the density of people's flow and the walking speed with reference to statistics about people's flow that you have trained from the videos of the monitoring cameras at the station in the morning.”
[0270] <6-7. Varied Simulation Data> As described in each of the above examples, it is possible to generate the metadata as the natural simulation data MD by inputting an interactive prompt to the large-scale language model M.
[0271] Note that the simulation data MD created focusing only on “naturalness” tends to have little change.
[0272] For example, consider a monitoring camera in a warehouse. Since neither a human nor an animal enters the warehouse, the change in the image becomes poor. A change in brightness in the morning, the daytime, and the evening is conceivable, but there is no change in brightness in the warehouse without a window. Since loading and unloading are performed only in a limited time zone, there is no change in the captured video in most of the time zones.
[0273] Meanwhile, the purpose of generating the simulation data MD is to develop or verify the algorithm that processes the metadata obtained as a result of recognition of the monitoring camera. Data with a change is necessary for the development and verification of the algorithm.
[0274] Therefore, it is conceivable to extract only the time zone with a change from the metadata once created in pursuit of naturalness. The large-scale language model M is also good at extracting such data. The large-scale language model M is originally good at a task of summarizing or clipping an important part.
[0275] Some examples of prompts input by the developer to extract the time zone with a change will be given.
[0276] A first example is a prompt to instruct clipping with a keyword of “the time zone with a change”. Specifically, the developer inputs a sentence such as “From a series of natural metadata created in the steps so far, please cut five scenes in a length that does not exceed three minutes in the time zone in which the bounding box moves or changes.”, as the second prompt P2.
[0277] The second example is an example in which passage of time per second of playback time is made variable depending on a scene. Specifically, the developer inputs a sentence such as “Temporally compress the series of natural metadata created in the steps so far. Therefore, please shorten the time so that the playback speed becomes five times in the time zone in which the position and size of the bounding box do not change. On the other hand, please do not change the playback speed in the time zone in which the position and size of the bounding box change. Please compress the length so that the entire playback time is about 15 minutes.”, as the second prompt P2.
[0278] By inputting such a second prompt P2, it is possible to acquire data having high density of change suitable for the development of the algorithm as the simulation data MD while the scene is natural.
[0279] Furthermore, there is also a method of specifying the way of changing a scene in a natural language and causing the large-scale language model M to generate the way of changing a scene.
[0280] For example, the developer inputs a sentence such as “Please add a scene of loading and unloading to the metadata of the video of the monitoring camera in the warehouse created in the steps so far. Please add a scene where a person picks up a load on a carriage and taking out the load and a scene where an automatic conveyance robot picks up a load.”, as the second prompt P2. Thereby, the developer can also specify a normal loading / unloading scene.
[0281] Furthermore, the developer can specify the specific action of the subject by inputting a sentence such as “Add a scene where a person puts a small package in a pocket and takes out the small package to the metadata of the video of the monitoring camera in the warehouse created in the steps so far.”, as the second prompt P2.
[0282] <6-8. Simulation Data Generated over Long Period of Time> In each of the above-described examples, examples of obtaining the desired simulation data MD by repeating the input of the prompt for a relatively short time of about several minutes have been described.
[0283] The method of obtaining the simulation data MD using the large-scale language model M can be further developed. For example, an example in which an elapsed time from the input of the prompt to the acquisition of the simulation data MD is several days is conceivable. That is, technically, more suitable simulation data MD is obtained by “increasing a duration of a session”.
[0284] By increasing the duration of a session, the developer can input, for example, a sentence such as “Assume a house where a mother over 80 years old lives alone. A camera is installed in a living room of this house. Processing of detecting a person is performed by performing image recognition for a video of the camera. Person detection is performed at intervals of 10 seconds. Please output metadata of a person detection result, assuming that the mother spends one day. After that, please output the metadata twenty four hours later in a case where an instruction such as “Please start data generation” is given.”, as the second prompt P1.
[0285] Furthermore, the developer can further specify an imaging device or an angle of view that actually serves as a reference by inputting “Here, for a video of a camera in a living room of a house of an elderly person living alone, please refer to a video of XXX.”, as an additional prompt. Note that “XXX” is connection information of a network camera or the like that can be used by the developer without any right problem.
[0286] The developer can give an instruction to start generation of the simulation data MD by the large-scale language model M by inputting the sentence such as “Please start data generation.” as the first prompt P1 after specifying the preconditions and the like by the above prompts.
[0287] By waiting without disconnecting the session with the large-scale language model M, the developer can generate the simulation data MD as instructed by the large-scale language model M.
[0288] Note that, in the large-scale language model M, the simulation data MD may be generated using trained knowledge, but the simulation data MD for one day may be generated by analyzing a video for twenty four hours actually captured for the specified “XXX”.
[0289] According to such a mode, the developer is not able to determine whether or not the MD is generated on the basis of the knowledge acquired by the large-scale language model M or whether or not the simulation data MD is generated by performing image processing for the video of an actually installed network camera. That is, the actual method of creating the simulation data MD is in a state of being entrusted to the large-scale language model M.
[0290] Since it is sufficient for the developer to acquire appropriate simulation data MD, any method of generating the simulation data MD by the large-scale language model M is fine.
[0291] In an environment in which an agent such as the large-scale language model M or a chat system is highly developed, the input of the prompt of the developer is performed by repeating the steps described in the above examples, but the method of generating the simulation data MD by the large-scale language model M may not be recognized by the developer.
[0292] Even if the metadata is generated with the knowledge acquired by the large-scale language model M or the metadata is generated by performing image processing for an image obtained from the real world through the network camera, the second prompt P2 or the third prompt P3 is interactively input similarly to the other examples thereafter, so that the simulation data MD can be corrected to more natural and easy-to-use data as data for algorithm development (verification).
[0293] An example of a prompt input by the developer to further improve the simulation data MD obtained twenty four hours after the input of the first prompt P1 is illustrated.
[0294] For example, in a case where a sentence such as “the metadata for one day has been generated. Please specify a folder. A file is stored” is presented as a response sentence from the large-scale language model M, the developer inputs a sentence such as “Please compress a time length of the metadata before storing the file. Please delete a scene where no person is present.”, as the second prompt P2, to increase density of a change in the metadata.
[0295] Furthermore, the developer may further increase the density of a change in the metadata by inputting a sentence such as “Please delete a time zone in which an electric light is off since a person is not detectable.” or a sentence such as “Please set a time interval to one minute to increase a playback speed to six times in a scene where a person is stationary.”, as the second prompt P2.
[0296] For the developer who is a human, it is extremely difficult to give an instruction to obtain the perfect simulation data MD with one prompt input from a zero state where no instruction is given to the large-scale language model M. However, it is possible to confirm the first simulation data MD obtained as a base and notice a sense of discomfort or notice points different from the image of the developer. It is easy to divide and input a prompt for making a desired change about these sense of discomfort and different points a plurality of times, whereby the developer can finally obtain the suitable desired simulation data MD. Then, these instructions are not quantitative but qualitative, and use of natural language makes the work easy for humans.
[0297] To improve the simulation data MD as the metadata using such an interactive method, it is necessary to obtain the first metadata as the base, but the metadata may be generated with knowledge acquired by the large-scale language model M, or may be generated by performing image processing for a video obtained from a monitoring camera or the like in the real world.
[0298] In other words, it may be difficult to use the metadata generated on the basis of knowledge or the metadata obtained by performing image processing for a video of the real world as it is to develop or verify the algorithm. Therefore, as described in each example, the method of gradually correcting the simulation data MD as the metadata by interactively inputting the prompt is very effective.
[0299] <7. Implementation of Functions> The processing circuit 51 is implemented by circuitry that performs operation. That is, the circuit as the processing circuit 51 of the simulation data generation device 2 performs a predetermined operation to implement a function as the prompt acquisition unit F1, a function as the answer information acquisition unit F2, a function as the presentation processing unit F3, a function as the information provision processing unit F4, and a function as the auxiliary information acquisition unit F5.
[0300] In a case where the processing circuit 51 is a control unit implemented by an arithmetic circuit such as a central processing unit (CPU), a graphics processing unit (GPU), or a tensor processing unit (TPU), the various functions are implemented by the processing circuit 51 executing a program for performing a predetermined operation using a storage area such as various memories. That is, a predetermined program, a calculation result obtained during execution of the program, and the like are appropriately written in the storage area such as a memory provided inside or outside the processing circuit 51.
[0301] Furthermore, in a case where the processing circuit 51 is implemented by a circuit such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA), the circuit is designed and constructed in a mode of implementing a predetermined function, whereby the various functions are implemented without the necessity of the storage area such as a memory provided outside the processing circuit 51. Note that, in implementing the various functions by the ASIC, the FPGA, or the like, the memory or the like provided inside the processing circuit 51 may be used.
[0302] Note that these examples of the implementation modes can be applied to the developer terminal 1, the first server device 3, the second server device 4, and the like.
[0303] <8. Summary> As described above, the simulation data generation device 2 according to the present technology includes the prompt acquisition unit F1 that acquires, as the first prompt P1, the user input (input by the developer) that instructs generation of time-series data of the simulation data MD simulating the metadata output from the imaging device as the analysis result of the captured image, and the answer information acquisition unit F2 that inputs the first prompt P1 to the large-scale language model M and obtains the time-series data of the simulation data MD as the answer information. Furthermore, the time-series data of the simulation data MD is data that enables confirmation of the motion of the subject captured in the captured image by visualization, and the prompt acquisition unit F1 acquires the user input instructing correction of the answer information as the second prompt P2 and the answer information acquisition unit F2 inputs the second prompt P2 to the large-scale language model M to obtain the corrected time-series data of the simulation data MD as the answer information. For the development of the application using the metadata output from the imaging device, the metadata to serve as input or the simulation data MD simulating the metadata is necessary. However, in the method of actually generating the metadata, an image capture environment and a subject are necessary. In particular, in a case where the subject is a person, it is conceivable to employ an actor or the like. However, since the motion in a case where the actor acts is different from the motion of a person who is not conscious of the imaging device such as a camera, unnatural metadata is generated. Furthermore, as the method of generating the simulation data MD, it is conceivable to generate the subject of a three-dimensional model, arrange a virtual imaging device and the three-dimensional model in the space SP to be sensed, move the three-dimensional model, generate a two-dimensional image of the angle of view of the imaging device, and then output the simulation data MD simulating the metadata. However, it is difficult to cause the three-dimensional model to naturally move, and even with this method, unnatural simulation data MD is obtained, which is different from the metadata obtained in a case where a person who naturally moves is imaged. The simulation data generation device 2 having the present configuration directly generates the simulation data MD simulating the metadata without performing actual image capture and without generating a three-dimensional model. Thereby, it is possible to eliminate the unnaturalness of the simulation data MD caused by the unnaturalness of the motion of the three-dimensional model. Furthermore, it is possible to save time and effort to actually install the imaging device in the space SP to be sensed and to reduce a development cost of the application. Note that, although there is a possibility that the mode of the metadata changes if the imaging device changes, creation of a main motion portion of the application to be developed can be completed without preparing the imaging device according to the present configuration. Then, in a case where the application is made to correspond to a new type of imaging device, it is only necessary to create only a portion of an end point to serve as an interface, and thus, it is possible to improve work efficiency and reduce the number of development steps. Furthermore, the generated simulation data MD is the time-series data and can be visualized, so that the user can visually confirm the unnaturalness of the movement of the simulation data MD. Therefore, it is easy to generate the prompt (second prompt P2) for correcting the time-series data of the simulation data MD, and it is possible to contribute to the generation of more natural simulation data MD. Moreover, in a case where the metadata is obtained by programming the motion of the subject, only the metadata of a case where the motions of patterns conceived by the developer are made into the subject or the three-dimensional model can be obtained, and situations that the created application can cope with are limited, and an ability of coping is lowered. However, by inputting the first prompt P1 to cause the large-scale language model M to generate the time-series data of the simulation data MD, it is possible to obtain the simulation data MD based on the motion that is not conceivable by the user and is based on a natural motion. Therefore, it is possible to widen the variation of the input data in the creation of the application, and it is possible to create the application for enabling execution of appropriate processing according to various situations. Furthermore, since it is possible to correct the simulation data MD only by inputting the second prompt P2 for correcting the simulation data MD, it is possible to greatly reduce labor and time necessary for preparing the metadata or the simulation data MD used for developing the application. Moreover, for example, consider an application in which the metadata obtained from the imaging device that monitors a pool is input and an alert is output in a case where a person is drowning. In this case, by capturing the situation where a person is drowning by the imaging device, and preparing the metadata obtained at that time, it is possible to develop an appropriate application. However, it is not appropriate and not easy to cause a person to drown and capture an image. Furthermore, even if an actor is employed and performs a drowning state, it is not possible to reproduce a motion of a person who is actually drowning, and it is not possible to individually reproduce the way of drowning that varies depending on attributes of the person such as gender, age, and muscle mass. Also in this regard, by using the large-scale language model M actually trained on the basis of an accident video or the like regarding drowning, it is possible to generate the simulation data MD appropriately simulating the metadata output from the captured image when a person has drowned, and it is possible to contribute to improvement in performance of the application.
[0304] As described with reference to FIGS. 4, 13, and the like, in the simulation data generation device 2, the acquisition of the second prompt P2 by the prompt acquisition unit F1 and the acquisition of the corrected time-series data of the simulation data MD by the answer information acquisition unit F2 may be performed a plurality of times. Therefore, the input of the second prompt P2 and the acquisition of the answer are repeated, and the obtained corrected simulation data MD can be made more natural.
[0305] As described with reference to Figs. 3 and 4 and the like, the prompt acquisition unit F1 in the simulation data generation device 2 may acquire the third prompt P3 instructing setting of information regarding the captured image, the information being not output as the answer information, and the answer information acquisition unit F2 may obtain the time-series data of the simulation data MD by inputting the third prompt P3 together with the first prompt P1 to the large-scale language model M. For example, in a case where a person is detected and the coordinate information of the bounding box BB is output as the metadata, there are various elements that affect the motion of the person. Although there are many pieces of information not included in the output metadata, it is possible to generate more appropriate and more natural simulation data MD by providing the information to the large-scale language model M as an input.
[0306] As described with reference to Fig. 3 and the like, in the simulation data generation device 2, the information regarding the captured image may be information regarding the subject. The information regarding the subject is, for example, attribute information of the person such as gender, age, height, and weight. Furthermore, information indicating that the target person uses a wheelchair, information indicating that the target person rides on a bicycle, information indicating that the target person uses a skateboard, or information indicating that the right hand is disabled is also regarded as the information regarding the subject. By inputting these pieces of information to the large-scale language model M, it is possible to give variations to the generated simulation data MD, and it is possible to achieve high functionality of the application.
[0307] As described with reference to Fig. 3 and the like, in the simulation data generation device 2, the information regarding the subject may be information regarding the personal space of the subject. There is a concept that a person has its own personal space, and the person instinctively avoids other people from entering the personal space. Therefore, for example, it is usually unlikely that a person acts so as to invade the personal space in an uncrowded space, and by providing information regarding the personal space to the large-scale language model, it is possible to generate the simulation data MD based on a more natural human action. Therefore, it is possible to develop a high-performance application using the simulation data MD.
[0308] As described with reference to Fig. 3 and the like, in the simulation data generation device 2, the information regarding the subject may be the attribute information of the subject. The attribute information of the subject is information such as gender, age, height, weight, dominant arm, and stride length of the subject. If the attribute information of the subject is different, the motion of the subject is naturally different. Therefore, by inputting the attribute information of the subject to the large-scale language model M, it is possible to obtain more natural simulation data MD. Moreover, by setting the personal space in consideration of these pieces of attribute information, it is possible to bring the motion of each subject closer to a more natural motion. Note that the subject is not limited to a person, and may be an animal, a vehicle, food, or the like other than a person. In a case where the subject is an animal, the type, size, weight, and the like of a dog, a cat, or the like are considered as the attribute information. Furthermore, information indicating whether or not the animal is a pet or information indicating whether or not the animal is a wild animal may also be the attribute information. In a case where the subject is a vehicle, a vehicle type, a color of the vehicle, information regarding whether or not a person is on board, information regarding whether or not the vehicle is traveling or stopping, a total length, a vehicle width, a weight, and the like of the vehicle are considered as the attribute information. In a case where the subject is food, type information for roughly dividing food such as vegetables, fruits, and meat, type information slightly finely dividing food such as potatoes and leaves, and type information more finely dividing food such as cucumber and carrot, color, size, and the like are considered as the attribute information. By setting these pieces of attribute information for each subject, it is possible to impart a finer difference to the generated simulation data MD, and it is possible to appropriately create various behaviors of the application.
[0309] As described with reference to Fig. 3 and the like, in the simulation data generation device 2, the information regarding the subject may be information of direction of the subject. By requesting the large-scale language model M to set the direction of the subject therein, the time-series data of the simulation data MD output from the large-scale language model M becomes close to the metadata output in a case where the subject performs a natural motion. For example, it is quite unlikely that a person walks backward, and in a case where the person wants to travel backward, it is usually assumed that the person turns around and walks forward. Therefore, in the simulation data MD in which the direction of the subject is taken into consideration, even in a case of changing the traveling direction, the simulation data MD is data in which traveling after a motion of changing the direction of the body is taken into consideration. Therefore, it is possible to bring the simulation data MD used for developing the application closer to more natural data.
[0310] As described with reference to Fig. 3 and the like, in the simulation data generation device 2, the information regarding the captured image may be information regarding the space SP to be sensed by the imaging device. The space SP to be sensed refers to the space SP in which the imaging device is installed. Then, the information regarding the space SP to be sensed can include not only information such as the size of the space SP and geographical conditions of the space SP, but also weather information, time information, season information, and the like. Furthermore, position information of an object arranged in the space SP, for example, the type, size, position, and the like of furniture and fixtures in the case of a retail store are also the information regarding the space SP to be sensed. By inputting these pieces of information to the large-scale language model M, it is possible to generate the simulation data MD affected by these pieces of information, and it is possible to develop an application having a high ability of coping, such as an application that can be used not only in the summer but also in the winter. Note that the information for identifying the type of the space SP, for example, whether or not the space SP is the space SP in a theme park, the space SP in a stadium, or the space SP in a store is an example of the information regarding the captured image. The motion of a person walking on a road and the motion of a person walking on a road in a theme park may be different, and by inputting these pieces of information to the large-scale language model M, it is possible to generate the simulation data MD in consideration of the motion of a person according to the situation.
[0311] As described with reference to Fig. 4 and the like, in the simulation data generation device 2, the time-series data of the simulation data MD may include data related to the specific action for the subject captured in the captured image. For example, as the simulation data MD that simulates the metadata obtained from the imaging device that monitors the store, the simulation data MD corresponding not only to an action performed in a case where the customer to store is shopping but also to other specific actions, specifically, an action performed in a case where the salesclerk is preparing before opening the store is output from the large-scale language model M.
[0312] As described with reference to Fig. 4 and the like, in the simulation data generation device 2, the specific action may be an action that can be caused by an abnormal subject. That is, the specific action is not an action performed by a normal customer in a case of shopping in the store nor an action of a person swimming in a pool, but is an action of a criminal who steals in the store, an action of a person drowning in the pool, or the like, and is an unsteady action. It is difficult to obtain natural metadata for the action that can be caused by the abnormal subject even if an actor or the like acts in front of the imaging device. Therefore, by generating the simulation data MD related to these specific actions using the large-scale language model M, it is possible to generate the application capable of appropriately detecting and coping with a case where the abnormal subject actually performs the specific action.
[0313] As described with reference to Fig. 11 and the like, in the simulation data generation device 2, the simulation data MD may include data simulating the metadata in a case where an error is included as an analysis result of the captured image. For example, consider the imaging device that sets the bounding box BB to a person detected within the angle of view and outputs the coordinate information of the bounding box BB as the metadata. As a result of detecting two persons, the coordinate information of two bounding boxes BB is output from the imaging device. Thereafter, in a case where the detected two persons move to positions overlapping each other within the angle of view, only one person may be temporarily detected, and the coordinate information about one bounding box BB may be output as the metadata. In the generation of the simulation data MD, when a scene in which two persons are positioned within the angle of view is considered, even if the two persons overlap within the angle of view, the presence of the two persons is known in a simulation or the like, and thus the coordinate information of the two bounding boxes BB is output. However, in the present configuration, the simulation data MD that simulates an erroneous analysis result at a moment when one subject is not able to be recognized in the imaging device is output. Thereby, it is possible to generate the simulation data MD that appropriately and highly accurately simulates the metadata that can be output from the actual imaging device, and it is possible to contribute to development of the application that can appropriately handle the metadata output as the erroneous analysis result.
[0314] As described with reference to Figs. 3, 4, 7, and the like, the answer information acquisition unit F2 in the simulation data generation device 2 may input the auxiliary information SD used for generating the time-series data of the simulation data MD together with the first prompt P1 to the large-scale language model M. By inputting the auxiliary information SD together with the first prompt P1 to the large-scale language model M, it is possible to make the generated simulation data MD more natural.
[0315] As described with reference to Fig. 3 and the like, in the simulation data generation device 2, the auxiliary information SD may be information regarding the data format of the simulation data MD. Thereby, it is possible to generate the simulation data MD according to an output format of the imaging device actually used in the space SP to be sensed. Note that the information regarding the data format may be information by which the output format can be estimated, such as model number information or manufacturer information of the imaging device to be used.
[0316] As described with reference to Fig. 3 and the like, in the simulation data generation device 2, the auxiliary information SD may be information related to the imaging device. The information regarding the imaging device may include, for example, the model number information and the manufacturer information of the imaging device used in the space SP to be sensed, the pixel number information, the angle of view information, and the like. By using these pieces of information, it is possible to generate the simulation data MD in consideration of a difference in the metadata caused by the imaging device.
[0317] As described with reference to Fig. 3 and the like, in the simulation data generation device 2, the auxiliary information SD may be information regarding the number and positions of the imaging devices arranged in the space SP to be sensed by the imaging device. In a case where the space SP to be sensed is monitored by a plurality of imaging devices, it is possible to output, as the simulation data MD, data that simulates the metadata output from the plurality of imaging devices and that ensures consistency between the pieces of metadata in consideration of a positional relationship between the imaging devices. Therefore, it is possible to provide the simulation data MD suitable for the development of the application to be analyzed using the metadata output from the plurality of imaging devices.
[0318] As described with reference to Figs. 3 and 7 and the like, in the simulation data generation device 2, the auxiliary information SD may be a search result obtained by searching the vector database VD using the first prompt P1. The vector database VD stores various types of information expressed by vectors. Furthermore, the vector is assumed to be a high-dimensional vector including a large number of elements. That is, the vector database can be said to be a database in which the similarity as a meaning or a concept is expressed for each piece of information. Therefore, it is possible to obtain the simulation data MD suitable for a situation by inputting a searched result extracted from the vector database VD to the large-scale language model M.
[0319] As described in the modifications and the like, in the simulation data generation device 2, the auxiliary information SD may be the class information for specifying the subject to be sensed of the imaging device. Thereby, it is possible to generate the simulation data MD related to the limited subject to be sensed.
[0320] As described with reference to Fig. 3 and the like, in the simulation data generation device 2, the simulation data MD may be the coordinate information about the bounding box BB surrounding the subject to be detected in the imaging device. The bounding box BB has a rectangular shape, and the position and the size in the angle of view can be specified by the coordinate information of at least two points. Therefore, the coordinate information of the bounding box BB can be said to be simplified information. By outputting the simplified data as the simulation data MD, it is possible to reduce a burden of processing using the large-scale language model M. Furthermore, the developer can easily view the simulation data MD when the simulation data MD is visualized, and can easily issue a correction instruction of the simulation data MD.
[0321] In the simulation data generation method according to the present technology, an information processing apparatus executes: processing of acquiring, as the first prompt P1, the user input (the input by the developer) instructing generation of data that is the time-series data of the simulation data MD simulating the metadata output from the imaging device as the analysis result of the captured image and enables confirmation of the motion of the subject captured in the captured image by visualization; processing of inputting the first prompt P1 to the large-scale language model M to obtain the time-series data of the simulation data MD as the answer information; processing of acquiring, as the second prompt P2, the user input instructing correction of the answer information; and processing of inputting the second prompt P2 to the large-scale language model M to obtain the corrected time-series data of the simulation data MD as the answer information.
[0322] The simulation data generation system according to the present technology includes: a storage unit in which the large-scale language model M is stored; the prompt acquisition unit F1 configured to acquire, as the first prompt P1, the user input (the input by the developer) instructing generation of the time-series data of the simulation data MD simulating the metadata output from the imaging device as the analysis result of the captured image; the answer information acquisition unit F2 configured to input the first prompt P1 to the large-scale language model M to obtain the time-series data of the simulation data MD as the answer information; and the presentation processing unit F3 configured to present the information that visualizes the time-series data of the simulation data MD and enables confirmation of the motion of the subject captured in the captured image, in which the prompt acquisition unit F1 acquires, as the second prompt P2, the user input instructing correction of the answer information, the answer information acquisition unit F2 inputs the second prompt P2 to the large-scale language model M to acquire the corrected time-series data of the simulation data MD as the answer information, and the presentation processing unit F3 presents the information that visualizes the corrected time-series data of the simulation data MD.
[0323] The various functions and effects described above can also be obtained by such a simulation data generation method and simulation data generation system.
[0324] Note that the program for implementing the above-described functions can be recorded in advance in a hard disk drive (HDD) as a recording medium built in a device such as a computer device, a ROM in a microcomputer having a CPU, or the like. Alternatively, the program can be temporarily or permanently stored (recorded) in a removable recording medium such as a flexible disk, a compact disc read only memory (CD-ROM), a magneto optical (MO) disk, a digital versatile disc (DVD), a Blu-ray disc (registered trademark), a magnetic disk, a semiconductor memory, or a memory card. Such a removable recording medium can be provided as what is called package software. Furthermore, such a program can be installed from the removable recording medium into a personal computer or the like, or can be downloaded from a download site via a network such as a LAN or the Internet.
[0325] Note that, the effects described in the present specification are merely examples and are not limited, and other effects may be provided.
[0326] Furthermore, the above-described examples may be combined in any way, and the above-described various functions and effects can be obtained even in a case where various combinations are used.
[0327] <9. Present Technology> The present technology can also adopt the following configurations. (1) A simulation data generation device comprising: processing circuitry configured to: input a first prompt to a large-scale language model, wherein the first prompt includes instructions for generation of time-series simulation data simulating metadata of a captured image output from an imaging device, wherein the time-series data enables confirmation of motion of a subject in the captured image through visualization, obtain the time-series simulation data as answer information from the large-scale language model, input a second prompt to the large-scale language model instructing modification of the answer information, obtain modified time-series simulation data as answer information from the large-scale language model, and output the modified time-series simulation data, the output being a visualized presentation to a user. (2) The simulation data generation device according to (1), wherein the processing circuitry is further configured to perform input of the second prompt and acquisition of the modified time-series simulation data multiple times. (3) The simulation data generation device according to (1) or (2), wherein the processing circuitry is further configured to input a third prompt including an instruction for setting information regarding the captured image that is not output as the answer information. (4) The simulation data generation device according to any of (1) to (3), wherein the third prompt is input to the large-scale language model together with the first prompt. (5) The simulation data generation device according to (3) or (4), wherein the information regarding the captured image includes one or more of information regarding the subject, information regarding a personal space of the subject, or information regarding a space to be sensed by the imaging device. (6) The simulation data generation device according to (5), wherein the information regarding the subject comprises attribute information of the subject including one or more of gender, age, height, weight, dominant arm, or stride length. (7) The simulation data generation device according to any of (1) to (6), wherein the time-series simulation data includes data related to a specific action of the subject captured in the captured image, wherein the specific action is an action caused by an abnormal subject. (8) The simulation data generation device according to any of (1) to (7), wherein the time-series simulation data includes simulating metadata when an error is included as an analysis result of the captured image. (9) The simulation data generation device according to any of (1) to (8), wherein the processing circuitry is further configured to input auxiliary information for use in generating the time-series simulation data together with the first prompt to the large-scale language model. (10) The simulation data generation device according to (9), wherein the auxiliary information includes one or more of information regarding a data format of the simulation data, information regarding the imaging device, or information regarding number and positions of imaging devices arranged in a space to be sensed. (11) The simulation data generation device according to (9) or (10), wherein the auxiliary information includes a search result obtained by searching a vector database using the first prompt. (12) The simulation data generation device according to any of (1) to (11), wherein the processing circuitry is further configured to generate the simulation data to include data patterns representing one or more shoplifting behaviors including a subject concealing an item, a subject moving toward an exit without passing through a checkout area, or a subject exhibiting abnormal movement patterns indicative of theft, and output the modified time-series simulation data in a format configured for use in developing or testing an application that generates alerts when detecting the shoplifting behavior in actual imaging device metadata. (13) A simulation data generation method executed by an information processing apparatus, the method comprising: inputting a first prompt to a large-scale language model, wherein the first prompt includes instructions for generation of time-series simulation data simulating metadata of a captured image output from an imaging device, wherein the time-series data enables confirmation of motion of a subject in the captured image through visualization; obtaining the time-series simulation data as answer information from the large-scale language model; inputting a second prompt to the large-scale language model instructing modification of the answer information; obtaining modified time-series simulation data as answer information from the large-scale language model; and outputting the modified time-series simulation data, the output being a visualized presentation to a user. (14) The simulation data generation method according to (13), further comprising: performing input of the second prompt and the obtaining of the modified time-series simulation data multiple times in an interactive manner. (15) The simulation data generation method according to (13) or (14), further comprising: inputting a third prompt including an instruction for setting information regarding the captured image that is not output as the answer information. (16) The simulation data generation method according to (15), further comprising: inputting the third prompt with the first prompt to the large-scale language model. (17) The simulation data generation method according to any of (13) to (16), wherein the information regarding the captured image comprises at least one of information regarding the subject, information regarding a personal space of the subject, or information regarding a space to be sensed by the imaging device. (18) The simulation data generation method according to any of (13) to (17), further comprising: searching a vector database using the first prompt to obtain auxiliary information; and inputting the auxiliary information with the first prompt to the large-scale language model. (19) The simulation data generation method according to any of (13) to (18), further comprising: visualizing the time-series simulation data for presentation to a user; and visualizing the modified time-series simulation data for presentation to the user. (20) A non-transitory computer-readable medium storing instructions that, when executed by processing circuitry, cause the processing circuitry to perform a method, the method comprising: inputting a first prompt to a large-scale language model, wherein the first prompt includes instructions for generation of time-series simulation data simulating metadata of a captured image output from an imaging device, wherein the time-series data enables confirmation of motion of a subject in the captured image through visualization; obtaining the time-series simulation data as answer information from the large-scale language model; inputting a second prompt to the large-scale language model instructing modification of the answer information; obtaining modified time-series simulation data as answer information from the large-scale language model; and outputting the modified time-series simulation data, the output being a visualized presentation to a user.
[0328] It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and alterations may occur depending on design requirements and other factors insofar as they are within the scope of the appended claims or the equivalents thereof.
[0329] 2 Simulation data generation device BB Bounding box BB1 Bounding box F1 Prompt acquisition unit F2 Answer information acquisition unit F3 Presentation processing unit M Large-scale language model MD Simulation data P1 First prompt P2 Second prompt P3 Third prompt S2 Simulation data generation system SD Auxiliary information SP Space VD Vector database
Claims
1. A simulation data generation device comprising: processing circuitry configured to: input a first prompt to a large-scale language model, wherein the first prompt includes instructions for generation of time-series simulation data simulating metadata of a captured image output from an imaging device, wherein the time-series data enables confirmation of motion of a subject in the captured image through visualization, obtain the time-series simulation data as answer information from the large-scale language model, input a second prompt to the large-scale language model instructing modification of the answer information, obtain modified time-series simulation data as answer information from the large-scale language model, and output the modified time-series simulation data, the output being a visualized presentation to a user.
2. The simulation data generation device according to claim 1, wherein the processing circuitry is further configured to perform input of the second prompt and acquisition of the modified time-series simulation data multiple times.
3. The simulation data generation device according to claim 1, wherein the processing circuitry is further configured to input a third prompt including an instruction for setting information regarding the captured image that is not output as the answer information.
4. The simulation data generation device according to claim 3, wherein the third prompt is input to the large-scale language model together with the first prompt.
5. The simulation data generation device according to claim 3, wherein the information regarding the captured image includes one or more of information regarding the subject, information regarding a personal space of the subject, or information regarding a space to be sensed by the imaging device.
6. The simulation data generation device according to claim 5, wherein the information regarding the subject comprises attribute information of the subject including one or more of gender, age, height, weight, dominant arm, or stride length.
7. The simulation data generation device according to claim 1, wherein the time-series simulation data includes data related to a specific action of the subject captured in the captured image, wherein the specific action is an action caused by an abnormal subject.
8. The simulation data generation device according to claim 1, wherein the time-series simulation data includes simulating metadata when an error is included as an analysis result of the captured image.
9. The simulation data generation device according to claim 1, wherein the processing circuitry is further configured to input auxiliary information for use in generating the time-series simulation data together with the first prompt to the large-scale language model.
10. The simulation data generation device according to claim 9, wherein the auxiliary information includes one or more of information regarding a data format of the simulation data, information regarding the imaging device, or information regarding number and positions of imaging devices arranged in a space to be sensed.
11. The simulation data generation device according to claim 9, wherein the auxiliary information includes a search result obtained by searching a vector database using the first prompt.
12. The simulation data generation device according to claim 1, wherein the processing circuitry is further configured to generate the simulation data to include data patterns representing one or more shoplifting behaviors including a subject concealing an item, a subject moving toward an exit without passing through a checkout area, or a subject exhibiting abnormal movement patterns indicative of theft, and output the modified time-series simulation data in a format configured for use in developing or testing an application that generates alerts when detecting the shoplifting behavior in actual imaging device metadata.
13. A simulation data generation method executed by an information processing apparatus, the method comprising: inputting a first prompt to a large-scale language model, wherein the first prompt includes instructions for generation of time-series simulation data simulating metadata of a captured image output from an imaging device, wherein the time-series data enables confirmation of motion of a subject in the captured image through visualization; obtaining the time-series simulation data as answer information from the large-scale language model; inputting a second prompt to the large-scale language model instructing modification of the answer information; obtaining modified time-series simulation data as answer information from the large-scale language model; and outputting the modified time-series simulation data, the output being a visualized presentation to a user.
14. The simulation data generation method according to claim 13, further comprising: performing input of the second prompt and the obtaining of the modified time-series simulation data multiple times in an interactive manner.
15. The simulation data generation method according to claim 12, further comprising: input a third prompt including an instruction for setting information regarding the captured image that is not output as the answer information.
16. The simulation data generation method according to claim 15, further comprising: inputting the third prompt with the first prompt to the large-scale language model.
17. The simulation data generation method according to claim 13, wherein the information regarding the captured image comprises at least one of information regarding the subject, information regarding a personal space of the subject, or information regarding a space to be sensed by the imaging device.
18. The simulation data generation method according to claim 13, further comprising: searching a vector database using the first prompt to obtain auxiliary information; and inputting the auxiliary information with the first prompt to the large-scale language model.
19. The simulation data generation method according to claim 13, further comprising: visualizing the time-series simulation data for presentation to a user; and visualizing the modified time-series simulation data for presentation to the user.
20. A non-transitory computer-readable medium storing instructions that, when executed by processing circuitry, cause the processing circuitry to perform a method, the method comprising: inputting a first prompt to a large-scale language model, wherein the first prompt includes instructions for generation of time-series simulation data simulating metadata of a captured image output from an imaging device, wherein the time-series data enables confirmation of motion of a subject in the captured image through visualization; obtaining the time-series simulation data as answer information from the large-scale language model; inputting a second prompt to the large-scale language model instructing modification of the answer information; obtaining modified time-series simulation data as answer information from the large-scale language model; and outputting the modified time-series simulation data, the output being a visualized presentation to a user.
Citation Information
Patent Citations
Monitoring system, imaging apparatus, analysis apparatus and monitoring method
JP2010273125A
Machine learning based dynamic composing in enhanced standard dynamic range video (SDR+)
CN113228660A
Computer-readable recording medium storing information processing program, information processing method, and information processing apparatus
US20240321009A1