Simulated data generation device, simulated data generation method, simulated data generation system

The simulated data generation system using a large-scale language model addresses the inefficiencies of physical setups and 3D modeling by generating and iteratively correcting metadata, ensuring natural-looking data for imaging system applications.

JP2026079230APending Publication Date: 2026-05-15SONY SEMICON SOLUTIONS CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SONY SEMICON SOLUTIONS CORP
Filing Date
2024-10-30
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing methods for generating metadata for application development in imaging systems are laborious, costly, and prone to producing unnatural data due to the need for physical setups and 3D modeling, which are time-consuming and may not accurately represent natural human movements.

Method used

A simulated data generation system using a large-scale language model to generate time-series metadata without physical setups, allowing for interactive correction and visualization of simulated data to ensure natural movement representation.

Benefits of technology

Enables efficient and cost-effective generation of natural-looking metadata by iteratively correcting and visualizing simulated data, reducing the need for physical setups and 3D modeling, and improving the accuracy of human movement representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026079230000001_ABST
    Figure 2026079230000001_ABST
Patent Text Reader

Abstract

It is data that simulates metadata and generates natural-looking simulated data. [Solution] The simulated data generation device includes a prompt acquisition unit that acquires a user input as a first prompt instructing the generation of time-series data of simulated data that simulates metadata output from the imaging device as an analysis result of the captured image, and a response information acquisition unit that inputs the first prompt into a large-scale language model and obtains time-series data of simulated data as response information. The time-series data of simulated data is data that allows the movement of the subject captured in the captured image to be confirmed by visualization. The prompt acquisition unit acquires a user input as a second prompt instructing the correction of the response information, and the response information acquisition unit inputs the second prompt into a large-scale language model and obtains the corrected time-series data of simulated data as response information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present technology relates to a simulated data generation device, a simulated data generation method, and a simulated data generation system for generating simulated data that simulates metadata output from an imaging device.

Background Art

[0002] There is known a system that analyzes an event that has occurred in a space to be sensed using a monitoring camera or the like. Some of these systems use metadata output from an imaging device as a monitoring camera (see, for example, Patent Document 1 below).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] For the development of an application used in such a system, appropriate metadata is required. For development metadata, for example, an imaging device is actually placed in a specific space, and an actor is made to perform a predetermined movement within the imaging angle of the imaging device, thereby causing the imaging device to output metadata. Based on this metadata, the development of the application is advanced so that the operation of the detection target can be appropriately estimated. However, the method of actually placing an imaging device in a specific space and taking pictures has a problem that it is too laborious and costly. In addition, the performance by an actor may be different from the actions of ordinary people who are not aware of the imaging device, and there is a high possibility that unnatural metadata will be output.

[0005] Furthermore, it is conceivable to place a three-dimensional model of a person in a virtual space, program the three-dimensional model to perform predetermined actions, and then place a virtual imaging device in the virtual space, generate an image estimated to be obtained within its field of view, and output simulated data that mimics metadata. While this method has the advantage of reducing the cost of hiring actors and the time spent on filming, it has the problem of requiring a great deal of time to create 3D models and program them.

[0006] This technology was developed in light of these problems, and aims to generate natural-looking simulated data by simulating metadata. [Means for solving the problem]

[0007] The simulated data generation device according to this technology includes a prompt acquisition unit that acquires a user input as a first prompt instructing the generation of time-series data of simulated data that simulates metadata output from an imaging device as an analysis result of an image capture device, and a response information acquisition unit that inputs the first prompt into a large-scale language model and obtains the time-series data of the simulated data as response information. The time-series data of the simulated data is data that allows the movement of a subject captured in the image capture to be confirmed by visualization. The prompt acquisition unit acquires a user input as a second prompt instructing the correction of the response information, and the response information acquisition unit inputs the second prompt into the large-scale language model and obtains the corrected time-series data of the simulated data as response information. This simulated data generation device generates simulated data that mimics metadata without actually taking photographs or generating a three-dimensional model. Furthermore, the generated simulated data is presented as time-series data and is visualized. Therefore, users can visually confirm any unnatural movements in the simulated data.

[0008] The simulated data generation method of this technology involves an information processing device executing the following steps: first, obtaining a user input as a first prompt to generate time-series data of simulated data that simulates metadata output from an imaging device as a result of analyzing an image, and which allows the movement of a subject captured in the image to be confirmed by visualization; second, inputting the first prompt into a large-scale language model to obtain the time-series data of the simulated data as response information; third, obtaining a user input as a second prompt to instruct the correction of the response information; and fourth, inputting the second prompt into the large-scale language model to obtain the corrected time-series data of the simulated data as response information.

[0009] The simulated data generation system of this technology comprises a storage unit storing a large-scale language model, a prompt acquisition unit that acquires a first prompt user input instructing the generation of time-series data of simulated data that simulates metadata output from an imaging device as an analysis result of an image capture device, a response information acquisition unit that inputs the first prompt to the large-scale language model and obtains the time-series data of the simulated data as response information, and a presentation processing unit that presents information that visualizes the time-series data of the simulated data and allows the movement of a subject captured in the image capture to be confirmed. The prompt acquisition unit acquires a second prompt user input instructing the correction of the response information, the response information acquisition unit inputs the second prompt to the large-scale language model and obtains the corrected time-series data of the simulated data as response information, and the presentation processing unit presents information that visualizes the corrected time-series data of the simulated data. [Brief explanation of the drawing]

[0010] [Figure 1] This block diagram shows a schematic configuration of a development system equipped with a simulated data generation system in this embodiment. [Figure 2] This is a block diagram showing the hardware configuration of computer devices such as information processing equipment. [Figure 3] This is a block diagram showing the functional configuration of a simulated data generation device. [Figure 4] This diagram illustrates the general flow of processing performed by the developer terminal and the simulated data generation device. [Figure 5] This diagram shows an example of a dialogue screen and represents an initial screen presented to the developer. [Figure 6] This diagram shows an example of a dialogue screen, specifically the screen after the developer has entered the first prompt. [Figure 7] This diagram illustrates the specific flow of the process for setting auxiliary information. [Figure 8] This figure shows an example of visualized simulated data. [Figure 9] This figure shows another example of visualized simulated data. [Figure 10] This figure shows yet another example of visualized simulated data. [Figure 11] This diagram shows an example of a dialogue screen, illustrating a state where the developer is presented with response information obtained from a large-scale language model M. [Figure 12] This diagram shows an example of a dialogue screen, specifically the screen after the developer has entered a second prompt. [Figure 13] This is a schematic diagram illustrating an example of interaction between a simulated data generation device and a large-scale language model. [Figure 14] This figure shows the corrected simulated data visualized. [Figure 15] This diagram shows an example of a dialogue screen, specifically a screen where the first and third prompts have been entered by the developer. [Figure 16] This figure shows an example of how bounding boxes related to a subject performing a specific action are displayed. [Modes for carrying out the invention]

[0011] The embodiments of the system related to this technology will be described below in the following order, with reference to the attached drawings. <1. Development System Configuration> <2. Hardware Configuration of Each Device> <3. Functions of the Simulation Data Generation Device> <4. Process Flow> <5. Variation Example> <6. Variation> <6-1. Specific Action> <6-2. Example of Interactive Prompt> <6-3. Effective Use Example of User Database> <6-4. Setting of Information Not Output as Answer Information> <6-5. Simulation Data Suitable for Application Development> <6-6. Simulation Data Containing Unsteady but Natural Changes> <6-7. Simulation Data with Changes> <6-8. Simulation Data Generated over a Long Period of Time> <7. Realization of Functions> <8. Summary> <9. This Technology>

[0012] <1. Configuration of Development System> The configuration of the development system S1 of this technology will be described with reference to FIG. 1.

[0013] The development system S1 is a system used by developers to develop applications. The development system S1 includes a simulation data generation system S2 that generates simulation data MD used for application development, and a developer terminal 1 used by developers.

[0014] The simulation data generation system S2 includes a simulation data generation device 2, a first server device 3, and a second server device 4, which are server devices used by the simulation data generation device 2 when generating simulation data MD.

[0015] The developer terminal 1, the simulation data generation device 2, the first server device 3, and the second server device 4 included in the development system S1 are connected via a communication network NW so that they can communicate with each other.

[0016] This section describes the simulated data MD generated by the simulated data generation system S2. The applications being developed include those that analyze data output from imaging devices such as surveillance cameras and issue alerts or notifications based on the analysis results, as well as applications that visualize the analysis results.

[0017] Some of these applications analyze not only the image data output from the imaging device, but also the metadata related to that image data. Developing these applications requires metadata output from the imaging device.

[0018] However, it can be difficult to prepare appropriate metadata output from the imaging device during the application development phase. For example, an imaging device can be installed in the spatial area (SP) to be sensed by the imaging device, and images of a person acting, such as shopping, can be captured within the imaging device's field of view. This makes it possible to actually output image data and metadata from the imaging device.

[0019] However, this method incurs various costs, including the financial cost and effort of actually preparing the spatial SP and imaging equipment, the effort of securing people to act in front of the imaging equipment, and the time effort of actually taking 10 hours of footage if 10 hours of data is needed. In addition, if the spatial SP to be sensed is a spatial SP inside a store, there are cases where the construction of the store itself has not yet been completed.

[0020] Furthermore, if actors are hired to perform predetermined movements in front of the imaging device, there will be additional financial costs involved.

[0021] Furthermore, there is a problem in that a discrepancy can occur between the movements of a person acting out predetermined movements in front of an imaging device and the movements of an ordinary person who is unaware of the imaging device and is trying to achieve their own objective, resulting in the inability to obtain natural and appropriate metadata.

[0022] Another possible approach involves placing a virtual imaging device and a 3D model of a person in a virtual space, programming the movement of the 3D model, generating image data obtained within the imaging device's field of view using techniques such as ray tracing, and generating metadata including the results of the image data analysis.

[0023] However, with this method, it is difficult to make a three-dimensional model of a person move naturally through programming, and the movements of the three-dimensional model are limited to only those that a human can conceive.

[0024] Furthermore, even when using Generative Artificial Intelligence (AI) to create videos, while the quality of still images has improved, there are still many unnatural aspects to videos intended to represent movement, and it is not possible to generate metadata obtained by capturing natural human movement.

[0025] The simulated data generation system S2 is a system that takes these problems into consideration. The simulated data generation system S2 uses a large language model (LLM) M stored in the first server device 3 to generate simulated data MD that directly simulates metadata without generating image data.

[0026] As is well known, the large-scale language model M generates response information according to the content of the input prompt and is an AI (Artificial Intelligence) model capable of performing various natural language processing tasks. For example, when a question is input as a prompt, the large-scale language model M generates answer information to the question by performing data retrieval processing according to the content of the question and text generation processing according to the search results. In addition to the task of generating answer information to questions, the large-scale language model M can also perform various tasks that generate response information according to the content of the input prompt, such as creating a computer program that performs the processing specified by the prompt, or generating images or music that satisfy the conditions specified by the prompt.

[0027] Examples of large-scale language models M that can be used by the simulated data generation device 2 in this embodiment include the following:

[0028] ·GPT(Generative Pre-trained Transformer) ·BERT(Bidirectional Encoder Representations from Transformers) ·T5(Text-To-Text Transfer Transformer) · XLNet ·ERNIE(Enhanced Representation through Knowledge Integration) ·ELECTRA(Efficiently Learning an Encoder that Classifies Token Replacements Accurately) ·BLOOM(BigScience Large Open-science Open-access Multilingual Language Model) Mistral ·OPT(Open Pre-trained Transformer) ·Topher ·LaMDA(Language Model for Dialogue Applications) ·Turing-NLG (Turing Natural Language Generation) ·PaLM(Pathways Language Mode) ·Llama(Language Large Models Meta AI)

[0029] Furthermore, new large-scale language models M are emerging daily, and the large-scale language models M usable in this embodiment are not limited to these models.

[0030] Furthermore, the simulated data generation system S2 generates simulated data MD based on more natural human movement by using the user database, which is the user-side database provided by the second server device 4. The user database may be a relational database (RDB) or any other type of database. In the following explanation, a vector database VD is given as an example of a user database.

[0031] A vector database (VD) stores information with a vector associated with it. In other words, a vector database (VD) is a database that stores information in vector format. The information stored in a vector database (VD) could be, for example, a document itself.

[0032] The vectors stored in the vector database VD represent the meaning and concepts of the information they represent, and can have hundreds to thousands of dimensions. In other words, the vector database VD can be described as a database that represents the similarity of meaning and concepts of each piece of information.

[0033] The simulated data generation device 2 displays input screens, presentation screens, etc., to the developer terminal 1. The developer using the developer terminal 1 can interact with the large-scale language model M provided by the first server device 3 via the simulated data generation device 2 by viewing the screens and performing input operations on the screens.

[0034] The simulated data generation device 2 inputs the developer's input as a prompt into the large-scale language model M and obtains response information. Furthermore, the simulated data generation device 2 uses the vector database VD provided by the second server device 4 to make the response information more appropriate.

[0035] The simulated data generation device 2 visualizes the time-series data of the simulated data MD obtained from the large-scale language model M as response information and presents it to the developer terminal 1. A specific example will be described later.

[0036] <2. Hardware configuration of each device> Figure 2 shows an example of the hardware configuration of developer terminal 1, simulated data generation device 2, first server device 3, and second server device 4. Here, if developer terminal 1, simulated data generation device 2, first server device 3, and second server device 4 are not distinguished, they will be referred to as computer device Com.

[0037] The computer device Com comprises a processing circuit 51, a ROM (Read Only Memory) 52, and a RAM (Random Access Memory) 53.

[0038] The processing circuit 51, ROM 52, and RAM 53 are capable of communicating data with each other via the bus 54.

[0039] An input / output interface (I / F) 55 is further connected to bus 54.

[0040] The input / output interface 55 is connected to an input device 56, an output device 57, a storage unit 58, a communication interface 59, and a drive 60, respectively.

[0041] The processing circuit 51 is, for example, a circuit that performs various calculations, such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit). The processing circuit 51 may also consist of multiple circuits, such as one that includes both a CPU and a GPU.

[0042] The processing circuit 51 includes arithmetic circuits that perform calculations, control circuits that control these arithmetic circuits, storage circuits such as registers and caches, and an internal bus used as a data transmission path.

[0043] The processing circuit 51 executes various processes according to the program stored in the ROM 52 or the program loaded from the storage unit 58 into the RAM 53. The RAM 53 also stores data necessary for the processing circuit 51 to execute various processes as appropriate.

[0044] The input device 56 could be, for example, a pointing device 56a such as a mouse, a keyboard 56b, a camera 56c, a microphone 56d, or various other controls and operating devices such as keys, dials, touch panels, touchpads, and remote controllers. The input device 56 detects user operations and transmits a signal corresponding to the input operation to the processing circuit 51. The processing circuit 51 interprets the signal and executes the corresponding processing.

[0045] The output device 57 may be, for example, a speaker 57a or a display device 57b consisting of an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) panel.

[0046] The display device 57b is used for displaying various types of information and consists of, for example, a device provided on the casing of the computer device Com, or a separate device connected to the computer device Com.

[0047] The display device 57b displays images for various image processing tasks and videos to be processed on the display screen based on instructions from the processing circuit 51. The display device 57b also displays various operation menus, icons, messages, etc., i.e., a GUI (Graphical User Interface), based on instructions from the processing circuit 51.

[0048] For example, the display device 57b of the developer terminal 1 displays the dialogue screen G1 and confirmation screen G2, which will be described later.

[0049] The storage unit 58 consists of components such as an HDD (Hard Disk Drive) and solid memory. The storage unit 58 of the first server device 3, shown in Figure 1, stores the large-scale language model M. The storage unit 58 of the second server device 4 functions as a vector database VD by storing various types of information along with vectors. The large-scale language model M and the vector database VD only need to be present in some of the computer devices Com, and are therefore shown with dashed lines in Figure 2. Furthermore, it is not necessary for each computer device Com to possess all of the other components shown in Figure 2.

[0050] The communication interface 59 is comprised of a transmitter and receiver for wired or wireless communication between other computer devices Com. The communication interface 59 is composed of interface circuits for communication using various communication standards such as USB (Universal Serial Bus), Bluetooth (registered trademark), and Wi-Fi (registered trademark). Alternatively, the communication interface 59 may be configured as a circuit found in a NIC (Network Interface Card), modem, or Wi-Fi adapter.

[0051] The drive 60 is a device on which a removable recording medium 61, such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory, is appropriately mounted.

[0052] The drive 60 can read data files such as programs used for various processes from the installed removable recording medium 61. The read data files are stored in the storage unit 58, images contained in the data files are displayed on the display device 57b, and sound is output from the speaker 57a. Furthermore, computer programs and the like read from the removable recording medium 61 by the drive 60 are installed in the storage unit 58 as needed.

[0053] In a computer device Com having the hardware configuration described above, for example, the software for processing in this embodiment can be installed via network communication using the communication interface 59 or via the removable recording medium 61. Alternatively, the software may be pre-stored in the ROM 52 or storage unit 58, etc. In the computer device Com, the processing circuit 51 performs processing operations based on various programs, thereby executing the necessary information processing and communication processing for the developer terminal 1, the simulated data generation device 2, the first server device 3, and the second server device 4.

[0054] Furthermore, the computer device Com, such as developer terminal 1 and simulated data generation device 2, is not limited to a single device as shown in Figure 3, but may be configured as a system of multiple computer devices. These multiple computer devices may be systematized via a LAN (Local Area Network), or they may be located remotely via a VPN (Virtual Private Network) using the Internet, etc. The multiple computer devices may also include computer devices that function as a group of servers (cloud) available through cloud computing services.

[0055] <3. Functions of the simulated data generation device> The functions realized by the execution of a program by the processing circuit 51 of the simulated data generation device 2 will be explained with reference to Figure 3.

[0056] The processing circuit 51 functions as a prompt acquisition unit F1, a response information acquisition unit F2, a presentation processing unit F3, an information provision processing unit F4, and an auxiliary information acquisition unit F5 by executing a predetermined program.

[0057] The prompt acquisition unit F1 performs the process of acquiring prompts to be input to the large-scale language model M.

[0058] Specifically, the prompt acquisition unit F1 acquires a statement entered by the developer requesting the generation of simulated data MD as the first prompt P1. The prompt acquisition unit F1 also acquires a statement entered by the developer requesting the modification or correction of the generated simulated data MD as the second prompt P2.

[0059] Furthermore, the prompt acquisition unit F1 acquires a third prompt P3, which is a prompt that instructs the setting of information that is not output as response information, in order to make the simulated data MD obtained as response information from the large-scale language model M more natural. In the following explanation, the information set by the input of the third prompt P3 will be referred to as "internal parameters".

[0060] The third prompt P3 may be obtained by acquiring a statement entered by the developer, or by acquiring an instruction statement generated by the simulated data generation device 2.

[0061] Here, we will give specific examples of each prompt. In this example, the spatial SP where the imaging device is placed and the spatial SP to be sensed is the sales floor space of a retail store, the subjects to be detected are people such as customers and store employees, and the simulated metadata MD output from the imaging device is information including the coordinates of the bounding box BB related to the detected person.

[0062] The first prompt P1 could be a sentence such as, "Create metadata output from an imaging device placed in a retail store to detect people. The metadata should include the coordinate information of the bounding box of the detected person."

[0063] The second prompt, P2, is said to be a message such as, "It appears there are too few people. Please set the maximum number of people in the store to 5 and recreate the metadata."

[0064] The third prompt P3 is a statement that instructs the system to set information other than the coordinates of the bounding box BB, which is output as metadata. For example, the third prompt P3 is a statement that instructs the system to set information about the subject to be detected, the subject's personal space, or the space SP to be sensed. The information set here can be said to be information about the captured image captured by the imaging device.

[0065] Note that the setting of information that is not output as response information in response to the input of the third prompt P3 is not necessarily performed. In other words, the developer may obtain the desired simulated data MD by inputting only the first prompt P1 and the second prompt P2.

[0066] Personal space is the area one does not want others to enter, and it can be set for each individual. Generally, men tend to have a larger personal space than women, and adults tend to have a larger personal space than children. Personal space also depends on the relationship between people; it tends to be smaller between two men, two women, and larger between people of opposite sexes. Furthermore, in the case of a married couple or parent and child, personal space may be smaller than that of strangers.

[0067] Another concept similar to personal space is following distance between vehicles. Personal space is set to make the position of the bounding box BB more natural when the subject is a person. Following distance is set to make the position of the bounding box BB more natural when the subject is a vehicle.

[0068] Automatically generating vehicle bounding boxes (BB) is likely to produce data where vehicles are extremely close together. However, by setting "inter-vehicle distance" as an internal parameter, the relative positions of vehicles will maintain an appropriate distance, increasing the likelihood of generating appropriate simulated data (MD). Furthermore, since the inter-vehicle distance increases proportionally with vehicle speed, it is also effective to set "speed" information as an internal parameter. Another similar example is transport vehicles in a warehouse. Considering the footage from warehouse surveillance cameras, it is highly likely that not only people but also transport vehicles such as forklifts (AMRs: Autonomous Mobile Robots) are present inside the warehouse. In this case, a distance from the surroundings is set for transport vehicles to ensure safety. Therefore, a "safe distance" may be set as an internal parameter for transport vehicles. On the other hand, since transport vehicles in a warehouse may travel in linked formation, it may not be necessary to consider inter-vehicle distances that are proportional to speed, as is the case with vehicles.

[0069] Thus, the perceived distance between subjects varies depending on the type of subject and the situation. By taking these factors into consideration and setting internal parameters such as personal space and vehicle distance, it is possible to generate natural-looking simulated data (MD). It should be noted that the perceived distance between subjects is a three-dimensional distance, and differs from the distance on the bounding box BB image projected in two dimensions. For example, the bounding boxes BB for subjects aligned along the optical axis of the imaging device may overlap without appearing unnatural. However, the bounding boxes BB may overlap when two people pass each other in a store aisle.

[0070] Thus, even when there is a three-dimensional distance between people, the distances between bounding boxes BB on the projected image vary. Humans have the ability to perceive three-dimensional space from two-dimensional images, and can imagine what kind of three-dimensional space is projected onto the visualized two-dimensional image. Therefore, we can understand the three-dimensional positional relationship in which the overlapping bounding boxes BB occur, and judge whether or not there is any sense of incongruity.

[0071] On the other hand, since the large-scale language model M has also learned from a large amount of data, it may not have acquired the ability to grasp three-dimensional space like a human. However, it has learned empirically that "the distance between subjects in the depth direction of an image corresponds to the depth direction of the imaging device and may overlap," "the distance between subjects in the lateral direction of an image corresponds to the distance between subjects and it is natural for there to be gaps," and "the bounding box BB is larger closer to the foreground of the image." Therefore, it can be said that it has already learned the natural arrangement of the bounding box BB. Therefore, when a human instructs on the setting of personal space, the large-scale language model M can generate simulated data MD that reflects that intent.

[0072] The third prompt, P3, is a statement such as, "Assume that each customer has a specific purpose for visiting the store, and that each customer is moving around the store in order to achieve that purpose."

[0073] Information related to the subject being detected includes, for example, attribute information such as the subject's gender, age, height, weight, dominant hand, and stride length, as well as information indicating whether the person is a wheelchair user or a customer who arrived by bicycle, and information about the direction their body is facing.

[0074] By setting information related to the subject in this way, the large-scale language model M can make the changes in the coordinates of the bounding box BB based on the customer's movements more natural. In other words, the simulated data MD output from the large-scale language model M can be made more natural.

[0075] For example, by setting the direction the body is facing as the subject, the direction of change of the bounding box BB becomes more natural. Also, when the subject changes direction of movement, the motion of the subject changing its body orientation is taken into consideration, and it becomes possible to generate more natural simulated data MD, such as ensuring that the position of the bounding box BB does not change much while the subject is changing orientation.

[0076] Information about a subject's personal space can vary depending on the subject's gender, age, build, personality, race, etc.

[0077] By setting personal space, it becomes possible to reflect the movement of the bounding box (BB) coordinates in the simulated data (MD) when a customer moves along a natural path.

[0078] Information related to the spatial SP to be sensed includes the size and location of the spatial SP, as well as information such as weather, time of day, and season. Location conditions include information about other stores in the surrounding area and the density of buildings. Alternatively, information related to the spatial SP may also include information about the type of spatial SP. Information that identifies whether a spatial SP is located within a theme park, a baseball stadium, or a store is an example of information related to the spatial SP to be sensed. For example, the movement of a person walking on the street may differ from the movement of a person walking on a road within a theme park. By inputting this information into a large-scale language model M, it becomes possible to reflect situation-appropriate human movement in simulated data MD.

[0079] The third prompt, P3, is a statement that instructs the system to configure these various pieces of information. For example, it might be a statement like, "Please configure attribute information for each customer and then generate metadata."

[0080] The response information acquisition unit F2 performs the process of acquiring the output of the large-scale language model M as response information.

[0081] The presentation processing unit F3 performs the process of presenting various screens, which will be described later, to the developer. On the screens presented to the developer, the developer can input text, and response information from the large-scale language model M obtained by the response information acquisition unit F2 is presented.

[0082] The information provision processing unit F4 performs processing to provide information that helps generate simulated data MD for the large-scale language model M. Here, this information is referred to as "auxiliary information SD".

[0083] The supplementary information (SD) is, for example, information related to the data format of the simulated data (MD), such as the JSON format.

[0084] Auxiliary information SD may, for example, be information relating to an imaging device placed in spatial SP. Specifically, auxiliary information SD may include the model number and manufacturer information of the imaging device, or information such as the number of pixels and field of view. Another example of auxiliary information SD is sensor type information that can identify the type of image sensor equipped in the imaging device, such as a monochrome sensor, an RGB sensor, a color sensor such as a CMYK sensor, a distance measuring sensor that outputs a distance image, a temperature sensor that outputs a heat map image, or an EVS sensor that outputs an image consisting of event data.

[0085] The auxiliary information SD may include information on the number and location of imaging devices placed in the spatial area SP to be sensed.

[0086] The auxiliary information SD may be information obtained using the vector database VD. For example, the simulated data generation device 2 uses the first prompt P1 input by the developer to search the vector database VD and obtains the search results. The search results obtained from the vector database VD can be called the auxiliary information SD.

[0087] The information provision processing unit F4, by inputting this auxiliary information SD, makes the simulated data MD, which is the response information obtained from the large-scale language model M, more natural.

[0088] The auxiliary information acquisition unit F5 performs the process of acquiring the auxiliary information SD mentioned above. The acquisition of auxiliary information SD may be performed, for example, using the vector database VD of the second server device 4, or in response to the developer's input on the developer terminal 1.

[0089] The simulated data generation device 2 acquires the text input by the developer as the various prompts mentioned above and inputs it into the large-scale language model M to obtain the simulated data MD desired by the developer. In other words, the developer interacts with the large-scale language model M, bringing the generated simulated data MD closer to the desired result. This eliminates the need for developers to consider the perfect first prompt P1 to generate the desired simulated data MD in a single step. Furthermore, by repeatedly entering the second prompt P2 to modify the generated simulated data MD, the desired simulated data MD can be obtained efficiently and effectively.

[0090] Furthermore, by appropriately inputting the third prompt P3 and auxiliary information SD during this process, it is possible to efficiently obtain natural simulated data MD.

[0091] Here is an example of how to obtain the desired simulated data (MD) through a series of interactive exchanges.

[0092] Please note that several preconditions will be set for the explanation. The developers are requesting simulated data (MD) as metadata that simulates the movement paths of customers visiting a convenience store.

[0093] Furthermore, the developers desire this simulated MD data for verifying the algorithms to be implemented in the application under development.

[0094] Furthermore, the requested simulated data MD is a simulated data MD of surveillance camera footage taken inside a convenience store during the daytime in August. However, the vector database VD, which serves as the user database of the second server device 4, stores video data of surveillance camera footage taken inside a convenience store in the morning in February, but does not store similar video data for the daytime in August.

[0095] First, the developer receives the following input along with the first prompt P1, which states, "Using a video taken inside a convenience store in February of this year as a reference, we will create metadata showing the movement paths of people inside the convenience store. Please generate the metadata in the data format specified below, which will include the coordinates of the people's bounding boxes as output data."

[0096] { "descriptions": { { "category": "person", "bbox": { 40, 30, 100, 200}, "score": 0.80 }, { "category": "person", "bbox": { 34, 28, 90, 165}, "score": 0.60 } } }

[0097] The developer views the visualized presentation data, which is simulated data MD generated by the large-scale language model M based on the aforementioned input sentences.

[0098] Furthermore, the developers identify areas for improvement based on the presented visualization data. For example, they might notice that a large number of bounding boxes is causing queues and hindering human movement in a convenience store aisle.

[0099] The developer then felt that the video specified in the first prompt P1 was from the morning of February, and therefore the resulting data was unnatural for a simulated MD from the daytime in August.

[0100] In that case, the developer can then enter a second prompt P2, such as, "There seem to be too many people. Please limit the number of people inside the convenience store to a maximum of about 5."

[0101] As a result, the large-scale language model M generates a new simulated data MD, which is a modified version of the previous simulated data MD. The newly generated simulated data MD is visualized and presented to the developer.

[0102] The developers who reviewed the visualized data of this simulated data MD further noted that "while the queues of people inside the store had disappeared, people were still concentrated in specific areas within the convenience store, picking up items," which they found unnatural. The developers then realized that this unnaturalness stemmed from the fact that the video specified in the first prompt P1 was from the morning of February, a time when many customers were buying sandwiches, rice balls, and other items on their way to work.

[0103] In response, the developer can input a second prompt P2 such as, "Consider that customers at convenience stores often buy cold drinks and ice cream."

[0104] Furthermore, developers can enter a second prompt P2 as appropriate each time they review a newly visualized simulated data MD.

[0105] "This is a summer scene. It's hot, so please keep a little more distance between you two." "Please increase the proportion of children, taking into consideration that they don't need to attend school during summer vacation." "Please add some people who walk a little faster, or otherwise create some variation in people's walking speeds." "It's unnatural for everyone to spend the same amount of time standing in front of the shelves. Please create a variety of patterns, such as some people walking quickly and grabbing items immediately, while others walking slowly and carefully selecting their products."

[0106] By having the developer sequentially input these second prompts P2, the simulated data MD can be gradually modified to align with the developer's intentions.

[0107] Furthermore, after the coordinate changes of the people (bounding boxes BB) appearing in the visualized simulated data MD have been corrected to the desired values, the developer may enter a second prompt P2 such as, "Please change the output data format from the current JSON format to the CSV format specified below."

[0108] In addition, along with the second prompt P2, the developer will enter the following information.

[0109] { object, id, x1, y1, x2, y2 }

[0110] Here, "id" refers to a unique number assigned to a human subject located within the field of view. "x1, y1, x2, y2" indicates the coordinate information of two points diagonally opposite each other in the bounding box BB for the detected human.

[0111] As explained above, the reason why it is possible to obtain the desired simulated data MD by accumulating interactive exchanges is as follows:

[0112] It is difficult for a developer to specify all conditions perfectly from scratch. If the developer knew in advance what modifications or unnatural elements were included in the simulated data MD generated by the large-scale language model M, they could add conditions to the first prompt P1, which is the initial instruction, to prevent the simulated data MD from becoming such. However, it is difficult for the developer to know in advance what the simulated data MD generated by the large-scale language model M will be like.

[0113] Furthermore, developers can sense inconsistencies by examining the simulated MD data, even if it is incomplete, as long as it has been generated and visualized.

[0114] Furthermore, it is difficult for developers to quantitatively explain the discrepancies by using the amount of change in coordinate data included in the simulated MD data to resolve them. On the other hand, it is relatively easy for developers to qualitatively express in natural language how to correct the coordinate data included in the simulated MD data in order to resolve any inconsistencies.

[0115] Developers can compare the presented data and determine which is closer to the desired simulated data (MD). However, it is difficult for developers to explain the reasoning behind their judgment in comparing the presented data.

[0116] Furthermore, regarding the simulated data MD generated by the large-scale language model M, there is no need for developers to provide further instructions on aspects that appear natural. For example, if the number of people in a store is appropriate from the beginning for the simulated data MD generated by the large-scale language model M, the developers do not need to intentionally provide prompts regarding that number from start to finish.

[0117] Due to these characteristics of humans or large-scale language models M, developers can easily obtain simulated data MD for natural scenes, compared to when they perform simulations from scratch and have to intentionally and completely specify all conditions. In other words, developers can obtain natural simulated data MD simply by using natural language, which humans are good at, to input prompts as needed; in other words, by pointing out inconsistencies in the simulated data MD as prompts.

[0118] In the example shown above, the developer initially extracted template metadata data from a user database and provided it to the large-scale language model M. The model then repeatedly entered prompts interactively to improve the simulated data MD. However, the large-scale language model M may also use metadata templates that it has acquired without referring to other databases.

[0119] Furthermore, while the correction policy was indicated in natural language for the input of the second prompt P2, if there is relevant data in the user database or the trained data of the large-scale language model M, the same effect can be achieved by inputting a second prompt P2 such as, "Please correct the metadata using the video taken at the convenience store in xxx around xxx year x month as a reference."

[0120] <4. Processing Flow> The general flow of processing performed in Developer Terminal 1 and Simulated Data Generation Device 2 will be explained with reference to Figure 4. Note that each process shown in Figure 4 is executed by the processing circuit 51 of each device. Here, the "processing circuit 51 of Developer Terminal 1," which is the entity executing the processing, will simply be referred to as "Developer Terminal 1." The same applies to Simulated Data Generation Device 2.

[0121] In response to the developer's operation to launch the application for generating simulated data MD on developer terminal 1, developer terminal 1 launches the simulated data generation application AP in step S101. This launch process, for example, involves accessing the cloud application provided by the simulated data generation device 2 and requesting the dialogue screen G1.

[0122] In step S201, the simulated data generation device 2 performs the process of presenting the dialogue screen G1 in response to a request from the dialogue screen G1.

[0123] Accordingly, in step S102, the developer terminal 1 performs the process of displaying the dialogue screen G1.

[0124] An example of the dialogue screen G1 is shown in Figure 5. The dialogue screen G1 includes a title bar 21, a dialogue field 22, and an input field 23.

[0125] The title bar 21 displays the name of the service used by the developer and provided by the simulated data generation device 2. In this example, since the service generates and provides simulated data MD, which simulates the metadata output from the imaging device, the title "Simulated Data Generation Application" is displayed in the title bar 21.

[0126] The dialogue box 22 displays, in chronological order, the input text entered by the developer and the corresponding response text output from the simulated data generation device 2, as will be described in more detail later. Note that the dialogue box 22 may also display buttons, choices, and other elements in addition to text.

[0127] In addition, in the dialogue panel 22 of the dialogue screen G1 shown in Figure 5, the initial display shows a description of the data generated by using the application and text prompting the user to input the conditions for the simulated data MD to be generated.

[0128] Input field 23 is provided for the developer to enter text. The text being entered is displayed in input field 23. In the illustrated dialogue screen G1, buttons such as the submit button for the entered text are omitted, but such buttons may be placed in or near input field 23.

[0129] When the developer enters a sentence instructing the generation of simulated data MD into input field 23, the developer terminal 1 processes the input sentence as the first prompt P1 in step S103 of Figure 4.

[0130] Figure 6 shows an example of the interactive screen G1 during the execution of the process in step S103. As shown in the diagram, the dialogue field 22 of the dialogue screen G1 displays the sentence entered by the developer, which reads, "I want the data output from the retail store's surveillance camera." This sentence can be treated as the first prompt P1 to be input to the large-scale language model M. If the sentence displayed in the dialogue field 22 of the dialogue screen G1 is not suitable as the first prompt P1 for the large-scale language model M, the simulated data generator 2 may make appropriate modifications or additions.

[0131] The first prompt P1, entered by the developer, is acquired by the simulated data generation device 2 in step S202 of Figure 4.

[0132] In addition, the developer terminal 1 may send the sentence entered by the developer to the simulated data generation device 2 without determining whether or not it is the first prompt P1, and the simulated data generation device 2 may determine whether or not the sentence corresponds to the first prompt P1, that is, whether or not it is a sentence that requests the large-scale language model M to generate response information.

[0133] Next, the simulated data generation device 2 performs information setting processing in step S203. The information set here is the information set by inputting the aforementioned third prompt P3 into the large-scale language model M. In other words, the processing in step S203 can be described as the process of inputting the third prompt P3 to the large-scale language model M.

[0134] This allows information related to the captured image, specifically information about the subject to be detected, information about the subject's personal space, and information about the spatial area (SP) to be sensed, to be set prior to the generation of the simulated data (MD).

[0135] Next, in step S204, the simulated data generation device 2 executes the process of setting the auxiliary information SD.

[0136] Here, the specific details of the auxiliary information SD setting process are shown in Figure 7. The auxiliary information SD setting process is realized, for example, through the cooperation of the simulated data generation device 2 and the second server device 4.

[0137] In step S221, the simulated data generation device 2 sends a first prompt P1 to the second server device 4. This transmission process can be rephrased as sending a search query to the vector database VD provided by the second server device 4.

[0138] In step S401, the second server device 4 receives the first prompt P1.

[0139] In step S402, the second server device 4 performs the process of adding a vector to the statement that is the first prompt P1.

[0140] The process of adding vectors may be performed on the entire sentence as the first prompt P1, or it may be performed on each of the multiple keywords extracted from the first prompt P1.

[0141] In step S403, the second server device 4 performs a search using the vector database VD. Through this process, the second server device 4 performs a search for the search query based on the first prompt P1 and obtains the search results.

[0142] In step S404, the second server device 4 transmits the search results extracted from the vector database VD to the simulated data generation device 2.

[0143] In step S222, the simulated data generation device 2 receives search results from the second server device 4. The information received by the simulated data generation device 2 from the second server device 4 is referred to as auxiliary information SD.

[0144] Returning to the explanation of Figure 4. In step S205, the simulated data generation device 2 obtains response information from the large-scale language model M. The response information from the large-scale language model M is referred to as the simulated data MD described above. Furthermore, obtaining response information from the large-scale language model M is achieved by inputting the first prompt P1, as well as the third prompt P3 and auxiliary information SD as needed, to the large-scale language model M.

[0145] In step S206, the simulated data generation device 2 performs visualization processing. The simulated data MD is data that simulates metadata, and may be difficult for developers to understand as is. The processing in step S206 is the process of processing the simulated data MD output from the large-scale language model M to make it easier to understand. Note that the visualization processing may also be implemented by instructing the large-scale language model M to perform processing to make it easier to understand, such as graphing.

[0146] The simulated data generation device 2 obtains data such as graphs, videos, and tables as a result of the visualization process.

[0147] An example is shown in Figure 8. Note that Figure 8 is an example of a video, and shows extracted images from the nth frame and the (n+1)th frame of the video.

[0148] Each frame's image consists of an outer frame 31 that mimics the field of view of the imaging device, and multiple bounding boxes BB arranged inside the outer frame 31.

[0149] Furthermore, in order to facilitate understanding, each frame's image may be shown in a way that indicates the location of furniture, obstacles, walls, windows, doors, etc., that are placed within the imaging device's field of view. For example, in the example shown in Figure 9, the developer is presented with a video showing not only the bounding box BB but also the furniture 32 placed within the imaging device's field of view.

[0150] Another example of data obtained through visualization is shown in Figure 10. Figure 10 is a graph showing the count results of the number of subjects within the field of view, and is considered an example of confirmation screen G2 for developers to check the results of the visualization process.

[0151] On confirmation screen G2, as shown in the diagram, the number of people (or average number of people) within the field of view at predetermined time intervals, such as 10 minutes, is represented in a graph through visualization processing. By checking this graph, developers can confirm the recognition results of the imaging device, and consequently, an overview of the simulated data MD.

[0152] Returning to the explanation of Figure 4. In step S207, the simulated data generation device 2 performs a presentation process. In this process, it notifies the developer on the interactive screen G1 that the simulated data MD has been generated, and presents the developer with a means to download the various generated data.

[0153] Figure 11 shows an example of the dialogue screen G1 presented to the developer through the presentation process. The interactive screen G1 presented to the developer by the processing in step S207 displays an overview of various pieces of information set by the third prompt P3 that are not output as simulated data MD. Specifically, it notifies the developer that the size of the store, the number of cameras, and attribute information for each customer have been set when generating the simulated data MD.

[0154] Furthermore, while the imaging device actually installed in the spatial SP recognizes the subject captured in the image for each imaging frame, the recognition results are not always correct.

[0155] For example, consider the process of detecting a person and setting a bounding box (BB). Even if two people were detected up to the frame immediately preceding a certain frame, it is possible that in that frame, the two people overlap so that they are aligned along the optical axis of the imaging device, and only one person can be detected. In this case, in an actual imaging device, one bounding box (BB) would be set as the detection result. However, in a simulation, since the presence of two people is known, two bounding boxes (BB) are set to overlap.

[0156] In this example, considering that misrecognition in such imaging devices can occur with a certain probability, the developers have been notified that the large-scale language model M has been instructed to be configured to produce incorrect recognition results with a certain probability. This allows us to obtain simulated data MD that mimics two people being mistakenly identified as one person with a certain probability. In other words, it becomes possible to obtain simulated data MD that appropriately simulates the metadata output from an actual imaging device.

[0157] Furthermore, various other factors could be contributing to false positives. For example, there are cases where some subjects cannot be detected due to backlighting, or where subjects that repeatedly move in and out of the field of view at the edge of the imaging device's field of view cannot be detected, or in vehicle detection, occlusion occurs when a passenger car overtakes a bus, making detection impossible, or the passenger car and bus are detected as a single vehicle.

[0158] It is possible to obtain simulated data (MD) that appropriately simulates cases where such false positives occur.

[0159] Furthermore, actions that are not normally expected may occur within the field of view of the imaging device. For example, shoplifting in a retail store. Other examples include the behavior of a vehicle entering a store due to mistaking the accelerator for the brake, or drowning in a river. These actions can also be described as actions that an abnormal subject might perform.

[0160] In the following explanation, these actions that are not normally anticipated will be referred to as "specific actions." Note that "unanticipated" means actions that are not anticipated by the subject performing the action or by the object receiving the action, such as a store clerk or manager; it does not mean that they are not anticipated for surveillance purposes. Specifically, while shoplifting instead of shopping is not something a normal person would anticipate, it is an action that can be anticipated for the purpose of surveillance using an imaging device.

[0161] The developers are notified that they have been instructed to generate simulated data MD related to metadata that would be output from the imaging device in the event of such "specific actions" occurring without their explicit instructions. These instructions are implemented by inputting predetermined prompts into the large-scale language model M in step S205. Specifically, this is achieved by inputting sentences such as "Please also generate simulated data for the case where there is a person shoplifting," "Please also generate simulated data considering the case where a vehicle has entered the store," and "Please generate simulated data for the case where there is a person drowning in the river" into the large-scale language model M along with the first prompt P1 and the third prompt P3.

[0162] Thus, in the presentation process of step S207, not only is the simulated data MD obtained from the large-scale language model M presented to the developer, but various conditions and settings used to generate the simulated data MD may also be notified as appropriate.

[0163] In the example shown in Figure 11, the simulated data MD is not presented directly, but rather by presenting a button Btn1 with a link to download the simulated data MD. Similarly, the visualized data is presented by presenting a button Btn2 with a link to download a video file.

[0164] Returning to the explanation of Figure 4. Developer terminal 1 performs the display process of simulated data MD as appropriate in step S104 based on the developer's operation.

[0165] For example, when a developer operates button Btn2, video data visualizing the simulated data MD is downloaded, and the downloaded video data is played back.

[0166] Developers can check the contents of the simulated data MD by playing back the downloaded video data and make corrections to the simulated data MD as needed. When making corrections, developers enter the text instructing the corrections in input field 23 of the dialogue screen G1.

[0167] When the developer enters a sentence instructing the modification of the simulated data MD into the input field 23, the developer terminal 1 processes the input sentence as the second prompt P2 in step S105 of Figure 4. The determination of whether the input sentence is the second prompt P2 may be performed by the developer terminal 1 or by the simulated data generation device 2.

[0168] There are many possible sentences for the second prompt P2. For example, as shown in Figure 12, examples include sentences such as "The number of people seems low. Please recreate the metadata with a maximum of 5 people in the store," "Please make personal space more diverse," "Please consider the behavior of customers when they view product advertisements posted on the wall and purchase those products," "Please set the number of customers visiting the store taking into account the time of day," and "Please consider customers who come in the summer just to cool off without the intention of purchasing anything." The examples are endless.

[0169] In response to the developer terminal 1 receiving input for the second prompt P2 in step S105 of Figure 4 and transmitting it to the simulated data generation device 2, the simulated data generation device 2 performs the process of acquiring the second prompt P2 in step S208.

[0170] Next, in step S209, the simulated data generation device 2 obtains response information from the large-scale language model M. Obtaining response information from the large-scale language model M is achieved by inputting a second prompt P2 to the large-scale language model M. Additionally, additional auxiliary information SD may be input to the large-scale language model M as needed, or prompts indicating additional settings may be input.

[0171] The process in step S209 is the same as the process in step S205, and involves inputting prompts to the large-scale language model M to obtain response information.

[0172] After step S209, steps S206 and S207 are performed in the simulated data generation device 2. Accordingly, steps S104 and S105 are performed in the developer terminal 1.

[0173] In other words, in the developer terminal 1 and the simulated data generation device 2, each of the processes in steps S205, S206, S207, S104, S105, and S208 may be repeated multiple times.

[0174] Next, the prompts input from the simulated data generation device 2 to the large-scale language model M stored in the first server device 3, and the response information output from the large-scale language model M to the simulated data generation device 2, will be explained with reference to Figure 13. In Figure 13, the information is presented in chronological order from top to bottom.

[0175] For the large-scale language model M, first, the first prompt P1 and the third prompt P3 The auxiliary information SD is then entered.

[0176] The large-scale language model M outputs simulated data MD as response information.

[0177] Furthermore, when a second prompt P2 is input to the large-scale language model M, the large-scale language model M outputs the modified simulated data MD as response information.

[0178] The input for the second prompt P2 and the output of the modified simulated data MD can be performed multiple times.

[0179] Figure 14 shows an example of confirmation screen G2, which is used to check the information generated as a result of visualizing the corrected simulated data MD.

[0180] The example confirmation screen G2 shown in Figure 14 is a screen example for checking the visualization of the graph generated as a result of improvements made to the example shown in Figure 10, which graphed the simulated data MD before correction. Furthermore, the correction in this example was made in response to the input of the sentence "Please set the number of customers visiting the store taking the time of day into consideration." as the second prompt P2 into the large-scale language model M.

[0181] As shown in each figure, in the graph shown in Figure 10, the detected number of people showed no bias regardless of the time of day, and the pattern of increase and decrease was uniform, whereas in the graph shown in Figure 14, the maximum value of the detected number of people differed depending on the time of day, resulting in a bias in the number of people. In other words, the corrected simulated data MD takes into account the changes in the level of store congestion depending on the time of day, resulting in a more natural simulated data MD.

[0182] <5. Variation> The example shown illustrates how the developer's input to the third prompt P3 involves entering a sentence that provides instructions for the large-scale language model M. Specifically, the developer might enter a sentence such as, "Please generate simulated data after setting attribute information for each customer."

[0183] This is not the only option; developers may simply input their evaluation of the simulated data MD generated by the large-scale language model M. The evaluations entered by the developer are used by the simulated data generator 2 to generate prompts as appropriate. For example, if the developer enters "The variation in the number of customers depending on the time of day is small," the simulated data generator 2 may generate a prompt as the third prompt P3, "Please generate simulated data MD considering the variation in the number of customers depending on the time of day," and input this into the large-scale language model M in step S209 of Figure 4.

[0184] In the example shown in Figure 6, the developer only entered the statement corresponding to the first prompt P1. However, as shown in Figure 15, the developer may also enter a statement specifying auxiliary information SD along with the statement corresponding to the first prompt P1, or a statement corresponding to the third prompt P3.

[0185] In the example shown in Figure 15, the developer specifies that the detection target is "customers." By specifying the class of the detection target in this way, the simulated data (MD) desired by the developer can be appropriately generated. The class information that specifies the class of the detection target represents the category of the subject, such as distinguishing between "people," "cars," "airplanes," "ships," "trucks," "birds," "cats," "dogs," "deer," "frogs," and "horses." In other words, the class information that identifies the class of the subject to be detected may also be the auxiliary information (SD) mentioned above.

[0186] Figures 8 and 9 illustrate how to show developers an image in which multiple bounding boxes BB are placed inside an outer frame 31 that mimics the field of view of an imaging device. The bounding box BB related to the aforementioned unforeseen behavior may be presented to developers in a modified manner.

[0187] For example, Figure 16 shows a frame in a video presentation of simulated data MD, which simulates metadata output from an imaging device that captures tourists playing in a river. In other words, it simulates a state where a number of tourists equal to the number of bounding boxes BB are detected. In Figure 16, bounding box BB1, which corresponds to the detection result of a drowning person, is shown with a thicker line than the other bounding boxes BB.

[0188] In addition to the above, you may also indicate by changing the line type or line color of the bounding box BB for individuals who have performed specific actions that are not normally expected.

[0189] Furthermore, in the display on the video, the frames before drowning may be displayed in the same manner as other bounding box BBs, while the frames from the moment of drowning onwards may be displayed in a different manner than other bounding box BBs.

[0190] These embodiments enable developers to recognize the bounding box (BB) of a person performing a specific action, making it easier to determine whether the movement of the bounding box (BB) is natural or not.

[0191] It should be noted that these changes to the bounding box BB line type and color are merely a technique to streamline the interactive modification of simulated data MD between the large-scale language model M and the developers, and that the metadata for the final output simulated data MD can be the usual bounding box. In other words, it is possible to add a flag to the metadata indicating a specific action, but it is not required.

[0192] <6. Variations> <6-1. Specific actions> In addition to those mentioned above, various other types of specific behaviors that are not normally anticipated are also possible. For example, certain behaviors can be included in actions performed unconsciously or natural movements. Specifically, various examples include "a person swaying their body while standing," "a person standing with their weight shifted to their right foot instead of standing straight," "a person alternating between having their weight on their right foot and their left foot," "a person sitting and writing, looking to the upper left when thinking about something, or to the lower right when trying to remember something," and "a person readjusting their posture in a chair."

[0193] Furthermore, although it may seem surprising, certain behaviors can also be found among the movements we commonly observe. Specifically, various scenarios are possible, such as "a person tripping," "a person stopping because their clothes get caught on an obstacle," "a person meeting an acquaintance and starting a conversation," "a vehicle suddenly parking in front of the camera," "a vehicle entering an intersection immediately after the traffic light changes from yellow to red," and "a vehicle turning left in a no-left-turn zone."

[0194] Furthermore, even among unusual movements, specific behaviors may be included. Specifically, if the subject is a "person," possible scenarios include "being involved in an accident," "fighting on the street," or "two subjects bumping shoulders." If the subject is a "vehicle," possible scenarios include "colliding with another vehicle or obstacle," "hitting a person," or "aggressive driving." Furthermore, if the subject is a "dog," possible reasons include "the dog is not around its owner" or "it bites people." These are just a few examples.

[0195] <6-2. Examples of interactive prompts> This section explains examples of inputs for various prompts.

[0196] The first prompt P1 is, for example, a statement that specifies a file after it has been attached, which is formed in the sample data format. Specifically, the first prompt P1 is entered as follows: "aaa1234.json is metadata that outputs the coordinates of the bounding box BB surrounding a person inside a convenience store. Please create metadata with a time length of 20 minutes in the same data format."

[0197] As explained earlier, the second prompt P2 can occur multiple times, but here we will show an example where the input for the second prompt P2 is given three times. The first second prompt P2 is the message, "It appears there are too few people. Please recreate the metadata with a maximum of 5 people in the store."

[0198] The second prompt, P2, reads: "Human movement is monotonous. Create metadata that shows a variety of people, such as those who stop to look for and pick up items from the shelves, and those who leave immediately because they don't find what they want on the shelves."

[0199] The third second prompt, P2, is the sentence, "Create metadata assuming that of the five people, one is a child and two are female."

[0200] By repeatedly inputting the second prompt P2 in this way, the generated simulated data MD can be gradually improved to match the desired result. Furthermore, the generated simulated data MD may be visualized each time it is modified. This allows the developer to identify inconsistencies and areas for improvement in the presented simulated data MD, enabling them to input the second prompt P2 for correction. Thus, the generated simulated data MD can be gradually improved to become more natural.

[0201] Furthermore, in such interactive exchanges, it is difficult to input a second prompt P2 that explicitly indicates an implicit assumption. For example, there are infinitely many second prompts P2 that indicate an implicit assumption, such as "Please ensure that humans do not float in the air," and it is difficult to input all of them. However, the large-scale language model M naturally learns implicit assumptions during its training process, such as the fact that humans do not float in the air, and in most cases, a second prompt P2 to instruct such corrections is unnecessary. In other words, developers can obtain natural simulated data MD with simple instructions using natural language.

[0202] Here are a few more examples of variations of the second prompt P2.

[0203] "The camera used to film inside the convenience store appears to be positioned too low. Please try setting the camera to a slightly higher position and taking the footage again." "The data format has been improved. Please increase the data length from 20 minutes to 30 minutes." • "Please change the metadata output interval from 10 seconds to 5 seconds. Note that the amount of movement on the screen will be smaller because the human movement speed will not change." "In convenience stores, shelves are arranged vertically. Please redesign the movement paths, taking into account that people can only move horizontally through the back or front aisles."

[0204] These second prompts P2 are entered to make further corrections if there are any inconsistencies in the presented simulated data MD.

[0205] Furthermore, a second prompt P2 could be considered not to resolve the discrepancy, but to generate new simulated data MD. That is, the second prompt P2 exemplified below is an example of a second prompt P2 for generating new simulated data MD by changing the conditions.

[0206] "Please change the data from the convenience store to the surveillance camera footage from the office building lobby." "Please double the depth of the convenience store."

[0207] <6-3. Examples of Effective Use of User Databases> The user database provided by the second server device 4 will now be described. The vector database VD, which serves as the user database mentioned above, has the function of searching for data similar to the first prompt P1 entered by the developer.

[0208] For example, in response to the first prompt P1, "Create data on human movement trajectories within a convenience store," past cases are searched using a vector database VD, and the search results are input into a large-scale language model M along with the first prompt P1.

[0209] This allows the large-scale language model M to present simulated data MD based on past examples.

[0210] The user database provided by the second server device 4 is not limited to a vector database VD, but may be other types of databases such as RDBs. In other words, the user database provided by the second server device 4 can be considered a database that stores auxiliary information SD (hint information) that can be provided to the large-scale language model M.

[0211] For example, a developer can search for the data format of metadata they have created in the past and provide it to a large-scale language model M as supplementary information (SD). The metadata output from imaging devices is typically in JSON or CSV format. For example, the JSON data format is complex and difficult to express in natural language. In other words, it is difficult for developers to explain and provide the JSON format in appropriate language.

[0212] In such cases, developers can assist the large-scale language model M in generating MDs by obtaining metadata from past JSON formats and providing it as auxiliary information SD. Since JSON format metadata is not generally publicly available, the large-scale language model M does not possess knowledge of it. However, by utilizing the user database provided by the second server device 4, if the developer can retrieve JSON format metadata acquired in the past, it becomes possible to provide the retrieved JSON format metadata itself, along with the first prompt P1, to the large-scale language model M. This allows the large-scale language model M to present the developer with simulated data MD created in the appropriate JSON format.

[0213] Alternatively, instead of inputting JSON-format metadata as auxiliary information SD into the large-scale language model M along with the first prompt P1, the auxiliary information SD may be input into the large-scale language model M in response to the second prompt P2. In other words, the user database has the function of searching for data similar to the second prompt P2 and third prompt P3 that the developer inputs.

[0214] For example, after obtaining simulated data MD in a format other than JSON from the large-scale language model M by entering the first prompt P1, if the second prompt P2 is entered as "Please generate metadata in the same JSON format as the person detection in past AAA projects," the second server device 4, upon receiving this input, can use its user database to obtain past metadata in JSON format and input it into the large-scale language model M along with the second prompt P2.

[0215] Similarly, by entering "Generate metadata in the CSV format used for person detection in past BBB projects" as the second prompt P2, the second server device 4 can use its user database to retrieve past metadata in CSV format and input it into the large-scale language model M along with the second prompt P2.

[0216] Furthermore, developers can further enhance the simulated MD data based on various formats.

[0217] Specifically, by entering the second prompt P2, which states, "Add posture information for each bounding box BB to the JSON format used for person detection in past AAA projects. Add the posture information to the JSON format as a posture variable, such as {"posture":"standing"} or {"posture":"sitting"}," you can enrich the simulated data MD by including more information.

[0218] Furthermore, developers can modify the conditions of past projects to generate the desired simulated MD data. These changes and modifications can be implemented interactively by repeatedly entering various prompts and acquiring simulated data (MD).

[0219] Furthermore, it is possible to obtain metadata for simulated data (MD) not as a JSON-formatted text file, but as serialized data encoded in Base64 format.

[0220] For example, a developer can obtain simulated data MD in the desired format by entering a prompt such as, "Serialize and encode the simulated data MD using the same method used in the CCC project."

[0221] <6-4. Settings for information that will not be output as response information> This section explains the variations of the third prompt, P3.

[0222] As mentioned earlier, the third prompt P3 is a prompt that instructs the user to set information that is not output as response information from the large-scale language model M, and is a prompt for obtaining more natural simulated data MD.

[0223] In other words, internal parameters are information used to generate natural simulated data MD, and are not included in the simulated data MD output from the large-scale language model M.

[0224] Here are some examples of internal parameters.

[0225] The first example of an internal parameter is a parameter related to mass or orientation. For example, consider internal parameters to make the vehicle's movement trajectory more natural. Because trucks and buses have a large mass, the energy required for acceleration and deceleration is greater than that for passenger cars. Therefore, sudden stops due to sudden braking are less likely, and their acceleration and deceleration are smaller compared to passenger cars.

[0226] Furthermore, if the expected metadata is a bounding box (BB), then the difference between trucks and buses does not need to be considered, and the simulated data (MD) can be generated by dividing the vehicles as subjects into two classes: passenger cars and buses.

[0227] In other words, the developer enters a third prompt P3, which means, "Please set the vehicles to be detected to two classes: passenger cars and buses."

[0228] Alternatively, you may enter the second prompt P2 in addition to the third prompt P3. For example, the developer might enter the following sentence as the second prompt P2: "A bus has twice the mass of a passenger car. Its acceleration is roughly half."

[0229] This allows the large-scale language model M to generate simulated data MD that includes both heavy and light vehicles, and whose movements are natural.

[0230] Furthermore, developers can make the modified simulated data MD more natural by inputting a second prompt P2, which states, "Buses and passenger cars have a direction. They move forward 90% of the time and turn right or left 10% of the time. Reverse is not considered."

[0231] The second example of an internal parameter is a parameter related to attributes such as gender (female or male) or age (child or adult).

[0232] These internal parameters are useful when creating movement trajectories for person detection.

[0233] Generally, men are more likely than women to have a larger bounding box (BB), and adults tend to have a larger bounding box (BB) than children.

[0234] Therefore, developers can obtain natural simulated MD data that includes both adults and children by entering a third prompt P3 such as, "Please divide the people to be detected into two classes: adults and children." It is also possible to enter a third prompt P3 along with a location.

[0235] For example, by inputting a third prompt P3 such as, "The camera is being installed in a children's recreation facility, so please include more children," the developer can cause the large-scale language model M to generate appropriate simulated data MD tailored to the location.

[0236] Furthermore, developers may input information about a character as a third prompt P3, such as "carrying heavy luggage," "carrying large luggage," or "being injured and therefore able to walk slowly."

[0237] Furthermore, for heavy luggage or containers, by entering a third prompt P3 to set the "weight of the contents inside," it is possible to generate simulated data MD that takes into account the ease of movement, mobility, and vibration of the detected subject.

[0238] Thus, the third prompt P3 can be used in various ways, such as to express differences in the subjects being detected.

[0239] The third example of an internal parameter is related to personal space, as mentioned earlier. This internal parameter is intended to make the spatial relationships between multiple people appear natural.

[0240] By inputting the third prompt P3 such as "Set a personal space for each detected person.", the developer can obtain simulation data MD that takes into account the personal space.

[0241] Also, the developer may input the third prompt P3 such as "Set the distance of the personal space for each bounding box and generate data where adjacent bounding boxes move while maintaining the personal space." More specifically. After inputting a prompt requesting the setting of the personal space once, it can also be regarded as the second prompt P2, which is a prompt for modifying the setting of the personal space.

[0242] Also, the developer may input a sentence such as "When taking a product in front of the shelf, please modify the data so that the distance of the personal space becomes smaller." as the second prompt P2.

[0243] By inputting these prompts, natural simulation data MD that takes into account the appropriate width of the personal space according to the situation can be obtained.

[0244] An example of the fourth internal parameter is about human habits.

[0245] For example, there are individual habits in the way of doing sports such as running. Taking running as an example, there are people who start from the right foot and people who start from the left foot. Also, when taking a break while running, there are people who stop and look down to recover their strength, and there are people who recover their strength while walking slowly.

[0246] Also, for a person sitting on a chair, there are people who sit cross-legged, and there are also people whose knees shake like jiggling (foot shaking).

[0247] These individual differences in movement are not random occurrences but rather personal traits. That is, if a person is sitting cross-legged in a chair, they are likely to move similarly again after walking around and then sitting down again.

[0248] Therefore, by having the developer input these personal quirks for each person as a third prompt P3, the large-scale language model M can generate simulated data MD that simulates the movement of bounding boxes BB, which appropriately reflects the individuality.

[0249] <6-5. Simulated data suitable for application development> One case in which simulated data MD generated by a large-scale language model M is used in application development is when this simulated data MD is used to verify the application's algorithms.

[0250] This section describes an example of prompts for generating appropriate simulated MD data for algorithm verification.

[0251] A large-scale language model M, having already learned software testing techniques, can generate simulated data MD suitable for testing. Specifically, the large-scale language model M is assumed to have learned concepts such as boundary value testing and equivalence partition testing.

[0252] For such a large-scale language model M, the developer enters the following as the first prompt P1: "Use software testing techniques to create metadata for testing. First, create metadata so that boundary value testing and equivalence partitioning testing can be performed for each of the four corner coordinates of the bounding box."

[0253] This allows developers to obtain simulated MD data with software testing in mind.

[0254] Furthermore, developers can obtain more precisely-intentioned simulated data MD by entering a second prompt P2 such as, "Set the boundaries to 0 and 10 for the X and Y axes respectively," or "Create two more bounding boxes."

[0255] In other words, the second prompt P2, instead of including multiple instructions at once, breaks them down into individual instructions, allowing the large-scale language model M to generate simulated data MD that aligns with the developer's intentions.

[0256] Another example of prompts for generating appropriate simulated MD data for algorithm verification is described.

[0257] By using simulated MD data for testing purposes, which may be outside the developer's original thinking, in algorithm verification, it is possible to further improve the completeness of the application.

[0258] To obtain simulated data MD that they could not conceive themselves, the developer enters a first prompt P1, such as, "Remove the physical constraints on the coordinates of the four corners of the bounding box to create test data that would not normally occur."

[0259] This allows developers to obtain simulated data MDs that are not normally expected, such as simulated data MDs where the left-right or up-down coordinates of the four corners of the bounding box BB are swapped, simulated data MDs where the width or height of the bounding box BB changes abruptly, or simulated data MDs where some of the coordinates of the four corners of the bounding box BB are missing.

[0260] Furthermore, instead of generating simulated data MD with a single input of the first prompt P1, the desired simulated data MD may be obtained by inputting a preparatory prompt before the first prompt P1.

[0261] For example, the developer first inputs a preparatory prompt such as "Please come up with 10 test cases that are normally impossible for the metadata containing the coordinates of the bounding box."

[0262] Subsequently, the developer inputs the first prompt P1 such as "For each of the 10 cases above, please create metadata containing the actual coordinate information of the bounding box as test data that is normally impossible."

[0263] By using such a method, it is possible to prevent the generation of simulated data MD that humans are prone to creating cases based on their own knowledge and the occurrence of omissions or leaks in the cases, and to create highly comprehensive test data.

[0264] <6-6. Simulated data including non-steady but natural changes> The large language model M is trained using various videos available via communication networks such as the Internet. And those videos naturally include the images of surveillance cameras installed in various locations.

[0265] However, most of the images of surveillance cameras are likely not suitable for the development of metadata analysis algorithms in applications. That is because the images of surveillance cameras tend to be images with little change in the viewing angle, such as images of passers-by moving left and right and passing through.

[0266] Also, since the shooting positions (camera positions) of surveillance cameras and the surrounding environments are different for each monitored space, each image is positioned as a different new learning material, and a large number of images with little change are used for learning. Also, since the data desired by the developer is the simulated data MD itself as metadata rather than images, these large amounts of images with little change are not essential.

[0267] It should be noted that these videos are not entirely useless; they are extremely useful as statistical data showing human movement trends. For example, footage from surveillance cameras installed in train stations shows differences in passenger density and walking direction between morning, noon, and evening, as well as differences in passenger density, distance between people, and walking speed between rainy and sunny days. Using a large-scale language model M constructed by learning from such a large amount of statistical video data is more likely to yield more natural and appropriate simulated data MD than setting statistical parameters based on human imagination.

[0268] In other words, by stepping through prompts such as the first prompt P1, "Generate metadata that includes the coordinates of bounding boxes for passengers photographed at the ticket gates of train stations in the morning," and the second prompt P2, "Generate the density and walking speed of people based on the statistics you learned about people flow from surveillance camera footage at train stations in the morning," the developer can generate simulated data MD that incorporates the natural movements of passengers based on past footage used for training by the large-scale language model M.

[0269] <6-7. Simulated data with variations> As explained in the examples above, metadata can be generated as natural simulated data MD by inputting interactive prompts to a large-scale language model M.

[0270] However, simulated data MDs created with only "naturalness" as the focus tend to show little change.

[0271] For example, consider a security camera inside a warehouse. Since very few people or animals enter the warehouse, the changes in the footage will be minimal. While changes in brightness between morning, noon, and night are possible, if the warehouse has no windows, even these changes will not occur. Since loading and unloading of goods only takes place during limited time slots, the captured footage will remain unchanged for most of the time.

[0272] On the other hand, the purpose of generating simulated data (MD) is to develop or verify algorithms for processing metadata obtained as a result of surveillance camera recognition. Developing and verifying these algorithms requires data that is constantly changing.

[0273] Therefore, it is conceivable to extract only the time periods in which changes occur from metadata that has been created by pursuing naturalness. This kind of data extraction is also something that large-scale language models M excel at. Large-scale language models M are inherently good at tasks such as summarization and extracting important parts.

[0274] Here are some examples of prompts developers might enter to extract time periods of change.

[0275] The first example is a prompt that instructs the user to extract data using the keyword "period of change". Specifically, the developer enters a message as the second prompt P2 that reads, "From the sequence of natural metadata created in the previous steps, extract five scenes, each no longer than 3 minutes in length, that show a movement or change in the bounding box."

[0276] The second example involves making the elapsed time per second of playback variable depending on the scene. Specifically, the developer would enter a message as the second prompt P2 that reads: "Compress the sequence of natural metadata created in the previous steps over time. To do this, shorten the playback speed by 5 times during periods when the position and size of the bounding boxes do not change. On the other hand, do not change the playback speed during periods when the position and size of the bounding boxes do change. Compress the length so that the total playback time is approximately 15 minutes."

[0277] By inputting this second prompt P2, it is possible to obtain simulated data MD that is a natural scene yet has a high density of change suitable for algorithm development.

[0278] Another method involves specifying how the scene changes using natural language and having a large-scale language model M generate the changes.

[0279] For example, the developer might enter a message as the second prompt P2 that reads, "Add loading and unloading scenes to the metadata of the warehouse surveillance camera footage created in the previous steps. Add scenes of a person loading items onto a cart and scenes of an automated guided vehicle (AGV) unloading items." This also allows developers to specify typical loading and unloading scenarios.

[0280] Furthermore, developers can specify a particular action of the subject by entering a second prompt P2, such as, "Add a scene of a person putting a small package into their pocket and carrying it out to the metadata of the warehouse surveillance camera footage created in the previous steps."

[0281] <6-8. Simulated data generated over a long period of time> The examples described above illustrate how to obtain the desired simulated data (MD) by repeatedly entering prompts over a relatively short period of time, typically a few minutes.

[0282] The method of obtaining simulated data MD using a large-scale language model M can be further developed. For example, consider cases where the elapsed time from prompt input to obtaining simulated data MD spans several days. In other words, technically, a more suitable simulated data MD can be obtained by "increasing the session duration."

[0283] By increasing the session duration, developers can enter a second prompt P1 such as, for example, "Imagine a household where an 80-year-old mother lives alone. A camera is installed in the living room of this house. The system will perform image recognition on the camera's video feed to detect people. Person detection will be performed at 10-second intervals. Assume the mother spends a day and output metadata for the person detection results. After this, if instructed to 'start data generation,' output the metadata 24 hours later."

[0284] Furthermore, developers can specify the actual imaging device and field of view to use as a reference by providing additional prompts such as, "Here, please refer to the video from XXX for footage from a living room camera in the home of an elderly person living alone." "XXX" refers to connection information for a network camera that the developer can use without any copyright issues.

[0285] After specifying prerequisites and other conditions via these prompts, developers can instruct the large-scale language model M to begin generating simulated data MD by entering a sentence such as "Please start data generation" as the first prompt P1.

[0286] By waiting without disconnecting the session with the large-scale language model M, the developer can have the large-scale language model M generate the simulated data MD as instructed.

[0287] In the case of the large-scale language model M, simulated data MD may be generated using pre-trained knowledge, or one day's worth of simulated data MD may be generated by analyzing 24 hours of actual video footage of the specified "XXX".

[0288] In this manner, it is impossible to determine whether the developer generated the MD based on knowledge acquired by the large-scale language model M, or whether they generated simulated data MD by image processing of video footage from actually installed network cameras. In other words, the method for creating the actual simulated data MD is left to the large-scale language model M.

[0289] For developers, the only requirement is obtaining appropriate simulated data MD, so the method used to generate the simulated data MD using the large-scale language model M is not a problem.

[0290] In environments where agents such as large-scale language models M or chat systems are highly developed, the developer's prompt input is performed by repeating the steps described in the examples above, but the method of generating simulated data MD by the large-scale language model M may not require the developer to be aware of it.

[0291] Whether the metadata is generated using knowledge acquired by a large-scale language model M, or generated by image processing of images obtained from the real world via a network camera, the subsequent interactive inputs to the second prompt P2 and third prompt P3, as in other examples, can be modified into more natural and user-friendly simulated data MD for algorithm development (verification).

[0292] This shows an example of prompts that a developer might enter to further improve the simulated data MD obtained 24 hours after the input of the first prompt P1.

[0293] For example, if the large-scale language model M provides the response, "I have generated one day's worth of metadata. Please specify a folder. I will store the files," the developer can increase the density of metadata changes by entering a second prompt P2 such as, "Before storing the files, compress the metadata time length. Please remove scenes where no people are present."

[0294] Furthermore, developers may increase the density of metadata changes by inputting a second prompt P2 such as, "Please remove scenes where the lights are off as people cannot be detected during that time," or "For scenes where people are still, please set the time interval to 1 minute and the playback speed to 6 times the normal speed."

[0295] For a human developer, it is extremely difficult to provide instructions to a large-scale language model M from a state of zero, where no instructions have been given, in order to obtain a perfect simulated data set MD with a single prompt input. However, it is possible to notice inconsistencies or discrepancies between the initial simulated MD data obtained as a base and the developer's expectations. It is easy to input multiple prompts for desired changes regarding these inconsistencies and discrepancies, thereby ultimately obtaining the suitable simulated MD data desired by the developer. Furthermore, by making these instructions qualitative rather than quantitative and using natural language, the process becomes easy for humans as well.

[0296] In order to improve the simulated data MD as metadata using such interactive methods, it is necessary to have the initial base metadata, which may be generated from knowledge acquired by a large-scale language model M, or it may be generated by performing image processing on video obtained from real-world surveillance cameras, etc.

[0297] In other words, metadata generated based on knowledge, or metadata obtained by image processing of real-world images, may be difficult to use directly for algorithm development and verification. Therefore, the method of gradually modifying the simulated data MD as metadata by interactively inputting prompts, as explained in the examples above, is extremely effective.

[0298] <7. Implementation of Functions> The processing circuit 51 is implemented by a circuit that performs calculations. Specifically, the circuit acting as the processing circuit 51 of the simulated data generation device 2 performs predetermined calculations to realize the functions of a prompt acquisition unit F1, a response information acquisition unit F2, a presentation processing unit F3, an information provision processing unit F4, and an auxiliary information acquisition unit F5.

[0299] When the processing circuit 51 is a control unit implemented by an arithmetic circuit such as a CPU (Central Processing Unit), GPU (Graphics Processing Unit), or TPU (Tensor Processing Unit), various functions are realized by the processing circuit 51 executing a program to perform predetermined calculations using various memory or other storage areas. That is, predetermined programs and calculation results obtained during the execution of the program are appropriately written to memory or other storage areas provided inside or outside the processing circuit 51.

[0300] Furthermore, when the processing circuit 51 is implemented using a circuit such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array), various functions can be realized without requiring memory or other storage areas provided outside the processing circuit 51, by designing and constructing the circuit in a manner that realizes predetermined functions. However, in the implementation of various functions using an ASIC or FPGA, the memory provided inside the processing circuit 51 may be utilized.

[0301] These implementation examples can be applied to the developer terminal 1, the first server device 3, the second server device 4, and so on.

[0302] <8. Summary> As described above, the simulated data generation device 2 related to this technology includes a prompt acquisition unit F1 that acquires a user input (input by the developer) as a first prompt P1 that instructs the generation of time-series data of simulated data MD, which simulates metadata output from the imaging device as an analysis result of the captured image, and a response information acquisition unit F2 that inputs the first prompt P1 to a large-scale language model M and obtains time-series data of simulated data MD as response information. Furthermore, the time-series data of the simulated data MD is made into data that allows the movement of the subject captured in the captured image to be confirmed by visualization. The prompt acquisition unit F1 acquires user input instructing correction of the response information as a second prompt P2, and the response information acquisition unit F2 inputs the second prompt P2 into the large-scale language model M to obtain the corrected time-series data of the simulated data MD as response information. Developing applications that utilize metadata output from imaging devices requires input metadata, or simulated metadata (MD) that mimics the metadata. However, methods for actually generating metadata require a shooting environment and a subject. In particular, if the subject is a person, it may be possible to hire an actor, but the movements of an actor when acting are different from the movements of a person who is not aware of the camera or other imaging device, resulting in the generation of unnatural metadata. Another possible method for generating simulated data (MD) involves creating a three-dimensional model of the subject, placing a virtual imaging device and the three-dimensional model in the spatial SP of the sensing target, moving the three-dimensional model, generating a two-dimensional image of the imaging device's field of view, and then outputting simulated data (MD) that simulates metadata. However, it is difficult to make the three-dimensional model move naturally, and even with this method, unnatural simulated data (MD) is obtained that differs from the metadata obtained when imaging a person moving naturally. The simulated data generation device 2, equipped with this configuration, directly generates simulated data MD that simulates metadata without actually performing imaging or generating a three-dimensional model. This eliminates the unnaturalness of the simulated data MD caused by the unnatural movement of the three-dimensional model. Furthermore, it eliminates the need to actually install an imaging device in the spatial SP of the sensing target, thereby reducing the development cost of the application. While the metadata structure may change depending on the imaging device, this configuration allows for the creation of the main functional parts of the application being developed without needing to prepare an imaging device. Furthermore, when adapting the application to a new type of imaging device, only the interface endpoint needs to be created, thus improving work efficiency and reducing development time. Furthermore, the generated simulated data MD is presented as time-series data and is made visualizeable, allowing users to visually confirm any unnatural behavior in the simulated data MD. Consequently, it becomes easier to generate prompts (second prompt P2) to correct the time-series data of the simulated data MD, contributing to the generation of more natural-looking simulated data MD. Furthermore, when obtaining metadata by programming the movement of a subject, the metadata obtained is limited to movement patterns conceived by the developer for the subject or 3D model, thus limiting the situations that the created application can handle and reducing its adaptability. However, by inputting the first prompt P1, the large-scale language model M generates time-series data of simulated data MD, which allows for the acquisition of simulated data MD based on natural movements that the user would not have conceived. This expands the variety of input data when creating applications, enabling the creation of applications that perform appropriate processing according to various situations. Furthermore, since the simulated data MD can be modified simply by entering the second prompt P2, the effort and time required to prepare metadata or simulated data MD used in application development can be significantly reduced. Furthermore, consider an application that receives metadata from an imaging device monitoring a swimming pool and outputs an alert if a person is drowning. In this case, an appropriate application can be developed by capturing images of the drowning situation with the imaging device and preparing the metadata obtained at that time. However, actually drowning people for filming is neither appropriate nor easy. Furthermore, even if actors are hired to act out the process of drowning, it is impossible to accurately reproduce the movements of someone actually drowning, nor can it accurately replicate the different ways in which people drown, varying in gender, age, muscle mass, and other attributes. In this regard, by using a large-scale language model M trained on actual drowning accident footage, it is possible to generate simulated data MD that appropriately simulates the metadata output from captured images when a person drowns, thereby contributing to improved application performance.

[0303] As explained with reference to Figures 4 and 13, the simulated data generation device 2 may perform the acquisition of the second prompt P2 by the prompt acquisition unit F1 and the acquisition of the time-series data of the modified simulated data MD by the response information acquisition unit F2 multiple times each. This allows for repeated input and response acquisition for the second prompt P2, making the resulting modified simulated data MD more natural.

[0304] As explained with reference to Figures 3 and 4, the prompt acquisition unit F1 in the simulated data generation device 2 may acquire a third prompt P3 that instructs the setting of information related to the captured image that is not output as response information, and the response information acquisition unit F2 may obtain time-series data of the simulated data MD by inputting the third prompt P3 together with the first prompt P1 into the large-scale language model M. For example, when detecting a person and outputting the coordinate information of the bounding box BB as metadata, various factors influence the person's movement. Much of this information is not included in the output metadata, but by providing it as input to a large-scale language model M, it becomes possible to generate more appropriate and natural simulated data MD.

[0305] As explained with reference to Figure 3, etc., in the simulated data generation device 2, the information relating to the captured image may be information relating to the subject. Information relating to the subject includes, for example, personal attributes such as gender, age, height, and weight. Information such as whether the subject uses a wheelchair, rides a bicycle, uses a skateboard, or has a disability in their right hand is also considered relevant to the subject. By inputting this information into a large-scale language model M, it becomes possible to introduce variations into the generated simulated data MD, thereby enhancing the functionality of the application.

[0306] As explained with reference to Figure 3, etc., in the simulated data generation device 2, the information relating to the subject may be information relating to the subject's personal space. Each person has a concept of personal space, and they instinctively try to avoid others entering that space. Therefore, it is generally unlikely that a person would behave in a way that invades personal space in an uncrowded space, for example. By providing information related to personal space to a large-scale language model, it is possible to generate simulated data (MD) that is based on more natural human behavior. Therefore, it becomes possible to develop high-performance applications using this simulated data (MD).

[0307] As explained with reference to Figure 3, etc., in the simulated data generation device 2, the information relating to the subject may be attribute information of the subject. Subject attribute information includes information such as the subject's gender, age, height, weight, dominant hand, and stride length. If the subject's attribute information differs, the subject's movements will naturally differ as well. Therefore, by inputting the subject's attribute information into a large-scale language model M, it becomes possible to obtain more natural simulated data MD. Furthermore, by taking this attribute information into account when setting personal space, it becomes possible to make the movements of each subject appear more natural. Furthermore, the subject is not limited to people; it may also be animals, vehicles, food, or anything other than people. If the subject is an animal, attribute information could include its species (dog, cat, etc.), size, weight, etc. Information such as whether it is a pet or a wild animal can also be considered attribute information. If the subject is a vehicle, attribute information can include the make and model, color, whether or not there are people inside, whether it is moving or stationary, the vehicle's length, width, and weight. If the subject is food, attribute information can include broad category information such as vegetables, fruits, and meat, or more specific category information such as potatoes and leafy greens, and even more specific category information such as cucumbers and carrots, as well as color and size. By setting these attribute details for each subject, it becomes possible to add subtle differences to the generated simulated data (MD), allowing for the appropriate creation of diverse application behaviors.

[0308] As explained with reference to Figure 3, etc., in the simulated data generation device 2, the information relating to the subject may be information relating to the orientation of the subject. By requesting the large-scale language model M to internally set the orientation of the subject, the time-series data of the simulated data MD output from the large-scale language model M will be closer to the metadata that would be output if the subject were performing natural movements. For example, it is unlikely that a person would walk backward; if they wanted to move backward, it is usually assumed that they would turn their body around and then walk forward. Therefore, in simulated data MD that takes into account the orientation of the subject, the data will take into account that when changing the direction of movement, the movement will be preceded by a change in the body's orientation. Therefore, it becomes possible to make the simulated data (MD) used in application development closer to natural data.

[0309] As explained with reference to Figure 3, etc., in the simulated data generation device 2, the information relating to the captured image may be information relating to the spatial SP that is the object of sensing by the imaging device. The spatial SP to be sensed refers to the spatial SP on which the imaging device is installed. The information related to the spatial SP to be sensed may include not only information such as the size and location of the spatial SP, but also weather information, time information, and seasonal information. In addition, the location information of objects placed in the spatial SP, such as the type, size, and location of fixtures in a retail store, is also considered information related to the spatial SP to be sensed. By inputting this information into a large-scale language model M, it becomes possible to generate simulated data MD influenced by that information, enabling the development of highly adaptable applications, such as applications usable not only in summer but also in winter. Furthermore, information that identifies the type of spatial SP, such as whether it is a spatial SP within a theme park, a spatial SP within a baseball stadium, or a spatial SP within a store, is considered an example of information related to the captured image. The movement of a person walking on the street may differ from the movement of a person walking on a road within a theme park. By inputting this information into a large-scale language model M, it becomes possible to generate simulated data MD that takes into account the movement of people in different situations.

[0310] As explained with reference to Figure 4, etc., in the simulated data generation device 2, the time-series data of the simulated data MD may include data relating to specific actions of the subject captured in the captured image. For example, the large-scale language model M outputs simulated data MD that simulates metadata obtained from imaging devices that monitor stores. This simulated data MD corresponds not only to the actions taken by customers while they are shopping, but also to other specific actions, such as the actions taken by store employees while they are preparing the store before it opens.

[0311] As explained with reference to Figure 4, etc., in the simulated data generation device 2, the specific behavior may be a behavior that an abnormal subject might exhibit. In other words, specific behaviors are not the actions of a normal customer shopping in a store or a person swimming in a pool, but rather the actions of a criminal stealing in a store or a person drowning in a pool; they are non-stationary behaviors. Even if actors or other performers are instructed to act in front of the imaging device, it is difficult to obtain natural metadata about the abnormal behaviors that these subjects may exhibit. Therefore, by generating simulated data MD related to these specific behaviors using a large-scale language model M, it is possible to create an application that can appropriately detect and respond to cases where a subject that is actually abnormal performs these specific behaviors.

[0312] As explained with reference to Figure 11, etc., in the simulated data generation device 2, the simulated data MD may include data that simulates metadata in cases where the analysis results of the captured image contain errors. For example, consider an imaging device that sets a bounding box BB for a person detected within the field of view and outputs the coordinate information of that bounding box BB as metadata. If the imaging device detects two people, the coordinate information of two bounding box BBs will be output. Subsequently, if the two detected people move to a position where they overlap within the field of view, the device may temporarily only be able to detect one person, and the coordinate information of one bounding box BB may be output as metadata. In the generation of simulated data MD, when considering a scene where two people are located within the field of view, even if the two people overlap within the field of view, the simulation will output the coordinate information of two bounding box BBs because it is known that two people exist. However, in this configuration, simulated data MD is output that simulates an incorrect analysis result for a moment when the imaging device cannot recognize one subject. This makes it possible to generate simulated data (MD) that accurately and precisely simulates the metadata that can be output from an actual imaging device, and contributes to the development of applications that can properly handle metadata output as erroneous analysis results.

[0313] As explained with reference to Figures 3, 4, and 7, the response information acquisition unit F2 in the simulated data generation device 2 may input auxiliary information SD, which is used to generate time-series data of the simulated data MD, to the large-scale language model M along with the first prompt P1. By inputting auxiliary information SD along with the first prompt P1 into the large-scale language model M, the generated simulated data MD can be made more natural.

[0314] As explained with reference to Figure 3, etc., in the simulated data generation device 2, the auxiliary information SD may be information relating to the data format of the simulated data MD. This makes it possible to generate simulated data (MD) that conforms to the output format of the imaging device actually used in the spatial SP being sensed. Furthermore, the information related to the data format may include information that allows the output format to be estimated, such as the model number and manufacturer information of the imaging device used.

[0315] As explained with reference to Figure 3, etc., in the simulated data generation device 2, the auxiliary information SD may be information related to the imaging device. Information related to the imaging device may include, for example, the model number and manufacturer information of the imaging device used in the spatial SP to be sensed, or information such as the number of pixels and field of view. By using this information, it becomes possible to generate simulated MD data that takes into account differences in metadata caused by the imaging device.

[0316] As explained with reference to Figure 3, etc., in the simulated data generation device 2, the auxiliary information SD may be information relating to the number and position of imaging devices arranged in the spatial SP that is the target of sensing by the imaging devices. When monitoring a spatial SP (Sensing Target) with multiple imaging devices, it is possible to output simulated data (MD) which simulates the metadata output from multiple imaging devices, and ensures consistency between the metadata by considering the positional relationships between the imaging devices. Therefore, it is possible to provide simulated data (MD) suitable for developing applications that analyze metadata output from multiple imaging devices.

[0317] As explained with reference to Figures 3 and 7, etc., in the simulated data generation device 2, the auxiliary information SD may be the search result obtained by performing a search on the vector database VD using the first prompt P1. Vector databases (VD) store various types of information represented as vectors. Furthermore, vectors are considered high-dimensional entities containing a very large number of elements. In other words, it can be described as a database in which the degree of similarity in meaning and concept is expressed for each piece of information. Therefore, by inputting the results retrieved and extracted from the vector database VD into the large-scale language model M, it becomes possible to obtain situation-appropriate simulated data MD.

[0318] As explained in the modified examples, in the simulated data generation device 2, the auxiliary information SD may be class information that identifies the subject being sensed by the imaging device. This makes it possible to generate simulated data (MD) related to a limited set of sensing targets.

[0319] As explained with reference to Figure 3, etc., in the simulated data generation device 2, the simulated data MD may be coordinate information for the bounding box BB surrounding the subject to be detected in the imaging device. The bounding box BB is rectangular, and its position and size within the field of view can be determined by the coordinate information of at least two points. Therefore, the coordinate information of the bounding box BB can be considered simplified information. By outputting simplified data as simulated data MD, the processing burden using the large-scale language model M can be reduced. Furthermore, the visualization is easier for developers to understand, and it is easier to issue instructions for modifying the simulated data MD.

[0320] The simulated data generation method related to this technology involves an information processing device executing the following steps: first, a user input (input by the developer) as a first prompt P1 that instructs the generation of time-series data of simulated data MD, which simulates metadata output from the imaging device as a result of analyzing captured images, and which allows the movement of the subject captured in the captured image to be confirmed by visualization; second, a user input as a second prompt P2 that instructs the correction of the response information; and third, a user input as a second prompt P2 that instructs the correction of the response information.

[0321] The simulated data generation system related to this technology comprises a storage unit storing a large-scale language model M, a prompt acquisition unit F1 that acquires a first prompt P1 as a user input (input by the developer) instructing the generation of time-series data of simulated data MD, which simulates metadata output from the imaging device as an analysis result of the captured image, an answer information acquisition unit F2 that inputs the first prompt P1 to the large-scale language model M and obtains time-series data of simulated data MD as answer information, and a presentation processing unit F3 that presents information that visualizes the time-series data of simulated data MD and allows the movement of the subject captured in the captured image to be confirmed. The prompt acquisition unit F1 acquires a second prompt P2 as a user input instructing the correction of the answer information, the answer information acquisition unit F2 inputs the second prompt P2 to the large-scale language model M and obtains the corrected time-series data of simulated data MD as answer information, and the presentation processing unit F3 presents information that visualizes the corrected time-series data of simulated data MD.

[0322] The various effects and benefits described above can also be obtained through such simulated data generation methods and systems.

[0323] The programs for realizing the aforementioned functions can be pre-recorded on recording media such as HDDs (Hard Disk Drives) built into computer devices, or on ROMs within microcomputers with CPUs. Alternatively, the programs can be temporarily or permanently stored (recorded) on removable recording media such as flexible disks, CD-ROMs (Compact Disk Read Only Memory), MO (Magneto Optical) disks, DVDs (Digital Versatile Discs), Blu-ray Discs (registered trademark), magnetic disks, semiconductor memory, and memory cards. Such removable recording media can be provided as so-called packaged software. In addition to installing such programs from removable storage media to personal computers, they can also be downloaded from download sites via networks such as LANs and the Internet.

[0324] Furthermore, the effects described herein are merely illustrative and not limited to those described herein, and other effects may also occur.

[0325] Furthermore, the examples mentioned above can be combined in any way, and it is possible to obtain the various effects and benefits described above even when using various combinations.

[0326] <9. This Technology> This technology can also be configured as follows: (1) A prompt acquisition unit acquires a user input as the first prompt that instructs the generation of time-series data of simulated data, which simulates the metadata output from the imaging device as a result of analyzing the captured image. The system includes a response information acquisition unit that inputs the first prompt into a large-scale language model and obtains time-series data of the simulated data as response information, The time-series data of the simulated data is made available for visualization, allowing the movement of the subject captured in the image to be confirmed. The prompt acquisition unit acquires user input instructing the correction of the response information as a second prompt, The response information acquisition unit inputs the second prompt into the large-scale language model and obtains the time-series data of the modified simulated data as response information. A simulated data generation device. (2) The process of obtaining the second prompt by the prompt acquisition unit and obtaining the time-series data of the modified simulated data by the response information acquisition unit is performed multiple times each. The simulated data generation device described in (1) above. (3) The prompt acquisition unit acquires a third prompt that instructs the setting of information related to the captured image that is not output as the response information, The response information acquisition unit obtains the time-series data of the simulated data by inputting the third prompt along with the first prompt into the large-scale language model. A simulated data generation device as described in either (1) or (2) above. (4) The information relating to the captured image was defined as information relating to the subject. The simulated data generation device described in (3) above. (5) The information relating to the subject was defined as information relating to the subject's personal space. The simulated data generation device described in (4) above. (6) The information relating to the subject was defined as attribute information of the subject. A simulated data generation device as described in any of (4) to (5) above. (7) The information relating to the subject was defined as information regarding the orientation of the subject. A simulated data generation device as described in any of (4) to (6) above. (8) The information relating to the captured image was defined as information relating to the space being sensed by the imaging device. A simulated data generation device as described in any of (3) through (7) above. (9) The time-series data of the simulated data includes data relating to specific actions of the subject captured in the captured image. A simulated data generation device as described in any of the above (1) to (8). (10) The aforementioned specific behaviors were defined as behaviors that an abnormal subject might exhibit. The simulated data generation device described in (9) above. (11) The simulated data includes data that simulates metadata in cases where the analysis results of the captured image contain errors. A simulated data generation device as described in any of (1) to (10) above. (12) The response information acquisition unit inputs auxiliary information used to generate the time-series data of the simulated data, along with the first prompt, into the large-scale language model. A simulated data generation device as described in any of (1) through (11) above. (13) The aforementioned auxiliary information was defined as information relating to the data format of the simulated data. The simulated data generation device described in (12) above. (14) The aforementioned auxiliary information was defined as information relating to the imaging device. A simulated data generation device as described in either (12) or (13) above. (15) The aforementioned auxiliary information is information relating to the number and location of the imaging devices arranged in the space to be sensed by the imaging devices. A simulated data generation device as described in any of (12) to (14) above. (16) The aforementioned auxiliary information was obtained as search results by performing a search on the vector database using the first prompt. A simulated data generation device as described in any of (12) to (15) above. (17) The aforementioned auxiliary information is defined as class information that identifies the subject being sensed by the imaging device. A simulated data generation device as described in any of (12) to (16) above. (18) The aforementioned simulated data was defined as coordinate information for the bounding box surrounding the subject to be detected in the imaging device. A simulated data generation device as described in any of the above (1) to (17). (19) A process to obtain, as the first prompt, user input instructing the generation of time-series data of simulated data that simulates metadata output from the imaging device as a result of analyzing the captured image, and which allows the movement of the subject captured in the image to be confirmed by visualization; The process involves inputting the first prompt into a large-scale language model and obtaining time-series data of the simulated data as response information, The process involves obtaining user input as a second prompt to instruct the user to correct the aforementioned response information, The information processing device performs the following steps: input the second prompt into the large-scale language model and obtain the time-series data of the modified simulated data as response information. A method for generating simulated data. (20) A memory unit that stores a large-scale language model, A prompt acquisition unit acquires a user input as the first prompt that instructs the generation of time-series data of simulated data, which simulates the metadata output from the imaging device as a result of analyzing the captured image. A response information acquisition unit inputs the first prompt into the large-scale language model and obtains time-series data of the simulated data as response information, The system includes a presentation processing unit that visualizes the time-series data of the simulated data and presents information that allows the movement of the subject captured in the captured image to be confirmed, The prompt acquisition unit acquires user input instructing the correction of the response information as a second prompt, The response information acquisition unit inputs the second prompt to the large-scale language model and acquires the time-series data of the modified simulated data as response information. The presentation processing unit presents information that visualizes the time-series data of the modified simulated data. A system for generating simulated data. [Explanation of Symbols]

[0327] 2. Simulated Data Generation Device BB Bounding Box BB1 Bounding Box F1 prompt acquisition unit F2 Answer information acquisition part F3 Presentation Processing Unit M Large-scale language models MD simulated data P1 First Prompt P2 Second Prompt P3 Third Prompt S2 Simulated Data Generation System SD supplementary information SP space VD Vector Database

Claims

1. A prompt acquisition unit acquires a user input as the first prompt that instructs the generation of time-series data of simulated data, which simulates the metadata output from the imaging device as a result of analyzing the captured image. The system includes a response information acquisition unit that inputs the first prompt into a large-scale language model and obtains time-series data of the simulated data as response information, The time-series data of the simulated data is made available for visualization, allowing the movement of the subject captured in the image to be confirmed. The prompt acquisition unit acquires user input instructing the correction of the response information as a second prompt, The response information acquisition unit inputs the second prompt into the large-scale language model and obtains the time-series data of the modified simulated data as response information. A simulated data generation device.

2. The process of obtaining the second prompt by the prompt acquisition unit and obtaining the time-series data of the modified simulated data by the response information acquisition unit is performed multiple times each. A simulated data generation device according to claim 1.

3. The prompt acquisition unit acquires a third prompt that instructs the setting of information related to the captured image that is not output as the response information, The response information acquisition unit obtains the time-series data of the simulated data by inputting the third prompt along with the first prompt into the large-scale language model. A simulated data generation device according to claim 1.

4. The information relating to the captured image was defined as information relating to the subject. The simulated data generation device according to claim 3.

5. The information relating to the subject was defined as information relating to the subject's personal space. The simulated data generation device according to claim 4.

6. The information relating to the subject was defined as attribute information of the subject. The simulated data generation device according to claim 4.

7. The information relating to the subject was defined as information regarding the orientation of the subject. The simulated data generation device according to claim 4.

8. The information relating to the captured image was defined as information relating to the space being sensed by the imaging device. The simulated data generation device according to claim 3.

9. The time-series data of the simulated data includes data relating to specific actions of the subject captured in the captured image. A simulated data generation device according to claim 1.

10. The aforementioned specific behaviors were defined as behaviors that an abnormal subject might exhibit. The simulated data generation device according to claim 9.

11. The simulated data includes data that simulates metadata in cases where the analysis results of the captured image contain errors. A simulated data generation device according to claim 1.

12. The response information acquisition unit inputs auxiliary information used to generate time-series data of the simulated data, along with the first prompt, into the large-scale language model. A simulated data generation device according to claim 1.

13. The aforementioned auxiliary information was defined as information relating to the data format of the simulated data. The simulated data generation device according to claim 12.

14. The aforementioned auxiliary information was defined as information relating to the imaging device. The simulated data generation device according to claim 12.

15. The aforementioned auxiliary information is information relating to the number and location of the imaging devices arranged in the space to be sensed by the imaging devices. The simulated data generation device according to claim 12.

16. The aforementioned auxiliary information was obtained as search results by performing a search on the vector database using the first prompt. The simulated data generation device according to claim 12.

17. The aforementioned auxiliary information is defined as class information that identifies the subject being sensed by the imaging device. The simulated data generation device according to claim 12.

18. The aforementioned simulated data was defined as coordinate information for the bounding box surrounding the subject to be detected in the imaging device. A simulated data generation device according to claim 1.

19. A process to obtain, as the first prompt, user input instructing the generation of time-series data of simulated data that simulates metadata output from the imaging device as a result of analyzing the captured image, and which allows the movement of the subject captured in the image to be confirmed by visualization; The process involves inputting the first prompt into a large-scale language model and obtaining time-series data of the simulated data as response information, The process involves obtaining user input as a second prompt to instruct the user to correct the aforementioned response information, The information processing device performs the following steps: input the second prompt into the large-scale language model and obtain the time-series data of the modified simulated data as response information. Method for generating simulated data.

20. A memory unit that stores a large-scale language model, A prompt acquisition unit acquires a user input as the first prompt that instructs the generation of time-series data of simulated data, which simulates the metadata output from the imaging device as a result of analyzing the captured image. A response information acquisition unit inputs the first prompt into the large-scale language model and obtains time-series data of the simulated data as response information, The system includes a presentation processing unit that visualizes the time-series data of the simulated data and presents information that allows the movement of the subject captured in the captured image to be confirmed, The prompt acquisition unit acquires user input instructing the correction of the response information as a second prompt, The response information acquisition unit inputs the second prompt to the large-scale language model and acquires the time-series data of the modified simulated data as response information. The presentation processing unit presents information that visualizes the time-series data of the modified simulated data. A system for generating simulated data.