Information processing device, information processing method, and information processing program
The information processing device generates text focusing on specific events in a monitored space by integrating sensor, guide, and environmental information, addressing the limitations of existing text creation methods.
Patent Information
- Authority / Receiving Office
- JP Β· JP
- Patent Type
- Applications
- Current Assignee / Owner
- NEC CORP
- Filing Date
- 2024-10-29
- Publication Date
- 2026-05-15
AI Technical Summary
Existing text creation techniques fail to generate texts that focus on specific events within a monitored space, lacking the ability to highlight important areas or events as intended by the user.
An information processing device and method that acquires sensor and environmental information, guides information about the space, and generates text using these inputs to focus on specific events, including object shapes, movements, and spatial boundaries.
Generates text that accurately highlights specific events and areas of interest in a monitored space, enhancing the relevance and accuracy of the generated content.
Smart Images

Figure 2026078941000001_ABST
Abstract
Description
Technical Field
[0006] , ,
[0001] The present disclosure relates to an information processing apparatus, an information processing method, and an information processing program.
Background Art
[0002] Techniques for assisting in the creation of texts such as reports have been proposed. As a technique for assisting in the creation of texts, for example, the technique described in Patent Document 1 can be cited. Patent Document 1 discloses an information processing system that causes a first user terminal to display a selection screen for receiving a selection operation of the type of application document, instructs an artificial intelligence having a text creation function to create a text of the type of application document selected by the selection operation, causes the first user terminal to display the application document created by the instruction in an editable manner, and causes a second user terminal of a proofreader familiar with the type of the selected application document to display a request for proofreading the application document edited on the first user terminal.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] By the way, a user who creates a text such as a report may want to create a text focusing on a specific area or a specific event in a space to be monitored. The technique described in Patent Document 1 has a problem that it cannot generate a text focusing on an event that a user wants to focus on.
[0005] The present disclosure has been made in view of the above problems, and an exemplary object thereof is to provide a technique for generating a text focusing on a specific event in a space to be monitored.
Means for Solving the Problems
[0006] An information processing device relating to an illustrative aspect of this disclosure includes: acquisition means for acquiring sensor information indicating sensing results from a sensor that senses a space and environmental information relating to the environment of the space; guide information acquisition means for acquiring guide information indicating at least one of the following: the shape of an object included in the space, the trajectory of the object's movement, the relationship between objects, and the boundary in the space; text generation means for generating text describing the space using the sensor information, the guide information, and the environmental information; and text output means for outputting the text generated by the text generation means.
[0007] An information processing method relating to an exemplary aspect of this disclosure includes: an acquisition process in which at least one processor acquires sensor information indicating sensing results from a sensor that senses a space and environmental information relating to the environment of the space; a guide information acquisition process in which the at least one processor acquires guide information indicating at least one of the following: the shape of an object included in the space, the trajectory of the object's movement, the relationship between objects, and the boundary in the space; a text generation process in which the at least one processor generates text describing the space using the sensor information, the guide information, and the environmental information; and a text output process in which the at least one processor outputs the text generated in the text generation process.
[0008] An illustrative aspect of the present disclosure is an information processing program for causing a computer to function as an information processing device, wherein the computer functions as: an acquisition means for acquiring sensor information indicating sensing results from a sensor that senses a space and environmental information relating to the environment of the space; a guide information acquisition means for acquiring guide information indicating at least one of the following: the shape of an object included in the space, the trajectory of movement of an object, the relationship between objects, and the boundary in the space; a text generation means for generating text describing the space using the sensor information, the guide information, and the environmental information; and a text output means for outputting the text generated by the text generation means. [Effects of the Invention]
[0009] One illustrative aspect of this disclosure is that it provides a technology that generates text focusing on specific events within a monitored space. [Brief explanation of the drawing]
[0010] [Figure 1] This is a block diagram showing the configuration of the information processing device related to this disclosure. [Figure 2] This is a flowchart showing the flow of the information processing method related to this disclosure. [Figure 3] This is a block diagram showing the configuration of the information processing device related to this disclosure. [Figure 4] This is a block diagram showing an example of the functional configuration of the information processing device related to this disclosure. [Figure 5] This figure shows an example of an image displayed by the display control unit related to this disclosure. [Figure 6] This figure shows specific examples of the guide information related to this disclosure. [Figure 7] This is a block diagram showing the configuration of a computer that functions as an information processing device related to this disclosure. [Modes for carrying out the invention]
[0011] The following are examples of embodiments of the present invention. However, the present invention is not limited to the exemplary embodiments shown below, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining some or all of the technologies (things or methods) employed in each of the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, embodiments obtained by appropriately omitting some of the technologies employed in each of the exemplary embodiments shown below may also be included in the scope of the present invention. In addition, the effects mentioned in each of the exemplary embodiments shown below are examples of effects that can be expected in that exemplary embodiment and do not define the scope of the present invention. That is, embodiments that do not produce the effects mentioned in each of the exemplary embodiments shown below may also be included in the scope of the present invention.
[0012] [First Exemplary Embodiment] A first exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. This exemplary embodiment is the basic form for each of the exemplary embodiments described later. The scope of application of each technology adopted in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technology adopted in this exemplary embodiment can also be adopted in other exemplary embodiments included in this disclosure, to the extent that no particular technical problems occur. Furthermore, each technology shown in the drawings referenced to explain this exemplary embodiment can also be adopted in other exemplary embodiments included in this disclosure, to the extent that no particular technical problems occur.
[0013] (Configuration of information processing device) The configuration of the information processing device 1 will be described with reference to Figure 1. Figure 1 is a block diagram showing the configuration of the information processing device 1. As shown in Figure 1, the information processing device 1 comprises an acquisition unit 11, a guide information acquisition unit 12, a text generation unit 13, and a text output unit 14. The acquisition unit 11 acquires sensor information indicating sensing results from a sensor that senses space, and environmental information relating to the environment of the space. The guide information acquisition unit 12 acquires guide information indicating at least one of the following: the shape of an object included in the space, the trajectory of the object's movement, the relationship between objects, and the boundary in the space. The text generation unit 13 generates text describing the space using the sensor information, the guide information, and the environmental information. The text output unit 14 outputs the text generated by the text generation unit 13.
[0014] (Effects of information processing equipment) As described above, the information processing device 1 employs a configuration comprising: an acquisition unit 11 that acquires sensor information indicating sensing results from a sensor that senses space and environmental information relating to the environment of the space; a guide information acquisition unit 12 that acquires guide information indicating at least one of the following: the shape of an object included in the space, the trajectory of the object's movement, the relationship between objects, and the boundary in the space; a text generation unit 13 that generates text describing the space using the sensor information, the guide information, and the environmental information; and a text output unit 14 that outputs the text generated by the text generation unit 13. Therefore, the information processing device 1 has the effect of generating text that focuses on specific events in the space being monitored.
[0015] (Information processing flow) The flow of the information processing method S1 will be explained with reference to Figure 2. Figure 2 is a flowchart showing the flow of the information processing method S1. As shown in Figure 2, the information processing method S1 includes an acquisition process S11, a guide information acquisition process S12, a text generation process S13, and a text output process S14.
[0016] In the acquisition process S11, at least one processor acquires sensor information indicating a sensing result by a sensor that senses a space, and environmental information regarding the environment of the space. In the guide information acquisition process S12, the at least one processor acquires guide information indicating at least any one of the shape of an object included in the space, the trajectory of the movement of the object, the relationship between objects, and the boundary in the space. In the text generation process S13, the at least one processor generates text for explaining the space using the sensor information, the guide information, and the environmental information. In the text output process S14, the at least one processor outputs the text generated in the text generation process S13.
[0017] (Effect of the information processing method) As described above, in the information processing method S1, there is adopted a configuration including an acquisition process S11 in which at least one processor acquires sensor information indicating a sensing result by a sensor that senses a space and environmental information regarding the environment of the space, a guide information acquisition process S12 in which the at least one processor acquires guide information indicating at least any one of the shape of an object included in the space, the trajectory of the movement of the object, the relationship between objects, and the boundary in the space, a text generation process S13 in which the at least one processor generates text for explaining the space using the sensor information, the guide information, and the environmental information, and a text output process S14 in which the at least one processor outputs the text generated in the text generation process S13. Therefore, according to the information processing method S1, an effect of generating text focusing on a specific event in a space that is a monitoring target can be obtained.
[0018] [Second exemplary embodiment] A second exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same functions as those described in the above exemplary embodiments are denoted by the same reference numerals, and the description thereof will be omitted as appropriate. Note that the scope of application of each technique adopted in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technique adopted in this exemplary embodiment can be adopted in other exemplary embodiments included in the present disclosure as long as there are no particular technical obstacles. In addition, each technique shown in each drawing referred to in order to explain this exemplary embodiment can be adopted in other exemplary embodiments included in the present disclosure as long as there are no particular technical obstacles.
[0019] (Configuration of Information Processing Apparatus) The configuration of the information processing apparatus 1A according to the present disclosure will be described with reference to FIG. 3. FIG. 3 is a block diagram showing the configuration of the information processing apparatus 1A. The information processing apparatus 1A includes a control unit 10A, a storage unit 20A, a communication unit 30A, an input unit 40A, and an output unit 50A. The communication unit 30A communicates with a device outside the information processing apparatus 1A via a communication line N. The communication unit 30A transmits data supplied from the control unit 10A to other devices, or supplies data received from other devices to the control unit 10A.
[0020] (Input Unit and Output Unit) The input unit 40A is a configuration for receiving an input to the information processing apparatus 1A. As an example, the input unit 40A includes input devices such as a keyboard, a mouse, a touch panel, a camera, and a microphone. Further, the input unit 40A may be configured to receive data from the input device via an interface such as USB (Universal Serial Bus). The output unit 50A is a configuration for outputting from the information processing apparatus 1A. As an example, the output unit 50A includes output devices such as a display, a printer, a touch panel, and a speaker. The output unit 50A may be configured to include an interface such as USB and output data to the output device via the interface.
[0021] (Storage Unit) The memory unit 20A stores various types of information referenced by the control unit 10A. The memory unit 20A particularly includes a data storage unit 201A. The data storage unit 201A stores various types of data, such as guide information acquired by the distance sensor information guide acquisition unit 105A (described later), guide information acquired by the image sensor information guide acquisition unit 106A (described later), guide information acquired by the spatial information guide acquisition unit 107A (described later), and reference text generated by the reference text generation unit 111A (described later).
[0022] (Control Unit) The control unit 10A includes a distance sensor information acquisition unit 101A, an image sensor information acquisition unit 102A, a spatial information acquisition unit 103A, an environmental information acquisition unit 104A, a distance sensor information guide acquisition unit 105A, an image sensor information guide acquisition unit 106A, a spatial information guide acquisition unit 107A, a region division unit 108A, an object detection unit 109A, an object tracking unit 110A, a reference text generation unit 111A, a text generation unit 112A, an output adjustment unit 113A, a text output unit 114A, and a display control unit 115A.
[0023] The distance sensor information acquisition unit 101A, the image sensor information acquisition unit 102A, the spatial information acquisition unit 103A, and the environmental information acquisition unit 104A are examples of acquisition means related to this disclosure. The distance sensor information guide acquisition unit 105A, the image sensor information guide acquisition unit 106A, and the spatial information guide acquisition unit 107A are examples of guide information acquisition means related to this disclosure. The area division unit 108A, the target detection unit 109A, the target tracking unit 110A, and the reference text generation unit 111A are examples of area division means, target detection means, target tracking means, and reference text generation means related to this disclosure, respectively. The text generation unit 112A, the output adjustment unit 113A, the text output unit 114A, and the display control unit 115A are examples of text generation means, output adjustment means, text output means, and display control means related to this disclosure, respectively.
[0024] (Distance measurement sensor information acquisition unit) Figure 4 is a block diagram showing an example of the functional configuration of the information processing device 1A. The distance sensor information acquisition unit 101A acquires distance sensor information that shows the sensing results from the distance sensor. Examples of distance sensors include radar or LIDAR (Laser Imaging Detection and Ranging). Examples of distance sensor information include information showing the measurement results from radar or LIDAR.
[0025] As an example, the distance sensor information acquisition unit 101A acquires distance sensor information input to the input unit 40A. The distance sensor information acquisition unit 101A may also receive distance sensor information from other devices via the communication unit 30A. The distance sensor information acquisition unit 101A may also acquire distance sensor information by reading it from a storage location specified by the user of the information processing device 1A (which may be a storage device within the information processing device 1A or a storage device outside the information processing device 1A). The distance sensor information acquisition unit 101A may also perform preprocessing such as noise reduction on the distance sensor information.
[0026] (Image sensor information acquisition unit) The image sensor information acquisition unit 102A acquires image sensor information that indicates the sensing results from the image sensor. Examples of image sensors include event cameras, infrared cameras, surveillance cameras, and in-vehicle cameras. Examples of image sensor information include image data (multispectral images, SAR (Synthetic Aperture Radar) images, infrared images, surveillance images, in-vehicle images, etc.). The distance measuring sensor and image sensor are examples of sensors related to this disclosure, and the distance measuring sensor information and image sensor information are examples of sensor information related to this disclosure. In other words, the sensor information related to this disclosure can also be said to include distance measuring sensor information acquired from the distance measuring sensor and image sensor information acquired from the image sensor.
[0027] The image sensor information acquisition unit 102A acquires image sensor information input to the input unit 40A, for example. The image sensor information acquisition unit 102A may also receive image sensor information from other devices via the communication unit 30A. The image sensor information acquisition unit 102A may also acquire image sensor information by reading it from a storage location specified by the user of the information processing device 1A (which may be a storage device within the information processing device 1A or a storage device outside the information processing device 1A). The image sensor information acquisition unit 102A may also perform preprocessing such as noise reduction on the image sensor information.
[0028] (Spatial information acquisition unit) The spatial information acquisition unit 103A acquires environmental information relating to the spatial environment. Examples of spatial information include information representing maps, satellite images, or aerial images. The spatial information may also include information representing the geography of the space (for example, information indicating latitude and longitude). As an example, the spatial information acquisition unit 103A acquires spatial information input to the input unit 40A. The spatial information acquisition unit 103A may also receive spatial information from other devices via the communication unit 30A. The spatial information acquisition unit 103A may also acquire spatial information by reading spatial information from a storage location specified by the user of the information processing device 1A (which may be a storage device within the information processing device 1A or a storage device outside the information processing device 1A).
[0029] (Environmental Information Acquisition Department) The environmental information acquisition unit 104A acquires environmental information about the target environment. Examples of environmental information include temperature, climate, current events (external news such as an aircraft taking off from xx airport), observation information from other locations (such as sensor information from another location), and the date and time when the sensor information was acquired. Examples of observation information from other locations include satellite data from other locations and information indicating the weather in the surrounding environment.
[0030] As an example, the environmental information acquisition unit 104A acquires environmental information input to the input unit 40A. The environmental information acquisition unit 104A may also receive environmental information from other devices via the communication unit 30A. Alternatively, the environmental information acquisition unit 104A may acquire environmental information by reading it from a storage location specified by the user of the information processing device 1A (which may be a storage device within the information processing device 1A or a storage device outside the information processing device 1A).
[0031] (Display Control Unit) The display control unit 115A displays a first image representing distance sensor information, a second image representing image sensor information, and a third image representing spatial information representing space on the display device. For example, the display control unit 115A outputs data representing the first image, data representing the second image, and data representing the third image to the display device connected to the output unit 50A, causing the images to be displayed on the display device. Alternatively, the display control unit 115A may transmit the image data to another device connected via the communication unit 30A, causing the image represented by the image data to be displayed on the display of the other device. Hereinafter, the first image, second image, and third image are displayed for the input of guide information, which will be described later. Hereinafter, when it is not necessary to distinguish between the first image, second image, and third image, they will be referred to as "guide input images".
[0032] Figure 5 shows an example of an image displayed by the display control unit 115A. In the example in Figure 5, image A11 is an example of a first image represented by distance measurement sensor information, more specifically an image showing the detection result of an object by radar. Image A12 is an example of a second image represented by image sensor information, more specifically an aerial photograph or satellite photograph. Image A13 is an example of a third image represented by spatial information, more specifically an image representing a map. Images A11 and A13 display objects detected by the object detection unit 109A, which will be described later.
[0033] (Guide information) On the screen displaying the image shown in Figure 5, the user can input guide information. Guide information is information that indicates the events the user wants to focus on in the guide input image. As an example, guide information may include information that specifies a part of a space (rectangle, etc.), the shape of an object contained in the space, the trajectory of the object's movement, the relationship between objects, and the boundary of the space (no entry, etc.).
[0034] For example, the user can provide guides for important objects or people using rectangles, or they can provide the past trajectories of people or objects as guides on the sensor data. Alternatively, they can provide assistance such as the relationship between objects and people. Or they can provide important boundaries or areas (areas where entry is prohibited, areas where passage is prohibited) on the sensor data as guides.
[0035] Figure 6 shows a specific example of a guide input image with superimposed guide information. In the example in Figure 6, image A21 is an image in which shapes A211 and A212, representing user-inputted guide information, are superimposed on a guide input image representing distance measurement sensor information. Shapes A211 and A212 are rectangles surrounding the object the user wants to focus on. Image A22 is an image in which shapes A221 and A222, representing user-inputted guide information, are superimposed on a guide input image representing image sensor information. Shape A221 is a rectangle indicating the area the user wants to focus on, and shape A222 is an arrow indicating the trajectory of the object the user wants to focus on. Image A23 is an image in which shapes A231 to A234, representing user-inputted guide information, are superimposed on a guide input image representing spatial information. Shapes A231 to A234 are arrows indicating the movement trajectory of the object the user wants to focus on.
[0036] (Distance measurement sensor information guide acquisition unit, image sensor information guide acquisition unit, spatial information guide acquisition unit) The distance sensor information guide acquisition unit 105A acquires guide information corresponding to the first image represented by the distance sensor information. The image sensor information guide acquisition unit 106A acquires guide information corresponding to the second image represented by the image sensor. The spatial information guide acquisition unit 107A acquires guide information corresponding to the third image represented by the spatial information.
[0037] The guide information acquired by the distance sensor information guide acquisition unit 105A, the image sensor information guide acquisition unit 106A, and the spatial information guide acquisition unit 107A may be guide information entered by the user, or it may be other information. In other words, when acquiring user-entered guide information, the distance sensor information guide acquisition unit 105A, the image sensor information guide acquisition unit 106A, and the spatial information guide acquisition unit 107A may acquire user-entered guide information for at least one of the first image, second image, and third image displayed by the display control unit 115A.
[0038] Furthermore, as another example of guide information, the guide information may include information indicating at least one of the following: the detection result by the target detection unit 109A (described later), the tracking result by the target tracking unit 110A (described later), and the division result by the area division unit 108A. In addition, the user may correct this information, and the corrected information may be acquired as guide information.
[0039] The distance sensor information guide acquisition unit 105A, the image sensor information guide acquisition unit 106A, and the spatial information guide acquisition unit 107A each store the acquired guide information and the guide input image in the data storage unit 201A. At this time, the distance sensor information guide acquisition unit 105A, the image sensor information guide acquisition unit 106A, and the spatial information guide acquisition unit 107A may each generate a superimposed image by superimposing the image representing the acquired guide information onto the guide input image, and store the generated superimposed image as guide information in the data storage unit 201A. In other words, the guide information can also be said to represent information that represents a superimposed image obtained by superimposing a figure that shows at least one of the following: the shape of an object contained in space, the trajectory of the object's movement, the relationship between objects, and the boundary in space onto the guide input image (at least one of the image representing space and the image representing sensor information).
[0040] (area division part) The region division unit 108A performs region division on the sensor information. For example, the region division unit 108A divides the image into multiple regions based on the feature quantities of the image represented by the sensor information.
[0041] (Target detection unit) The object detection unit 109A uses sensor information to detect objects present in space and generates data indicating the detection result. Examples of objects include aircraft, ships, drones, automobiles, robots, people, and animals. However, the objects are not limited to these. One example of the data generated by the object detection unit 109A is coordinate data that shows the area of ββthe object as a rectangle.
[0042] For example, if the sensor information is information indicating measurement results from radar or LiDAR, the target detection unit 109A detects the target based on the measurement results from radar or LiDAR. Also, if the sensor information is image data (multispectral image, infrared image, etc.), the target detection unit 109A detects the target using, for example, a method employing an object detection model such as YOLOX. The method employing an object detection model is not limited to YOLOX, and the target detection unit 109A may also detect the target using other methods such as YOLO (You Only Look Once), ViT (Vision Transformer), Faster R-CNN (Regions with CNN features), SSD (Single Shot MultiBox Detector).
[0043] (Target tracking unit) The object tracking unit 110A tracks the object detected by the object detection unit 109A by correlating the objects in a time series and generates data indicating the tracking results. The data indicating the tracking results is, for example, the time-series coordinates of each trajectory obtained by the tracking of the object tracking unit 110A. For example, if the sensor information is information indicating measurement results from radar or LiDAR, the object tracking unit 110A tracks the object using a method such as a Kalman filter. If the sensor information is image data, the object tracking unit 110A tracks the object using, for example, the ByteTrack method.
[0044] (Reference text generation unit) The reference text generation unit 111A generates reference text using sensor information. For example, the reference text is text describing the image represented by the sensor information. For example, the reference text generation unit 111A generates reference text based on output data obtained by inputting sensor information into a generation model. As a generation model, for example, a model generated using the BLIP-2 (Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models) method can be used. When the reference text generation unit 111A generates reference text using a generation model, the input data input to the generation model includes, for example, at least one of sensor information, spatial information, and environmental information. The output of the generation model includes reference text corresponding to the input information (image data, etc.).
[0045] (Text generation unit) The text generation unit 112A generates text describing the space using the sensor information, guide information, and environmental information. The text describing the space is, for example, a descriptive text focusing on an object or area corresponding to the guide information. This text may be used, for example, as a report in space monitoring operations. Alternatively, this text may be used, for example, as instructions related to operations in space monitoring operations.
[0046] As an example, the text generation unit 112A generates the above text based on output data obtained by inputting input data, including the above sensor information, the above guide information, and the above environmental information, into a large-scale language model. Examples of large-scale language models include, but are not limited to, generation AIs such as ChatGPT (Chat Generative Pre-trained Transformer), GPT-4 (Generative Pre-trained Transformer 4), and GPT-4o, or generation AIs that have been fine-tuned using environmental information and spatial information.
[0047] The large-scale language model may be stored in the storage unit 20A of the information processing device 1A, or it may be stored in a device other than the information processing device 1A. Here, when we say that the large-scale language model is stored in a memory device (such as the storage unit 20A), we mean that the parameters that define the large-scale language model are stored in the memory device. If the large-scale language model is stored in a device other than the information processing device 1A, the text generation unit 112A transmits input data to the device via the communication unit 30A, receives output data transmitted from the device, and generates the above text based on the received output data.
[0048] (Input to large-scale language models) The input data provided to the large-scale language model may include, in addition to sensor information, guide information, and environmental information, reference text generated by the reference text generation unit 111A. In other words, the text generation unit 112A can also generate the text using the reference text in addition to the sensor information, guide information, and environmental information.
[0049] Furthermore, the input data may include the division results from the region division unit 108A. In other words, the text generation unit 112A can generate the text using the division results from the region division unit 108A in addition to the sensor information, guide information, and environment information.
[0050] Furthermore, the input data may include at least one of the detection results from the target detection unit 109A and the tracking results from the target tracking unit 110A. The input data may also include accumulated past data. Here, the accumulated past data may include, as an example, at least one of the following: distance sensor information acquired by the distance sensor information acquisition unit 101A, image sensor information acquired by the image sensor information acquisition unit 102A, spatial information acquired by the spatial information acquisition unit 103A, environmental information acquired by the environmental information acquisition unit 104A, guide information acquired by the distance sensor information guide acquisition unit 105A, guide information acquired by the image sensor information guide acquisition unit 106A, spatial information acquired by the spatial information guide acquisition unit 107A, division results obtained by the area division unit 108A, detection results detected by the target detection unit 109A, trajectory obtained by the target tracking unit 110A, and reference text generated by the reference text generation unit 111A.
[0051] Furthermore, the input data may include instructional text. For example, instructional text may be: "Below are images with the tracked object superimposed, the date, the tracked object's trajectory, the user-entered guide, and environmental information (such as temperature). Please summarize these using past response text as a reference."
[0052] (Output of a large-scale language model) The output of a large-scale language model may include, as an example, text that provides an explanation (report, instructions regarding work, etc.) focusing on the events indicated by the guide information in the guide input image. An example of such text might be: "At xx / yy:zz, bb (target object) passed near point aa with a speed of cc (velocity, etc.). It may pass through dd in the future. For reference, a similar past instance occurred at ee / ff / gg / hh:jj, where it passed through point kk and then point mm. In that instance, the decision was made to take action nn (e.g., rescue)." The output of a large-scale language model may also include data other than text (image data, audio data, etc.).
[0053] (Output adjustment section) The output adjustment unit 113A uses guide information to tune at least one of the large-scale language model and the input data. More specifically, as an example, the output adjustment unit 113A acquires training text corresponding to the guide information and uses the guide information and training text to fine-tune or instruction-tune the large-scale language model. Alternatively, the output adjustment unit 113A may perform prompt tuning using the guide information and training text. Here, the training text is, as an example, text previously output by the large-scale language model.
[0054] As an example, the training data used for tuning includes multiple sets of first data, each set containing at least one of the following: image sensor information acquired by the image sensor information acquisition unit 102A, distance sensor information acquired by the distance sensor information acquisition unit 101A, spatial information acquired by the spatial information acquisition unit 103A, environmental information acquired by the environmental information acquisition unit 104A, guide information acquired by the distance sensor information guide acquisition unit 105A, guide information acquired by the image sensor information guide acquisition unit 106A, guide information acquired by the spatial information guide acquisition unit 107A, division results obtained by the region division unit 108A, detection results by the target detection unit 109A, tracking results by the target tracking unit 110A, and reference text generated by the reference text generation unit 111A, and text corresponding to the first data (training text).
[0055] (Text output section) The text output unit 114A outputs the text generated by the text generation unit 112A. For example, the text output unit 114A outputs the above text to a display connected to the output unit 50A, and displays the text on the display. Alternatively, the text output unit 114A may transmit the above text to another device connected via the communication unit 30A, and display the text on the display of that other device. Text A24 in Figure 6 is an example of text output by the text output unit 114A. In the example in Figure 6, text A24 is displayed on the display along with images A21 to A23, which have guide information superimposed on guide input images.
[0056] Furthermore, the text output unit 114A may output the text by writing it to a storage location specified by the user of the information processing device 1A (which may be a storage device within the information processing device 1A or a storage device outside the information processing device 1A). Alternatively, the text output unit 114A may output the text to an output device such as a speaker or a printer.
[0057] (Examples of practical applications) The information processing device 1A relating to this disclosure is applicable to various technological fields such as robotics, logistics systems, and drone control. For example, in the case of robotics, the guide information relating to this disclosure includes, as an example, information showing a figure (e.g., a connecting line) indicating the correspondence between a robot that moves an object and the object that the robot moves, and information showing a figure indicating a restricted area for the robot. In this case, the text generated by the text generation unit 112A may include, as an example, a report explaining the work performed by the robot or instructions related to the work.
[0058] Furthermore, in the case of a logistics system, for example, the guide information may include, as an example, information showing a diagram (e.g., a connecting line) that indicates the relationship between the delivery vehicle that delivers the products and the storage location of the products being delivered, and information indicating areas where passage is prohibited due to accidents, etc. In this case, as an example, the text generated by the text generation unit 112A may include work reports and instructions regarding delivery operations using delivery vehicles.
[0059] (Effects of information processing equipment) However, conventional technologies have the problem that even if text (such as reports) is created from spatial information, it does not result in explanatory text that focuses on important parts. Furthermore, in large-scale image / video data / radar information processing, it is difficult to output explanatory text (such as reports) that focuses on a specific area simply by controlling text prompts. In addition, it is difficult to input prompt engineering to correspond to small areas in the target data, as well as relationships between objects and people (context) and spatiotemporal correlations in text. In contrast, the information processing device 1A according to this disclosure can generate text that focuses on specific events in the monitored space by using guide information in addition to sensor information and environmental information.
[0060] Furthermore, the information processing device 1A includes a display control unit 115A that displays a first image represented by the distance measuring sensor information, a second image represented by the image sensor information, and a third image represented by spatial information representing space on a display device. The distance measuring sensor information guide acquisition unit 105A, the image sensor information guide acquisition unit 106A, and the spatial information guide acquisition unit 107A are configured to acquire user-inputted guide information for at least one of the first, second, and third images displayed by the display control unit 115A.
[0061] Conventional technologies, even when generating text (such as reports) from spatial information, may not result in explanatory text that focuses on the important parts. In contrast, the information processing device 1A related to this disclosure can generate text that reflects the user's intent by using guide information entered by the user.
[0062] Furthermore, the information processing device 1A includes an object detection unit 109A that detects objects in space using sensor information, and an object tracking unit 110A that tracks the objects detected by the object detection unit 109A. The guide information includes information indicating at least one of the detection results by the object detection unit 109A and the tracking results by the object tracking unit 110A. Therefore, the information processing device 1A can generate text that focuses on the object detection results or the object tracking results.
[0063] Furthermore, the information processing device 1A is further equipped with a reference text generation means that generates reference text based on output data obtained by inputting sensor information into a generation model, and the text generation unit 112A is configured to generate the text using the reference text in addition to the sensor information, guide information, and environmental information. Therefore, the information processing device 1A can generate text with greater accuracy by using the reference text.
[0064] Furthermore, the information processing device 1A includes a region division unit 108A that performs region division on sensor information, and the text generation unit 112A generates text using the sensor information, guide information, and environmental information, as well as the division results from the region division unit 108A. Therefore, the information processing device 1A can generate text with greater accuracy by using the regions obtained through region division.
[0065] Furthermore, the information processing device 1A employs a configuration in which the text generation unit 112A generates text based on output data obtained by inputting input data including sensor information, guide information, and environmental information into a large-scale language model, and the output adjustment unit 113A tunes at least one of the large-scale language model and the input data using the guide information. Therefore, with the information processing device 1A, text can be generated with greater accuracy by using the large-scale language model updated by the output adjustment unit 113A.
[0066] Furthermore, in the information processing device 1A, the output adjustment unit 113A acquires training text corresponding to guide information and uses the guide information and training text to fine-tune or instruction-tune the large-scale language model. Therefore, with the information processing device 1A, by using the large-scale language model updated by the output adjustment unit 113A, text can be generated with greater accuracy.
[0067] Furthermore, in the information processing device 1A, the guide information is configured to represent a superimposed image, in which a figure indicating at least one of the followingβthe shape of an object contained in the space, the trajectory of the object's movement, the relationship between objects, and the boundary in the spaceβis superimposed on at least one of the images representing the space and the image representing the sensor information. Therefore, the information processing device 1A can generate text that focuses on specific events in the space being monitored.
[0068] [Examples of implementation using software] Some or all of the functions of the information processing devices 1, 1A, and 2 (hereinafter also referred to as "each of the above devices") may be implemented by hardware such as integrated circuits (IC chips) or by software.
[0069] In the latter case, each of the above devices is implemented, for example, by a computer that executes instructions for a program, which is software that realizes each function. An example of such a computer (hereinafter referred to as Computer C) is shown in Figure 7. Figure 7 is a block diagram showing the hardware configuration of Computer C, which functions as each of the above devices.
[0070] Computer C comprises at least one processor C1 and at least one memory C2. Memory C2 stores a program P that causes computer C to operate as each of the above-mentioned devices. In computer C, processor C1 reads program P from memory C2 and executes it, thereby realizing each of the above-mentioned devices.
[0071] For processor C1, for example, a CPU (Central Processing Unit), GPU (Graphic Processing Unit), DSP (Digital Signal Processor), MPU (Micro Processing Unit), FPU (Floating Point Number Processing Unit), PPU (Physics Processing Unit), TPU (Tensor Processing Unit), quantum processor, microcontroller, or a combination thereof can be used. For memory C2, for example, flash memory, HDD (Hard Disk Drive), SSD (Solid State Drive), or a combination thereof can be used.
[0072] Computer C may also be equipped with RAM (Random Access Memory) for loading program P at runtime and for temporarily storing various data. Furthermore, computer C may be equipped with communication interfaces for sending and receiving data with other devices. Additionally, computer C may be equipped with input / output interfaces for connecting input / output devices such as keyboards, mice, displays, and printers.
[0073] Furthermore, program P can be recorded on a non-temporary, tangible recording medium M that is readable by computer C. Such a recording medium M could be, for example, tape, disk, card, semiconductor memory, or programmable logic circuitry. Computer C can acquire program P via such a recording medium M. Program P can also be transmitted via a transmission medium. Such a transmission medium could be, for example, a communication network or broadcast waves. Computer C can also acquire program P via such a transmission medium.
[0074] Furthermore, each of the above functions of each of the above devices may be implemented by a single processor in a single computer, by multiple processors in a single computer working together, or by multiple processors in each of multiple computers working together. In addition, the programs for implementing each of the above functions in each of the above devices may be stored in a single memory in a single computer, distributed and stored in multiple memories in a single computer, or distributed and stored in multiple memories in each of multiple computers.
[0075] [Additional Note A] This disclosure includes the technologies described in the following appendices. However, the present invention is not limited to the technologies described in the following appendices, and various modifications are possible within the scope of the claims.
[0076] (Note A1) An acquisition means for acquiring sensor information showing the sensing results from a sensor that senses space, and environmental information regarding the environment of the space, A guide information acquisition means that acquires guide information indicating at least one of the following: the shape of an object contained in the space, the trajectory of the object's movement, the relationship between objects, and the boundary in the space; A text generation means that generates text describing the space using the sensor information, the guide information, and the environmental information, A text output means that outputs the text generated by the text generation means, An information processing device equipped with the following features.
[0077] (Appendix A2) The aforementioned sensor information includes distance measurement sensor information acquired from radar and image sensor information acquired from an image sensor. The device further includes a display control means for displaying a first image represented by the distance measurement sensor information, a second image represented by the image sensor information, and a third image represented by spatial information representing the space, The guide information acquisition means acquires the guide information entered by the user for at least one of the first image, the second image, and the third image displayed by the display control means. The information processing device described in Appendix A1.
[0078] (Note A3) A target detection means for detecting an object present in the space using the aforementioned sensor information, The object detection means further comprises an object tracking means for tracking an object detected by the object detection means, The guide information includes information indicating at least one of the detection results by the target detection means and the tracking results by the target tracking means. The information processing device described in Appendix A1 or A2.
[0079] (Note A4) The system further comprises a reference text generation means that generates reference text based on output data obtained by inputting the aforementioned sensor information into a generation model, The text generation means generates the text using the reference text in addition to the sensor information, the guide information, and the environmental information. An information processing device as described in any one of the appendices A1 to A3.
[0080] (Note A5) The system further comprises region division means for performing region division on the sensor information, The text generation means generates the text using the sensor information, the guide information, and the environmental information, as well as the division results by the region division means. An information processing device as described in one of the appendices A1 through A4.
[0081] (Note A6) The text generation means generates the text based on output data obtained by inputting input data, including the sensor information, the guide information, and the environmental information, into a large-scale language model. The system further comprises output adjustment means for tuning at least one of the large-scale language model and the input data using the guide information. An information processing device as described in one of the appendices A1 to A5.
[0082] (Note A7) The output adjustment means is The training text corresponding to the guide information is obtained, and the large-scale language model is fine-tuned or instruction-tuned using the guide information and the training text. The information processing device described in Appendix A6.
[0083] (Note A8) The guide information is information representing a superimposed image obtained by superimposing a figure showing at least one of the following onto at least one of the images representing the space and the image representing the sensor information: the shape of an object included in the space, the trajectory of the object's movement, the relationship between objects, and the boundary in the space. An information processing device as described in any one of the appendices A1 to A7.
[0084] [Additional Note B] This disclosure includes the technologies described in the following appendices. However, the present invention is not limited to the technologies described in the following appendices, and various modifications are possible within the scope of the claims.
[0085] (Note B1) At least one processor performs an acquisition process to acquire sensor information indicating sensing results from a sensor that senses space, and environmental information regarding the environment of the space. The at least one processor performs a guide information acquisition process to acquire guide information indicating at least one of the following: the shape of an object included in the space, the trajectory of the object's movement, the relationship between objects, and the boundary in the space. The at least one processor performs a text generation process that generates text describing the space using the sensor information, the guide information and the environment information, The at least one processor performs a text output process that outputs the text generated in the text generation process, Information processing methods including
[0086] (Note B2) The aforementioned sensor information includes distance measurement sensor information acquired from radar and image sensor information acquired from an image sensor. The at least one processor further includes display control processing for displaying a first image represented by the distance sensor information, a second image represented by the image sensor information, and a third image represented by the spatial information representing the space on a display device, In the guide information acquisition process, the at least one processor acquires the guide information entered by the user for at least one of the first image, the second image, and the third image displayed in the display control process. The information processing method described in Appendix B1.
[0087] (Note B3) The at least one processor performs an object detection process that uses the sensor information to detect an object present in the space, The at least one processor further includes an object tracking process that tracks the object detected in the object detection process, The guide information includes information indicating at least one of the detection results from the target detection process and the tracking results from the target tracking process. The information processing method described in Appendix B1 or B2.
[0088] (Note B4) The at least one processor further includes a reference text generation process that generates reference text based on output data obtained by inputting the sensor information into a generation model, In the text generation process, the at least one processor generates the text using the reference text in addition to the sensor information, the guide information, and the environment information. The information processing method described in any one of the appendices B1 to B3.
[0089] (Note B5) The at least one processor further includes a region division process that divides the sensor information into regions, In the text generation process, the at least one processor generates the text using the sensor information, the guide information, and the environment information, as well as the division results from the region division process. The information processing method described in any one of the appendices B1 to B4.
[0090] (Note B6) In the text generation process, the at least one processor generates the text based on output data obtained by inputting the input data, including the sensor information, the guide information, and the environmental information, into a large-scale language model. The at least one processor further includes an output adjustment process that uses the guide information to tune at least one of the large-scale language model and the input data. The information processing method described in one of the appendices B1 through B5.
[0091] (Note B7) In the output adjustment process, at least one processor, The training text corresponding to the guide information is obtained, and the large-scale language model is fine-tuned or instruction-tuned using the guide information and the training text. The information processing method described in Appendix B6.
[0092] (Note B8) The guide information is information representing a superimposed image obtained by superimposing a figure showing at least one of the following onto at least one of the images representing the space and the image representing the sensor information: the shape of an object included in the space, the trajectory of the object's movement, the relationship between objects, and the boundary in the space. The information processing method described in any one of the appendices B1 to B7.
[0093] [Additional Note C] This disclosure includes the technologies described in the following appendices. However, the present invention is not limited to the technologies described in the following appendices, and various modifications are possible within the scope of the claims.
[0094] (Note C1) A program for causing a computer to function as an information processing device, wherein the computer, An acquisition means for acquiring sensor information showing the sensing results from a sensor that senses space, and environmental information regarding the environment of the space, A guide information acquisition means that acquires guide information indicating at least one of the following: the shape of an object contained in the space, the trajectory of the object's movement, the relationship between objects, and the boundary in the space; A text generation means that generates text describing the space using the sensor information, the guide information, and the environmental information, A text output means that outputs the text generated by the text generation means, An information processing program that functions as such.
[0095] (Note C2) The aforementioned sensor information includes distance measurement sensor information acquired from radar and image sensor information acquired from an image sensor. The aforementioned computer, The display control means further functions to display the first image represented by the distance measurement sensor information, the second image represented by the image sensor information, and the third image represented by the spatial information representing the space on a display device, The guide information acquisition means acquires the guide information entered by the user for at least one of the first image, the second image, and the third image displayed by the display control means. The information processing program described in Appendix C1.
[0096] (Note C3) The aforementioned computer, A target detection means for detecting an object present in the space using the aforementioned sensor information, The object detection means further functions as an object tracking means that tracks the object detected by the object detection means, The guide information includes information indicating at least one of the detection results by the target detection means and the tracking results by the target tracking means. The information processing program described in Appendix C1 or C2.
[0097] (Note C4) The aforementioned computer, The aforementioned sensor information is input to a generation model and, based on the output data obtained, is further configured to function as a reference text generation means, which generates reference text based on the output data. The text generation means generates the text using the reference text in addition to the sensor information, the guide information, and the environmental information. An information processing program described in any one of the appendices C1 to C3.
[0098] (Note C5) The aforementioned computer, Furthermore, it functions as a region division means that performs region division on the aforementioned sensor information, The text generation means generates the text using the sensor information, the guide information, and the environmental information, as well as the division results by the region division means. An information processing program described in any one of the appendices C1 to C4.
[0099] (Appendix C6) The text generation means generates the text based on output data obtained by inputting input data, including the sensor information, the guide information, and the environmental information, into a large-scale language model. The aforementioned computer, The aforementioned guide information is used to further function as an output adjustment means for tuning at least one of the large-scale language model and the input data. An information processing program described in any one of the appendices C1 to C5.
[0100] (Note C7) The output adjustment means is The training text corresponding to the guide information is obtained, and the large-scale language model is fine-tuned or instruction-tuned using the guide information and the training text. The information processing program described in Appendix C6.
[0101] (Note C8) The guide information is information representing a superimposed image obtained by superimposing a figure showing at least one of the following onto at least one of the images representing the space and the image representing the sensor information: the shape of an object included in the space, the trajectory of the object's movement, the relationship between objects, and the boundary in the space. An information processing program described in any one of the appendices C1 to C7.
[0102] [Additional Note D] This disclosure includes the technologies described in the following appendices. However, the present invention is not limited to the technologies described in the following appendices, and various modifications are possible within the scope of the claims.
[0103] (Note D1) It comprises at least one processor, and the at least one processor is An acquisition process that acquires sensor information showing the sensing results from a sensor that senses the space, and environmental information regarding the environment of that space. A guide information acquisition process that acquires guide information indicating at least one of the following: the shape of an object included in the space, the trajectory of the object's movement, the relationship between objects, and the boundary in the space; A text generation process that generates text describing the space using the sensor information, the guide information, and the environmental information, A text output process that outputs the text generated in the text generation process, An information processing device that performs the following actions.
[0104] The information processing device may also include memory. Furthermore, the memory may store a program that causes at least one processor to execute each of the aforementioned processes.
[0105] (Note D2) The aforementioned sensor information includes distance measurement sensor information acquired from radar and image sensor information acquired from an image sensor. The aforementioned at least one processor, Further, display control processing is performed to display the first image represented by the distance measurement sensor information, the second image represented by the image sensor information, and the third image represented by the spatial information representing the space on a display device. In the guide information acquisition process, the at least one processor acquires the guide information entered by the user for at least one of the first image, the second image, and the third image displayed in the display control process. The information processing device described in Appendix D1.
[0106] (Note D3) The aforementioned at least one processor, A target detection process that uses the aforementioned sensor information to detect an object present in the space, Further, the object tracking process is performed to track the object detected in the object detection process, The guide information includes information indicating at least one of the detection results from the target detection process and the tracking results from the target tracking process. The information processing device described in Appendix D1 or D2.
[0107] (Note D4) The aforementioned at least one processor, A reference text generation process is further executed, which generates reference text based on the output data obtained by inputting the aforementioned sensor information into the generation model. In the text generation process, the at least one processor generates the text using the reference text in addition to the sensor information, the guide information, and the environment information. An information processing device as described in any one of the appendices D1 to D3.
[0108] (Note D5) The aforementioned at least one processor, A region segmentation process is further performed on the aforementioned sensor information, In the text generation process, the at least one processor generates the text using the sensor information, the guide information, and the environment information, as well as the division results from the region division process. An information processing device as described in any one of the appendices D1 to D4.
[0109] (Note D6) In the text generation process, the at least one processor generates the text based on output data obtained by inputting the input data, including the sensor information, the guide information, and the environmental information, into a large-scale language model. The aforementioned at least one processor, Further output adjustment processing is performed to tune at least one of the large-scale language model and the input data using the aforementioned guide information. An information processing device as described in any one of the appendices D1 to D5.
[0110] (Note D7) In the output adjustment process, at least one processor, The training text corresponding to the guide information is obtained, and the large-scale language model is fine-tuned or instruction-tuned using the guide information and the training text. The information processing device described in Appendix D6.
[0111] (Note D8) The guide information is information representing a superimposed image obtained by superimposing a figure showing at least one of the following onto at least one of the images representing the space and the image representing the sensor information: the shape of an object included in the space, the trajectory of the object's movement, the relationship between objects, and the boundary in the space. An information processing device as described in any one of the appendices D1 to D7.
[0112] [Additional Note E] This disclosure includes the technologies described in the following appendices. However, the present invention is not limited to the technologies described in the following appendices, and various modifications are possible within the scope of the claims.
[0113] (Note E1) A program for causing a computer to function as an information processing device, wherein the computer, An acquisition process that acquires sensor information showing the sensing results from a sensor that senses the space, and environmental information regarding the environment of that space. A guide information acquisition process that acquires guide information indicating at least one of the following: the shape of an object included in the space, the trajectory of the object's movement, the relationship between objects, and the boundary in the space; A text generation process that generates text describing the space using the sensor information, the guide information, and the environmental information, A text output process that outputs the text generated in the text generation process, A non-temporary recording medium that stores an information processing program for executing [the specified action]. [Explanation of Symbols]
[0114] 1. 1A Information Processing Device 11 Acquisition Department 12 Guide Information Acquisition Section 13, 112A Text generation unit 14, 114A Text output section 101A Distance measurement sensor information acquisition unit 102A Image sensor information acquisition unit 103A Spatial information acquisition unit 104A Environmental Information Acquisition Department 105A Distance measurement sensor information guide acquisition unit 106A Image sensor information guide acquisition unit 107A Spatial Information Guide Acquisition Unit 108A Area division part 109A Target detection unit 110A Target Tracking Unit 111A Reference Text Generation Unit 113A output adjustment section 115A Display Control Unit 201A Data Storage Unit
Claims
1. An acquisition means for acquiring sensor information showing the sensing results from a sensor that senses space, and environmental information regarding the environment of the space, A guide information acquisition means that acquires guide information indicating at least one of the following: the shape of an object contained in the space, the trajectory of the object's movement, the relationship between objects, and the boundary in the space; A text generation means that generates text describing the space using the sensor information, the guide information, and the environmental information, A text output means that outputs the text generated by the text generation means, An information processing device equipped with the following features.
2. The sensor information includes distance sensor information acquired from a distance sensor and image sensor information acquired from an image sensor. The system further includes a display control means for displaying a first image represented by the distance measurement sensor information, a second image represented by the image sensor information, and a third image represented by the spatial information representing the space on a display device, The guide information acquisition means acquires the guide information entered by the user for at least one of the first image, the second image, and the third image displayed by the display control means. The information processing apparatus according to claim 1.
3. A target detection means for detecting an object present in the space using the aforementioned sensor information, The object detection means further comprises an object tracking means for tracking an object detected by the object detection means, The guide information includes information indicating at least one of the detection results by the target detection means and the tracking results by the target tracking means. The information processing apparatus according to claim 1 or 2.
4. The system further comprises a reference text generation means that generates reference text based on output data obtained by inputting the aforementioned sensor information into a generation model, The text generation means generates the text using the reference text in addition to the sensor information, the guide information, and the environmental information. The information processing apparatus according to claim 1 or 2.
5. The system further comprises region division means for performing region division on the sensor information, The text generation means generates the text using the sensor information, the guide information, and the environmental information, as well as the division results by the region division means. The information processing apparatus according to claim 1 or 2.
6. The text generation means generates the text based on output data obtained by inputting input data, including the sensor information, the guide information, and the environmental information, into a large-scale language model. The system further comprises output adjustment means for tuning at least one of the large-scale language model and the input data using the guide information. The information processing apparatus according to claim 1 or 2.
7. The output adjustment means is The training text corresponding to the guide information is obtained, and the large-scale language model is fine-tuned or instruction-tuned using the guide information and the training text. The information processing apparatus according to claim 6.
8. The guide information is information representing a superimposed image obtained by superimposing a figure showing at least one of the following onto at least one of the images representing the space and the image representing the sensor information: the shape of an object included in the space, the trajectory of the object's movement, the relationship between objects, and the boundary in the space. The information processing apparatus according to claim 1 or 2.
9. At least one processor performs an acquisition process to acquire sensor information indicating sensing results from a sensor that senses space, and environmental information regarding the environment of the space. The at least one processor performs a guide information acquisition process to acquire guide information indicating at least one of the following: the shape of an object included in the space, the trajectory of the object's movement, the relationship between objects, and the boundary in the space. The at least one processor performs a text generation process that generates text describing the space using the sensor information, the guide information and the environment information, The at least one processor performs a text output process that outputs the text generated in the text generation process, Information processing methods including
10. An information processing program for causing a computer to function as an information processing device, wherein the computer, An acquisition means for acquiring sensor information showing the sensing results from a sensor that senses space, and environmental information regarding the environment of the space, A guide information acquisition means that acquires guide information indicating at least one of the following: the shape of an object contained in the space, the trajectory of the object's movement, the relationship between objects, and the boundary in the space; A text generation means that generates text describing the space using the sensor information, the guide information, and the environmental information, A text output means that outputs the text generated by the text generation means, An information processing program that functions as such.