Information Processing Apparatus, Information Processing Method, and Program
The information processing apparatus improves position measurement accuracy by aligning movement trajectories using language records to identify shared locations, addressing the limitation of lacking a geographical database.
Patent Information
- Application Number
- JP2025012208
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-01-28
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-01-28
AI Technical Summary
Existing methods for correcting position information are inadequate when a geographical database is not provided, leading to inaccuracies in position measurement.
An information processing apparatus that corrects movement trajectories using language records to identify phrases or sets of phrases indicating the same location, aligning positions at those times to improve accuracy without relying on a geographical database.
Enhances the accuracy of position information by aligning movement trajectories based on language records, even when a geographical database is unavailable.
Smart Images

Figure 0007713607000001_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to an information processing device, an information processing method, and a program. [Background technology]
[0002] Various methods for measuring a position have been proposed. For example, Patent Document 1 proposes a method for measuring a position using measurement data from an inertial sensor. Other methods for measuring a position include a method using a satellite positioning module (GPS sensor, etc.), a method using wireless communication (Bluetooth (registered) Examples of such methods include using beacons (registered trademarks) and using SLAM. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2014-167461 A [Non-patent literature]
[0004] [Non-Patent Document 1] Smitha Sheshadri, et al. “Conversational Localization: Indoor Human Localization through Intelligent Conversation”, [online], [Retrieved January 16, 2025], Internet<URL:https: / / dl.acm.org / doi / pdf / 10.1145 / <3631404> Summary of the Invention [Problem to be solved by the invention]
[0005] Errors can occur in position measurement. Therefore, various methods have been proposed to reduce errors. For example, Non-Patent Document 1 proposes a method of correcting position information by using conversation information and a geographical database (a database of geographically indexed entities). Specifically, the method proposed in Non-Patent Document 1 extracts words indicating entities in the real space from the conversation information, collates the entities indicated by the extracted words with the geographical database, and corrects the measured position information based on the positions of the entities identified from the geographical database as a result of the collation. As a result, the accuracy of the measured position information can be improved. However, in the method proposed in Non-Patent Document 1, it is difficult to correct the position information if the geographical database is not provided.
[0006] In one aspect, the present disclosure has been made in view of such circumstances. One of its purposes is to provide a technique for correcting position information even when a geographical database is not provided in advance.
Means for Solving the Problems
[0007] In order to solve the above-described problems, the present disclosure adopts the following configuration. Note that the following configurations can be combined as appropriate.
[0008] An information processing apparatus according to one aspect of the present disclosure includes a control unit. The control unit is configured to acquire one or more movement trajectories respectively measured in time series by a positioning sensor in each of one or more agents, acquire one or more language records respectively measured in synchronization with each of the one or more movement trajectories in each of the one or more agents, extract one or more combinations of times related to the same location by identifying a phrase or a set of phrases indicating the same location in the acquired one or more language records, correct the one or more movement trajectories so that the positions at the times of each of the extracted one or more combinations approach each other, and output the corrected one or more movement trajectories.
[0009] For example, assume that the language record of one agent contains a phrase "Returned to the location five minutes ago". This phrase "Returned to the location five minutes ago" serves as a hint that the agent was located at the same place at the time when this phrase was uttered and five minutes before that time. Thus, there is a possibility that a phrase serving as a hint that the agent was located at the same place at multiple times can be obtained from the language record. Also, there is a possibility that a set of phrases serving as hints that the agent was located at the same place at multiple times, such as "Came to make coffee" and "Came to get the made coffee", can be obtained from the language record. Similarly, when there are multiple agents, there is a possibility that a phrase serving as a hint that at least a part of the multiple agents was located at the same place at the target time (combination of times), such as "Just joined BBB (the second agent) now" and "BBB was using this room until five minutes ago", can be obtained from the language record. There is a possibility that a set of phrases serving as hints that at least a part of the multiple agents was located at the same place at the target time, such as the utterance of the first agent "Took out pork from the refrigerator", the utterance of the second agent "Took out beef from the refrigerator", the utterance of the first agent "Arrived at the refrigerator", and the utterance of the second agent "AAA (the first agent) is still taking out luggage from the refrigerator", can be obtained from the language record. In the said configuration, by using such a phrase or set of phrases indicating the same place as a hint that the agent was located at the same place, the movement trajectory of one or more agents is corrected. That is, by specifying in the language record a phrase or set of phrases indicating the same place, one or more combinations of times related to the same place are extracted. Then, the movement trajectories of one or more agents are corrected so that the positions at the times of each of the one or more combinations approach each other. As a result, according to this configuration, although a geographical database may be used, it is possible to correct the movement trajectories (position information) of one or more agents without using a geographical database. As a result, it is possible to improve the accuracy of the measured position information (movement trajectory).
[0010] In the information processing apparatus according to the above aspect, the one or more agents may include a first agent. The one or more movement trajectories may include a first movement trajectory of the first agent. The one or more language records may include a first language record of the first agent. Extracting one or more combinations of times related to the same location may include extracting one or more combinations of times related to the same location by specifying a phrase or a set of phrases included in the first language record. Correcting the one or more movement trajectories may include correcting the first movement trajectory so that the positions at the times of the extracted one or more combinations approach each other. According to this configuration, it is possible to improve the accuracy of the position information (first movement trajectory) within one agent (the first agent).
[0011] In the information processing apparatus according to the above aspect, specifying a phrase or a set of phrases included in the first language record may include specifying a first phrase indicating being located at the same location at a plurality of times. The one or more combinations to be extracted may include a first combination constituted by a plurality of times corresponding to the specified first phrase. According to this configuration, by using the first phrase indicating being located at the same location, the first movement trajectory of the first agent can be appropriately corrected.
[0012] In the information processing apparatus according to the above aspect, specifying a phrase or a set of phrases included in the first language record may include specifying a second phrase and a third phrase indicating the same location as each other. The one or more combinations to be extracted may include a second combination constituted by the time corresponding to the specified second phrase and the time corresponding to the third phrase. According to this configuration, by using the second phrase and the third phrase (set of phrases) indicating the same location as each other, the first movement trajectory of the first agent can be appropriately corrected.
[0013] In the information processing apparatus according to the above aspect, the second phrase may indicate that it is located at the same position as the third phrase using a word different from the third phrase. For example, it is possible to refer to the same position with different words such as "opened the refrigerator" and "lit the stove" (both referring to being in the kitchen). According to this configuration, by using this, the first movement trajectory of the first agent can be appropriately corrected.
[0014] In the information processing apparatus according to the above aspect, the one or more agents may include the first agent and the second agent. The one or more movement trajectories may include the first movement trajectory of the first agent and the second movement trajectory of the second agent. The one or more language records may include the first language record of the first agent and the second language record of the second agent. Extracting one or more combinations of times related to the same location may include identifying a phrase or a set of phrases included in at least one of the first language record and the second language record to extract one or more combinations of times related to the same location. Correcting the one or more movement trajectories may include correcting at least one of the first movement trajectory and the second movement trajectory so that the positions at the times of the extracted one or more combinations approach each other. According to this configuration, it is possible to improve the accuracy of the position information (the first movement trajectory, the second movement trajectory) among a plurality of agents (the first agent, the second agent).
[0015] In the information processing apparatus according to the above aspect, specifying a phrase or a set of phrases included in at least one of the first language record and the second language record may include specifying, from the first language record, a fourth phrase indicating that the first agent is in the same location as the second agent, or specifying, from the second language record, a fifth phrase indicating that the second agent is in the same location as the first agent. One or more combinations to be extracted may include a third combination constituted by the time of the first agent and the time of the second agent corresponding to the specified fourth phrase or fifth phrase. According to this configuration, by using the fourth phrase or the fifth phrase, at least one of the first movement trajectory of the first agent and the second movement trajectory of the second agent can be appropriately corrected.
[0016] In the information processing apparatus according to the above aspect, specifying a phrase or a set of phrases included in at least one of the first language record and the second language record may include specifying a sixth phrase from the first language record and specifying, from the second language record, a seventh phrase indicating the same location as the sixth phrase. One or more combinations to be extracted may include a fourth combination constituted by the time corresponding to the sixth phrase and the time corresponding to the seventh phrase. According to this configuration, by using the sixth phrase and the seventh phrase (a set of phrases), at least one of the first movement trajectory of the first agent and the second movement trajectory of the second agent can be appropriately corrected.
[0017] In the information processing apparatus according to the above aspect, the sixth phrase may indicate the same location as the seventh phrase using words different from those of the seventh phrase. It is possible to refer to the same location even with different words. According to this configuration, by using this, at least one of the first movement trajectory of the first agent and the second movement trajectory of the second agent can be appropriately corrected.
[0018] In the information processing apparatus according to the above aspect, the positioning sensor may include an inertial sensor. Each of the one or more movement trajectories may be measured by analyzing the measurement data of the inertial sensor. According to the inertial sensor, the position can be measured even indoors. Therefore, according to this configuration, the movement trajectory (position information) can be obtained in a scene where the agent exists either indoors or outdoors.
[0019] In the information processing apparatus according to the above aspect, the positioning sensor may include at least one of a camera and LiDAR (light detection and ranging). Each of the one or more movement trajectories may be measured by performing SLAM (simultaneous localization and mapping) analysis on the measurement data of at least one of the camera and LiDAR. According to at least one of the camera and LiDAR, the position can be measured even indoors. Therefore, according to this configuration, the movement trajectory (position information) can be obtained in a scene where the agent exists either indoors or outdoors.
[0020] Note that the embodiments of the present disclosure are not necessarily limited to the above information processing apparatus (correction apparatus). As another aspect of the information processing apparatus according to each of the above aspects, one aspect of the present disclosure may be an information processing method (correction method) that implements all or part of the above configurations, or a program, or a machine-readable storage medium such as a computer that stores such a program. Here, the machine-readable storage medium may be a non-temporary medium that stores information such as a program by an electrical, magnetic, optical, mechanical, or chemical action. The non-temporary storage medium may include storage media (CD, DVD, semiconductor memory, etc.), auxiliary storage devices of a computer, external storage devices connected to the computer, and the like.
[0021] For example, an information processing method according to an aspect of the present disclosure may be executed by a computer. The information processing method includes: obtaining one or more movement trajectories respectively measured in time series by a positioning sensor in each of one or more agents; obtaining one or more language records respectively measured in synchronization with each of the one or more movement trajectories in each of the one or more agents; extracting one or more combinations of times related to the same location by identifying a phrase or a set of phrases indicating the same location in the obtained one or more language records; correcting the one or more movement trajectories so that the positions of the times of each of the extracted one or more combinations approach each other; and outputting the corrected one or more movement trajectories.
[0022] Also, for example, a program according to an aspect of the present disclosure may be a program for causing a computer to execute the information processing method. The information processing method includes: obtaining one or more movement trajectories respectively measured in time series by a positioning sensor in each of one or more agents; obtaining one or more language records respectively measured in synchronization with each of the one or more movement trajectories in each of the one or more agents; extracting one or more combinations of times related to the same location by identifying a phrase or a set of phrases indicating the same location in the obtained one or more language records; correcting the one or more movement trajectories so that the positions of the times of each of the extracted one or more combinations approach each other; and outputting the corrected one or more movement trajectories.
Advantages of the Invention
[0023] According to an aspect of the present disclosure, it is possible to provide a technique for correcting position information even when a geographical database is not provided in advance.
Brief Description of the Drawings
[0024]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Mode for Carrying Out the Invention
[0025] Hereinafter, embodiments according to one aspect of the present disclosure will be described with reference to the drawings. However, the embodiments described below are merely examples of the present disclosure in all respects. Various improvements or modifications may be made without departing from the scope of the present disclosure. In implementing the present disclosure, a specific configuration according to the embodiment may be appropriately adopted. Note that the data appearing in this embodiment is described in natural language, but more specifically, it is specified by a pseudo language, command, parameter, machine language, electrical signal, etc. that can be recognized by a machine such as a computer.
[0026] §1 Application Example Figure 1 schematically shows an example of a scene to which the present disclosure is applied. The information processing apparatus 1 according to this embodiment is one or more computers configured to correct the movement trajectory 20 (position information) of the agent Z.
[0027] The information processing apparatus 1 according to the present embodiment acquires one or more movement trajectories 20 respectively measured in time series by the positioning sensor M1 in each of one or more agents Z. The information processing apparatus 1 acquires one or more language records 25 measured in synchronization with each of the one or more movement trajectories 20 in each of the one or more agents Z. The information processing apparatus 1 extracts one or more combinations 40 of times related to the same location by specifying a phrase 30 or a set 35 of phrases indicating the same location in the acquired one or more language records 25. The information processing apparatus 1 corrects the one or more movement trajectories 20 so that the positions of the times of each of the extracted one or more combinations 40 approach each other. Then, the information processing apparatus 1 outputs the corrected one or more movement trajectories 20.
[0028] According to the present embodiment, by using the phrase 30 or the set 35 of phrases that appear in the language record 25 synchronized with the movement trajectory 20 as a hint that they are located at the same location, the movement trajectory 20 (position information) of the agent Z is corrected. Thereby, the movement trajectory 20 (position information) of the agent Z can be corrected without using a geographical database. As a result, the accuracy of the measured movement trajectory 20 (position information) can be improved. Note that, in the present embodiment, the use of the geographical database may not be prohibited. The information processing apparatus 1 may further use the geographical database to correct the movement trajectory 20 (position information). Thereby, further improvement in the accuracy of the measured movement trajectory 20 (position information) may be achieved.
[0029] [Agent] Agent Z may be any object that can move manually or automatically. Agent Z may be capable of autonomous movement. For example, Agent Z may be composed of at least one of any living organism and any device. Any living organism may include a human. Any device may include a movable device. When composed of a movable device, Agent Z may be provided with known components used for movement such as a propeller, wheels, and a driving device. In one example, Agent Z may be a user (which may include a worker, etc.), a moving body, etc. The moving body may be a device configured to be operable or autonomously movable. In one example, the moving body may include an aircraft, a vehicle, a mobile robot, etc. The aircraft may include an unmanned aircraft such as a drone. The vehicle may include a manually driven vehicle and an autonomously driven vehicle. The mobile robot may include movable working robots such as a transport robot (including a delivery robot and a food delivery robot), a guiding robot, and a cleaning robot. The moving body may be used for any purpose such as manufacturing, logistics and transportation (warehousing, food delivery, distribution, etc.), service industry (restaurants, retail, etc.), medical care and nursing care, guiding and security (stations, streets, commercial facilities, etc.), and agriculture, forestry, and fisheries.
[0030] [Positioning Sensor] The positioning sensor M1 is used for measuring the position. As long as the position can be measured, the type of the positioning sensor M1 may not be particularly limited and may be appropriately selected according to the embodiment.
[0031] (I) Inertial Sensor As shown in FIG. 1, in one example, the positioning sensor M1 may include an inertial sensor M10. The type of the inertial sensor M10 may be arbitrarily selected. For example, the inertial sensor M10 may be configured to detect three-dimensional inertial motion (translational motion and rotational motion in three axial directions) by including an acceleration sensor and a gyro sensor (angular velocity sensor). The inertial sensor M10 may be an inertial measurement unit (IMU).
[0032] When the positioning sensor M1 includes the inertial sensor M10, each of the one or more movement trajectories 20 may be measured by analyzing the measurement data of the inertial sensor M10. The method for analyzing the measurement data of the inertial sensor M10 may not be particularly limited and may be appropriately selected according to the embodiment. For the analysis method, for example, known methods such as inertial navigation and pedestrian dead reckoning (PDR) may be adopted. For the analysis of the measurement data, a trained machine learning model may be used. The machine learning model is configured to have one or more arithmetic parameters adjustable by machine learning. The one or more arithmetic parameters are used for the arithmetic of the target inference (estimation of position). The machine learning model may be constituted by, for example, a neural network, a regression model, a decision tree model, a support vector machine, or other functional expressions. The method of machine learning may be appropriately selected according to the machine learning model adopted (for example, the error backpropagation method, etc.). In machine learning, the value of each arithmetic parameter of the machine learning model may be appropriately adjusted (optimized) using training samples so as to acquire the ability to derive position information from the measured values of the inertial sensor M10.
[0033] According to the inertial sensor M10, the position can be measured regardless of whether it is outdoors or indoors. Therefore, according to an example of this embodiment, the movement trajectory 20 (position information) can be obtained in a scene where the agent Z exists either indoors or outdoors. Also, compared with other sensors such as a camera, the inertial sensor M10 consumes less energy. Therefore, according to an example of this embodiment, energy consumption savings can be expected when obtaining the movement trajectory 20. Also, in the position measurement (position estimation) by the inertial sensor M10, the error tends to accumulate as the measurement time becomes longer (that is, the measurement accuracy of the position becomes lower). On the other hand, in an example of this embodiment, the estimation accuracy of the movement trajectory 20 can be preferably improved by correcting the movement trajectory 20 based on the language record 25.
[0034] (II) SLAM In another example, the positioning sensor M1 may include at least one of a camera M15 and a LiDAR (light detection and ranging) M16. Each of the one or more movement trajectories 20 is measured by performing SLAM (Simultaneous Localization And Mapping) analysis on the measurement data of at least one of the camera M15 and the LiDAR M16 may be. SLAM analysis is a method that simultaneously performs self-position estimation and mapping. The movement trajectory 20 may be generated as a result of self-position estimation for the measurement data. The result of mapping may or may not be used for any purpose. Known methods may be adopted for the calculation method of SLAM analysis. Similarly to the above, a trained machine learning model may be used for the calculation of SLAM analysis. If position measurement by SLAM analysis is possible, the types of the camera M15 and the LiDAR M16 do not particularly have to be limited and may be appropriately selected according to the implementation form. According to at least one of the camera M15 and the LiDAR M16, the position can be measured regardless of whether it is outdoors or indoors. Therefore, according to an example of this embodiment, the agent Z can
[0035] obtain the movement trajectory 20 (position information) in any scene where it exists indoors or outdoors. Also, similar to the case of using the inertial sensor M10, in SLAM analysis, the longer the measurement time, the more the error accumulates (that is, the lower the measurement accuracy of the position). In contrast, in an example of this embodiment, the estimation accuracy of the movement trajectory 20 can be preferably improved by correcting the movement trajectory 20 based on the language record 25.
[0036] (III) Other positioning sensors The type of the positioning sensor M1 is not limited to the example of FIG. 1 and may be appropriately changed according to the embodiment. In another example, the positioning sensor M1 may include a satellite positioning module. The satellite positioning module may include a GPS (Global Positioning System) sensor, a GNSS (Global Navigation Satellite System) sensor, etc. In this case, each of the one or more movement trajectories 20 may be generated as a measurement result of the position by the satellite positioning module.
[0037] Also, in another example, the positioning sensor M1 may include a positioning module configured to measure the position by using wireless communication with a wireless access point such as a beacon or a base station. The standard of the wireless communication may be arbitrarily selected (for example, Bluetooth (registered trademark), etc.). The positioning module by wireless communication may include, for example, a processor resource, a communication module (receiver, etc.). In one example, the positioning module by wireless communication may be configured to estimate the distance to each wireless access point according to the intensity of the signal received from each wireless access point. In addition, the positioning module by wireless communication may be configured to measure the position by any method (such as triangulation) from the estimated distances to each wireless access point. In this case, each of the one or more movement trajectories 20 may be generated as a measurement result of the position by using the positioning module by wireless communication.
[0038] (Relationship with the agent) Note that when there are a plurality of agents Z, the type of the positioning sensor M1 of each agent Z may be appropriately selected according to the embodiment. The types of the positioning sensors M1 of each agent Z may be the same or at least partially different.
[0039] Also, the method of deploying the positioning sensor M1 for the agent Z is not particularly limited and may be appropriately selected according to the embodiment. In one example, the positioning sensor M1 may be appropriately deployed on the agent Z. The positioning sensor M1 may be built into the agent Z, may be directly held by the agent Z, or may be indirectly held by being deployed on a device (such as a terminal) held by the agent Z. The terminal held by the agent Z may be constituted by any computer. In one example, the information processing device 1 may also serve as the terminal held by the agent Z. In another example, another computer other than the information processing device 1 may be used as the terminal held by the agent Z.
[0040] For example, when the agent Z is a user, the positioning sensor M1 may be directly held by the user. When the user holds a user terminal, by the user terminal including the positioning sensor M1, the positioning sensor M1 may be indirectly held by the user. The method of connecting the positioning sensor M1 to the user terminal may be arbitrarily selected. In one example, the positioning sensor M1 may be built into the user terminal. In another example, the positioning sensor M1 may be connected to the user terminal from the outside by wire or wirelessly. The user terminal is an example of the terminal held by the agent Z. The user terminal may include a smartphone, a tablet terminal, etc.
[0041] Also, for example, when the agent Z is a moving body, the positioning sensor M1 may be appropriately deployed on the moving body. When the moving body includes a terminal (such as a control device, an in-vehicle device, etc.), by the terminal including the positioning sensor M1, the positioning sensor M1 may be deployed on the moving body. Similar to the form of the above user terminal, the method of connecting the positioning sensor M1 to the terminal of the moving body may be arbitrarily selected. The positioning sensor M1 may be built into the terminal of the moving body, or may be connected to the terminal of the moving body from the outside by wire or wirelessly. The terminal of the moving body is an example of the terminal held by the agent Z.
[0042] In another example, the positioning sensor M1 may be configured to externally observe the position of the agent Z, such as a plurality of cameras or the like arranged at different locations. In this case, the positioning sensor M1 may be appropriately arranged outside the agent Z (that is, the environment in which the agent Z operates), rather than being arranged on the agent Z. The positioning sensor M1 may be connected to an arbitrary computer such as a computer (terminal) or a server device arranged in the environment, and at least a part of the movement trajectory 20 may be generated by an arbitrary computer. The information processing device 1 may also serve as an arbitrary computer connected to the positioning sensor M1.
[0043] [Movement trajectory] The movement trajectory 20 may be configured to show the progress of the movement of the agent in time series by including measurement values of the position at one or more times (time points). For example, the movement trajectory 20 may be composed of one or more combinations of the measurement time and the measurement value of the position. The position may be arbitrarily expressed. In one example, the position may be expressed in two or three dimensions. The measurement value of the position may be directly obtained as the sensing data of the positioning sensor M1, or may be obtained as a result of arithmetic processing (analysis, etc.) on the sensing data of the positioning sensor M1. The measurement value of the position may also be referred to as position information. The measurement conditions such as the measurement interval of the position, the time resolution, and the type of coordinate system may not be particularly limited and may be appropriately determined according to the embodiment.
[0044] The measurement value at each time may indicate the position relatively or absolutely. For example, in the case of using the inertial sensor M10, the case of using SLAM analysis, etc., the position indicated by the measurement value at each time may be a relative position from the start point of the measurement. Also, for example, in the case of using the satellite positioning module, the case of using wireless communication, etc., the position indicated by the measurement value at each time may be an absolute position.
[0045] The movement trajectory 20 may further include any physical information other than the position information. For example, when an inertial sensor (including the inertial sensor M10) is deployed on the agent Z, the movement trajectory 20 may further include the measured values of the attitude and speed of the agent Z (the positioning sensor M1) at one or more times (time points). The attitude may include the orientation. The measured value of the attitude may also be referred to as attitude information. The measured value of the speed may also be referred to as speed information or displacement information. Similarly, when using the SLAM analysis by at least one of the camera M15 and the LiDAR M16 for the generation of the movement trajectory 20, the movement trajectory 20 may further include the measured values of the attitude and speed of the agent Z at one or more times (time points). Similarly, when using the SLAM analysis by at least one of the camera M15 and the LiDAR M16 for the generation of the movement trajectory 20, the movement trajectory 20 may further include the measured values of the attitude and speed of the agent Z at one or more times (time points).
[0046] Also, the position at each time in the movement trajectory 20 may typically be treated as the position of the corresponding agent Z itself. However, the treatment of the position at each time in the movement trajectory 20 is not limited to such an example. The position at each time in the movement trajectory 20 may be arbitrarily treated according to the usage purpose in addition to the position of the corresponding agent Z itself. In one example, the position at each time in the movement trajectory 20 may be treated as at least one of the position of the corresponding agent Z itself and the position of an object related to the agent Z. For example, the position at at least some of the times included in the movement trajectory 20 may be treated as the position of an object (the target object that appears at the same time as the time of the position of the target) that appears at the corresponding time in the language record 25 of the corresponding agent Z.
[0047] The type of the object is not particularly limited and may be appropriately selected according to the embodiment. The object may include, for example, products and goods (foodstuffs, beverages, daily necessities, clothing, electrical appliances, etc.), tools (including household furniture such as desks, chairs, shelves, etc.), facilities (kitchen facilities, refrigeration facilities, registers, etc.), architectural structures (entrances, doors, windows, pillars, etc.), monuments, artworks (statues, paintings etc.), and parts thereof. The type of the object may be appropriately defined according to the embodiment such as the operation scenario. The object may also be referred to as a landmark.
[0048] Even under the condition that the positioning sensor M1 measures the position of the agent Z, the position at the target time included in the movement trajectory 20 may be treated as the position of the object. In this case, the position at the target time may be appropriately corrected according to being treated as the position of the object. For example, information (such as a phrase) indicating the positional relationship between the agent Z and the object, such as "beef (an example of an object) exists in front", may be detected from the language record 25. The position at the target time may be corrected according to the detected information. Also, for example, a correction rule such as correcting n[m] ahead of the measured value of the position (n is an arbitrary value) may be given in advance. The position at the target time may be corrected according to the given rule. When the movement trajectory 20 further includes physical information such as the measured value of the posture of the agent Z, the physical information may be used for the correction of the position at the target time (for example, in the above example of correction, specifying the forward direction from the measured value of the posture, etc.). When sensors such as a camera (including camera M15) and LiDAR (LiDAR M16) are used, the position at the target time may be corrected according to the observation result of the object by the sensor. ject.
[0049] Note that, in an example of the present embodiment, the correction method of the present disclosure (the correction method of the movement trajectory 20 based on the language record 25) may be used in combination with an existing correction method for the measured value of the position. For example, as described above, the correction method of the present disclosure may be used in combination with the correction method of Non-Patent Document 1 that uses a geographical database. Also, for example, when measuring the position by the above SLAM analysis, other sensors other than the camera M15 and LiDAR M16 may be further installed in the agent Z. Other sensors may include, for example, an inertial measurement unit (IMU), a satellite positioning module, an encoder that measures the rotation speed of the wheels, and the like. The position (position information) measured by the SLAM analysis may be corrected by known methods such as a method using the measured value of the inertial measurement unit, a method using the measured value of the satellite positioning module, and wheel odometry (a method using the rotation speed of the wheels).
[0050] [Language Recording] The language recording 25 may be configured to include language information (language data) at one or more times (time points). The language recording 25 may be measured in time series. For example, the language recording 25 may be composed of one or more combinations of measurement times and language information. The data format of the language information may not be particularly limited and may be appropriately selected according to the embodiment (such as text, etc.).
[0051] (Generation Method) The measurement method of the language recording 25 may be arbitrarily selected. That is, the language information included in the language recording 25 may be generated by any method. In one example, for the generation of at least a part of the language information included in the language recording 25, the input device M2 may be used. The type of the input device M2 may be arbitrarily selected. The input device M2 may include, for example, a camera, a microphone, a keyboard, a touch panel, an operator, etc.
[0052] When the input device M2 includes a camera, the language information may be generated as a result of image analysis (captioning) on the image obtained by the camera. The result of the image analysis may include the detection result of an object, the identification result of sign language, the identification result of body language, etc. The method of image analysis may not be particularly limited and may be appropriately selected according to the embodiment. Known methods may be adopted for image analysis. A trained machine learning model generated by machine learning may also be used for image analysis. The trained machine learning model may include a large-scale generation model such as a large-scale vision-language model. Note that when the positioning sensor M1 includes the camera M15, the camera M15 may also be used as the input device M2. Alternatively, a camera different from the camera M15 may be used as the input device M2.
[0053] When the input device M2 includes a microphone, the language information may be generated as a result of speech analysis of the sound data (voice data) obtained by the microphone. The result of the speech analysis may include the result of the analysis of the utterance by voice. When there is an object that emits sound, the result of the speech analysis may further include the detection result of the object. The method of speech analysis may not be particularly limited and may be appropriately selected according to the embodiment. A known method may be adopted for speech analysis. A trained machine learning model generated by machine learning may be used for speech analysis. The trained machine learning model may include a large-scale generation model such as a large-scale speech model. When the input device M2 includes a keyboard, a touch panel, an operator, etc., the language information may be generated according to an input operation such as chat input.
[0054] The input device M2 may be appropriately deployed. Similar to the positioning sensor M1, in one example, the input device M2 may be deployed in the agent Z. The input device M2 may be built into the agent Z, may be directly held by the agent Z, or may be indirectly held by being deployed in a device (such as a terminal) held by the agent Z. In another example, the input device M2 may be appropriately deployed in the environment where the agent Z operates so as to observe the language information of the agent Z from the outside. In this case, the input device M2 may be connected to any computer such as a computer (terminal) or a server device deployed in the environment, and at least a part of the language record 25 may be generated by any computer. The information processing device 1 may also serve as any computer connected to the input device M2.
[0055] In another example, the language information may be generated by computer processing without using the input device M2. The computer processing for generating the language information may not be particularly limited and may be appropriately selected according to the embodiment. Known techniques such as an automatic operator or a chatbot may be adopted for the computer processing. A trained machine learning model generated by machine learning may be used for generating the language information. The trained machine learning model may include large-scale generation models such as a large language model, a large vision-language model, or a large speech model. In one example, part of the language information included in the language record 25 may be generated via the input device M2, and the rest may be generated by computer processing.
[0056] (Type) The language information included in the language record 25 may include any type of language expression observed in at least either the agent Z or its surroundings.
[0057] For example, the language information may include the utterances of the agent Z itself. When the agent Z is interacting with an operator, the language information may further include the utterances of the operator. The operator may be a human or a computer (automatic operator). The utterances of the agent Z and the operator are an example of the language expressions observed by the agent Z. The utterances may include sign languages such as sign language and body languages such as gestures. The language information from the utterances of the agent Z and the operator may be obtained by the above image analysis, speech analysis, input operation, computer processing, etc. Also, when the agent Z is a device such as a mobile body, the language information may include any information (utterances, etc.) input by the operation of the person using the agent Z. For example, when the agent Z is a mobile robot, the language information may include the utterances input by the operation of the terminal of the mobile robot. The utterances input to the terminal may include, for example, requests to the mobile robot, task instructions, and other operations.
[0058] Also, for example, the language information may include utterances observed around Agent Z. The utterances observed around Agent Z may be obtained by the above-described image analysis, voice analysis, etc. In one example, the utterances observed around Agent Z may include the utterances received by Agent Z. For example, when Agent Z is a mobile robot and a microphone is installed, the mobile robot may receive any utterance from a person using the mobile robot and it is okay. The utterances received by the mobile robot may include instructions to the mobile robot, task instructions, etc. The utterances observed around Agent Z may include the utterances received by this mobile robot. The utterances observed around Agent Z may also include utterances other than the utterances directed to Agent Z. For example, when a microphone is installed, the utterances observed around Agent Z may include the linguistic expressions of any voice (e.g., announcements, etc.) observed by the microphone in the environment where Agent Z operates.
[0059] Also, for example, the language information may include the linguistic expressions of the characteristics of the objects observed around Agent Z. The linguistic expressions of the characteristics may include, for example, the names of the objects, type names, linguistic expressions of the external characteristics, etc. The external characteristics may include, for example, color, shape, size, quantity, etc. The linguistic expressions of the characteristics may be obtained by the above-described image analysis, etc.
[0060] (Synchronization) The language record 25 is synchronized with the corresponding movement trajectory 20 (i.e., the movement trajectory 20 of the same Agent Z). The synchronization method may not be particularly limited and may be appropriately selected according to the embodiment. In one example, when the movement trajectory 20 and the language record 25 are measured by the same computer, the movement trajectory 20 and the language record 25 may be synchronized in real time. In another example, after the movement trajectory 20 and the language record 25 are generated, the time ranges of the movement trajectory 20 and the language record 25 may be matched. Thereby, the movement trajectory 20 and the language record 25 may be synchronized retrospectively.
[0061] The synchronization of the movement trajectory 20 and the language record 25 may be processed by the information processing apparatus 1 or may be processed by another computer other than the information processing apparatus 1. The information processing apparatus 1 may directly obtain the movement trajectory 20 from the positioning sensor M1 or may indirectly obtain it via an external computer. Obtaining via an external computer may include obtaining via a storage medium (such as the storage medium 91 in FIG. 7). Similarly, the information processing apparatus 1 may directly obtain the language record 25 from the input device M2 or the like, or may indirectly obtain it via an external computer. The acquisition path for each of the movement trajectory 20 and the language record 25 may be appropriately selected according to the embodiment.
[0062] One or more computers may be involved in the generation of the movement trajectory 20. The one or more computers involved in the generation of the movement trajectory 20 may or may not include the information processing apparatus 1. That is, at least a part of the movement trajectory 20 may be generated on the information processing apparatus 1 or may be generated on another computer other than the information processing apparatus 1. Similarly, one or more computers may be involved in the generation of the language information included in the language record 25. The one or more computers involved in the generation of the language information may or may not include the information processing apparatus 1. At least a part of the language record 25 may be generated on the information processing apparatus 1 or may be generated on another computer other than the information processing apparatus 1. The information processing apparatus 1 may obtain the movement trajectory 20 and the language record 25 that have been synchronized in advance by another computer. After obtaining the movement trajectory 20 and the language record 25, the information processing apparatus 1 may synchronize the movement trajectory 20 and the language record 25. The information processing apparatus 1 may be connected to the positioning sensor M1, the input device M2, etc., and obtain the movement trajectory 20 and the language record 25 while synchronizing them.
[0063] [Phrase or set of phrases] By identifying at least one of phrase 30 and phrase set 35 from one or more language records 25, a combination of times 40 related to the same location can be extracted. By extracting one or more combinations 40, at least one of one or more movement trajectories 20 can be corrected. Therefore, the information processing apparatus 1 may identify at least one of phrase 30 and phrase set 35 from one or more language records 25. That is, identifying phrase 30 or phrase set 35 may include identifying only phrase 30, identifying only phrase set 35, and identifying both phrase 30 and phrase set 35. The number of phrases 30 or phrase sets 35 to be identified depends on the language record 25. The greater the number of phrases 30 or phrase sets 35 identified from one or more language records 25, the more hints of having been located at the same location are obtained, so an improvement in the accuracy of one or more movement trajectories 20 by correction can be expected. The number of phrases 30 or phrase sets 35 to be identified depends on the language record 25. The greater the number of phrases 30 or phrase sets 35 identified from one or more language records 25, the more hints of having been located at the same location are obtained, so an improvement in the accuracy of one or more movement trajectories 20 by correction can be expected.
[0064] Phrase 30 and phrase set 35 are identified (detected) from the language record 25 as hints of having been located at the same location. Being located at the same location may include, in addition to being located at the same point (being completely located at the same location), being located within a nearby range, such as the relationship between a kitchen workbench and a stove. That is, the same location may be defined as a space having a certain range. The index for evaluating the nearby range may be appropriately defined according to the embodiment. The nearby range may be defined invariantly or variably according to the characteristics of the environment (the place to be evaluated, the object, etc.). Note that identifying phrase 30 or phrase set 35 is synonymous with detecting phrase 30 or phrase set 35. In the following description, the expression "detect" is also adopted.
[0065] (Unit of phrase) The unit of the phrases to be detected (the phrases included in phrase 30 and set 35) may not be particularly limited and may be appropriately defined according to the embodiment. In one example, a phrase may be composed of one or more words. A word may be a morpheme. One phrase may correspond to, for example, a continuous utterance such as "I came back to the place five minutes ago", or may not correspond. The time interval of one phrase may be arbitrarily defined. One phrase may constitute one or more sentences, or may not constitute a sentence. One phrase may be composed within a range shorter than one sentence. A phrase may not be a grammatically correct sentence, such as a list of words, for example.
[0066] In one example, the language information included in the language record 25 may be pre-segmented into phrase units. In this case, detecting phrase 30 or set 35 of phrases may be constituted by selecting phrase 30 or set 35 of phrases from one or more phrases included in the language record 25. In another example, the language information may not be segmented into phrase units. In this case, one or more words detected by the detection process of phrase 30 or set 35 of phrases may be treated as a phrase.
[0067] (Detection method) If it is possible to detect one or more words referring to the same location, the method of detecting phrase 30 and set 35 of phrases may not be particularly limited and may be appropriately selected according to the embodiment. Known methods such as co-reference resolution and entity linking may be adopted for the method of detecting phrase 30 and set 35 of phrases. Any detection model may be used for the detection of phrase 30 and set 35 of phrases. The detection model may be constituted by at least one of a rule-based model and a trained machine learning model.
[0068] The rule-based model is configured to derive the result of inference (detection of phrase 30 and phrase set 35) from a given input according to rules. The rules may be set as appropriate. The machine learning model may be configured in the same manner as each of the above examples. The trained machine learning model may include a large-scale generation model such as a large language model.
[0069] When using a large-scale generation model as a detection model, an instruction to detect phrase 30 or phrase set 35 may be given to the large-scale generation model together with the language record 25. In one example, the detection instruction may be configured to directly instruct the detection of phrase 30 or phrase set 35, such as "Please detect words and word sets that refer to the same location". In another example, the detection instruction may be configured to instruct the performance of any task related to the detection of the same location (i.e., indirectly instruct the detection of phrase 30 or phrase set 35). For example, assuming a scenario where agent Z is operated within a store and the products displayed in the store (an example of an object) are utilized as hints for the same location, the detection instruction may be configured to instruct solving the task of identifying the shelves of the product sales area. In this case, based on references to products displayed on the same shelf, phrase 30 or phrase set 35 can be detected. The same task can be performed in a warehouse to detect phrase 30 or phrase set 35 indicating the same location.
[0070] Note that the detection model may be deployed in the information processing device 1 or in another computer other than the information processing device 1. When the detection model is deployed in the information processing device 1, the information processing device 1 may execute the process of detecting phrase 30 or phrase set 35 on its own device. When the detection model is deployed in another computer, the information processing device 1 may instruct the other computer to execute the detection process, and as a response, obtain the result of detecting phrase 30 or phrase set 35 from the other computer. Detecting phrase 30 or phrase set 35 may include causing such a detection process to be executed on another computer and obtaining the detection result.
[0071] In addition, any information other than the language record 25 may be further used for the detection of the phrase 30 or the phrase set 35. In one example, prerequisite information indicating the prerequisite conditions of the environment in which Agent Z operates may be used for the detection of the phrase 30 or the phrase set 35. The prerequisite conditions may be appropriately defined according to the implementation forms such as the operation scenario. In one example, the prerequisite conditions may include a list of objects existing in the target environment. For example, when assuming a scenario of operating the above Agent Z in a store, the prerequisite conditions may include information for identifying products displayed on the same shelf such as a product list of the store and a shelf list. When assuming a scenario of operating Agent Z in a shopping mall having a plurality of stores, the prerequisite conditions may further include a store list in the shopping mall. The product list and the shelf list may be provided for each store.
[0072] (Detection range) For at least any one of one or more Agents Z, one or more language records 25 may be appropriately referred to for detecting the phrase 30 or the phrase set 35 indicating the same location.
[0073] For example, the target Agent Z may refer to its own location, such as "I came back to the place five minutes ago", "I came to make coffee", "I came to pick up the made coffee". Therefore, in one example, within the language record 25 of the target Agent Z (one Agent), a phrase 30 or a phrase set 35 indicating the same location may be detected for the target Agent Z. Accordingly, by using the detected phrase 30 or phrase set 35 as a hint that it is located at the same location, the movement trajectory 20 of the target Agent Z may be corrected.
[0074] Also, for example, an agent Z may refer to the location of another agent Z, such as "I have now joined BBB (the second agent)". In addition, for example, multiple agents Z may refer to the same location, such as the first agent "took out pork from the refrigerator" and the second agent "took out beef from the refrigerator". Therefore, in one example, when multiple agents Z are given, among the language records 25 of the multiple agents Z, for at least any one of the multiple agents Z, a phrase 30 or a set 35 of phrases indicating the same location may be detected. By using the detected phrase 30 or set 35 of phrases as a hint that they are located at the same location, the movement trajectory 20 of at least any one of the multiple agents Z may be corrected.
[0075] (1) Position correction within one agent FIG. 2 schematically shows an example of a scene for correcting the movement trajectory 20 (position information) within one agent. In one example, one or more agents Z may include the first agent Z1. One or more movement trajectories 20 may include the first movement trajectory 201 of the first agent Z1. One or more language records 25 may include the first language record 251 of the first agent Z1. Extracting one or more combinations 40 of times related to the same location may include identifying a phrase 30 or a set 35 of phrases included in the first language record 251 to extract one or more combinations 40 of times related to the same location. Correcting one or more movement trajectories 20 may include correcting the first movement trajectory 201 so that the positions at the times of the extracted one or more combinations 40 approach each other. According to an example of this embodiment, within one agent (the first agent Z1), by correcting the position information (the first movement trajectory 201), the accuracy of the position information (the first movement trajectory 201) can be improved.
[0076] If the first agent Z1 is located at the same location at multiple times, each of the phrases 30 and phrase sets 35 included in the first language record 251 may not be particularly limited and may be appropriately determined according to the embodiment. The information processing apparatus 1 may identify (detect) at least any one of the phrases 30 and phrase sets 35 regarding the fact that the first agent Z1 is located at the same location at multiple times from the first language record 251.
[0077] (An example of a phrase) In one example, identifying the phrase 30 or the phrase set 35 included in the first language record 251 may include identifying a first phrase 301 indicating being located at the same location at multiple times. The first phrase 301 is an example of the phrase 30. The first phrase 301 may indicate being located at the same location with a single phrase. If it mentions being located at the same location at multiple times, the content of the first phrase 301 may not be particularly limited and may be appropriately determined according to the embodiment. In one example, the first phrase 301 may include phrases that mention being located at the same location in the past, such as "I came back to the place five minutes ago", "I am at the same place as 15 minutes ago", "I arrived at the place where I was an hour ago again", "I returned to the same place as at first", etc.
[0078] When the first phrase 301 is identified from the first language record 251, the one or more combinations 40 extracted may include a first combination 401 composed of a plurality of times corresponding to the identified first phrase 301. The plurality of times constituting the first combination 401 may be appropriately determined according to the identified first phrase 301 and the measurement time of the first phrase 301. For example, when a phrase such as "I came back to the place five minutes ago" is detected, the measurement time (utterance time) of this phrase and the time five minutes before the measurement time may be an example of the plurality of times constituting the first combination 401. According to an example of this embodiment, by using the first phrase 301 indicating being located at the same location, the first movement trajectory 201 of the first agent Z1 can be appropriately corrected.
[0079] Typically, the first combination 401 may be composed of two times (the first time and the second time). In the above example, the measurement time of the phrase "returned to the place five minutes ago" is one of the first time and the second time, and the time five minutes before the measurement time is the other. However, the number of times located at the same place is not limited to two, and may be three or more. For example, the first phrase 301 may refer to being located at the same place at three or more times, such as "passed the same place five minutes ago and ten minutes ago" (in this example, three times: the measurement time, the time five minutes before, and the time ten minutes before). In this case, the first combination 401 may be composed of three or more times (the first combination 401 may further include times after the third time). That is, the number of times constituting the first combination 401 is not limited to two, and may be three or more. Alternatively, the first combination 401 is composed of two times, and three or more times located at the same place may be composed of two or more first combinations 401. The same may apply to the combination 40. That is, when located at the same place at three or more times, the corresponding combination 40 may be composed of three or more times. Alternatively, the combination 40 is composed of two times, and three or more times located at the same place may be composed of two or more combinations 40.
[0080] Also, in a typical example, the time (the first time, the second time, etc.) may be defined to refer to one point in time. However, there may be cases where the time corresponding to the first phrase 301 is not determined to be one point in time. For example, when the phrase "returned to the place five minutes ago" is detected as the first phrase 301, since this "five minutes ago" is a rough time, it is not necessarily located at the exact "five minutes ago" at the same place. Therefore, the time may be defined as a time having a certain width. In the above example, the time of "five minutes ago" may be defined within the range of the point in time of "five minutes ago" and its vicinity, such as a one-minute period belonging to "five minutes ago". The measurement time (speech time) of the phrase may be defined within the time from the start point to the end point of the phrase and its vicinity. The width of the time may be appropriately specified according to the embodiment.
[0081] Furthermore, the time corresponding to the first phrase 301 may be adjusted according to the movement status of the first agent Z1 appearing in the first movement trajectory 201. In one example, at least one of the plurality of times corresponding to the first phrase 301 may be the time specified from the first movement trajectory 201 when the first agent Z1 is stopped in the vicinity of the time specified from the first phrase 301. For example, when a phrase such as "returned to the place five minutes ago" is detected as the first phrase 301, the time specified from the first movement trajectory 201 when the first agent Z1 is stopped in the vicinity of the time five minutes before the measurement time may be adopted as the time five minutes ago. In the vicinity of the measurement time of the above phrase, the time specified from the first movement trajectory 201 when the first agent Z1 is stopped may be adopted as the measurement time.
[0082] The same may apply to the times constituting the combination 40. That is, the times constituting the combination 40 may be defined to indicate a single point in time, or may be defined as a time having a certain width. Further, the times constituting the combination 40 may be adjusted according to the movement status of the agent Z appearing in the corresponding movement trajectory 20. In one example, the time corresponding to a phrase (a phrase included in the phrase 30 and the set 35) may be the time specified from the corresponding movement trajectory 20 when the corresponding agent Z is stopped in the vicinity of the time specified from the phrase.
[0083] (An example of a set of phrases) In another example, identifying phrases 30 and phrase sets 35 included in the first language record 251 may include identifying a second phrase 352 and a third phrase 353 that indicate the same location as each other. The set of the second phrase 352 and the third phrase 353 is an example of the phrase set 35. When the second phrase 352 and the third phrase 353 are identified from the first language record 251, the one or more combinations 40 extracted may include a second combination 402 composed of the time corresponding to the identified second phrase 352 and the time corresponding to the third phrase 353. According to an example of this embodiment, by using the second phrase 352 and the third phrase 353 (phrase set 35) that indicate the same location as each other, the first movement trajectory 201 of the first agent Z1 can be appropriately corrected.
[0084] Note that if it is mentioned that they are located at the same location as each other, the content of each phrase (352, 353) may not be particularly limited and may be appropriately determined according to the embodiment.
[0085] In one example, "came to make coffee" · "came to get the made coffee", "opened the refrigerator" · "is in front of the refrigerator", language of the characteristics of the same object The second phrase 352 may indicate that it is located at the same location as the third phrase 353 by using the same word as the third phrase 353, as in the case where expressions of the characteristics of the same object are included at a plurality of times. In this case, in addition to the above detection method, the second phrase 352 and the third phrase 353 may be detected by searching for the same word that appears at a plurality of times in the first language record 251.
[0086] In another example, the second phrase 352 may indicate being located at the same location as the third phrase 353 using different words, such as "opened the refrigerator" and "lit the stove" (both referring to being in the kitchen), "going out for a while" and "returned to the previous room" (both referring to being in the target room), and when the linguistic expressions of the features of different objects arranged at the same location are included at multiple times. In one example, the second phrase 352 and the third phrase 353 may be detected by the above detection method. Also, in one example, whether each object is arranged at the same location may be specified from the prerequisite information. When using a large-scale generation model for detecting the set of phrases 35, the information processing device 1 may provide the prerequisite information to the large-scale generation model together with the first language record 251, so that the large-scale generation model detects from the first language record 251 the set of phrases 35 (the second phrase 352, the third phrase 353) referring to different objects arranged at the same location. Alternatively, it is known that the large-scale generation model can acquire common sense by learning a large amount of data. Therefore, the information processing device 1 may cause the large-scale generation model to detect from the first language record 251 the set of phrases 35 referring to different objects arranged at the same location without providing the prerequisite information. Whether each object is arranged at the same location may be specified based on rules.
[0087] As in the above example, it is possible to refer to the same location using different words. According to an example of the present embodiment, by using this, the first movement trajectory 201 of the first agent Z1 can be appropriately corrected. Also, by not limiting the search range for the set of the second phrase 352 and the third phrase 353 to the same word, a set of the second phrase 352 and the third phrase 353 that hints at being located at the same location can be widely acquired from the first language record 251. Thereby, an improvement in the accuracy of the first movement trajectory 201 can be expected.
[0088] In addition, when the first language record 251 includes the language record of the dialogue between the first agent Z1 and the operator or the like, the dialogue may not be completed in two phrases, namely the second phrase 352 and the third phrase 353. For example, in a scenario where the operator instructs the first agent Z1 to transport goods, it is assumed that the first language record 251 includes the language information of the first agent Z1 "I have taken out the goods from the shelf in the backyard", the operator "Please transport the goods to the designated location", and the first agent Z1 "I have transported the goods to the designated location and returned to the backyard". Also, for example, in a scenario where the operator instructs the first agent Z1 to move, it is assumed that the first language record 251 includes the language information of the operator "Please move to location XXX", the first agent Z1 "I have moved to location XXX", the operator "Please return to the original location", and the first agent Z1 "I have returned to the original location". In each of the above examples, "I have taken out the goods from the shelf in the backyard" / "Please move to location XXX" is one of the second phrase 352 and the third phrase 353, and "I have transported the goods to the designated location and returned to the backyard" / "I have returned to the original location" is the other of the second phrase 352 and the third phrase 353. As in these examples, the dialogue may be composed of three or more phrases. The second phrase 352 and the third phrase 353 may be detected from the dialogue composed of three or more phrases.
[0089] Also, the time corresponding to the second phrase 352 may be appropriately determined according to at least one of the measurement times of the second phrase 352 and the second phrase 352. In one example, the time corresponding to the second phrase 352 may be the measurement time (such as the speaking time) of the second phrase 352. Similar to the time of the first phrase 301 above, the time corresponding to the second phrase 352 is 1 It may be defined to indicate a time point, or it may be defined as a time having a certain width. The time corresponding to the second phase 352 may be adjusted according to the movement situation of the first agent Z1 appearing in the first movement locus 201. In one example, the time corresponding to the second phase 352 may be the time specified from the first movement locus 201 when the first agent Z1 is stopped in the vicinity of the time specified from the second phase 352 (such as the measurement time). Except for replacing the second phase 352 with the third phase 353, the time corresponding to the third phase 353 may be the same as the time corresponding to the second phase 352.
[0090] Also, typically, a set of the second phase 352 and the third phase 353 may constitute a phrase set 35. However, the number of phrases referring to the same location is not limited to two and may be three or more. For example, like "opened the refrigerator", "standing in front of the refrigerator", "took out beef from the refrigerator" (in this example, three phrases), in addition to the first phrase (one of the second phase 352 and the third phase 353) and the second phrase (the other of the second phase 352 and the third phase 353), there may be phrases after the third. In this case, the phrase set 35 may be composed of three or more phrases including the second phase 352 and the third phase 353. Alternatively, the phrase set 35 may be composed of two phrases (the second phase 352, the third phase 353), and three or more phrases indicating the same location may be composed of two or more phrase sets 35. Accordingly, similar to the first combination 401 above, the number of times located at the same location is not limited to two and may be three or more. The second combination 402 may be composed of three or more times (times corresponding to each phrase). Alternatively, the second combination 402 may be composed of two times, and three or more times located at the same location may be composed of two or more second combinations 402.
[0091] In one example, the second phrase 352 may indicate that it is located at the same position as the third phrase 353 using at least one of the same words and different words as the third phrase 353. Also, in one example, the information processing apparatus 1 may detect both the phrase 30 (the first phrase 301) and the phrase set 35 (the second phrase 352 and the third phrase 353) from the first language record 251.
[0092] (Correction method) It does not have to be questioned which of the positions at each time of the extracted combination 40 is the true value. As long as the positions at each time are corrected so as to approach each other, the method of correcting the movement trajectory 20 does not have to be particularly limited and may be appropriately selected according to the embodiment. A known method may be adopted as the method of correcting the movement trajectory 20. Approaching the positions at each time may include making the positions at each time coincide. When the target time is defined as a time having a certain width, the position of the target time may be arbitrarily determined from the position information included in the corresponding time. For example, the position of the target time may be constituted by the position of a representative time point, or may be constituted by a statistic (average value, median, etc.) of the positions included in the corresponding time. The method of determining the representative time point does not have to be particularly limited and may be appropriately selected according to the embodiment. The above-mentioned stopped time is an example of a representative time point.
[0093] The first movement trajectory 201 may be corrected so as to approach the positions at each time included in each of one or more combinations 40. In one example, when the first phrase 301 is detected, the first movement trajectory 20 may be corrected so as to approach the positions at each time included in the first combination 401. When the second phrase 352 and the third phrase 353 are detected, the first movement trajectory 201 may be corrected so as to approach the positions at each time included in the second combination 402. In one example, when the phrase (the first phrase 301, the second phrase 352, the third phrase 353) includes the word of the object, the position to be corrected is the position of the corresponding agent Z (the first agent Z1) and the object related to the agent Z It may be at least any one of the positions.
[0094] In addition, when speed information (displacement information) is obtained, such as in the case where the inertial sensor is deployed or the case where SLAM analysis is used, for the method of correcting the movement trajectory 20 (including the first movement trajectory 201), it is preferable to adopt an optimization method that corrects to approximate the positions at each time of the same location while suppressing the change of the speed information (that is, saving the speed information as much as possible). For such a correction method, the type of the optimization method may not be particularly limited and may be appropriately selected according to the embodiment. Known methods such as loop closure and time-series smoothing (Kalman smoother, etc.) may be adopted as the optimization method.
[0095] FIG. 3 schematically shows an example of correcting the movement trajectory 20 (the first movement trajectory 201) within one agent (the first agent Z1). In an example of FIG. 3, the first movement trajectory 201 includes position information from time t1 to time t6, and a scene is assumed in which time t2 and time t5 are extracted as the times of the first combination 401 or the second combination 402. Time t2 and time t5 are examples of the first time and the second time. When the combination 40 (including the first combination 401 and the second combination 402) is extracted within the movement trajectory 20 of one agent, each time included in the combination 40 is different. In an example of FIG. 3, time t2 and time t5 are different from each other. The information processing device 1 may appropriately correct the first movement trajectory 201 so as to approximate the position at time t2 and the position at time t5.
[0096] In an example of FIG. 3, both the position at time t2 and the position at time t5 are corrected. Thus, in one example, the information processing apparatus 1 may correct the positions at all times included in the combination 40. However, the method of correcting the movement locus 20 is not limited to such an example. In another example, the positions at some of the multiple times included in the combination 40 may be fixed without being corrected. For example, in an example of FIG. 3, either the position at time t2 or the position at time t5 may be fixed without being corrected. That is, the information processing apparatus 1 may treat the position at any time included in the combination 40 as the true value.
[0097] Also, in one example, a priority may be set for the position at each time included in the combination 40. The information processing apparatus 1 may adjust the correction amount of the position at each time according to the set priority. For example, the information processing apparatus 1 may decrease the correction amount of the position at the target time as the set priority is higher, and increase the correction amount of the position at the target time as the priority is lower. The priority of the position at each time may be set by any method. In one example, the priority may be set according to the measurement accuracy of the position information. For example, in the case of using the inertial sensor M10, the case of using at least one of the camera M15 and the LiDAR M16 (using SLAM analysis), etc., the error may accumulate as the measurement time becomes longer. Therefore, in an example of a case where the measurement accuracy may deteriorate according to such a measurement time, the priority of the position at each time may be set such that the earlier the measurement time, the higher the priority, and the later the measurement time, the lower the priority. Also, for example, in the case of using the satellite positioning module, etc., the value of the measurement accuracy may be directly obtained together with the measurement of the position by the positioning sensor M1. In the case of using the satellite positioning module, the DOP (Dilution of Precision) value indicating the degree of deterioration of the positioning accuracy is available. This is an example of a directly obtained value of measurement accuracy. In an example of a case where such a value of measurement accuracy is directly obtained, the priority of the position at each time may be set so that the higher the measurement accuracy indicated by the obtained value, the higher the priority, and the lower the measurement accuracy, the lower the priority. According to this example of the present embodiment, it is expected that the accuracy of the obtained movement trajectory 20 (including the first movement trajectory 201) will be further improved.
[0098] (2) Position correction among multiple agents FIG. 4 is a schematic diagram showing an example of a scene in which a movement trajectory 20 (position information) is corrected between two agents. In this example, one or more agents Z include a first agent Z1 and a second agent Z2. The second agent Z2 may be an individual different from the first agent Z1. The type of the second agent Z2 may be the same as or different from that of the first agent Z1. For example, the first agent Z1 may be a user, whereas the second agent Z2 may be a mobile robot. The one or more movement trajectories 20 may include a first movement trajectory 201 of the first agent Z1 and a second movement trajectory 202 of the second agent Z2. The one or more language records 25 may include a first language record 251 of the first agent Z1 and a second language record 252 of the second agent Z2. Extracting one or more combinations 40 of times related to the same location may include extracting one or more combinations 40 of times related to the same location by identifying a phrase 30 or a set of phrases 35 included in at least one of the first language record 251 and the second language record 252. Correcting the one or more movement trajectories 20 may include correcting at least one of the first movement trajectory 201 and the second movement trajectory 202 so that the positions at the time of each of the one or more extracted combinations 40 become closer to each other. According to one example of the present embodiment, it is possible to improve the accuracy of position information (at least one of the first movement trajectory 201 and the second movement trajectory 202) between multiple agents (the first agent Z1, the second agent Z2).
[0099] If at least one of the first agent Z1 and the second agent Z2 is located at the same location at a plurality of times, each of the phrase 30 and the phrase set 35 included in at least one of the first language record 251 and the second language record 252 may not be particularly limited and may be appropriately determined according to the embodiment. The plurality of times located at the same location may include that the first time of the first agent Z1 and the second time of the second agent Z2 are the same. The information processing apparatus 1 may specify (detect) at least one of the phrase 30 and the phrase set 35 regarding that at least one of the first agent Z1 and the second agent Z2 is located at the same location at a plurality of times from at least one of the first language record 251 and the second language record 252.
[0100] (An example of a phrase) In one example, specifying the phrase 30 and the phrase set 35 included in at least one of the first language record 251 and the second language record 252 may include specifying the fourth phrase 304 indicating that the first agent Z1 is at the same location as the second agent Z2 from the first language record 251, or specifying the fifth phrase 305 indicating that the second agent Z2 is at the same location as the first agent Z1 from the second language record 252. Each phrase (304, 305) is an example of the phrase 30. In one example, the information processing apparatus 1 may specify both the fourth phrase 304 and the fifth phrase 305. Each phrase (304, 305) may indicate being located at the same location with a single phrase.
[0101] If the first agent Z1 mentions being in the same location as the second agent Z2, the content of the fourth phrase 304 may not be particularly limited and may be appropriately determined according to the embodiment. In one example, the fourth phrase 304 may include language expressions corresponding to the detection result of the second agent Z2, such as "I have now merged with BBB (the second agent Z2)", "BBB was using this room until 5 minutes ago", "I passed by BBB 10 minutes ago", and phrases that mention being in the same location as the second agent Z2 at the measurement time or in the past from the measurement time.
[0102] Similarly, if the second agent Z2 mentions being in the same location as the first agent Z1, the content of the fifth phrase 305 may not be particularly limited and may be appropriately determined according to the embodiment. In one example, the fifth phrase 305 may include language expressions corresponding to the detection result of the first agent Z1, such as "I am with AAA (the first agent Z1)", "I saw AAA 5 minutes ago", and phrases that mention being in the same location as the first agent Z1 at the measurement time or in the past from the measurement time.
[0103] When the fourth phrase 304 is specified from the first language record 251 or the fifth phrase 305 is specified from the second language record 252, the one or more combinations 40 extracted may include a third combination 403 composed of the time of the first agent Z1 and the time of the second agent Z2 corresponding to the specified fourth phrase 304 or fifth phrase 305. When the fourth phrase 304 is specified, the one or more combinations 40 may include a third combination 403 composed of the time of the first agent Z1 and the time of the second agent Z2 corresponding to the fourth phrase 304. When the fifth phrase 305 is specified, the one or more combinations 40 may include a third combination 403 composed of the time of the first agent Z1 and the time of the second agent Z2 corresponding to the fifth phrase 305. Hereinafter, for convenience of explanation, the time of the first agent Z1 is also referred to as the first time, and the time of the second agent Z2 is also referred to as the second time.
[0104] The times of each agent (Z1, Z2) constituting the third combination 403 extracted from the fourth phrase 304 may be appropriately determined according to the specified fourth phrase 304 and the measurement time of the fourth phrase 304. For example, when a phrase such as "I have just merged with BBB" is detected as the fourth phrase 304, the measurement time (utterance time) of this phrase may be an example of the times of each agent (Z1, Z2) constituting the third combination 403. Also, for example, when a phrase such as "BBB was using this room until 5 minutes ago" is detected as the fourth phrase 304, the measurement time of this phrase may be an example of the time of the first agent Z1 included in the third combination 403, and the time 5 minutes before the measurement time may be an example of the time of the second agent Z2 included in the third combination 403. Similarly, the times of each agent (Z1, Z2) constituting the third combination 403 extracted from the fifth phrase 305 may be appropriately determined according to the specified fifth phrase 305 and the measurement time of the fifth phrase 305. For example, when a phrase such as "I am with AAA" is detected as the fifth phrase 305, the measurement time (utterance time) of this phrase may be an example of the times of each agent (Z1, Z2) constituting the third combination 403. Also, for example, when a phrase such as "I saw AAA 5 minutes ago" is detected as the fifth phrase 305, the time 5 minutes before the measurement time of this phrase may be an example of the times of each agent (Z1, Z2) constituting the third combination 403. According to an example of the present embodiment, by using at least one of the fourth phrase 304 and the fifth phrase 305, at least one of the first movement trajectory 201 of the first agent Z1 and the second movement trajectory 202 of the second agent Z2 can be appropriately corrected.
[0105] Note that, similar to the first phase 301 of one agent, the number of times located at the same location does not have to be limited to two, and may be three or more. For example, "I passed by BBB 5 minutes ago and 10 minutes ago" (in this example, the time 5 minutes before the measurement time as the first time, the time 5 minutes before the measurement time as the second time, the time 10 minutes before the measurement time as the first time, and the time 10 minutes before the measurement time as the second time, a total of four times), etc. Each phrase (304, 305) may refer to being located at the same location at three or more times. In this case, the number of times constituting the third combination 403 does not have to be limited to two, and may be three or more. Alternatively, the third combination 403 is composed of two times, and three or more times located at the same location may be composed of two or more third combinations 403.
[0106] Also, each time constituting the third combination 403 may be defined to refer to one point in time, or may be defined as a time having a certain width. Furthermore, the times of each agent (Z1, Z2) constituting the third combination 403 may be adjusted according to the movement status of each agent (Z1, Z2) appearing in each movement trajectory (201, 202). The first time included in the third combination 403 may be adjusted according to the movement status of the first agent Z1 appearing in the first movement trajectory 201. The second time included in the third combination 403 may be adjusted according to the movement status of the second agent Z2 appearing in the second movement trajectory 202. In one example, the first time of the first agent Z1 corresponding to each phrase (304, 305) may be the time specified from the first movement trajectory 201 when the first agent Z1 is stopped near the time specified from each phrase (304, 305). The second time of the second agent Z2 corresponding to each phrase (304, 305) may be the time specified from the second movement trajectory 202 when the second agent Z2 is stopped near the time specified from each phrase (304, 305).
[0107] Note that the phrase 30 specified from at least one of the first language record 251 and the second language record 252 does not have to be limited to the fourth phrase 304 and the fifth phrase 305.
[0108] In another example, specifying the phrase 30 and the phrase set 35 included in at least one of the first language record 251 and the second language record 252 may include specifying, from the second language record 252, a phrase (the eighth phrase) indicating that the first agent Z1 was located at the same location at multiple times. The eighth phrase may include a phrase referring to the fact that the first agent Z1 was located at the same location in the past, such as, for example, "AAA (the first agent Z1) returned to the place where it was 10 minutes ago". The one or more combinations 40 extracted may include a combination (the fifth combination) constituted by multiple times of the first agent Z1 corresponding to the specified eighth phrase. The multiple times constituting the fifth combination may be appropriately determined according to the specified eighth phrase and the measurement time of the eighth phrase. Correcting at least one of the first movement trajectory 201 and the second movement trajectory 202 may include correcting the first movement trajectory 201 so that the positions at each time included in the fifth combination approach each other. Note that the eighth phrase is the same as the first phrase 301 except that the information source (language record 25) is different. Therefore, including the above points, the eighth phrase may be treated in the same manner as the first phrase 301.
[0109] In another example, identifying phrases 30 and phrase sets 35 included in at least one of the first language record 251 and the second language record 252 may include identifying, from the first language record 251, a phrase (ninth phrase) indicating that the second agent Z2 was located at the same location at multiple times. The ninth phrase may include a phrase referring to the fact that the second agent Z2 was located at the same location in the past, such as, for example, "BBB (second agent Z2) is at the same location as five minutes ago". One or more combinations 40 to be extracted may include a combination (sixth combination) constituted by multiple times of the second agent Z2 corresponding to the identified ninth phrase. The multiple times constituting the sixth combination may be appropriately determined according to the identified ninth phrase and the measurement time of the ninth phrase. Correcting at least one of the first movement trajectory 201 and the second movement trajectory 202 may include correcting the second movement trajectory 202 so that the positions at each time included in the sixth combination approach each other. Note that the ninth phrase is the same as the first phrase 301 except that the first agent Z1 is the second agent Z2 and the information source (language record 25) is different. Therefore, by replacing the first agent Z1 with the second agent Z2, the ninth phrase may be treated in the same manner as the first phrase 301.
[0110] (An example of a phrase set) In another example, identifying phrases 30 and phrase sets 35 included in at least one of the first language record 251 and the second language record 252 may include identifying the sixth phrase 356 from the first language record 251 and identifying, from the second language record 252, a seventh phrase 357 indicating the same location as the sixth phrase 356. The set of the sixth phrase 356 and the seventh phrase 357 is an example of the phrase set 35. When the sixth phrase 356 and the seventh phrase 357 are identified, one or more combinations 40 to be extracted are the fourth combination constituted by the time corresponding to the sixth phrase 356 and the time corresponding to the seventh phrase 357. It may include the combination 404. The time corresponding to each phrase (356, 357) may be the time of at least one of the first agent Z1 and the second agent Z2. According to an example of this embodiment, by using the sixth phrase 356 and the seventh phrase 357 (phrase set 35), at least one of the first movement trajectory 201 of the first agent Z1 and the second movement trajectory 202 of the second agent Z2 can be appropriately corrected.
[0111] In addition, if it is mentioned that they are located at the same place, the content of each phrase (356, 357) may not be particularly limited and may be appropriately determined according to the embodiment.
[0112] In one example, the sixth phrase 356 may indicate that it is located at the same place as the seventh phrase 357 using the same word as the seventh phrase 357, such as the first agent Z1 "took out pork from the refrigerator" · the second agent Z2 "took out beef from the refrigerator", the first agent Z1 "arrived at the refrigerator" · the second agent Z2 "AAA (the first agent Z1) is still taking out luggage from the refrigerator", and the first language record 251 and the second language record 252 each contain a linguistic expression of the characteristics of the same object. In this case, in addition to the above detection method, the sixth phrase 356 and the seventh phrase 357 may be detected by searching for the same word that commonly appears in the first language record 251 and the second language record 252.
[0113] In another example, the sixth phrase 356 may indicate being located at the same location as the seventh phrase 357 using different words from the seventh phrase 357. For example, the first agent Z1 "lit the stove", the second agent Z2 "is stir-frying beef" (both referring to being in the kitchen), the linguistic expression of the characteristics of the first object is included in the first language record 251, and the linguistic expression of the characteristics of the second object (an object different from the first object) arranged at the same location as the first object is included in the second language record 252. In one example, the sixth phrase 356 and the seventh phrase 357 may be detected by the above detection method. Also, in one example, similar to the second phrase 352 and the third phrase 353, whether each object is arranged at the same location may be specified from the premise information. When using a large-scale generation model for the detection of the phrase set 35, the information processing device 1 may provide the premise information to the large-scale generation model together with the first language record 251 and the second language record 252, so that the large-scale generation model detects the phrase set 35 (the sixth phrase 356, the seventh phrase 357) referring to different objects arranged at the same location from each language record (251, 252). Alternatively, the information processing device 1 may cause the large-scale generation model to detect the phrase set 35 referring to different objects arranged at the same location from each language record (251, 252) without providing the premise information. Whether each object is arranged at the same location may be specified based on rules.
[0114] As in the above example, different words can refer to the same location. According to an example of the present embodiment, by using this, at least one of the first movement trajectory 201 of the first agent Z1 and the second movement trajectory 202 of the second agent Z2 can be appropriately corrected. Also, by not limiting the search range of the set of the sixth phrase 356 and the seventh phrase 357 to the same word, a set of the sixth phrase 356 and the seventh phrase 357 that hints at being located at the same location can be widely obtained from the first language record 251 and the second language record 252. Thereby, an improvement in the accuracy of at least one of the first movement trajectory 201 and the second movement trajectory 202 can be expected.
[0115] Note that the time corresponding to the sixth phrase 356 may be appropriately determined according to at least one of the sixth phrase 356 and the measurement time of the sixth phrase 356. In one example, the time corresponding to the sixth phrase 356 may be the measurement time (such as the speaking time) of the sixth phrase 356. Similar to the time of the first phrase 301 described above, the time corresponding to the sixth phrase 356 may be defined to indicate a single point in time, or may be defined as a time having a certain width. The time corresponding to the sixth phrase 356 may be adjusted according to the movement status of the first agent Z1 that appears in the first movement trajectory 201. In one example, the time corresponding to the sixth phrase 356 may be the time specified from the first movement trajectory 201 when the first agent Z1 is stopped in the vicinity of the time (such as the measurement time) specified from the sixth phrase 356. Also, in one example, when the sixth phrase 356 is related to the first agent Z1 itself, such as "lit the stove", the time corresponding to the sixth phrase 356 may be treated as the time of the first agent Z1. On the other hand, when the sixth phrase 356 is related to the second agent Z2, such as "BBB (the second agent Z2) is still taking luggage out of the refrigerator", the time corresponding to the sixth phrase 356 may be treated as the time of the second agent Z2. Except for the point where the first agent Z1 and the second agent Z2 are interchanged and the sixth phrase 356 is replaced with the seventh phrase 357, the time corresponding to the seventh phrase 357 may be the same as the time corresponding to the sixth phrase 356.
[0116] Also, in one example, the fourth combination 404 may be composed of the time of the first agent Z1 and the time of the second agent Z2. For example, this corresponds to a case where "Turned on the stove" is detected as the sixth phrase 356 and "Stir-frying beef" is detected as the seventh phrase 357. In this case, the time corresponding to the sixth phrase 356 may be treated as the time of the first agent Z1, and the time corresponding to the seventh phrase 357 may be treated as the time of the second agent Z2. In another example, the fourth combination 404 may be composed of multiple times of either the first agent Z1 or the second agent Z2. For example, this corresponds to a case where "Arrived at the refrigerator" is detected as the sixth phrase 356 and "AAA (the first agent Z1) is still taking luggage out of the refrigerator" is detected as the seventh phrase 357. In this case, both the time corresponding to the sixth phrase 356 and the time corresponding to the seventh phrase 357 may be treated as the time of the first agent Z1.
[0117] Also, similar to the cases of the second phrase 352 and the third phrase 353, typically, a set of the sixth phrase 356 and the seventh phrase 357 may constitute a phrase set 35. However, the number of phrases referring to the same location is not limited to two and may be three or more. In addition to the first phrase (one of the sixth phrase 356 and the seventh phrase 357) and the second phrase (the other of the sixth phrase 356 and the seventh phrase 357), there may be phrases after the third phrase. In this case, the phrase set 35 may be composed of three or more phrases including the sixth phrase 356 and the seventh phrase 357. Alternatively, the phrase set 35 may be composed of two phrases (the sixth phrase 356, the seventh phrase 357), and three or more phrases indicating the same location may be composed of two or more phrase sets 35. Accordingly, similar to the first combination 401 etc., the number of times located at the same location is not limited to two and may be three or more. The fourth combination 404 may be composed of three or more times (times corresponding to each phrase). Alternatively, the fourth combination 404 may be composed of two times, and three or more times located at the same location may be composed of two or more fourth combinations 404.
[0118] In one example, the sixth phrase 356 may indicate being located at the same location as the seventh phrase 357 using at least either the same word or a different word as the seventh phrase 357. In one example, the information processing apparatus 1 may detect both the phrase 30 (the fourth phrase 304, the fifth phrase 305) and the phrase set 35 (the sixth phrase 356 and the seventh phrase 357) from the first language record 251 and the second language record 252. Also, in one example, the information processing apparatus 1 may use both the correction process of FIG. 2 (correction within one agent) and the correction process of FIG. 4 (correction between a plurality of agents). The information processing apparatus 1 may execute the correction process of FIG. 2 also for the first agent Z1 of FIG. 4. The information processing apparatus 1 may replace the first agent Z1 with the second agent Z2 and execute the correction process of FIG. 2 also for the second agent Z2.
[0119] (Correction method) The correction method within the above-mentioned one agent may also be adopted as a correction method among a plurality of agents. That is, it is not necessary to question which position at each time of the extracted combination 40 is the true value. As long as it is a method of correcting the position at each time so as to bring the positions at each time closer, the method of correcting the movement trajectory 20 (the first movement trajectory 201, the second movement trajectory 202, etc.) is not particularly limited and may be appropriately selected according to the embodiment. A known method may be adopted as the correction method of the movement trajectory 20. At least one of the first movement trajectory 201 and the second movement trajectory 202 may be corrected so as to bring the positions at each time included in each of the one or more combinations 40 closer. In one example, when at least one of the fourth phrase 304 and the fifth phrase 305 is detected, at least one of the first movement trajectory 201 and the second movement trajectory 202 may be corrected so as to bring the positions at each time included in the third combination 403 closer. When the sixth phrase 356 and the seventh phrase 357 are detected, at least one of the first movement trajectory 201 and the second movement trajectory 202 may be corrected so as to bring the positions at each time included in the fourth combination 404 closer. In one example, when the phrase (the fourth phrase 304, the fifth phrase 305, the sixth phrase 356, the seventh phrase 357, the eighth phrase, the ninth phrase) includes the word of the object, the position to be corrected may be at least either the position of the corresponding agent Z (the first agent Z1, the second agent Z2) or the position of the object related to the agent Z.
[0120] FIG. 5 schematically shows an example of correcting movement trajectories 20 (a first movement trajectory 201 and a second movement trajectory 202) between two agents (a first agent Z1 and a second agent Z2). In an example of FIG. 5, a scene is assumed in which the first movement trajectory 201 includes position information at times t11, t12, and t13, and the second movement trajectory 202 includes position information at times t21, t22, and t23. In addition, a scene is assumed in which the time t12 of the first agent Z1 and the time t22 of the second agent Z2 are extracted as the time of the third combination 403 or the fourth combination 404. When combinations 40 (including the third combination 403 and the fourth combination 404) are extracted between the movement trajectories 20 of a plurality of agents, the combination 40 may include the time of the first agent Z1 and the time of the second agent Z2. In this case, the time of the first agent Z1 and the time of the second agent Z2 may be the same as each other or different from each other. In an example of FIG. 5, the times t12 and t22 may be the same as each other or different from each other. The information processing apparatus 1 may appropriately correct at least one of the first movement trajectory 201 and the second movement trajectory 202 so as to bring the position at time t12 and the position at time t22 closer to each other.
[0121] In an example of FIG. 5, the positions at time t12 of the first movement trajectory 201 and at time t22 of the second movement trajectory 202 are both corrected. In one example, when the combination 40 includes the times of both the first agent Z1 and the second agent Z2, correcting at least one of the first movement trajectory 201 and the second movement trajectory 202 may be constituted by correcting the first movement trajectory 201 and the second movement trajectory 202. However, the method of correcting the movement trajectory 20 need not be limited to such an example. In another example, the positions at some of the plurality of times included in the combination 40 may be fixed without being corrected. For example, in an example of FIG. 5, either the position at time t12 of the first movement trajectory 201 or the position at time t22 of the second movement trajectory 202 may be fixed without being corrected. In this case, even if the combination 40 includes the times of both the first agent Z1 and the second agent Z2, either one of the first movement trajectory 201 and the second movement trajectory 202 may not be corrected and only the other may be corrected.
[0122] The above "Arrived at the refrigerator" is detected as the sixth phrase 356, and "AAA (the first agent Z1) is still taking out luggage from the refrigerator" is detected as the seventh phrase 357. Cases like this, the combination 40 (the fourth combination 404, the fifth combination, the sixth combination, etc.) may be constituted by only the time of either the first agent Z1 or the second agent Z2. When the combination 40 includes only the time of the first agent Z1, correcting either one of the first movement trajectory 201 and the second movement trajectory 202 may be constituted by correcting the first movement trajectory 201. When the combination 40 includes only the time of the second agent Z2, correcting either one of the first movement trajectory 201 and the second movement trajectory 202 may be constituted by correcting the second movement trajectory 202. When the combination 40 includes only the time of either the first agent Z1 or the second agent Z2, each time included in the combination 40 may be different.
[0123] Also, in one example, similar to the case of the above-mentioned one agent, priorities may be set for the positions at each time included in the combination 40 (the third combination 403, the fourth combination 404). The information processing apparatus 1 may adjust the correction amount of the position at each time according to the set priority. For example, the information processing apparatus 1 may decrease the correction amount of the position at the target time as the set priority is higher, and increase the correction amount of the position at the target time as the priority is lower. The priority of the position at each time may be set by any method. In one example, the priority may be set according to the measurement accuracy of the position information.
[0124] Also, in one example, in the case of using the above-mentioned inertial sensor M10, the case of using at least one of the camera M15 and the LiDAR M16 (using SLAM analysis), etc., each movement trajectory 20 may be constituted by relative position information. In this case, due to the difference in the measurement start point, the coordinate systems of the respective movement trajectories 20 may be different. When the coordinate systems of the respective movement trajectories 20 are different, the collation between the respective movement trajectories 20 (between the first movement trajectory 201 and the second movement trajectory 202) is not obvious. Regarding this, in one example of the present embodiment, the combination 40 (the third combination 403, the fourth combination 404) may include the times of both the first agent Z1 and the second agent Z2. Thereby, by bringing the corresponding positions of the first movement trajectory 201 and the corresponding positions of the second movement trajectory 202 closer, the first movement trajectory 201 and the second movement trajectory 202 can be collated. Correcting at least one of the first movement trajectory 201 and the second movement trajectory 202 may include collating such first movement trajectory 201 and second movement trajectory 202. Also in this respect, it is possible to suitably improve the accuracy of the movement trajectory 20 (the first movement trajectory 201, the second movement trajectory 202).
[0125] Also, similar to the case of the above-mentioned single agent, when speed information (displacement information) is obtained, such as in the case where an inertial sensor is deployed or in the case where SLAM analysis is utilized, for the method of correcting each movement trajectory 20 (the first movement trajectory 201, the second movement trajectory 202), it is preferable to adopt an optimization method that corrects to bring the positions at each time of the same location closer while suppressing changes in the speed information. For such a correction method, the type of the optimization method does not particularly need to be limited and may be appropriately selected according to the embodiment. Known methods such as loop closure and time-series smoothing (Kalman smoother, etc.) may be adopted as the optimization method. Further, in the case of correcting between a plurality of agents, the following correction method may be adopted as the optimization method.
[0126] FIG. 6 schematically shows an example of a method for correcting the movement trajectory 20 between a plurality of agents. In the example of FIG. 6, it is assumed a scenario where the movement trajectory 20i includes the positions of each object (OA, OB, OC) due to the detection of the object OA, the object OB, and the object OC from the language record 25 of the corresponding agent Z. It is assumed a scenario where the movement trajectory 20j includes the positions of each object (OB, OC, OD) due to the detection of the object OB, the object OC, and the object OD from the language record 25 of the corresponding agent Z. In addition, it is assumed a scenario where a two-dimensional coordinate system is adopted for the coordinate system of each movement trajectory (20i, 20j). The movement trajectory 20i is an example of the i-th movement trajectory 20 and the movement trajectory 20j is an example of the j-th movement trajectory 20. One of the movement trajectory 20i and the movement trajectory 20j may be the first movement trajectory 201 and the other may be the second movement trajectory 202.
[0127] The correction method in FIG. 6 optimizes the parameters of rotation and translation of each coordinate system and corrects each movement trajectory (201, 202) so that the (identical or similar) objects evaluated to be located at the same location are arranged closer. (i) k and v (i) k indicate the position and name of the k-th object in the i-th coordinate system. R(i) and t (i) represents the i-th rotation matrix and translation vector. These are the unknown parameters to be optimized. The transformed position can be expressed as "q (i) k =R (i) ·p (i) k +t (i) ". In one example, the optimization goal can be defined as Equation 1 below.
[0128]
Equation
[0129] Here, Q is the set of all objects. S(v (i) k , v (j) l ) represents the similarity between two objects (v (i) k , v (j) l ). The more it is evaluated to be located at the same position, the larger S(v (i) k , v (j) l ) may be set (in the example of Figure 6, the combinations of object OB - object OB, object OC - object OC). On the other hand, the less it is evaluated to be located at the same position, the smaller S(v (i) k , v (j) l ) may be set (in the example of Figure 6, combinations other than the above). By minimizing L with respect to {R (i) , t (i)} (where i ranges from 1 to N), each movement trajectory 20 can be corrected while being aligned. N is the number of movement trajectories 20 (coordinate systems). N can be 2 or more. This minimization of L can be performed by the gradient descent method. A known optimizer such as Adam may be used for the calculation of the gradient descent method to minimize L.
[0130] In an example of the present embodiment, when each movement trajectory 20 is constituted by relative position information, the origin of the coordinate system of each movement trajectory 20 and the directions of each axis are set by the position and orientation of the positioning sensor M1 (inertial sensor M10, camera M15, LiDAR M16, etc.) at the time when the measurement of the position starts. This may be the case. As described above, the collation of the plurality of movement trajectories 20 may not be obvious. On the other hand, by using the correction method of FIG. 6, a plurality of objects in different coordinate systems can be aligned, and they can be described by unified map information.
[0131] Note that when a plurality of agents Z are given, the number of agents Z does not have to be limited to two, and may be three or more. When three or more agents Z are given, the above correction method may be appropriately extended so as to be applied to three or more agents Z. The information processing apparatus 1 may correct at least any one of the three or more movement trajectories 20 among the three or more agents Z by the extended correction method. Further, when three or more agents Z are given, two agents Z may be appropriately selected from the three or more agents Z by any method such as exhaustive search. By treating one of the selected two agents Z as the first agent Z1 and the other as the second agent Z2, the above correction method may be directly applied to three or more agents Z.
[0132] Further, the information processing apparatus 1 may integrate two or more movement trajectories 20 and two or more language records 25 obtained from two or more agents Z. The information processing apparatus 1 may treat the integration result of the movement trajectory 20 and the integration result of the language record 25 as the movement trajectory 20 and the language record 25 of one agent Z.
[0133] [Output] In a typical example, outputting one or more corrected movement trajectories 20 may be configured by directly outputting at least a part of the one or more corrected movement trajectories 20. However, regarding the corrected movement trajectory 20, the content of the information to be output is not limited to such an example and may be appropriately determined according to the embodiment. In another example, the information processing apparatus 1 may perform arbitrary information processing on the movement trajectory 20 during or after correction. Outputting one or more corrected movement trajectories 20 may be configured by outputting the result of this information processing.
[0134] The information processing for the movement trajectory 20 may be appropriately selected according to the embodiment. In one example, the information processing apparatus 1 may generate map information in the activity range of one or more agents Z from one or more corrected movement trajectories 20. When a plurality of agents Z are provided, the information processing apparatus 1 may generate map information by combining a plurality of movement trajectories 20. The information processing apparatus 1 may output the generated map information. That is, outputting one or more corrected movement trajectories 20 may include generating map information from one or more movement trajectories 20 and outputting the generated map information. When the language expression of an object is detected from the language record 25, the information processing apparatus 1 may generate map information indicating the positional relationship of the detected object.
[0135] FIG. 7 schematically shows an example of a scene for generating map information of an object. The method for generating map information is not particularly limited and may be appropriately selected according to the embodiment. In one example, the information processing apparatus 1 may extract the language expression of an object from each of one or more language records 25. The information processing apparatus 1 may extract the position (position information) at the same time as the time of the extracted language expression from each of one or more corrected movement trajectories 20. The information processing apparatus 1 may generate map information indicating the position of the object by combining the extracted language expression of the object and the position (position information).
[0136] In an example of FIG. 7, it is assumed that language expressions of objects (OO1 to OO5) are detected at times t1 to t6, and language expressions of the same object (OO2) are detected at times t2 and t5. By combining the language expressions of the objects (OO1 to OO5) and the position information at each of the times t1 to t6, map information indicating the positional relationship of the five objects (OO1 to OO5) can be generated.
[0137] Note that when a plurality of positions (position information) are extracted from one or more movement trajectories 20 as in the object OO2 in FIG. 7, the position of the corresponding object may be arbitrarily expressed using the plurality of extracted positions. For example, the plurality of extracted positions may be used as they are as the positions of the corresponding object. For example, the information processing apparatus 1 may determine a representative position from among the plurality of positions. The representative position may be determined by any method. The determined representative position may be used as the position of the corresponding object. Also, for example, the information processing apparatus 1 may calculate a statistical quantity (average value, median, etc.) of the plurality of positions. The position of the calculated statistical quantity may be used as the position of the corresponding object.
[0138] Also, when extracting the language expression of an object from one or more language records 25, the information processing apparatus 1 may further extract from the one or more language records 25 the language expression regarding the attribute of the extracted object. The attribute of the object may include, for example, dynamic attributes such as the number of objects and usage status. If the language expression of the attribute can be extracted, the information processing apparatus 1 may include, in the map information, the attribute indicated by the extracted language expression after associating it with the corresponding object. Thereby, map information that enables grasping of the position and attribute of the object
[0139] The generated map information may be used for any task. In one example, the generated map information may be used for object search. For example, when agent Z is operating in a store and the object is a product, the map information may be used to search for the location of the product. The generated map information may be used on information processing device 1 or on other computers other than information processing device 1.
[0140] In one example of this embodiment, cases where the inertial sensor M10 is used, cases where at least one of the camera M15 and the LiDAR M16 is used (using SLAM analysis), etc. The movement trajectory 20 may be constituted by relative position information. In this case, due to differences in the measurement start points, the coordinate systems of the respective movement trajectories 20 may be different. In contrast, in one example of this embodiment, as described above, by using at least one of the phrase 30 and the phrase set 35 indicating the same location (such as the linguistic expression of an object arranged at the same location) as a clue, a plurality of movement trajectories 20 can be appropriately collated. Thereby, the information processing device 1 can integrate a plurality of movement trajectories 20 obtained from a plurality of agents Z to generate map information of the object. Since the position (position information) of the object can be obtained from each agent Z, it is possible to expect facilitation of the accumulation of map information.
[0141] §2 Configuration Example [Hardware Configuration] FIG. 8 schematically shows an example of the hardware configuration of the information processing device 1 according to this embodiment. In one example, the information processing device 1 may be configured as a computer to which a control unit 11, a storage unit 12, a communication module 13, an input device 14, and an output device 15 are electrically connected.
[0142] The control unit 11 is configured to execute information processing based on programs and various data. For example, the control unit 11 includes a CPU (Central Processing Unit), which is a hardware processor, a RAM (Random Access Memory), a ROM (Read Only Memory), etc. That's right. The control unit 11 (CPU) is an example of a processor resource.
[0143] The storage unit 12 is configured to hold arbitrary data. For example, the storage unit 12 may include a hard disk drive, a solid state drive, a semiconductor memory, etc. The storage unit 12, the RAM, and the ROM are examples of the memory resources of the information processing apparatus 1. In an example of the present embodiment, the storage unit 12 may store various information such as a correction program 81, one or more movement trajectories 20, one or more language records 25, etc.
[0144] The correction program 81 is a program for causing the information processing apparatus 1 to execute information processing (FIG. 10 described later) related to the correction of the movement trajectory 20. The correction program 81 includes a series of instructions for the said information processing. The correction program 81 is an example of the program of the present disclosure.
[0145] In one example, at least any one of the correction program 81, at least a part of the movement trajectory 20, and at least a part of the language record 25 may be stored in the storage medium 91 instead of or together with the storage unit 12. The storage medium 91 is configured to store the various information (stored programs, etc.) by an electrical, magnetic, optical, mechanical, or chemical action so that a machine such as a computer can read the information. The storage unit 12 and the storage medium 91 are an example of a non-temporary storage medium. The information processing apparatus 1 may acquire at least any one of the correction program 81, at least a part of the movement trajectory 20, and at least a part of the language record 25 from the storage medium 91. The storage medium 91 may be a disk-type storage medium (CD, DVD, etc.), or may be a storage medium other than the disk-type such as a semiconductor memory (flash memory, etc.). For reading the information stored in the storage medium 91, any drive device may be used. The type of the drive device may be selected according to the storage medium 91. The drive device may be connected to the information processing apparatus 1 by any method. The storage medium 91 may include an external storage device.
[0146] The communication module 13 is configured to perform wired or wireless communication via a network. The communication module 13 may be configured by, for example, a wired LAN (Local Area Network) module, a wireless LAN module, or the like. The standard of the network may not be particularly limited and may be appropriately selected according to the embodiment. For example, the type of the network may be appropriately selected from the Internet, a wireless communication network, a mobile communication network, a telephone network, a dedicated network, or the like. The information processing apparatus 1 may execute data communication with another computer via the communication module 13. In one example, the information processing apparatus 1 may acquire at least one of at least a part of the movement trajectory 20 and at least a part of the language record 25 via the network by using the communication module 13. In one example, when the information processing apparatus 1 is involved in generating at least a part of the movement trajectory 20, the communication module 13 may constitute a positioning sensor M1 (a positioning module by wireless communication) together with the control unit 11.
[0147] The input device 14 is configured to receive input of information. The input device 14 may be composed of, for example, a camera, a microphone, a mouse, a keyboard, a touch panel, an operator, etc. The output device 15 is configured to output information. The output device 15 may be composed of, for example, a display, a speaker, etc. The information processing device 1 may be operated by using the input device 14 and the output device 15. The input device 14 and the output device 15 may be directly connected to the information processing device 1, or may be indirectly connected via at least one of the communication module 13 and the external interface. The external interface may be appropriately configured to be connected to an external device, either wired or wirelessly, by, for example, a USB (Universal Serial Bus) port, a dedicated port, etc. The input device 14 and the output device 15 may be at least partially integrally configured by a touch panel display or the like. In one example, the input device 14 may be used as the input device M2.
[0148] Regarding the specific hardware configuration of the information processing device 1, depending on the embodiment, omission, replacement, and addition of components are possible as appropriate. For example, the control unit 11 may include a plurality of hardware processors. The hardware processor may be composed of a microprocessor, an FPGA (field-programmable gate array), a DSP (digital signal processor), a GP U (Graphics Processing Unit), an ASIC (application specific integrated circuit), etc. At least one of the communication module 13, the input device 14, and the output device 15 At least one of them may be omitted. In one example, when the information processing apparatus 1 is involved in the generation of the movement locus 20, the positioning sensor M1 may be connected to the information processing apparatus 1. At least one of the correction program 81, at least a part of the movement locus 20, and at least a part of the language record 25 may be stored in an external storage device such as a NAS (Network Attached Storage). The external storage device is also an example of a non-transitory storage medium. The information processing apparatus 1 may be configured by a plurality of computers. In this case, the hardware configurations of the respective computers may or may not match. The information processing apparatus 1 may be, in addition to a computer designed specifically for the provided service, a general-purpose server device, a general-purpose PC (Personal Computer), a notebook PC, a terminal device, or the like. The terminal device may include user terminals such as smartphones and tablet terminals. The terminal device may include a control device, an in-vehicle device, or the like.
[0149] [Software Configuration] FIG. 9 schematically shows an example of the software configuration of the information processing apparatus 1 according to the present embodiment. The control unit 11 of the information processing apparatus 1 executes instructions included in the correction program 81 stored in the storage unit 12 by the CPU. Thereby, the information processing apparatus 1 operates as a computer including the first acquisition unit 111, the second acquisition unit 112, the language analysis unit 113, the correction unit 114, and the output processing unit 115 as software modules. That is, in one example, each software module of the information processing apparatus 1 may be realized by the control unit 11 (CPU).
[0150] The first acquisition unit 111 is configured to acquire one or more movement trajectories 20 respectively measured in time series by the positioning sensor M1 for each of one or more agents Z. The second acquisition unit 112 is configured to acquire one or more language records 25 respectively measured in synchronization with each of the one or more movement trajectories 20 for each of one or more agents Z. The language analysis unit 113 is configured to extract one or more combinations 40 of times related to the same location by identifying phrases 30 or sets 35 of phrases indicating the same location in the acquired one or more language records 25. The correction unit 114 is configured to correct one or more movement trajectories 20 so that the positions of the times of each of the extracted one or more combinations 40 approach each other. The output processing unit 115 is configured to output the corrected one or more movement trajectories 20.
[0151] Note that in one example of the above-described embodiment, each software module of the information processing apparatus 1 is realized by a general-purpose CPU. However, the method for realizing each of the above modules is not limited to such an example and may be appropriately changed according to the embodiment. Part or all of the above software modules may be realized by one or more dedicated processors or chip sets. Each of the above modules may be realized as a hardware module. Regarding the software configuration of the information processing apparatus 1, omission, replacement, and addition of modules may be appropriately performed according to the embodiment.
[0152] §3 Operation Example FIG. 10 is a flowchart showing an example of a processing procedure for correcting the movement trajectory 20 by the information processing apparatus 1 according to the present embodiment. The following processing procedure is an example of an information processing method (correction method) executed by a computer. The following processing procedure is only an example, and each step may be changed as much as possible. Also, regarding the following processing procedure, omission, replacement, and addition of steps are possible as appropriate according to the embodiment.
[0153] (Step S101) In step S101, the control unit 11 operates as a first acquisition unit 111, and acquires one or more movement trajectories 20 respectively measured in time series by the positioning sensor M1 in each of the one or more agents Z.
[0154] In one example, the positioning sensor M1 may include an inertial sensor M10. In this case, each of the one or more movement trajectories 20 may be measured by analyzing the measurement data of the inertial sensor M10. Also, in one example, the positioning sensor M1 may include at least one of a camera M15 and a LiDAR M16. In this case, each of the one or more movement trajectories 20 may be measured by performing SLAM analysis on the measurement data of at least one of the camera M15 and the LiDAR M16. It may be measured by doing so.
[0155] In one example, the one or more agents Z may include a first agent Z1. The one or more movement trajectories 20 may include a first movement trajectory 201 of the first agent Z1. In one example, the one or more agents Z may include a first agent Z1 and a second agent Z2. The one or more movement trajectories 20 may include a first movement trajectory 201 of the first agent Z1 and a second movement trajectory 202 of the second agent Z2. Note that the timing of executing the process of step S101 is not limited to such an example. Step S101 may be executed at any timing before the process of step S106 described later. When acquiring the one or more movement trajectories 20, the control unit 11 proceeds to the next step S102 with the process.
[0156] (Step S102) In step S102, the control unit 11 operates as a second acquisition unit 112, and acquires one or more language records 25 respectively measured in synchronization with each of the one or more movement trajectories 20 in each of the one or more agents Z.
[0157] In one example, the one or more language records 25 may include the first language record 251 of the first agent Z1. In one example, the one or more language records 25 may include the first language record 251 of the first agent Z1 and the second language record 252 of the second agent Z2. Note that the timing for executing the process of step S102 is not limited to such an example. Step S102 may be executed before step S101. When acquiring the one or more language records 25, the control unit 11 proceeds to the next step S103.
[0158] (Steps S103 to S105) In step S103, the control unit 11 operates as the language analysis unit 113 and analyzes the acquired one or more language records 25. Thereby, the control unit 11 searches for a phrase 30 or a set of phrases 35 indicating the same location in the acquired one or more language records 25. The method of language analysis is not particularly limited and may be appropriately selected according to the embodiment. Known methods may be adopted for language analysis.
[0159] In step S104, the control unit 11 determines the branch destination of the process according to the search result of step S103. When a phrase 30 or a set of phrases 35 indicating the same location is identified (detected), the control unit 11 proceeds to the next step S105. On the other hand, when a phrase 30 or a set of phrases 35 indicating the same location is not identified (detected), the control unit 11 ends the processing procedure according to this operation example.
[0160] In step S105, the control unit 11 operates as the language analysis unit 113 and extracts one or more combinations 40 of times related to the same location according to the identified phrase 30 or set of phrases 35.
[0161] In one example, in step S103 and step S105, the control unit 11 may extract one or more combinations 40 of times related to the same location by identifying the phrase 30 or the set of phrases 35 included in the first language record 251. For example, the control unit 11 may identify a first phrase 301 indicating that it was located at the same location at a plurality of times from the first language record 251, and extract a first combination 401 composed of the plurality of times corresponding to the identified first phrase 301. Also, for example, the control unit 11 may identify a second phrase 352 and a third phrase 353 indicating the same location from the first language record 251, and extract a second combination 402 composed of the time corresponding to the identified second phrase 352 and the time corresponding to the third phrase 353. The second phrase 352 may indicate that it was located at the same location as the third phrase 353 using at least one of the same word and a different word as the third phrase 353.
[0162] Also, in one example, in step S103 and step S105, the control unit 11 may extract one or more combinations 40 of times related to the same location by identifying the phrase 30 or the set of phrases 35 included in at least one of the first language record 251 and the second language record 252. For example, the control unit 11 may identify a fourth phrase 304 indicating that the first agent Z1 is at the same location as the second agent Z2 from the first language record 251, and extract a third combination 403 composed of the time of the first agent Z1 and the time of the second agent Z2 corresponding to the identified fourth phrase 304. For example, the control unit 11 may identify a fifth phrase 305 indicating that the second agent Z2 is at the same location as the first agent Z1 from the second language record 252, and the third set composed of the time of the first agent Z1 and the time of the second agent Z2 corresponding to the identified fifth phrase 305 The matching 403 may be extracted. Also, for example, the control unit 11 may identify the sixth phrase 356 from the first language record 251 and identify the seventh phrase 357 indicating the same location as the sixth phrase 356 from the second language record 252. Thereby, the control unit 11 may extract the fourth combination 404 composed of the time corresponding to the sixth phrase 356 and the time corresponding to the seventh phrase 357. In one example, the sixth phrase 356 may indicate that it is located at the same location as the seventh phrase 357 using at least one of the same word and different words as the seventh phrase 357. When one or more combinations 40 of times related to the same location are extracted, the control unit 11 proceeds to the next step S106.
[0163] (Step S106) In step S106, the control unit 11 operates as the correction unit 114 and corrects one or more movement trajectories 20 so that the positions of the times of each of the one or more extracted combinations 40 approach each other.
[0164] In one example, the control unit 11 may correct the first movement trajectory 201 so that the positions of the times of each of the one or more extracted combinations 40 approach each other. For example, the control unit 11 may correct the first movement trajectory 201 so that the positions of the times included in each of the one or more first combinations 401 approach each other. Also, for example, the control unit 11 may correct the first movement trajectory 201 so that the positions of the times included in each of the one or more second combinations 402 approach each other.
[0165] In one example, the control unit 11 may correct at least one of the first movement trajectory 201 and the second movement trajectory 202 so that the positions of the times of each of the extracted one or more combinations 40 approach each other. For example, the control unit 11 may correct at least one of the first movement trajectory 201 and the second movement trajectory 202 so that the positions of the times included in each of the extracted one or more third combinations 403 approach each other. Further, for example, the control unit 11 may correct at least one of the first movement trajectory 201 and the second movement trajectory 202 so that the positions of the times included in each of the extracted one or more fourth combinations 404 approach each other. When correcting the one or more movement trajectories 20, the control unit 11 proceeds to the next step S107 for processing.
[0166] (Step S107) In step S107, the control unit 11 operates as the output processing unit 115 and outputs the corrected one or more movement trajectories 20.
[0167] The content and output destination of the information to be output may be appropriately selected according to the embodiments. In one example, the control unit 11 may output the one or more corrected movement trajectories 20 as they are. In one example, the control unit 11 may perform arbitrary information processing on the movement trajectory 20 during or after correction and output the execution result of the information processing. For example, the control unit 11 may generate map information in the activity range of one or more agents Z from the one or more corrected movement trajectories 20. Also, for example, the control unit 11 may extract the language expression of an object from one or more language records 25. The control unit 11 may extract the position (position information) at the same time as the time of the extracted language expression from the one or more corrected movement trajectories 20. The control unit 11 may generate map information indicating the position of the object by combining the extracted language expression of the object and the position (position information). Also, the control unit 11 may further extract from the one or more language records 25 the language expression regarding the attribute of the extracted object. The control unit 11 may include the attribute (attribute information) indicated by the extracted language expression in the map information after associating it with the corresponding object. The output destination is not particularly limited and may be appropriately selected according to the embodiment. The output destination may be, for example, a RAM, the storage unit 12, the output device 15, the storage medium 91, an external computer, an external storage device, etc. When the output of the one or more corrected movement trajectories 20 is completed, the control unit 11 ends the processing procedure according to this operation example.
[0168] §4 Modification As described above, the embodiments of the present disclosure have been described in detail, but the description so far is merely an exemplification of the present disclosure in all respects. The processes and means described in the present disclosure can be freely combined and implemented as long as no technical contradiction occurs. Also, in the above embodiments, various improvements or modifications may be appropriately made.
Description of Reference Numerals
[0169] 1... Information processing apparatus, 20... Movement trajectory, 25... Language record, 30... Phrase, 35... Set (of phrases), 40... combination (of times) M1... positioning sensor, M2... input device Z... agent
Claims
1. Obtaining one or more movement trajectories respectively measured in time series by a positioning sensor for each of the agents of 1 or more; For each of the one or more agents, obtaining one or more language records measured in synchronization with each of the one or more movement trajectories; In the obtained one or more language records, extracting one or more combinations of times related to the same location by identifying a phrase or a set of phrases indicating the same location; Correcting the one or more movement trajectories so that the positions of the times of each of the extracted one or more combinations approach each other; and Outputting the corrected one or more movement trajectories; A control unit configured to execute the above is provided; An information processing apparatus.
2. The one or more agents include a first agent; The one or more movement trajectories include a first movement trajectory of the first agent; The one or more language records include a first language record of the first agent; Extracting one or more combinations of times related to the same location includes extracting one or more combinations of times related to the same location by identifying a phrase or a set of phrases included in the first language record; Correcting the one or more movement trajectories includes correcting the first movement trajectory so that the positions of the times of each of the extracted one or more combinations approach each other; The information processing apparatus according to claim 1.
3. Identifying a phrase or a set of phrases included in the first language record includes identifying a first phrase indicating being located at the same location at a plurality of times; The one or more combinations to be extracted include a first combination composed of the plurality of times corresponding to the identified first phrase; The information processing apparatus according to claim 2.
4. Identifying a phrase or a set of phrases included in the first language record includes identifying a second phrase and a third phrase indicating the same location as each other; The one or more combinations to be extracted include a second combination composed of the time corresponding to the identified second phrase and the time corresponding to the third phrase; The information processing apparatus according to claim 2.
5. The second phrase uses words different from those of the third phrase and indicates being located at the same location as the third phrase; The information processing apparatus according to claim 4.
6. The one or more agents include a first agent and a second agent, The one or more movement trajectories include a first movement trajectory of the first agent and a second movement trajectory of the second agent, The one or more language records include a first language record of the first agent and a second language record of the second agent, Extracting one or more combinations of times related to the same location includes identifying a phrase or a set of phrases included in at least one of the first language record and the second language record to extract one or more combinations of times related to the same location, Correcting the one or more movement trajectories includes correcting at least one of the first movement trajectory and the second movement trajectory so that the positions at the times of the one or more extracted combinations approach each other, including correcting, The information processing apparatus according to claim 1.
7. Identifying a phrase or a set of phrases included in at least one of the first language record and the second language record includes identifying a fourth phrase indicating that the first agent is at the same location as the second agent from the first language record, or identifying a fifth phrase indicating that the second agent is at the same location as the first agent from the second language record, The one or more combinations to be extracted include a third combination composed of the time of the first agent and the time of the second agent corresponding to the identified fourth phrase or the fifth phrase, The information processing apparatus according to claim 6.
8. Identifying a phrase or a set of phrases included in at least one of the first language record and the second language record includes identifying a sixth phrase from the first language record and identifying a seventh phrase indicating the same location as the sixth phrase from the second language record, The one or more combinations to be extracted include a fourth combination composed of the time corresponding to the sixth phrase and the time corresponding to the seventh phrase, The information processing apparatus according to claim 6.
9. The sixth phrase indicates the same location as the seventh phrase using words different from those of the seventh phrase, The information processing apparatus according to claim 8.
10. The positioning sensor includes an inertial sensor, Each of the one or more movement trajectories is measured by analyzing measurement data of the inertial sensor. The information processing apparatus according to claim 1.
11. The positioning sensor includes at least one of a camera and LiDAR (light detection and ranging), and each of the one or more movement trajectories is measured by performing SLAM (simultaneous localization and mapping) analysis on measurement data of at least one of the camera and the LiDAR.
12. An information processing method executed by a computer, the method comprising: obtaining, for each of one or more agents, one or more movement trajectories respectively measured in time series by a positioning sensor; obtaining, for each of the one or more agents, one or more language records respectively measured in synchronization with each of the one or more movement trajectories; extracting, from the obtained one or more language records, one or more combinations of times related to the same location by identifying a phrase or a set of phrases indicating the same location; correcting the one or more movement trajectories so that the positions of the times of each of the extracted one or more combinations approach each other; and outputting the corrected one or more movement trajectories.
13. A program for causing a computer to execute an information processing method, the information processing method comprising: obtaining, for each of one or more agents, one or more movement trajectories respectively measured in time series by a positioning sensor; obtaining, for each of the one or more agents, one or more language records respectively measured in synchronization with each of the one or more movement trajectories; extracting, from the obtained one or more language records, one or more combinations of times related to the same location by identifying a phrase or a set of phrases indicating the same location; correcting the one or more movement trajectories so that the positions of the times of each of the extracted one or more combinations approach each other; and outputting the corrected one or more movement trajectories.
Citation Information
Patent Citations
Information display device, route setting method, and program
JP2011038983A
Inertial navigation device and program
JP2014013202A
Inertial system, method, and program
JP2014167461A
Information processing device, information processing method, and program
JP2016006611A
Information processing apparatus, information processing method, and information processing program
JP2021189116A