Information processing device, information processing method, generation method, learning model, program, and storage medium
A machine learning model trained with a loss function for point cloud coordinates and vectors improves image prediction accuracy for specific regions, addressing computational efficiency and accuracy challenges.
Patent Information
- Application Number
- JP2024028896
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-28
- Publication Date
- 2025-09-09
AI Technical Summary
Existing image prediction technologies face challenges in achieving high accuracy while minimizing computational costs and processing times, particularly in predicting specific regions within images using pixel classification or point cloud surrounding areas.
A machine learning model trained using a loss function that considers the difference between predicted and correct coordinates and vectors of a point cloud surrounding a specific area, improving prediction accuracy by constraining the relationship between points in the cloud.
Enhances prediction accuracy for specific regions in images by accurately determining the position and relationship of points in the point cloud, reducing computational costs and processing times.
Smart Images

Figure 2025131265000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device, an information processing method, a generation method, a learning model, a program, and a storage medium. [Background technology]
[0002] In recent years, techniques for predicting specific regions within an image have become known, such as segmentation, which predicts the region of a subject contained in an image, and visual grounding, which predicts specific regions within an image that correspond to a query given in natural language.
[0003] In Non-Patent Document 1, image features obtained from an input image are combined with a prompt generated from a natural language sentence, and a transformer encoder is used to classify each pixel in the image using the combined information, thereby predicting the area in the image that corresponds to the natural language. Non-Patent Document 2 proposes a technology that predicts a point cloud surrounding the area in the image that corresponds to the query (a point cloud on the periphery of the area), instead of predicting each pixel in the area in the image that corresponds to the query in the natural language sentence. [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] Bin Yan, 6 others, "Universal Instance Perception as Object Discovery and Retrieval", arXiv:2303.06674v2 [cs.CV] August 17, 2023 [Non-patent document 2] Chaoyang Zhu, 9 others, "SeqTR: A Simple yet Universal Network for Visual Grounding", arXiv:2203.16265v2 [cs.CV] July 24, 2022 Summary of the Invention [Problem to be solved by the invention]
[0005] The technology proposed in Non-Patent Document 1 performs class classification for each pixel in an image, which results in high accuracy, but has the drawback of high computational costs and long processing times. On the other hand, the technology proposed in Non-Patent Document 2 predicts only the point cloud surrounding an area, which reduces computational costs and processing times compared to when classifying for each pixel, but the accuracy of the predicted area becomes an issue.
[0006] The present invention has been made in consideration of the above-mentioned problems, and its purpose is to realize a technology that can improve prediction accuracy when predicting a specific area in an image using a point group surrounding the specific area. [Means for solving the problem]
[0007] According to the present invention, an acquisition means for acquiring an image as input information; a prediction means configured with one or more machine learning models that extracts features from the input information and predicts a specific region within the image based on the extracted features; processing means; the prediction means outputs a prediction result indicating the specific area, including coordinates of a plurality of points surrounding the specific area and information indicating a point next to each of the plurality of points; The information processing device is characterized in that the processing means trains the one or more machine learning models using a loss function based on the difference between the prediction result, which includes coordinates of multiple points surrounding the specific area and information indicating the next point for each of the multiple points, and correct data for the prediction result. [Effects of the Invention]
[0008] According to the present invention, it is possible to improve the prediction accuracy when predicting a specific region in an image using a point group surrounding the specific region. [Brief explanation of the drawings]
[0009] [Figure 1A] FIG. 1 shows an example of the configuration of a moving body according to an embodiment. [Figure 1B] FIG. 2 shows an example of the configuration of a moving body according to an embodiment. [Figure 2] FIG. 1 is a block diagram showing an example of the configuration of a control system of a moving body according to an embodiment; [Figure 3] FIG. 1 is a diagram showing an example of the functional configuration of a control unit 130 according to an embodiment. [Figure 4] FIG. 1 is a diagram illustrating a model used in region prediction processing according to an embodiment. [Figure 5] FIG. 10 is a diagram illustrating another example of a prediction result according to the embodiment. [Figure 6] 1 is a flowchart showing a series of operations for training a machine learning model used in region prediction processing according to an embodiment. [Figure 7] 10 is a flowchart showing another series of operations for training a machine learning model used in region prediction processing according to an embodiment. [Figure 8] 1 is a flowchart showing a series of operations in an inference stage of a region prediction process according to an embodiment; DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, the embodiments will be described in detail with reference to the accompanying drawings. Note that the following embodiments do not limit the scope of the invention as claimed, and not all combinations of features described in the embodiments are necessarily essential to the invention. Two or more of the features described in the embodiments may be arbitrarily combined. Furthermore, the same reference numerals are used for the same or similar components, and redundant explanations will be omitted.
[0011] In the following embodiment, a neural network model as a machine learning model, which will be described later, is executed on a mobile object such as micromobility, which is an example of an information processing device. However, the machine learning model according to this embodiment is not limited to being executed on a mobile object such as micromobility, and may also be executed on an information processing server on the cloud, which is another example of an information processing device.
[0012] Furthermore, the processes of the learning stage and the inference stage of the machine learning model according to this embodiment may both be executed on the same information processing device, or may each be executed on a separate information processing device. For example, both the process of the learning stage and the process of the inference stage of machine learning may be executed on a mobile body such as micromobility or an information processing server, or the process of the learning stage may be executed on an information processing server and the process of the inference stage may be executed on a mobile body such as micromobility.
[0013] In the following embodiments, an ultra-compact electric vehicle with a passenger capacity of about one person will be described as an example of a mobile body that is micromobility. However, micromobility may also include vehicles that travel together with people while carrying luggage instead of a person on board. Furthermore, this embodiment is not limited to examples in which the mobile body is an electric vehicle, and can also be applied to mobile bodies other than electric vehicles.
[0014] Mobile bodies such as the micromobility described above do not necessarily travel along specific, fixed routes. Furthermore, in order to be able to travel in both areas where automobiles and pedestrians are mobile, they may travel in areas where high-precision maps are not available. Therefore, the mobile body 100 according to this embodiment recognizes the travel area and generates a route using images captured by the mobile body itself, without using high-precision maps, and autonomously travels according to the generated route. In this case, the mobile body 100 according to this embodiment executes a machine learning model that appropriately predicts the area on the image corresponding to the location specified by the user's utterance, for example, in order to appropriately move to the location specified by the user's utterance.
[0015] <Configuration of moving body> An example of the configuration of the moving body 100 will be described with reference to Figures 1A and 1B. Figure 1A shows a side view of the moving body 100 according to this embodiment, and Figure 1B shows the internal configuration of the moving body 100. In the figures, arrow X indicates the front-to-rear direction of the moving body 100, with F indicating the front and R indicating the rear. Arrows Y and Z indicate the width direction (left-right direction) and up-down direction of the moving body 100.
[0016] The mobile object 100 is an electric autonomous vehicle that includes a propulsion unit 112 and uses a battery 113 as a main power source. The battery 113 is, for example, a secondary battery such as a lithium-ion battery, and the mobile object 100 is self-propelled by the propulsion unit 112 using power supplied from the battery 113. The propulsion unit 112 includes a pair of left and right drive wheels 120 that are front wheels, and one driven wheel 121 that is a rear wheel. Note that the example of the propulsion unit 112 shown in FIG. 1A is just one example, and the propulsion unit 112 may have other forms, such as a four-wheeled vehicle. Furthermore, the rear wheels are not limited to being driven wheels, and may be driven by a drive mechanism. The mobile object 100 includes, for example, a single seat 111, but may also include multiple seats.
[0017] The traveling unit 112 includes a drive mechanism 122. The drive mechanism 122 is a mechanism that uses motors 122a and 122b as drive sources to rotate the corresponding drive wheels 120. The drive mechanism 122 can move the mobile body 100 forward or backward by rotating each of the drive wheels 120. The drive mechanism 122 can also change the traveling direction of the mobile body 100 by generating a rotation difference between the motors 122a and 122b. The driven wheels 121 can rotate around the Z direction as a rotation axis.
[0018] The moving body 100 is equipped with detection units 114 to 116 that detect targets around the moving body 100. The detection units 114 to 116 are a group of external sensors that monitor the periphery of the moving body 100. In this embodiment, the detection units 114 to 116 are all imaging devices that capture images of the periphery of the moving body 100, and include, for example, an optical system such as a lens and an image sensor. However, in addition to the imaging devices, radar or lidar (Light Detection and Ranging) may also be used.
[0019] The detection units 114 are arranged, for example, in pairs at the front of the moving body 100, spaced apart in the Y direction, and are mainly used to acquire images of the area in front of the moving body 100. The detection units 114 may be configured as a single imaging device. The detection units 115 are arranged on the left and right sides of the moving body 100, respectively, and are mainly used to acquire images of the areas to the sides of the moving body 100. The detection unit 116 is arranged at the rear of the moving body 100, and is mainly used to acquire images of the area behind the moving body 100. The moving body 100 does not necessarily have to include the detection units 115 and 116.
[0020] <Configuration example of a control system for a moving object> 2 is a block diagram of a control system of the mobile object 100. The mobile object 100 includes a control unit (ECU) 130. The control unit 130 includes one or more processors including a CPU or a GPU, a memory device such as a semiconductor memory, an interface with an external device, and the like. The memory device stores programs executed by the processor and various data used by the processor for processing (e.g., weight parameters of a trained model, etc.). Multiple sets of processors, memory devices, and interfaces may be provided for different functions of the mobile object 100 and configured to be able to communicate with each other.
[0021] The control unit 130 acquires the outputs (e.g., images) of the detection units 114 to 116, input information from the operation unit 131, and audio information input from the audio input device 133, and executes various processes. The control unit 130, for example, controls the motors 122a and 122b (controls the driving of the driving unit 112), controls the display of a display panel included in the operation unit 131, and outputs audio notifications and information to the occupants of the moving object 100. In addition, as will be described later, the control unit 130 executes a process (area prediction process) of predicting an area in an image corresponding to a location specified by linguistic information. The area prediction process is executed using a machine learning model (e.g., a deep neural network).
[0022] The voice input device 133 includes, for example, a microphone, and collects voices such as speeches of passengers (users) of the moving object 100. The GNSS (Global Navigation Satellite system) sensor 134 receives GNSS signals and detects the current position of the moving object 100.
[0023] The storage device 135 includes a non-volatile recording medium that stores various data. The storage device 135 may also store programs executed by the processor, data used by the processor for processing, etc. The storage device 135 may also store various parameters of the machine learning model executed by the control unit 130 (for example, trained weight parameters and hyperparameters of a deep neural network, etc.).
[0024] The communication device 136 is a communication device that can communicate with an external device (for example, a communication terminal 140 owned by a user or an information processing server) via wireless communication such as Wi-Fi or fifth generation mobile communication.
[0025] <Example of control unit functional configuration> Next, an example of the functional configuration of the control unit 130 according to this embodiment will be described with reference to FIG. 3. The functions of the components of the control unit 130 shown in FIG. 3 are realized, for example, by one or more processors of the control unit 130 executing a program stored in a memory or the like. Note that the example shown in FIG. 3 illustrates a case in which the control unit 130 includes both a target area prediction unit 303 and a learning processing unit 304. That is, the example shown in FIG. 3 illustrates a case in which the control unit 130 is capable of executing both an inference stage process using a trained machine learning model and a learning stage process for training the machine learning model. However, if the control unit 130 only performs an inference stage process using a trained machine learning model, the control unit 130 does not need to include the learning processing unit 304. In this case, the learning stage process for training the machine learning model is executed by another information processing device.
[0026] When performing processing in the inference stage, the instruction acquisition unit 301 acquires user instructions input via the operation unit 131 or the voice input device 133. User instructions input via the voice input device 133 may be converted into spoken sentences written in natural language by voice recognition, or may be acquired as voice information including utterances in natural language. Furthermore, the user instructions may be text written in natural language input via the operation unit 131. In either case, the user instructions are acquired as linguistic information including a location specification expressed in natural language. When performing processing in the learning stage, the instruction acquisition unit 301 acquires user instructions included in training data (which will be described later) (i.e., linguistic information including a location specification expressed in natural language).
[0027] When performing processing at the inference stage, the image information acquisition unit 302 acquires the outputs (images) of the detection units 114 to 116. When performing processing at the learning stage, the image information acquisition unit 302 acquires images included in training data, which will be described later.
[0028] The target area prediction unit 303 performs area prediction processing using a machine learning model, using linguistic information specifying a location from the instruction acquisition unit 301 and an image from the image information acquisition unit 302. The machine learning model may be configured with one or more machine learning models. When performing processing at the inference stage, the target area prediction unit 303 is executed using parameters of the trained machine learning model (for example, weight parameters of an optimized neural network).
[0029] In addition to the processing by the target area prediction unit 303, the control unit 130 can recognize the position and shape of an obstacle, the driving area, etc., using image information. Recognition of the position and shape of an obstacle ahead of the moving object 100, the driving area, road structure, etc. may be performed, for example, by applying a pre-trained machine learning model for image recognition (different from the model used for area prediction processing) to images obtained from the two detection units 114. In addition, the process may include estimating the depth from the moving object 100 using the images obtained from the two detection units 114 as stereo images.
[0030] The learning processing unit 304 trains the machine learning model used in the target area prediction unit 303 to generate a trained machine learning model. The learning processing unit 304 calculates the value of a loss function based on the difference between the prediction result by the target area prediction unit 303 and the correct data for the prediction result. At this time, the machine learning model of the target area prediction unit 303 outputs the prediction result using parameters of the machine learning model in the middle of the learning stage (e.g., weight parameters of a neural network). The learning processing unit 304 changes the parameters of the machine learning model so as to reduce the value of the loss function. The learning processing unit 304 controls the learning stage processing so as to repeat prediction by the target area prediction unit 303, calculation of the value of the loss function, and change of the parameters of the machine learning model using training data.
[0031] The training data consists of multiple sets of data sets, each set consisting of an image, linguistic information expressed in natural language that includes a location specification within the image, and ground truth data that indicates an area within the image. Details of the ground truth data will be described later.
[0032] The driving control unit 305 determines a driving route to the specified location based on the area in the image corresponding to the instruction predicted by the target area prediction unit 303 and the driving area recognized using the image information, and determines a control amount for the mobile object according to the determined driving route. The driving control unit 305 is executed only when the target area prediction unit 303 performs processing in the inference stage. Any method can be used to determine a driving route using an area in the image as a target area, and known methods may be used. The driving control unit 305 further controls the driving of the mobile object 100 according to the determined control amount (for example, controls the motors 122a and 122b).
[0033] <Operation of machine learning model used in region prediction processing> With reference to FIG. 4, a machine learning model used in the region prediction process according to this embodiment will be described.
[0034] The image 401 is an image acquired by the image information acquisition unit 302, and is a photographed image or an image included in the training data. The image feature extraction unit 403 inputs the image 401 and extracts image features. The image feature extraction unit 403 can extract image features by, for example, convolution or pooling processing, but the image feature extraction unit 403 may also extract image features by other configurations such as a transformer.
[0035] The linguistic information 402 is linguistic information acquired by the instruction acquisition unit 301, and is linguistic information from the voice input device 133 or the operation unit 131, or linguistic information included in the training data. The linguistic information 402 includes a location specification expressed in natural language, such as "Pull over next to the women at the corner." The location specification corresponds to a specific area within the image.
[0036] The language feature extraction unit 404 may be configured with a machine learning model using a transformer such as BERT, or a recursive machine learning model such as GRU, etc. The language feature extraction unit 404 extracts language features included in the language information 402.
[0037] The feature fusion unit 405 fuses the image features extracted by the image feature extraction unit 403 and the language features extracted by the language feature extraction unit 404 to generate fusion features (multimodal features). The feature fusion unit 405 can generate fusion features using any configuration. The feature fusion unit 405 may include, for example, a Pixel-Word Attention Module (PWAM). For example, the PWAM inputs image features as queries to an attention mechanism and language features as keys and values to the attention mechanism, and generates fusion features by fusing the image features with the language features. The feature fusion unit 405 may combine the image features extracted by the image feature extraction unit 403 with the fusion features to generate new fusion features. Alternatively, the feature fusion unit 405 may further combine the new fusion features with language features from the language feature extraction unit 404 to generate further new fusion features.
[0038] The transformer encoder 406 is an encoder that receives the fused features generated by the feature fusion unit 405 and further encodes the multimodal features. That is, it extracts features that are more effective for the task from the input multimodal features. In the example shown in FIG. 4, the encoder that encodes the fused features is a transformer. However, the encoder may be configured using a model other than a transformer.
[0039] The transform decoder 407 receives the features encoded by the transform encoder 406 and predicts a specific region within an image. Using a transformer as a decoder makes it possible to output highly accurate prediction results. The example shown in FIG. 4 illustrates a case in which the encoded features are decoded using a transformer. However, the decoder may be configured using a machine learning model other than a transformer, such as a recurrent machine learning model. For example, when using a recurrent machine learning model such as an LSTM or GRU, it is possible to speed up processing by using a model that is relatively smaller in scale than a transformer.
[0040] The transformer decoder 407 outputs a prediction result indicating a specific region, including coordinates of a plurality of points surrounding the specific region in the image and information indicating the next point of each of the plurality of points. The information indicating the next point of each of the plurality of points is represented by a vector from each of the plurality of points to the next point.
[0041] Image 410 in FIG. 4 schematically illustrates an example in which the prediction result is superimposed on the input image. Region 411 indicates the predicted specific region. More specifically, region 411 indicates a target region in the image corresponding to the location specified in the linguistic information. In the example shown in FIG. 4, if the linguistic information 402 is "Pull over next to the women at the corner," the specific region corresponds to the location in front of the women at the corner. Points 412-1, 412-2, ..., 412-n indicate multiple points (point cloud) surrounding the specific region. The point cloud is the vertices of the specific region represented by a polygon. Furthermore, vectors 413-1, 413-2, ..., 413-n are information indicating the point next to each point in the point cloud. Each vector is represented, for example, by the difference from the coordinates of each point to the next point, and represents the relationship between points in the point cloud. In this embodiment, by including a point cloud and a vector in the prediction result, a trajectory surrounding a region with a directional point cloud is predicted. In this way, in region prediction, not only the position of the point cloud is predicted but also the relationship between the points in the point cloud is predicted. This constrains the relationship between the point clouds to be appropriate in region prediction, thereby enabling region prediction with higher accuracy.
[0042] When the transform decoder 407 outputs a prediction result including coordinates and vectors of a point group surrounding a specific region in an image, the correct answer data indicating the region in the image included in the training data includes coordinates and vectors of the point group surrounding the specific region. For example, the prediction result r for a polygon having n vertices is p is (x1, y1, Δx 12 , Δy 12 , x2, y2, , x n , y n , Δxn1 , Δy n1 ), where (x1, y1) are the coordinates of the first point in the point cloud, and (Δx 12 , Δy 12 ) represents the vector from the first point to the second point. g has a similar format, (x1, y1, Δx 12 , Δy 12 , x2, y2, , x n , y n , Δx n1 , Δy n1 ) is included.
[0043] To train one or more machine learning models described in FIG. 4, the learning processing unit 304 calculates a loss function value based on the difference between a prediction result including coordinates and vectors of a point cloud surrounding a specific region and the correct data for the prediction result. The loss function includes, for example, a loss based on the sum of differences in coordinates of the point cloud surrounding a specific region and a loss based on the sum of dissimilarities between vectors (e.g., 1-vector similarity). The dissimilarity between vectors can be obtained, for example, by calculating 1-cosine similarity. The learning processing unit 304 changes parameters of one or more machine learning models to reduce the loss function.
[0044] The learning processing unit 304 can use a loss function based on the optimal transportation cost for the point cloud surrounding a specific area in the prediction result and the point cloud surrounding a specific area in the correct data. The optimal transportation algorithm is an algorithm that calculates a transportation plan for moving a vertex set using the minimum transportation cost. In other words, it can use the cost required to minimize the transportation cost for moving the vertex set in the prediction result to the vertex set in the correct data. By using the optimal transportation cost as the loss function, the value of the loss function can be calculated using the difference in the positions of the corresponding points with the smallest transportation cost. In other words, when the prediction result and the correct data are similar n-gons, the problem of large losses due to differences in the order of the vertices can be resolved even if the vertex positions are similar. For example, the Sinkhorn algorithm can be used as the optimal transportation algorithm.
[0045] Note that the prediction result of machine learning in this embodiment is not limited to the above example and may include further information. For example, FIG. 5 shows another example of the prediction result in this embodiment. The prediction result shown in FIG. 5 includes the coordinates of a center point 501 (target position) within the specific region in addition to the coordinates (412-1, . . . , 412-n) of a point cloud surrounding a specific region and vectors (413-1, . . . , 413-n) indicating the next point of each of the multiple points. The coordinates of the center point 501 within the specific region may be represented by a vector 502 from the position of a target 503 in the image to the coordinates of the center point 501.
[0046] By predicting the center point 501 in the specific area separately from the coordinates of the point cloud and adding a constraint based on the center position, it is possible to output the area of the prediction result around the area that can be the target position. In addition, by using a vector from the target position, it is possible to take into account the positional relationship with the target, and it is possible to reduce the number of areas that are not suitable as a specified location (for example, areas corresponding to places where a mobile object cannot stop) being predicted as the result.
[0047] When the machine learning model (transformer decoder 407) used in the target region prediction unit 303 outputs a prediction result that includes the center point 501, the learning processing unit 304 trains one or more machine learning models using a loss function that takes the center point 501 into consideration. Specifically, the learning processing unit 304 trains one or more machine learning models using a single loss function that is based on the difference between a prediction result that includes the coordinates of a point cloud surrounding a specific region, a vector indicating the next point of each point in the point cloud, and the coordinates of the center point within the specific region, and the correct answer data for the prediction result. In this case, the machine learning model can be optimized while reducing processing costs through the relatively simple calculation of calculating one loss function.
[0048] The learning processing unit 304 may perform multi-task learning to predict the specific area and the target position as separate tasks. Specifically, the learning processing unit 304 calculates a first loss function value based on the difference between a first portion of the prediction result, including the coordinates of a point cloud surrounding the specific area and a vector indicating the next point of each point in the point cloud, and the correct answer for the first portion of the prediction result in the correct answer data. The learning processing unit 304 also calculates a second loss function value based on the difference between a second portion of the prediction result, including the target position (the coordinates of the center point within the specific area), and the correct answer for the second portion of the prediction result in the correct answer data. The learning processing unit 304 trains one or more machine learning models using the values of the first loss function and the second loss function. This allows the multi-task learning to improve the prediction accuracy of the specific area and the prediction accuracy of the target position.
[0049] Next, a series of operations for training a machine learning model used in the region prediction process will be described with reference to Fig. 6. This process is realized by the control unit 130 expanding a program stored in the storage device 135 into the memory device of the control unit 130 and executing it. If the control unit 130 does not include the learning processing unit 304, the following process may be realized, for example, by one or more processors in an information processing server separate from the mobile object 100 executing the program. In this case, the information processing server realizes the operations of the instruction acquisition unit 301, image information acquisition unit 302, target region prediction unit 303, and learning processing unit 304 by executing the program in one or more processors.
[0050] In S601, for example, the instruction acquisition unit 301 and the image information acquisition unit 302 acquire the linguistic information of the training data (i.e., information including location designation expressed in natural language) and the image of the training data, respectively. Also, the learning processing unit 304 acquires the correct answer data corresponding to the training data (e.g., coordinates and vectors of a point cloud surrounding a specific area in the image).
[0051] In S602, the target area prediction unit 303 extracts image features and language features using the machine learning model (image feature extraction unit 403, language feature extraction unit 404, feature fusion unit 405) as described above, and fuses the image features and language features.
[0052] In S603, the target region prediction unit 303 predicts a specific region in the image that indicates a location specified by linguistic information using the fused feature amount by the machine learning model (transformer encoder 406, transform decoder 407) as described above. The machine learning model outputs a prediction result using machine learning parameters in the learning stage. At this time, the target region prediction unit 303 outputs, for example, a prediction result r for a polygon having n vertices. p (x1, y1, Δx 12 , Δy 12 , x2, y2, , x n , y n , Δx n1 , Δy n1 ) is output.
[0053] In S604, the learning processing unit 304 calculates the loss based on the coordinates of the point cloud surrounding the specific region in the prediction result and the correct answer data, as described above. g (x1, y1, Δx 12 , Δy 12 , x2, y2, , x n , y n , Δx n1 , Δy n1 ) is used.
[0054] In S605, the learning processing unit 304 calculates a loss based on the vector from each point group to the next point in the prediction result and the correct data, as described above. In S606, the learning processing unit 304 calculates the value of a loss function based on both losses, as described above. For example, the learning processing unit 304 calculates each of the losses as a separate loss function through multi-task learning.
[0055] In S607, the learning processing unit 304 determines whether processing for a group of data among the training data has been completed. If the learning processing unit 304 determines that processing for the group of data has not been completed, the process returns to S601 and the calculation of the loss function value using other data is repeated. If the learning processing unit 304 determines that processing for the group of data has been completed, the process proceeds to S608.
[0056] In S608, the learning processing unit 304 determines whether an optimization termination condition is met. The optimization processing condition may be any condition, but may include a predetermined number of iterations being repeated, the value of the loss function being reduced to a predetermined value or less, etc. If the learning processing unit 304 determines that the optimization termination condition is met, it ends this series of operations; if not, it proceeds to S609.
[0057] In S609, the learning processing unit 304 changes the machine learning parameters so that the value of the loss function decreases (for example, based on the calculation result of backpropagation). The learning processing unit 304 then returns the process to S601.
[0058] In this way, it is possible to generate machine learning models that improve the accuracy of predictions when predicting specific regions in an image using the point cloud that surrounds that region.
[0059] <Other series of operations for training a model used in region prediction processing> Next, another series of operations for training a machine learning model used in the region prediction process will be described with reference to Fig. 7. Note that this process is realized by the control unit 130 expanding a program stored in the storage device 135 into the memory device of the control unit 130 and executing it. Also, the information processing server may realize the operations of the instruction acquisition unit 301, image information acquisition unit 302, target region prediction unit 303, and learning processing unit 304 by executing the program with one or more processors. Note that the series of operations shown in Fig. 7 differs from the example of Fig. 6 in that an optimal transportation cost is used as the cost function.
[0060] The control unit 130 executes the processes of S601 to S603 in the same manner as the processes shown in FIG. 6, and outputs the prediction result obtained by predicting a specific region in the image using the machine learning model.
[0061] In S701, the learning processing unit 304 calculates the optimal transportation cost for a plurality of points (point cloud) surrounding a specific area in the prediction result and a plurality of points (point cloud) in the correct answer data, as described above.
[0062] In S702, the learning processing unit 304 applies the calculated optimal transportation cost as the loss of the loss function. Note that, as described above, the learning processing unit 304 may use the optimal transportation cost as a loss based on the coordinates of the point cloud surrounding a specific area in the prediction result and the ground truth data, and may separately calculate a loss based on the vector from each point cloud to the next point in the prediction result and the vector in the ground truth data.
[0063] Thereafter, the learning processing unit 304 executes the processes of S607 to S609, and if it determines in S608 that the optimization termination condition is satisfied, the learning processing unit 304 terminates this series of operations.
[0064] In this way, it is possible to generate a machine learning model that improves prediction accuracy when predicting a specific region using a point cloud surrounding that region in an image.
[0065] Next, a series of operations in the inference stage of the machine learning model will be described with reference to Fig. 8. This process is realized by the control unit 130 expanding a program stored in the storage device 135 into the memory device of the control unit 130 and executing it.
[0066] In S801, for example, the instruction acquisition unit 301 acquires, as language information (i.e., information including a location specification expressed in natural language), a user instruction input via the operation unit 131 or the voice input device 133. In addition, the image information acquisition unit 302 acquires an image from the detection unit 114.
[0067] In S802, the target area prediction unit 303 extracts image features and language features using the machine learning model (image feature extraction unit 403, language feature extraction unit 404, feature fusion unit 405) as described above, and fuses the image features and language features.
[0068] In S803, the target area prediction unit 303 predicts a specific area in the image that indicates a location specified by linguistic information using the fused feature amount by the machine learning model (transformer encoder 406, transform decoder 407) as described above. At this time, the target area prediction unit 303 predicts a prediction result r for a polygon having n vertices, for example. p (x1, y1, Δx 12 , Δy 12 , x2, y2, , x n , y n , Δx n1 , Δy n1 ) After outputting the prediction result, the control unit 130 ends this series of operations.
[0069] As described above, in the above-described embodiment, one or more machine learning models are used to extract features from input information and predict a specific region in an image based on the extracted features. In this case, the one or more machine learning models output a prediction result indicating the specific region, including the coordinates of a point cloud surrounding the specific region and vectors indicating the next point of each point in the point cloud. To train such a machine learning model, a loss function is used based on the difference between the prediction result, including the coordinates of a point cloud surrounding the specific region and vectors indicating the next point of each point in the point cloud, and the ground truth data for the prediction result. This makes it possible to improve prediction accuracy when predicting a specific region in an image using a point cloud surrounding the specific region.
[0070] The above-described machine learning model may be executed by various types of information processing devices. For example, the information processing device may be the mobile object 100, or may be configured to be incorporated inside the mobile object 100 (i.e., the control unit 130). The information processing device may also be an information processing server that acquires images and sounds captured by the mobile object 100 and executes the above-described machine learning model.
[0071] <Summary of the embodiment> The above-described embodiments include an information processing device, an information processing method, a generation method, a learning model, a program, and a storage medium shown in the following items.
[0072] (Item 1) an acquisition means for acquiring an image as input information; a prediction means configured with one or more machine learning models that extracts features from the input information and predicts a specific region within the image based on the extracted features; processing means; the prediction means outputs a prediction result indicating the specific area, including coordinates of a plurality of points surrounding the specific area and information indicating a point next to each of the plurality of points; The information processing device is characterized in that the processing means trains the one or more machine learning models using a loss function based on the difference between the prediction result, which includes coordinates of multiple points surrounding the specific area and information indicating the next point for each of the multiple points, and ground truth data for the prediction result.
[0073] According to this embodiment, it is possible to provide an information processing device that improves prediction accuracy when predicting a specific region in an image using a point cloud surrounding the specific region. More specifically, in predicting the region, not only the position of the point cloud but also the relationship between the points in the point cloud is predicted, thereby constraining the relationship between the point clouds to be appropriate in the region prediction, and generating a machine learning model that predicts the region with higher accuracy.
[0074] (Item 2) the acquiring means further acquires, as the input information, language information including a location specification expressed in a natural language; 2. The information processing device according to item 1, wherein the prediction means predicts a target area in the image corresponding to the location specification as the specific area based on image features extracted from the image and language features extracted from the language information.
[0075] According to this embodiment, it is possible to improve the prediction accuracy when predicting a specific region in an image specified by language information.
[0076] (Item 3) Item 1. The information processing device according to item 1, characterized in that the loss function is based on an optimal transportation cost for the plurality of points surrounding the specific area in the prediction result and the plurality of points surrounding the specific area in the ground truth data.
[0077] According to this embodiment, it is possible to solve the problem that the loss increases due to the difference in the order of the vertices even though the positions of the vertices in the prediction result and the correct answer data are similar to each other.
[0078] (Item 4) Item 1. The information processing device according to item 1, characterized in that the loss function includes a loss based on the coordinates of a plurality of points surrounding the specific region and a loss based on the similarity between vectors indicating the next point of each of the plurality of points.
[0079] According to this embodiment, the loss of both the coordinates and vectors of the point cloud surrounding a specific region in an image can be calculated using a simple calculation method.
[0080] (Item 5) Item 1. The information processing device according to item 1, characterized in that the processing means outputs the prediction result including coordinates of a plurality of points surrounding the specific area, information indicating the next point of each of the plurality of points, and coordinates of a center point within the specific area.
[0081] According to this embodiment, by adding a constraint based on the central position, it is possible to output the region of the prediction result around the region that can be the target position.
[0082] (Item 6) 6. The information processing device according to item 5, wherein the processing means trains the one or more machine learning models using one loss function based on the difference between the prediction result, which includes coordinates of a plurality of points surrounding the specific region, information indicating the next point of each of the plurality of points, and coordinates of a center point within the specific region, and ground truth data for the prediction result.
[0083] According to this embodiment, the machine learning model can be optimized by suppressing processing costs through relatively simple calculations for calculating one loss function.
[0084] (Item 7) Item 6. The information processing device according to item 5, wherein the processing means trains the one or more machine learning models using a first portion of the prediction result including coordinates of a plurality of points surrounding the specific region and information indicating the next point of each of the plurality of points, a first loss function based on a difference between the correct answer for the first portion of the prediction result in the correct answer data, and a second portion of the prediction result including coordinates of a center point within the specific region and a second loss function based on a difference between the correct answer for the second portion of the prediction result in the correct answer data.
[0085] According to this embodiment, the prediction accuracy of a specific region and the prediction accuracy of a target position can be improved by multitask learning.
[0086] (Item 8) 6. The information processing device according to item 5, wherein the coordinates of the center point within the specific area are represented by vector information from the position of a target within the image to the coordinates of the center point within the specific area.
[0087] According to this embodiment, the positional relationship with the target can be taken into consideration, and it is possible to reduce the number of areas that are not suitable as specified locations (for example, areas that correspond to places where a moving body cannot stop) being predicted as results.
[0088] (Item 9) 2. The information processing device according to item 1, wherein the information indicating the next point of each of the plurality of points is expressed as a vector from each of the plurality of points to the next point.
[0089] According to this embodiment, the information representing the orientation of the point cloud can be described in a simple manner, and the calculation of the loss function can be facilitated.
[0090] (Item 10) 3. The information processing device according to item 2, wherein the prediction means generates fusion features by fusing the image features and the language features based on the input information, and predicts a specific region within the image based on the fusion features.
[0091] According to this embodiment, it is possible to predict an area in which image features and language features are appropriately associated with each other.
[0092] (Item 11) Item 11. The information processing device according to item 10, wherein the one or more machine learning models include a recursive machine learning model that receives the generated fusion features as input and outputs the prediction result indicating the specific region including coordinates of a plurality of points surrounding the specific region and information indicating a point next to each of the plurality of points.
[0093] According to this embodiment, when a recursive machine learning model is used, it is possible to speed up processing by using a model that is relatively smaller in scale than a transformer, and it is possible to generate a trajectory that takes into account the context of multiple points.
[0094] (Item 12) Item 11. The information processing device according to item 10, wherein the one or more machine learning models include a transform decoder that receives the generated fusion features as input and outputs the prediction result indicating the specific region including coordinates of a plurality of points surrounding the specific region and information indicating a point next to each of the plurality of points.
[0095] According to this embodiment, by using a transformer as a decoder, it is possible to output highly accurate prediction results.
[0096] (Item 13) Item 11. The information processing device according to item 10, wherein the one or more machine learning models include an encoder that receives the generated fusion features as input and encodes the features, and a decoder that receives the encoded features as input and outputs the prediction result indicating the specific region including coordinates of a plurality of points surrounding the specific region and information indicating a point next to each of the plurality of points.
[0097] According to this embodiment, it is possible to further extract features that are effective for prediction from the fusion features, and to output highly accurate prediction results based on the encoded effective features.
[0098] (Item 14) Item 14. The information processing device according to item 13, wherein the encoder is a transformer encoder and the decoder is a transformer decoder.
[0099] According to this embodiment, by using a transformer in both the encoder and the decoder, it is possible to extract effective features from fused features with high accuracy, and to output highly accurate prediction results from the encoded effective features.
[0100] (Item 15) An information processing method executed in an information processing device, an acquisition step of acquiring an image as input information; a prediction step of extracting features from the input information acquired in the acquisition step using one or more machine learning models, and predicting a specific region within the image based on the extracted features; and a processing step, In the prediction step, the one or more machine learning models output a prediction result indicating the specific region, the prediction result including coordinates of a plurality of points surrounding the specific region and information indicating a next point for each of the plurality of points; In the processing step, the one or more machine learning models are trained using a loss function based on the difference between the prediction result, which includes coordinates of multiple points surrounding the specific area and information indicating the next point for each of the multiple points, and ground truth data for the prediction result.
[0101] According to this embodiment, it is possible to provide an information processing method that improves prediction accuracy when predicting a specific area in an image using a point group surrounding the specific area.
[0102] (Item 16) A method for generating one or more machine learning models executed in an information processing device, comprising: an acquisition step of acquiring an image as input information; a prediction step of extracting features from the input information acquired in the acquisition step using one or more machine learning models, and predicting a specific region within the image based on the extracted features; and generating the one or more machine learning models; In the prediction step, the one or more machine learning models output a prediction result indicating the specific region, the prediction result including coordinates of a plurality of points surrounding the specific region and information indicating a next point for each of the plurality of points; The processing step generates the one or more machine learning models by training the prediction results, which include coordinates of multiple points surrounding the specific area and information indicating the next point for each of the multiple points, using a loss function based on the difference between the prediction results and correct data.
[0103] According to this embodiment, it is possible to generate a machine learning model that improves prediction accuracy when predicting a specific region using a point cloud surrounding the specific region in an image.
[0104] (Item 17) A machine learning model executed in an information processing device, a prediction means for extracting a feature amount from an image input as input information and predicting a specific region in the image based on the extracted feature amount; the prediction means outputs a prediction result indicating the specific area, including coordinates of a plurality of points surrounding the specific area and information indicating a point next to each of the plurality of points; The machine learning model is characterized in that it is trained using a loss function based on the difference between the prediction result, which includes coordinates of multiple points surrounding the specific area and information indicating the next point for each of the multiple points, and correct data for the prediction result.
[0105] According to this embodiment, it is possible to improve the prediction accuracy when predicting a specific region in an image using a point group surrounding the specific region.
[0106] (Item 18) 15. A program that causes a computer to function as each of the means of the information processing device according to any one of items 1 to 14.
[0107] According to this embodiment, it is possible to provide a program for generating a machine learning model that improves prediction accuracy when predicting a specific area using a point cloud surrounding the specific area in an image.
[0108] (Item 19) A storage medium that stores a program that causes a computer to function as each of the means of the information processing device described in any one of items 1 to 14.
[0109] According to this embodiment, a storage medium can be provided that generates a machine learning model that improves prediction accuracy when predicting a specific area using a point cloud surrounding the specific area in an image.
[0110] The invention is not limited to the above-described embodiment, and various modifications and variations are possible within the scope of the gist of the invention. [Explanation of symbols]
[0111] 100... moving body, 130... control unit, 302... image information acquisition unit, 303... target area prediction unit, 304... learning processing unit
Claims
1. an acquisition means for acquiring an image as input information; a prediction means configured with one or more machine learning models that extracts features from the input information and predicts a specific region within the image based on the extracted features; processing means; the prediction means outputs a prediction result indicating the specific area, including coordinates of a plurality of points surrounding the specific area and information indicating a point next to each of the plurality of points; The information processing device is characterized in that the processing means trains the one or more machine learning models using a loss function based on the difference between the prediction result, which includes coordinates of multiple points surrounding the specific area and information indicating the next point for each of the multiple points, and correct data for the prediction result.
2. the acquiring means further acquires, as the input information, language information including a location specification expressed in a natural language; 2. The information processing device according to claim 1, wherein the prediction means predicts a target area in the image corresponding to the location specification as the specific area based on image features extracted from the image and language features extracted from the language information.
3. 2. The information processing device according to claim 1, wherein the loss function is based on an optimal transportation cost for the plurality of points surrounding the specific region in the prediction result and the plurality of points surrounding the specific region in the ground truth data.
4. 2. The information processing device according to claim 1, wherein the loss function includes a loss based on coordinates of a plurality of points surrounding the specific region, and a loss based on a similarity between vectors indicating a next point of each of the plurality of points.
5. 2. The information processing device according to claim 1, wherein the processing means outputs the prediction result including coordinates of a plurality of points surrounding the specific area, information indicating a next point for each of the plurality of points, and coordinates of a center point within the specific area.
6. 6. The information processing device according to claim 5, wherein the processing means trains the one or more machine learning models using a loss function based on a difference between the prediction result, which includes coordinates of a plurality of points surrounding the specific region, information indicating a next point for each of the plurality of points, and coordinates of a center point within the specific region, and ground truth data for the prediction result.
7. 6. The information processing device according to claim 5, wherein the processing means trains the one or more machine learning models using a first portion of the prediction result including coordinates of a plurality of points surrounding the specific region and information indicating a next point for each of the plurality of points, a first loss function based on a difference between a correct answer for the first portion of the prediction result in correct answer data, and a second portion of the prediction result including coordinates of a center point within the specific region and a second loss function based on a difference between a correct answer for the second portion of the prediction result in correct answer data.
8. 6. The information processing apparatus according to claim 5, wherein the coordinates of the center point within the specific area are expressed by information of a vector from the position of a target within the image to the coordinates of the center point within the specific area.
9. 2. The information processing apparatus according to claim 1, wherein the information indicating the next point of each of the plurality of points is expressed as a vector from each of the plurality of points to the next point.
10. 3. The information processing device according to claim 2, wherein the prediction means generates fusion features by fusing the image features and the language features based on the input information, and predicts a specific region in the image based on the fusion features.
11. 11. The information processing device according to claim 10, wherein the one or more machine learning models include a recursive machine learning model that receives the generated fusion features as input and outputs the prediction result indicating the specific region including coordinates of a plurality of points surrounding the specific region and information indicating a point next to each of the plurality of points.
12. 11. The information processing device according to claim 10, wherein the one or more machine learning models include a transform decoder that receives the generated fusion features as input and outputs the prediction result indicating the specific region including coordinates of a plurality of points surrounding the specific region and information indicating a point next to each of the plurality of points.
13. 11. The information processing device according to claim 10, wherein the one or more machine learning models include: an encoder that receives the generated fusion features as input and encodes the features; and a decoder that receives the encoded features as input and outputs the prediction result indicating the specific region including coordinates of a plurality of points surrounding the specific region and information indicating a point next to each of the plurality of points.
14. 14. The information processing apparatus according to claim 13, wherein the encoder is a transformer encoder, and the decoder is a transformer decoder.
15. An information processing method executed in an information processing device, an acquisition step of acquiring an image as input information; a prediction step of extracting features from the input information acquired in the acquisition step using one or more machine learning models, and predicting a specific region within the image based on the extracted features; and a processing step, In the prediction step, the one or more machine learning models output a prediction result indicating the specific area, the prediction result including coordinates of a plurality of points surrounding the specific area and information indicating a point next to each of the plurality of points; In the processing step, the one or more machine learning models are trained using a loss function based on the difference between the prediction result, which includes the coordinates of multiple points surrounding the specific area and information indicating the next point for each of the multiple points, and correct data for the prediction result.
16. A method for generating one or more machine learning models executed in an information processing device, comprising: an acquisition step of acquiring an image as input information; a prediction step of extracting features from the input information acquired in the acquisition step using one or more machine learning models, and predicting a specific region within the image based on the extracted features; and generating the one or more machine learning models; In the prediction step, the one or more machine learning models output a prediction result indicating the specific area, the prediction result including coordinates of a plurality of points surrounding the specific area and information indicating a point next to each of the plurality of points; The processing step generates the one or more machine learning models by training the prediction results, which include coordinates of multiple points surrounding the specific area and information indicating the next point for each of the multiple points, using a loss function based on the difference between the prediction results and correct data.
17. A machine learning model executed in an information processing device, a prediction means for extracting a feature amount from an image input as input information and predicting a specific region in the image based on the extracted feature amount; the prediction means outputs a prediction result indicating the specific area, including coordinates of a plurality of points surrounding the specific area and information indicating a point next to each of the plurality of points; The machine learning model is characterized in that it is trained using a loss function based on the difference between the prediction result, which includes the coordinates of multiple points surrounding the specific area and information indicating the next point for each of the multiple points, and correct data for the prediction result.
18. A program that causes a computer to function as each of the means of the information processing device according to any one of claims 1 to 14.
19. A storage medium for storing a program that causes a computer to function as each of the means of the information processing device according to any one of claims 1 to 14.
Citation Information
Cited By
Information processing device, control method for information processing device, and control program for information processing device
JP7796291B1