Model generation device, model generation method, and program
The model generation device improves object recognition accuracy in sparse measurement data by integrating situation information into the prediction model, addressing the limitations of existing techniques.
Patent Information
- Application Number
- JP2023205005
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-17
AI Technical Summary
Existing object recognition techniques using point cloud data from sparse measurements suffer from decreased recognition accuracy.
A model generation device and method that generates object recognition information from measurement data, adds situation information to create prediction information, and uses machine learning to generate a model that improves object recognition accuracy in sparse data scenarios.
The proposed solution enhances object recognition accuracy by incorporating situation information into the prediction model, effectively addressing the limitations of sparse point cloud data.
Smart Images

Figure 2025090044000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a model generation device, a model generation method, and a program.
Background Art
[0002] As described in Patent Document 1, object recognition is performed using measurement data obtained by measuring a space. Specifically, Patent Document 1 describes recognizing the position of a package using distance data obtained by measuring inside a warehouse.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the technique described in Patent Document 1 above, there is a case where the point cloud data, which is measurement data, is sparse, and in such a case, there arises a problem that the recognition accuracy of an object decreases.
[0005] Therefore, an object of the present disclosure is to solve the above-described problem that the recognition accuracy decreases when recognizing an object from measurement data obtained by measuring a space.
Means for Solving the Problems
[0006] A model generation device according to one aspect of the present disclosure a recognition unit that generates object recognition information representing information of an object recognized for each frame from measurement data including a plurality of frames obtained by measuring a space with a sensor; a prediction information generation unit that generates prediction information in which situation information representing a situation at the time of measuring the space is added to the object recognition information for each frame; A model generation unit that takes as input a plurality of pieces of the prediction information corresponding to the plurality of the frames and generates, by performing machine learning using the input plurality of pieces of the prediction information, the output object recognition result in the space, and the correct data of the object recognition result; is provided with and has the following configuration. Also, a model generation method according to an aspect of the present disclosure generates object recognition information representing information on an object recognized for each frame from measurement data including a plurality of frames obtained by measuring a space with a sensor, generates prediction information in which situation information representing a situation at the time of measuring the space is added to the object recognition information for each frame, and generates, by performing machine learning using the input plurality of pieces of the prediction information corresponding to the plurality of the frames, the output object recognition result in the space, and the correct data of the object recognition result, a model that takes as input the plurality of pieces of the prediction information corresponding to the plurality of the frames and outputs the object recognition result in the space. and has the following configuration. Also, a program according to an aspect of the present disclosure generates object recognition information representing information on an object recognized for each frame from measurement data including a plurality of frames obtained by measuring a space with a sensor, generates prediction information in which situation information representing a situation at the time of measuring the space is added to the object recognition information for each frame, and generates, by performing machine learning using the input plurality of pieces of the prediction information corresponding to the plurality of the frames, the output object recognition result in the space, and the correct data of the object recognition result, a model that takes as input the plurality of pieces of the prediction information corresponding to the plurality of the frames and outputs the object recognition result in the space. causes a computer to execute the processing and has the following configuration.
Advantages of the Invention
[0007] With the present disclosure configured as described above, it is possible to improve the recognition accuracy when recognizing an object from measurement data obtained by measuring a space.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Mode for Carrying Out the Invention
[0009] <First Embodiment> The first embodiment of the present disclosure will be described with reference to the drawings. Note that the drawings may be relevant to any of the embodiments.
[0010] [Configuration] The model generation device 10 in this embodiment generates a model for recognizing an object from measurement data obtained by measuring the space. In this embodiment, as an example, a model for recognizing objects such as packages, heavy machinery, and people from measurement data obtained by measuring the space in a warehouse will be generated. However, the objects to be recognized are not limited to packages in the warehouse or the like, and any object may be used. Also, the space is not limited to the inside of the warehouse, and it may be any location.
[0011] The model generation device 10 is composed of one or more information processing devices including an arithmetic device and a storage device. And, as shown in FIG. 1, the model generation device 10 includes a frame sequence generation unit 11, an object recognition unit 12, a prediction information generation unit 13, and a model generation unit 14. Each function of the frame sequence generation unit 11, the object recognition unit 12, the prediction information generation unit 13, and the model generation unit 14 can be realized by the arithmetic device executing a program for realizing each function stored in the storage device. Also, the model generation device 10 includes a learning data storage unit 16 composed of a storage device. Hereinafter, each configuration will be described in detail.
[0012] The learning data storage unit 16 stores learning data used for generating the above-described model. The learning data consists of measurement data obtained by measuring the inside of the warehouse and correct answer data which is information on the objects existing in the warehouse. The measurement data is, for example, measurement data obtained by measuring the inside of the warehouse with a predetermined sensor, and in particular, in this embodiment, it is distance data composed of point cloud data representing the distance to an object at each position measured by 3D LiDAR (Light Detection And Ranging). However, the measurement data is not limited to distance data composed of a distance image measured by 3D LiDAR, and any distance data measured by any sensor may be used.
[0013] And the measurement data, which is the training data, is composed of a plurality of frames measured at the same location at the same time. That is, the measurement data consists of a plurality of frames F obtained by measuring the same location multiple times along the time series. For example, the measurement data consists of a plurality of frames obtained at intervals of 10 frames or 5 frames per second.
[0014] Also, the correct answer data, which is the training data, consists of information for identifying the objects actually existing in the warehouse, that is, the objects existing in the measurement data. For example, the correct answer data consists of vector information and includes information such as the presence or absence of an object, the type of the object, the position of the object, the size of the object, and the posture of the object. As an example, the presence or absence of an object (obj) is represented by "1 (present), 0 (absent)", and the type of the object (cls) is represented by "package, heavy machinery, human", etc. Also, the position of the object (x, y, z) is represented by the three-dimensional coordinates where the reference point of the object is located in the space inside the warehouse, and the size of the object (dx, dy, dz) is represented by the three-dimensional length representing the size of the object. Also, the posture of the object (θ) is represented by the inclination of the object with respect to the reference direction. Thus, the correct answer data includes the presence or absence of an object and the information of a rectangle (cuboid shape) specified by the position, size, and posture of the object. However, the correct answer data is not limited to including the above-mentioned information, and may be a part of it, or may further include other information.
[0015] The frame sequence generation unit 11 extracts the measurement data composed of a plurality of frames F measured at the same location from the training data and generates it as the object of learning. For example, the frame sequence generation unit 11 extracts the measurement data consisting of 10 frames measured at the same location for 1 second. Note that FIG. 2 shows an example of the frame F of the measurement data. Note that the frame sequence generation unit 11 is not limited to extracting and generating the measurement data with the above-mentioned number of frames as the object of learning, and any number of frames may be extracted. Also, the frame sequence generation unit 11 is not necessarily limited to extracting a plurality of consecutive frames, and a plurality of frames with intervals may be extracted.
[0016] The object recognition unit 12 (recognition unit) performs a process of recognizing an object from a plurality of frames F generated as learning targets as described above, and generates object recognition information. At this time, the object recognition unit 12 generates object recognition information representing the result of recognizing an object for each frame F. In the present embodiment, the object recognition unit 12 inputs distance data, which is point cloud data composed of one frame F, to an object recognition model Ma that has been machine-learned in advance, and thereby obtains an object prediction vector Va output from the object recognition model Ma, so as to generate object recognition information for each frame F. Here, the object prediction vector Va, which is object recognition information, consists of, for example, the presence or absence of an object, the type of the object, the position of the object, the size of the object, and the posture of the object, similar to the above-described correct answer data. In this way, the object prediction vector Va includes the presence or absence of an object and information on a rectangle (cuboid shape) specified by the position, size, and posture of the object. Note that the object recognition unit 12 is not necessarily limited to generating all of the above-described object recognition information, and may be a part of it, or may generate object recognition information including other information.
[0017] Note that the object recognition model Ma is generated by machine-learning learning data consisting of a set of distance data, which is point cloud data composed of a prepared frame in advance, and an object prediction vector, which is object recognition information of an object existing in such a frame. However, the object recognition unit 12 is not necessarily limited to generating the object prediction vector Va, which is object recognition information, using the object recognition model Ma, and may generate the object prediction vector Va, which is object recognition information, by a method such as performing an arithmetic process for object recognition prepared in advance using the distance data in the frame.
[0018] The prediction information generation unit 13 generates a prediction profile (prediction information) by adding environmental information (situation information) representing the situation at the time of measurement of the measured space to the above-described object recognition information. At this time, similar to the above-described object recognition unit 12, the prediction information generation unit 13 generates a prediction profile for each frame F. Specifically, in the present embodiment, the prediction information generation unit 13 generates a feature vector Vb composed of feature amounts of the distance data as environmental information from the distance data which is point cloud data composed of the frame F. For example, the prediction information generation unit 13 acquires the position of the recognized object included in the object prediction vector Va which is the object recognition information generated by the object recognition unit 12, and generates a feature vector Vb which is a feature amount of the distance data which is point cloud data at the position of the object in the frame F as environmental information. In the example shown in FIG. 2, a feature vector Vb which is a feature amount of the distance data in a partial frame f shown by a white rectangle specified by the position, size, and orientation of the object recognized in the frame F is generated.
[0019] At this time, the prediction information generation unit 13 inputs distance data which is point cloud data composed of the partial frame f to a feature extraction model Mb which has been machine-learned in advance, and thereby acquires a feature vector Vb which is a feature amount output from the feature extraction model Mb, to generate environmental information for each frame F. Note that the feature extraction model Mb is generated by machine-learning distance data which is point cloud data composed of a prepared partial frame. However, the prediction information generation unit 13 is not necessarily limited to generating a feature amount from the partial frame f using the feature extraction model Mb, and a feature amount may be generated by a method such as performing an arithmetic process for feature extraction prepared in advance using the distance data in the partial frame.
[0020] Then, as shown in FIG. 2, the prediction information generation unit 13 combines the feature vector Vb, which is the environmental information generated as described above, and the object prediction vector Va generated by the object recognition unit 12 to generate a prediction profile Vc, which is vector information. At this time, since the prediction information generation unit 13 generates a prediction profile Vc for each frame F, as shown in FIG. 2, a plurality of prediction profiles Vc corresponding to each frame F are generated, and a prediction profile group Vg is generated. That is, as shown in FIG. 3, first, the object recognition unit 12 described above generates an object prediction vector Va that specifies the position of a rectangle and the like for each frame of a frame group Fg composed of a plurality of frames F1, F2,.... Then, the prediction information generation unit 13 adds the feature vector Vb generated for each frame to each object prediction vector Va, and a prediction profile group Vg composed of each prediction profile Vc is generated.
[0021] Here, in addition to the feature vector Vb described above, the prediction information generation unit 13 may generate a prediction profile Vc by adding other environmental information representing the situation at the time of measurement of the measured space to the object prediction vector Va. Specifically, in the present embodiment, the prediction information generation unit 13 acquires the position of the recognized object included in the object prediction vector Va, which is the object recognition information generated by the object recognition unit 12, and uses the distance of the sensor (3D LiDAR) that measures the measurement data for the position of the object in the frame F as environmental information. Then, as shown in FIG. 4, the prediction information generation unit 13 generates a new object prediction vector Va' obtained by adding the distance of the sensor, which is environmental information, to the object prediction vector Va, and generates a prediction profile Vc by adding the feature vector Vb described above to such an object prediction vector Va'.
[0022] In this way, the prediction information generation unit 13 may generate a prediction profile Vc by adding the feature vector Vb and the sensor distance as environmental information to the object prediction vector Va. However, the prediction information generation unit 13 may use only the sensor distance as the environmental information to be added to the object prediction vector Va, or only the feature vector Vb. Further, the prediction information generation unit 13 may obtain other information representing the state of the sensor, not limited to the sensor distance, as environmental information and add it to the object prediction vector Va. For example, the prediction information generation unit 13 may add information representing the environment around the location where the 3D LiDAR, which is a sensor, is installed, particularly information representing the environment that affects the operation of the sensor (e.g., temperature, weather, presence or absence of laser-reflecting objects, etc.) to the object prediction vector Va as the above-described environmental information.
[0023] Here, the prediction information generation unit 13 generates a prediction profile group Vg composed of a plurality of prediction profiles Vc based on the position of the object specified from the object prediction vector Va included in each prediction profile Vc. For example, the prediction information generation unit 13 includes those in which the positions of the objects specified by the object prediction vector Va included in the prediction profile Vc overlap, that is, those in which the rectangles corresponding to the position, size, and orientation of the object, which are the recognition results represented by the object prediction vector Va included in the prediction profile Vc, overlap, in the same prediction profile group Vg. However, at this time, the prediction information generation unit 13 is configured to generate the prediction profile group Vg such that one or less prediction profiles Vc generated from one same frame F are included in one prediction profile group Vg. In other words, when the model generation unit 14 generates a plurality of prediction profiles Vc from one same frame F, the prediction profile group Vg is generated such that these plurality of prediction profiles Vc are not included in one same prediction profile group Vg. As an example, as shown in FIG. 5, it is assumed that two rectangular object prediction vectors V1a, which are prediction results of the position of the object indicated by the dotted line, are generated from one frame F1 included in the frame group Fg. In this case, as shown in FIG. 5, the prediction profile group Vg is generated such that the two rectangular object prediction vectors V1a belong to different groups G1 and G2, that is, each prediction profile generated from each object prediction vector V1a is included only one in each of the prediction profile groups G1 and G2.
[0024] As shown in FIG. 6, the model generation unit 14 generates a prediction model M that takes as input a prediction profile group Vg composed of a set of prediction profiles Vc generated for each frame F, and outputs a prediction result vector V representing the prediction result of object recognition in the space predicted from the prediction profile group Vg. Specifically, the model generation unit 14 sets a loss according to the difference between the prediction result vector V, which is the prediction result output from the prediction model M by inputting the prediction profile group Vg, and the correct answer data corresponding to the frame group Fg that constitutes the learning data that is the generation source of the prediction profile group Vg, and generates the prediction model M by machine learning and adjusting the parameters of the prediction model M that minimize such loss.
[0025] Here, the loss L used for machine learning by the model generation unit 14 will be described with reference to FIG. 7. First, the prediction model M machine-learned by the model generation unit 14 is configured to output, as the prediction result of object recognition, a prediction result vector V including information on the presence or absence of an object, the type of the object, the position of the object, the size of the object, and the posture of the object, as shown in Case 1 of FIG. 7. Then, the model generation unit 14 calculates a loss for each type of information included in the prediction result and the correct answer data as a loss according to the difference between the prediction result and the correct answer data. Specifically, as shown in Case 1 of FIG. 7, an existence loss L E which is a loss according to the difference in the presence or absence of an object, and an object type loss L C which is a loss according to the difference in the type of the object, and a rectangle loss L B which is a loss according to the differences in the position, size, and posture of the object are calculated. At this time, as an example, the existence loss L E is the Cross-entropy Loss of the presence or absence of an object, the object type loss L C is the Cross-entropy Loss of the type of the object, and the rectangle loss L B is the Smooth L1 Loss of the rectangle position, size, and posture information.
[0026] Then, the model generation unit 14 calculates the loss L used for machine learning based on the above three types of losses as shown in the following Equation (1). [Number]
[0027] As shown in Figure 1, when the information of "object present (obj = 1)" is included in the prediction result by the prediction model M, as shown in Case 1 of Figure 7, all three types of losses L mentioned above E , L C , L B are all integrated to calculate the loss L. On the other hand, when the information of "no object (obj = 0)" is included in the prediction result by the prediction model M, as shown in Case 2 of Figure 7, the loss L consisting only of the object presence / absence loss L E among the three types mentioned above is calculated.
[0028] The model generation unit 14 stores the prediction model M generated as described above. As a result, by inputting measurement data consisting of a plurality of frames F that measure a space where the object recognition result is unknown to the prediction model M, an output of a prediction result vector V that is the result of object recognition can be obtained, and object recognition can be performed. At this time, for the prediction model M, as described above, object prediction vectors Va are generated from a plurality of frames F respectively, and environmental information such as a feature vector Vb is added to generate a plurality of prediction profiles Vc, and these groups Vg are input.
[0029] [Operation] Next, the operation of the model generation device 10 described above will be explained. It is assumed that the learning data described above is stored in the model generation device 10 in advance.
[0030] The model generation device 10 acquires measurement data consisting of a plurality of frames F that measure the same location from the learning data (step S1 in Figure 8). For example, the model generation device 10 acquires measurement data consisting of 10 frames that measure the same location for 1 second.
[0031] Subsequently, the model generation device 10 performs a process of recognizing an object from a plurality of frames F and generates object recognition information (step S2 in FIG. 8). At this time, the model generation device 10 generates object recognition information representing the result of recognizing the object for each frame F. For example, the model generation device 10 inputs distance data, which is point cloud data composed of each frame F, to the object recognition model Ma, and thereby obtains an object prediction vector Va output from the object recognition model Ma, so as to generate object recognition information for each frame F.
[0032] Subsequently, the model generation device 10 generates a prediction profile (prediction information) in which environmental information representing the situation at the time of measurement of the measured space is added to the generated object recognition information (step S3 in FIG. 8). At this time, the model generation device 10 generates a prediction profile for each frame F. Specifically, the model generation device 10 generates a feature vector Vb composed of feature amounts of distance data corresponding to the position of the recognized object from the distance data, which is point cloud data composed of the frame F, as environmental information. Then, the model generation device 10 generates a prediction profile Vc, which is vector information obtained by combining the object prediction vector Va and the feature vector Vb, which is environmental information. Note that the model generation device 10 may use the distance of the sensor that measured the measurement data for the position of the object in the frame F as environmental information, and add the distance of such a sensor to the object prediction vector Va to generate the prediction profile Vc.
[0033] In this way, by the model generation device 10 generating the prediction profiles Vc from the plurality of frames F respectively, a prediction profile group Vg composed of the plurality of prediction profiles Vc is generated. At this time, the model generation device 10 generates the prediction profile group Vg according to the positions of the objects specified by the object prediction vectors Va included in the respective prediction profiles Vc generated from the plurality of frames F. For example, the model generation device 10 sets the prediction profiles Vc in which the positions of the objects recognized from each frame F overlap as the same group. On the other hand, when a plurality of objects are recognized from one same frame F, the prediction profiles Vc corresponding to the plurality of objects are not made to belong to the same group.
[0034] Then, the model generation device 10 generates a prediction model M that takes the prediction profile group Vg as an input and outputs a prediction result vector V representing the prediction result of object recognition in the space predicted from the prediction profile group Vg (step S4 in FIG. 8). Specifically, the model generation device 10 sets a loss L according to the difference between the prediction result vector V, which is the prediction result output from the prediction model M by inputting the prediction profile group Vg, and the correct data corresponding to the frame group Fg that constitutes the learning data that is the generation source of the prediction profile group Vg, and generates the prediction model M by performing machine learning to adjust the parameters of the prediction model M such that such loss is minimized.
[0035] Note that, as an example, as shown in Case 1 of FIG. 7, the model generation device 10 calculates an existence loss L according to the difference in the presence or absence of an object E and an object type loss L according to the difference in the type of the object C and a rectangle loss L according to the differences in the position, size, and posture of the object. B Then, when the information of "object exists (obj = 1)" is included in the prediction result by the prediction model M, the model generation device 10 calculates all three types of losses L described above E , L C , L BIntegrate all of them to calculate the loss L. On the other hand, when the information of "no object (obj = 0)" is included in the prediction result by the prediction model M, as shown in Case 2 of FIG. 7, among the above three types, the object presence / absence loss L E Calculate the loss L consisting only of
[0036] As described above, in this embodiment, machine learning is performed with a prediction profile obtained by adding environment information representing the situation at the time of measurement to a plurality of measurement data as an input, and a prediction model is generated. Thereby, it is possible to improve the accuracy of object recognition from the measurement data using the generated prediction model.
[0037] <Second Embodiment> Next, a second embodiment of the present disclosure will be described with reference to the drawings. In this embodiment, a schematic configuration of the model generation device described in the above embodiment is shown. Note that FIGS. 9 to 10 are diagrams for explaining the configuration, and such drawings may be related to any of the embodiments.
[0038] First, with reference to FIG. 9, the hardware configuration of the model generation device 100 will be described. The model generation device 100 is configured by a general information processing device, and as an example, is equipped with the following hardware configuration. ·CPU (Central Processing Unit) 101 (arithmetic unit) ·ROM (Read Only Memory) 102 (storage device) ·RAM (Random Access Memory) 103 (storage device) ·Program group 104 loaded into RAM 103 ·Storage device 105 that stores the program group 104 ·Drive device 106 that reads and writes to the storage medium 110 outside the information processing device ·Communication interface 107 that connects to the communication network 111 outside the information processing device ·Input / output interface 108 that performs input / output of data ·Bus 109 that connects each component
[0039] Note that FIG. 9 shows an example of the hardware configuration of the information processing apparatus which is the model generation apparatus 100, and the hardware configuration of the information processing apparatus is not limited to the case described above. For example, the information processing apparatus may be configured from a part of the above-described configuration such as not having the drive device 106. Further, instead of the above-described CPU, the information processing apparatus may use a GPU (Graphic Processing Unit), a DSP (Digital Signal Processor), an MPU (Micro Processing Unit), an FPU (Floating point number Processing Unit), a PPU (Physics Processing Unit), a TPU (TensorProcessingUnit), a quantum processor, a microcontroller, or a combination thereof.
[0040] Then, the model generation apparatus 100 can construct and equip the recognition unit 121, the prediction information generation unit 122, and the model generation unit 123 shown in FIG. 10 by the CPU 101 acquiring the program group 104 and the CPU 101 executing it. Note that the program group 104 is stored in the storage device 105 or the ROM 102 in advance, for example, and is loaded into the RAM 103 and executed by the CPU 101 as necessary. Further, the program group 104 may be supplied to the CPU 101 via the communication network 111, or may be stored in the storage medium 110 in advance, and the drive device 106 may read the program and supply it to the CPU 101. However, the above-described recognition unit 121, prediction information generation unit 122, and model generation unit 123 may be constructed by a dedicated electronic circuit for realizing such means.
[0041] The recognition unit 121 generates object recognition information representing information of an object recognized for each of the frames from measurement data including a plurality of frames obtained by measuring a space with a sensor. The prediction information generation unit 122 generates prediction information in which situation information representing the situation at the time of measuring the space is added to the object recognition information for each of the frames. The model generation unit 123 generates a model that takes as input a plurality of pieces of the prediction information corresponding to the plurality of frames and outputs an object recognition result in the space, by performing machine learning using the input plurality of pieces of the prediction information, the output object recognition result, and correct data of the object recognition result.
[0042] As configured above, the present disclosure generates a prediction model by performing machine learning using, as input, a prediction profile in which environment information representing the situation at the time of measurement is added to measurement data including a plurality of frames. Thereby, it is possible to improve the accuracy of object recognition from the measurement data using the generated prediction model.
[0043] Note that at least one or more of the functions of the recognition unit 121, the prediction information generation unit 122, and the model generation unit 123 described above may be executed by an information processing apparatus installed and connected at any location on a network, that is, may be executed by so-called cloud computing.
[0044] In addition, the above-described program can be stored using various types of non-transitory computer readable media and supplied to a computer. Non-transitory computer readable media include various types of tangible storage media. Examples of non-transitory computer readable media include magnetic recording media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical recording media (e.g., magneto-optical disks), CD-ROM (Read Only Memory), CD-R, CD-R / W, semiconductor memories (e.g., mask ROM, PROM (Programmable ROM), EPROM (Erasable PROM), flash ROM, RAM (Random Access Memory)). Also, the program may be supplied to the computer by various types of transitory computer readable media. Examples of transitory computer readable media include electrical signals, optical signals, and electromagnetic waves. The transitory computer readable media can supply the program to the computer via wired communication paths such as electric wires and optical fibers, or wireless communication paths.
[0045] As described above, the present disclosure has been described with reference to the above embodiments and the like, but the present disclosure is not limited to the above-described embodiments. Various changes that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. And each of the above-described embodiments can be combined with other embodiments as appropriate.
[0046] <Supplementary Note> Some or all of the above embodiments can also be described as follows. Hereinafter, an outline of the configuration of the model generation device, model generation method, and program in the present disclosure will be described. However, the present disclosure is not limited to the following configuration. (Supplementary Note 1) A recognition unit that generates object recognition information representing information on an object recognized for each frame from measurement data composed of a plurality of frames obtained by measuring a space with a sensor, For each of the frames, a prediction information generation unit that generates prediction information obtained by adding situation information representing the situation at the time of measurement of the space to the object recognition information; A model generation unit that generates a model that takes as input a plurality of pieces of the prediction information corresponding to the plurality of frames and outputs an object recognition result in the space by performing machine learning using the plurality of pieces of the input prediction information, the output object recognition result, and correct answer data of the object recognition result; A model generation device including the above. (Appendix 2) The model generation device according to Appendix 1, wherein the prediction information generation unit generates, for each of the frames, information based on the measurement data as the situation information, and generates the prediction information obtained by adding the situation information to the object recognition information; A model generation device. (Appendix 3) The model generation device according to Appendix 2, wherein the recognition unit generates, for each of the frames, object recognition information including the position of the recognized object; the prediction information generation unit generates, for each of the frames, a feature amount of the measurement data at the position of the recognized object as the situation information, and generates the prediction information obtained by adding the situation information to the object recognition information; A model generation device. (Appendix 4) The model generation device according to Appendix 1, wherein the prediction information generation unit generates, for each of the frames, information representing the state of the sensor in the space as the situation information, and generates the prediction information obtained by adding the situation information to the object recognition information; A model generation device. (Appendix 5) The model generation device according to Appendix 4, wherein the recognition unit generates, for each of the frames, object recognition information including the position of the recognized object; the prediction information generation unit generates, for each of the frames, the distance of the sensor with respect to the position of the recognized object as the situation information, and generates the prediction information obtained by adding the situation information to the object recognition information; Model generation device. (Appendix 6) The model generation device according to Appendix 1, the recognition unit generates the object recognition information including the position of the recognized object for each frame, the model generation unit machine-learns the model using a loss corresponding to the difference between the object recognition result output from the model including at least information on whether the object exists and information representing the position of the object, and the correct data of the object recognition result, Model generation device. (Appendix 7) The model generation device according to Appendix 6, when the object recognition result output from the model includes information indicating that the object does not exist, the model generation unit machine-learns the model using a loss corresponding to the difference between the object recognition result including only information on whether the object exists and the correct data of the object recognition result, Model generation device. (Appendix 8) The model generation device according to Appendix 1, the recognition unit generates the object recognition information including the position of the recognized object for each frame, the prediction information generation unit generates a group of the prediction information generated from the object recognition information based on the position of the recognized object, the model generation unit machine-learns the model using the prediction information belonging to the same group as an input to the model, Model generation device. (Appendix 8-1) The model generation device according to Appendix 8, the prediction information generation unit generates the group so as to include the prediction information generated from one or less pieces of the object recognition information recognized from one frame in one group, Model generation device. (Appendix 9) Generate object recognition information representing information on an object recognized for each frame from measurement data composed of a plurality of frames obtained by measuring a space with a sensor, generate prediction information in which situation information representing the situation at the time of measurement of the space is added to the object recognition information for each frame, generate a model that takes as input a plurality of pieces of the prediction information corresponding to a plurality of the frames and outputs an object recognition result in the space by performing machine learning using the plurality of pieces of the input prediction information, the output object recognition result, and correct answer data of the object recognition result, Model generation method. (Appendix 9-1) The model generation method according to Appendix 9-1, generate information based on the measurement data as the situation information for each frame, and generate the prediction information in which the situation information is added to the object recognition information, Model generation device. (Appendix 9-2) The model generation method according to Appendix 9-1, generate the object recognition information including the position of the recognized object for each frame, generate a feature amount of the measurement data at the position of the recognized object as the situation information for each frame, and generate the prediction information in which the situation information is added to the object recognition information, Model generation method. (Appendix 9-3) The model generation method according to Appendix 9, generate information representing the state of the sensor in the space as the situation information for each frame, and generate the prediction information in which the situation information is added to the object recognition information, Model generation method. (Appendix 9-4) The model generation method according to Appendix 9-3, generate the object recognition information including the position of the recognized object for each frame, generate the distance of the sensor with respect to the position of the recognized object as the situation information for each frame, and generate the prediction information in which the situation information is added to the object recognition information, Model generation method. (Appendix 9-5) A model generation method according to Appendix 9, generating the object recognition information including the position of the recognized object for each frame, using a loss corresponding to the difference between the object recognition result output from the model including at least information on whether the object exists and information representing the position of the object, and the correct answer data of the object recognition result, to perform machine learning on the model, Model generation device. (Appendix 9-6) A model generation method according to Appendix 9-5, when the information indicating that the object does not exist is included in the object recognition result output from the model, using a loss corresponding to the difference between the object recognition result including only the information on whether the object exists and the correct answer data of the object recognition result, to perform machine learning on the model, Model generation device. (Appendix 9-7) A model generation method according to Appendix 9, generating the object recognition information including the position of the recognized object for each frame, generating a group of prediction information generated from the object recognition information based on the recognized position of the object, using the prediction information belonging to the same group as an input to the model to perform machine learning, Model generation method. (Appendix 9-8) A model generation method according to Appendix 9-7, generating the group so that the group includes prediction information generated from one or less pieces of the object recognition information recognized from one frame, Model generation method. (Appendix 10) generating object recognition information representing information on an object recognized for each frame from measurement data including a plurality of frames obtained by measuring a space with a sensor, For each of the frames, prediction information is generated by adding situation information representing the situation at the time of measurement of the space to the object recognition information. A model that takes, as input, a plurality of pieces of the prediction information corresponding to a plurality of the frames and outputs an object recognition result in the space is generated by performing machine learning using the plurality of pieces of the input prediction information, the output object recognition result, and correct answer data of the object recognition result. A program for causing a computer to execute the processing.
Explanation of Signs
[0047] 10 Model generation device 11 Frame sequence generation unit 12 Object recognition unit 13 Prediction information generation unit 14 Model generation unit 16 Learning data storage unit 100 Model generation device 101 CPU 102 ROM 103 RAM 104 Program group 105 Storage device 106 Drive device 107 Communication interface 108 Input / output interface 109 Bus 110 Storage medium 111 Communication network 121 Recognition unit 122 Prediction information generation unit 123 Model generation unit
Claims
1. A recognition unit that generates object recognition information representing information of an object recognized for each frame from measurement data including a plurality of frames obtained by measuring a space with a sensor; A prediction information generation unit that generates prediction information in which situation information representing a situation at the time of measurement of the space is added to the object recognition information for each frame; A model generation unit that generates a model that takes as input a plurality of pieces of the prediction information corresponding to a plurality of the frames and outputs an object recognition result in the space, by performing machine learning using the input plurality of pieces of the prediction information, the output object recognition result, and correct answer data of the object recognition result; A model generation device comprising the above.
2. The model generation device according to claim 1, wherein the prediction information generation unit generates information based on the measurement data as the situation information for each frame, and generates the prediction information in which the situation information is added to the object recognition information; A model generation device.
3. The model generation device according to claim 2, wherein the recognition unit generates the object recognition information including the position of the recognized object for each frame, and the prediction information generation unit generates, as the situation information, a feature amount of the measurement data at the position of the recognized object for each frame, and generates the prediction information in which the situation information is added to the object recognition information; A model generation device.
4. The model generation device according to claim 1, wherein the prediction information generation unit generates, as the situation information, information representing the state of the sensor in the space for each frame, and generates the prediction information in which the situation information is added to the object recognition information; A model generation device.
5. The model generation device according to claim 4, The recognition unit generates the object recognition information including the position of the recognized object for each frame. The prediction information generation unit generates, for each frame, the distance of the sensor with respect to the position of the recognized object as the situation information, and generates the prediction information obtained by adding the situation information to the object recognition information. Model generation device. **Claim 6** The model generation device according to claim 1, The recognition unit generates the object recognition information including the position of the recognized object for each frame. The model generation unit performs machine learning on the model using a loss corresponding to the difference between the object recognition result output from the model including at least information on whether the object exists and information representing the position of the object, and the correct data of the object recognition result. Model generation device. **Claim 7** The model generation device according to claim 6, When the object recognition result output from the model includes information indicating that the object does not exist, the model generation unit performs machine learning on the model using a loss corresponding to the difference between the object recognition result including only information on whether the object exists and the correct data of the object recognition result. Model generation device. **Claim 8** The model generation device according to claim 1, The recognition unit generates the object recognition information including the position of the recognized object for each frame. The prediction information generation unit generates a group of prediction information generated from the object recognition information based on the recognized position of the object. The model generation unit performs machine learning using the prediction information belonging to the same group as the input to the model. Model generation device. **Claim 9** Generate object recognition information representing information on an object recognized for each frame from measurement data composed of a plurality of frames obtained by measuring a space with a sensor, generate prediction information in which situation information representing the situation at the time of measurement of the space is added to the object recognition information for each frame, generate a model that takes as input a plurality of pieces of the prediction information corresponding to the plurality of frames and outputs an object recognition result in the space by performing machine learning using the plurality of pieces of the input prediction information, the output object recognition result, and correct answer data for the object recognition result, Model generation method.
10. Generate object recognition information representing information on an object recognized for each frame from measurement data composed of a plurality of frames obtained by measuring a space with a sensor, generate prediction information in which situation information representing the situation at the time of measurement of the space is added to the object recognition information for each frame, generate a model that takes as input a plurality of pieces of the prediction information corresponding to the plurality of frames and outputs an object recognition result in the space by performing machine learning using the plurality of pieces of the input prediction information, the output object recognition result, and correct answer data for the object recognition result, A program for causing a computer to execute the processing.
Citation Information
Patent Citations
Three-dimensional object recognition system and inventory system using the same
JP2010023950A