Data processor, learning data generation device, data processing method, and learning data generation method
By incorporating a learning model to correct the target position, the data processing device and training data generation device improve recognition accuracy by aligning the position with the requirements of subsequent recognition processes.
Patent Information
- Application Number
- JP2024048205
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-25
- Publication Date
- 2025-10-07
AI Technical Summary
The recognition accuracy of a predetermined recognition process is compromised when the determined position of a recognition target in input data is not suitable for the subsequent processing.
A data processing device and training data generation device that include a detection unit to determine a target position, a correction unit to correct this position using a learning model, and a processing unit to perform the recognition process based on the corrected position, thereby improving accuracy.
The recognition accuracy is enhanced by using a learning model to correct the target position, ensuring suitability for subsequent recognition processes.
Smart Images

Figure 2025147786000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a data processing device, a training data generating device, a data processing method, and a training data generating method. [Background technology]
[0002] 2. Description of the Related Art A technique is known in which the position of a recognition target within input data is determined, and a predetermined recognition process is performed on the recognition target based on the determined position. For example, the tracking device described in Patent Document 1 below detects a target area in which the tracking target appears from input images acquired sequentially at multiple different times, and tracks the same tracking target by associating target areas of the same tracking target with each other. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2023-147846 Summary of the Invention [Problem to be solved by the invention]
[0004] However, if the result of determining the position of the recognition target in the input data is not a position suitable for the subsequent recognition process, there is a risk that the recognition accuracy of the recognition process will decrease. The present invention has been made in consideration of the above problems, and has as its object to improve the recognition accuracy when a predetermined recognition process is performed based on the determination result of the position of a recognition target within input data. [Means for solving the problem]
[0005] A data processing device according to one aspect of the present invention includes a detection unit that outputs a target position, which is the position of a recognition target within input data, a correction unit that outputs a corrected position obtained by correcting the target position based on the input data and the target position, and a processing unit that executes a predetermined recognition process for the recognition target based on at least the corrected position. The correction unit is a learning model that learns the corrected position for the input data and the target position so as to improve the accuracy of the recognition process in the processing unit.
[0006] According to another aspect of the present invention, there is provided a training data generation device that generates training data for a position detector that outputs a target position, which is the position of a recognition target within input data. The training data generation device includes a detection unit that outputs a target position, which is the position of a recognition target within the input data, a correction unit that outputs a corrected position obtained by correcting the target position based on the input data and the target position, and a training data output unit that outputs training data from the corrected position output by the correction unit and the input data. The correction unit is a learning model that learns the corrected position for the input data and the target position based on the input data and the corrected position output by the correction unit so as to improve recognition accuracy when a predetermined recognition process is performed on a recognition target included in the input data. [Effects of the Invention]
[0007] According to the present invention, it is possible to improve the recognition accuracy when performing a predetermined recognition process based on the determination result of the position of the recognition target within input data. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a schematic diagram illustrating an example of the overall configuration of a recognition system including a data processing device according to an embodiment. [Figure 2] 1 is a block diagram illustrating an example of a functional configuration of a data processing apparatus according to an embodiment. [Figure 3] 10(a) to 10(c) are schematic diagrams of a first example of input data, target positions, and correction positions. [Figure 4] 10(a) to 10(c) are schematic diagrams of a second example of input data, target positions, and correction positions. [Figure 5] 10(a) to 10(c) are schematic diagrams of a third example of input data, target positions, and correction positions. [Figure 6] FIG. 1 is a schematic diagram illustrating an example of the overall configuration of a learning device used for learning a correction unit. [Figure 7] FIG. 2 is a block diagram showing an example of the functional configuration of a learning device used for learning of a correction unit. [Figure 8] 1A is a flowchart illustrating an example of a learning method for a correction unit, and FIG. 1B is a flowchart illustrating an example of a data processing method according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the embodiments of the present invention shown below are merely examples of devices and methods for embodying the technical concept of the present invention, and the technical concept of the present invention does not limit the structure, arrangement, etc. of the components to those described below. The technical concept of the present invention can be modified in various ways within the technical scope defined by the claims.
[0010] A data processing device according to an embodiment of the present invention determines the position of a recognition target within input data, and executes a predetermined recognition process for the recognition target based on the determined position. For example, the recognition processing by the data processing device of the embodiment may be processing for recognizing some information about a recognition target based on image data of the recognition target. For example, the recognition processing by the data processing device of the embodiment may be processing for re-identification (ReID) of an object (e.g., a person, a vehicle, etc.) in the image data or processing for estimating attributes of the object. Furthermore, for example, the recognition processing by the data processing device of the embodiment may be processing for estimating the pose of an articulated body (e.g., a person, an animal, an articulated robot, etc.) in the image data.
[0011] Furthermore, for example, the recognition process by the data processing device of the embodiment may be a process of recognizing some information about the recognition target based on voice data obtained by acquiring a sound emitted by the recognition target, such as a voice recognition process of the speech content included in the voice data, or an attribute estimation process of the recognition target. FIG. 1 is a schematic diagram showing an example of the overall configuration of a recognition system 1 including a data processing device 10 according to an embodiment.
[0012] The recognition system 1 comprises a sensor 3 for observing the target of recognition processing by the recognition system 1 (hereinafter referred to as "recognition target 2"), and a data processing device 10 for executing a predetermined recognition processing for the recognition target 2 based on the output of the sensor 3. The data processing device 10 includes an input unit 12, a communication unit 13, a storage unit 14, a control unit 15, and an output unit 16. Of these, the storage unit 14 and the control unit 15 can be realized by a so-called computer, and the input unit 12, the communication unit 13, and the output unit 16 can be realized as peripheral devices of the computer.
[0013] The input unit 12 includes a user interface such as a keyboard, a mouse, etc. that is operated by a user to input data, etc. The input unit 12 is connected to the control unit 15, converts the user's operation into an operation signal, and outputs it to the control unit 15. Input unit 12 may also include a DVD (Digital Versatile Disc) drive and a USB (Universal Serial Bus) interface. Input unit 12 inputs data to control unit 15 as a file and outputs data from control unit 15 as a file.
[0014] The communication unit 13 includes a network interface and the like that transmits and receives data between the data processing device 10 and an external device via wired or wireless communication. The storage unit 14 is a memory device such as a ROM (Read Only Memory) or a RAM (Random Access Memory), and stores various programs and various data. The control unit 15 is a controller that acquires the output of the sensor 3 as input data and executes predetermined recognition processing for the recognition target 2 based on the input data, and is composed of arithmetic devices (processors) such as a CPU (Central Processing Unit), a DSP (Digital Signal Processor), an MCU (Micro Control Unit), and a GPU (Graphics Processing Unit).
[0015] The control unit 15 is connected to the storage unit 14, and operates as various processing units by reading and executing computer programs from the storage unit 14, and stores and reads various data in the storage unit 14. The functions of the data processing device 10 described below are realized by the control unit 15 executing the computer programs stored in the storage unit 14. The control unit 15 is also connected to the input unit 12 to accept operation inputs to the data processing device 10 by the user.
[0016] The control unit 15 is connected to the output unit 16, and outputs the processing results of the control unit 15 to the output unit 16. The control unit 15 is also connected to the communication unit 13, and transmits and receives various data to and from an external device via the communication unit 13. For example, the control unit 15 may transmit the processing results of the control unit 15 to the external device via the communication unit 13, or may acquire learning models and rule-based artificial intelligence parameters used in predetermined recognition processing from an external learning device.
[0017] 2 is a block diagram showing an example of a functional configuration of the data processing device 10 according to the embodiment. The data processing device 10 includes an input data acquisition unit 20, a detection unit 21, a correction unit 22, and a processing unit . The input data acquisition unit 20 acquires input data to be used for recognition processing by the data processing device 10. The input data may be, for example, image data 30 obtained by capturing an image of a recognition target 31 as shown in Fig. 3(a) or Fig. 4(a), or may be audio data 32 as shown in Fig. 5(a).
[0018] For example, the input data acquiring unit 20 may acquire, as input data, output data of the sensor 3. Furthermore, for example, the input data acquiring unit 20 may acquire the input data via the input unit 12 or the communication unit 13. See Fig. 2. The detection unit 21 detects a target position, which is the position of a recognition target within input data, and outputs the detected target position to the correction unit 22. The detection logic used by the detection unit 21 may be, for example, a learning model (such as a neural network) constructed by machine learning, or a rule-based artificial intelligence, or a combination of these.
[0019] 3(b) is an example of the target position PT1 output by the detection unit 21. For example, when the recognition process by the data processing device 10 is a re-identification process or attribute estimation process of the target 31 on the input data (image data 30), the target position PT1 may be a target region representing a region in which the target 31 exists on the image data 30. For example, the target position PT1 may be a region represented by a figure surrounding the target 31 on the image data 30.
[0020] For example, the target position PT1 may be a rectangular area as shown in FIG. T , B B , B L , B R are the upper boundary, lower boundary, left boundary, and right boundary of the target position PT1, respectively. The shape of the target position PT1 is not limited to the rectangle shown in FIG. 3(b), but may be another polygon, circle, or ellipse.
[0021] Furthermore, for example, when the recognition target 31 is a person, the range of the target position PT1 is not limited to the whole body as shown in Fig. 3(b), but may be a range in which parts of the body such as the face, upper body, lower body, arms, legs, etc. The range of the target position PT1 may also be a range in which belongings of the person of the recognition target 31 (for example, a bag, hat, etc.) are present.
[0022] Fig. 4(b) is an example of the target position PT2 output by the detection unit 21. For example, if the recognition process by the data processing device 10 is a process of estimating the posture of an articulated body that is the recognition target 31 in the input data (image data 30), the target position PT2 may be a keypoint that represents the position of the joint point of the articulated body. Each of the multiple circular plots in Fig. 4(b) represents a keypoint of the recognition target 31, and the straight line L represents a link that is connection information connecting the keypoints together.
[0023] Fig. 5(b) is an example of a target position PT3 output by the detection unit 21. For example, when the recognition process by the data processing device 10 is a voice recognition process of the speech content of the recognition target included in the input data (voice data 32) or an attribute estimation process of the recognition target, the target position PT3 may be an extraction range for extracting a portion to be used in the recognition process from the input voice data 32. The target position PT3 in Fig. 5(b) is an extraction range that starts from a start position (start time) Tst and ends at an end position (end time) Tet. In the following description, the target positions PT1 to PT3 may be collectively referred to as "target positions PT."
[0024] See Fig. 2. The correction unit 22 outputs a corrected position obtained by correcting the target position based on the input data and the target position. For example, the correction unit 22 may extract feature amounts from the input data and the target position using a means such as a neural network, and calculate the amount of correction from the target position to the corrected position based on the extracted feature amounts. In the example of FIG. 3(c), the correction unit 22 outputs a corrected position PC1 obtained by correcting the target position PT1 based on the image data 30 and the target position PT1. For example, the correction unit 22 outputs an offset vector O=(o T ,o B ,o L ,o R ) may be output.
[0025] scalar value o T , o B , o L , o Rare the upper boundary B of the target position PT1, respectively. T , lower boundary B B , left boundary B L , right boundary B R The signs of the correction amounts in the directions of enlarging or shrinking the range indicated by the target position PT1 may be defined as positive or negative. Instead, the correction unit 22 may output the value of each boundary position itself of the correction position PC1 after correction.
[0026] 4(c), the correction unit 22 outputs a corrected position (key point) PC2 obtained by correcting the target position PT2 based on the image data 30 and the target position PT2. For example, the correction unit 22 may output an offset vector O indicating the amount of movement of each target position PT2 to each correction position PC2. Alternatively, the correction unit 22 may output the value of each position of the corrected correction position PC2 itself.
[0027] In the example of FIG. 5(c), the correction unit 22 outputs a corrected position PC3 obtained by correcting the target position PT3 based on the audio data 32 and the target position PT3. The corrected position PC3 is an extracted range that starts from a start position (start time) Tsc and ends at an end position (end time) Tec. For example, the correction unit 22 may output an offset vector O indicating the amount of correction for correcting the start position Tst and end position Tet of the target position PT3 to the start position Tsc and end position Tec of the corrected position PC3, respectively. Alternatively, the correction unit 22 may output the values of the start position Tsc and end position Tec themselves of the corrected corrected position PC3. In the following description, the correction positions PC1 to PC3 may be collectively referred to as "correction positions PC."
[0028] See Fig. 2. The correction unit 22 is a learning model that learns a correction position PC for the input data and the target position PT so as to improve the accuracy of the recognition process in the subsequent processing unit 23. The learning method of the correction unit 22 will be described later. The processing unit 23 executes a predetermined recognition process for the recognition target based on at least the correction position PC.
[0029] For example, the recognition process by the processing unit 23 may be a class recognition process for the recognition target 31. For example, the processing unit 23 may perform a posture estimation process for the recognition target 31 based on the corrected position PC2 (a key point of the articulated body that is the recognition target 31) corrected by the correction unit 22. Furthermore, for example, the processing unit 23 may execute a predetermined recognition process for the recognition target based on the input data and the correction position PC. For example, the processing unit 23 may execute a re-identification process or an attribute estimation process based on the image data 30 that is the input data and the correction position PC1 (the range on the image where the recognition target 31 exists) corrected by the correction unit 22.
[0030] Furthermore, for example, the processing unit 23 may perform speech recognition processing of the speech content of the recognition target or attribute estimation processing of the recognition target based on the portion of the input data, that is, the speech data 32, at the correction position PC3 (cut-out range) corrected by the correction unit 22. For example, the processing unit 23 may be a learning model that learns classes related to recognition targets in advance and recognizes classes related to recognition targets included in input data. In the following description, the classes learned by the processing unit 23 may be referred to as "learned classes." The processing unit 23 outputs the results of the recognition processing via the output unit 16 and the communication unit 13.
[0031] Next, a description will be given of a learning method for the correction unit 22. Fig. 6 is a schematic diagram showing an example of the overall configuration of a learning device 40 used for learning by the correction unit 22. The learning device 40 includes an input unit 42, a communication unit 43, a storage unit 44, a control unit 45, and an output unit 46. Of these, the storage unit 44 and the control unit 45 can be realized by a so-called computer, and the input unit 42, the communication unit 43, and the output unit 46 can be realized as peripheral devices of the computer.
[0032] The input unit 42 includes a user interface such as a keyboard, a mouse, etc. that is operated by the user to input data, etc. The input unit 42 is connected to the control unit 45, converts the user's operation into an operation signal, and outputs it to the control unit 45. The input unit 42 may also include a DVD drive and a USB interface. The input unit 42 inputs data to the control unit 45 as a file, and outputs data from the control unit 45 as a file.
[0033] The communication unit 43 includes a network interface and the like that transmits and receives data between the learning device 40 and an external device via wired or wireless communication. The storage unit 44 is a memory device such as a ROM or RAM, and stores various programs and various data. The control unit 45 is a controller that executes the learning process of the learning model of the correction unit 22 of the data processing device 10, and is composed of an arithmetic unit such as a CPU, a DSP, an MCU, or a GPU.
[0034] The control unit 45 is connected to the storage unit 44, and operates as various processing units by reading and executing computer programs from the storage unit 44, and stores and reads various data in the storage unit 44. The functions of the learning device 40 described below are realized by the control unit 45 executing the computer programs stored in the storage unit 44. The control unit 45 is also connected to the input unit 42 to accept operational inputs to the learning device 40 by the user.
[0035] The control unit 45 is connected to the output unit 46 and outputs the processing results of the control unit 45 to the output unit 46 . The control unit 45 is also connected to the communication unit 43, and transmits and receives various data to and from external devices via the communication unit 43. Parameters of the learning model of the correction unit 22 acquired by the learning process of the control unit 45 may be transmitted to the data processing device 10 via the communication unit 43. The learning device 40 and the data processing device 10 may be configured as the same hardware. For example, the data processing device 10 may have a learning function performed by the learning device 40 described below.
[0036] 7 is a block diagram showing an example of the functional configuration of learning device 40. In addition to the above-mentioned storage unit 44, learning device 40 includes a correction unit 50, a processing unit 51, a learning data acquisition unit 52, a target position setting unit 53, and a learning unit 54. The correction unit 50 is a learning model having the same configuration as the correction unit 22 of the data processing device 10. For example, the correction unit 50 may have the same hardware configuration as the correction unit 22.
[0037] The learning device 40 generates parameters of the trained learning model obtained by training the learning model of the correction unit 50 as parameters of the learning model of the correction unit 22 of the data processing device 10. When the learning device 40 and the data processing device 10 are configured using the same hardware, the correction unit 22 of the data processing device 10 may be used as the correction unit 50 of the learning device 40.
[0038] The processing unit 51 is a learning model that performs the same recognition processing as the processing unit 23 of the data processing device 10 (i.e., performs the recognition processing with the same parameters as the processing unit 23). That is, the processing unit 51 may be a learning model that recognizes a learned class included in input data. When the learning device 40 and the data processing device 10 are configured using the same hardware, the processing unit 23 of the data processing device 10 may be used as the processing unit 51 of the learning device 40 . The learning data acquisition unit 52 acquires learning data used for learning by the correction unit 50 from the storage unit 44 .
[0039] The learning data used for learning by the correction unit 50 may be, for example, data used for learning the learning model of the processing unit 51. For example, when the recognition processing by the processing unit 51 is a re-identification process of a recognition target, an attribute estimation process, or a voice recognition process, data including a learned class learned by the processing unit 51 may be acquired as the learning data. Furthermore, for example, if the recognition processing of the processing unit 51 is a processing for estimating the posture of the recognition target, the processing unit 51 may acquire, as training data, image data from which key points are extracted when learning the posture of the learned class.
[0040] The target position setting unit 53 detects the target position, which is the position of the recognition target in the learning data, and outputs the detected target position to the correction unit 50. For example, the target position setting unit 53 may detect the target position using a method similar to that of the detection unit 21 of the data processing device 10, add a perturbation to the detected target position, and output the perturbation to the correction unit 50. When the learning device 40 and the data processing device 10 are configured using the same hardware, the target position setting unit 53 may use the detection unit 21 of the data processing device 10. Also, for example, the target position setting unit 53 may set the target position randomly.
[0041] The correction unit 50 corrects the target position based on the learning data and the target position, and outputs the corrected position to the processing unit 51. The processing unit 51 executes a predetermined recognition process for the recognition target based on at least the corrected position. For example, the processing unit 51 may execute a posture estimation process for the recognition target based on the corrected position corrected by the correcting unit 50.
[0042] Furthermore, for example, the processing unit 51 may perform a predetermined recognition process for the recognition target based on the training data and the correction position. For example, the processing unit 51 may perform a re-identification process or an attribute estimation process based on the image data of the training data and the correction position. For example, the processing unit 51 may perform a voice recognition process or an attribute estimation process based on the voice data of the training data and the correction position.
[0043] The learning unit 54 calculates the error (e.g., a loss function) between the processing result of the processing unit 51 and the learned class (correct class), and adjusts the parameters of the learning class of the correction unit 50 so as to reduce the error. For example, the learning unit 54 calculates the parameter update amount of the learning model to reduce the error using a gradient method or a coordinate descent method with the error as an energy function, updates the learning model by the update amount, and then has the processing unit 51 perform the recognition process again to evaluate the error. This process is repeated until an iteration termination condition is met. Note that the learning unit 54 may calculate the error between the processing result of the processing unit 51 and the learned class, as well as the error between the learned class and a class different from the learned class, and adjust the parameters of the learning class of the correction unit 50 taking each error into consideration. For example, the learning unit 54 calculates the error between the learned class and a class different from the learned class, calculates the parameter update amount of the learning model to reduce the sum or average error of each error, updates the learning model by the update amount, and has the processing unit 51 perform the recognition process again to evaluate the error. This process is repeated until an iteration termination condition is met.
[0044] As described above, by training the correction unit 50 (i.e., by training the correction unit 22), the correction unit 50 can calculate a correction position that matches the characteristics of the processing unit 23. For example, depending on whether the recognition processing of the processing unit 23 is a re-identification processing based on an image of the upper body of the person to be recognized or a re-identification processing based on the clothing of the person in the image data, a correction position that is suitable for each re-identification processing can be calculated.
[0045] In addition, when the detection unit 21 is constructed using a learning model, the correction position output by the learned correction unit 22 and the input data may be used as learning data (teacher data) to learn the learning model of the detection unit 21. For example, a learning data generation device may be provided that generates learning data for the detection unit 21 using the correction position output by the correction unit 22. That is, the learning data generation device may include a learning data output unit that outputs learning data for the detection unit 21 from the correction position output by the correction unit 22 and the input data. By training the detection unit 21 using the correction position output by the correction unit 22, the detection unit 21 can be trained to output a target position PT suitable for recognition processing in the processing unit 23.
[0046] Furthermore, the learning device 40 may use the correction positions output by the trained correction unit 50 as learning data for training the processing unit 23 of the data processing device 10. For example, a learning data generation device may be provided that generates learning data for the processing unit 23 using the correction positions output by the correction unit 22. That is, the learning data generation device may include a learning data output unit that outputs learning data for the detection unit 21 based on the correction positions output by the correction unit 22 and the input data. For example, when training the processing unit 23, the range in which the recognition target exists in the image data or audio data input as training data may be specified by the correction positions output by the trained correction unit 50, and the processing unit 23 may be trained using the specified portion of data. This allows training to be performed using a portion of the training data that is appropriate for recognition processing, thereby improving the accuracy of the processing unit 23.
[0047] (operation) FIG. 8(a) is a flowchart of an example of a learning method for the corrector 50 (corrector 22). In step S1, the learning data acquisition unit 52 acquires learning data to be used for learning by the correction unit 50. In step S2, the target position setting unit 53 sets the target position, which is the position of the recognition target in the learning data.
[0048] In step S3, a corrected position is determined by correcting the target position based on the learning data and the target position. In step S4, the processing unit 51 executes a predetermined recognition process for the recognition target based on at least the correction position. In step S5, the learning unit 54 learns the correction unit 50 so as to reduce the error between the processing result of the processing unit 51 and the learned class, and then the process ends.
[0049] FIG. 8B is a flowchart of an example of the data processing method according to the embodiment. In step S11, the input data acquisition unit 20 acquires input data to be used in the recognition process by the data processing device 10. In step S12, the detection unit 21 determines the target position, which is the position of the recognition target within the input data.
[0050] In step S13, the correction unit 22 corrects the target position based on the input data and the target position, and outputs the corrected position. In step S14, the processing unit 23 executes a predetermined recognition process for the recognition target based on at least the correction position, and then the process ends.
[0051] (Effects of the embodiment) (1) The data processing device 10 includes a detection unit 21 that outputs a target position, which is the position of a recognition target within input data, a correction unit 22 that outputs a corrected position obtained by correcting the target position based on the input data and the target position, and a processing unit 23 that executes a predetermined recognition process for the recognition target based on at least the corrected position. The correction unit 22 is a learning model that learns the corrected position for the input data and the target position so as to improve the accuracy of the recognition process in the processing unit 23. This allows the recognition process to be performed based on the position of the recognition target that is suitable for the recognition process in the processing unit 23. As a result, the recognition accuracy of the recognition process in the processing unit 23 can be improved.
[0052] (2) The processing unit 23 may be a learning model that recognizes a learned class that has been learned in advance for the recognition target. The correction unit 22 may be a learning model that has been trained so as to reduce an error between the processing result by the processing unit 23 and the learned class when the processing unit 23 recognizes the learned class based on the correction position. This makes it possible to improve the recognition accuracy of the recognition process executed by the learning model that recognizes classes related to the recognition target.
[0053] (3) The processing unit 23 may execute a predetermined recognition process for the recognition target based on the input data and the correction position. The processing unit 23 may be a learning model that recognizes a learned class that has been learned in advance for the recognition target. The correction unit 22 may be a learning model that has been trained so as to reduce an error between the processing result by the processing unit 23 and the learned class when the processing unit 23 recognizes the learned class based on the correction position from input data including the learned class learned by the processing unit 23. This can improve the recognition accuracy of the predetermined recognition process that is performed based on the input data and the correction position.
[0054] (4) The input data may be image data, and the target position may be the image area occupied by the target in the image data. This can improve the recognition accuracy of, for example, the re-identification process and attribute estimation process of the target. (5) The input data may be image data, the recognition target may be an articulated body, and the target position may be an indirect position of the recognition target on the image data. This improves the recognition accuracy of the posture estimation process for the articulated body to be recognized.
[0055] (6) The input data may be speech data, and the target position may be the start and end positions of the range of the recognition target in the speech data. This improves the recognition accuracy of speech recognition processing based on speech data and attribute estimation processing of the recognition target. (7) A training data generation device generates training data for a position detector that outputs a target position, which is the position of a recognition target within input data. The training data generation device may include a detection unit that outputs a target position, which is the position of a recognition target within the input data; a correction unit that outputs a corrected position obtained by correcting the target position based on the input data and the target position; and a training data output unit that outputs training data from the corrected position output by the correction unit and the input data. The correction unit may be a learning model that learns the corrected position for the input data and the target position based on the input data and the corrected position output by the correction unit so as to improve recognition accuracy when a predetermined recognition process is performed on the recognition target included in the input data. This allows the position detector to be trained to output a target position suitable for the predetermined recognition process.
[0056] (8) A training data generation device generates training data for a recognition model that performs a predetermined recognition process on a recognition target included in input data. The training data generation device includes a detection unit that outputs a target position, which is the position of the recognition target within the input data, a correction unit that outputs a corrected position obtained by correcting the target position based on the input data and the target position, and a training data output unit that outputs training data from the corrected position output by the correction unit and the input data. The correction unit is a training model that has learned the corrected position for the input data and the target position so as to improve recognition accuracy when a predetermined recognition process is performed based on the input data and the corrected position output by the correction unit. This allows learning to be performed using data from the input data that is appropriate for recognition processing, thereby improving the accuracy of the processing unit. [Explanation of symbols]
[0057] 1...recognition system, 2...recognition target, 3...sensor, 10...data processing device, 12, 42...input unit, 13, 43...communication unit, 14, 44...storage unit, 15, 45...control unit, 16, 46...output unit, 20...input data acquisition unit, 21...detection unit, 22, 50...correction unit, 23, 51...processing unit, 30...image data, 31...recognition target, 32...audio data, 40...learning device, 52...learning data acquisition unit, 53...target position setting unit, 54...learning unit
Claims
1. a detection unit that outputs a target position, which is the position of a recognition target within input data; a correction unit that corrects the target position based on the input data and the target position and outputs a corrected position; a processing unit that executes a predetermined recognition process for the recognition target based on at least the correction position, the correction unit is a learning model that learns the correction position for the input data and the target position so as to improve accuracy of the recognition processing in the processing unit. A data processing device comprising:
2. the processing unit is a learning model that recognizes a learned class that has been learned in advance regarding the recognition target, the correction unit is a learning model that is trained so as to reduce an error between a processing result by the processing unit and the learned class when the processing unit recognizes the learned class based on the correction position.
2. The data processing device according to claim 1.
3. 2. The data processing apparatus according to claim 1, wherein the processing unit executes a predetermined recognition process for the recognition target based on the input data and the correction position.
4. the processing unit is a learning model that recognizes a learned class that has been learned in advance regarding the recognition target, the correction unit is a learning model that is trained so as to reduce an error between a processing result by the processing unit and the learned class when the processing unit recognizes the learned class based on the correction position from the input data including the learned class.
4. The data processing device according to claim 3.
5. 2. The data processing apparatus according to claim 1, wherein the input data is image data, and the target position is an image range occupied by the recognition target in the image data.
6. 2. The data processing apparatus according to claim 1, wherein the input data is image data, the recognition target is an articulated body, and the target position is an indirect position of the recognition target on the image data.
7. 2. The data processing apparatus according to claim 1, wherein the input data is voice data, and the target position is a start position and an end position of a range occupied by the recognition target in the voice data.
8. A training data generation device for generating training data for a position detector that outputs a target position, which is a position of a recognition target in input data, comprising: a detection unit that outputs a target position that is a position of the recognition target within the input data; a correction unit that corrects the target position based on the input data and the target position and outputs a corrected position; a learning data output unit that outputs the learning data based on the correction position output by the correction unit and the input data; Equipped with the correction unit is a learning model that learns the correction position for the input data and the target position based on the input data and the correction position output by the correction unit so as to improve recognition accuracy when a predetermined recognition process is executed on a recognition target included in the input data. A training data generation device characterized by:
9. determining a target position, which is the position of a recognition target within the input data; determining a corrected position obtained by correcting the target position based on the input data and the target position using a learning model; Executing a predetermined recognition process for the recognition target based on at least the corrected position; The learning model learns the corrected position to be output for the input data and the target position so that accuracy of the recognition process is improved. A data processing method comprising:
10. A learning data generation method for generating learning data for a position detector that outputs a target position, which is a position of a recognition target in input data, comprising: determining a target position, which is a position of the recognition target within the input data; outputting a corrected position obtained by correcting the target position using a learning model based on the input data and the target position; outputting the learning data from the correction position and the input data; learning the correction position output by the learning model for the input data and the target position, based on the input data and the correction position output by the learning model, so as to improve recognition accuracy when a predetermined recognition process is executed on the recognition target included in the input data; A training data generation method comprising:
Citation Information
Patent Citations
Association device and association method
JP2023147846A