Information processing device, information processing method, and program
By generating scene images and labels with recognition rates lower than a predetermined value, automatically associating and repeating the process, the problem of insufficient recognition rate is solved, the burden of manual annotation is reduced, and the recognition rate and accuracy of the recognition model are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-09
- Publication Date
- 2026-04-21
AI Technical Summary
In existing technologies, the recognition model has insufficient recognition rate in specific scenarios, and requires a large number of manually labeled images, which increases labor and cost, and the number of images available for learning is limited.
By generating scene images with a recognition rate lower than a predetermined value along with labels, an image generation unit is constructed using image generation technology. This unit automatically associates images with labels and repeats the process until the recognition rate reaches the predetermined value, reducing the burden of manual annotation and increasing the amount of learning data.
It has improved the recognition rate, reduced labor and costs, and can generate a large number of images and labels under different conditions, ensuring sufficient learning data and improving the recognition accuracy of the recognition model.
Smart Images

Figure CN121909496A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to information processing apparatus, information processing methods and procedures, and in particular to information processing apparatus, information processing methods and procedures that can easily improve the recognition rate of recognition models. Background Technology
[0002] The technology of using machine learning to generate image-based information to identify the categories and states of people and objects has become widely used.
[0003] The recognition model is generated by machine learning using learning data, but for certain scenarios, it is sometimes impossible to learn sufficiently and the recognition rate is insufficient.
[0004] Therefore, it is often necessary to perform training to improve the recognition rate of images in specific scenes with low recognition rates.
[0005] To this end, the following technology has been proposed: extracting images similar to those of specific scenes with recognition rates below a predetermined value from the verification images used to verify the recognition rate of the recognition model and using them for learning, thereby improving the recognition rate (see Patent Document 1).
[0006] The following technology is also proposed: pre-collecting images actually taken in multiple countries, extracting images similar to images of specific scenes with low recognition rates from the collected images and using them for learning, thereby improving the recognition rate (see Patent Document 2).
[0007] Existing technical documents
[0008] Patent documents
[0009] Patent Document 1: Japanese Patent Application Publication No. 2019-109924
[0010] Patent Document 2: Japanese Patent Application Publication No. 2019-021201 Summary of the Invention
[0011] The technical problem that the invention aims to solve
[0012] However, in the technologies of patent documents 1 and 2, manual annotation of the extracted images is required, and the more images extracted, the greater the labor and burden of annotation.
[0013] Furthermore, in the technologies of Patent Documents 1 and 2, the scene images used for learning that have low recognition rates are extracted from images that are pre-generated as verification images or from pre-collected images, thus limiting the number of images available for learning. Consequently, it may be impossible to ensure the necessary number of learning images, thus failing to adequately improve the recognition rate.
[0014] This disclosure is made in view of the following situation, in particular by generating scene images with labels together with the recognition model, thereby easily improving the recognition accuracy of the recognition model.
[0015] Technical solutions for solving technical problems
[0016] As one aspect of this disclosure, the information processing apparatus and program include: an evaluation unit that evaluates the recognition rate of a recognition model on verification data consisting of images and labels generated based on multiple conditions, and determines conditions in the recognition model that result in a low recognition rate lower than a predetermined recognition rate; and a low recognition rate data generation model that generates low recognition rate data of the recognition model consisting of the images and labels based on the determined low recognition rate conditions, wherein the program enables a computer to function as the evaluation unit and the low recognition rate data generation model.
[0017] One aspect of the information processing method disclosed herein includes: evaluating the recognition rate of a recognition model for verification data consisting of images and labels generated based on multiple conditions; determining a condition in the recognition model that has a low recognition rate lower than a predetermined recognition rate; and generating low recognition rate data of the recognition model consisting of the images and labels based on the determined low recognition rate condition.
[0018] According to one aspect of this disclosure, an evaluation of the recognition model for verification data consisting of images and labels generated based on multiple conditions is performed, a condition for a low recognition rate in the recognition model that is lower than a predetermined recognition rate is determined, and low recognition rate data of the recognition model consisting of the images and labels is generated based on the determined low recognition rate condition. Attached Figure Description
[0019] Figure 1 This is a diagram illustrating the summary of this disclosure.
[0020] Figure 2 This is a diagram illustrating a structural example of the first embodiment of the identification model improvement system of this disclosure.
[0021] Figure 3 This is a diagram illustrating the functions achieved by improving the system through the recognition model.
[0022] Figure 4 This is a diagram illustrating the data generation model for validation and the data generation model for low recognition rate.
[0023] Figure 5 This is a diagram illustrating the model of this disclosure obtained by applying Unidiffuser.
[0024] Figure 6This diagram illustrates the method by which the identification model evaluation department evaluates the identification model.
[0025] Figure 7 This is an example diagram illustrating the generated image and label when the condition of low recognition rate is a label.
[0026] Figure 8 This is a diagram illustrating an example of the image and label generated when the condition of low recognition rate is an image.
[0027] Figure 9 This is an example diagram illustrating the generated image and labels when the condition for low recognition rate is text.
[0028] Figure 10 This is an example diagram illustrating the generated image and labels when the condition for low recognition rate is text.
[0029] Figure 11 This is an explanation of the reason. Figure 3 The flowchart shows the recognition model improvement process performed by the recognition model improvement system.
[0030] Figure 12 This is a diagram illustrating a structural example of a second embodiment of the identification model improvement system of this disclosure.
[0031] Figure 13 This is a diagram illustrating a structural example of a third embodiment of the identification model improvement system of this disclosure.
[0032] Figure 14 This is a diagram illustrating an application example of this disclosure.
[0033] Figure 15 This is a diagram illustrating an example of the structure of a general-purpose computer. Detailed Implementation
[0034] The preferred embodiments of this disclosure will now be described in detail with reference to the accompanying drawings. Furthermore, in this specification and the drawings, constituent elements having substantially the same functional structure are given the same reference numerals, thereby omitting redundant descriptions.
[0035] The following describes the methods used to implement this technology. The descriptions will proceed in the following order.
[0036] 1. Summary of this disclosure
[0037] 2. First Implementation Method
[0038] 3. Second Implementation Method
[0039] 4. Third Implementation Method
[0040] 5. Application Examples
[0041] 6. Examples executed by software
[0042] <<1. Summary of this Disclosure>>
[0043] This disclosure specifically improves the recognition rate of a recognition model by generating scene images with labels together with the recognition model, where the recognition rate is lower than a predetermined value.
[0044] First, let me state the general outline of this disclosure. The technology of using machine learning to generate recognition models that identify the categories and states of people and objects within images based on image information has become widespread.
[0045] The recognition model is generated by machine learning using learning data. However, due to the limited categories of images used as learning data, sometimes for images under specific conditions (scenes), the learning is insufficient and the recognition rate (the success rate of recognition processing performed by the recognition model) is lower than the predetermined value, resulting in inappropriate recognition.
[0046] Therefore, it is often necessary to perform learning (relearning) to improve the recognition accuracy of images under specific conditions (scenes) with low recognition rates.
[0047] To address this, the following technique has been proposed: extract images similar to those in specific scenes with recognition rates below a predetermined value, attach correct labels, and learn from them to improve recognition accuracy.
[0048] However, in this case, the extracted images need to be manually labeled correctly, and the more images extracted, the greater the labor and burden of labeling them.
[0049] In addition, the scene images with low recognition rates are extracted from pre-generated images or pre-collected images rather than newly created images. Therefore, it is sometimes impossible to ensure that the number of images that should be needed to improve the recognition rate is sufficient to improve the recognition rate.
[0050] Therefore, in this disclosure, an image generation unit is constructed by applying existing image generation techniques that generate images by specifying text or images, thereby generating images by associating them with labels based on conditions, labels, etc.
[0051] Furthermore, the image generation unit of this disclosure is repeatedly used to perform the above-described process. Figure 1 The processing steps S1 to S3 shown improve the recognition rate of the recognition model.
[0052] exist Figure 1In the processing, firstly in step S1, the image generation unit generates images for various conditions (scenes) and associates them with corresponding labels. Then, using the images generated by the image generation unit, the recognition model performs recognition processing, calculates the recognition rate by comparing the recognition results with the corresponding labels, and determines the conditions (scenes) that are lower than the predetermined recognition rate.
[0053] Here, for example, when the user has pre-determined a condition (scenario) where the recognition rate is lower than a predetermined recognition rate using a recognition model, an image of the predetermined condition (scenario) can be generated.
[0054] Furthermore, even if the recognition model is used by the user in advance, there may still be conditions (scenarios) where the recognition rate is lower than the predetermined recognition rate that the user may not notice. Therefore, in the initial processing, images under predetermined conditions (scenarios) can be generated along with the corresponding labels, as well as images under various different conditions (scenarios).
[0055] Next, in step S2, the image generation unit generates a large number of images and labels for conditions (scenes) that are determined to be below the predetermined recognition rate.
[0056] Then, through the processing in step S3, the recognition model is trained using images and labels that meet the conditions (scenes) below the predetermined recognition rate.
[0057] After the processing in step S3, return to the processing in step S1, and repeat the processing from step S1 to S3 until the condition (scene) below the predetermined recognition rate is no longer identified (does not exist).
[0058] Furthermore, in the processing after the second time in step S1, it is also possible to determine the conditions (scenes) in which the recognition rate of the learned recognition model is still lower than the predetermined value in the images that were identified as having a recognition rate lower than the predetermined value in the first processing.
[0059] Through the above processing, the conditions with low recognition rates can be identified, and images and labels for the identified low recognition rate conditions (scenes) can be generated. The process of learning using the generated low recognition rate condition (scene) images and labels is repeated until the recognition rate is higher than the predetermined value, thereby improving the recognition rate of the recognition model.
[0060] At this time, since the image generation unit of this disclosure is used to generate images and labels for conditions (scenes) where the recognition rate is lower than a predetermined value, it is possible to reduce the labor and cost of manually labeling each image one by one.
[0061] Furthermore, it can generate images and labels for various conditions (scenes) with reduced labor and costs, thus enabling the easy generation of large amounts of learning data. Consequently, based on the recognition results of the recognition model using images and labels for various conditions (scenes), it is easy and appropriate to determine the conditions (scenes) where the recognition rate is lower than a predetermined value.
[0062] Furthermore, because it is possible to generate a large number of images and labels for conditions (scenes) where the recognition rate is lower than the predetermined value, it is easy to prepare a sufficient number of images and labels to achieve a state where the recognition rate of the recognition model is higher than the predetermined value, resulting in a high-precision improvement in the recognition rate of the recognition model.
[0063] <<2. First Implementation>>
[0064] The following will refer to Figure 2 An example of the structure of the identification model improvement system that applies the technology disclosed herein is described.
[0065] Figure 2 The identification model improvement system 11 includes smartphones 31-1 to 31-n, a cloud server 32, and a network 33. The smartphones 31-1 to 31-n and the cloud server 32 are structures that can communicate with each other via a network 33, such as a public line or the Internet, and are, for example, structures that can send and receive data and programs.
[0066] Furthermore, without special distinction Figure 2 In the case of smartphones 31-1 to 31-n, they are simply referred to as smartphone 31, and the same applies to other structures.
[0067] The smartphone 31 is carried by the user and has a so-called recognition function, which is to identify people and objects as subjects in images taken by cameras (not shown) or images provided by other smartphones 31 via the network 33, or to identify categories and states.
[0068] The so-called recognition function of people and objects as subjects in an image refers to the function of identifying individuals and states by using images of people, faces, etc. as subjects in an image, or the function of identifying the category and state of objects as subjects.
[0069] More specifically, the smartphone 31 has a recognition model improvement request unit 51 and a recognition model 52.
[0070] The recognition model 52 includes DNN (Deep Neural Network) and other components, which are constructed through machine learning and have the aforementioned recognition functions of identifying people and objects in images, or identifying categories and states based on images.
[0071] When using the recognition model 52, if the user reports that the recognition rate of the image for a specific condition (scene) is low in the user's perception, or if the recognition rate calculated based on the actual recognition result is lower than the predetermined value, the recognition model improvement request unit 51 requests an improvement from the cloud server 32 via the network 33.
[0072] When a request for improvement of the recognition model 52 is received from the smartphone 31, the cloud server 32 determines the conditions (scenarios) in the recognition model 52 where the recognition rate is low, and relearns the recognition model in a way that improves the recognition rate under the determined conditions, so as to update the recognition model 52 of the smartphone 31 that made the improvement request.
[0073] More specifically, the cloud server 32 has a recognition model improvement processing unit 71 and a recognition model update unit 72.
[0074] When a request to improve the recognition model 52 is received from the smartphone 31, the recognition model improvement processing unit 71 determines the conditions (scenes) where the recognition rate of the recognition model 52 is low. More specifically, the recognition model improvement processing unit 71 generates images and labels for various different conditions (scenes).
[0075] In addition, the recognition model improvement processing unit 71 has the same recognition model as the recognition model 52 built into the smartphone 31. It uses images and labels generated under various conditions (scenes) to calculate the recognition rate of the recognition model 52, and determines the conditions (scenes) with low recognition rates based on the recognition rate results.
[0076] Then, the recognition model improvement processing unit 71 generates images and labels for the identified low recognition rate conditions (scenes), and uses the generated images and labels for the identified low recognition rate conditions (scenes) to enable the recognition model 52 to relearn.
[0077] In addition, the recognition model improvement processing unit 71 calculates the recognition result obtained by relearning the recognition model using the image and label of the determined low recognition rate condition (scene), and determines whether the decline in the recognition rate of the recognition model has been improved based on the recognition result.
[0078] Furthermore, if the recognition result does not indicate that the recognition rate of the recognition model has decreased and thus improved, the same process is repeated until the recognition rate exceeds a predetermined value.
[0079] Then, when it is determined from the recognition results that the recognition rate of the recognition model has decreased and has been improved, the recognition model update unit 72 updates the recognition model 52 of the smartphone 31 that made the improvement request via the network 33 with the recognition model that has been improved due to the decrease in recognition rate.
[0080] By Figure 2 The functions implemented by the improved recognition model system >
[0081] Next, refer to Figure 3 , for Figure 2 The functions implemented by the improved recognition model system will be explained here. In addition, the functions implemented by the cloud server 32 will be explained in particular.
[0082] As described above, the cloud server 32 includes a recognition model improvement processing unit 71 and a recognition model update unit 72.
[0083] The recognition model improvement processing unit 71 includes a condition generation unit 91, a verification data generation model 92, a recognition model 93, a recognition model evaluation unit 94, a low recognition rate data generation model 95, and a learning unit 96.
[0084] When a request to improve the recognition rate of the recognition model 52 is received from the smartphone 31, the condition generation unit 91 generates conditions for generating verification data in the subsequent verification data generation model 92.
[0085] When an improvement request is provided from the recognition model improvement request unit 51, such as a request to improve the recognition rate drop under specific conditions, the condition generation unit 91 provides the specific conditions contained in the improvement request to the verification data generation model 92.
[0086] In addition, when the improvement request provided by the recognition model improvement request unit 51 is an improvement request for the overall decline in recognition rate, the condition generation unit 91 generates various different conditions and provides them to the verification data generation model 92 in order to improve the overall decline in recognition rate.
[0087] The verification data generation model 92 generates images and corresponding correct labels as verification data in pairs based on the conditions provided by the condition generation unit 91, and provides them to the recognition model 93. Further details regarding the verification data generation model 92 will be explained later.
[0088] The recognition model 93 is generated based on machine learning, such as a DNN (Deep Neural Network), by the learning unit 96, and is eventually updated to the recognition model 52 of the smartphone 31.
[0089] The recognition model 93 recognizes the image based on the pairing of the image and its correct label provided by the verification data generation model 92 from the verification data, and outputs the recognition result and the correct label to the recognition model evaluation unit 94.
[0090] Furthermore, it is assumed that the recognition model 93 is exactly the same as the recognition model 52 of the smartphone 31 before the improvement request is received. However, the recognition model 93 is improved after receiving the improvement request and updated as the recognition model 52 when it is confirmed that the recognition rate has been sufficiently improved. Therefore, it is different from the recognition model 52 before the improvement request is received and it is updated.
[0091] The recognition model evaluation unit 94 evaluates the recognition model 93 by comparing the recognition result obtained by the recognition model 93, which uses the image and label as verification data, with the label. The recognition model evaluation unit 94 determines a low recognition rate condition (where the recognition rate is below a predetermined threshold) from the evaluation results of the recognition model 93 and outputs it to the low recognition rate data generation model 95.
[0092] The low recognition rate data generation model 95 generates images as low recognition rate data and their correct labels based on the low recognition rate conditions provided by the recognition model evaluation unit 94, and provides them to the learning unit 96.
[0093] The low recognition rate data generation model 95 is basically the same as the verification data generation model 92, except that when generating image and label data, the image and label are generated as low recognition rate data based on the low recognition rate conditions determined by the evaluation results of the recognition model evaluation unit 94.
[0094] Furthermore, details regarding the low-recognition-rate data generation model 95 will be provided later, along with the verification data generation model 92.
[0095] The learning unit 96 uses the images provided by the low recognition rate data generation model 95 as low recognition rate data and their correct labels to enable the recognition model 93 to learn (relearn), thereby improving the decline in recognition rate.
[0096] After the recognition model 93 has been trained by the learning unit 96, the recognition model 93 performs recognition processing again using the verification data. The recognition model evaluation unit 94 determines the conditions for low recognition rate based on the recognition results. Then, if the conditions for low recognition rate exist in the recognition results of the recognition model 93 and are determined, the improvement of the recognition model 93 is insufficient. Therefore, the determined low recognition rate conditions are output to the low recognition rate data generation model 95. The low recognition rate data generation model 95 generates low recognition rate data, and the learning unit 96 retrains the recognition model 93 based on the low recognition rate data.
[0097] That is, the same process is repeated until the recognition model evaluation unit 94 can no longer extract the conditions that cause the low recognition rate from the recognition results of the recognition model 93, that is, the decline in the recognition rate is improved.
[0098] Then, if no condition with a low recognition rate can be extracted from the recognition results of recognition model 93, the recognition rate of recognition model 93 is considered to have been sufficiently improved. Therefore, recognition model update unit 72 uses recognition model 93 in the state of improved recognition rate to update recognition model 52 of smartphone 31.
[0099] <Regarding data generation models for validation and low-recognition-rate data generation models>
[0100] Next, we will provide a detailed explanation of the data generation model 92 for verification and the data generation model 95 for low recognition rate.
[0101] Both the validation data generation model 92 and the low-recognition-rate data generation model 95 are based on conditionally generating images and corresponding labels. Furthermore, in Figure 4 In the text, the data generation model 92 for validation and the data generation model 95 for low recognition rate are described as generation models 92 and 95.
[0102] like Figure 4 As shown, the validation data generation model 92 and the low recognition rate data generation model 95 generate images and corresponding labels in pairs based on conditions specified by various labels and text, including key points, facebboxes, and semantic segmentation.
[0103] Figure 4 The following example is shown: Only the key point KP is specified as a condition. The verification data generation model 92 and the low recognition rate data generation model 95 generate key points (CKP), face bounding boxes (CFBB), and semantic segments (Seg) as corresponding labels when generating the image CP.
[0104] In addition, although Figure 4 Only the key point KP can be specified as a condition, but at least one of other text or labels can also be specified as a condition. If multiple labels or texts are specified as conditions, it is a combination of multiple conditions. Additionally, the validation data generation model 92 and the low-recognition-rate data generation model 95 can also generate an image and at least one other label.
[0105] <Methods for generating models using validation data and models generated using low-recognition-rate data>
[0106] Next, refer to Figure 5 The generation methods for the validation data generation model 92 and the low recognition rate data generation model 95 are explained.
[0107] As for the data generation model 92 for validation and the data generation model 95 for low recognition rate, the results obtained by applying a technique called UniDiffuser can be utilized, for example. UniDiffuser generates images, text, and images and text based on input conditions, and generates text from images and images from text.
[0108] Additionally, for details about UniDiffuser, please refer to "One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale" by Fan Bao, Shen Nie, Kaiwen Xue, Chongxuan Li, Shi Pu, Yaole Wang, Gang Yue, Yue Cao, Hang Su, and Jun Zhu (https: / / arxiv.org / pdf / 2303.06555.pdf).
[0109] The verification data generation model 92 and the low-recognition-rate data generation model 95 of this disclosure can, for example, be extended by UniDiffuser to generate labels in addition to images and text. The labels referred to herein include various types of labels, such as bounding boxes (faces, human bodies, etc.), key points (human body models), and semantic segmentation.
[0110] More specifically, such as Figure 5 As shown, during learning, different noise strengths (t) are assigned to each type of data (by image, text, and label). x , t y , t c The model M11, shown in the dashed box in the figure, is trained based on the feature group (x) of the noise generated according to the noise intensity described above. tx y ty c tc To infer the noise (E) assigned to each data point. θ x E θ y E θ c ).
[0111] The verification data generation model 92 and the low recognition rate data generation model 95 disclosed herein are formed based on model M11.
[0112] That is, by learning the labels like model M11 to recursively repeat the noise inference process according to the appropriate noise intensity set according to the input conditions, it is possible to generate a matching image with the correct label that meets the input conditions.
[0113] in addition, Figure 5 Model M11 is a structure formed by adding the supplementary learning part model M2 (dashed box in the figure) disclosed herein to the learned part model M1 (solid box in the figure) used to construct the original UniDiffuser, so as to make it adaptable to the labels.
[0114] That is, the original UniDiffuser uses a partially learned model M1 as a basis to form a structure to handle text and images.
[0115] In contrast, the verification data generation model 92 and the low recognition rate data generation model 95 of this disclosure are based on model M11, which is formed by adding corresponding labels to the learned partial model M1.
[0116] Therefore, when forming model M11, the structure is adopted which uses the learned partial model M1 obtained by pairing existing images and text and adds an additional learned partial model M2 for corresponding labels, thus reducing the learning cost.
[0117] However, regarding model M11, it can be formed in the same way as the learned partial model M1 used in the learning of the original UniDiffuser, through learning that can generate pairs of images and labels.
[0118] Furthermore, this specification is intended to illustrate the following example: the verification data generation model 92 and the low recognition rate data generation model 95 use models that, by extending UniDiffuser, can generate labels in addition to images and text. However, the verification data generation model 92 and the low recognition rate data generation model 95 are not limited to models obtained by extending UniDiffuser, as long as they possess the same functionality as the models obtained by extending UniDiffuser; they can also be other models with the same functionality.
[0119] <Examples of conditions leading to low recognition rates, as determined by the recognition model evaluation department>
[0120] Next, refer to Figure 6 This section describes specific evaluation examples from the evaluation section 94 of the identification model.
[0121] For example, consider the following situation: the condition generation unit 91 generates the text "sunny day" and "bad weather day" as conditions.
[0122] In this case, the verification data generation model 92 generates images and correct labels that satisfy various conditions such as "sunny day" and "bad weather day" composed of text as verification data and provides them to the recognition model 93. Here, the verification data generation model 92 generates multiple pairs consisting of images and correct labels that satisfy various conditions such as "sunny day" and "bad weather day" composed of text as verification data and provides them to the recognition model 93.
[0123] The recognition model 93 performs predetermined recognition processing based on images that satisfy various conditions such as "sunny day" and "bad weather day" composed of text, which are provided by the verification data generation model 92 from the verification data. The recognition result and the corresponding label are then provided to the recognition model evaluation unit 94.
[0124] When the recognition result and correct label are obtained from the recognition model 93, the recognition model evaluation unit 94 calculates the recognition rate (recognition success rate) for each condition, identifies the conditions below the predetermined threshold as low recognition rate conditions, and outputs the identified low recognition rate conditions to the low recognition rate data generation model 95.
[0125] Figure 6 The example shown has a 97% success rate in recognizing images under "sunny" conditions and a 65% success rate in recognizing images under "bad weather" conditions.
[0126] For example, when the threshold for recognition success rate is set to 90%, the recognition model evaluation unit 94 determines the condition of "bad weather day" as the condition of low recognition rate and outputs it to the low recognition rate data generation model 95.
[0127] In addition, Figure 6 In addition to the conditions shown, when the recognition success rate of a specific label is lower than the threshold, the recognition model evaluation unit 94 may also use the specific label as a low recognition rate condition.
[0128] For example, if the image is of a "bad weather day", and there are key points that identify a person as the subject posing in a specific pose, the recognition model evaluation unit 94 can use the label of "bad weather day" and key points indicating a specific pose as a low recognition rate condition.
[0129] In addition, the recognition model evaluation unit 94 can extract at least one of the following as a low recognition rate condition: the label, text, and image (the image itself with a low recognition rate in the evaluation result of the recognition model 93) associated with the image whose recognition success rate is lower than a predetermined value, and output it to the low recognition rate data generation model 95.
[0130] <Example of low-recognition-rate data generation in a low-recognition-rate data generation model>
[0131] Next, refer to Figures 7 to 10 This paper explains the generation example of low recognition rate data in the low recognition rate data generation model 95.
[0132] (Firstly)
[0133] For example, consider the following situation: Figure 7 As shown, the low recognition rate condition consists of the text “bad weather day” and the label L1, which consists of key points as a specific pose.
[0134] also, Figure 7 The label L1 is a key point indicating that a person is standing side by side on the left side of the picture, and another person shorter than the first person is standing on the right side.
[0135] The low recognition rate data generation model 95 generates, for example, the low recognition rate data shown on the right, based on the condition of low recognition rate, and outputs it as learning data to the learning unit 96.
[0136] exist Figure 7 In the middle, the low-recognition-rate data generation model 95 generates images P1 to P3 of a bad weather day, where the person on the right is shorter than the person on the left, based on the low-recognition-rate condition consisting of the text "bad weather day" and the label L1, which is a key point representing a specific pose. At this time, the low-recognition-rate data generation model 95 also outputs the label L1 as a shared label for images P1 to P3. Furthermore, although... Figure 7 The example shown is of generating 3 images, but in reality, more images will be generated for learning purposes.
[0137] The learning unit 96 uses low-recognition-rate data, which is formed by merging images P1 to P3 with label L1, as learning data to enable the recognition model 93 to relearn. As a result, the recognition model 93 can improve the recognition success rate for images with low recognition-rate conditions, such as those consisting of text "bad weather day" and labels L1 consisting of key points representing specific postures.
[0138] (Second)
[0139] The above illustrates an example of a low recognition rate condition consisting of the text "bad weather day" and the label L1 consisting of keypoints representing a specific pose. However, a low recognition rate condition can also be the image itself where the recognition rate of recognition model 93 is below a threshold. Here, we consider a low recognition rate condition such as... Figure 8 The image shown is the case of image P11.
[0140] also, Figure 8 Image P11 is an image of a man standing side by side on the left and a woman on the right on a rainy day (a day with bad weather).
[0141] The low recognition rate data generation model 95 generates, for example, low recognition rate data as shown on the right, based on the image P11 which is the condition for low recognition rate, and outputs it as learning data to the learning unit 96.
[0142] exist Figure 8 In the process, the low-recognition-rate data generation model 95 generates images P21 and P22 based on the low-recognition-rate conditions formed by image P11, showing a man standing side-by-side on the left and a woman on the right on a rainy day and a snowy day (bad weather days), respectively. At this time, the low-recognition-rate data generation model 95 also outputs key points KP21 and KP22 as labels for images P21 and P22.
[0143] The learning unit 96 uses learning data, which combines images P21 and P22 with key points KP21 and KP22 used as labels, to enable the recognition model 93 to relearn. As a result, the recognition model 93 can improve the recognition success rate for images with low recognition rates, such as those composed of image P11, where the man stands side by side on the left and the woman on the right, and it is a bad weather day.
[0144] (Thirdly)
[0145] The above illustrates an example of generating an image and a label based on a low recognition rate condition.
[0146] However, it is also possible to generate an image along with multiple corresponding labels based on a low recognition rate.
[0147] For example, such as Figure 9 As shown, when the low recognition rate condition is "an image that only reflects one person" consisting of text, the low recognition rate data generation model 95 can generate low recognition rate data consisting of multiple labels for each image based on the text that is "an image that only reflects one person" as the low recognition rate condition, for example, together with the multiple images shown on the right, and output it to the learning unit 96 as learning data.
[0148] exist Figure 9 In this process, the low recognition rate data generation model 95 generates low recognition rate data consisting of images P31 to P33 that only show one person, based on the low recognition rate condition of the text "images showing only one person". At this time, the low recognition rate data generation model 95 also generates key points KP31 to KP33 as labels for images P31 to P33 and human body bounding boxes (boxes representing the area where the human body exists) BB31 to BB33 as low recognition rate data and outputs them.
[0149] The learning unit 96 uses the learning data formed by merging images P31 to P33, key points KP31 to KP33 as labels, and human bounding boxes BB31 to BB33 to enable the recognition model 93 to relearn.
[0150] The recognition model 93 not only learns the images P31 to P33 generated based on the condition of "only one person is shown" which is a low recognition rate condition, but also learns the labels formed by the corresponding key points KP31 to KP33 and the human body bounding boxes BB31 to BB33, thereby further improving the recognition success rate of "only one person is shown" which is a low recognition rate condition.
[0151] That is, for example, when the recognition model 93 is a model that infers key points from an image, as a preprocessing step for inferring key points, sometimes the bounding box of a human body is inferred from within the image to infer the key points of the pose of the person within the inferred bounding box of the human body.
[0152] In this case, by also using the correct position of the human bounding box needed to infer key points as a label for learning, key points can be inferred with a higher recognition rate.
[0153] (4)
[0154] The above illustrates an example of generating an RGB image based on a low recognition rate. However, images in various data formats can also be generated depending on the processing content of the recognition model and the data format of the image to be processed.
[0155] For example, Figure 10 As shown, when the low recognition rate condition is "images of two people reflected in a dark environment", the low recognition rate data generation model 95 can generate labels composed of key points for each image based on the text "images of two people reflected in a dark environment" as the low recognition rate condition, for example, together with images in multiple data formats as shown on the right, and output them as learning data to the learning unit 96.
[0156] exist Figure 10 In this process, the low-recognition-rate data generation model 95 generates, based on the low-recognition-rate condition of the text "an image of two people projected in a dark environment," an image P41 composed of RAW data and an image P42 composed of black and white image data, which are images of two people projected in a dark environment, as low-recognition-rate data. At the same time, the low-recognition-rate data generation model 95 also generates and outputs the key points KP41 and KP42 as labels for images P41 and P42 as low-recognition-rate data.
[0157] The learning unit 96 uses learning data, which is composed of image P41 made of RAW data, image P42 made of black and white image, and key points KP41 and KP42 as corresponding labels, to enable the recognition model 93 to relearn.
[0158] Recognition model 93 learns by using images P41 and P42, generated with different data formats based on the low recognition rate condition "images of two people reflected in a dark environment". Figure 10 In the example, it can improve the recognition success rate of "images of two people appearing in a dark environment" in at least RAW data and black and white image data.
[0159] Furthermore, when dealing with images in dark environments, such as when virtually attaching labels like key points to an image, the accuracy may not be sufficient due to the darkness. If the image is associated with the virtually attached labels for learning, the recognition success rate may not be adequately improved.
[0160] However, as Figure 10 As shown, the low recognition rate data generation model 95 adds key points to images of different data formats to generate low recognition rate data. Based on the generated low recognition rate data as learning data, the recognition model 93 learns, thereby improving the recognition success rate of the recognition model 93 in images in dark environments by combining key point labels.
[0161] <Improved Recognition Model Processing>
[0162] Next, refer to Figure 11 The flowchart illustrates... Figure 3 The recognition model improvement system 11 improves the recognition model processing.
[0163] In step S31, the recognition model improvement request unit 51 determines whether there is a request to improve the recognition model 52 due to a user's request. For example, when a user makes a request to improve the recognition model 52 by operating an operation panel (not shown), the recognition model improvement request unit 51 determines that there is a request to improve the recognition model 52 based on the operation signal corresponding to the operation content of the operation panel, and proceeds to step S32.
[0164] Furthermore, if at this time, information that can be considered as a low recognition rate condition is input from the user's haptic feedback when using the recognition model 52, or the statistical processing results of the recognition success rate of the recognition results of the recognition model 52, the input of information that can be considered as a low recognition rate condition will be accepted at the same time as the improvement request.
[0165] In step S32, the recognition model improvement request unit 51 sends an improvement request for the recognition model 52 to the cloud server 32 via the network 33. At this time, when there is information that can be regarded as a low recognition rate condition, the recognition model improvement request unit 51 also sends the information that can be regarded as a low recognition rate condition along with the improvement request to the cloud server 32.
[0166] In step S51, the condition generation unit 91 determines whether an improvement request for the recognition model 52 has been sent from the smartphone 31 via the network 33. If an improvement request for the recognition model 52 has been sent from the smartphone 31 in step S51, the process proceeds to step S52.
[0167] In step S52, the condition generation unit 91 generates various conditions required for generating verification data in the verification data generation model 92 and provides them to the verification data generation model 92. At this time, if information that can be considered a low recognition rate condition is sent along with an improvement request, the low recognition rate condition sent along with the improvement request can also be output to the verification data generation model 92.
[0168] In step S53, the verification data generation model 92 generates images and labels as verification data in pairs based on the conditions provided from the condition generation unit 91 and provides them to the recognition model 93.
[0169] In step S54, the recognition model 93 performs recognition processing based on the pairing of images and labels used as verification data, and outputs the recognition result and corresponding label to the recognition model evaluation unit 94.
[0170] In step S55, the recognition model evaluation unit 94 calculates the recognition success rate based on the recognition results and corresponding labels of the recognition model 93 to evaluate the recognition model 93 and generates an evaluation result composed of the recognition success rate.
[0171] In step S56, the recognition model evaluation unit 94 determines, based on the evaluation results, whether there are recognition results with a recognition success rate lower than the predetermined recognition rate.
[0172] In step S56, when there is a recognition result with a recognition success rate lower than the predetermined recognition rate, the process proceeds to step S57.
[0173] In step S57, the recognition model evaluation unit 94 determines the low recognition rate condition based on the recognition result which is lower than the predetermined recognition rate, and provides it to the low recognition rate data generation model 95.
[0174] In step S58, the low recognition rate data generation model 95 generates low recognition rate data consisting of images and labels based on the condition of low recognition rate, and outputs it to the learning unit 96 as learning data.
[0175] In step S59, the learning unit 96 uses the images and labels that constitute the low recognition rate data as learning data to enable the recognition model 93 to learn, and the process returns to step S54.
[0176] That is, in step S56, the processing of steps S54 to S59 is repeated until there are no more recognition results with a recognition success rate lower than the predetermined recognition rate, low recognition rate data is generated and the recognition model 93 is relearned repeatedly.
[0177] Then, in step S56, when there are no more recognition results with a recognition success rate lower than the predetermined recognition rate, the process proceeds to step S60.
[0178] In step S60, the recognition model update unit 72 provides the recognition model 93, which has been relearned and no longer has low recognition success rate, to the smartphone 31 via the network 33 for updating the recognition model 52 of the smartphone 31.
[0179] In step S33, the recognition model 52 of the smartphone 31 is updated from the recognition model 93 provided by the cloud server 32.
[0180] Furthermore, in step S31, if there is no request for improvement of the recognition model, the processing in steps S32 and S33 is skipped, and the process proceeds to step S34.
[0181] In addition, in step S51, if there is no request for improvement of the recognition model, the processing of steps S52 to S60 is skipped, and the process proceeds to step S61.
[0182] In steps S34 and S61, it is determined whether the process has been terminated. If the process has not been terminated, the process returns to steps S31 and S51 respectively, and the subsequent processes are repeated.
[0183] Then, in steps S34 and S61, the process ends when the end of the process is indicated.
[0184] Furthermore, in step S56, if there are no recognition results with a success rate lower than the predetermined recognition rate in the initial processing, no improvement is needed even if there is a request to improve the recognition model. Therefore, the smartphone 31 can be notified that no improvement is needed, and the processing in steps S60 and S33 can be skipped. Alternatively, the unimproved recognition model 93 can be provided to the smartphone 31 as is, and the recognition model 52 can be updated by the provided (unimproved) recognition model 93 while remaining in the unimproved state.
[0185] Through the above processing, for example, when the user feels that the recognition success rate is low during the process of using the recognition model 52, or when the actual recognition success rate is low, if the improvement request is provided to the cloud server 32, images and labels will be generated as verification data under various different conditions.
[0186] Next, the recognition process is performed by the recognition model 93 (which corresponds to the current recognition model 52) obtained using the generated verification data. The recognition success rate is calculated based on the recognition results. The condition of low recognition rate, which is lower than a predetermined threshold, is determined. Low recognition rate data consisting of images and labels under the determined low recognition rate condition is generated.
[0187] Furthermore, by repeatedly using low-recognition-rate data as learning data to relearn the recognition model 93 until the recognition success rate is higher than a predetermined value, the recognition success rate of the recognition model 93 can be improved, thereby achieving the improvement of the recognition model 52.
[0188] At this point, when generating verification data and low recognition rate data, multiple images and labels are generated in pairs and associated with each other, which can reduce the labor and cost of manually labeling each image one by one.
[0189] In addition, by reducing labor and costs, it is possible to generate a large number of images and labels under various conditions (scenes) as verification data. Therefore, it is possible to easily and appropriately determine the conditions with low recognition rates based on the recognition results of the recognition model when using images and labels under various conditions.
[0190] Furthermore, it is possible to generate a large amount of low-recognition-rate data consisting of images and labels under low-recognition-rate conditions as learning data. Therefore, it is possible to use more images and labels under low-recognition-rate conditions to enable the recognition model 93 to learn and appropriately improve the recognition success rate under low-recognition-rate conditions.
[0191] As a result, it is possible to easily determine the low recognition rate conditions, efficiently generate low recognition rate data based on the determined low recognition rate conditions, and then enable the recognition model 93 to learn on this basis. Therefore, it is possible to easily and appropriately improve the recognition success rate of images corresponding to the low recognition rate conditions.
[0192] Furthermore, the above illustrates the following example: The smartphone 31, which constitutes the recognition model improvement system 11, requests the cloud server 32 to improve the recognition rate of the recognition model 52 it carries. The cloud server 32 determines the low recognition rate conditions of the recognition model 93 corresponding to the recognition model 52, generates low recognition rate data based on the determined low recognition rate conditions, and uses it as learning data to enable the recognition model 93 to learn and update the recognition model 52 of the smartphone 31.
[0193] However, the recognition rate of the recognition model 52 can also be improved independently by equipping the smartphone 31 with a structure corresponding to the recognition model improvement processing unit 71 in the cloud server 32. That is, the smartphone 31 can also independently determine the low recognition rate conditions of the recognition model 52, and repeatedly perform the process of generating low recognition rate data and learning by the recognition model 52 based on the determined low recognition rate conditions, thereby improving the recognition rate of the recognition model 52.
[0194] <<3. Second Implementation Method>>
[0195] The above describes an example of generating verification data upon receiving an improvement request, using the verification data to determine low recognition rate conditions of the recognition model, generating low recognition rate data based on the low recognition rate conditions, and improving the recognition success rate by using the low recognition rate data as learning data for the recognition model.
[0196] However, it can also be done in the following way: when generating verification data periodically based on a large number of conditions required for evaluating the recognition model 93, there is no need to wait for improvement requests from the smartphone 31. Whenever the verification data is updated, the recognition model 93 is evaluated with the updated verification data to determine low recognition rate conditions, and low recognition rate data is generated for use as learning data, thereby enabling the recognition model 93 to learn and improve, and updating the recognition model 52 of the smartphone 31.
[0197] Figure 12 This is an example of the structure of the recognition model improvement system 11', in which a structure other than the cloud server 32 (not shown) periodically updates the verification data based on a large number of conditions required for the evaluation of the recognition model 93, evaluates the recognition model 93 with the updated verification data to determine low recognition rate conditions, generates low recognition rate data to be used as learning data to enable the recognition model 93 to learn and improve, and updates the recognition model 52 of the smartphone 31.
[0198] In addition, Figure 12 In the improved recognition model system 11', for those possessing the characteristics of... Figure 3 The identification model improvement system 11 has the same structure with the same attached figures and its description is omitted as appropriate.
[0199] That is, in Figure 12 In the improved recognition model system 11', with Figure 3 The different structure of the recognition model improvement system 11 is as follows: for example, the smartphone 31 becomes the smartphone 31', and in the cloud server 32, the recognition model improvement processing unit 71 becomes the recognition model improvement processing unit 71'.
[0200] In addition, the smartphone 31' adopts a structure that omits the recognition model improvement request unit 51 in the smartphone 31.
[0201] Furthermore, the identification model improvement processing unit 71' adopts a structure in which a verification data acquisition unit 131 is provided instead of the condition generation unit 91 and the verification data generation model 92 in the identification model improvement processing unit 71.
[0202] The verification data acquisition unit 131 acquires verification data 121 updated by a structure other than cloud server 32 (not shown) and provides it to the recognition model 93.
[0203] The verification data 121 is generated based on images and labels obtained from conditions reported by users of other smartphones 31' as having low recognition success rates, or low recognition rate conditions determined in other recognition model improvement processing units 71'. It is updated when new low recognition rate conditions are obtained from other smartphones 31' or other recognition model improvement processing units 71'. By using the verification data 121 to evaluate the recognition model 93, low recognition rate conditions for the recognition model 93 can be selectively determined.
[0204] Using this structure, whenever the verification data 121 is updated based on a large number of conditions required for evaluating the recognition model 93, the verification data acquisition unit 131 acquires the verification data 121 and provides it to the recognition model 93. By evaluating the recognition model 93, determining low recognition rate conditions, generating low recognition rate data, and using the low recognition rate data as learning data to learn the recognition model 93, the recognition success rate of the recognition model 93 can be improved efficiently.
[0205] In addition, regarding Figure 12 The identification model improvement process of the identification model improvement system 11' differs from the reference model improvement process in that the process of generating verification data whenever an improvement request is sent is changed to the process of obtaining verification data 121 whenever verification data 121 is updated. Figure 11 The process described in the flowchart is the same, so its description is omitted.
[0206] <<4. Third Implementation>>
[0207] The above illustrates an example of how the recognition model evaluation unit 94 calculates the recognition success rate by comparing the recognition result of the recognition model 93 with the label, and determines the low recognition rate condition based on the recognition success rate.
[0208] However, based on the recognition results and label information of recognition model 93, low recognition rate conditions can also be determined using, for example, large language models. Low recognition rate data generation model 95 generates low recognition rate data consisting of images and labels based on the determined low recognition rate conditions.
[0209] Figure 13For example, the structure of the recognition model improvement system 11” is to use large language models to determine the recognition model conditions with low recognition rate based on the recognition results of recognition model 93 and the information of the labels.
[0210] In addition, Figure 13 In the "Improved Recognition Model System 11", the system has the following characteristics: Figure 3 The same reference numerals are attached to the same structure as the identification model improvement system 11, and their descriptions are omitted as appropriate.
[0211] That is, in Figure 13 In the "Improved Recognition Model System 11", with Figure 3 The difference in structure of the recognition model improvement system 11 is as follows: In the cloud server 32, the recognition model improvement processing unit 71 is changed to "recognition model improvement processing unit 71".
[0212] In addition, the recognition model improvement processing unit 71” adopts a structure in which a recognition model evaluation unit 94 is replaced by a recognition model evaluation unit 94, which is provided with a recognition model evaluation LLM (Large Language Models) 151.
[0213] The recognition model evaluation LLM 151 adopts a structure that utilizes a large language model. Based on the recognition results of the recognition model 93 obtained using validation data and the information of the correct labels, it generates information for determining low recognition rate conditions and outputs it to the low recognition rate data generation model 95.
[0214] At this point, the recognition model evaluation using LLM 151 can generate not only information for determining low recognition rate conditions, but also information for determining low recognition rate data used as learning data, and the low recognition rate data generation model 95 should generate labels along with the images according to the low recognition rate conditions.
[0215] In addition, regarding Figure 13 The recognition model improvement process of the recognition model improvement system 11, except that the recognition model evaluation unit 94 is replaced by the recognition model evaluation unit 94 by the recognition model evaluation LLM 151 to evaluate the recognition model 93 and determine the low recognition rate conditions, is similar to the reference... Figure 11 The process described in the flowchart is the same, so its description is omitted.
[0216] <<5. Application Examples>>
[0217] (Used for contrastive learning)
[0218] There is a learning method called contrastive learning, which can be considered for application to the learning of recognition model 93. In general contrastive learning, data obtained by performing different data augmentations on the same image are used as positive pairs, and all other data are used as negative pairs for learning.
[0219] As a result, through contrastive learning, it is possible to learn that images similar to those augmented as positive data pairs are easily identified as originating from the same image, while images similar to those augmented as negative data pairs are easily identified as originating from different images.
[0220] Therefore, by using contrastive learning in the learning of models that can generate images and labels based on conditions, such as the verification data generation model 92 and the low recognition rate data generation model 95 disclosed herein, more targeted learning for downstream tasks is possible.
[0221] For example, consider the application of contrastive learning in the learning process that enables recognition model 93 to infer key points (human pose). In this case, for example, low-recognition-rate data generates model 95, such as... Figure 14 As shown, multiple conditions are generated, consisting of various correct labels from low recognition rate conditions, and images matching those labels.
[0222] Figure 14 The following example is shown: Low recognition rate data generation model 95 generates corresponding images P51 to P53 and images P61 to P63 for key points KP51 and KP61 of a human body pose as a condition of low recognition rate.
[0223] At this point, the learning unit 96 enables comparative learning by using images with the same correct label as positive sample pairs and other images as negative sample pairs, thereby potentially enabling learning to infer human posture based on human posture.
[0224] That is, for the case where images P51 to P53 and images P61 to P63 are treated as a group of images, by setting positive and negative sample pairs according to human pose, it is expected to achieve learning for inferences corresponding to human pose.
[0225] More specifically, in Figure 14 In the case of P51 to P53 generated based on the same key point KP51, and P61 to P63 generated based on the same key point KP61, each are used as positive sample pairs for learning.
[0226] In addition, images P51 to P53 generated based on the same keypoint KP51 and images P61 to P63 generated based on the same keypoint KP61 are used as negative sample pairs for learning.
[0227] As a result, images P51 to P53 and images P61 to P63 are only learned as a set of images based on the condition of a person's human posture. However, by comparing and learning the images generated for the key points KP51 and KP61, which are labels composed of different human postures, with the corresponding labels, it is hoped that learning for human posture inference can be carried out.
[0228] <<6. Examples of Software Execution>>
[0229] Furthermore, the aforementioned series of processes can be executed by either hardware or software. When the series of processes are executed by software, the program constituting the software is installed from a recording medium into a computer embedded in dedicated hardware, or, for example, into a general-purpose computer capable of performing various functions by installing various programs.
[0230] Figure 15 This diagram illustrates the structure of a general-purpose computer. The computer has a built-in CPU (Central Processing Unit) 1001. Input / output interfaces 1005 are connected to the CPU 1001 via a bus 1004. A ROM (Read Only Memory) 1002 and a RAM (Random Access Memory) 1003 are connected to the bus 1004.
[0231] The input / output interface 1005 is connected to an input unit 1006, an output unit 1007, a storage unit 1008, and a communication unit 1009. The input unit 1006 includes input devices such as a keyboard and mouse for users to input operation commands. The output unit 1007 outputs the processing screen and the image of the processing result to a display device. The storage unit 1008 includes a hard disk drive for storing programs and various data. The communication unit 1009 includes a LAN (Local Area Network) adapter and performs communication processing via a network such as the Internet. A drive 1010 is also connected, which reads and writes data to removable storage media 1011 such as disks (including floppy disks), optical disks (including CD-ROMs (Compact Disc-Read Only Memory) and DVDs (Digital Versatile Discs)), magneto-optical disks (including Mini Discs (MDs)), or semiconductor memory.
[0232] The CPU 1001 performs various processes according to programs stored in the ROM 1002, or programs read from a removable storage medium 1011 such as a disk, optical disk, magneto-optical disk, or semiconductor memory and installed into the storage unit 1008, and loaded from the storage unit 1008 into the RAM 1003. Additionally, the RAM 1003 also stores, as appropriate, data required by the CPU 1001 during the execution of various processes.
[0233] In the computer configured as described above, the CPU 1001 loads the program stored in the storage unit 1008 into the RAM 1003 and executes it via, for example, the input / output interface 1005 and the bus 1004, thereby performing the series of processes described above.
[0234] The program executed by the computer (CPU 1001) can be provided, for example, by recording it on a removable storage medium 1011, such as a packaging medium. Alternatively, the program can be provided via wired or wireless transmission media such as a local area network, the Internet, or digital satellite broadcasting.
[0235] In a computer, by installing the removable storage medium 1011 onto the drive 1010, a program can be installed into the storage unit 1008 via the input / output interface 1005. Alternatively, the program can be received by the communication unit 1009 and installed into the storage unit 1008 via a wired or wireless transmission medium. Furthermore, the program can be pre-installed in the ROM 1002 or the storage unit 1008.
[0236] Furthermore, the programs executed by the computer can be either programs that are processed sequentially in the order described in this specification, or programs that are processed in parallel or at timed intervals as required, such as when called.
[0237] also, Figure 15 CPU 1001 in Figure 3 Recognition Model Improvement Processing Unit 71 Figure 12 The recognition model improvement processing unit 71' and Figure 13 The function of the recognition model improvement processing unit 71 was implemented.
[0238] Furthermore, in this specification, a system refers to a collection of multiple constituent elements (devices, modules (components), etc.), regardless of whether all constituent elements are in the same housing. Therefore, multiple devices housed in a single housing and connected via a network, as well as a device housing multiple modules in a single housing, are both systems.
[0239] Furthermore, the embodiments of this disclosure are not limited to the above-described embodiments, and various modifications can be made without departing from the spirit of this disclosure.
[0240] For example, this disclosure may employ a cloud computing architecture where multiple devices share and jointly process a function via a network.
[0241] In addition, the steps described in the flowchart above can be performed by multiple devices in addition to being executed by one device.
[0242] Furthermore, when a step includes multiple processes, these processes can be performed by multiple devices in addition to being executed by one device.
[0243] In addition, this disclosure can also adopt the following structure.
[0244] <1> An information processing device, comprising:
[0245] The evaluation department evaluates the recognition rate of the recognition model on verification data including images and labels generated based on multiple conditions, and determines the conditions in the recognition model where the recognition rate is lower than a predetermined recognition rate; and
[0246] A low recognition rate data generation model generates low recognition rate data of the recognition model, including the image and the label, based on the determined low recognition rate conditions.
[0247] <2> According to the information processing device described in <1>, wherein,
[0248] The low recognition rate data generation model uses a speculative model to recursively repeat the process of setting noise intensity that meets the conditions to speculate on the noise, thereby generating a pair of images and labels that meet the conditions. The speculative model defines different noise intensity for each of the images, text and labels, and infers the noise assigned to each of the images, text and labels based on the image feature groups, text feature groups and label feature groups that are assigned noise according to the noise intensity.
[0249] <3> According to the information processing device described in <2>, wherein,
[0250] The inference model is a structure obtained by adding an additional inference model to the basic inference model, wherein...
[0251] The basic inference model defines different noise intensities for the image and the text respectively, and infers the noise assigned to the image and the text respectively based on the image feature groups and text feature groups that are assigned noise according to the noise intensities.
[0252] The additional inference model defines different noise intensities for the image and the label respectively, and infers the noise assigned to the image and the label respectively based on the image feature group and label feature group that are assigned the noise according to the noise intensities.
[0253] <4> The information processing device according to <3>, wherein,
[0254] The low recognition rate data generation model is an extension of UniDiffuser.
[0255] The UniDiffuser uses the basic inference model to recursively repeat the process of assigning intensity to noise that meets the conditions in order to infer the noise, thereby generating a pair of images and text that meet the conditions.
[0256] <5> According to the information processing device described in <1>, wherein,
[0257] The condition for low recognition rate is at least one of the following: an image with low recognition rate in the recognition model, text representing the condition for generating an image with low recognition rate, and a label that serves as the condition for generating an image with low recognition rate.
[0258] <6> According to the information processing apparatus described in <1>, wherein,
[0259] The low recognition rate data generation model generates the image and corresponding multiple labels based on the low recognition rate condition.
[0260] <7> The information processing apparatus according to <1> further includes:
[0261] The learning department uses the low recognition rate data as learning data to enable the recognition model to learn.
[0262] <8> The information processing apparatus according to <7>, wherein,
[0263] The evaluation unit determines the conditions for the low recognition rate in the recognition model obtained by the learning unit.
[0264] The low recognition rate data generation model generates low recognition rate data again, including the image and the label, based on the low recognition rate conditions in the recognition model determined by the evaluation unit and learned by the learning unit.
[0265] The learning unit uses the low recognition rate data regenerated by the low recognition rate data generation model as learning data to enable the recognition model to learn.
[0266] <9> According to the information processing apparatus described in <8>, wherein,
[0267] The evaluation unit determines the condition of low recognition rate, the low recognition rate data is regenerated by the low recognition rate data generation model, and the learning unit enables the recognition model to learn. This process is repeated until the evaluation unit can no longer determine the condition of low recognition rate.
[0268] <10> The information processing apparatus according to <7>, wherein,
[0269] The learning unit uses comparative learning, which uses the low-recognition-rate data as learning data, to enable the recognition model to learn.
[0270] <11> The information processing apparatus according to <1> further includes:
[0271] The condition generation unit generates the plurality of conditions; and
[0272] The verification data generation model generates verification data, including the pairing of the image and the label, based on the plurality of conditions generated by the condition generation unit.
[0273] <12> According to the information processing apparatus described in <11>, wherein,
[0274] The verification data generation model uses a speculative model to recursively repeat the process of setting the noise intensity that meets the conditions to speculate on the noise, thereby generating a pair of images and labels that meet the conditions. The speculative model defines different noise intensity for each of the images, text and labels, and speculates on the noise assigned to each of the images, text and labels based on the image feature groups, text feature groups and label feature groups that are assigned noise according to the noise intensity.
[0275] <13> According to the information processing apparatus described in <1>, wherein,
[0276] The verification data includes a pairing of the image and the label generated in another information processing device equipped with the evaluation unit and the low recognition rate data generation model, based on the condition of the low recognition rate.
[0277] <14> The information processing apparatus according to any one of <1> to <13>, wherein,
[0278] The evaluation department consists of a large language model.
[0279] <15> The information processing apparatus according to any one of <1> to <14>, wherein,
[0280] The label includes at least one of the following: bounding box (face, human body, etc.) corresponding to the image, key points (human body model), and semantic segmentation.
[0281] <16> An information processing method, comprising the following steps:
[0282] The evaluation method assesses the recognition rate of the recognition model on verification data including images and labels generated based on multiple conditions, and determines the conditions in the recognition model where the recognition rate is lower than a predetermined recognition rate; and
[0283] Based on the determined low recognition rate conditions, low recognition rate data of the recognition model including the image and the label is generated.
[0284] <17> A program that enables a computer to function as a unit:
[0285] The evaluation department evaluates the recognition rate of the recognition model on verification data including images and labels generated based on multiple conditions, and determines the conditions in the recognition model where the recognition rate is lower than a predetermined recognition rate; and
[0286] A low recognition rate data generation model generates low recognition rate data of the recognition model, including the image and the label, based on the determined low recognition rate conditions.
[0287] Figure Labels
[0288] 11, 11', 11'': Recognition model improvement system; 31, 31': Smartphone; 32: Cloud server; 51: Recognition model improvement request unit; 52: Recognition model; 71, 71': Recognition model improvement processing unit; 72: Recognition model update unit; 91: Condition generation unit; 92: Model generation for validation data; 93: Recognition model; 94: Recognition model evaluation unit; 95: Model generation for low recognition rate data; 96: Learning unit; 121: Validation data; 131: Validation data acquisition unit; 151: LLM (Large Language Model) for recognition model evaluation.
Claims
1. An information processing device, comprising: The evaluation department evaluates the recognition rate of the recognition model on verification data including images and labels generated based on multiple conditions, and determines the conditions in the recognition model where the recognition rate is lower than a predetermined recognition rate; and A low recognition rate data generation model generates low recognition rate data of the recognition model, including the image and the label, based on the determined low recognition rate conditions.
2. The information processing apparatus according to claim 1, wherein, The low recognition rate data generation model uses a speculative model to recursively repeat the process of setting noise intensity that meets the conditions to speculate on the noise, thereby generating a pair of images and labels that meet the conditions. The speculative model defines different noise intensity for each of the images, text and labels, and infers the noise assigned to each of the images, text and labels based on the image feature groups, text feature groups and label feature groups that are assigned noise according to the noise intensity.
3. The information processing apparatus according to claim 2, wherein, The inference model is a structure obtained by adding an additional inference model to the basic inference model, wherein... The basic inference model defines different noise intensities for the image and the text respectively, and infers the noise assigned to the image and the text respectively based on the image feature groups and text feature groups that are assigned noise according to the noise intensities. The additional inference model defines different noise intensities for the image and the label respectively, and infers the noise assigned to the image and the label respectively based on the image feature group and label feature group that are assigned the noise according to the noise intensities.
4. The information processing apparatus according to claim 3, wherein, The low recognition rate data generation model is an extension of UniDiffuser. The UniDiffuser uses the basic inference model to recursively repeat the process of assigning intensity to noise that meets the conditions in order to infer the noise, thereby generating a pair of images and text that meet the conditions.
5. The information processing apparatus according to claim 1, wherein, The condition for low recognition rate is at least one of the following: an image with low recognition rate in the recognition model, text representing the condition for generating an image with low recognition rate, and a label that serves as the condition for generating an image with low recognition rate.
6. The information processing apparatus according to claim 1, wherein, The low recognition rate data generation model generates the image and corresponding multiple labels based on the low recognition rate condition.
7. The information processing apparatus according to claim 1, further comprising: The learning department uses the low recognition rate data as learning data to enable the recognition model to learn.
8. The information processing apparatus according to claim 7, wherein, The evaluation unit determines the conditions for the low recognition rate in the recognition model obtained by the learning unit. The low recognition rate data generation model generates low recognition rate data again, including the image and the label, based on the low recognition rate conditions in the recognition model determined by the evaluation unit and learned by the learning unit. The learning unit uses the low recognition rate data regenerated by the low recognition rate data generation model as learning data to enable the recognition model to learn.
9. The information processing apparatus according to claim 8, wherein, The evaluation unit determines the condition of low recognition rate, the low recognition rate data is regenerated by the low recognition rate data generation model, and the learning unit enables the recognition model to learn. This process is repeated until the evaluation unit can no longer determine the condition of low recognition rate.
10. The information processing apparatus according to claim 7, wherein, The learning unit uses comparative learning, which uses the low-recognition-rate data as learning data, to enable the recognition model to learn.
11. The information processing apparatus according to claim 1, further comprising: The condition generation unit generates the multiple conditions; as well as The verification data generation model generates verification data, including the pairing of the image and the label, based on the plurality of conditions generated by the condition generation unit.
12. The information processing apparatus according to claim 11, wherein, The verification data generation model uses a speculative model to recursively repeat the process of setting the noise intensity that meets the conditions to speculate on the noise, thereby generating a pair of images and labels that meet the conditions. The speculative model defines different noise intensity for each of the images, text and labels, and speculates on the noise assigned to each of the images, text and labels based on the image feature groups, text feature groups and label feature groups that are assigned noise according to the noise intensity.
13. The information processing apparatus according to claim 1, wherein, The verification data includes a pairing of the image and the label generated in another information processing device equipped with the evaluation unit and the low recognition rate data generation model, based on the condition of the low recognition rate.
14. The information processing apparatus according to claim 1, wherein, The evaluation department includes a large language model.
15. The information processing apparatus according to claim 1, wherein, The label includes at least one of the following: bounding box (face, human body, etc.) corresponding to the image, key points (human body model), and semantic segmentation.
16. An information processing method, comprising the following steps: The evaluation method assesses the recognition rate of the recognition model on verification data including images and labels generated based on multiple conditions, and determines the conditions in the recognition model where the recognition rate is lower than a predetermined recognition rate; and Based on the determined low recognition rate conditions, low recognition rate data of the recognition model including the image and the label is generated.
17. A program that enables a computer to function as a unit: The evaluation department evaluates the recognition rate of the recognition model on verification data including images and labels generated based on multiple conditions, and determines the conditions in the recognition model where the recognition rate is lower than a predetermined recognition rate; and A low recognition rate data generation model generates low recognition rate data of the recognition model, including the image and the label, based on the determined low recognition rate conditions.
Citation Information
Patent Citations
Learning server, and assist system
JP2019021201A
Information processing system, information processing method, and program
JP2019109924A