Learning device, estimation device, learning method, estimation method, and computer program

By converting video data into pixel-based learning images and preprocessing, the computational cost and training time for estimating stroke order in paintings are reduced, enabling efficient estimation without high-performance computers.

JP2026006030APending Publication Date: 2026-01-16TOPPAN HOLDINGS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024104751
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-28
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing technologies for estimating the stroke order of paintings using machine learning are computationally expensive and require high-performance computers due to the processing of video data.

Method used

A learning device that converts video data into pixel-based learning images and performs preprocessing to generate training data, reducing computational cost by using these images for machine learning to estimate stroke order, and an estimation device that uses a trained model to output the stroke order as a pixel-based image or video.

Benefits of technology

Reduces computational cost and training time, allowing for efficient estimation of stroke order without the need for high-performance computers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026006030000001_ABST
    Figure 2026006030000001_ABST
Patent Text Reader

Abstract

To reduce a calculation cost in machine learning for estimating a stroke order of drawing a picture.SOLUTION: A learning device includes a moving image analysis unit that acquires pixel coordinates and a moving image time at which a colored pixel is detected from a moving image obtained by capturing a painting drawn by a person, a learning image generation unit that converts the moving image time into a pixel value of the pixel coordinates and generates a learning image using the pixel value for each pixel coordinate, and a learning unit that performs machine learning of a learning model so as to generate an output image representing a drawing order of the painting using the learning image and an image of a drawing result of the moving image as input data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a learning device, an estimation device, a learning method, an estimation method, and a computer program. [Background technology]

[0002] Patent Document 1 describes a technology that acquires a video of a user's handwriting when the user inputs handwriting, processes the acquired video using a neural network to detect multiple ink strokes, and recognizes the user's handwriting in real time using the detected multiple ink strokes. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2020-9442 Summary of the Invention [Problem to be solved by the invention]

[0004] Understanding the process of painting is gaining attention as a way to pass on painting techniques and add value to painting exhibitions. Traditionally, identifying the stroke order of a painting's creation process has been done primarily through visual inspection by experts, but this can leave the results dependent on the individual's ability. For this reason, it is possible to use artificial intelligence (AI) to estimate the stroke order of a painting. In this case, for example, a model for estimating the stroke order of a painting can be generated by performing machine learning using video footage of people actually painting as training data.

[0005] However, the technology described in the above-mentioned Patent Document 1 uses a neural network to process a user's handwritten video, but machine learning of video (moving images) is computationally expensive, so learning takes a long time and requires a high-performance computer.

[0006] The present invention has been made in consideration of the above circumstances, and its purpose is to reduce the computational cost in machine learning for estimating the stroke order of paintings. [Means for solving the problem]

[0007] One aspect of the present invention is a learning device that includes a video analysis unit that acquires pixel coordinates and video time when colored pixels are detected from a video captured of a painting being painted by a person; a learning image generation unit that converts the video time into pixel values ​​of the pixel coordinates and generates a learning image for each of the pixel coordinates using the pixel values; and a learning unit that uses the learning images and images resulting from the drawing of the video as input data to perform machine learning on a learning model so as to generate an output image that represents the stroke order of the painting. One aspect of the present invention is the learning device described above, wherein the output image has a pixel value representing the drawing time when each pixel is colored. One aspect of the present invention is a learning device, further comprising a stroke order video generation unit that generates a video in which each drawing pixel is colored at the drawing time from the output image and the image of the drawing result of the video. One aspect of the present invention is a learning device, further comprising a pre-processing unit that performs at least one of the following processing on the learning image and the image resulting from the video: rotation, enlargement, movement, brightness adjustment, and RGB channel change, and the learning unit further uses the results of the processing as input data.

[0008] One aspect of the present invention is an estimation device comprising a stroke order estimation unit that acquires pixel coordinates and video time at which painted pixels are detected from a video captured of a painting being painted by a person, converts the video time into pixel values ​​of the pixel coordinates, generates a training image using the pixel values ​​for each of the pixel coordinates, inputs an image of a painting for which stroke order estimation is to be estimated to a trained model that has undergone machine learning to generate an output image representing the stroke order of the painting using the training image and an image of the drawing result of the video as input data, and acquires an output image representing the stroke order of the painting for which stroke order estimation is to be estimated from the trained model. One aspect of the present invention is the estimation device described above, wherein the output image has a pixel value representing a drawing time when each drawing pixel is colored. One aspect of the present invention is the above-mentioned estimation device, further comprising a stroke order video generation unit that generates a video in which each drawing pixel is colored at the drawing time from the output image and an image of a painting that is a target for stroke order estimation.

[0009] One aspect of the present invention is a learning method executed by an information processing device, the learning method including: a video analysis step of acquiring pixel coordinates and video time of colored pixels detected from a video of a painting being painted by a person; a learning image generation step of converting the video time into pixel values ​​of the pixel coordinates and generating a learning image for each pixel coordinate using the pixel values; and a learning step of performing machine learning on a learning model using the learning images and images resulting from the drawing of the video as input data to generate an output image representing the stroke order of the painting.

[0010] One aspect of the present invention is an estimation method executed by an information processing device, the estimation method including a stroke order estimation step of acquiring pixel coordinates and video time at which colored pixels are detected from a video captured of a painting being painted by a person, converting the video time into pixel values ​​of the pixel coordinates, generating a training image using the pixel values ​​for each of the pixel coordinates, inputting an image of the painting for which stroke order estimation is to be estimated to a trained model that has been machine-learned to generate an output image representing the stroke order of the painting using the training image and an image of the drawing result of the video as input data, and acquiring an output image representing the stroke order of the painting for which stroke order estimation is to be estimated from the trained model.

[0011] One aspect of the present invention is a computer program for causing a computer to execute the following steps: a video analysis step of acquiring pixel coordinates and video time at which colored pixels are detected from a video of a painting being painted by a person; a training image generation step of converting the video time into pixel values ​​of the pixel coordinates and generating a training image for each pixel coordinate using the pixel values; and a learning step of performing machine learning on a learning model using the training images and images resulting from the video as input data to generate an output image representing the stroke order of the painting.

[0012] One aspect of the present invention is a computer program for causing a computer to execute a stroke order estimation step of acquiring pixel coordinates and video time at which painted pixels are detected from a video captured of a painting being painted by a person, converting the video time into pixel values ​​of the pixel coordinates, generating a training image using the pixel values ​​for each of the pixel coordinates, inputting an image of a painting for which stroke order estimation is to be estimated to a trained model that has been machine-learned to generate an output image representing the stroke order of the painting using the training image and an image of the drawing result of the video as input data, and acquiring an output image representing the stroke order of the painting for which stroke order estimation is to be estimated from the trained model. [Effects of the Invention]

[0013] According to the present invention, it is possible to obtain an effect of reducing the computational cost in machine learning for estimating the stroke order of a painting. [Brief explanation of the drawings]

[0014] [Figure 1] FIG. 1 is a block diagram illustrating an example of the configuration of a stroke order estimation system according to an embodiment. [Figure 2] 10 is a flowchart illustrating an example of a procedure of a learning method according to an embodiment. [Figure 3] FIG. 10 is a diagram illustrating learning data according to an embodiment. [Figure 4] 10 is a flowchart illustrating an example of a procedure of an estimation method according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0015] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. Fig. 1 is a block diagram showing an example of the configuration of a stroke order estimation system 1 according to an embodiment. In Fig. 1, the stroke order estimation system 1 includes a learning device 10 and an estimation device 20. The learning device 10 and the estimation device 20 may transmit and receive data online via communication, or may input and output data offline.

[0016] Video data A is input to the learning device 10. The video data A is data of a video of a person painting a picture. An example of a painting is a sumi-e painting. The painting may also be an oil painting.

[0017] One example of a method for acquiring video data A is to capture a photograph from behind the drawing surface of a piece of Japanese paper, such as calligraphy paper, on which the ink painting is drawn. This makes it possible to acquire video data A that captures the state of the ink painting being drawn as seen through the back of the drawing surface of the Japanese paper. In the case of an oil painting, by having the artist draw the oil painting on a transparent acrylic board, it is possible to acquire video data A that captures the state of the oil painting being drawn as seen through the back of the drawing surface of the transparent acrylic board, just like in the case of an ink painting.

[0018] The learning device 10 uses video data A to generate a model for estimating the order of strokes in a painting through machine learning. The model (trained model) MA generated by the learning device 10 is supplied to the estimation device 20. The estimation device 20 uses the trained model MA to estimate the order of strokes in the painting, which is the target of stroke order estimation, from an image G of the painting, which is the target of stroke order estimation (target painting image).

[0019] [Learning device] In FIG. 1, the learning device 10 includes a receiving unit 11, a video analysis unit 12, a learning image generation unit 13, a preprocessing unit 14, a learning unit 15, a stroke order video generation unit 16, and an output unit 17.

[0020] Each function of learning device 10 is realized by learning device 10 having computer hardware such as a CPU (Central Processing Unit) and memory, and by the CPU executing a computer program stored in the memory. Learning device 10 may be configured using a general-purpose computer device or a dedicated hardware device. For example, learning device 10 may be configured using a server computer connected to a communication network such as the Internet. Each function of learning device 10 may also be realized by cloud computing. Learning device 10 may also be realized by a single computer, or by distributing the functions of learning device 10 across multiple computers. Learning device 10 may also be configured to have a website established using, for example, a WWW system.

[0021] The receiving unit 11 receives the input of the video data A.

[0022] The video analysis unit 12 acquires pixel coordinates and video time at which painted pixels are detected from the video data A. If the video data A is video data obtained by filming the drawing of a watercolor painting, and if the video analysis unit 12 detects a pixel painted black in the video data A, it acquires the coordinates (pixel coordinates) of the painted pixel and the frame time (video time) at which the painted pixel was detected. If the video data A is video data obtained by filming the drawing of an oil painting, and if the video analysis unit 12 detects a pixel painted in any color in the video data A, it acquires the coordinates (pixel coordinates) of the painted pixel and the frame time (video time) at which the painted pixel was detected. The video analysis unit 12 executes the process of acquiring pixel coordinates and video time from the first frame to the last frame of the video data A.

[0023] The training image generation unit 13 converts the video time into pixel values ​​of pixel coordinates and generates a training image using the pixel values ​​for each pixel coordinate. The training image becomes a single still image in which the pixel values ​​of colored pixels in the video data A represent the frame times (video times) at which they were colored. In this way, the video data A is visualized into a single training image.

[0024] The learning image and the image resulting from drawing the video data A are used as learning data in the machine learning of a model (learning model) MB that estimates the stroke order of a painting. For example, the last frame (final frame) of the video data A may be used as the image resulting from drawing the video data A.

[0025] When the amount of training data is insufficient, the preprocessing unit 14 performs preprocessing to increase the amount of training data. Specifically, the preprocessing unit 14 performs at least one of processing operations of rotation, enlargement, and movement on the training image and the image (e.g., the final frame) of the rendering result of the video data A.

[0026] For example, the preprocessing unit 14 performs a rotation process to rotate the training images and the final frame of the video data A by a predetermined rotation amount. As a result, two data sets are obtained as training data: one data set consisting of the original training images before rotation and the final frame of the video data A, and another data set consisting of the training images after rotation and the final frame of the video data A. Furthermore, by performing a rotation process for each of a plurality of different rotation amounts, data sets consisting of the training images after rotation and the final frame of the video data A for a plurality of different rotation amounts may be generated.

[0027] For example, the preprocessing unit 14 performs an enlargement process to enlarge the training image and the final frame of the video data A by a predetermined enlargement amount. As a result, two data sets are obtained as training data: one data set consisting of the original training image before enlargement and the final frame of the video data A, and another data set consisting of the enlarged training image and the final frame of the video data A. Furthermore, by performing an enlargement process for each of a plurality of different enlargement amounts, data sets consisting of the enlarged training image and the final frame of the video data A for a plurality of different enlargement amounts may be generated.

[0028] For example, the preprocessing unit 14 performs a movement process that moves pixels within an image by a predetermined movement amount between the training image and the final frame of video data A. As a result, two data sets are obtained as training data: one data set consisting of the original training image before movement and the final frame of video data A, and another data set consisting of the training image after movement and the final frame of video data A. Furthermore, by performing movement processes for each of a plurality of different movement amounts, data sets consisting of the training image after movement and the final frame of video data A for a plurality of different movement amounts may be generated.

[0029] Note that rotation, enlargement, and translation may be performed independently, or a combination of these may be performed. Furthermore, if the painting is an oil painting, preprocessing may include brightness adjustment and RGB channel change in addition to rotation, enlargement, and translation. The preprocessing unit 14 may be configured so that the user can set what preprocessing to perform. Furthermore, the preprocessing unit 14 may also be configured so that the user can set whether or not to perform preprocessing.

[0030] The learning unit 15 uses the learning image and the image of the drawing result of the video data A as input data to perform machine learning of the learning model MB so as to generate an output image representing the stroke order of drawing the picture. The learning model MB is, for example, a neural network. Alternatively, a deep neural network may be used as the learning model MB to perform deep learning.

[0031] The input data for machine learning of the learning model MB may be only one dataset consisting of the original training images before preprocessing by the preprocessing unit 14 and the images resulting from rendering of the video data A. Alternatively, the input data for machine learning of the learning model MB may be one dataset consisting of the original training images before preprocessing by the preprocessing unit 14 and the images resulting from rendering of the video data A, and one or more datasets consisting of the training images after preprocessing by the preprocessing unit 14 and the images resulting from rendering of the video data A. The configuration is such that the user can set what dataset to use as input data for machine learning of the learning model MB.

[0032] As an example of this embodiment, the output image generated by the learning model MB through machine learning is an image (estimated stroke order time image) having pixel values ​​representing the drawing time at which each drawn pixel (drawing pixel) was colored. The drawing time is a relative time representing the drawing order between the drawing pixels. According to the estimated stroke order time image, the drawing time at which each drawing pixel was colored is indicated by the pixel value of each drawing pixel. Therefore, the stroke order of the drawing of a painting is represented by a single estimated stroke order time image, which is extremely efficient.

[0033] The stroke order video generation unit 16 generates a video (stroke order estimated playback video) in which each drawing pixel is colored at the drawing time from the stroke order estimated time image and the image of the drawing result of video data A. Specifically, the stroke order video generation unit 16 generates a video (stroke order estimated playback video) in which, for each drawing pixel of the stroke order estimated time image, the drawing pixel of the image of the drawing result of video data A is lit at the drawing time corresponding to the pixel value of the drawing pixel. With the stroke order estimated playback video, the estimated result of the stroke order of the drawing of the picture is played back as a video, so that the user can easily visually recognize the estimated result of the stroke order of the drawing of the picture.

[0034] The output unit 17 outputs the stroke order estimation playback video generated by the stroke order video generation unit 16, the trained model MA generated by the learning unit 15, etc. The trained model MA can be output, for example, by writing it to a computer-readable recording medium or by transmitting the data via communication. The trained model MA output by the output unit 17 is supplied to the estimation device 20.

[0035] Note that stroke order video generation unit 16 is not essential, and learning device 10 does not have to include stroke order video generation unit 16. If learning device 10 does not include stroke order video generation unit 16, output unit 17 outputs a stroke order estimated time image. The stroke order estimated time image allows the user to recognize the estimated stroke order for drawing a picture.

[0036] [Estimation device] In FIG. 1, the estimation device 20 includes a receiving unit 21, a stroke order estimation unit 22, a stroke order video generation unit 23, and an output unit 24.

[0037] Each function of the estimation device 20 is realized by the estimation device 20 including computer hardware such as a CPU and a memory, and the CPU executing a computer program stored in the memory. The estimation device 20 may be configured using a general-purpose computer device or a dedicated hardware device. For example, the estimation device 20 may be configured using a server computer connected to a communication network such as the Internet. Each function of the estimation device 20 may be realized by cloud computing. The estimation device 20 may be realized by a single computer, or may be realized by distributing the functions of the estimation device 20 across multiple computers. The estimation device 20 may also be configured to have a website set up using, for example, a WWW system.

[0038] The receiving unit 21 receives input of an image G of a painting (target painting image) for which the stroke order is to be estimated.

[0039] The stroke order estimation unit 22 includes a trained model MA generated by the learning device 10. The stroke order estimation unit 22 inputs a target painting image G to the trained model MA and obtains an output image representing the stroke order of the painting that is the target for stroke order estimation from the trained model MA. When the target painting image G is input, the trained model MA outputs an output image representing the stroke order of the painting in the target painting image G.

[0040] As an example of this embodiment, the output image generated by the trained model MA is an image (estimated stroke order time image) having pixel values ​​representing the drawing time at which each drawn pixel (drawing pixel) was colored. The drawing time is a relative time representing the drawing order between the drawing pixels. According to the estimated stroke order time image, the drawing time at which each drawing pixel was colored is indicated by the pixel value of each drawing pixel. Therefore, the stroke order of the drawing of a painting is represented by a single estimated stroke order time image, which is extremely efficient.

[0041] The stroke order video generation unit 23 generates a video (stroke order estimation playback video) in which each drawing pixel is colored at the drawing time from the stroke order estimation time image and the target painting image G. Specifically, the stroke order video generation unit 23 generates a video (stroke order estimation playback video) in which, for each drawing pixel in the stroke order estimation time image, the drawing pixel of the target painting image G is lit at the drawing time corresponding to the pixel value of the drawing pixel. With the stroke order estimation playback video, the estimated result of the stroke order of the painting is played back as a video, so that the user can easily visually recognize the estimated result of the stroke order of the painting.

[0042] The output unit 24 outputs an estimation result of the stroke order of the painting of the target painting image G (target painting stroke order estimation result). The target painting stroke order estimation result may be, for example, an output image of the trained model MA. The target painting stroke order estimation result may be, for example, a stroke order estimation time image that is an output image of the trained model MA. The target painting stroke order estimation result may be, for example, a stroke order estimation playback video generated from the stroke order estimation time image that is an output image of the trained model MA.

[0043] Examples of methods for outputting the stroke order estimation result of the target painting include displaying it on a display screen, printing it on a print medium, writing it on a computer-readable recording medium, and transmitting the data via communication. Note that the output unit 24 may simply write the stroke order estimation result of the target painting to a memory or recording medium built into the estimation device 20 without outputting it to an external device of the estimation device 20. The stroke order estimation result of the target painting written to a memory or recording medium built into the estimation device 20 may be configured to be accessible from a device external to the estimation device 20.

[0044] Next, the operation of the stroke order estimating system 1 according to this embodiment will be described with reference to FIGS.

[0045] <Learning stage> Fig. 2 is a flowchart showing an example of the procedure of the learning method according to this embodiment. The learning stage S100 according to this embodiment will be described with reference to Fig. 2. The learning stage S100 is a stage in which machine learning is performed to generate a trained model MA to be used in the estimation stage S200, which will be described later with reference to Fig. 4. The learning stage S100 is a stage executed by the learning device 10.

[0046] (Step S101) The receiving unit 11 receives input of video data A.

[0047] (Step S102) The video analysis unit 12 acquires the pixel coordinates and video time at which the colored pixels are detected from the video data A.

[0048] (Step S103) The training image generation unit 13 converts the video time into pixel values ​​of pixel coordinates and generates a training image using the pixel values ​​for each pixel coordinate. A data set consisting of the training images and the image resulting from rendering of the video data A is used as training data in the machine learning of the learning model MB. For example, the last frame (final frame) of the video data A may be used as the image resulting from rendering of the video data A.

[0049] FIG. 3 is a diagram illustrating learning data according to this embodiment. FIG. 3(1) shows the final frame AFe of video data A. The final frame AFe shows the drawing result PA of the video data A. FIG. 3(1) also shows the stroke order W of the drawing result PA. In the video data A of FIG. 3(1), the drawing result PA is drawn in the stroke order W. In the example of FIG. 3, machine learning of the learning model MB is performed to estimate the stroke order of the drawing result PA.

[0050] FIG. 3(2) shows training image B generated from video data A in FIG. 3(1). Training image B includes a drawing area PB with the same pixel coordinates as the drawing result PA. The pixel value of each pixel included in drawing area PB represents the frame time (video time) at which each pixel was colored in video data A. In the example of FIG. 3(2), the later the video time (the later the stroke order), the larger the pixel value (the brighter the pixel). Therefore, in training image B, the drawing area PB changes from dark pixels to bright pixels from the beginning to the end of the stroke order W.

[0051] 3(3) shows a data set C consisting of training image B and the final frame AFe of video data A. Data set C is used as training data in the machine learning of learning model MB.

[0052] When the preprocessing unit 14 performs preprocessing on the data set C, a data set C' is further generated after the preprocessing of the data set C.

[0053] (Step S104) The learning unit 15 uses the learning image and the image of the drawing result of the video data A as input data to perform machine learning of the learning model MB so as to generate an output image that represents the stroke order of the drawing of the picture. In the example of FIG. 3, the learning unit 15 uses the dataset C as input data to perform machine learning of the learning model MB so as to generate an output image that represents the stroke order of the drawing of the picture. Furthermore, when the dataset C' resulting from the execution of preprocessing by the preprocessing unit 14 is also used, the learning unit 15 uses the dataset C and the dataset C' as input data to perform machine learning of the learning model MB so as to generate an output image that represents the stroke order of the drawing of the picture.

[0054] Here, the output image generated by the learning model MB through machine learning is an image (stroke order estimated time image) having pixel values ​​representing the drawing time when each pixel (drawn pixel) was colored.

[0055] (Step S105) The learning unit 15 determines whether the loss function in the machine learning of the learning model MB converges within a predetermined range. If the result of this determination is that the loss function converges within the predetermined range (step S105, YES), the process proceeds to step S106. On the other hand, if the loss function does not converge within the predetermined range (step S105, NO), the machine learning of the learning model MB in step S104 continues.

[0056] (Step S106) The learning unit 15 estimates the stroke order for the video data A, which is the verification data, using the learning model MB. Specifically, the learning unit 15 inputs an image of the drawing result of the video data A to the learning model MB, and obtains an output image representing the stroke order of the painting in the image of the drawing result of the video data A from the learning model MB. When the image of the drawing result of the video data A is input, the learning model MB outputs an estimated stroke order time image, which is an output image representing the stroke order of the painting in the image of the drawing result of the video data A. The stroke order video generation unit 16 generates a video (estimated stroke order playback video) in which each drawing pixel is colored at the drawing time from the estimated stroke order time image output from the learning model MB and the drawing result image of the video data A. Specifically, the stroke order video generation unit 16 generates a video (estimated stroke order playback video) in which, for each drawing pixel in the estimated stroke order time image, the drawing pixel of the image of the drawing result of the video data A is lit at the drawing time corresponding to the pixel value of the drawing pixel.

[0057] The user recognizes the stroke order of the painting in video data A estimated by learning model MB by playing the stroke order estimation playback video. The user evaluates the accuracy of the learning model MB's estimation of the stroke order of the painting by comparing the actual stroke order of the painting in video data A with the learning model MB's stroke order estimation result. If the user determines that the learning model MB's estimation accuracy of the stroke order of the painting is sufficient, the user instructs the learning device 10 to end machine learning of the learning model MB.

[0058] (Step S107) If the user determines that the learning model MB has sufficient accuracy in estimating the stroke order of the painting (step S107, YES), the process in Fig. 2 ends. On the other hand, if the user determines that the learning model MB has insufficient accuracy in estimating the stroke order of the painting (step S107, NO), the process returns to step S101, and other video data A is acquired to continue machine learning of the learning model MB.

[0059] In the example of the learning method procedure in FIG. 2, if the user determines that the stroke order estimation accuracy of the painting by the learning model MB is insufficient (step S107, NO), other video data A is added and then machine learning of the learning model MB is continued; however, this is not limited to this. For example, if the user determines that the stroke order estimation accuracy of the painting by the learning model MB is insufficient (step S107, NO), the preprocessing unit 14 may perform additional preprocessing on the dataset C to increase the amount of learning data and then continue machine learning of the learning model MB. For example, if the user determines that the stroke order estimation accuracy of the painting by the learning model MB is insufficient (step S107, NO), the user may adjust the hyperparameters of the machine learning of the learning model MB and then continue machine learning of the learning model MB. If the user determines that the stroke order estimation accuracy of the painting by the learning model MB is insufficient (step S107, NO), the user may adjust the hyperparameters of the machine learning of the learning model MB and then continue machine learning of the learning model MB. The configuration is such that the user can set what action to take when the user determines that the stroke order estimation accuracy of the painting by the learning model MB is insufficient (step S107, NO).

[0060] The trained model MA generated in the training step S100 is supplied to the estimation device 20.

[0061] <Estimation stage> Fig. 4 is a flowchart showing an example of the procedure of the estimation method according to this embodiment. The estimation step S200 according to this embodiment will be described with reference to Fig. 4. The estimation step S200 is a step of estimating the stroke order of a painting that is a target of stroke order estimation, using the trained model MA generated in the learning step S100 described above with reference to Fig. 2. The estimation step S200 is a step executed by the estimation device 20.

[0062] (Step S201) The receiving unit 21 receives an input of an image G of a painting (target painting image) for which a stroke order is to be estimated.

[0063] (Step S202) The stroke order estimation unit 22 inputs the target painting image G to the trained model MA and obtains from the trained model MA an output image representing the stroke order of the painting in the target painting image G. When the target painting image G is input, the trained model MA outputs an output image representing the stroke order of the painting in the target painting image G. Here, the output image generated by the trained model MA is an image (stroke order estimated time image) having pixel values ​​representing the drawing time when each pixel (drawing pixel) was colored.

[0064] (Step S203) The stroke order video generation unit 23 generates a video (stroke order estimation playback video) in which each drawing pixel is colored at the drawing time from the stroke order estimation time image and the target painting image G. Specifically, the stroke order video generation unit 23 generates a video (stroke order estimation playback video) in which, for each drawing pixel of the stroke order estimation time image, the drawing pixel of the target painting image G is lit at the drawing time corresponding to the pixel value of the drawing pixel.

[0065] (Step S204) The output unit 24 outputs the stroke order estimation playback video. By playing back the stroke order estimation playback video, the user can recognize the stroke order of the drawing of the painting that is the subject of stroke order estimation estimated by the trained model MA. According to this stroke order estimation playback video, the estimation result of the stroke order of the drawing of the painting that is the subject of stroke order estimation is played back as a video, so that the user can easily visually recognize the estimation result of the stroke order of the drawing of the painting that is the subject of stroke order estimation.

[0066] According to the above-described embodiment, a video of a person painting a picture is converted into a single training image (still image) and used as training data for machine learning of a model that estimates the stroke order of painting. This reduces the calculation cost compared to using the video as training data as is, resulting in effects such as a reduction in training time and suppression of the need for high-performance computers.

[0067] This embodiment may be applied to a painting drawn in one stroke, or to a painting drawn in succession.

[0068] Furthermore, the learning device 10 and the estimation device 20 may be realized using the same information processing device, or may be realized using different information processing devices.

[0069] In addition, a computer program for realizing the functions of each of the above-described devices may be recorded on a computer-readable recording medium, and the program recorded on the recording medium may be read and executed by a computer system. Note that the "computer system" here may also include hardware such as an OS and peripheral devices. Furthermore, if a WWW system is used, the "computer system" also includes the homepage provision environment (or display environment). In addition, "computer-readable recording medium" refers to writable non-volatile memory such as a flexible disk, optical magnetic disk, ROM, or flash memory, portable media such as a DVD (Digital Versatile Disc), or a storage device such as a hard disk built into a computer system.

[0070] Furthermore, the term "computer-readable recording medium" also includes those that retain a program for a certain period of time, such as volatile memory (e.g., DRAM (Dynamic Random Access Memory)) within a computer system that serves as a server or client when a program is transmitted via a network such as the Internet or a communication line such as a telephone line. The program may be transmitted from a computer system storing the program in a storage device or the like to another computer system via a transmission medium or by transmission waves in the transmission medium. Here, the "transmission medium" that transmits the program refers to a medium that has the function of transmitting information, such as a network (communication network) such as the Internet or a communication line (communication line) such as a telephone line. The program may also be a program for implementing some of the functions described above, or may be a so-called differential file (differential program) that can implement the functions described above in combination with a program already stored in the computer system.

[0071] Although an embodiment of the present invention has been described in detail above with reference to the drawings, the specific configuration is not limited to this embodiment, and design changes and the like are also included within the scope that does not deviate from the gist of the present invention. [Explanation of symbols]

[0072] 1...stroke order estimation system, 10...learning device, 20...estimation device, 11, 21...reception unit, 12...video analysis unit, 13...learning image generation unit, 14...preprocessing unit, 15...learning unit, 16...stroke order video generation unit, 17, 24...output unit, 22...stroke order estimation unit, 23...stroke order video generation unit, MB...learning model, MA...learned model

Claims

1. a video analysis unit that acquires pixel coordinates and video time at which painted pixels are detected from a video of a painting being painted by a person; a training image generating unit that converts the video time into a pixel value of the pixel coordinate and generates a training image for each pixel coordinate using the pixel value; a learning unit that performs machine learning of a learning model using the learning image and the image of the drawing result of the video as input data to generate an output image that represents the stroke order of drawing a painting; A learning device comprising:

2. The output image has a pixel value representing the time at which each pixel is painted. The learning device according to claim 1 .

3. a stroke order animation generating unit that generates an animation in which each drawing pixel is colored at a drawing time from the output image and the image of the drawing result of the animation; The learning device according to claim 2 .

4. a pre-processing unit that performs at least one of the following processes on the learning image and the image resulting from the moving image: rotation, enlargement, movement, adjustment of brightness, and change of RGB channels; The learning unit further uses the processing result on the input data. The learning device according to any one of claims 1 to 3.

5. a stroke order estimation unit that acquires pixel coordinates and video time at which painted pixels are detected from a video captured of a painting being painted by a person, converts the video time into pixel values ​​of the pixel coordinates, generates training images using the pixel values ​​for each of the pixel coordinates, inputs an image of the painting that is the subject of stroke order estimation to a trained model that has been machine-learned to generate an output image that represents the stroke order of the painting using the training images and an image of the drawing result of the video as input data, and acquires an output image that represents the stroke order of the painting that is the subject of stroke order estimation from the trained model; An estimation device comprising:

6. The output image has a pixel value representing the time at which each pixel is painted. The estimation device according to claim 5 .

7. a stroke order animation generation unit that generates an animation in which each drawing pixel is colored at a drawing time from the output image and an image of a painting that is a target of stroke order estimation; The estimation device according to claim 6 .

8. A learning method executed by an information processing device, A video analysis step of acquiring pixel coordinates and video time at which painted pixels are detected from a video of a painting being painted by a person; a learning image generating step of converting the video time into pixel values ​​of the pixel coordinates and generating a learning image for each of the pixel coordinates using the pixel values; a learning step of performing machine learning of a learning model using the training image and the image of the drawing result of the video as input data to generate an output image representing the stroke order of drawing a painting; Learning methods including.

9. An estimation method executed by an information processing device, a stroke order estimation step of acquiring pixel coordinates and video time at which painted pixels are detected from a video captured of a painting being painted by a person, converting the video time into pixel values ​​of the pixel coordinates, generating a training image using the pixel values ​​for each of the pixel coordinates, inputting an image of the painting that is the subject of stroke order estimation to a trained model that has been machine-learned to generate an output image that represents the stroke order of the painting using the training image and an image of the drawing result of the video as input data, and acquiring an output image that represents the stroke order of the painting that is the subject of stroke order estimation from the trained model; Estimation methods including:

10. On the computer, A video analysis step of acquiring pixel coordinates and video time at which painted pixels are detected from a video of a painting being painted by a person; a learning image generating step of converting the video time into pixel values ​​of the pixel coordinates and generating a learning image for each of the pixel coordinates using the pixel values; a learning step of performing machine learning of a learning model using the training image and the image of the drawing result of the video as input data to generate an output image representing the stroke order of drawing a painting; A computer program for executing

11. On the computer, a stroke order estimation step of acquiring pixel coordinates and video time at which painted pixels are detected from a video captured of a painting being painted by a person, converting the video time into pixel values ​​of the pixel coordinates, generating a training image using the pixel values ​​for each of the pixel coordinates, inputting an image of the painting that is the subject of stroke order estimation to a trained model that has been machine-learned to generate an output image that represents the stroke order of the painting using the training image and an image of the drawing result of the video as input data, and acquiring an output image that represents the stroke order of the painting that is the subject of stroke order estimation from the trained model; A computer program for executing

Citation Information

Patent Citations

  • Systems, methods, and programs for real-time end-to-end capturing of ink strokes from video

    JP2020009442A