Information processing methods, apparatus, and programs

The information processing method enhances rehabilitation motivation by generating injury/illness level videos to help subjects visualize their recovery, addressing the challenge of maintaining motivation in painful treatments.

JP7842948B2Active Publication Date: 2026-04-09EXAWIZARDS INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-11-17
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Existing rehabilitation technologies face challenges in maintaining the motivation of subjects undergoing treatment due to the painful nature of the process.

Method used

An information processing method that includes acquiring a subject video, extracting skeletal information, training a generation model, and generating an injury/illness level video to help subjects visualize their recovery progress.

Benefits of technology

The method effectively suppresses a decline in motivation by allowing subjects to clearly visualize their recovery progress, thereby enhancing their treatment adherence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007842948000001
    Figure 0007842948000001
  • Figure 0007842948000002
    Figure 0007842948000002
  • Figure 0007842948000003
    Figure 0007842948000003
Patent Text Reader

Abstract

To suppress the reduction of a motivation of a person subjected to a medical treatment for improving a function.SOLUTION: An information processing method according to an embodiment is executed by an information processing device, and includes: an obtaining process of obtaining a subjected person motion image that is a motion image of the subjected person carrying out a first action; an extracting process of extracting a skeleton motion image of the subjected person from the subjected person motion image; a learning process of training a generating model so as to generate the subjected person motion image from the skeleton motion image; and a generating process of inputting, to the generating model, the skeleton motion image corresponding to a second action in accordance with an injury and disease level, and of generating an injury and disease motion image that is a motion image in which the subjected person performs the second action in accordance with the injury and disease level.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing method, apparatus, and program.

Background Art

[0002] There is known a technology for assisting rehabilitation, which is one of the treatments performed on a subject with an injury or disability. Patent Document 1 discloses a technology for automatically determining a rehabilitation menu based on the actions of a subject.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] For a subject with an injury or disability, treatment including rehabilitation is painful, so it is difficult to maintain motivation. There is room for improvement in the technology of Patent Document 1 from the viewpoint of maintaining the motivation of the subject.

[0005] An object of the present invention is to suppress a decrease in the motivation of a subject who undergoes treatment for improving the function of the body.

Means for Solving the Problems

[0006] An information processing method according to one embodiment is an information processing method executed by an information processing device, comprising: an acquisition process for acquiring a subject video, which is a video of a subject performing a first action; an extraction process for extracting a skeletal video of the subject from the subject video; a learning process for training a generation model to generate the subject video from the skeletal video; and a generation process for inputting a skeletal video corresponding to a second action according to the level of injury or illness into the generation model and generating an injury / illness level video, which is a video of the subject performing the second action according to the level of injury or illness. [Effects of the Invention]

[0007] According to one embodiment, it is possible to suppress a decline in motivation of individuals undergoing treatment to improve their physical function. [Brief explanation of the drawing]

[0008] [Figure 1] This figure shows an example of the configuration of the information processing system according to this embodiment. [Figure 2] This figure shows an example of the hardware configuration of an information processing device. [Figure 3] This figure shows an example of the functional configuration of a video generation device. [Figure 4] This flowchart shows an example of the process performed by a video generation device. [Figure 5] This diagram schematically illustrates an example of the processing performed by a video generation device. [Figure 6] This figure shows an example of a video display screen shown on a user's terminal. [Modes for carrying out the invention]

[0009] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In the description and drawings of each embodiment, components having substantially the same functional configuration will be denoted by the same reference numerals, and redundant descriptions will be omitted.

[0010] <System Configuration> First, an overview of the information processing system according to this embodiment will be described. The information processing system according to this embodiment is a system for suppressing a decline in the motivation of a subject with an injury or illness by generating a video of the subject after treatment and showing the video to the subject. Injuries and illnesses include diseases and traumas. Treatment includes surgery, medication, and rehabilitation. The information processing system is used, for example, in medical facilities, nursing care facilities, or rehabilitation facilities.

[0011] Figure 1 shows an example of the configuration of an information processing system according to this embodiment. As shown in Figure 1, the information processing system according to this embodiment comprises a video generation device 1 and a user terminal 2, which are connected to each other so as to be able to communicate via a network N. The network N is, for example, a wired LAN (Local Area Network), a wireless LAN, the Internet, a public telephone network, a mobile data communication network, or a combination thereof. In the example in Figure 1, the information processing system comprises one video generation device 1 and one user terminal 2, but it may comprise multiple units of each.

[0012] The video generation device 1 is an information processing device that generates injury / illness level videos 127 based on subject video 121. The video generation device 1 is, for example, a PC (Personal Computer), a smartphone, a tablet terminal, a server device, or a microcomputer, but is not limited to these. The video generation device 1, subject video 121, and injury / illness level videos 127 will be described in detail later.

[0013] User terminal 2 is an information processing device used by the user. The user is, for example, a doctor, caregiver, occupational therapist, physical therapist, or nurse, but is not limited to these. User terminal 2 is, for example, a PC, smartphone, tablet, server device, or microcomputer, but is not limited to these. User terminal 2 transmits the subject video 121 to the video generation device 1 and receives and displays the injury / illness level video 127 from the video generation device 1. User terminal 2 may also acquire the subject video 121 from an external device and transmit the acquired subject video 121 to the video generation device 1. Furthermore, user terminal 2 may be equipped with a camera and transmit the subject video 121 captured by the camera to the video generation device 1. Note that the user terminal 2 that transmits the subject video 121 and the user terminal 2 that displays the injury / illness level video 127 may be different information processing devices.

[0014] <Hardware Configuration> Next, the hardware configuration of the information processing device 100 will be described. Figure 2 shows an example of the hardware configuration of the information processing device 100. As shown in Figure 2, the information processing device 100 comprises a processor 101, a memory 102, a storage 103, a communication I / F 104, an input / output I / F 105, and a drive device 106, all interconnected via bus B.

[0015] The processor 101 controls the various components of the information processing device 100 and realizes the functions of the information processing device 100 by loading various programs, including the OS (Operating System), stored in the storage 103, into the memory 102 and executing them. The processor 101 is, for example, a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a GPU (Graphics Processing Unit), an ASIC (Application Specific Integrated Circuit), a DSP (Digital Signal Processor), or a combination thereof.

[0016] The memory 102 is, for example, a ROM (Read Only Memory), a RAM (Random Access Memory), or a combination thereof. The ROM is, for example, a PROM (Programmable ROM), an EPROM (Erasable Programmable ROM), an EEPROM (Electrically Erasable Programmable ROM), or a combination thereof. The RAM is, for example, a DRAM (Dynamic RAM), an SRAM (Static RAM), or a combination thereof.

[0017] The storage 103 stores various programs and data including the OS. The storage 103 is, for example, a flash memory, a HDD (Hard Disk Drive), a SSD (Solid State Drive), a SCM (Storage Class Memories), or a combination thereof.

[0018] The communication I / F 104 is an interface for connecting the information processing apparatus 100 to an external apparatus via the network N and controlling communication. The communication I / F 104 is, for example, an adapter compliant with Bluetooth (registered trademark), Wi-Fi (registered trademark), ZigBee (registered trademark), Ethernet (registered trademark), or optical communication, but is not limited thereto.

[0019] The input / output I / F 105 is an interface for connecting the input device 107 and the output device 108 to the video generation apparatus 1. The input device 107 is, for example, a mouse, a keyboard, a touch panel, a microphone, a scanner, a camera, various sensors, operation buttons, or a combination thereof. The output device 108 is, for example, a display, a projector, a printer, a speaker, a vibrator, or a combination thereof.

[0020] The drive device 106 reads and writes data on the disk medium 109. The drive device 106 is, for example, a magnetic disk drive, an optical disk drive, a magneto-optical disk drive, or a combination thereof. The disk medium 109 is, for example, a CD (Compact Disc), a DVD (Digital Versatile Disc), a FD (Floppy Disk), a MO (Magneto-Optical disk), a BD (Blu-ray (registered trademark) Disc), or a combination thereof.

[0021] Note that in this embodiment, the program may be written in the memory 102 or the storage 103 at the manufacturing stage of the information processing device 100, may be provided to the information processing device 100 via the network N, or may be provided to the information processing device 100 via a non-temporary computer-readable recording medium such as the disk medium 109.

[0022] <Functional Configuration> Next, the functional configuration of the video generation device 1 will be described. FIG. 3 is a diagram showing an example of the functional configuration of the video generation device 1. As shown in FIG. 3, the video generation device 1 includes a communication unit 11, a storage unit 12, and a control unit 13.

[0023] The communication unit 11 is realized by the communication I / F 104. The communication unit 11 transmits and receives information to and from the user terminal 2 via the network N. The communication unit 11 receives the target person video 121 from the user terminal 2 and transmits the injury level video 127 to the user terminal 2.

[0024] The storage unit 12 is realized by the memory 102 and the storage 103. The storage unit 12 stores the target person video 121, the target person skeleton video 122, the skeleton extraction model 123, the generation model 124, the identification model 125, the injury level skeleton video 126, and the injury level video 127.

[0025] Subject video 121 is a video of a subject with an injury or illness performing a first action. The first action may be, for example, walking, standing up, sitting down, or raising an arm, but is not limited to these.

[0026] The subject skeletal video 122 is a skeletal video extracted from subject video 121 by the skeletal extraction model 123. A skeletal video is a video consisting of frames containing skeletal information. Skeletal information is information indicating the coordinates of multiple pre-set skeletons. Each frame of subject skeletal video 122 corresponds to each frame of subject video 121. In other words, subject skeletal video 122 contains the chronological skeletal information of the subject extracted from subject video 121.

[0027] Skeleton extraction model 123 is a machine learning model that extracts human skeletal information (detects posture) from videos. Any model, such as open pose models, can be used as skeleton extraction model 123.

[0028] Generative model 124 is a machine learning model that generates videos containing people from skeletal videos. In other words, generative model 124 is a model that performs the reverse operation of skeletal extraction model 123. Any model can be used as generative model 124. Generative model 124 is used as a generator for GAN (Generative Adversarial Network).

[0029] The discriminative model 125 is a machine learning model that identifies whether an input video was filmed by a person or generated by the generative model 124. Any model can be used as the discriminative model 125. The discriminative model 125 is used as a classifier in a GAN.

[0030] The injury / illness level skeletal video 126 is a skeletal video extracted by the skeletal extraction model 123 from a video of a person with an injury or illness (injured person) performing a second movement. The second movement may be the same as or different from the first movement. The second movement may be, for example, walking, standing up, sitting down, or raising an arm, but is not limited to these. An injury / illness level skeletal video 126 is prepared for each injury / illness. Furthermore, an injury / illness level skeletal video 126 is prepared for each level of injury / illness. The injury / illness level is an index indicating the degree of injury or illness, prepared for each injury / illness, and can be set arbitrarily. For example, if the injury / illness level of a certain injury / illness is set to five stages from 1 to 5, then the injury / illness level skeletal video 126 for injury / illness level 1 of that injury / illness is extracted from a video of an injured person with injury / illness level 1 performing a second movement and stored in the memory unit 12. The same applies to injury / illness level skeletal videos 126 for injury / illness levels 2 to 5.

[0031] Injury / Illness Level Video 127 is a video in which the subject performs a second action corresponding to the injury / illness level. Injury / Illness Level Video 127 can be generated for any injury / illness level, regardless of the subject's first action in Subject Video 121 or the subject's actual injury / illness level.

[0032] The control unit 13 is realized by the processor 101 reading and executing a program from the memory 102 and cooperating with other hardware components. The control unit 13 controls the overall operation of the video generation device 1. The control unit 13 comprises an acquisition unit 131, an extraction unit 132, a learning unit 133, a generation unit 134, and a display control unit 135.

[0033] The acquisition unit 131 acquires the target person video 121 received by the communication unit 11 from the user terminal 2 and stores it in the storage unit 12.

[0034] The extraction unit 132 uses the skeletal extraction model 123 to extract the subject's skeletal video 122 from the subject's video 121 and saves it to the storage unit 12.

[0035] The learning unit 133 uses the subject video 121 and the subject skeleton video 122 as training data to train the generative model 124 to generate the subject video 121 from the subject skeleton video 122, and stores the trained parameters in the storage unit 12. As a result, the generative model 124 is trained to output a video of the subject corresponding to the skeleton video when a skeleton video is input. The learning method will be described in more detail later.

[0036] The generation unit 134 inputs the injury / illness level skeletal video 126 into the trained generation model 124 to generate an injury / illness level video 127 for the subject corresponding to the injury / illness level skeletal video 126, and stores it in the storage unit 12.

[0037] The display control unit 135 transmits the subject video 121 and the injury / illness level video 127 to the user terminal 2 via the communication unit 11, causing the subject video 121 and the injury / illness level video 127 to be displayed on the user terminal 2's display.

[0038] Furthermore, each functional configuration of the video generation device 1 may be implemented by software as described above, or by hardware such as an IC chip, SoC (System on Chip), LSI (Large Scale Integration), or microcomputer.

[0039] <Processing performed by the video generation device> Next, the processes performed by the video generation device 1 according to this embodiment will be described. Figure 4 is a flowchart showing an example of the processes performed by the video generation device 1. Figure 5 is a diagram illustrating the processes performed by the video generation device 1.

[0040] (Step S101) The acquisition unit 131 of the video generation device 1 acquires the subject video 121 from the user terminal 2 and stores it in the storage unit 12.

[0041] (Step S102) The extraction unit 132 uses the skeletal extraction model 123 to extract the subject's skeletal video 122 from the subject's video 121 acquired in step S101 and stores it in the storage unit 12. Specifically, as shown in Figure 5, the extraction unit 132 inputs the subject's video 121 into the skeletal extraction model 123 and acquires the skeletal information for each frame output by the skeletal extraction model 123 as the subject's skeletal video 122.

[0042] (Step S103) The learning unit 133 uses the subject video 121 acquired in step S101 and the subject skeleton video 122 extracted in step S102 as training data to train the generation model 124 so that when the subject skeleton video 122 is input, it outputs the subject video 121. Specifically, as shown in Figure 5, the learning unit 133 inputs the subject skeleton video 122 into the generation model 124 and obtains the video 121' output by the generation model 124. Next, the learning unit 133 inputs the subject skeleton video 122 and either the subject video 121 or video 121' into the identification model 125, causing the identification model 125 to identify whether the input video is the subject video 121 or video 121'. The learning unit 133 trains the generative model 124 so that video 121' is identified as subject video 121 by the discrimination model 125, and trains the discrimination model 125 so that video 121' is identified as video 121' by the discrimination model 125. In other words, the learning unit 133 trains the generative model 124 using a GAN with the generative model 124 as the generator and the discrimination model 125 as the discriminator. As a result, the generative model 124 becomes a model that, when inputting a subject skeleton video 122, outputs a video 121' that approximates the subject video 121. Consequently, the generative model 124 is trained to output a video of the subject corresponding to the skeleton video when inputting a skeleton video. The learning unit 133 stores the parameters of the generative model 124 obtained through training in the storage unit 12 as parameters for the generative model 124 for the subject.

[0043] (Step S104) The generation unit 134 receives a selection from the user terminal 2 for the injury / illness level for which to generate the injury / illness level video 127.

[0044] (Step S105) The generation unit 134 generates a subject's injury / illness level video 127 using the subject's generation model 124, which was trained in step S103, and the injury / illness level skeletal video 126, which corresponds to the subject's injury / illness and the injury / illness level received in step S104. Specifically, as shown in Figure 5, the generation unit 134 inputs the injury / illness level skeletal video 126, which corresponds to the subject's injury / illness and the injury / illness level received in step S104, into the generation model 124, and acquires the video output by the generation model 124 as the subject's injury / illness level video 127.

[0045] For example, if the subject's injury is osteoarthritis of injury level 3, and injury level 1 is selected in step S104, the generation unit 134 inputs the injury level skeletal video 126 corresponding to injury level 1 osteoarthritis into the generation model 124. As a result, a video of the subject with injury level 1 osteoarthritis performing the second movement is generated as injury level video 127.

[0046] (Step S106) The display control unit 135 transmits the subject video 121 acquired in step S101 and the subject injury / illness level video 127 generated in step S105 to the user terminal 2 via the communication unit 11, and displays them on the user terminal 2's display.

[0047] Here, Figure 6 shows an example of the video display screen D displayed on the user terminal 2. The video display screen D displays the subject video 121 and the injury / illness level video 127 side by side. It also displays a dropdown list d for the user to select the injury / illness level for which the injury / illness level video 127 will be generated, and a generate button b for requesting the video generation device 1 to generate the injury / illness level video 127 for the selected injury / illness level. When the user selects injury / illness level 1 from the dropdown list d and presses the generate button b, the video generation device 1 generates the injury / illness level video 127 for injury / illness level 1, and the generated injury / illness level video 127 is displayed on the video display screen D. Note that the video display screen D is not limited to the example in Figure 6.

[0048] <Summary> As explained above, the video generation device 1 extracts the subject's skeletal video 122 from the subject's video 121, trains the generation model 124 using the subject's video 121 and the subject's skeletal video 122, and inputs the injury / illness level skeletal video 126 into the trained generation model 124 to generate the subject's injury / illness level video 127. In other words, the video generation device 1 can generate an injury / illness level video 127 with an injury / illness level different from the subject's actual injury / illness level based on the subject's video 121. Furthermore, the video generation device 1 can generate an injury / illness level video 127 with a second action different from the subject's first video based on the subject's video 121.

[0049] Let's consider a scenario where the user is a physician and the subject is a patient with osteoarthritis at disease level 3. Generally, even if a physician verbally explains to a patient that treatment will restore them to disease level 1, the patient may not be able to visualize what disease level 1 looks like. Therefore, the longer the treatment lasts, or the greater the burden on the patient, the more likely their motivation for treatment is to decrease.

[0050] In contrast, according to this embodiment, a doctor can simply send a subject video 121 to a video generation device 1 and select injury / illness level 1 to display an injury / illness level video 127 of a patient at injury / illness level 1 on the user terminal 2. By showing the patient the injury / illness level video 127 of a patient at injury / illness level 1 and explaining the treatment, the patient can clearly visualize what it will be like to recover to injury / illness level 1 through treatment. This helps to suppress a decline in the patient's motivation for treatment.

[0051] <Note> This embodiment includes the following disclosures.

[0052] (Note 1) An information processing method performed by an information processing device, The acquisition process involves obtaining a video of the subject, which is a video of the subject performing their first action, and An extraction process to extract a video of the subject's skeleton from the subject's video, A learning process to train a generation model to generate the subject video from the skeletal video, A generation process that inputs a skeletal video corresponding to a second movement according to the level of injury or illness into the generation model and generates an injury / illness level video, which is a video of the subject performing the second movement according to the level of injury or illness; Information processing methods including

[0053] (Note 2) The aforementioned skeletal video includes chronological skeletal information. The information processing method described in Appendix 1.

[0054] (Note 3) The skeletal video corresponding to the second movement according to the injury level is generated based on a video of a person with the injury level performing the second movement. The information processing method described in Appendix 1 or Appendix 2.

[0055] (Note 4) The second action is walking. The information processing method described in any of the appendices 1 through 3.

[0056] (Note 5) An acquisition unit that acquires a subject video, which is a video of the subject performing the first action, An extraction unit that extracts a video of the subject's skeleton from the subject's video, A learning unit that trains a generation model to generate the subject video from the skeletal video, A generation unit inputs a skeletal video corresponding to a second movement according to the level of injury or illness into the generation model and generates an injury / illness level video, which is a video of the subject performing the second movement according to the level of injury or illness. An information processing device equipped with the following features.

[0057] (Note 6) An information processing method performed by an information processing device, The acquisition process involves obtaining a video of the subject, which is a video of the subject performing their first action, and An extraction process to extract a video of the subject's skeleton from the subject's video, A learning process to train a generation model to generate the subject video from the skeletal video, A generation process that inputs a skeletal video corresponding to a second movement according to the level of injury or illness into the generation model and generates an injury / illness level video, which is a video of the subject performing the second movement according to the level of injury or illness; A program that causes a computer to execute an information processing method that includes such methods.

[0058] The embodiments disclosed herein should be considered in all respects to be illustrative and not restrictive. The scope of the present invention is indicated by the claims, not in the sense described above, and is intended to include all modifications in the sense and scope equivalent to the claims. Furthermore, the present invention is not limited to the embodiments described above, and various modifications are possible within the scope of the claims, and embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention. [Explanation of Symbols]

[0059] 1: Evaluation device 2: User terminal 11: Communications Department 12: Storage part 13: Control Unit 101: Processor 102: Memory 103: Storage 104: Communication I / F 105: Input / Output Interface 106: Drive unit 107: Input device 108: Output device 109: Disc media 121: Target video 122: Subject's skeletal video 123: Skeleton extraction model 124: Generative Models 125: Discriminant Model 126: Skeletal video showing injury / illness level 127: Videos showing injury / illness levels 131: Acquisition Department 132: Extraction part 133: Learning Department 134: Generation part 135: Display Control Unit

Claims

1. An information processing method performed by an information processing device, The acquisition process involves obtaining a video of the subject, which is a video of the subject performing their first action, and An extraction process to extract a video of the subject's skeleton from the subject's video, A learning process that trains a generation model to generate the subject video from the skeletal video, and trains an identification model to identify whether the video generated by the generation model is the subject video or the video generated by the generation model when the skeletal video and the subject video or the video generated by the generation model are input, so that the video generated by the generation model is identified as the subject video. A generation process that, corresponding to the injury or illness of the subject and for each of the multiple injury / illness levels set, inputs a pre-prepared skeletal video corresponding to the second action, which represents the recovered state of the injury / illness and corresponds to the arbitrarily set injury / illness level, into the generation model, and generates an injury / illness level video which is a video of the subject performing the second action according to the injury / illness level, Display process for displaying the aforementioned injury / illness level video on the video display screen, Information processing methods including

2. The aforementioned skeletal video includes chronological skeletal information. The information processing method according to claim 1.

3. The skeletal video corresponding to the second movement according to the injury level is generated based on a video of a person with the injury level performing the second movement. The information processing method according to claim 1 or claim 2.

4. The second action is walking. The information processing method according to any one of claims 1 to 3.

5. An acquisition unit that acquires a subject video, which is a video of the subject performing the first action, An extraction unit that extracts a video of the subject's skeleton from the subject's video, A learning unit trains a generation model to generate the subject video from the skeletal video, and a recognition model that, when inputting the skeletal video and the subject video or a video generated by the generation model, identifies whether it is the subject video or a video generated by the generation model, to train the generation model so that the video generated by the generation model is identified as the subject video. A generation unit that, corresponding to the injury or illness of the subject and for each of the multiple injury / illness levels set, inputs a pre-prepared skeletal video corresponding to the second action, which represents the recovered state of the injury / illness and is arbitrarily set to the injury / illness level, into the generation model, and generates an injury / illness level video which is a video of the subject performing the second action according to the injury / illness level, A display control unit that displays the aforementioned injury / illness level video on a video display screen, An information processing device equipped with the following features.

6. The acquisition process involves obtaining a video of the subject, which is a video of the subject performing their first action, and An extraction process to extract a video of the subject's skeleton from the subject's video, A learning process that trains a generation model to generate the subject video from the skeletal video, and trains an identification model to identify whether the video generated by the generation model is the subject video or the video generated by the generation model when the skeletal video and the subject video or the video generated by the generation model are input, so that the video generated by the generation model is identified as the subject video. A generation process that, corresponding to the injury or illness of the subject and for each of the multiple injury / illness levels set, inputs a pre-prepared skeletal video corresponding to the second action, which represents the recovered state of the injury / illness and corresponds to the arbitrarily set injury / illness level, into the generation model, and generates an injury / illness level video which is a video of the subject performing the second action according to the injury / illness level, Display process for displaying the aforementioned injury / illness level video on the video display screen, A program that causes a computer to execute an information processing method that includes such methods.

Citation Information

Patent Citations

  • Character action video generation method and system based on human skeleton sequence information and storage medium

    CN112419455A

  • Rehabilitation support system, rehabilitation support method, and program

    JP2019024580A

  • Operation information processing device

    JP2019046481A

  • Video synthesis method, model training method, device, and storage medium

    US20210243383A1