Information processing method, information processing device, program, video content production system, video content production method, and movable robot device
The system addresses the limitations of pre-generated camerawork by generating and refining motion information for mobile robot cameras using user feedback and digital twin simulations, ensuring desired video compositions are achieved.
Patent Information
- Application Number
- PCT/JP2025/008706
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-28
- Filing Date
- 2025-03-10
- Publication Date
- 2025-10-02
AI Technical Summary
Existing video shooting systems, particularly those using mobile robots with cameras, often fail to capture desired video compositions due to limited subject and composition variety, as they rely on pre-generated camerawork that may not align with the photographer's intentions.
A system that generates multiple types of motion information for a mobile robot equipped with a camera, updates this information based on user feedback, and uses a digital twin to simulate and adjust camerawork in real time, allowing for dynamic adjustment and alignment with the photographer's preferences.
Enables the capture of desired video compositions by suggesting and refining camerawork based on user feedback, ensuring that the final video shooting meets the photographer's creative vision.
Smart Images

Figure JP2025008706_02102025_PF_FP_ABST
Abstract
Description
Information processing method, information processing device, program, video content production system, video content production method, and movable robot device
[0001] The present disclosure relates to an information processing method, an information processing device, a program, a video content production system, a video content production method, and a movable robot device, and in particular to an information processing method, an information processing device, a program, a video content production system, a video content production method, and a movable robot device that enable a photographer to achieve desired video shooting.
[0002] In an imaging device that automatically shoots videos, if there is little physical movement of the camera location or the location of the subject, the variety of subjects and compositions is limited, so it is desirable to shoot a wide variety of videos.
[0003] In response to this, Patent Document 1 discloses a method of determining the camerawork to be applied to an automatically shot video based on the results of subject detection, and changing the imaging composition and the focus of the imaging optical system in a predetermined pattern while shooting the video based on the type of camerawork determined.
[0004] In recent years, it has become common to use mobile robots equipped with cameras to capture a variety of videos by operating them according to pre-generated camerawork.
[0005] Japanese Patent Application Laid-Open No. 2021-57818
[0006] However, it is not always possible to shoot video as desired by the photographer.
[0007] The present disclosure has been made in consideration of such circumstances, and enables a photographer to achieve desired video shooting.
[0008] An information processing method according to a first aspect of the present disclosure is an information processing method that generates multiple types of motion information representing the motion of a mobile robot equipped with a camera, and updates the motion information suggested to a user based on feedback information regarding the motion information selected by the user.
[0009] An information processing device according to a first aspect of the present disclosure is an information processing device that includes a generation unit that generates multiple types of motion information representing the motion of a mobile robot equipped with a camera, and the generation unit updates the motion information suggested to a user based on feedback information regarding the motion information selected by the user.
[0010] A program according to a first aspect of the present disclosure is a program for causing a computer to execute a process of generating multiple types of motion information representing the motion of a mobile robot equipped with a camera, and updating the motion information suggested to a user based on feedback information regarding the motion information selected by the user.
[0011] A video content production system according to a second aspect of the present disclosure is a video content production system including: a movable robot equipped with a camera; a digital twin generation unit that generates a digital twin that models the environment captured by the camera; a generation AI unit that holds a generation AI model that generates the camerawork of the camera and at least one of the contents in the digital twin based on instructions from a user; and a simulation execution unit that executes a simulation of the camera's shooting of at least one of the contents and the camerawork in the digital twin.
[0012] A video content production method according to a second aspect of the present disclosure is a video content production method including: a video content production system generating a digital twin that models the environment captured by a camera mounted on a movable robot based on images or audio output from the camera; having a generative AI model generate at least one of the camerawork of the camera and content in the digital twin based on instructions from a user; running a simulation in the digital twin regarding at least one of the camerawork and the content; and changing at least one of the camerawork or the content generated by the generative AI model based on interaction with the user, thereby changing the simulation in real time.
[0013] A third aspect of the present disclosure provides a movable robot device equipped with a camera, the movable robot device including a control unit having control parameters of the movable robot device, the control unit transmitting images or audio to a digital twin that models the environment captured by the camera, and changing the control parameters in the digital twin in accordance with the learning results of a learning model that is based on simulation results relating to at least one of content generated by a generative AI model or camerawork of the camera.
[0014] In a first aspect of the present disclosure, multiple types of motion information representing the motion of a mobile robot equipped with a camera are generated, and the motion information suggested to a user is updated based on feedback information regarding the motion information selected by the user.
[0015] In a second aspect of the present disclosure, a digital twin is generated that models the environment photographed by a camera mounted on a movable robot, and a generative AI model is maintained that generates the camerawork of the camera and at least one of the content in the digital twin based on instructions from a user, and a simulation of the camera photographing the environment with respect to at least one of the camerawork and the content is performed in the digital twin.
[0016] In a third aspect of the present disclosure, a control unit having control parameters of a movable robotic device equipped with a camera transmits images or audio to a digital twin that models the environment captured by the camera, and in the digital twin, the control parameters are changed according to the learning results of a learning model that is based on simulation results regarding at least one of the content generated by a generative AI model or the camerawork of the camera.
[0017] 1 is a diagram illustrating an example configuration of a mobile robot control system according to an embodiment of the present disclosure; FIG. 2 is a block diagram illustrating an example hardware configuration of a mobile robot; FIG. 3 is a block diagram illustrating an example hardware configuration of a computer; FIG. 4 is a diagram illustrating an example functional configuration of an information processing unit; FIG. 5 is a flowchart illustrating a flow of proposing camerawork; FIG. 6 is a diagram illustrating an example of a request for proposing camerawork; FIG. 7 is a diagram illustrating an example of generating camerawork; FIG. 8 is a diagram illustrating an example of extracting proposed camerawork; FIG. 9 is a diagram illustrating downloading camerawork; FIG. 10 is a diagram illustrating an example of a presentation screen for proposed camerawork; FIG. 11 is a diagram illustrating an example of an evaluation screen for proposed camerawork; FIG. 12 is a diagram illustrating an example of a robot operation screen; FIG. 13 is a diagram illustrating examples of read information and upload information; FIG. 14 is a diagram illustrating an example adjustment screen for proposed camerawork; FIG. 15 is a block diagram illustrating an example functional configuration of a video content production system; FIG. 16 is a flowchart illustrating a flow of producing video content.
[0018] Modes for carrying out the present disclosure (hereinafter referred to as embodiments) will be described below in the following order.
[0019] 1. Background 2. Configuration of a mobile robot control system 3. Camera work proposal flow 4. UI presentation example 5. Configuration of a video content production system
[0020] 1. Background In recent years, it has become common to shoot a variety of videos by operating a mobile robot equipped with a camera according to a pre-generated camerawork. However, it has not always been possible to shoot video using the camerawork desired by the cameraman.
[0021] For example, there are camera robots that can move and shoot in accordance with people's movements, enabling new camerawork that only a robot can do. Such camera robots are used for live shooting on indoor stages at events and concerts.
[0022] Currently, in live shooting situations, creators involved in the video production, such as cinematographers and switchers, select the camerawork they want to use from a large number of pre-generated cameraworks and repeatedly try them out. When operating a camera robot in this process, sometimes footage is captured that the creator's experience and intuition could not have predicted.
[0023] On the other hand, if many cameraworks are proposed using generative AI (artificial intelligence), creators may need to spend a huge amount of time checking them, and may not be able to use them properly.
[0024] Therefore, in the technology disclosed herein, the camerawork suggested to the user is updated based on feedback information regarding the operation information (camerawork) selected by the user, thereby suitably suggesting camerawork that the creator would like to adopt, and ultimately realizing the video shooting desired by the photographer.
[0025] 2. Configuration of Mobile Robot Control System (Overall System Configuration) FIG. 1 is a diagram illustrating an example configuration of a mobile robot control system according to an embodiment of the present disclosure.
[0026] The mobile robot control system 1 shown in FIG. 1 can be applied to a system for controlling a camera robot that moves in accordance with the movements of people and takes pictures, for example, at the above-mentioned live shooting site.
[0027] The mobile robot control system 1 is configured to include a user terminal 10 , a mobile robot 20 , and a server 30 .
[0028] The user terminal 10 is composed of a PC (Personal Computer), a tablet terminal, a smartphone, and a controller that controls the mobile robot 20, which are operated by a user (creator CR). The user terminal 10 constitutes a platform that can control various types of mobile robots 20.
[0029] The user terminal 10 operates the mobile robot 20 in accordance with operation commands based on camerawork selected by operation input from the creator CR from among the proposed cameraworks downloaded from the server 30. The user terminal 10 also accepts prompt input from the creator CR, such as an input of a request for camerawork proposal, and transmits this as input information to the server 30. Furthermore, the user terminal 10 transmits to the server 30, as feedback information regarding the camerawork, the operation input from the creator CR regarding the proposed camerawork downloaded from the server 30 (whether the proposed camerawork was selected, whether it was actually used, etc.) and the operation results of the mobile robot 20 based on the operation commands (whether it operated according to the selected proposed camerawork, etc.).
[0030] The mobile robot 20 is a mobile robot equipped with a camera. In the embodiment of the present disclosure, the mobile robot 20 is configured as a camera robot that captures images while moving in accordance with the movements of people at a live shooting site or the like, but may also be configured as an inspection robot that autonomously patrols a construction site to collect images, a drone that takes aerial photographs, or the like.
[0031] The mobile robot 20 performs actions (movement and shooting) in accordance with the action commands from the user terminal 10, thereby realizing the camerawork selected by the creator CR on the user terminal 10. The mobile robot 20 transmits the results of the actions based on the camerawork to the user terminal 10. The mobile robot 20 also transmits the results of the actions to the server 30 as feedback information regarding the camerawork.
[0032] The server 30 is configured, for example, by one or more virtual servers built on the Internet (cloud).
[0033] The server 30 generates a plurality of cameraworks based on input information from the user terminal 10. Furthermore, the server 30 extracts proposed cameraworks to be presented to the user terminal 10 from the generated plurality of cameraworks based on the analysis results of their feature amounts, etc. The server 30 also updates the cameraworks proposed to the creator CR based on feedback information from the user terminal 10 and the mobile robot 20.
[0034] (Example of Hardware Configuration of Mobile Robot) FIG. 2 is a block diagram showing an example of the hardware configuration of the mobile robot 20. As shown in FIG.
[0035] As shown in FIG. 2, the mobile robot 20 is made up of sensors 21 a, 21 b, . . . , a camera 22, a gimbal 23, an actuator 24, a communication unit 25, a memory unit 26, and a control unit 27.
[0036] The sensors 21a, 21b, ... (hereinafter simply referred to as sensors 21) are composed of distance sensors, collision prevention sensors, speed sensors, acceleration sensors, etc. Some of the sensors 21 may not be mounted on the mobile robot 20, but may be set in the environment in which the mobile robot 20 moves. Sensor data acquired by each of the sensors 21 is supplied to the control unit 27.
[0037] The camera 22 is composed of an RGB camera mounted on the mobile robot 20 and captures images under the control of the control unit 27 .
[0038] The gimbal 23 connects the main body of the mobile robot 20 to the camera 22. Each axis of the gimbal 23 rotates under the control of the control unit 27, allowing the camera 22 to stably capture images even while the mobile robot 20 is moving.
[0039] The actuator 24 is configured as a motor or the like that rotates the wheels or the like of the mobile robot 20. The actuator 24 is driven under the control of the control unit 27, thereby enabling the mobile robot 20 to move.
[0040] The communication unit 25 is a wireless communication module that performs wireless communication with external devices such as the user terminal 10 and the server 30. The communication unit 25 supplies information received via wireless communication to the control unit 27, and transmits information supplied from the control unit 27 via wireless communication.
[0041] The storage unit 26 is configured by a volatile memory such as a dynamic random access memory (DRAM), etc. The storage unit 26 stores various data obtained by calculations performed by the control unit 27.
[0042] The control unit 27 is composed of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), etc. The control unit 27 controls each part of the mobile robot 20. Specifically, the control unit 27 executes programs stored in a memory (not shown) to realize various functional blocks such as a sensor data integration unit 31, a recognition unit 32, a path / posture planning unit 33, a camera control unit 34, and a drive control unit 35.
[0043] The sensor data integration unit 31 integrates the sensor data acquired by each of the sensors 21. The recognition unit 32 performs recognition processing of the surrounding environment of the mobile robot 20, etc., based on the sensor data integrated by the sensor data integration unit 31. The path / posture planning unit 33 creates a path plan and posture plan for the mobile robot 20 based on the results of the recognition processing by the recognition unit 32. The results of the recognition processing by the recognition unit 32 may be stored in the memory unit 26. The camera control unit 34 controls the shooting by the camera 22 and the rotation of each axis of the gimbal 23 based on the path plan and posture plan created by the path / posture planning unit 33. The drive control unit 35 controls the drive of the actuator 24 based on the path plan and posture plan created by the path / posture planning unit 33.
[0044] (Hardware Configuration of Computer) FIG. 3 is a block diagram showing an example of the hardware configuration of a computer 50 that constitutes the user terminal 10 or the server 30. As shown in FIG.
[0045] 3, a CPU 51 executes various processes in accordance with a program stored in a read-only memory (ROM) 52 or a program loaded into a random access memory (RAM) 53. The RAM 53 also stores data necessary for the CPU 51 to execute various processes as needed.
[0046] The CPU 51, ROM 52, and RAM 53 are interconnected via a bus 54. An input / output interface 55 is also connected to this bus 54.
[0047] The input / output interface 55 is connected to an input unit 56 , an output unit 57 , a storage unit 58 , and a communication unit 59 .
[0048] The input unit 56 is composed of a keyboard, mouse, touch panel, etc. The output unit 57 is composed of a display such as a liquid crystal or organic EL (Electro-Luminescence) display, and a speaker, etc. The storage unit 58 is composed of a hard disk, etc. The communication unit 59 is composed of a wired communication module for performing wired communication, a wireless communication module for performing wireless communication, etc. The communication unit 59 performs communication processing with external devices directly or via a network such as the Internet.
[0049] A drive 60 is also connected to the input / output interface 55 as needed, and removable media 61 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory is appropriately attached to the drive 60. Computer programs read from the drive 60 are installed in the storage unit 58 as needed.
[0050] In the following, the functional configuration of the information processing unit realized mainly in the server 30 will be described.
[0051] (Example of Functional Configuration of Information Processing Unit) FIG. 4 is a block diagram showing an example of the functional configuration of an information processing unit to which the technology according to the present disclosure is applied.
[0052] 4 is configured to include an information input unit 110, a camerawork generation unit 120, and a camerawork extraction unit 130. When the information processing unit 100 is realized in the server 30, each functional block constituting the information processing unit 100 may be realized by one virtual server, or may each be realized by a plurality of virtual servers.
[0053] The information input unit 110 inputs input information and feedback information from the user terminal 10 and feedback information from the mobile robot 20 to the camerawork generation unit 120 and camerawork extraction unit 130 .
[0054] The camerawork generation unit 120 generates multiple types of motion information (camerawork) that represent the motion of the mobile robot 20, and outputs it to the camerawork extraction unit 130. The camerawork includes, for example, the movement path of the mobile robot 20 or the camera 22 mounted on the mobile robot 20 when taking an image, the movement speed of the mobile robot 20 or the camera 22 mounted on the mobile robot 20 when taking an image, and the posture or shooting direction of the camera 22 mounted on the mobile robot 20 when taking an image.
[0055] For example, the camerawork generation unit 120 generates camerawork by simulating the movement of a virtual robot corresponding to the mobile robot 20 using a generation AI that learns using linguistic information input by a user (creator CR) in a digital twin corresponding to a filming location. The digital twin includes, for example, data modeling an environment captured by a camera 22 mounted on the mobile robot 20, or data modeling an environment captured by a camera other than the camera 22 mounted on the mobile robot 20. The linguistic information is input to the user terminal 10 as a prompt from the creator CR. The prompt from the creator CR may be input by text or by voice. The linguistic information input as a prompt includes the filming location, filming conditions, the footage to be filmed, filming time, the characteristics and number of subjects to be filmed, the target audience, etc.
[0056] The prompt is not limited to linguistic information, but may also include specifications of the mobile robot 20, the status of the shooting environment, and images showing the state of the shooting location. The specifications of the mobile robot 20 include the degrees of freedom of the posture that the mobile robot 20 can assume, maximum speed, maximum acceleration, and limit values of other control parameters. The status of the shooting environment includes, for example, the presence and status of instruments, lighting, speakers, smoke, etc. used in an event or concert. Note that the specifications of the mobile robot 20 may be input to the user terminal 10 by the user (creator CR) or may be acquired directly from the mobile robot 20.
[0057] The camerawork includes the moving path and moving speed of the mobile robot 20 (coordinates and speed for each time t), the shooting position and shooting direction of the camera 22, as well as the settings of the camera 22 such as the zoom value, focus value, exposure and contrast, and the position of the subject within the angle of view. The camerawork also includes predicted images, which are CG (Computer Graphics) images corresponding to the images that may be captured by the camera 22 as the mobile robot 20 moves.
[0058] Furthermore, the camerawork generation unit 120 updates the camerawork proposed to the user (creator CR) based on feedback information from the information input unit 110. The update here may be an improvement (enhancement) of the camerawork previously proposed to the creator CR, or the generation of new camerawork proposed to the creator CR. The updated camerawork is closer to the camerawork desired by the user. Furthermore, the feedback information from the information input unit 110 may include an interaction with the user (creator CR). The interaction with the user is considered to be a multimodal interaction. Multimodal interaction includes interactions using voice, images, video, handwritten input, text, gestures, or the like. For example, the multimodal interaction is converted into a corresponding prompt and input to the generative AI model.
[0059] The camerawork extraction unit 130 extracts proposed camerawork (proposed operation information) to be proposed to the creator CR from the multiple types of camerawork generated by the camerawork generation unit 120 based on feedback information from the information input unit 110. The extracted proposed camerawork is presented to the creator CR on the user terminal 10.
[0060] Note that some or all of the functional blocks constituting the information processing unit 100 may be implemented in the user terminal 10 or the mobile robot 20 .
[0061] 3. Flow of Camerawork Proposal The flow of camerawork proposal in the mobile robot control system 1 will be described with reference to the flowchart of FIG.
[0062] In step S1, the user terminal 10 receives an input of a request for a camerawork proposal.
[0063] For example, as shown in Figure 6, the creator CR inputs prompts such as the shooting location, the subject, and the purpose of the shooting as a request for camerawork proposals. In the example of Figure 6, the creator CR inputs that the shooting location is "Yepp Yokohama," the subject is "a band," and the purpose of the shooting is "live streaming for fans." In addition, in the example of Figure 6, the creator CR also inputs the desired atmosphere of the video ("a dark light atmosphere, with a screen that looks like you're moving your head"), the lighting environment ("white light and blue light coming in at an angle"), and the specifications of the mobile robot 20 ("a Sony camera robot that can move in all directions")
[0064] In this way, the request for camerawork suggestions that has been input and accepted at the user terminal 10 is transmitted to the server 30 as input information.
[0065] In step S2, the server 30 (camerawork generator 120) generates multiple (multiple types of) camerawork based on the input request. Specifically, the camerawork generator 120 generates camerawork based on the movement of a virtual mobile robot corresponding to the mobile robot 20 in a virtual space serving as a simulation environment corresponding to real space.
[0066] For example, as shown in Fig. 7 , in the camerawork generation unit 120, the generation AI generates a shooting trajectory for the mobile robot 20 in a digital twin live house constructed in digital space corresponding to the live house in real space that is the shooting location. In the example of Fig. 7 , position information and speed information for the mobile robot 20 at each time point are generated as the shooting trajectory for the mobile robot 20. The generation AI then generates a predicted image, which is a CG image of the footage shot by the mobile robot 20 according to the shooting trajectory. The generation AI can generate multiple types of camerawork in response to one request.
[0067] In step S3, the server 30 (camerawork extraction unit 130) extracts a proposed camerawork to be presented on the user terminal 10 from the multiple cameraworks that have been generated in response to the input request.
[0068] For example, assume that four cameraworks A, B, C, and D are generated as shown in Figure 8. Of these, cameraworks B and D are extracted as proposed cameraworks that meet the requirements of the creator CR. On the other hand, camerawork A is determined to have an image with camera shaking so much that it is unbearable for humans to watch, based on the analysis results of its feature amounts, and is therefore excluded from the proposed cameraworks. Furthermore, camerawork C is determined to have over-lighting, causing the image to be blown out, based on the analysis results of its feature amounts, and is therefore excluded from the proposed cameraworks.
[0069] The proposed camerawork extracted in this manner is presented to the creator CR on the user terminal 10.
[0070] In step S4, the user terminal 10 accepts input of feedback information regarding the camerawork selected by the creator CR from the proposed cameraworks presented.
[0071] For example, as shown in Figure 9, first, the creator CR accesses the camerawork suggestion system (server 30) on the user terminal 10. Second, the creator CR checks the predicted images for each proposed camerawork that is presented. And third, the creator CR downloads the camerawork that he or she wants to use.
[0072] That is, in step S4, before the mobile robot 20 is operated based on the camerawork selected by the user, input of feedback information is accepted. The feedback information here includes playback information indicating that predicted video corresponding to the camerawork selected by the user has been played back. Furthermore, if the proposed camerawork selected by the user has been downloaded for controlling the mobile robot 20 and the camera 22, the feedback information includes download information indicating that the camerawork has been downloaded for controlling the mobile robot 20 and the camera 22. Furthermore, the feedback information may include evaluation information indicating the user's evaluation of the camerawork selected by the user. The evaluation information may include a score assigned to the predicted video corresponding to the camerawork.
[0073] In step S5, the user terminal 10 controls the mobile robot 20 (camera 22) based on the downloaded camerawork. Software (application) for controlling the mobile robot 20 including the camera 22 is installed in the user terminal 10, thereby configuring a robot control system that serves as the basis for robot control. The robot control system may not only be configured within the user terminal 10, but may also be configured by other hardware and software working together. Note that the number of mobile robots 20 controlled by one camerawork is not limited to one, and multiple robots may also be used.
[0074] In step S6, the user terminal 10 receives input of feedback information regarding the downloaded camerawork (camerawork that realizes control of the mobile robot 20).
[0075] That is, in step S6, input of feedback information is accepted during or after the operation of the mobile robot 20 under control based on the camerawork selected by the user. The feedback information here includes operation result information that indicates the results of the operation of the mobile robot 20 under control based on the camerawork selected by the user. The operation result information may include real-time video captured by the camera 22 while the mobile robot 20 is operating under control based on the camerawork. The feedback information may also include linguistic information input by the user while the mobile robot 20 is operating.
[0076] In this way, the feedback information input and accepted at the user terminal 10 is transmitted to the server 30 .
[0077] Then, in step S7, the server 30 (camerawork generation unit 120) updates the camerawork proposed to the user (creator CR) based on the feedback information input to the user terminal 10.
[0078] According to the above processing, the camerawork suggested to the user is updated based on feedback information regarding the camerawork selected by the user, so that the creator can suitably suggest camerawork that they would like to use, and ultimately, it becomes possible to realize the video shooting that the photographer desires.
[0079] 4. UI Presentation Example The following describes a presentation example of a UI (User Interface) in the mobile robot control system 1. The UI may be presented mainly on the user terminal 10.
[0080] (Proposed Camerawork Presentation Screen) FIG. 10 is a diagram showing an example of a proposed camerawork presentation screen.
[0081] In the upper right corner of the presentation screen 200 shown in FIG. 10, a creator ID indicating the user who is logged in to the camerawork proposal system (server 30) is displayed.
[0082] On the presentation screen 200, thumbnails 221 of predicted images for a plurality of proposed cameraworks are displayed for each purpose and characteristic of the camerawork. In the example of Fig. 10, thumbnails 221 of camerawork preferred by the photographer and camerawork suitable for live streaming are displayed side by side in the horizontal direction of the screen. The presentation screen 200 is scrollable in the vertical direction. In addition to those shown in Fig. 10, the purposes and characteristics of camerawork may include, for example, camerawork suitable for recording media such as live DVDs, for music festivals, and for bands, as well as intense camerawork, mellow camerawork, and camerawork preferred by young people.
[0083] When the thumbnail 221 is selected (clicked) by the user, the predicted video of the proposed camerawork is enlarged and displayed. Also, when the thumbnail 221 is selected (clicked) by the user, playback information indicating that the predicted video corresponding to the proposed camerawork has been played back is transmitted to the server 30 as feedback information.
[0084] Below the thumbnail 221 of each proposed camerawork, match degree information 222 is displayed, and a download button 223 and a scoring button 224 are provided.
[0085] The match degree information 222 is a value indicating the degree of matching of the proposed camerawork with the purpose and characteristics of the camerawork.
[0086] The download button 223 is a button for downloading the proposed camerawork to the user terminal 10. When the download button 223 is selected (clicked) by the user, the file of the proposed camerawork is downloaded from the server 30 to the user terminal 10. Furthermore, when the download button 223 is selected (clicked) by the user, download information indicating that the proposed camerawork has been downloaded is transmitted to the server 30 as feedback information.
[0087] The scoring button 224 is a button for assigning a score as evaluation information indicating the user's evaluation of the camerawork. When the scoring button 224 is selected (clicked) by the user, an evaluation screen for the proposed camerawork is popped up on the presentation screen 200.
[0088] (Proposed Camerawork Evaluation Screen) FIG. 11 is a diagram showing an example of a proposed camerawork evaluation screen.
[0089] As shown in FIG. 11, the evaluation screen 230 is displayed superimposed on the presentation screen 200 that is displayed in gray.
[0090] An image display area 231 is provided at the top of the evaluation screen 230. In the image display area 231, a predicted image of the proposed camerawork corresponding to the selected scoring button 224 is displayed (played).
[0091] In the evaluation screen 230, an evaluation input UI 232 is provided for each evaluation item below the video display area 231. In the example of FIG. 11 , the evaluation input UI 232 displays a slide bar for evaluating (scoring) the evaluation items, such as whether the camerawork is to the creator's (the creator's) taste, whether the camerawork is something the creator would not have thought of, and whether the camerawork is usable. The evaluation screen 230 is scrollable up and down. In addition to the evaluation items shown in FIG. 11 , the evaluation items may include the target of the video, the filming location, the characteristics of the subject, the atmosphere of the music, and so on. The evaluation input UI 232 is not limited to a slide bar, and may be a checkbox or drop-down list for selecting the score for the evaluation item, a text box for entering an evaluation comment, or the like. Evaluation information (scores) may be input by voice input, in addition to operating the evaluation input UI 232.
[0092] In this way, the evaluation information input on the evaluation screen 230 is also transmitted to the server 30 as feedback information.
[0093] (Robot Operation Screen) FIG. 12 is a diagram showing an example of a robot operation screen for controlling the movement of the mobile robot 20 based on the downloaded camera work.
[0094] A movement path display area 310 is provided in the left two-thirds of the robot operation screen 300 shown in FIG. 12 . The movement path display area 310 displays a movement path MR (photography trajectory) along which the mobile robot 20 moves in real space shown in a top view. In the example of FIG. 12 , the movement path display area 310 displays, along with the movement path MR, an icon 20V indicating the current position of the mobile robot 20 and an arrow SD indicating the shooting direction of the camera 22 mounted on the mobile robot 20. This allows the user (creator CR) to confirm the current position of the mobile robot 20 and the shooting direction of the camera 22. For example, if the current position of the mobile robot 20 deviates from the movement path MR defined by the downloaded camerawork, warning information AL indicating the position of the mobile robot 20 may be presented, as shown in FIG. 12 .
[0095] On the robot operation screen 300, a predicted image display area 321 and a captured image display area 322 are arranged vertically to the right of the movement path display area 310. The predicted image display area 321 displays predicted images (CG images) of the downloaded camera work, and the captured image display area 322 displays captured images (real-time images) captured by the camera 22 while the mobile robot 20 is operating under control based on the downloaded camera work. This allows the user (creator CR) to understand the difference between the predicted images of the camera work and the real-time images captured by the camera 22.
[0096] On the robot operation screen 300, a camerawork load button 331 and an upload button 332 are provided above the predicted image display area 321. The camerawork load button 331 is a button for loading the camerawork (proposed camerawork) downloaded to the user terminal 10 into, for example, a robot control system configured in the user terminal 10. The upload button 332 is a button for uploading to the server 30 operation result information that indicates the results of the operation of the mobile robot 20 under control based on the camerawork.
[0097] FIG. 13 is a diagram showing an example of read information that is read into the robot control system by the camerawork read button 331 and an example of upload information that is uploaded by the upload button 332.
[0098] As shown in Figure 13A, the read information read into the robot control system includes camerawork downloaded from the server 30. The camerawork includes a camerawork ID unique to the camerawork, parameters such as the wheels, weight, and motor torque of the vehicle (virtual mobile robot) used, predicted images generated based on the digital twin, and the shooting trajectory of the simulated virtual mobile robot. The robot control system (user terminal 10) can control the mobile robot 20 by using this read information.
[0099] 13B, the upload information uploaded to the server 30 includes the camerawork downloaded from the server 30 and operation result information indicating the results of the operation of the mobile robot 20. In addition to the camerawork ID and parameters of the machine used (mobile robot 20), the operation result information includes the video (real-time video) captured according to the camerawork, the actual shooting trajectory of the mobile robot 20, and even the words (linguistic information) uttered by the creator CR while the mobile robot 20 is operating. The words uttered by the creator CR include the creator CR's desired composition, intention, areas for improvement, etc., and are obtained by being picked up by a microphone (not shown) and converted into text information.
[0100] Such upload information linking the operation result information and the camerawork is transmitted as feedback information to the server 30. The server 30 (camerawork generation unit 120) performs learning using the words and camerawork uttered by the creator CR that are linked to each other as learning data, thereby making it possible to generate camerawork that the creator CR is more likely to adopt.
[0101] (Adjustment Screen for Proposed Camerawork) An adjustment screen for accepting adjustments to the proposed camerawork may be presented to the user.
[0102] For example, on the presentation screen 200 described with reference to Figure 10, when a user selects (clicks) an adjustment button (not shown) for adjusting the proposed camerawork, an adjustment screen for the proposed camerawork pops up on the presentation screen 200.
[0103] FIG. 14 is a diagram showing an example of a proposed camerawork adjustment screen.
[0104] As shown in FIG. 14, the adjustment screen 410 is displayed superimposed on the presentation screen 200 that is displayed in gray.
[0105] An image display area 411 is provided in the upper left corner of the adjustment screen 410. In the image display area 411, a predicted image of the proposed camerawork corresponding to the selected adjustment button is displayed (played back).
[0106] A simulation image display area 412 is provided in the upper right section of the adjustment screen 410. A simulation image of the proposed camerawork corresponding to the selected adjustment button is displayed (played) in the simulation image display area 412. The simulation image may include the shooting trajectory of the virtual mobile robot in the digital twin.
[0107] On the adjustment screen 410, adjustment buttons 421 are provided below the simulation image display area 412, and a camerawork adjustment area 422 is provided in the lower half of the adjustment screen 410. The adjustment buttons 421 are buttons for adjusting, for example, the rhythm and tempo of a piece of music, such as a live performance, that is filmed using the camerawork. The camerawork adjustment area 422 displays time-series data on the movement path and movement speed of the mobile robot 20, and the shooting position and shooting direction of the camera 22. In the example of FIG. 14 , the camerawork adjustment area 422 displays time-series data on the point of gaze of the mobile robot 20, the zoom speed of the camera 22, and the speed and position of the mobile robot 20. The creator CR can adjust the angle of view of the filmed video, the position of the subject within the angle of view, and the like by operating the camerawork adjustment area 422 while adjusting the rhythm and tempo of the music using the adjustment buttons 421.
[0108] 5. Configuration of the video content production system The above describes a mobile robot control system that updates the camerawork suggested to the user based on feedback information regarding the camerawork selected by the user, thereby suitably suggesting camerawork that the creator would like to adopt.
[0109] Not limited to this, the technology disclosed herein can also realize the video shooting desired by the photographer by constructing a digital twin from the robot's sensor data and running a simulation on the constructed digital twin.
[0110] FIG. 15 is a diagram illustrating an example configuration of a video content production system according to an embodiment of the present disclosure.
[0111] The video content production system 501 shown in Figure 15 can be applied to a system that controls a camera robot that shoots from various viewpoints at the live shooting site described above, as well as at the shooting site of video content including movies and various videos.
[0112] Video content production system 501 is configured to include a mobile robot 510, a digital twin generation unit 520, a generation AI unit 530, and a simulation execution unit 540. Digital twin generation unit 520, generation AI unit 530, and simulation execution unit 540 may be realized in server 30 of FIG. 1 , or may each be realized in a separate server or computer.
[0113] Movable robot 510 is a mobile robot equipped with camera 511. In an embodiment of the present disclosure, movable robot 510 may be configured as a camera robot that captures images from various viewpoints at filming locations for various video content, or as a crane camera attached to a crane-type manipulator. Movable robot 510 includes control unit 510c having control parameters for movable robot 510. Control unit 510c can transmit images or audio to a digital twin that models the environment captured by camera 511. Furthermore, control unit 510c can change the control parameters in the digital twin in accordance with the learning results of a learning model based on simulation results related to at least one of the content generated by the generative AI model held by generation AI unit 530 or the camerawork.
[0114] The digital twin generation unit 520 generates a digital twin that models the environment captured by the camera 511. Specifically, the digital twin generation unit 520 constructs a digital twin of the filming location of the video content based on images captured by the camera 511 mounted on the mobile robot 510 and sensor data obtained from various other sensors.
[0115] The generation AI unit 530 holds a generation AI model that generates at least one of the camerawork of the camera 511 and the content in the digital twin generated by the digital twin generation unit 520 based on instructions from a user (creator CR such as a film director or video producer) via a UI (not shown). The generation AI model can create various photographed scenes in the digital twin in response to instructions from the user.
[0116] The content of a digital twin is a digital model represented by at least one of images and video corresponding to the real world. The video may include moving images with or without audio. The digital model also includes a virtual movable robot equipped with a virtual camera corresponding to the movable robot 510 equipped with the camera 511. Conceptually, the content may include the placement of subjects (actors, extras, vehicles, furniture, buildings, etc.) in a filmed scene of a movie, animation, game, live concert, or live sports broadcast, weather (sunny, cloudy, rainy, snowy, etc.), and time of day (morning, noon, evening, night, etc.). The generation AI unit 530 can also change the content in real time by changing the camerawork generated by the generation AI model based on interactions with the user (creator CR).
[0117] The simulation execution unit 540 executes a simulation of shooting by the camera 511 regarding at least one of the camerawork of the camera 511 and the content in the digital twin generated by the digital twin generation unit 520. That is, the simulation execution unit 540 executes a simulation of the shooting composition at the shooting location by moving the viewpoint of the virtual camera corresponding to the camera 511 and positioning the subject in the digital twin. At this time, the simulation execution unit 540 can also execute the simulation by selecting one of multiple cameraworks generated by the generation AI model held by the generation AI unit 530.
[0118] Then, the generative AI unit 530 modifies the generative AI model it holds, using the results of the simulation performed by the simulation execution unit 540 as learning data.
[0119] In the video content production system 501 configured in this manner, the creator CR can change the weather (sunny, cloudy, rainy, snowy, etc.) and time (morning, noon, evening, night, etc.) of the shooting scene in real time by making a request to the generative AI model, and simulate each scene in the digital twin. The creator CR can also change the shooting scene to the present, past, or future and simulate it. Furthermore, the creator CR can freely simulate multiple cameraworks generated by the generative AI model on the digital twin and select them in real time. The creator CR can also interact with the generative AI model in a multimodal manner, such as through voice, handwriting, gestures, and text. For example, the creator CR can generate prompts corresponding to multimodal interactions and input them into the generative AI, thereby changing the simulation of the shooting scene in real time.
[0120] The flow of video content production by the video content production system 500 will be described with reference to the flowchart of FIG.
[0121] In step S101, the digital twin generation unit 520 generates a digital twin that models the environment captured by the camera 511 based on the images or audio output from the camera 511 mounted on the mobile robot 510.
[0122] In step S102, the generation AI unit 530 causes the generation AI model to generate at least one of the camerawork of the camera 511 and the content in the digital twin based on instructions from the user (creator CR).
[0123] In step S103, the simulation execution unit 540 executes a simulation regarding at least one of the camerawork and content generated by the generation AI model in the digital twin generated by the digital twin generation unit 520.
[0124] Then, in step S104, the simulation execution unit 540 changes the camerawork or content generated by the generative AI model based on interaction with the user, and changes the simulation in real time.
[0125] With the above configuration and processing, creators can use the generative AI model on set to freely change actors, extras, furniture, buildings, scenery, and time of day, while simulating the digital twin to create scenes and camerawork in real time. Creators can also select and edit multiple cameraworks suggested by the generative AI via a UI. This allows them to pursue a shooting composition on set without being constrained by human resources, time, weather, and other factors, ultimately enabling them to achieve the video shoot they desire.
[0126] The embodiments of the present disclosure are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present disclosure.
[0127] In the embodiments of the present disclosure, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.
[0128] For example, the embodiment of the present disclosure can be configured as a cloud computing system in which a single function is shared and processed collaboratively by multiple devices via a network.
[0129] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.
[0130] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.
[0131] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.
[0132] Furthermore, the technology according to the present disclosure may have the following configurations. (1) An information processing method that generates multiple types of motion information representing motion of a mobile robot equipped with a camera, and updates the motion information proposed to the user based on feedback information related to the motion information selected by the user. (2) The information processing method described in (1), in which the motion information includes a movement path and a movement speed of the mobile robot, and a shooting position and a shooting direction of the camera. (3) The information processing method described in (2), in which the motion information further includes CG images corresponding to images that may be captured by the camera as the mobile robot moves. (4) The information processing method described in any of (1) to (3), in which input of the feedback information is accepted before the mobile robot moves under control based on the motion information selected by the user. (5) The information processing method described in (4), in which the feedback information includes playback information indicating that CG images corresponding to the motion information selected by the user have been played back. (6) The information processing method described in (4), in which the feedback information includes download information indicating that the motion information selected by the user has been downloaded for control of the mobile robot and the camera. (7) The information processing method according to (4), wherein the feedback information includes evaluation information indicating the user's evaluation of the operation information selected by the user. (8) The information processing method according to (7), wherein the evaluation information includes a score assigned to a CG image corresponding to the operation information. (9) The information processing method according to any of (1) to (8), wherein input of the feedback information is accepted during or after operation of the mobile robot under control based on the operation information selected by the user. (10) The information processing method according to (9), wherein the feedback information includes operation result information indicating a result of operation of the mobile robot under control based on the operation information selected by the user. (11) The information processing method according to (10), wherein the operation result information includes real-time video captured by the camera while the mobile robot is operating under control based on the operation information.(12) The information processing method according to (9), wherein the feedback information includes linguistic information input by the user while the mobile robot is moving. (13) The information processing method according to any of (1) to (12), wherein the motion information is generated using a generation AI that learns using linguistic information input by the user as input. (14) The information processing method according to any of (1) to (13), wherein proposed motion information that becomes the motion information proposed to the user is extracted based on the feedback information, and the extracted proposed motion information is presented to the user. (15) The information processing method according to (14), wherein an adjustment screen that accepts adjustment of the proposed motion information is presented to the user. (16) The information processing method according to any of (1) to (15), wherein the motion information is generated based on motion of a virtual mobile robot corresponding to the mobile robot in a virtual space corresponding to real space. (17) An information processing device comprising: a generation unit that generates multiple types of motion information representing the motion of a mobile robot equipped with a camera, wherein the generation unit updates the motion information proposed to a user based on feedback information regarding the motion information selected by the user. (18) A program that causes a computer to execute a process of generating multiple types of motion information representing the motion of a mobile robot equipped with a camera, and updating the motion information proposed to the user based on feedback information regarding the motion information selected by the user. (19) A video content production system comprising: a movable robot equipped with a camera; a digital twin generation unit that generates a digital twin modeling an environment captured by the camera; a generation AI unit that holds a generative AI model that generates camerawork of the camera and at least one of content in the digital twin based on instructions from a user; and a simulation execution unit that executes a simulation of shooting by the camera regarding at least the camerawork and the content in the digital twin. (20) The video content production system described in (19), wherein the generation AI unit uses a result of the simulation execution as learning data to change the generative AI model.(21) The video content production system described in (19) or (20), wherein the simulation execution unit executes the simulation by selecting one of the plurality of cameraworks generated by the generative AI model. (22) The video content production system described in any of (19) to (21), wherein the simulation execution unit changes the simulation in real time by changing the camerawork generated by the generative AI model based on an interaction with the user. (23) The video content production system described in (22), wherein the interaction with the user is a multimodal interaction. (24) The video content production system described in (23), wherein the multimodal interaction is an interaction using audio, images, video, handwritten input, text, or gestures. (25) The video content production system described in (24), wherein the multimodal interaction is converted into a corresponding prompt and input to the generative AI model. (26) The video content production system according to any one of (19) to (25), wherein the content is expressed by at least one of images and video. (27) The video content production system according to (26), wherein the content includes images or video related to the placement of subjects, weather, buildings, or time of day in a filmed scene of a movie, animation, game, live concert, or live sports broadcast. (28) The video content production system according to (27), wherein the subjects include actors and extras, the weather includes sunny, cloudy, rainy, and snowy, and the time of day includes morning, noon, evening, and night.(29) A video content production method comprising: a video content production system generating a digital twin modeling an environment captured by a camera mounted on a movable robot based on images or audio output from the camera; causing a generative AI model to generate at least one of camerawork of the camera and content in the digital twin based on instructions from a user; running a simulation of at least one of the camerawork and the content in the digital twin; and changing the camerawork or the content generated by the generative AI model based on interaction with the user to change the simulation in real time. (30) A mobile robot device equipped with a camera, comprising: a control unit having control parameters of the mobile robot device, the control unit transmitting images or audio to a digital twin modeling an environment captured by the camera, and changing the control parameters in the digital twin according to learning results of a learning model based on simulation results of at least one of the content generated by a generative AI model or the camerawork of the camera.
[0133] 1 Mobile robot control system, 10 User terminal, 20 Mobile robot, 30 Server, 22 Camera, 100 Information processing unit, 110 Information input unit, 120 Camera work generation unit, 130 Camera work extraction unit, 501 Video content production system, 510 Movable robot, 510c Control unit, 511 Camera, 520 Digital twin generation unit, 530 Generation AI unit, 540 Simulation execution unit
Claims
1. An information processing method for generating multiple types of motion information representing the motion of a mobile robot equipped with a camera, and updating the motion information suggested to a user based on feedback information regarding the motion information selected by the user.
2. The information processing method according to claim 1, wherein the operation information includes the moving path and moving speed of the mobile robot, and the shooting position and shooting direction of the camera.
3. The information processing method according to claim 2, wherein the operation information further includes CG images corresponding to images that may be captured by the camera as the mobile robot moves.
4. The information processing method according to claim 1, further comprising accepting input of the feedback information before the mobile robot operates under control based on the operation information selected by the user.
5. The information processing method according to claim 4, wherein the feedback information includes playback information indicating that a CG image corresponding to the action information selected by the user has been played back.
6. The information processing method according to claim 4, wherein the feedback information includes download information indicating that the operation information selected by the user has been downloaded for control of the mobile robot and the camera.
7. The information processing method according to claim 4, wherein the feedback information includes evaluation information indicating the user's evaluation of the action information selected by the user.
8. The information processing method according to claim 7, wherein the evaluation information includes a score given to a CG image corresponding to the action information.
9. The information processing method according to claim 1, wherein the input of the feedback information is accepted during or after the mobile robot is operating under control based on the operation information selected by the user.
10. The information processing method according to claim 9, wherein the feedback information includes operation result information indicating the result of the mobile robot operating under control based on the operation information selected by the user.
11. The information processing method according to claim 10, wherein the operation result information includes real-time video captured by the camera while the mobile robot is operating under control based on the operation information.
12. The information processing method according to claim 9, wherein the feedback information includes linguistic information input by the user while the mobile robot is operating.
13. The information processing method according to claim 1, wherein the action information is generated using a generation AI that learns using linguistic information input by the user as input.
14. The information processing method according to claim 1, further comprising: extracting suggested motion information that will be the motion information proposed to the user based on the feedback information; and presenting the extracted suggested motion information to the user.
15. The information processing method according to claim 14, further comprising presenting to the user an adjustment screen for accepting adjustment of the proposed action information.
16. The information processing method according to claim 1, wherein the movement information is generated based on the movement of a virtual mobile robot corresponding to the mobile robot in a virtual space corresponding to real space.
17. An information processing device comprising: a generation unit that generates multiple types of motion information representing the motion of a mobile robot equipped with a camera, wherein the generation unit updates the motion information suggested to a user based on feedback information regarding the motion information selected by the user.
18. A program for causing a computer to execute a process of generating multiple types of motion information representing the motion of a mobile robot equipped with a camera, and updating the motion information suggested to a user based on feedback information regarding the motion information selected by the user.
19. A video content production system comprising: a movable robot equipped with a camera; a digital twin generation unit that generates a digital twin that models the environment captured by the camera; a generation AI unit that holds a generation AI model that generates the camerawork of the camera and at least one of the contents in the digital twin based on instructions from a user; and a simulation execution unit that executes a simulation of the camera's shooting of at least one of the camerawork and the content in the digital twin.
20. A video content production system as described in claim 19, wherein the generation AI unit uses the results of the simulation as learning data to modify the generation AI model.
21. A video content production system as described in claim 19, wherein the simulation execution unit executes the simulation by selecting one of the multiple cameraworks generated by the generative AI model.
22. The video content production system of claim 19, wherein the simulation execution unit changes the simulation in real time by changing the camerawork generated by the generative AI model based on interaction with the user.
23. The video content production system according to claim 22, wherein the interaction with the user is a multimodal interaction.
24. The video content production system according to claim 23, wherein the multimodal interaction is an interaction using voice, images, video, handwriting input, text, or gestures.
25. The video content production system of claim 24, wherein the multimodal interactions are converted into corresponding prompts and input to the generative AI model.
26. The video content production system according to claim 19, wherein the content is expressed by at least one of an image and a video.
27. The video content production system of claim 26, wherein the content includes images or videos related to the placement of subjects, weather, buildings, or time of day in a filmed scene of a movie, animation, game, live concert, or live sports broadcast.
28. The video content production system according to claim 27, wherein the subjects include actors and extras, the weather includes sunny, cloudy, rainy, and snowy, and the time of day includes morning, noon, evening, and night.
29. A video content production method comprising: a video content production system generating a digital twin that models the environment captured by a camera mounted on a movable robot based on images or audio output from the camera; causing a generative AI model to generate at least one of the camera work of the camera and content in the digital twin based on instructions from a user; running a simulation of at least one of the camera work and the content in the digital twin; and changing the camera work or the content generated by the generative AI model based on interaction with the user, thereby changing the simulation in real time.
30. A movable robotic device equipped with a camera, comprising a control unit having control parameters of the movable robotic device, wherein the control unit transmits images or audio to a digital twin that models the environment photographed by the camera, and changes the control parameters in the digital twin according to the learning results of a learning model that is based on simulation results regarding at least one of content generated by a generative AI model or the camerawork of the camera.
Citation Information
Patent Citations
Workpiece overturning gripper structure
CN216140858U
Camera operation simulation device and program thereof, and camera image generation device and program thereof
JP2023070220A
Information processing device, information processing method, and program
WO2023286367A1