Virtual camera control method, device, equipment, storage medium and product

By acquiring the skeletal motion data and target scene size of virtual objects, the target camera position of the virtual camera is determined, solving the problem of low efficiency in switching camera scene sizes in virtual scenes and realizing rapid positioning and efficient image acquisition of virtual cameras.

CN116485879BActive Publication Date: 2026-04-24TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2023-01-03
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

In existing technologies, switching camera framing in virtual scenes requires repeated manual settings, resulting in high labor costs and low control efficiency.

Method used

By acquiring the skeletal motion data and target scene size of virtual objects in the virtual scene, the target camera position of the virtual camera is determined, and the virtual camera is controlled to capture images based on this, thus achieving rapid positioning and image capture of the virtual camera.

Benefits of technology

It enables rapid positioning and image acquisition of virtual cameras, improving control efficiency and reducing labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116485879B_ABST
    Figure CN116485879B_ABST
Patent Text Reader

Abstract

The application provides a virtual camera control method and device, electronic equipment, a computer readable storage medium and a computer program product, comprising: obtaining bone motion data corresponding to the bones of a virtual object in a virtual scene, and obtaining a target scene type corresponding to a virtual camera; the target scene type is used to indicate the proportion of different parts of the virtual object occupying in the collection picture of the virtual camera; based on the target scene type and the bone motion data, a target camera position corresponding to the target scene type of the virtual camera is determined; based on the target camera position, the virtual camera is controlled to collect a picture of the virtual object, and a target picture including the virtual object is obtained. Through the application, the rapid positioning of the virtual camera can be realized, so that the picture collection operation of the virtual camera is quickly performed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to computer vision technology, and more particularly to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for controlling a virtual camera. Background Technology

[0002] Computer vision (CV) is a type of artificial intelligence technology. It uses cameras and computers to replace human eyes in tasks such as target recognition, tracking, and measurement, and further performs image processing to make the computer-processed images more suitable for human observation or transmission to instruments for detection.

[0003] In related technologies, virtual scenes often contain a large number of camera shots with different framing and movement. Switching between camera framing often requires repetitive and simple manual settings, resulting in high labor costs and low camera control efficiency. Summary of the Invention

[0004] This application provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for controlling a virtual camera, which can not only achieve rapid positioning of the virtual camera, but also quickly perform image acquisition operations of the virtual camera.

[0005] The technical solution of this application embodiment is implemented as follows:

[0006] This application provides a method for controlling a virtual camera, including:

[0007] Obtain the skeletal motion data corresponding to the skeleton of the virtual object in the virtual scene, and obtain the target scene corresponding to the virtual camera;

[0008] The target scene size is used to indicate the proportion of different parts of the virtual object in the captured image of the virtual camera;

[0009] Based on the target scene size and the skeletal motion data, the target camera position of the virtual camera corresponding to the target scene size is determined;

[0010] Based on the target camera position, the virtual camera is controlled to capture images of the virtual object, thereby obtaining a target image including the virtual object.

[0011] This application provides a control device for a virtual camera, including:

[0012] The acquisition module is used to acquire skeletal motion data corresponding to the skeleton of a virtual object in a virtual scene, and to acquire the target scene corresponding to the virtual camera; wherein, the target scene is used to indicate the proportion occupied by different parts of the virtual object in the captured image of the virtual camera;

[0013] The determination module is used to determine the target camera position of the virtual camera corresponding to the target scene based on the target scene and the skeletal motion data;

[0014] The control module is used to control the virtual camera to capture images of the virtual object based on the target camera position, so as to obtain a target image including the virtual object.

[0015] This application provides an electronic device, including:

[0016] Memory, used to store executable instructions;

[0017] The processor, when executing executable instructions stored in the memory, implements the virtual camera control method provided in the embodiments of this application.

[0018] This application provides a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are executed by a processor, the processor will execute the virtual camera control method provided in this application.

[0019] This application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. The processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the virtual camera control method provided in this application.

[0020] The embodiments of this application have the following beneficial effects:

[0021] By applying the embodiments of this application, combining the skeletal motion data of virtual objects in a virtual scene with the target scene size of the virtual camera, the target camera position that matches the target scene size of the virtual camera is determined in real time, and the virtual camera is controlled to capture the target image including the virtual object based on the target camera position. In this way, not only can the virtual camera be quickly located, but the image capture operation of the virtual camera can also be performed quickly. Attached Figure Description

[0022] Figure 1 This is a schematic diagram of the architecture of the virtual camera control system 100 provided in an embodiment of this application;

[0023] Figure 2This is a schematic diagram of the structure of the electronic device 500 that implements the virtual camera control method according to an embodiment of this application;

[0024] Figure 3 This is a flowchart illustrating the virtual camera control method provided in an embodiment of this application;

[0025] Figure 4 This is a schematic diagram of the scene visualization interface provided in the embodiments of this application;

[0026] Figure 5 This is a flowchart of the skeletal motion data acquisition method provided in the embodiments of this application;

[0027] Figure 6 This is a schematic diagram of the detection of key points of the human skeleton provided in an embodiment of this application;

[0028] Figure 7 This is a flowchart of the method for determining the target camera position provided in an embodiment of this application;

[0029] Figure 8 This is another flowchart of the method for determining the target camera position provided in the embodiments of this application;

[0030] Figure 9 This is a flowchart illustrating the specific implementation method for determining the target camera position provided in the embodiments of this application;

[0031] Figure 10 This is a flowchart of a method for determining the position of a target camera when the target scene is panoramic, provided in an embodiment of this application.

[0032] Figure 11 This is a flowchart illustrating the target image acquisition method provided in the embodiments of this application;

[0033] Figure 12 This is a flowchart illustrating the method for creating an animation sequence for a virtual camera according to an embodiment of this application;

[0034] Figure 13 This is an overall flowchart of the virtual camera control method provided in the embodiments of this application;

[0035] Figure 14 This is a flowchart of the skeletal structure detection method provided in the embodiments of this application;

[0036] Figure 15 These are structural diagrams of different types of skeletal assets provided in the embodiments of this application;

[0037] Figure 16 This is a schematic diagram of the skeletal tree structure provided in an embodiment of this application;

[0038] Figure 17 This is a flowchart of virtual camera settings and scene selection provided in the embodiments of this application;

[0039] Figure 18 This is a flowchart of the animation sequence creation process for a virtual camera provided in an embodiment of this application;

[0040] Figure 19 This is a schematic diagram of the captured image from the virtual camera provided in this application embodiment. Detailed Implementation

[0041] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0042] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0043] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0044] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0045] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0046] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0047] 1) Virtual Human: A virtual digital human is a three-dimensional model synthesized by simulating the skeletal structure and appearance of a real human body through digital technology.

[0048] 2) Human Model: A virtual human body is a flexible and complex non-rigid object with many specific features such as motion structure, body shape, surface texture, and the position of body parts or joints. A mature human model does not necessarily need to include all human attributes, but should meet the requirements of the specific task of constructing and describing human posture.

[0049] 3) There are three commonly used human body models in human pose estimation: skeleton-based models, contour-based models, and volume-based models.

[0050] 4) Human pose estimation: Based on input data such as images and videos, it locates human body parts and establishes human representation (e.g., human skeleton), and is widely used in applications including human-computer interaction, motion analysis, augmented reality and virtual reality.

[0051] Based on the above explanation of the nouns and terms used in the embodiments of this application, the control system of the virtual camera provided in the embodiments of this application is described below. See also Figure 1 , Figure 1 This is a schematic diagram of the architecture of the virtual camera control system 100 provided in the embodiments of this application. In order to support an exemplary application, the terminal (terminal 400-1 and terminal 400-2 are shown as examples) connects to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two, and data transmission is achieved using wireless or wired links.

[0052] In some embodiments, the terminal (such as terminal 400-1 and terminal 400-2) is equipped with a virtual scene application (such as virtual human performance, virtual live broadcast, etc.) to send a scene data request for the virtual scene to the server and receive the animation sequence of the virtual camera in the virtual scene returned by the server, and render the virtual scene based on the animation sequence.

[0053] In some embodiments, server 200 is configured to acquire skeletal motion data corresponding to the skeleton of a virtual object in a virtual scene, and acquire the target scene size corresponding to the virtual camera; the target scene size indicates the proportion occupied by different parts of the virtual object in the captured image of the virtual camera; based on the target scene size and skeletal motion data, determine the target camera position corresponding to the target scene size of the virtual camera; based on the target camera position, control the virtual camera to capture images of the virtual object, obtaining a target image including the virtual object. Simultaneously, an animation sequence is created and saved for the entire process of the virtual camera switching from the initial camera position to the target camera position in the virtual scene. Upon receiving a scene data request for scene data of the virtual scene sent by the terminal, the server finds the animation sequence corresponding to the virtual scene and returns the animation sequence to the terminal.

[0054] In practical applications, server 200 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminals (such as terminals 400-1 and 400-2) can be smartphones, tablets, laptops, desktop computers, smart speakers, smart TVs, smartwatches, etc., but are not limited to these. Terminals (such as terminals 400-1 and 400-2) and server 200 can be directly or indirectly connected via wired or wireless communication; this application does not impose any restrictions on this connection.

[0055] The electronic device implementing the virtual camera control method provided in the embodiments of this application will now be described. See also Figure 2 , Figure 2 This is a schematic diagram of the structure of an electronic device 500 that implements a virtual camera control method according to an embodiment of this application. The electronic device 500 can be... Figure 1 The server 200 and electronic device 500 shown can also be terminals capable of implementing the virtual camera control method provided in this application, with electronic device 500 as the example. Figure 1 Taking the server shown as an example, an electronic device implementing the virtual camera control method of this application embodiment will be described. The electronic device 500 provided in this application embodiment includes: at least one processor 510, a memory 550, at least one network interface 520, and a user interface 530. The various components in the electronic device 500 are coupled together through a bus system 540. It is understood that the bus system 540 is used to realize the connection and communication between these components. In addition to a data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 2 The general labeled all buses as Bus System 540.

[0056] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0057] User interface 530 includes one or more output devices 531 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 530 also includes one or more input devices 532, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0058] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 550 may optionally include one or more storage devices physically located away from the processor 510.

[0059] The memory 550 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 550 described in this application embodiment is intended to include any suitable type of memory.

[0060] In some embodiments, memory 550 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0061] Operating system 551 includes system programs for handling various basic system services and performing hardware-related tasks, such as framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks; network communication module 552 is used to reach other computing devices via one or more (wired or wireless) network interfaces 520, exemplary network interfaces 520 include: Bluetooth, WiFi, and Universal Serial Bus (USB), etc.; presentation module 553 is used to enable the presentation of information (e.g., user interface for operating peripheral devices and displaying content and information) via one or more output devices 531 associated with user interface 530 (e.g., display screen, speaker, etc.); input processing module 554 is used to detect and translate one or more user inputs or interactions from one or more input devices 532.

[0062] In some embodiments, the virtual camera control device provided in this application can be implemented in software. Figure 2A control device 555 for a virtual camera stored in memory 550 is shown. It can be software in the form of programs and plug-ins, including the following software modules: acquisition module 5551, determination module 5552 and control module 5553. These modules are logical and can therefore be arbitrarily combined or further divided according to the functions they implement. The functions of each module will be described below.

[0063] In other embodiments, the virtual camera control device provided in this application can be implemented using a combination of hardware and software. As an example, the virtual camera control device provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the virtual camera control method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0064] In some embodiments, the terminal or server can implement the virtual camera control method provided in this application by running a computer program. For example, the computer program can be a native program or software module in an operating system; it can be a native application (APP), that is, a program that needs to be installed in the operating system to run, such as an instant messaging APP or a web browser APP; it can also be a mini-program, that is, a program that only needs to be downloaded into the browser environment to run; or it can be a mini-program that can be embedded in any APP. In short, the above-mentioned computer program can be any form of application, module or plugin.

[0065] Based on the above description of the control system and electronic device for the virtual camera provided in the embodiments of this application, the control method for the virtual camera provided in the embodiments of this application is described below. In actual implementation, the control method for the virtual camera provided in the embodiments of this application can be implemented by the terminal or the server alone, or by the terminal and the server working together, so that... Figure 1 The following description uses the example of server 200 executing the virtual camera control method provided in this application embodiment. See also... Figure 3 , Figure 3 This is a flowchart illustrating the virtual camera control method provided in the embodiments of this application, which will be combined with... Figure 3The steps shown are explained.

[0066] In step 101, the server obtains the skeletal motion data corresponding to the skeleton of the virtual object in the virtual scene, and obtains the target scene corresponding to the virtual camera. The target scene is used to indicate the proportion of different parts of the virtual object in the captured image of the virtual camera.

[0067] In actual implementation, the server is equipped with Unreal Engine for virtual scene development. Through visual operations in Unreal Engine, the control process of the virtual camera can be realized, and the movement process of the virtual camera in the virtual scene can be stored as an animation sequence. When a scene data request for the virtual scene is received from the terminal, the animation sequence is sent to the terminal so that the terminal can render and display the virtual scene based on the animation sequence.

[0068] In practice, the image captured by a virtual camera for a virtual object is related to the framing; different framings correspond to different captured images. When the focal length of the virtual camera is fixed, the varying distances between the virtual camera and the virtual object in the virtual scene result in different ranges of the virtual object appearing in the camera's view. These different ranges are called the framings of the virtual camera. In order of distance from the virtual object, from closest to farthest, the framings of a virtual camera include at least one of the following: close-up, medium shot, long shot, full shot, and long shot. See also... Figure 4 , Figure 4 This is a schematic diagram of the visualization interface for different shot types provided in the embodiments of this application. When the shot type is a close-up, the bottom of the virtual camera's captured image of the virtual object is positioned above the shoulder of the human model, meaning the captured image shows the human model corresponding to the virtual object above the shoulder. When the shot type is a medium shot, the bottom of the virtual camera's captured image of the virtual object is positioned above the chest of the human model, meaning the captured image shows the human model corresponding to the virtual object above the chest. When the shot type is a medium shot, the bottom of the virtual camera's captured image of the virtual object is positioned above the hip of the human model, meaning the captured image shows the human model corresponding to the virtual object above the hip. When the shot type is a wide shot, the captured image of the virtual camera includes the entire human body and the surrounding environment. When the shot type is a long shot, the captured image of the virtual camera mainly includes the environment in which the virtual object is located.

[0069] In some embodiments, the target shot can be determined by displaying at least one selectable shot, including at least one of the following: close-up, medium shot, full shot, and long shot; and in response to a shot selection operation for at least one shot, determining the selected shot as the target shot.

[0070] In practice, the server's Unreal Engine displays one or more scene types to choose from through the user interface. Upon receiving a user's selection of a scene type based on the current interface, the server uses the selected scene type as the target scene type. It should be noted that preset scene types can also be read from the application's configuration file for the virtual scene and used as the target scene type.

[0071] For example, the server's Unreal Engine can provide a human-computer interaction interface for selecting the scene type. In the human-computer interaction interface, multiple scene type selection options are displayed, each scene type selection option corresponds to a scene type. When a selection operation is received for the "close-up" selection option, the "close-up" scene type corresponding to the "close-up" selection option is used as the target scene type of the virtual camera in the current virtual scene.

[0072] In practice, once the target shot size of the virtual camera is determined, the distance between the virtual object being filmed and the virtual camera can be determined. Therefore, based on the skeletal motion data corresponding to the skeleton of the virtual object, the camera parameters of the virtual camera can be automatically adjusted so that the virtual camera is positioned at the target camera position corresponding to the target shot size. Since the virtual objects in the virtual scene are in motion, the server can obtain the skeletal motion data corresponding to the skeleton of the virtual object in the virtual scene through different acquisition methods, thereby determining the distance between the virtual object and the virtual camera in real time. This allows for real-time adjustment of the virtual camera's position, ultimately determining the target camera position corresponding to the target shot size. It should be noted that the skeletal motion data includes both displacement and rotation data of the skeletons.

[0073] The method for acquiring skeletal motion data is described below. In some embodiments, see [link to documentation]. Figure 5 , Figure 5 This is a flowchart of a skeletal motion data acquisition method provided in an embodiment of this application, based on... Figure 3 Step 101 can be implemented by steps 1011-1013, combined with Figure 5 The steps shown are explained.

[0074] Step 1011: The server detects the virtual resources in the virtual scene and obtains the detection results.

[0075] In practice, the Unreal Engine on the server can provide various types of virtual resources for creating virtual objects. These virtual resources can include skeletal resources for creating the human body model associated with the virtual object. Skeletal resources include the skeletal architecture of the human body model provided by the virtual engine, the number of bones in the skeletal structure, and the name of each bone. Different types of skeletal resources can create virtual objects with skeletons that can be driven by real-time or offline animation data.

[0076] For example, taking Unreal Engine as an example, three types of skeletal resources, namely Actor, Character, and BP, can be detected in the virtual scene rendered based on Unreal Engine.

[0077] Step 1012: When the detection result indicates that the virtual resources include bone resources for creating virtual objects, the skeletons of the virtual objects in the virtual scene are created based on the bone resources, and the bone motion data of the skeletons are initialized to obtain the bone motion data corresponding to the skeletons of the virtual objects.

[0078] For example, taking Unreal Engine as an example, in a virtual scene rendered by Unreal Engine, three types of skeletal resources can be detected: Actor, Character, and Backbone. Different skeletal resources have corresponding skeletal trees. Based on the skeletal resources, a human body model associated with the corresponding virtual object is constructed, and the movement of the skeletal movement is driven by assigning positional and rotational information to each skeletal point in the human body model.

[0079] Step 1013: When the detection result indicates that there is no bone resource in the virtual resource for creating the bone of the virtual object, perform human skeleton key point detection on the virtual object in the virtual scene to obtain the bone key points carrying the bone information, and determine the bone motion data of the corresponding bone based on the bone key points.

[0080] In practice, when the skeletal resources of the virtual objects provided by Unreal Engine cannot be detected, the human skeleton key point detection algorithm can be used to detect the human skeleton key points of the virtual objects in the virtual scene, obtain the skeletal key points carrying the skeletal information, and thus obtain the position and rotation data of different bones of the virtual object.

[0081] For example, see Figure 6 , Figure 6 This is a schematic diagram illustrating the detection of key points in the human skeleton provided in an embodiment of this application. The human model associated with the virtual object shown in Figure 2 is a virtual object built using a programming language (such as C++), meaning it is not created using virtual resources provided by the Unreal Engine. In this case, the skeletal resources used to create the virtual object cannot be obtained using the skeletal resource detection method built into the Unreal Engine. Therefore, a human skeleton key point detection algorithm (such as Open Pose) can be used to detect the key points in the virtual object, obtaining the skeletal key points carrying skeletal information. Figure 1 shows the standard skeletal points corresponding to the human skeleton key point detection algorithm. Then, based on the displacement and rotation data of the skeletal key points, the skeletal motion data of the corresponding bones is determined.

[0082] In step 102, the target camera position corresponding to the target scene is determined based on the target scene size and skeletal motion data.

[0083] In actual implementation, the server determines the distance between the virtual object and the virtual camera based on the received target scene and the skeletal motion data of the virtual object detected in the virtual scene, thereby determining the target camera position that matches the target scene.

[0084] In some embodiments, see Figure 7 , Figure 7 This is a flowchart of a method for determining the target camera position provided in an embodiment of this application, based on... Figure 3 Step 102 can be implemented by steps 1021 to 1023, combined with Figure 7 The steps shown are explained.

[0085] Step 1021: Obtain the position information of the virtual object in the virtual scene, and obtain the target height of the virtual camera and the focal length of the virtual camera.

[0086] In practice, under the target scene, the distance between the virtual camera and the virtual object is related to the camera's in-camera parameters and the virtual object's position in the virtual scene. The in-camera parameters are those related to the camera's own characteristics, such as the camera's focal length and target size (including the target's height and width).

[0087] Step 1022: Select the target skeleton from the skeleton of the virtual object that is compatible with the target scene.

[0088] The target skeleton refers to the skeleton of the lowest part of the virtual object captured when the virtual camera captures the image of the virtual object based on the target field of view.

[0089] In practice, the virtual camera will present different body parts of the virtual object in the captured image of the virtual object under different target scene conditions.

[0090] In some embodiments, see Figure 8 , Figure 8 This is another flowchart of the method for determining the target camera position provided in the embodiments of this application, based on Figure 7 Step 1022 can be implemented by steps 201 to 203, combined with Figure 8 The steps shown are explained.

[0091] Step 201: When the target shot is a close-up, select a target bone from at least one bone in the shoulder of the virtual object that is compatible with the close-up shot.

[0092] In actual implementation, when the target shot is a close-up, the lowest part of the virtual object captured in the virtual camera's view is the shoulder. At this time, the target skeleton that matches the close-up is the shoulder skeleton.

[0093] For example, see Figure 6 The key points of the human skeleton shown in the figure are the shoulder when the target shot is a close-up. The shoulder bone points shown in the figure are bone points 2 and 5. It can be determined that one of bone points 2 and 5 is the target bone when the shot is a close-up.

[0094] Step 202: When the target shot is a close-up, select a target bone from at least one bone in the chest of the virtual object that is compatible with the close-up shot.

[0095] In actual implementation, when the target shot is a close-up, the lowest part of the virtual object captured in the virtual camera's image is the chest. At this time, the target skeleton that matches the close-up is the chest skeleton.

[0096] Step 203: When the target shot is a medium shot, select a target bone that is compatible with the medium shot from at least one bone in the hip of the virtual object.

[0097] In actual implementation, when the target shot is a medium shot, the lowest part of the virtual object captured in the virtual camera's view is the hip. At this time, the target skeleton that matches the medium shot is the hip skeleton, etc.

[0098] For example, see Figure 6 The key points of the human skeleton shown in the figure are number 1. When the target scene is medium shot, the target part of the corresponding virtual object is the hip. The hip bone points shown in the figure are bone points 8, 9 and 12. Any one of the three can be used as the target bone when the scene is medium shot.

[0099] It should be noted that, Figure 7 There is no strict execution order between steps 201-203.

[0100] Step 1023: Based on the location information, the skeletal motion data of the target skeleton, the target surface height, and the focal length, determine the target camera position corresponding to the target scene of the virtual camera.

[0101] In practice, the server determines the target camera position for the corresponding target scene based on the virtual object's position information in the virtual scene, the target skeleton's skeletal motion data, the target surface height, and the focal length. Specifically, the virtual object's position information is typically determined by selecting the anchor point of the human model associated with the virtual object as the center point, which can be denoted as... And determine the skeletal motion data of the target bone. For example... Figure 6The anchor point of the human body model shown in number 1 can be bone point 8, or the center point between bone point 1 and bone point 8 can be used as the anchor point of the human body model associated with the virtual object.

[0102] In some embodiments, the server may specifically determine the target camera position of the virtual camera in the following ways: based on the skeletal motion data of the target skeleton, determine the height between the target skeleton and the skeleton of the virtual object's head; based on the height, target height, and focal length, determine the distance from the virtual camera to the virtual object; based on the position and distance, determine the target camera position of the virtual camera corresponding to the target scene.

[0103] In practice, the first step is to determine the height between the target skeleton, which is compatible with the target scene size, and the head skeleton of the human model associated with the virtual object. It should be noted that... (See...) Figure 6 The key points of the human skeleton shown in Figure 1 include bone points 15 and 16 for the head. Select one bone point from 15 and 16, and determine the height *l* between the target bone and that bone point. It should be noted that the selected head bone point is some distance from the top of the virtual object's head. Therefore, in practical applications, the height *l* between the target bone and that bone point can be increased by a known preset value of 5cm as the height between the target bone and the top of the head. Then, based on the target surface of the virtual camera... ,focal length The distance between the virtual camera and the virtual object is calculated. The position and rotation information of the virtual camera in the world coordinate system are calculated based on this distance, thereby determining the target camera position that matches the virtual camera with the target shot size.

[0104] In some embodiments, the server can determine the distance between the virtual object and the virtual camera by: obtaining the ratio between the target height and the focal length; determining the product between half the height and the ratio; and using the product as the distance from the virtual camera to the virtual object.

[0105] In actual implementation, the target surface of the virtual camera ,focal length Implementation method (the location of the virtual object is) The height of the human body model associated with the virtual object is The height of the target skeleton from the top of the head, which is appropriate for the target shot size, is The distance between the virtual object and the virtual camera is First, obtain the ratio between the target height and the focal length. Then, according to The ratio between them determines the distance d between the virtual object and the virtual camera, i.e. .

[0106] The height of the target skeleton from the top of the head, which is adapted to the target shot size, is To clarify, in actual implementation, when the target shot is a close-up, the height of the target skeleton from the top of the head... The height of the shoulder bones from the top of the head; when the target shot is close-up, the height of the target bones from the top of the head. The height of the sternum from the top of the head; when the target is in a medium shot, the height of the target skeleton from the top of the head. This refers to the height of the hip bones from the top of the head.

[0107] In some embodiments, see Figure 9 , Figure 9 This is a flowchart illustrating the specific implementation method for determining the target camera position provided in the embodiments of this application, combined with... Figure 9 The steps shown are explained.

[0108] Step 301: Obtain the first component of the position relative to the first direction of the world coordinate system, the second component of the position relative to the second direction of the world coordinate system, and the third component of the position relative to the third direction of the world coordinate system.

[0109] Among the first direction, the second direction, and the third direction, any two directions are perpendicular to each other.

[0110] For example, determining the position of a virtual object in a virtual scene. It should be noted that the position here is relative to the world coordinate system, where the position is relative to the first component of the first direction of the world coordinate system. The second component of the second direction relative to the world coordinate system And the third component in the third direction relative to the world coordinate system. .

[0111] Step 302: Perform a weighted summation of the second component and the distance to obtain the target component of the virtual camera in the second direction relative to the world coordinate system;

[0112] Continuing from the previous example, keeping the first and third components unchanged, what about the second component? The distance between the virtual object and the virtual camera is Perform a weighted summation, that is As the target component of the virtual camera in the second direction relative to the world coordinate system.

[0113] Step 303: Based on the first component, the target component, and the third component, determine the target camera position corresponding to the target scene of the virtual camera.

[0114] Continuing from the previous example, set the target camera position of the virtual camera to... ,in, With the first component equal, With target component equal, With the third component equal.

[0115] In some embodiments, see Figure 10 , Figure 10 This is a flowchart of a method for determining the target camera position when the target scene is panoramic, provided in an embodiment of this application. When the target scene is panoramic, the server can determine the target camera position through steps 401-403.

[0116] Step 401: The server obtains the position of the virtual object in the virtual scene and the height of the virtual object, and obtains the target height of the virtual camera and the focal length of the virtual camera.

[0117] In practice, when the target shot is panoramic, the virtual camera captures an image of the virtual object from head to toe. Since there is no fixed height limit for the virtual camera's captured image of the virtual object in panoramic mode, the target camera position can be determined based on the height h of the associated human model.

[0118] For example, a height between 1 and 1.5 times the model height h, i.e., a height between h and 1.5h, can be used as a reference height when calculating the panorama.

[0119] Step 402: Based on the height, determine a reference height that is compatible with the panorama, and obtain the ratio between the target surface height and the focal length.

[0120] In practice, the height h of the human body model associated with the virtual object, and a height of 1-1.5 times the height h, are used as reference heights to adapt to the panorama, based on the target surface size of the virtual camera. Target height in The focal length of a virtual camera Determine the target camera position for the virtual camera.

[0121] Step 403: Based on the reference height and ratio, determine the target camera position corresponding to the target scene of the virtual camera.

[0122] In actual implementation, the server obtains the target height. The focal length of a virtual camera The ratio between And using 1.2 times the model height h, 1.2h, as the reference height, through... Determine the distance between the virtual camera and the virtual object. And based on the location information of the virtual object in the virtual scene. Determine the target camera position as ,in, , , = That is, in the world coordinate system, .

[0123] It should be noted that when the target scene is a distant view, the virtual camera's captured image mainly includes the environment in which the virtual object is located. The distance d between the virtual camera and the virtual object is not fixed; it can obtain a height h that is 1-1.5 times the height h of the human model associated with the virtual object. Taking 1.5h as an example, according to... Determine the distance between the virtual camera and the virtual object. Based on preset coefficients and the distance between the virtual camera and the virtual object, Taking a preset coefficient of 5 as an example, the target camera position of the virtual camera can be determined. .

[0124] In step 103, based on the target camera position, the virtual camera is controlled to capture images of the virtual object, thereby obtaining a target image including the virtual object.

[0125] In actual implementation, once the target camera position of the virtual camera is determined, the virtual camera is controlled to capture images of the virtual object, thereby obtaining a target image that matches the target shot size.

[0126] The specific methods for acquiring the target image are described below. In some embodiments, see [link to documentation]. Figure 11 , Figure 11 This is a flowchart illustrating the target image acquisition method provided in the embodiments of this application, combined with... Figure 11 The steps shown are explained.

[0127] Step 1031: Obtain the initial camera position of the virtual camera and adjust the initial camera position of the virtual camera to the target camera position.

[0128] In practical implementation, within a virtual scene, the virtual camera has an initial camera position and an initial shot size. The virtual camera captures images of the virtual object from the initial camera position, obtaining a captured image adapted to the initial shot size. When the virtual camera's shot size changes, the camera position is adjusted from the initial camera position to the target camera position.

[0129] Step 1032: Based on the target camera position, control the virtual camera to capture images of the area between the target skeleton and the top of the head of the virtual object, and obtain a target image including the virtual object.

[0130] In actual implementation, the virtual camera is positioned at the target camera position, and the image of the virtual object is captured from the area between the top and bottom of the head of the target skeleton to obtain the target image including the virtual object.

[0131] In some embodiments, the server may also adjust the target camera position by receiving a scene switching instruction, adjusting the target camera position to obtain the adjusted camera position, and then, based on the adjusted camera position, re-controlling the virtual camera to capture images of the virtual object to obtain a new target image including the virtual object.

[0132] In actual implementation, the Unreal Engine on the server receives a scene-switching command, adjusts the target camera position of the current virtual camera to obtain an adjusted camera position that matches the scene-switching command. Then, it controls the virtual camera in the virtual scene to perform image capture operations on the virtual objects from the adjusted camera position.

[0133] In some embodiments, see Figure 12 , Figure 12 This is a flowchart illustrating the creation method of an animation sequence for a virtual camera provided in this application embodiment, combined with... Figure 12 The steps shown are explained below.

[0134] Step 501: Obtain the initial shot size of the virtual camera and create a first animation sequence of the virtual camera that matches the initial shot size.

[0135] The first animation sequence is used to characterize the movement of the virtual camera in the virtual scene during the process of the virtual camera capturing images of the virtual object based on the initial view scale.

[0136] In practice, Unreal Engine on the server can store the movement of a virtual camera in a virtual scene using animation sequences. In Unreal Engine's animation editor, an animation sequence adapted to the initial scene size is created and saved in the path used to store the virtual resources of the current virtual scene. The animation sequence records the trajectory of the virtual camera in the virtual scene; the camera's movement can be represented by changes in its focal length and rotation. Specifically, the camera's focal length trajectory is determined by the change in its focal length. After obtaining the distance between the virtual camera and the virtual object, the target camera position is determined based on the virtual camera's focal length, and the rotation direction of the virtual object and the rotation trajectory of the virtual camera are determined based on the bone rotation data in the virtual object's skeletal motion data.

[0137] Step 502: Determine the time point at which the virtual camera captures images of the virtual object based on the target camera position, and create a second animation sequence that matches the virtual camera with the target scene size starting from the time point.

[0138] In actual implementation, once the target camera position that matches the target shot size is determined, the virtual camera is controlled to start capturing images of the virtual object and record the current time point. Starting from the current time point, a second animation sequence of the virtual camera under the target shot size is created.

[0139] Step 503: Merge the first animation sequence and the second animation sequence to obtain the target animation sequence of the virtual camera. The target animation sequence is used to render the virtual scene.

[0140] In actual implementation, the Unreal Engine on the server merges the first and second animation sequences based on the merge instruction for the animation sequence to obtain the animation sequence of the entire process of the virtual camera from the initial view to the target view, that is, the target animation of the virtual camera. This target animation can be used to render and display the virtual scene during the rendering stage of the virtual scene.

[0141] It should be noted that when there are multiple virtual cameras in a virtual scene, each virtual camera has a corresponding animation sequence. When a switching command for a virtual camera is received, different virtual cameras can capture images of virtual objects in the virtual scene to obtain images of the same virtual object captured by different virtual cameras under the same target view.

[0142] By applying the above embodiments of this application, combining the virtual camera's focal length, target size, and target scene type, the target camera position of the virtual camera at the target scene type is automatically calculated, and a corresponding animation sequence of the virtual camera is generated. Thus, when the scene type of the virtual camera changes, there is no need to manually adjust the virtual camera's position and parameters based on the virtual camera's preview view. The target camera position and corresponding animation sequence are automatically generated based on the target size setting, focal length setting, and scene type selection, effectively reducing the operational complexity of the virtual camera and simplifying its control when the scene type changes. Furthermore, considering various forms of skeletal resources while automatically setting the scene type, it features high automation and strong applicability.

[0143] The following will describe an exemplary application of the embodiments of this application in a real-world application scenario.

[0144] In related technologies, the control of virtual cameras in game scenes is mainly based on the camera movement rules that control the player's perspective, and does not support the framing and camera movement functions of different shot sizes; while the control of virtual cameras in live streaming scenes is mainly through the control of camera directing rules, and can judge shot size information, but cannot control the lens. At the same time, the subjects being filmed are often real scenes and people, which are different from purely virtual environments.

[0145] Based on this, this application provides a virtual camera control method, which mainly uses the spatial information of a virtual human (i.e., the virtual object mentioned above) to automatically control the lens framing of the virtual camera. The spatial information of the virtual human refers to the skeletal motion information of the human body model associated with the virtual object in the virtual scene. This method can quickly set the camera position based on the virtual human model during virtual human performances and live broadcasts to complete image capture under different framing conditions. Furthermore, it is compatible with the composition methods of various virtual human skeletal structures and corresponding skeleton hierarchical structures in different Unreal Engines. It can automatically calculate the framing for various virtual human assets, including those that comply with skeletal resource rules and those that do not, enabling rapid virtual camera lens framing operations.

[0146] The technical aspects of the virtual camera control method provided in this application's embodiments are explained below; see [link to relevant documentation]. Figure 13 , Figure 13 This is an overall flowchart of the virtual camera control method provided in this application embodiment. The specific implementation process is as follows: 1. Input the target scene type, execute step 2, standard skeletal asset detection. When standard skeletal assets exist in the virtual scene, execute step 3.1, skeletal resource reading under standard skeletal assets. When standard skeletal assets do not exist in the virtual scene, execute step 3.2, non-standard skeletal asset switching and human pose recognition. After obtaining the corresponding virtual human skeletal information based on step 3.1 or step 3.2, execute step 4, obtain the corresponding skeletal information according to the target scene type. At this time, the skeletal information that matches the target scene type is obtained (i.e., the target skeleton mentioned above). Finally, execute step 5, create a camera animation sequence. Here, the camera is the virtual camera mentioned above. In actual implementation, the implementation process of the virtual camera control method is mainly divided into a skeletal structure detection module, a camera setting and scene type selection module, and an automated scene trajectory generation module.

[0147] First, the implementation of the skeletal structure detection module will be explained. (See [link / reference]). Figure 14 , Figure 14This is a flowchart of the skeletal structure detection method provided in this application embodiment. During implementation, the Unreal Engine used to create virtual scenes typically provides naming conventions for standard skeletal assets, their architecture, and the standardized names of each bone. Standard skeletal assets can be divided into two asset types: Skeletal Mesh Actor and Character. Non-standard skeletal assets may include BP objects built using Unreal Engine's Blueprint function, or other unknown virtual humans built using C++. All of these different types of skeletal assets can realize virtual humans (i.e., virtual objects mentioned earlier) with skeletons, driven by real-time or offline animation data, and these virtual humans are associated with human models. Therefore, when performing skeletal asset detection, the presence of the above four types of skeletal assets in the virtual scene is first detected. Here, skeletal assets refer to the skeletal resources mentioned earlier.

[0148] Taking Unreal Engine as an example, in a virtual scene rendered by Unreal Engine, three types of skeletal asset objects can be detected: Actor, Character, and Backbone (BP). These three types of assets are often composed of the following hierarchical structure, see [link to relevant documentation]. Figure 15 , Figure 15 This application provides different types of skeletal asset structure diagrams, showing a skeletal asset named TPose (Skeletal Mesh Actor). A skeletal MeshComponent is searched for among the three types of objects mentioned above; this component is the only one that can be used as a skeletal asset. For an explanation of the skeletal tree structure in Unreal Engine, see [link to documentation]. Figure 16 , Figure 16 This is a schematic diagram of the skeletal tree structure provided in an embodiment of this application. Figure 1 shows the skeletal tree structure of a Character type skeletal component, and figure 2 shows the skeletal tree structure of a Skeletal Mesh Actor component. The motion data of the virtual human drives the skeletal movement by assigning position and rotation information to each bone point in the human model. Therefore, the position and rotation data of each individual bone can be read and written.

[0149] During implementation, when Unreal Engine cannot detect the aforementioned different types of objects and skeletal components in a virtual scene, it can also use Open Pose, a computer vision algorithm, to detect key points on virtual human objects in the virtual scene, thereby obtaining the position and rotation data of different bones of the virtual object. The skeletal model formed by the Open Pose algorithm is as follows: Figure 6 As shown in number 1, the skeleton of the virtual human object is detected in real time in the virtual scene, such as... Figure 6 As shown in number 2, Unreal Engine can read the skeletal point information of non-standard objects based on the obtained skeleton.

[0150] This section explains the virtual camera settings and scene selection module. See [link / reference]. Figure 17 , Figure 17 This is a flowchart of virtual camera settings and scene selection provided in this application embodiment. The implementation process is as follows: Step 1, virtual camera settings, including camera object selection, camera target selection, and camera focal length settings; Step 2, scene selection, including selecting the target scene from close-up, medium shot, full shot, and long shot, and obtaining the corresponding skeleton position of the virtual person according to the selected target scene, such as the shoulder skeleton for close-up, the chest skeleton for medium shot, and the hip skeleton for medium shot; After determining the relevant settings of the virtual camera and the skeleton position corresponding to the target scene, Step 3, determining the distance between the camera and the virtual person being filmed; Step 4, obtaining the orientation of the virtual person being filmed in the world coordinate system; Step 5, calculating the position of the camera in the world coordinate system (i.e., the target camera position mentioned above). It should be noted that the execution of Step 1 and Step 2 is not sequential, and the execution of Step 3 and Step 4 is not sequential.

[0151] In practice, a virtual camera (i.e., the virtual camera mentioned earlier) is a simulation of a real camera in the computer world coordinate system. The parameters controlling the camera consist of two parts: intrinsic and extrinsic parameters. Intrinsic parameters include focus, focal length, camera target area, and camera distortion parameters, while extrinsic parameters refer to information such as camera rotation and displacement. The shot size is primarily distinguished by the proportion of the subject in the camera frame. Shot sizes can include close-up, medium shot, long shot, and extreme long shot, etc. (See [link to relevant documentation]). Figure 4 The diagram shows a visualization of the shot size.

[0152] In practice, the framing of the camera is determined by the camera's receiving surface, its focal length, and the distance between the camera and the subject. (The receiving surface of the camera...) ,focal length The system automatically calculates the distance between the camera and the subject using the three parameters mentioned above, including the target shot size. It then calculates the camera's position and rotation in the world coordinate system based on this distance, thereby determining the camera's external parameters and completing the automated framing operation.

[0153] This paper explains how to calculate the position of the camera in the world coordinate system under different shot sizes, assuming the model anchor point is the center and the position is... The model height is The corresponding bone distance from the top of the head is The distance between the camera and the model is The camera's position is When the shot is a close-up, the bottom of the frame should be positioned above the shoulders, with the shoulder bone a distance of 1 meter from the top of the head. Camera distance model The camera coordinates are = When the shot is a close-up, the bottom of the frame should be positioned at the chest level, with the sternum a distance of 1 meter from the top of the head. Camera distance model The camera coordinates are = When the shot is a medium shot, the bottom of the frame should be positioned at the hip area, with the hip bone spaced 1 meter from the top of the head. Camera distance model The camera coordinates are = When the shot is a full view, the subject is captured from head to toe. The model height *h* has no fixed range for full views, so any value between 1 and 1.5 times the model height can be used as the calculation standard. For example, 1.2h: Camera distance model The camera coordinates are = When the shot is a long shot, the camera distance is not fixed and is relatively large, for example, d=5h, etc., the camera coordinates are... = .

[0154] The explanation continues regarding the automated scene trajectory generation module. See [link / reference] Figure 18 , Figure 18This is a flowchart illustrating the creation process of a virtual camera animation sequence provided in this application embodiment. This module primarily creates camera animation assets and sets their parameters. First, step 1 is executed: creating an animation sequence asset. Using the asset tool, an animation sequence asset is selected and created as the main animation sequence. The save path and asset name are set. Next, step 2 is executed: selecting a camera. For the selected camera, step 2.1 is executed to create a camera animation asset sequence. When multiple cameras exist, step 2.2 is executed: creating a camera shot switching sequence. This asset and the main animation sequence are placed in the same file directory, and the sub-animation sequences are nested within the main animation sequence. When automatic shot setting is completed for multiple cameras, multiple camera sub-animation sequence assets are generated and included in the main animation sequence. Setting the start time of the sub-animation sequences allows for switching between different camera shot sizes. Step 3, defining camera motion trajectory parameters, mainly includes step 4.1, creating the camera focal length track, and step 4.2, creating the camera transformation track. In actual implementation, the automated camera sequence generated by the tool that automatically controls the virtual camera's lens framing primarily consists of the camera focal length track and the camera transformation track. The camera focal length track is determined by the user-inputted focal length, which is the default value for this track. Based on the distance between the camera and the subject calculated earlier, and the subject's rotation direction, the camera's position and orientation in the world coordinate system are calculated, setting the camera's transformation track. Finally, see... Figure 19 , Figure 19 This is a schematic diagram of the capture screen of the virtual camera provided in the embodiment of this application. The diagram shows the capture screen of the virtual camera on the virtual person when the target scene is a close-up.

[0155] Applying the above embodiments of this application has the following beneficial effects:

[0156] 1) Primarily focused on methods for quickly setting camera positions and designing different shot sizes within a virtual environment, such as virtual human performances and live streaming, using a virtual human model. It eliminates the need for manual adjustments to camera positions and parameters based on preview views; the corresponding camera animation sequences are automatically generated through simple camera selection, target setting, focal length setting, and shot size selection.

[0157] 2) There are no unified and fixed standards for the names and definitions of skeleton resources in different 3D production software and 3D rendering engines. It can be compatible with the structural composition methods and skeleton hierarchical structures of various virtual human objects under real-time 3D engines. It can automatically calculate the scene size for various virtual human assets, whether they comply with skeleton resource rules or not, and has high automation and strong applicability. At the same time, it can quickly perform virtual camera lens framing operations.

[0158] It should be noted that in the embodiments of this application, if the virtual scene involves data related to the user's attributes, when the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0159] The following description continues to illustrate the exemplary structure of the virtual camera control device 555 provided in the embodiments of this application as a software module. In some embodiments, such as... Figure 3 As shown, the software modules in the virtual camera control device 555 stored in the memory 550 may include:

[0160] The acquisition module 5551 is used to acquire the skeletal motion data corresponding to the skeleton of the virtual object in the virtual scene, and to acquire the target scene corresponding to the virtual camera; wherein, the target scene is used to indicate the proportion occupied by different parts of the virtual object in the captured image of the virtual camera;

[0161] The determining module 5552 is used to determine the target camera position of the virtual camera corresponding to the target scene based on the target scene and the skeletal motion data;

[0162] The control module 5553 is used to control the virtual camera to capture images of the virtual object based on the target camera position, so as to obtain a target image including the virtual object.

[0163] In some embodiments, the determining module is further configured to obtain the position of the virtual object in the virtual scene, and obtain the target height of the virtual camera and the focal length of the virtual camera; select a target skeleton from the skeleton of the virtual object that is compatible with the target scene; wherein, the target skeleton is the skeleton of the lowest part of the virtual object captured when the virtual camera captures the image of the virtual object based on the target scene; and determine the target camera position of the virtual camera corresponding to the target scene based on the position information, the skeletal motion data of the target skeleton, the target height, and the focal length.

[0164] In some embodiments, the determining module is further configured to determine the height between the target skeleton and the head skeleton of the virtual object based on the skeletal motion data of the target skeleton; determine the distance from the virtual camera to the virtual object based on the height, the target surface height, and the focal length; and determine the target camera position of the virtual camera corresponding to the target scene based on the position and the distance.

[0165] In some embodiments, the determining module is further configured to obtain the ratio between the target surface height and the focal length; determine the product between half of the height and the ratio, and use the product as the distance from the virtual camera to the virtual object.

[0166] In some embodiments, the determining module is further configured to obtain a first component of the position relative to a first direction of the world coordinate system, a second component of the position relative to a second direction of the world coordinate system, and a third component of the position relative to a third direction of the world coordinate system; wherein, any two of the first direction, the second direction, and the third direction are perpendicular to each other; the second component is weighted and summed with the distance to obtain the target component of the virtual camera relative to the world coordinate system in the second direction; and the target camera position of the virtual camera corresponding to the target scene is determined based on the first component, the target component, and the third component.

[0167] In some embodiments, the determining module is further configured to: when the target shot is a close-up, select a target bone from at least one bone of the shoulder of the virtual object that is compatible with the close-up; when the target shot is a medium shot, select a target bone from at least one bone of the chest of the virtual object that is compatible with the medium shot; and when the target shot is a medium shot, select a target bone from at least one bone of the hip of the virtual object that is compatible with the medium shot.

[0168] Accordingly, in some embodiments, the control module is further configured to control the virtual camera to capture images of the portion of the virtual object from the target skeleton to the top of the head, based on the target camera position, so as to obtain a target image including the virtual object.

[0169] In some embodiments, when the target scene is a panorama, the determining module is further configured to obtain the position of the virtual object in the virtual scene and the height of the virtual object, and obtain the target surface height of the virtual camera and the focal length of the virtual camera; determine a reference height adapted to the panorama based on the height; obtain the ratio between the target surface height and the focal length; and determine the target camera position of the virtual camera corresponding to the target scene based on the reference height and the ratio.

[0170] In some embodiments, the acquisition module is further configured to display at least one selectable shot type, the shot type including at least one of the following: close-up, medium shot, long shot, wide shot, and long shot; and in response to a shot type selection operation for the at least one shot type, to determine the selected shot type as the target shot type.

[0171] In some embodiments, the acquisition module is further configured to detect virtual resources in the virtual scene and obtain detection results; when the detection results indicate that the virtual resources include bone resources for creating the skeleton of the virtual object, the skeleton of the virtual object in the virtual scene is created based on the bone resources, and the bone motion data of the skeleton is initialized to obtain the bone motion data corresponding to the skeleton of the virtual object.

[0172] In some embodiments, the acquisition module is further configured to perform human skeleton key point detection on the virtual object in the virtual scene when the detection result indicates that there is no skeleton resource in the virtual resource for creating the skeleton of the virtual object, obtain skeleton key points carrying skeleton information, and determine the skeleton motion data of the corresponding skeleton based on the skeleton key points.

[0173] In some embodiments, the control of the virtual camera further includes a creation module. Before controlling the virtual camera to capture images of the virtual object, the creation module is used to obtain the initial shot size of the virtual camera and create a first animation sequence of the virtual camera that is adapted to the initial shot size. The first animation sequence is used to characterize the movement process of the virtual camera in the virtual scene during the process of the virtual camera capturing images of the virtual object based on the initial shot size.

[0174] Accordingly, in some embodiments, the creation module is further configured to determine the time point from which the virtual camera is controlled to capture images of the virtual object based on the target camera position, and from the time point, create a second animation sequence that is adapted to the target scene of the virtual camera; merge the first animation sequence and the second animation sequence to obtain the target animation sequence of the virtual camera, and the target animation sequence is used to render the virtual scene.

[0175] In some embodiments, the adjustment module is further configured to receive a scene switching instruction, adjust the target camera position to obtain an adjusted camera position, and based on the adjusted camera position, re-control the virtual camera to capture images of the virtual object to obtain a new target image including the virtual object.

[0176] This application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. The processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the virtual camera control method described above in this application.

[0177] This application provides a computer-readable storage medium storing executable instructions. When these executable instructions are executed by a processor, they cause the processor to execute the virtual camera control method provided in this application. For example, ... Figure 3 The method for controlling the virtual camera is shown.

[0178] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0179] In some embodiments, executable instructions may take the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0180] As an example, executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple collaborating files (e.g., a file that stores one or more modules, subroutines, or code sections).

[0181] As an example, executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.

[0182] In summary, the embodiments of this application offer the following advantages: by combining the virtual camera's focal length, target size, and target shot size, the target camera position of the virtual camera at the target shot size is automatically calculated, and a corresponding animation sequence for the virtual camera is generated. Thus, when the virtual camera's shot size changes, there is no need to manually adjust the virtual camera's position and parameters based on the virtual camera's preview view. The target camera position and corresponding animation sequence are automatically generated based on the target size setting, focal length setting, and shot size selection, effectively reducing the operational complexity of the virtual camera and simplifying its control when the shot size changes. Furthermore, considering various forms of skeletal resources while automatically setting the shot size, it exhibits high automation and strong applicability.

[0183] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A method for controlling a virtual camera, characterized in that, The method includes: Obtain the skeletal motion data corresponding to the skeleton of the virtual object in the virtual scene, and obtain the target scene corresponding to the virtual camera; The target scene size is used to indicate the proportion of different parts of the virtual object in the captured image of the virtual camera; The position information of the virtual object in the virtual scene is obtained, and the target height and focal length of the virtual camera are obtained. Select a target skeleton from the skeleton of the virtual object that is compatible with the target scene size; The target skeleton is the skeleton of the lowest part of the virtual object captured when the virtual camera captures the image of the virtual object based on the target scene. Based on the location information, the skeletal motion data of the target skeleton, the target surface height, and the focal length, the target camera position corresponding to the target scene of the virtual camera is determined; Based on the target camera position, the virtual camera is controlled to capture images of the virtual object, thereby obtaining a target image including the virtual object.

2. The method as described in claim 1, characterized in that, The step of determining the target camera position of the virtual camera corresponding to the target scene based on the location information, the skeletal motion data of the target skeleton, the target surface height, and the focal length includes: Based on the skeletal motion data of the target skeleton, the height between the target skeleton and the bones of the virtual object's head is determined; Based on the height, the target surface height, and the focal length, the distance from the virtual camera to the virtual object is determined; Based on the location information and the distance, the target camera position corresponding to the target scene is determined.

3. The method as described in claim 2, characterized in that, Determining the distance from the virtual camera to the virtual object based on the height, the target surface height, and the focal length includes: Obtain the ratio between the target surface height and the focal length; Determine the product between half of the height and the ratio, and use the product as the distance from the virtual camera to the virtual object.

4. The method as described in claim 2, characterized in that, Determining the target camera position of the virtual camera corresponding to the target scene based on the location information and the distance includes: Obtain the first component of the position information relative to the world coordinate system in a first direction, the second component of the position information relative to the world coordinate system in a second direction, and the third component of the position information relative to the world coordinate system in a third direction. Among the first direction, the second direction and the third direction, any two directions are perpendicular to each other; The second component and the distance are weighted and summed to obtain the target component of the virtual camera in the second direction relative to the world coordinate system; Based on the first component, the target component, and the third component, the target camera position corresponding to the target scene of the virtual camera is determined.

5. The method as described in claim 1, characterized in that, Selecting a target skeleton from the skeleton of the virtual object that is compatible with the target scene includes: When the target shot is a close-up, select a target bone from at least one bone in the shoulder of the virtual object that is compatible with the close-up shot; When the target scene is a close-up, select a target bone from at least one bone in the chest of the virtual object that is compatible with the close-up scene; When the target shot is a medium shot, select a target bone that is compatible with the medium shot from at least one bone in the hip of the virtual object; The step of controlling the virtual camera to capture images of the virtual object based on the target camera position, thereby obtaining a target image including the virtual object, includes: Based on the target camera position, the virtual camera is controlled to capture images of the area between the target skeleton and the top of the head of the virtual object, thereby obtaining a target image including the virtual object.

6. The method as described in claim 1, characterized in that, When the target scene is a panoramic view, the method further includes: Get the height of the virtual object; Based on the height, a reference height adapted to the panorama is determined; Obtain the ratio between the target surface height and the focal length; Based on the reference height and the ratio, the target camera position corresponding to the target scene is determined.

7. The method as described in claim 1, characterized in that, The step of obtaining the target scene corresponding to the virtual camera includes: Display at least one shot type to choose from, said shot type including at least one of the following: close-up, medium shot, long shot, and long shot; In response to the shot selection operation for the at least one shot type, the selected shot type is determined as the target shot type.

8. The method as described in claim 1, characterized in that, The process of acquiring the skeletal motion data corresponding to the skeleton of a virtual object in a virtual scene includes: The virtual resources in the virtual scene are detected, and the detection results are obtained; When the detection result indicates that the virtual resource includes a skeleton resource for creating the skeleton of the virtual object, the skeleton of the virtual object in the virtual scene is created based on the skeleton resource, and the skeleton motion data of the skeleton is initialized to obtain the skeleton motion data corresponding to the skeleton of the virtual object.

9. The method as described in claim 8, characterized in that, The method further includes: When the detection result indicates that there is no skeletal resource in the virtual resource for creating the skeleton of the virtual object, human skeleton key point detection is performed on the virtual object in the virtual scene to obtain skeletal key points carrying skeletal information, and based on the skeletal key points, the skeletal motion data of the corresponding skeleton is determined.

10. The method as described in claim 1, characterized in that, Before controlling the virtual camera to capture images of the virtual object, the method further includes: Obtain the initial shot size of the virtual camera and create a first animation sequence of the virtual camera that is adapted to the initial shot size; The first animation sequence is used to characterize the movement of the virtual camera in the virtual scene during the process of the virtual camera capturing images of the virtual object based on the initial scene size. The method further includes: Determine the time point from which the virtual camera captures images of the virtual object based on the target camera position, and create a second animation sequence from the time point that adapts the virtual camera to the target scene size; The first animation sequence and the second animation sequence are merged to obtain the target animation sequence of the virtual camera, which is used to render the virtual scene.

11. The method as described in claim 1, characterized in that, The method further includes: Upon receiving a shot switching command, the target camera position is adjusted to obtain the adjusted camera position; Based on the adjusted camera position, the virtual camera is controlled again to capture images of the virtual object, resulting in a new target image including the virtual object.

12. A control device for a virtual camera, characterized in that, The device includes: The acquisition module is used to acquire skeletal motion data corresponding to the skeleton of a virtual object in a virtual scene, and to acquire the target scene corresponding to the virtual camera; wherein, the target scene is used to indicate the proportion occupied by different parts of the virtual object in the captured image of the virtual camera; A determining module is used to acquire the position information of the virtual object in the virtual scene, and to acquire the target surface height and focal length of the virtual camera; to select a target skeleton that matches the target scene from the skeleton of the virtual object; wherein, the target skeleton is the skeleton of the lowest part of the virtual object captured when the virtual camera captures the image of the virtual object based on the target scene; and to determine the target camera position of the virtual camera corresponding to the target scene based on the position information, the skeletal motion data of the target skeleton, the target surface height, and the focal length. The control module is used to control the virtual camera to capture images of the virtual object based on the target camera position, so as to obtain a target image including the virtual object.

13. The apparatus according to claim 12, characterized in that, The determining module is further configured to determine the height between the target skeleton and the head skeleton of the virtual object based on the skeletal motion data of the target skeleton; determine the distance from the virtual camera to the virtual object based on the height, the target surface height, and the focal length; and determine the target camera position of the virtual camera corresponding to the target scene based on the position information and the distance.

14. The apparatus according to claim 13, characterized in that, The determining module is further configured to obtain the ratio between the target surface height and the focal length; determine the product between half of the height and the ratio, and use the product as the distance from the virtual camera to the virtual object.

15. The apparatus according to claim 13, characterized in that, The determining module is further configured to acquire a first component of the position information relative to a first direction of the world coordinate system, a second component of the position information relative to a second direction of the world coordinate system, and a third component of the position information relative to a third direction of the world coordinate system; wherein, any two of the first direction, the second direction, and the third direction are perpendicular to each other; the second component is weighted and summed with the distance to obtain the target component of the virtual camera relative to the world coordinate system in the second direction; and the target camera position of the virtual camera corresponding to the target scene is determined based on the first component, the target component, and the third component.

16. The apparatus according to claim 12, characterized in that, The determining module is further configured to: when the target shot is a close-up, select a target bone from at least one bone of the shoulder of the virtual object that is compatible with the close-up; when the target shot is a medium shot, select a target bone from at least one bone of the chest of the virtual object that is compatible with the medium shot; and when the target shot is a medium shot, select a target bone from at least one bone of the hip of the virtual object that is compatible with the medium shot. The control module is also used to control the virtual camera to capture images of the portion of the virtual object from the target skeleton to the top of the head, based on the target camera position, so as to obtain a target image including the virtual object.

17. The apparatus according to claim 12, characterized in that, When the target shot is a panoramic view. The determining module is further configured to obtain the height of the virtual object; determine a reference height adapted to the panorama based on the height; and obtain the ratio between the target surface height and the focal length. Based on the reference height and the ratio, the target camera position corresponding to the target scene is determined.

18. The apparatus according to claim 12, characterized in that, The acquisition module is further configured to display at least one selectable shot type, the shot type including at least one of the following: close-up, medium shot, long shot, wide shot, and long shot; and in response to a shot type selection operation for the at least one shot type, to determine the selected shot type as the target shot type.

19. The apparatus according to claim 12, characterized in that, The acquisition module is further configured to detect virtual resources in the virtual scene and obtain detection results; when the detection results indicate that the virtual resources include bone resources for creating the skeleton of the virtual object, the skeleton of the virtual object in the virtual scene is created based on the bone resources, and the bone motion data of the skeleton is initialized to obtain the bone motion data corresponding to the skeleton of the virtual object.

20. The apparatus according to claim 19, characterized in that, The acquisition module is further configured to perform human skeleton key point detection on the virtual object in the virtual scene when the detection result indicates that there is no skeleton resource in the virtual resource for creating the virtual object, obtain the skeleton key points carrying the skeleton information, and determine the skeleton motion data of the corresponding skeleton based on the skeleton key points.

21. The apparatus according to claim 12, characterized in that, The device further includes: A creation module is used to obtain the initial view of the virtual camera and create a first animation sequence of the virtual camera that is adapted to the initial view; wherein, the first animation sequence is used to characterize the movement process of the virtual camera in the virtual scene during the process of the virtual camera capturing images of the virtual object based on the initial view; The creation module is further configured to determine the time point at which the virtual camera captures images of the virtual object based on the target camera position, and to create a second animation sequence that matches the virtual camera with the target scene from the time point; and to merge the first animation sequence and the second animation sequence to obtain the target animation sequence of the virtual camera, which is used to render the virtual scene.

22. The apparatus according to claim 12, characterized in that, The device further includes: The adjustment module is used to receive a scene switching command, adjust the target camera position to obtain the adjusted camera position, and based on the adjusted camera position, re-control the virtual camera to capture images of the virtual object to obtain a new target image including the virtual object.

23. An electronic device, characterized in that, include: Memory, used to store executable instructions; A processor, when executing executable instructions stored in the memory, implements the virtual camera control method according to any one of claims 1 to 11.

24. A computer-readable storage medium storing computer-executable instructions, characterized in that, When the computer-executable instructions are executed by the processor, they implement the virtual camera control method according to any one of claims 1 to 11.

25. A computer program product comprising a computer program or computer-executable instructions, characterized in that, When the computer program or computer-executable instructions are executed by the processor, the virtual camera control method according to any one of claims 1 to 11 is implemented.

Citation Information

Patent Citations

  • Video-based attitude data capture method and system

    CN109145788A

  • Photographing method and device, computer equipment and storage medium

    CN109788191A