Video processing method, device, electronic device, storage medium and program product
By integrating the video data and virtual elements of recommended information in the virtual scene, the problem of separation of recommended information and video data in the virtual scene is solved, and the user perception efficiency and recommendation efficiency are improved.
Patent Information
- Application Number
- CN202211006770.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-22
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2042-08-22
AI Technical Summary
In the prior art, there is content separation between the recommendation information in the virtual scene and the video data, making it difficult to effectively improve the user's perception of the recommended information.
By seamlessly fusing the video data of the recommended information with the virtual elements in the virtual scene, the second video data is generated, and after the virtual scene is loaded, the virtual scene is switched to the third video data to display the running screen of the virtual scene, and rendering and displaying is used with graphics processing hardware.
It improves users' interest when playing videos, improves the recommendation efficiency of recommendation information, and avoids the sense of content splitting.
Smart Images

Figure CN115393554B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to video processing technology, and in particular to a method, device, electronic device, computer-readable storage medium, and computer program product for processing a video of a virtual scene. Background Art
[0002] Display technology based on graphics processing hardware has expanded the channels for perceiving the environment and obtaining information, especially multimedia technology for virtual scenes. With the help of human-computer interaction engine technology, it can realize diversified interactions between virtual objects controlled by users or artificial intelligence according to actual application needs. It has various typical application scenarios. For example, in virtual scenes such as games, it can simulate the real battle process between virtual objects.
[0003] Related technologies expose recommended information through video data in virtual scenes, hoping to improve users' perception efficiency of recommended information with the help of virtual scenes and video data. However, the video data including recommended information in related technologies is often separated from the virtual scenes in terms of content, making it difficult to truly associate the recommended information with the virtual scenes. Therefore, it is difficult for users to perceive the recommended information in the same way as they perceive the virtual scenes, and the perception efficiency of the recommended information cannot be improved. Summary of the Invention
[0004] The embodiments of the present application provide a video processing method, device, electronic device, computer-readable storage medium and computer program product for a virtual scene, which can seamlessly integrate the video data of recommended information with the virtual scene, thereby increasing the user's interest in playing the video, thereby improving the recommendation efficiency for the recommended information.
[0005] The technical solution of the embodiment of the present application is implemented as follows:
[0006] The present invention provides a method for processing a video of a virtual scene, including:
[0007] Acquiring first video data, wherein the first video data includes recommendation information;
[0008] Acquire virtual elements in the virtual scene;
[0009] fusing the first video data with the virtual element to obtain second video data;
[0010] Acquiring third video data of the virtual scene, wherein the third video data includes a running screen of the virtual scene;
[0011] The second video data and the third video data are sent to a terminal so that the terminal plays them.
[0012] The embodiment of the present application provides a video processing device for a virtual scene, comprising:
[0013] A first acquisition module is configured to acquire first video data, wherein the first video data includes recommendation information;
[0014] A second acquisition module, configured to acquire virtual elements in the virtual scene;
[0015] a fusion module, configured to fuse the first video data with the virtual element to obtain second video data;
[0016] a third acquisition module, configured to acquire third video data of the virtual scene, wherein the third video data includes a running screen of the virtual scene;
[0017] The sending module is used to send the second video data and the third video data to the terminal so that the terminal can play them.
[0018] In the above scheme, before obtaining the first video data, the first acquisition module is also used to obtain the account characteristics of the target account and the information characteristics of multiple candidate information, and determine the first feature distance between the information characteristics of each candidate information and the account characteristics, wherein the target account is the account that controls the virtual object in the virtual scene; based on the first feature distance, the multiple candidate information are sorted in ascending order, and the candidate information ranked first is used as the recommended information.
[0019] In the above scheme, before obtaining the first video data, the first acquisition module is also used to obtain the real-time scene features of the virtual scene and the information features of multiple candidate information, and determine the second feature distance between the information features of each candidate information and the real-time scene features; based on the second feature distance, the multiple candidate information are sorted in ascending order, and the candidate information ranked first is used as the recommended information.
[0020] In the above scheme, the fusion module is also used to: obtain the virtual object of the virtual scene and the background image of the virtual scene as the virtual elements; blur the background image of the virtual scene to obtain the background image to be synthesized; synthesize the background image to be synthesized and the first video data to obtain background synthesized video data; add the virtual object to the background synthesized video data to obtain the second video data.
[0021] In the above scheme, the fusion module is also used to: obtain the background pixels to be synthesized at each background position in the background image to be synthesized; perform any one of the following processing on each first video frame in the first video data: obtain multiple background pixels of the background picture of the first video frame, keep the multiple foreground pixels of the foreground picture of the first video frame unchanged, and use the background pixels to be synthesized to replace the background pixels at the same background position in the first video frame to obtain a background synthesized video frame; obtain multiple background pixels of the background picture of the first video frame, keep the multiple foreground pixels of the foreground picture of the first video frame unchanged, and synthesize the background pixels to be synthesized with the background pixels at the same background position in the first video frame to obtain a background synthesized video frame; encode multiple background synthesized video frames to obtain the background synthesized video data.
[0022] In the above scheme, the fusion module is also used to: obtain the first audio data corresponding to the first video data, and decode the first audio data to obtain each audio frame in the first audio data; obtain the facial features and action features of the virtual object, wherein the facial features and the action features match each of the audio frames in the first audio data; perform the following processing on each of the audio frames: perform special effects rendering processing using the facial features and action features matching the audio frame to obtain an object video frame including the virtual object; extract the background synthetic video frame corresponding to the audio frame from the background synthetic video data, and fuse the object video frame with the background synthetic video frame to obtain an object synthetic video frame corresponding to the audio frame; encode the object synthetic video frame corresponding to multiple audio frames one by one to obtain the second video data.
[0023] In the above scheme, the second acquisition module is also used to: use a virtual object that meets at least one of the following conditions as a virtual element of the virtual scene: the virtual object is a virtual object that the target account has controlled in the virtual scene; the virtual object is the virtual object that has been controlled the most times in the virtual scene; the virtual object is the virtual object that has been controlled the most times in the virtual scene.
[0024] In the above scheme, the sending module is also used to send the second video data to the terminal when the virtual scene is in the loading stage; when the virtual scene is loaded, stop sending the second video data to the terminal and send the third video data to the terminal; during the operation of the virtual scene, when a pause instruction for the third video data sent by the terminal is received, stop sending the third video data to the terminal and continue sending the second video data to the terminal.
[0025] The present invention provides a method for processing a video of a virtual scene, including:
[0026] In response to a start-up operation on the virtual scene, displaying a loading interface of the virtual scene;
[0027] Playing second video data in the loading interface, wherein the second video data includes recommendation information and at least one virtual element in the virtual scene;
[0028] In response to completion of loading of the virtual scene, the playing of the second video data is canceled, and the third video data is played, wherein the third video data includes a running screen of the virtual scene.
[0029] The embodiment of the present application provides a video processing device for a virtual scene, comprising:
[0030] A display module, configured to display a loading interface of the virtual scene in response to a start-up operation on the virtual scene;
[0031] a playing module, configured to play second video data in the loading interface, wherein the second video data includes recommendation information and at least one virtual element in the virtual scene;
[0032] The playback module is further configured to cancel playback of the second video data and play the third video data in response to completion of loading of the virtual scene, wherein the third video data includes a running screen of the virtual scene.
[0033] In the above scheme, the virtual elements include the background of the virtual scene and the virtual objects of the virtual scene, the second video data includes the virtual objects in the background, and the playback module is also used to: display the virtual objects performing object actions and broadcast first audio data in the loading interface; wherein, the first audio data is audio data used to introduce the recommended information, and the object action is the action of the virtual object in the virtual scene.
[0034] In the above solution, after playing the third video data, the playing module is further used to: in response to satisfying a condition for canceling the playing of the third video data, cancel the playing of the third video data and continue playing the second video data.
[0035] In the above scheme, the playback module is also used to: when the second video data played most recently from the current moment has unplayed subsequent content, play the second video data corresponding to the subsequent content, wherein the subsequent content includes the recommendation information, and the current moment is the moment that satisfies the playback cancellation condition; when the second video data played most recently from the current moment does not have unplayed subsequent content, play the second video data including another recommendation information.
[0036] In the above scheme, the playback module is also used to: perform any one of the following processes: playing the second video data in the human-computer interaction interface; displaying the second video data playback entrance, and playing the second video data in the human-computer interaction interface in response to a trigger operation on the second video data playback entrance; playing the second video data in the human-computer interaction interface in response to meeting the automatic playback condition of the second video data.
[0037] In the above scheme, in response to meeting the automatic playback condition of the second video data, before playing the second video data in the human-computer interaction interface, the playback module is also used to: obtain decision reference data, wherein the decision reference data includes at least one of the following: historical operation data of the target account, account attribute data of the target account, and scene data of the virtual scene, and the target account is the account that controls the virtual object in the virtual scene; call the first neural network model to perform the following processing: extract decision reference features from the decision reference data, and predict the target account's preference level for the second video data based on the decision reference features; when the preference level is less than the preference level threshold, determine that the automatic playback condition of the second video data is met.
[0038] In the above scheme, in response to meeting the automatic playback condition of the second video data, before playing the second video data in the human-computer interaction interface, the playback module is also used to: obtain the historical moment of each playback of the second video data from the historical playback record, and extract the historical scene features of the historical moment; fuse the historical scene features of multiple historical moments to obtain a fused scene feature; obtain the current scene data of the virtual scene, and extract the current scene feature of the current scene data; when the similarity between the current scene feature and the fused scene feature is greater than the similarity threshold, it is determined that the automatic playback condition of the second video data is met.
[0039] In the above scheme, the playback module is also used to: perform any one of the following processes: when the third video data is in the buffering stage, determine that the playback cancellation condition is met; in response to the target account's pause operation on the third video data, determine that the playback cancellation condition is met; in response to the target account's setting operation on the virtual scene, determine that the playback cancellation condition is met; wherein, the target account is the account that controls the virtual object in the virtual scene.
[0040] In the above scheme, the prominence of the display style of the second video data is negatively correlated with the first characteristic parameter; wherein, the first characteristic parameter includes at least one of the following: the activity level of the target account, the level of the target account, the historical performance of the target account in the virtual scene, and the number of times the second video data is played in the virtual scene; the target account is the account that controls the virtual object in the virtual scene.
[0041] An embodiment of the present application provides an electronic device, including:
[0042] a memory for storing computer-executable instructions;
[0043] The processor is used to implement the video processing method of the virtual scene provided in the embodiment of the present application when executing the computer executable instructions stored in the memory.
[0044] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions for implementing a video processing method for a virtual scene provided in an embodiment of the present application when executed by a processor.
[0045] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the video processing method of the virtual scene provided in the embodiment of the present application is implemented.
[0046] The embodiments of the present application have the following beneficial effects:
[0047] Through the embodiment of the present application, the virtual elements in the virtual scene are integrated into the first video data including the recommendation information, and the obtained second video data includes elements related to the virtual scene, so that the recommendation information is naturally integrated with the third video data through the virtual scene elements. Therefore, no matter whether the second video data or the third video data is played, there will be no sense of separation at the content level, which can enhance the user's interest in playing the video, thereby improving the recommendation efficiency for the recommendation information. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1A Schematic diagram of an application mode of the video processing method for a virtual scene provided in an embodiment of the present application;
[0049] Figure 1B Schematic diagram of an application mode of the video processing method for a virtual scene provided in an embodiment of the present application;
[0050] Figure 2 is a structural diagram of an electronic device provided in an embodiment of the present application;
[0051] Figure 3 is a structural diagram of an electronic device provided in an embodiment of the present application;
[0052] Figures 4A-4D 1 is a flow chart of a method for processing a video of a virtual scene provided in an embodiment of the present application;
[0053] Figure 5 Schematic diagram of the interface of the video processing method of the virtual scene provided by the embodiment of the present application;
[0054] Figure 6 Schematic diagram of the interface of the video processing method of the virtual scene provided by the embodiment of the present application;
[0055] Figure 7 Schematic diagram of the interface of the video processing method of the virtual scene provided by the embodiment of the present application;
[0056] Figure 8 is a flowchart of a video processing method for a virtual scene provided by an embodiment of the present application;
[0057] Figure 9 is a flowchart of a video processing method for a virtual scene provided by an embodiment of the present application;
[0058] Figure 10 is a flowchart of a video processing method for a virtual scene provided by an embodiment of the present application;
[0059] Figure 11 This is a flowchart of a video processing method for a virtual scene provided in an embodiment of the present application. DETAILED DESCRIPTION
[0060] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0061] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0062] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0064] Before further explaining the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0065] 1) Cloud gaming: The game runs on the server, which transmits the compressed and encoded game screen to the client, which then displays it.
[0066] 2) Gaussian blur: A technique that reduces image clarity, reduces image layers and details, and blurs the image.
[0067] 3) Virtual scenes: Utilizing the output of devices that are different from the real world, visual perception of virtual scenes can be formed with the naked eye or with the assistance of devices, such as two-dimensional images output by display screens, and three-dimensional images output by stereoscopic display technologies such as stereo projection, virtual reality, and augmented reality. In addition, various possible hardware can also be used to form various perceptions that simulate the real world, such as auditory perception, tactile perception, olfactory perception, and motion perception.
[0068] 4) Virtual objects: Objects that interact in virtual scenes, are controlled by users or robot programs (e.g., artificial intelligence-based robot programs), and can remain stationary, move, and perform various behaviors in the virtual scene, such as various characters in games.
[0069] Related technologies expose recommended information through video data in virtual scenes, hoping to improve users' perception efficiency of recommended information with the help of virtual scenes and video data. However, the video data of recommended information in related technologies often have content separation from the virtual scenes, making it difficult to truly associate the recommended information with the virtual scenes. There is a strong sense of separation between the video data of the recommended information and the content of the virtual scenes. Therefore, it is difficult for users to perceive the recommended information as they perceive the virtual scenes, and it is impossible to improve the perception efficiency of the recommended information, and thus it is impossible to improve the recommendation efficiency of the recommended information.
[0070] In the related art, the client configures a display position at a specific location. After the client is started, recommended information is exposed at the display position, or the user is required to manually slide to a specified page before the video data of the recommended information can be played on the specified page. Usually, users are not very willing to browse or click, resulting in low exposure of recommended information.
[0071] In response to the above technical problems, the embodiments of the present application provide a video processing method, device, electronic device, computer-readable storage medium and computer program product for a virtual scene, which can reduce the sense of disconnection between the video data of the recommended information and the virtual scene, and can increase the user's interest in playing the video, thereby improving the recommendation efficiency of the recommended information.
[0072] The following describes exemplary applications of the electronic device provided in the embodiments of the present application. The electronic device provided in the embodiments of the present application can be implemented as a terminal or as a server.
[0073] To facilitate understanding of the video processing method for a virtual scene provided in an embodiment of the present application, an exemplary implementation scenario of the video processing method for a virtual scene provided in an embodiment of the present application is first described. A terminal runs a client, and during the operation of the client, a virtual scene including a role-playing game is output. The virtual scene is an environment for game characters to interact, such as a combat game environment on a city street. The virtual scene includes virtual objects, which can be game characters controlled by a user (or player). That is, the virtual objects are controlled by a real player and will move in the virtual scene in response to the real player's operations on a controller (including a touch screen, voice-activated switch, keyboard, mouse, joystick, etc.). For example, the virtual object is controlled by the user to move in the virtual scene and perform a series of operations such as attacking and defending.
[0074] In one implementation scenario, see Figure 1A , Figure 1A This is a schematic diagram of an application mode of the video processing method for a virtual scene provided in an embodiment of the present application, which is applicable to cloud gaming scenarios where some games are completely run on server 200. Generally, it is applicable to an application mode that relies on the computing power of server 200 to complete virtual scene calculations and output the virtual scene on terminal 400.
[0075] In some embodiments, in response to the terminal 400 receiving the user's operation to run the client, the terminal 400 sends the client startup data to the server 200. The server 200 obtains the first video data including the recommendation information from the advertising provider 500 (for example, the native advertising video pulled from the advertising provider). The server 200 obtains the virtual elements in the virtual scene and fuses the first video data with the virtual elements to obtain the second video data. The server 200 also obtains the third video data including the running screen of the virtual scene through the rendering calculation of the game engine. Specifically, the server 200 calculates the data required for display through the graphics computing hardware, and completes the loading, parsing and rendering of the display data. When the graphics output hardware outputs the video frame that can form a visual perception of the virtual scene, during the process of the server rendering and calculating the third video data, the server 200 sends the second video data fused with the virtual elements to the terminal 400 through the network 300, and the terminal 400 plays the second video data. When the server calculates the third video data, the server 200 stops sending the second video data to the terminal 400 and sends the third video data to the terminal 400, and the terminal 400 plays the third video data.
[0076] In some embodiments, the terminal 400 can implement the video processing method of the virtual scene provided in the embodiments of the present application by running a computer program. For example, the computer program can be a native program or software module in the operating system; it can be a native application (APP, Application), that is, a program that needs to be installed in the operating system to run, such as a game APP (i.e., the above-mentioned client); it can also be a mini-program, that is, a program that can be run by simply downloading it into a browser environment; it can also be a game mini-program that can be embedded in any APP. In short, the above-mentioned computer program can be an application, module, or plug-in in any form.
[0077] The embodiments of the present application can be implemented with the help of cloud technology. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and network within a wide area network or a local area network to realize data computing, storage, processing, and sharing.
[0078] Cloud technology is a general term for network, information, integration, management platform, and application technologies used in the cloud computing business model. It can form a resource pool that can be used flexibly and conveniently on demand. Cloud computing technology will become a key support. The backend services of technical network systems require a large amount of computing and storage resources.
[0079] As an example, the server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal 400 can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, smart voice interaction device, smart home appliance, vehicle terminal, aircraft, etc., but is not limited to this. The terminal 400 and the server 200 can be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present application.
[0080] In some embodiments, see Figure 1B , Figure 1B This is a structural diagram of the image processing system of the virtual scene provided by the embodiment of the present application. The following describes an exemplary application of the embodiment of the present application based on the blockchain network. Figure 1B , including a blockchain network 600 (exemplarily showing nodes 610 - 1 and 610 - 2 included in the blockchain network 600 ), a server 200 , and a terminal 400 , which are described below respectively.
[0081] The server 200 (mapped as node 610-2) and the terminal 400 (mapped as node 610-1) can both join the blockchain network 600 and become nodes therein. Figure 1B The figure exemplarily shows that the terminal 400 is mapped as a node 610-1 of the blockchain network 600, and each node (such as node 610-1, node 610-2) has a consensus function and a bookkeeping function (i.e., maintaining a state database library, such as a key-value database).
[0082] The state database of each node (eg, node 610 - 1 ) records the second video data (advertisement video integrated with virtual elements) pulled by the server 400 , so that the terminal 400 can query the second video data recorded in the state database.
[0083] In some embodiments, in response to receiving a request to load a virtual scene, multiple servers 200 (each mapped as a node in the blockchain network) obtain second video data (an advertising video incorporating virtual elements). When the number of nodes that have reached consensus on the second video data exceeds a threshold, consensus is determined to have been achieved. Server 200 (mapped as node 610-2) then transmits the second video data to terminal 400 (mapped as node 610-1), presenting it on the human-computer interaction interface of terminal 400 and storing the second video data on the blockchain. Because the second video data is obtained through consensus across multiple servers, the legitimacy and reliability of the native advertisement can be ensured. Furthermore, due to the tamper-resistant nature of the blockchain network, the stored second video data cannot be maliciously tampered with.
[0084] See also Figure 2 , Figure 2 1 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application, and is described by taking a terminal as an example. Figure 2 The terminal 400 shown includes: at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. The various components in the terminal 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the bus system 440 is not shown in FIG. Figure 2 Various buses are labeled as bus system 440 .
[0085] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0086] The user interface 430 includes one or more output devices 431 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0087] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.
[0088] Memory 450 includes volatile memory or nonvolatile memory, or may include both volatile and nonvolatile memory. Nonvolatile memory may be read-only memory (ROM), and volatile memory may be random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0089] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.
[0090] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;
[0091] A network communication module 452 for reaching other computing devices via one or more (wired or wireless) network interfaces 420 , exemplary network interfaces 420 including Bluetooth, WiFi, and USB;
[0092] a presentation module 453 for enabling presentation of information via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with the user interface 430 (e.g., a user interface for operating peripheral devices and displaying content and information);
[0093] The input processing module 454 is configured to detect one or more user inputs or interactions from one of the one or more input devices 432 and to translate the detected inputs or interactions.
[0094] In some embodiments, the video processing device for the virtual scene provided in the embodiments of the present application can be implemented in software. Figure 2 A video processing device 455 of a virtual scene stored in a memory 450 is shown, which can be software in the form of programs and plug-ins, including the following software modules: a display module 4551 and a playback module 4552. These modules are logical, and therefore can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be explained below.
[0095] See also Figure 3 , Figure 3 2 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application, taking the electronic device as a server 200 as an example, Figure 3The server 200 shown includes: at least one processor 210, a memory 250, and at least one network interface 220. The various components in the server 200 are coupled together via a bus system 240. It is understood that the bus system 240 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 240 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the bus system 240 is not described in detail. Figure 3 Various buses are labeled as bus system 240 .
[0096] The processor 210 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0097] The memory 250 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 250 may optionally include one or more storage devices that are physically remote from the processor 210.
[0098] The memory 250 includes volatile memory or nonvolatile memory, or may include both volatile and nonvolatile memory. The nonvolatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 250 described in the embodiments of the present application is intended to include any suitable type of memory.
[0099] In some embodiments, the memory 250 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.
[0100] Operating system 251, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;
[0101] The network communication module 252 is used to reach other computing devices via one or more (wired or wireless) network interfaces 220 . Exemplary network interfaces 220 include Bluetooth, WiFi, and Universal Serial Bus (USB).
[0102] In some embodiments, the video processing device for the virtual scene provided in the embodiments of the present application can be implemented in software. Figure 3 An artificial intelligence-based data processing device 255 stored in the memory 250 is shown, which can be software in the form of programs and plug-ins, including the following software modules: a first acquisition module 2551, a second acquisition module 2552, a fusion module 2553, a third acquisition module 2554 and a sending module 2555. These modules are logical, and therefore can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be explained below.
[0103] The video processing method of the virtual scene provided in the embodiment of the present application can be executed by the server 200. Figure 4A , Figure 4A This is a flow chart of the video processing method of the virtual scene provided by the embodiment of the present application, which will be combined with Figure 4A It should be noted that, Figure 4A The illustrated method may be executed by various forms of computer programs executed by the server 200 .
[0104] In step 101, the server obtains first video data.
[0105] As an example, the video data involved in the embodiments of the present application includes audio data by default, and the first video data includes recommendation information. The recommendation information can be advertising information or other information that you want to expose to users. Before obtaining the first video data, you can perform the following processing to obtain the recommendation information from the candidate information, so that you can obtain the complete advertising video data including the recommendation information. For example, the complete advertising video data is a 15-second video, and the first video data can be part of the video data or all of the video data in the complete advertising video data. The account characteristics of the target account and the information characteristics of multiple candidate information are obtained. The target account is the account that controls the virtual object in the virtual scene, that is, the target account is the account used by the user who currently uses the terminal to experience the virtual scene. The account characteristics include but are not limited to at least the following: One of: features that characterize the browsing history of an account, features that characterize the attribute data of an account, the information features of the candidate information include but are not limited to at least one of the following: features that characterize the type of candidate information, features that characterize the price of candidate information, features that characterize the promotion population of candidate information, and determine the first feature distance between the information features of each candidate information and the account features. Since the features can be represented in the form of vectors, the first feature distance can be represented by the distance between two vectors. Based on the first feature distance, multiple candidate information are sorted in ascending order, and the candidate information ranked first is used as recommended information. Each time the first video data is obtained, at least one advertising video of a recommended information can be obtained, and multiple advertising videos of multiple recommended information ranked first can also be obtained. Through the embodiment of the present application, first video data that meets the user's interests can be obtained, thereby improving the recommendation efficiency for recommended information.
[0106] As an example, the first video data includes recommendation information, which may be advertising information or other information that is desired to be exposed to the user. Before obtaining the first video data, the following processing may be performed to obtain recommendation information from the candidate information, thereby obtaining complete advertising video data including the recommendation information. For example, the complete advertising video data is a 15-second video, and the first video data may be partial video data or all video data in the complete advertising video data. The real-time scene features of the virtual scene and information features of multiple candidate information are obtained. The scene features include but are not limited to at least one of the following: features characterizing the environmental data of the virtual scene, features characterizing the crowd-oriented data of the virtual scene, and the information features of the candidate information include but are not limited to at least one of the following: features characterizing the type of the candidate information, features characterizing the price of the candidate information, and features characterizing the promotion crowd of the candidate information. The second feature distance between the information feature of each candidate information and the real-time scene feature is determined. Since the feature can be represented in the form of a vector, the second feature distance can be represented by the distance between two vectors. The multiple candidate information is sorted in ascending order based on the second feature distance, and the candidate information ranked first is used as the recommendation information. Each time the first video data is obtained, at least one advertising video of one recommended information can be obtained, and multiple advertising videos of multiple recommended information ranked first can also be obtained. Through the embodiments of the present application, first video data matching the virtual scene can be obtained, thereby obtaining first video data that meets the user's interests, thereby improving the recommendation efficiency for recommendation information.
[0107] In step 102 , virtual elements in a virtual scene are obtained.
[0108] As an example, virtual elements in a virtual scene are identified to obtain multiple candidate virtual elements, and the candidate virtual elements are screened to obtain virtual elements used in subsequent steps. The virtual elements include at least one of the following: a background image of the virtual scene, and a virtual object in the virtual scene.
[0109] In some embodiments, obtaining the virtual elements in the virtual scene in step 102 can be achieved through the following technical solution: a virtual object that meets at least one of the following conditions is used as a virtual element of the virtual scene: the virtual object is a virtual object that the target account has controlled in the virtual scene, for example, the virtual object that the target account most recently controlled in the virtual scene; the virtual object is the virtual object that has been controlled the most times in the virtual scene, for example, there are 10 virtual objects in the virtual scene, among which virtual object A is selected by users for interaction in the virtual scene 80 times in 100 virtual scene operation records, and is the virtual object that has been controlled the most times in total. As long as virtual object A is selected by any user, the control count record increases by 1; the virtual object is the virtual object that has been controlled the most times in the virtual scene, for example, there are 10 virtual objects in the virtual scene, among which virtual object A is selected by the target account for interaction in the virtual scene 40 times in 100 virtual scene operation records, and is the virtual object that has been controlled the most times. As long as virtual object A is selected by the target account, the control count record increases by 1.
[0110] As an example, in addition to obtaining virtual objects in the above-mentioned manner, virtual objects can also be screened according to their frequency of appearance in the second video data. For example, if virtual object A is integrated into the second video data 10 times, virtual objects with a frequency exceeding the frequency threshold are used as virtual elements of the virtual scene. Screening can also be performed based on the account characteristics of the target account. The account characteristics are used to characterize the target account's preference for virtual objects. Virtual objects with a feature distance between the object characteristics of the virtual object and the account characteristics that is less than the feature distance threshold are used as virtual elements.
[0111] Through the embodiments of the present application, virtual objects as virtual elements can be filtered out in a diversified manner, so that the advertising video can contain more content that meets the user's interests while improving the correlation between the advertising video and the virtual scene, thereby improving the recommendation efficiency of the recommended information.
[0112] In step 103, the first video data is fused with the virtual element to obtain second video data.
[0113] As an example, the original virtual element is directly fused with the first video data, or the original virtual element is changed based on a set dimension to obtain a change result, and the change result is fused with the first video data, the set dimension includes at least one of the following: the size of the virtual element, the color of the virtual element, the clothing of the virtual element, the action of the virtual element, the expression of the virtual element, or the original virtual element is combined to obtain a combination result, and the combination result is fused with the first video data; or, after the first video data is fused with the virtual element, other decorative elements are added to the fusion result to obtain the second video data; or, after the first video data is fused with the virtual element, the fusion result is subjected to style migration processing to the recommended information to obtain the second video data.
[0114] In some embodiments, see Figure 4B , Figure 4B The flowchart of the video processing method of the virtual scene provided by the embodiment of the present application is as follows: the virtual elements include the virtual objects of the virtual scene and the background image of the virtual scene. In step 103, the first video data is fused with the virtual elements to obtain the second video data. Figure 4B The steps shown are implemented.
[0115] In step 1031, the background image of the virtual scene is blurred to obtain a background image to be synthesized.
[0116] As an example, Gaussian blur processing is performed on the background image of the virtual scene to obtain a background image to be synthesized, and the background image to be synthesized is an image to be fused into the first video data.
[0117] In step 1032, the background image to be synthesized and the first video data are synthesized to obtain background synthesized video data.
[0118] In some embodiments, step 1032 synthesizes the background image to be synthesized and the first video data to obtain background synthesized video data, which can be achieved by the following technical solutions: obtaining the background pixels to be synthesized at each background position in the background image to be synthesized, the size of the background image to be synthesized is the same as the size of the first video frame of the first video data, the first video frame is obtained by decoding the first video data, and any one of the following processes is performed on each first video frame in the first video data: obtaining multiple background pixels of the background picture of the first video frame, which is equivalent to obtaining the background pixels corresponding to each background position of the first video frame, the background position is the position in the first video frame that belongs to the background part, the background pixel is the pixel at the background position, and the foreground of the first video frame is kept The multiple foreground pixels of the picture remain unchanged, and the background pixels at the same background position in the first video frame are replaced with the background pixels to be synthesized to obtain a background synthesized video frame. That is, the background pixels at the background position in the background image to be synthesized are replaced with the background pixels at the background position in the first video frame to obtain multiple background pixels of the background picture of the first video frame. The multiple foreground pixels of the foreground picture of the first video frame remain unchanged, and the background pixels to be synthesized are synthesized with the background pixels at the same background position in the first video frame to obtain a background synthesized video frame. The synthesis process can be to average the color values of the two background pixels, and replace the background pixels at the background position in the first video frame with the pixels of the average processing result; synthesize and encode the multiple background synthesized video frames to obtain background synthesized video data. Through the embodiment of the present application, the background image to be synthesized can be integrated into the first video data, so that the background of the second video data is associated with the virtual scene, thereby increasing the user's interest in the second video data, and thus improving the recommendation efficiency of the recommended information.
[0119] In step 1033 , the virtual object is added to the background composite video data to obtain second video data.
[0120] In some embodiments, in step 1033, the virtual object is added to the background synthetic video data to obtain the second video data, which can be achieved by the following technical solutions: obtaining the first audio data corresponding to the first video data, and decoding the first audio data to obtain each audio frame in the first audio data; obtaining the facial features and action features of the virtual object, wherein the facial features and action features match each audio frame in the first audio data; performing the following processing on each audio frame: performing special effects rendering processing using the facial features and action features matching the audio frame to obtain an object video frame including the virtual object; extracting the background synthetic video frame corresponding to the audio frame from the background synthetic video data, and fusing the object video frame with the background synthetic video frame to obtain the object synthetic video frame corresponding to the audio frame; encoding the object synthetic video frame corresponding to multiple audio frames one by one to obtain the second video data.
[0121] In step 104 , third video data of the virtual scene is acquired.
[0122] As an example, the third video data includes the running screen of the virtual scene, which is rendered by the game engine. The user controls the virtual object through the client to interact with other virtual objects in the virtual scene. The real-time action data generated during the interaction is sent by the client to the server. The server will render the real-time action data to the environmental screen of the virtual scene. Since the environmental screen of the virtual scene does not change as frequently as the interactive action, the rendering of the running screen usually only needs to render the real-time action of the virtual object, so that the running screen of the virtual scene can be obtained in real time, and the real-time obtained running screen is encoded into the third video data.
[0123] In step 105, the second video data and the third video data are sent to the terminal so that the terminal plays them.
[0124] In some embodiments, see Figure 4C , Figure 4C This is a flow chart of the video processing method for a virtual scene provided by an embodiment of the present application. In step 105, the second video data and the third video data are sent to the terminal. Figure 4C The steps shown are implemented.
[0125] In step 1051, when the virtual scene is in the loading stage, second video data is sent to the terminal.
[0126] In step 1052, when the virtual scene is loaded, the second video data is stopped from being sent to the terminal, and the third video data is sent to the terminal.
[0127] In step 1053, during the operation of the virtual scene, when a pause instruction for the third video data is received from the terminal, the third video data is stopped from being sent to the terminal, and the second video data is continued to be sent to the terminal.
[0128] For example, see Figure 9 , Figure 9This is a flowchart of a video processing method for a virtual scene provided by an embodiment of the present application. In step 901, during the game loading phase, the server sends an advertising video clip to the client. In step 902, the server receives a game start command from the client. In step 903, the server sends the cloud game video (third video data) to the client. Specifically, after the user's game video finishes loading, the cloud game server pauses sending the advertising video (second video data) and sends the normal game video (third video data). In step 904, the server receives a pause game command from the client. In step 905, the server sends the advertising video (second video data) to the client. When the user manually pauses the game or performs actions such as menu bar operations, character changes, or skin changes, the client sends a pause command to the cloud game server, which then continues to send the advertising video to the client. When the user resumes the game, the server pauses sending the advertising video and begins sending the game video. While the user is playing the game, the advertising video continues to be sent to the client during the period when the user manually pauses the game, allowing the user to be exposed to the advertising video while making game-related settings.
[0129] Through the embodiment of the present application, the virtual elements in the virtual scene are integrated into the first video data including the recommendation information, and the obtained second video data includes elements related to the virtual scene. Subsequently, the second video data and the third video data of the virtual scene are sent to the terminal for playback, thereby reducing the sense of separation between the video data of the recommendation information and the virtual scene, and can increase the user's interest in playing the video, thereby improving the recommendation efficiency for the recommendation information.
[0130] The video processing method of the virtual scene provided in the embodiment of the present application can be executed by the terminal 400. Figure 4D , Figure 4D This is a flow chart of the video processing method of the virtual scene provided by the embodiment of the present application, which will be combined with Figure 4D The steps shown are explained.
[0131] It should be noted that Figure 4D The method shown can be executed by various forms of computer programs running on the terminal 400 and is not limited to the above-mentioned client, such as the operating system 451, software modules and scripts mentioned above. Therefore, the client should not be regarded as a limitation on the embodiments of the present application.
[0132] In step 201 , in response to a start-up operation on a virtual scene, a loading interface of the virtual scene is displayed.
[0133] In step 202 , second video data is played in the loading interface, wherein the second video data includes recommendation information and at least one virtual element in a virtual scene.
[0134] In some embodiments, the virtual elements include a background of a virtual scene and virtual objects in the virtual scene, and the second video data includes the virtual objects in the background. Playing the second video data in the loading interface in step 202 can be achieved by the following technical solution: displaying the virtual object performing an object action and broadcasting first audio data in the loading interface; wherein the first audio data is audio data used to introduce the recommended information, and the object action is the action of the virtual object in the virtual scene. Through embodiments of the present application, the degree of association between the second video data and the virtual scene can be increased, thereby effectively increasing the user's interest in the second video data and improving the perception efficiency of the recommended information.
[0135] As an example, the virtual element includes at least one of the background of the virtual scene and the virtual object of the virtual scene. When the virtual element includes the background, the second video data is an advertising video integrated into the background of the virtual scene. When the virtual element includes the virtual object, the second video data is an advertising video with the virtual object added. The virtual object will perform object actions, and the virtual object's lip shape is adapted to the first audio data, which is equivalent to forming a visual effect of introducing and recommending information about the virtual object. The virtual object can be the virtual object that the user who initiated the startup operation last controlled in the virtual scene, and so on.
[0136] In step 203 , in response to completion of loading of the virtual scene, the playing of the second video data is canceled, and the third video data is played, wherein the third video data includes a running screen of the virtual scene.
[0137] In some embodiments, after the third video data is played in step 203, in response to a condition for canceling the playback of the third video data being met, the playback of the third video data is canceled, and the playback of the second video data continues.
[0138] In some embodiments, any one of the following processes is performed: when the third video data is in the buffering stage, determining that the condition for canceling the playback is met; in response to the target account's pause operation on the third video data, determining that the condition for canceling the playback is met; in response to the target account's setting operation on the virtual scene, determining that the condition for canceling the playback is met; wherein, the target account is the account that controls the virtual object in the virtual scene.
[0139] As an example, when the game client is in the initialization phase, the third video data is buffering. At this time, the cancel play condition is met and the third video data is in an unplayed state. In order to effectively utilize the user's waiting time, the second video data can be played during this time period. The game client is initialized and the virtual scene starts running. At this time, the cancel play condition is met and the third video data is played in the human-computer interaction interface to ensure the user's normal gaming experience. In response to the user's pause operation on the third video data, such as the user actively pausing the game, the cancel play condition is met at this time and the third video data will not be played (in an unplayed state), so the second video data will be played. Alternatively, in response to the user's setting operation on the virtual scene, such as setting the control layout or setting the game resolution, since there is no game interaction during the operation setting process, the cancel play condition is met and the third video data is in an unplayed state. Through the embodiments of the present application, various time gaps that do not affect the player's normal gaming experience can be utilized to improve the display resource utilization and the exposure rate of recommended information.
[0140] In some embodiments, the above-mentioned continued playback of the second video data can be achieved through the following technical solution: when the second video data played most recently before the current moment has unplayed subsequent content, the second video data corresponding to the subsequent content is played, wherein the subsequent content includes recommendation information, and the current moment satisfies the playback cancellation condition; when the second video data played most recently before the current moment does not have unplayed subsequent content, the second video data including another recommendation information is played. Through the embodiments of the present application, the intermittently played second video data can be correlated, thereby improving the perception efficiency of the recommendation information.
[0141] As an example, advertising video A is an advertising video (second video data) that incorporates virtual elements for recommended information A, advertising video B is an advertising video (second video data) that incorporates virtual elements for recommended information B, and advertising video C is another advertising video (second video data) that incorporates virtual elements for recommended information A. For example, the second video data that was played most recently is the first 5 seconds of the complete advertising video A, then the last 10 seconds of advertising video A is the continuation content, and the second video data that continues to be played is the continuation content. When the second video data that was played most recently does not have continuation content, for example, the second video data that was played most recently is the last 10 seconds of advertising video A, the second video data that continues to be played is an advertising video for another recommended information, for example, advertising video B, or the second video data that continues to be played can also be another advertising video for the recommended information, for example, advertising video C.
[0142] In some embodiments, the above-mentioned continued playback of the second video data can be achieved through the following technical solutions: performing any one of the following processes: playing the second video data in the human-computer interaction interface; displaying a second video data playback entry, and playing the second video data in the human-computer interaction interface in response to a triggering operation on the second video data playback entry; playing the second video data in the human-computer interaction interface in response to meeting the automatic playback conditions of the second video data. The manual method provides space for user selection, while the automatic method can improve the efficiency of human-computer interaction.
[0143] In some embodiments, in response to satisfying the automatic playback condition of the second video data, before playing the second video data in the human-computer interaction interface, decision reference data is obtained, wherein the decision reference data includes at least one of the following: historical operation data of the target account, the operation data can represent the user's interest preference), account attribute data of the target account, the attribute data can also represent the user's interest preference, scene data of the virtual scene, the scene data can represent the progress of the game, such as whether it is in a certain summary stage, etc. The target account is the account that controls the virtual object in the virtual scene; calling the first neural network model to perform the following processing: extracting decision reference features from the decision reference data, and predicting the target account based on the decision reference features. The target account's preference for the second video data; when the preference is less than the preference threshold, it is determined that the automatic playback condition of the second video data is met. The first neural network model is obtained through training, and the training sample is the historical playback record of the second video data. When the historical second video data is closed during automatic playback, the data is marked as 0. When the historical second video data is not closed during automatic playback, the data is marked as 1. Combined with the decision parameter data sample obtained from the historical playback record, the predicted preference for the historical second video data A is predicted to be 0.8. Since the historical second video data A is marked as 1, there is an error of 0.2. The error is converged to the minimum value by updating the model parameters. The target account's preference for the second video data is determined by artificial intelligence, so that intelligent automatic playback can be achieved, and the played second video data meets the user's preferences, which can improve the recommendation efficiency of recommended information.
[0144] In some embodiments, in response to the automatic playback condition of the second video data being met, before playing the second video data in the human-computer interaction interface, the historical moment at each playback of the second video data is obtained from the historical playback record, and the historical scene features of the historical moment are extracted; the historical scene features of multiple historical moments are fused to obtain a fused scene feature; the current scene data of the virtual scene is obtained, and the current scene features of the current scene data are extracted; when the similarity between the current scene feature and the fused scene feature is greater than a similarity threshold, it is determined that the automatic playback condition of the second video data is met. By using artificial intelligence to determine whether the current virtual scene is suitable for playing the second video data, intelligent automatic playback can be achieved, effectively improving the degree of intelligence and the efficiency of human-computer interaction.
[0145] In some embodiments, the prominence of the display style of the second video data is negatively correlated with the first characteristic parameter; wherein the first characteristic parameter includes at least one of the following: the activity level of the target account, the level of the target account, the historical performance of the target account in the virtual scene, and the number of times the second video data is played in the virtual scene; the target account is the account that controls the virtual object in the virtual scene.
[0146] As an example, the prominence includes at least one of the following: clarity (the higher the clarity, the higher the prominence), color saturation (the higher the color saturation, the higher the prominence), brightness (the higher the brightness, the higher the prominence), etc. Taking the level of the target account as an example, the higher the level of the target account, the higher the prominence of the second video data, thereby encouraging users to have higher participation and better participation results in the virtual scene.
[0147] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.
[0148] In some embodiments, in response to the terminal receiving the user's operation of running the game client, the terminal sends the game client startup data to the server, and the server obtains first video data including recommendation information (for example, a native advertising video pulled from the advertising provider), and the server obtains the virtual elements in the virtual scene, and fuses the first video data with the virtual elements to obtain second video data. The server also obtains third video data including the running screen of the virtual scene through the rendering calculation of the game engine. Specifically, the server calculates the data required for display through graphics computing hardware, and completes the loading, parsing and rendering of the display data. The graphics output hardware outputs a video frame that can form a visual perception of the virtual scene. The server sends the second video data and the third video data to the terminal, and the terminal decodes and plays the second video data and the third video data. It can be played simultaneously and in parallel or serially. The decoding and playback process can be to present a two-dimensional video frame on the display screen of a smartphone, or to project a video frame on the lenses of augmented reality / virtual reality glasses to achieve a three-dimensional display effect.
[0149] The embodiment of the present application divides and reorganizes the video and audio streams of the original advertising video and the cloud game video, integrates the game background, game characters, etc. into the advertising video, and sends it to the client through compression encoding on the server side. The client decodes and displays it, thereby reducing the time taken by the client to pull the advertising video from the third-party advertising platform, inserting the advertising video into the video stream of the cloud game, and the server flexibly controls the style of the advertising video, splicing the advertising video and the virtual character form of the cloud game, thereby increasing the interesting exposure rate of the advertisement.
[0150] In some embodiments, the server pulls specified advertisements (for example, advertisements that meet the user's interests) from a third-party advertising platform based on the current user's characteristics, segments the advertisement video, and merges the segmented advertisement video data with the virtual elements of the current game to obtain a new advertisement video that incorporates virtual elements. The new advertisement video and the cloud game video are sent to the client in the form of a video stream, and the sending time period of the new advertisement video is embedded in the pause or non-game phase of the cloud game video. The received video data is decoded and displayed on the client. In response to the user clicking to enter the cloud game page, the video data of the new advertisement video is received during the buffering phase of the game. When the game is loaded, the cloud game video sent by the server is received. In response to the user setting operation or archiving operation, the video data of the new advertisement video continues to be received, thereby achieving the continued playback of the new advertisement video. The embodiments of the present application can improve the exposure rate of the advertisement video and the user experience.
[0151] In some embodiments, see Figure 5 , Figure 5This is an interface diagram of the video processing method for a virtual scene provided in an embodiment of the present application. The cloud game video data during the game is displayed in the human-computer interaction interface 501. In response to the user's setting operation or archiving operation, the cloud game server will push a new advertising video 502 including virtual elements, so that the user can operate the game menu (or change the skin) and watch the new advertising video at the same time. In response to the user's operation of continuing the game, the cloud game server suspends the video data push of the new advertising video and instead pushes the cloud game video data.
[0152] In some embodiments, see Figure 6 , Figure 6 This is a schematic diagram of the interface of the video processing method for a virtual scene provided by an embodiment of the present application. A setting menu 602 is displayed in a human-computer interaction interface 601. In response to a user's setting operation on the setting menu 602, the client sends a pause instruction to the server. The server synthesizes a new advertising video (incorporating real-time virtual elements) in real time or sends a pre-synthesized new advertising video (incorporating virtual elements) to the client. When the user operates the setting menu, a new advertising video 603 incorporating virtual elements is played, such as a new advertising video with a virtual character playing a voice. Figure 7 , Figure 7 It is an interface diagram of the video processing method of the virtual scene provided in an embodiment of the present application. In response to the user ending the setting operation, for example, the human-computer interaction interface 701 displays the end setting control 702. The end setting operation is a trigger operation for the end setting control 702. The client sends a continue instruction to the server, and the server continues to push the cloud game video.
[0153] In some embodiments, after the user starts the cloud game, when the cloud game is initialized to render, the cloud game server prioritizes merging the virtual character selected by the user last time with the advertising video. Before or after the merger, the advertising video can be cut into small segments that take up the time required for cloud game rendering, so that the advertisement can be displayed during the initial rendering phase of the cloud game. After the user starts to enter the cloud game, the server pushes the cloud game stream and no longer exposes the advertising video. When the user needs to pause the cloud game or operate the menu bar, the client sends a pause command to the server, and the server continues to send the unfinished advertising video to the client for display. After the user completes the operation, the client sends a continue game command to the server, and the server pauses the advertising video push and continues the cloud game push. During this process, the user's operation is not affected in any way. Only when the user pauses the game, the cloud game video push is changed to the advertising video push. Since the background of the advertising video is obtained by Gaussian blurring the game video background by the server, adding the animation of the virtual character to the advertising video for merging can greatly eliminate the visual difference between the advertising video and the cloud game video, enhancing the fun of the advertising video and the user's willingness to watch.
[0154] In some embodiments, see Figure 8 , Figure 8 This is a flowchart of the video processing method for the virtual scene provided by the embodiment of the present application. It is explained by taking the real-time acquisition of the advertising video for secondary processing as an example. In step 801, the server pulls the targeted advertising video from the third-party advertising platform according to the user's login account. In step 802, the client will start the cloud game after logging in, and the server will perform cloud game video rendering processing. In step 803, after pulling the targeted advertising video, the server segments the advertising video. Specifically, first, the video clip is obtained according to the buffering time required by the user before entering the game, and then the manual tags such as the user manually pausing the game, the user operating the menu bar, switching characters, changing skins, etc. are used as cutting points. After receiving the user's cutting point tag, the timestamp corresponding to the cutting point tag is used as the starting point of the segment cutting and the timestamp of re-entering the game is used as the end point of the segment cutting, that is, the advertising video for secondary processing is obtained in real time, and the rendering and pushing stream starts from the advertising video frame at the starting point until the rendering and pushing stream of the advertising video frame at the end point. In step 804, the server captures the current game's general background, applies a Gaussian blur to it, and superimposes the resulting background with the background of the advertising video, thereby eliminating the visual difference between the advertising video and the cloud game video. In step 805, the server selects the user's corresponding virtual game character from the current game's character database and generates a game character image. In step 806, the game character image is configured with standard actions to generate an action video. Combining the action video with the advertising video creates the effect of the game character directing the playback of the advertising video. In step 807, the client displays the synthesized advertising video. Specifically, the server sends the synthesized advertising video to the client. For example, while the user is loading a game, the server sends the synthesized advertising video to the client to prioritize the playback of the advertising video clip. In step 808, in response to the client entering the game, the client performs the game display. Specifically, after the cloud game video rendering is completed, the client performs the game display. In step 809, in response to the user performing other non-gaming operations, the client sends a pause command to the server, transitioning to step 803, thereby also playing the advertising video clip during non-gaming phases. In step 810, in response to receiving the user's "continue game" action, the client displays the game. In addition to using the universal game background image for Gaussian blurring, different background images can be dynamically captured at different cloud game stages and embedded into the ad video. This ensures that even if the user pauses the cloud game at any time, the ad video and the current game background remain visually consistent when the client plays the ad video in the background.
[0155] In some embodiments, see Figure 9 , Figure 9This is a flowchart of the video processing method for a virtual scene provided by an embodiment of the present application. During the game loading phase, the server sends an advertising video clip to the client, receives a game start instruction sent by the client, and then sends a cloud game video to the client. Specifically, after the user's game video is loaded, the cloud game server pauses sending the advertising video and sends a normal game video. It receives a pause game instruction sent by the client, and then sends an advertising video clip to the client. When the user manually pauses the game, or performs actions such as menu bar operations, character changes, or skin changes, the client sends a pause instruction to the cloud game server, and the cloud game server continues to send the advertising video to the client. When the user continues the game, the server pauses sending the advertising video and starts sending the game video. While the user is playing the game, the advertising video continues to be sent to the client according to the above logic while the user manually pauses the game, so as to achieve the purpose of allowing the user to make game-related settings while being exposed to the advertising video.
[0156] In some embodiments, see Figure 11 , Figure 11 This is a flowchart of a video processing method for a virtual scene provided in an embodiment of the present application.
[0157] In step 1101, the client sends a game start request to the cloud server.
[0158] In step 1102 , the cloud server pulls first video data from an advertisement provider, where the first video data is advertisement video data.
[0159] In step 1103, the cloud server integrates virtual elements into the first video data to obtain second video data.
[0160] In step 1104, the cloud server performs rendering processing in the buffering stage to obtain third video data, which carries the game screen.
[0161] In step 1105 , when the rendering process has not yet ended, the cloud server sends the second video data to the client.
[0162] In step 1106 , the client plays the second video data.
[0163] In step 1107, when the rendering process of the buffering stage is completed, the cloud server sends the third video data to the client and stops sending the second video data.
[0164] In step 1108 , the client plays the third video data.
[0165] In step 1109, the client receives the user's game pause operation and sends the game pause instruction to the cloud server.
[0166] In step 1110 , the cloud server sends the second video data to the client and stops sending the third video data.
[0167] In step 1111 , the client plays the second video data.
[0168] In step 1112, the client receives the user's game continue operation and sends the game continue instruction to the cloud server.
[0169] In step 1113, the cloud server sends the third video data to the client and stops sending the second video data.
[0170] In some embodiments, see Figure 10 , Figure 10 This is a flowchart of the video processing method of the virtual scene provided by the embodiment of the present application. It is explained by taking the non-real-time acquisition of advertising videos for secondary processing as an example. In step 1001, the server pulls the advertising video. In step 1002, the server renders the cloud game video. In step 1003, the server performs Gaussian blur processing on the general background of the game. In step 1004, the server fills the general background of the game with the advertising video background, thereby eliminating the sense of separation between the advertising video and the content displayed on the client user interface. In step 1005, the server selects the virtual character of the current user, generates a three-dimensional animated image, and then generates an animated video of the three-dimensional animated image with general behavior. In step 1006, the animation video of the current virtual character and the advertising video are superimposed on the animation and video to obtain an advertising video integrated with virtual elements, and finally a video effect of playing advertisements using virtual characters can be obtained. This is equivalent to starting to process the advertising video in advance after pulling it. The advertising video in the game buffer period can be understood as being rendered in real time. During the subsequent game operation, the entire advertising video is processed, so real-time rendering is no longer required. The advertising video can be continued to be pushed or stopped directly according to the client's instructions.
[0171] Through the embodiments of the present application, the virtual characters of the game and the general background of the game are embedded in the advertising video, which reduces the sense of separation between the advertising video and the game video, and embeds the advertising video in the form of clips into the game video according to user instructions, which can increase the correlation between the advertising video and the game video and the exposure rate of the advertising video.
[0172] It is understandable that in the embodiments of the present application, when user information and other related data are involved, when the embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0173] The following continues to describe the exemplary structure of the virtual scene video processing device 255 provided in the embodiment of the present application as a software module. In some embodiments, such as Figure 3 As shown, the software modules in the video processing device 255 of the virtual scene stored in the memory 250 may include: a first acquisition module 2551, used to acquire first video data, wherein the first video data includes recommendation information; a second acquisition module 2552, used to acquire virtual elements in the virtual scene; a fusion module 2553, used to fuse the first video data with the virtual elements to obtain second video data; a third acquisition module 2554, used to acquire third video data of the virtual scene, wherein the third video data includes the running screen of the virtual scene; a sending module 2555, used to send the second video data and the third video data to the terminal so that the terminal can play them.
[0174] In some embodiments, before obtaining the first video data, the first acquisition module 2551 is also used to obtain the account characteristics of the target account and the information characteristics of multiple candidate information, and determine the first feature distance between the information characteristics and the account characteristics of each candidate information, wherein the target account is the account that controls the virtual object in the virtual scene; based on the first feature distance, the multiple candidate information are sorted in ascending order, and the candidate information ranked first is used as recommended information.
[0175] In some embodiments, before obtaining the first video data, the first acquisition module 2551 is also used to obtain the real-time scene features of the virtual scene and the information features of multiple candidate information, and determine the second feature distance between the information features of each candidate information and the real-time scene features; sort the multiple candidate information in ascending order based on the second feature distance, and use the candidate information ranked first as recommended information.
[0176] In some embodiments, the fusion module 2553 is also used to: obtain virtual objects of the virtual scene and the background image of the virtual scene as virtual elements; blur the background image of the virtual scene to obtain the background image to be synthesized; synthesize the background image to be synthesized and the first video data to obtain background synthesized video data; add the virtual object to the background synthesized video data to obtain second video data.
[0177] In some embodiments, the fusion module 2553 is also used to: obtain the background pixels to be synthesized at each background position in the background image to be synthesized; perform any one of the following processing on each first video frame in the first video data: obtain multiple background pixels of the background picture of the first video frame, keep the multiple foreground pixels of the foreground picture of the first video frame unchanged, and replace the background pixels at the same background position in the first video frame with the background pixels to be synthesized to obtain a background synthesized video frame; obtain multiple background pixels of the background picture of the first video frame, keep the multiple foreground pixels of the foreground picture of the first video frame unchanged, and synthesize the background pixels to be synthesized with the background pixels at the same background position in the first video frame to obtain a background synthesized video frame; encode the multiple background synthesized video frames to obtain background synthesized video data.
[0178] In some embodiments, the fusion module 2553 is also used to: obtain first audio data corresponding to the first video data, and decode the first audio data to obtain each audio frame in the first audio data; obtain facial features and motion features of the virtual object, wherein the facial features and motion features match each audio frame in the first audio data; perform the following processing on each audio frame: perform special effects rendering processing using the facial features and motion features matching the audio frame to obtain an object video frame including the virtual object; extract background synthetic video frames corresponding to the audio frames from the background synthetic video data, and fuse the object video frames with the background synthetic video frames to obtain object synthetic video frames corresponding to the audio frames; encode the object synthetic video frames corresponding to multiple audio frames one by one to obtain second video data.
[0179] In some embodiments, the second acquisition module 2552 is further used to: use a virtual object that meets at least one of the following conditions as a virtual element of the virtual scene: the virtual object is a virtual object that the target account has ever controlled in the virtual scene; the virtual object is the virtual object that has been controlled the most times in the virtual scene; the virtual object is the virtual object that has been controlled the most times in the virtual scene.
[0180] In some embodiments, the sending module 2555 is also used to send the second video data to the terminal when the virtual scene is in the loading stage; when the virtual scene is loaded, stop sending the second video data to the terminal and send the third video data to the terminal; during the operation of the virtual scene, when a pause instruction for the third video data sent by the terminal is received, stop sending the third video data to the terminal and continue sending the second video data to the terminal.
[0181] The following continues to describe the exemplary structure of the virtual scene video processing device 455 provided in the embodiment of the present application as a software module. In some embodiments, such as Figure 2As shown, the software modules in the video processing device 455 of the virtual scene stored in the memory 450 may include: a display module 4551, which is used to display the loading interface of the virtual scene in response to the startup operation for the virtual scene; a playback module 4552, which is used to play the second video data in the loading interface, wherein the second video data includes recommendation information and at least one virtual element in the virtual scene; the playback module 4552 is also used to cancel the playback of the second video data in response to the completion of the loading of the virtual scene, and play the third video data, wherein the third video data includes the running screen of the virtual scene.
[0182] In some embodiments, the virtual elements include the background of the virtual scene and the virtual objects of the virtual scene, the second video data includes the virtual objects in the background, and the playback module 4552 is also used to: display the virtual objects performing object actions and broadcast the first audio data in the loading interface; wherein the first audio data is audio data used to introduce the recommended information, and the object action is the action of the virtual object in the virtual scene.
[0183] In some embodiments, after playing the third video data, the playback module 4552 is further configured to: in response to a condition for canceling the playback of the third video data being met, cancel the playback of the third video data and continue to play the second video data.
[0184] In some embodiments, the playback module 4552 is also used to: when the second video data played most recently from the current moment has unplayed subsequent content, play the second video data corresponding to the subsequent content, wherein the subsequent content includes recommendation information, and the current moment is a moment that satisfies the playback cancellation condition; when the second video data played most recently from the current moment does not have unplayed subsequent content, play the second video data including another recommendation information.
[0185] In some embodiments, the playback module 4552 is also used to: perform any one of the following processes: playing the second video data in the human-computer interaction interface; displaying the second video data playback entrance, and playing the second video data in the human-computer interaction interface in response to a trigger operation on the second video data playback entrance; playing the second video data in the human-computer interaction interface in response to meeting the automatic playback condition of the second video data.
[0186] In some embodiments, in response to satisfying the automatic playback condition of the second video data, before playing the second video data in the human-computer interaction interface, the playback module 4552 is further used to: obtain decision reference data, wherein the decision reference data includes at least one of the following: historical operation data of the target account, account attribute data of the target account, and scene data of the virtual scene, and the target account is the account that controls the virtual object in the virtual scene; call the first neural network model to perform the following processing: extract decision reference features from the decision reference data, and predict the target account's preference level for the second video data based on the decision reference features; when the preference level is less than the preference level threshold, determine that the automatic playback condition of the second video data is satisfied.
[0187] In some embodiments, in response to the automatic playback condition of the second video data being met, before playing the second video data in the human-computer interaction interface, the playback module 4552 is also used to: obtain the historical moment of each playback of the second video data from the historical playback record, and extract the historical scene features of the historical moment; fuse the historical scene features of multiple historical moments to obtain a fused scene feature; obtain the current scene data of the virtual scene, and extract the current scene features of the current scene data; when the similarity between the current scene feature and the fused scene feature is greater than the similarity threshold, it is determined that the automatic playback condition of the second video data is met.
[0188] In some embodiments, the playback module 4552 is also used to: perform any one of the following processes: when the third video data is in the buffering stage, determining that the condition for canceling the playback is met; in response to the target account's pause operation on the third video data, determining that the condition for canceling the playback is met; in response to the target account's setting operation on the virtual scene, determining that the condition for canceling the playback is met; wherein, the target account is the account that controls the virtual object in the virtual scene.
[0189] In some embodiments, the prominence of the display style of the second video data is negatively correlated with the first characteristic parameter; wherein the first characteristic parameter includes at least one of the following: the activity level of the target account, the level of the target account, the historical performance of the target account in the virtual scene, and the number of times the second video data is played in the virtual scene; the target account is the account that controls the virtual object in the virtual scene.
[0190] The present invention provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the virtual scene video processing method described in the present invention.
[0191] The embodiment of the present application provides a computer-readable storage medium in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor will execute the video processing method of the virtual scene provided by the embodiment of the present application, for example, Figures 4A-4D The video processing method of the virtual scene is shown.
[0192] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or various devices including one or any combination of the above memories.
[0193] In some embodiments, executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0194] As an example, executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file storing other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).
[0195] By way of example, executable instructions may be deployed to be executed on one computing device, or on multiple computing devices at one site, or on multiple computing devices distributed across multiple sites and interconnected by a communication network.
[0196] To sum up, through the embodiment of the present application, the virtual elements in the virtual scene are integrated into the first video data including the recommended information, and the obtained second video data includes elements related to the virtual scene. Subsequently, the second video data and the third video data of the virtual scene are sent to the terminal for playback, thereby avoiding a strong sense of separation between the video data of the recommended information and the virtual scene, and improving the information perception efficiency when playing the video.
[0197] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.
Claims
1. A video processing method for a virtual scene, characterized in that: The method comprises: Acquire first video data, wherein the first video data includes recommendation information; Acquire virtual elements in the virtual scene, where the virtual elements include virtual objects in the virtual scene and a background image of the virtual scene; Performing blur processing on the background image to obtain a background image to be synthesized, and performing synthesis processing on the background image to be synthesized and the first video data to obtain background synthesized video data; For each audio frame in the first video data, extracting a background synthetic video frame corresponding to the audio frame from the background synthetic video data, and fusing the object video frame with the background synthetic video frame to obtain an object synthetic video frame corresponding to the audio frame, wherein the object video frame is obtained by performing special effects rendering processing using facial features and action features of the virtual object that matches the audio frame; encoding the object-synthesized video frames corresponding one-to-one to the plurality of audio frames to obtain second video data; Acquiring third video data of the virtual scene, wherein the third video data includes a running screen of the virtual scene; The second video data and the third video data are sent to a terminal so that the terminal plays them.
2. The method according to claim 1, characterized in that Before acquiring the first video data, the method further includes: Obtaining an account feature of a target account and information features of a plurality of candidate information, and determining a first feature distance between the information feature of each candidate information and the account feature, wherein the target account is an account that controls a virtual object in the virtual scene; The plurality of candidate information are sorted in ascending order based on the first feature distance, and the candidate information ranked first is used as the recommended information.
3. The method according to claim 1, characterized in that Before acquiring the first video data, the method further includes: Acquire a real-time scene feature of the virtual scene and information features of a plurality of candidate information, and determine a second feature distance between the information feature of each candidate information and the real-time scene feature; The plurality of candidate information are sorted in ascending order based on the second feature distance, and the candidate information ranked first is used as the recommended information.
4. The method according to claim 1, wherein The synthesizing process of the background image to be synthesized and the first video data to obtain background synthesized video data includes: Obtaining the background pixels to be synthesized at each background position in the background image to be synthesized; Perform any one of the following processing on each first video frame in the first video data: Acquire multiple background pixels of the background picture of the first video frame, keep multiple foreground pixels of the foreground picture of the first video frame unchanged, and replace background pixels at the same background position in the first video frame with the background pixels to be synthesized to obtain a background synthesized video frame; Acquire multiple background pixels of the background picture of the first video frame, keep multiple foreground pixels of the foreground picture of the first video frame unchanged, and synthesize the background pixels to be synthesized with background pixels at the same background position in the first video frame to obtain a background synthesized video frame; The plurality of background composite video frames are encoded to obtain the background composite video data.
5. The method according to claim 1, wherein The obtaining of the virtual elements in the virtual scene includes: A virtual object that meets at least one of the following conditions is used as a virtual element of the virtual scene: The virtual object is a virtual object that the target account has previously controlled in the virtual scene; The virtual object is the virtual object that has been controlled the most times in the virtual scene; The virtual object is the virtual object that is controlled the most times by the target account in the virtual scene.
6. The method according to claim 1, characterized in that The sending the second video data and the third video data to the terminal includes: When the virtual scene is in the loading stage, sending the second video data to the terminal; When the virtual scene is loaded, stop sending the second video data to the terminal and send the third video data to the terminal; During the operation of the virtual scene, when a pause instruction for the third video data sent by the terminal is received, the sending of the third video data to the terminal is stopped, and the sending of the second video data to the terminal is continued.
7. A video processing method for a virtual scene, characterized in that: The method comprises: In response to a start-up operation on the virtual scene, displaying a loading interface of the virtual scene; Playing second video data in the loading interface, wherein the second video data includes recommendation information and at least one virtual element in the virtual scene, the virtual element includes a virtual object of the virtual scene and a background image of the virtual scene, and the second video data includes a plurality of object-synthesized video frames, the object-synthesized video frames being synthesized by synthesizing object video frames and background-synthesized video frames, the object video frames being obtained by performing special effects rendering processing on the virtual objects that match the audio frames in the first video data, and the background-synthesized video frames being synthesized by synthesizing the blurred background image and the first video data; In response to completion of loading of the virtual scene, the playing of the second video data is canceled, and the third video data is played, wherein the third video data includes a running screen of the virtual scene.
8. The method according to claim 7, characterized in that The second video data includes a virtual object in the background, and playing the second video data in the loading interface includes: Displaying, in the loading interface, the virtual object performing an object action and broadcasting first audio data; The first audio data is audio data for introducing the recommendation information, and the object action is an action of the virtual object in the virtual scene.
9. The method according to claim 7, characterized in that After playing the third video data, the method further includes: In response to the condition for canceling the playback of the third video data being met, the playback of the third video data is canceled, and the playback of the second video data continues.
10. The method according to claim 9, characterized in that The continuing to play the second video data includes: When the second video data played most recently from the current moment has unplayed subsequent content, playing the second video data corresponding to the subsequent content, wherein the subsequent content includes the recommendation information, and the current moment is a moment that satisfies the playback cancellation condition; When the second video data played most recently from the current moment has no unplayed subsequent content, the second video data including another piece of recommendation information is played.
11. The method according to claim 9, characterized in that The continuing to play the second video data includes: Perform any of the following: Playing the second video data in the human-computer interaction interface; displaying a second video data playback entrance, and playing the second video data in the human-computer interaction interface in response to a triggering operation on the second video data playback entrance; In response to satisfying an automatic play condition of the second video data, the second video data is played in the human-computer interaction interface.
12. The method according to claim 11, characterized in that In response to the automatic playing condition of the second video data being met, before playing the second video data in the human-computer interaction interface, the method further includes: Acquiring decision reference data, wherein the decision reference data includes at least one of the following: historical operation data of a target account, account attribute data of the target account, and scene data of the virtual scene, wherein the target account is an account that controls a virtual object in the virtual scene; Calling the first neural network model to perform the following processing: extracting decision reference features from the decision reference data, and predicting the target account's preference for the second video data based on the decision reference features; When the preference level is less than a preference level threshold, it is determined that an automatic playback condition for the second video data is met.
13. The method according to claim 11, characterized in that In response to the automatic playing condition of the second video data being met, before playing the second video data in the human-computer interaction interface, the method further includes: Acquire the historical moment of each playback of the second video data from the historical playback record, and extract the historical scene features of the historical moment; fusing the historical scene features of the plurality of historical moments to obtain a fused scene feature; Acquiring current scene data of the virtual scene, and extracting current scene features of the current scene data; When the similarity between the current scene feature and the fused scene feature is greater than a similarity threshold, it is determined that the automatic playback condition of the second video data is met.
14. The method according to claim 9, characterized in that The method further comprises: Perform any of the following processing: When the third video data is in a buffering stage, determining that the playback cancellation condition is met; In response to a pause operation of the target account on the third video data, determining that the playback cancellation condition is met; In response to a setting operation of the target account for the virtual scene, determining that the playback cancellation condition is satisfied; The target account is an account that controls a virtual object in the virtual scene.
15. The method according to claim 7, characterized in that The prominence of the display style of the second video data is negatively correlated with the first characteristic parameter; Among them, the first characteristic parameter includes at least one of the following: the activity level of the target account, the level of the target account, the historical performance of the target account in the virtual scene, and the number of times the second video data is played in the virtual scene; the target account is the account that controls the virtual object in the virtual scene.
16. A video processing device for a virtual scene, characterized in that: The device comprises: A first acquisition module is configured to acquire first video data, wherein the first video data includes recommendation information; A second acquisition module is configured to acquire virtual elements in the virtual scene, where the virtual elements include virtual objects in the virtual scene and a background image of the virtual scene; a fusion module configured to blur the background image to obtain a background image to be synthesized, synthesize the background image to be synthesized and the first video data to obtain background synthesized video data; extract, for each audio frame in the first video data, a background synthesized video frame corresponding to the audio frame from the background synthesized video data, and fuse the object video frame with the background synthesized video frame to obtain an object synthesized video frame corresponding to the audio frame, wherein the object video frame is obtained by performing special effects rendering processing using facial features and animation features of the virtual object that matches the audio frame; and encode the object synthesized video frames corresponding one-to-one to the multiple audio frames to obtain second video data; a third acquisition module, configured to acquire third video data of the virtual scene, wherein the third video data includes a running screen of the virtual scene; The sending module is used to send the second video data and the third video data to the terminal so that the terminal can play them.
17. The device according to claim 16, characterized in that The first acquisition module is also used to obtain the account characteristics of the target account and the information characteristics of multiple candidate information, and determine the first feature distance between the information characteristics of each candidate information and the account characteristics, wherein the target account is the account that controls the virtual object in the virtual scene; based on the first feature distance, the multiple candidate information are sorted in ascending order, and the candidate information ranked first is used as the recommended information.
18. The device according to claim 16, characterized in that The first acquisition module is further configured to acquire the real-time scene feature of the virtual scene and information features of a plurality of candidate information, and determine a second feature distance between the information feature of each candidate information and the real-time scene feature; The plurality of candidate information are sorted in ascending order based on the second feature distance, and the candidate information ranked first is used as the recommended information.
19. A video processing device for a virtual scene, characterized in that: The device comprises: A display module, configured to display a loading interface of the virtual scene in response to a start-up operation on the virtual scene; a playback module, configured to play second video data in the loading interface, wherein the second video data includes recommendation information and at least one virtual element in the virtual scene, the virtual element includes a virtual object of the virtual scene and a background image of the virtual scene, and the second video data includes a plurality of object-synthesized video frames, the object-synthesized video frames are synthesized by synthesizing object video frames and background-synthesized video frames, the object video frames are obtained by performing special effects rendering processing on the virtual objects that match the audio frames in the first video data, and the background-synthesized video frames are synthesized by synthesizing the blurred background image and the first video data; The playback module is further configured to cancel playback of the second video data and play third video data in response to completion of loading of the virtual scene, wherein the third video data includes a running screen of the virtual scene.
20. An electronic device, characterized in that: The electronic device comprises: a memory for storing executable instructions; The processor is configured to implement the video processing method of the virtual scene according to any one of claims 1 to 6 or 7 to 15 when executing the computer executable instructions stored in the memory.
21. A computer-readable storage medium storing computer-executable instructions, characterized in that: When the computer executable instructions are executed by a processor, the video processing method of the virtual scene according to any one of claims 1 to 6 or 7 to 15 is implemented.
22. A computer program product comprising a computer program or computer executable instructions, characterized in that When the computer program or computer executable instructions are executed by a processor, the video processing method of the virtual scene according to any one of claims 1 to 6 or 7 to 15 is implemented.
Citation Information
Patent Citations
Game data processing method and apparatus, and computer and readable storage medium
WO2022022281A1
KR20220021076A