Game image reconstruction network training and reconstruction method, device and storage medium
By training the game image reconstruction network, using the generative adversarial network to reconstruct the real-world image data into game-style image data, and adding stroke information to the reconstructed image data, it solves the problem that it is difficult to efficiently convert the video data style in the prior art, and realizes efficient and low-threshold game-style video data production.
Patent Information
- Application Number
- CN202210937965.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-05
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-08-05
AI Technical Summary
The prior art is difficult to efficiently convert the style of video data into game style, and it is difficult to achieve the real game style using multiple filters superimposed, which increases the threshold and time-consuming of producing video data.
By training the game image reconstruction network, the real-world image data is reconstructed into game-style image data using the generative adversarial network, and stroke information is added to the reconstructed image data to enhance the game style.
It realizes efficient conversion of video data style into game style, lowers the threshold and time-consuming to produce video data, and improves the efficiency of making game style video data.
Smart Images

Figure CN115829828B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a training and reconstruction method, device and storage medium for a game image reconstruction network. Background Art
[0002] In scenarios such as short videos and advertisements, users will produce various types of video data. After recording the original video data, the video data is usually post-processed to improve the quality of the video data.
[0003] Due to certain business needs, part of the post-processing is to convert the style of the video data to the style of a certain game. The commonly used post-processing is to add filters to the video data to convert the entire video data to other styles, such as retro, film, sunset, etc.
[0004] However, filters usually adjust the color values of pixels and add other decorative elements, resulting in a relatively simple effect. It is difficult to achieve the style of the game by using multiple filters. If the video data is designed according to the style of the game, this will greatly increase the threshold for producing video data, resulting in a much longer time for producing video data and low efficiency in producing video data. Summary of the invention
[0005] The present invention provides a method, device and storage medium for training and reconstructing a game image reconstruction network to solve the problem of how to efficiently realize the style of the game on the screen.
[0006] According to one aspect of the present invention, a method for training a game image reconstruction network is provided, comprising:
[0007] Collect and record first sample image data of the real world;
[0008] Identify sample games;
[0009] Collecting sample video data of the sample game recorded when the user controls the sample game;
[0010] extracting second sample image data from the sample video data;
[0011] The generative adversarial network is trained into a game image reconstruction network using the first sample image data as a source of content and the second sample image data as a source of game style.
[0012] According to another aspect of the present invention, there is provided an image reconstruction method, comprising:
[0013] Loading the game image reconstruction network trained according to the method of any embodiment of the present invention;
[0014] Acquire original image data to be reconstructed;
[0015] Inputting the original image data into the game image reconstruction network to reconstruct it into candidate image data containing a game style;
[0016] Add stroke information to the candidate image data to obtain target image data.
[0017] According to another aspect of the present invention, there is provided a video reconstruction method, comprising:
[0018] Loading the game image reconstruction network trained according to the method of any embodiment of the present invention;
[0019] The acquired content is original video data introducing a target game, wherein the original video data contains multiple frames of original image data;
[0020] Inputting the original image data into the game image reconstruction network to reconstruct it into candidate image data containing a game style;
[0021] Adding stroke information to the candidate image data to obtain target image data;
[0022] The target image data is used to replace the original image data in the original video data to obtain the target video data.
[0023] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0024] at least one processor; and
[0025] a memory communicatively connected to the at least one processor; wherein,
[0026] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the training method of the game image reconstruction network or the image reconstruction method or the video reconstruction method described in any embodiment of the present invention.
[0027] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, wherein the computer program is used to enable a processor to implement the training method of a game image reconstruction network, the image reconstruction method, or the video reconstruction method described in any embodiment of the present invention when executed.
[0028] In this embodiment, first sample image data of the real world is collected and recorded; a sample game is determined; sample video data of the sample game recorded when the user controls the sample game is collected; second sample image data is extracted from the sample video data; and the generative adversarial network is trained as a game image reconstruction network with the first sample image data as the source of the content and the second sample image data as the source of the game style. The sample video data recorded when the user controls the sample game is used to collect the screen of the sample game, thereby improving the efficiency of making samples, and the generative adversarial network is trained in a non-paired data manner, so that the game image reconstruction network can reconstruct the image data into the game style. Reconstructing the game style belongs to post-processing, which can maintain the threshold for making video data and the time consumption for making video data, thereby greatly improving the efficiency of making video data with a game style.
[0029] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0031] Figure 1 is a flowchart of a method for training a game image reconstruction network according to Embodiment 1 of the present invention;
[0032] Figure 2 is a flow chart of an image reconstruction method provided according to Embodiment 2 of the present invention;
[0033] FIG. 3A to FIG. 3E This is an example diagram of image reconstruction provided according to the first embodiment of the present invention;
[0034] Figure 4 is a flowchart of a video reconstruction method provided according to Embodiment 3 of the present invention;
[0035] Figure 5 is a structural schematic diagram of a training device for a game image reconstruction network provided according to a fourth embodiment of the present invention;
[0036] Figure 6 is a structural schematic diagram of an image reconstruction device provided according to Embodiment 5 of the present invention;
[0037] Figure 7is a structural schematic diagram of a video reconstruction device provided according to Embodiment 6 of the present invention;
[0038] Figure 8 It is a schematic diagram of the structure of an electronic device provided by Embodiment 7 of the present invention. DETAILED DESCRIPTION
[0039] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0040] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0041] Embodiment 1
[0042] Figure 1 This is a flowchart of a method for training a game image reconstruction network provided in the first embodiment of the present invention. This embodiment is applicable to the case of training a game image reconstruction network to realize a game style. The method can be executed by a training device for a game image reconstruction network. The training device for the game image reconstruction network can be implemented in the form of hardware and / or software. The training device for the game image reconstruction network can be configured in an electronic device. Figure 1 As shown, the method includes:
[0043] Step 101: collect and record first sample image data of the real world.
[0044] In this embodiment, multiple frames of image data can be collected through authorized use, public data sets, self-recording, etc. The content of these image data records the real world and can be used as samples for training the game image reconstruction network, and can be recorded as the first sample image data.
[0045] The real world may include a real natural environment, real buildings, real people, animals, etc., which is not limited in this embodiment.
[0046] Step 102: Determine sample games.
[0047] In this embodiment, a game can be selected according to business needs as a source of style when training the game image reconstruction network, which can be recorded as a sample game.
[0048] Among them, the types of games may include MOBA (Multiplayer Online Battle Arena), RPG (Role-playing game), SLG (Simulation Game), etc., which are not limited in this embodiment.
[0049] For example, an RPG game specifically designed to run on a game console, personal computer or other platform can be selected as a sample game. In this game, the player will play the role of a lost teenager who embarks on an adventure in another country in a different world and experience the teenager's growth process.
[0050] Generally speaking, the same sample game is developed by the same team and has a relatively unified and typical style in art design.
[0051] Step 103: Collect sample video data of the sample game, which is recorded when the user controls the sample game.
[0052] For sample games such as RPG, the screen changes as the plot progresses, which takes a lot of time and energy. Training the game image reconstruction network requires a large number of diverse samples. If the sample game is started directly and the screen is collected as samples in the sample game, the efficiency may be low.
[0053] Therefore, in this embodiment, sample video data recorded when a user controls a sample game can be collected by inviting the user to control the sample game, applying to the copyright holder, downloading from an open video platform, etc., wherein the sample video data content includes the sample game, that is, the screen of the sample video data is the screen of the sample game.
[0054] The picture clarity of the sample video data may be lower than that of the sample game, but in general, the picture clarity of the sample video data is sufficient for training the game image reconstruction network. The methods for obtaining sample video data are diversified, and a large amount of sample video data can be easily obtained. Different users have enough time to advance the plot of the sample game, which makes the picture clarity of the sample video data similar, which can meet the needs of training the game image reconstruction network and greatly improve the efficiency of sample collection.
[0055] In practical applications, the screening and input of sample video data may be the responsibility of artists and / or technicians. When the artists and / or technicians are familiar with the sample games, the screening efficiency is higher.
[0056] However, considering that the sample games have many screens, artists and / or technicians are not necessarily familiar with the screens of each sample game, and the candidate video data is relatively long. It will take a long time for artists and / or technicians to browse and familiarize themselves with the sample games and screen the candidate video data. Therefore, it is possible to analyze the characteristics of the sample games, automatically collect sample video data, and further provide it to artists and / or technicians for verification. This can greatly improve the efficiency of collecting sample video data.
[0057] In one embodiment of the present invention, step 103 may include the following steps:
[0058] Step 1031: Collect candidate video data recorded when the user controls the sample game.
[0059] In this embodiment, video data recorded when a user operates a sample game can be collected by inviting the user to operate the sample game, applying to the copyright holder, downloading from an open video platform, etc., and recorded as candidate video data.
[0060] In actual applications, the situations in which users manipulate sample games are different, so that the candidate video data may contain other pictures in addition to the pictures of the sample games.
[0061] For example, in game live broadcasts, game introductions and other scenarios, in addition to the sample game screen, the candidate video data may also contain video data collected for users. The video data collected for users may occupy an area of the screen or the entire screen.
[0062] For another example, in a scene introducing a game, the candidate video data may include not only the screen of the sample game but also the screen of other games, that is, the candidate video data is a collection of multiple games.
[0063] In one method of collecting candidate video data, the name of a sample game can be determined, and the name and keywords representing the manipulation (such as "play", "learn", etc.) can be used to retrieve video data recorded when a user manipulates the sample game in a preset video library as candidate video data.
[0064] In addition, the title of the candidate video data may be queried and semantic analysis may be performed on the title.
[0065] If the title contains keywords representing a collection, command line tools, library files, etc. can be used to use independent scenes as segmentation nodes, and the candidate video data can be divided into multiple segments according to the scenes, recorded as video segments, where each video segment has one or more independent scenes.
[0066] Furthermore, there are two ways to detect scenes:
[0067] 1. Threshold mode
[0068] For candidate video data with obvious scene boundaries, a threshold mode is applied to compare each frame of image data with the set black level. Based on the detection results, it is determined whether it is a scene boundary such as fade-in, fade-out, or cut to black, thereby dividing each scene in the candidate video data.
[0069] 2. Content Mode
[0070] For candidate video data that switches quickly between scenes, a content model is applied to compare each frame of image data, and image data with large content changes are sequentially searched as nodes for segmentation, thereby dividing each scene in the candidate video data.
[0071] Generally, candidate video data containing an independent scene can be divided into a video segment. Considering that some candidate video data containing an independent scene are short in duration, the scene can be merged with other adjacent scenes, thereby dividing candidate video data containing two or more connected scenes into a video segment. This embodiment does not impose any restrictions on this.
[0072] The duration of scene switching is detected at the head or tail of each video clip, wherein the scene switching is characterized by a white screen, a black screen, a still screen, etc., and can be detected by color, pixel difference, etc.
[0073] In some game introduction collections, the time for switching between different games is usually longer than the time for scene changes within the same game. Therefore, each interval duration can be compared with a preset threshold. If the interval duration exceeds (greater than or equal to) the preset threshold, it can be considered that a dividing point between different games appears, that is, the interval duration corresponding to the dividing point exceeds the preset first threshold.
[0074] The beginning and the end can both be used as a segmentation point, and the video clips between the segmentation points can be merged into a game clip. A game clip usually introduces the content of the same game.
[0075] Generally, the head of each game clip will display game introduction information, such as the game logo (icon), name, etc., then, in this embodiment, optical character recognition (OCR) can be performed on the head of each game clip to obtain text information.
[0076] If the text information contains the name of the sample game, indicating that the game clip is used to introduce the sample game, the game clip is retained in the candidate video data.
[0077] If the text information does not contain the name of the sample game, it means that the game clip is not used to introduce the sample game, and the game clip is filtered out from the candidate video data.
[0078] In this embodiment, for the case of a collection of introduced games, candidate video data for introducing sample games are screened out by switching time length and header information, and interference from other games is filtered out. This can reduce the number of candidate video data and the amount of computation required for subsequent screening of sample video data.
[0079] Step 1032: Identify in the candidate video data a first time point at which a character in the sample game is located and a second time point at which a user appears.
[0080] In this embodiment, a sampling test is performed on the candidate video data to detect characters and users in the sample games that appear in the candidate video data, wherein the users are generally real people, and people with specific identities are not required.
[0081] On the time axis of the candidate video data, the first time point where the character in the sample game is located and the second time point where the user is located are marked respectively.
[0082] In a specific implementation, the characters controllable by the user in the sample game are queried in the database of the sample game. The characters controllable by the user in the sample game are generally the protagonists and main supporting roles in the sample game. These characters appear frequently, which can ensure the accuracy of identifying the screen of the sample game.
[0083] The character's information is recorded in the database, such as name, portrait data (including avatar data), experience, etc.
[0084] Frames are uniformly sampled from the candidate video data, face detection is performed in each frame, face data is obtained, and the face data is matched with the portrait data.
[0085] Among them, face detection is also called face key point detection, positioning or face alignment. It means locating the key areas of the face, including eyebrows, eyes, nose, mouth, facial contours, etc., given face data.
[0086] In this embodiment, face detection can use the following method:
[0087] 1. Use artificially extracted features, such as Haar features, use the features to train classifiers, and use the classifiers for face detection.
[0088] 2. Inherit face detection from general object detection algorithms, for example, using Faster R-CNN to detect faces.
[0089] 3. Use cascaded convolutional neural networks, for example, Cascade CNN (Cascaded Convolutional Neural Networks), MTCNN (Multi-task Cascaded Convolutional Networks).
[0090] Considering that simply marking facial data is sufficient to distinguish between roles and users, the requirements for face detection algorithms are relatively low, noise is allowed, and general convolutional neural networks such as MTCNN can be used for face detection.
[0091] If the face data matches the portrait data successfully, the time point at which the face data is located (ie, the time point at which the frame at which the face data is detected is located) is marked as the first time point.
[0092] If the facial data fails to match the portrait data, the preset face identifier is called to identify whether the facial data is a real human face. The face identifier is a binary classifier that is trained using real human faces and in-game character faces to distinguish whether the facial data is a real human face.
[0093] If the facial data is a real human face, the time point at which the facial data is located (ie, the time point at which the frame where the facial data is detected is located) is marked as the second time point.
[0094] Step 1033: In the candidate video data, at least part of the first time points are connected to form a continuous first time range, and at least part of the second time points are connected to form a continuous second time range.
[0095] On the one hand, each first time point in the candidate video data may be traversed, and for two first time points that are adjacent in position, the first time interval between the two first time points may be calculated.
[0096] If the first time interval is less than or equal to the first interval threshold, it means that the first time interval is short and the probability of belonging to a coherent scene is high, and the two first time points can be connected to obtain a first candidate range.
[0097] If the first time interval is greater than the first interval threshold, it means that the first time interval is long, the probability of belonging to a coherent scene is low, and the two first time points can be ignored.
[0098] After traversing all first time points, in order to ensure the validity of the first time range and reduce the impact of noise when detecting and recognizing face data, a first candidate range whose length exceeds the first time threshold can be screened out as the first time range.
[0099] On the other hand, each second time point in the candidate video data may be traversed, and for two second time points that are adjacent in position, the second time interval between the two second time points may be calculated.
[0100] If the second time interval is less than or equal to the second interval threshold, it means that the second time interval is short and the probability of belonging to a coherent scene is high, and the two second time points can be connected to obtain a second candidate range.
[0101] If the second time interval is greater than the second interval threshold, it means that the second time interval is long, and the probability of belonging to a coherent scene is low, and the two second time points can be ignored.
[0102] After traversing all second time points, in order to ensure the validity of the second time range and reduce the impact of noise when detecting and recognizing face data, a second candidate range whose length exceeds the second time threshold can be screened out as the second time range.
[0103] Step 1034: Delete the partial area intersecting with the second time range within the first time range to obtain a third time range.
[0104] The first time range is compared with the second time range, and a partial area where the first time range and the second time range intersect (ie overlap) is screened out, and this partial area is deleted from the first time range, and the remaining partial area is recorded as the third time range.
[0105] Step 1035: extract candidate video data within the third time range as sample video data whose content is a sample game.
[0106] In the candidate video data, frames within the third time range are extracted respectively to obtain sample video data whose content is a sample game, so that the sample video data reduces or avoids capturing pictures containing users.
[0107] The screen containing users does not have the style of the sample game. Using samples of screens containing users to train the game image reconstruction network may have a certain impact on the performance of the game image reconstruction network. This embodiment can reduce or avoid collecting screens containing users, and can improve the uniformity of the style of the samples, thereby improving the performance of the game image reconstruction network.
[0108] Furthermore, in order to ensure the effectiveness of the first time range and mitigate the influence of noise when detecting and recognizing facial data, the number of third time ranges may be counted separately, and the total length of all third time ranges may be counted.
[0109] The quantity and the total length are linearly fused into the confidence of the third time range, where the confidence is positively correlated with the total length, that is, the larger the total length, the higher the confidence, and vice versa, the smaller the total length, the lower the confidence, and the confidence is negatively correlated with the quantity, that is, the smaller the quantity, the higher the confidence, and vice versa, the larger the quantity, the lower the confidence.
[0110] If the confidence level is greater than or equal to the preset second threshold, indicating that the confidence level is high, the candidate video data within the third time range may be extracted as sample video data whose content is a sample game.
[0111] When some users are live streaming games or introducing games, they provide game screens most of the time and provide their own screens a small part of the time. This can ensure the experience of browsing game content. In this case, the third time range is small in number and large in total length. After excluding the user's own screen, a relatively pure game screen can be obtained.
[0112] When some users are live streaming or introducing games, they continuously display their own images on the game screen. In this case, the number of third time ranges is large and the total length is small, so it is troublesome to remove the user's own image.
[0113] Step 104: extract second sample image data from the sample video data.
[0114] In this embodiment, part or all of the image data may be extracted from the sample video data by random frame sampling, uniform frame sampling, etc., and recorded as the second sample image data.
[0115] In one embodiment of the present invention, step 104 may include the following steps:
[0116] Step 1041: extract a frame of second sample image data from the sample video data at every preset time interval.
[0117] In this embodiment, an interval time may be preset, such as 10 ms, and in the process of traversing the sample video data, a frame of image data may be extracted at each interval of the time and recorded as the second sample image data.
[0118] Step 1042: Detect the target object in the second sample image data.
[0119] In this embodiment, one or more elements may be selected from the sample game as target objects according to business requirements.
[0120] Furthermore, the target object may be a specific element, such as a certain character, a certain mecha, etc., or a broad element, such as any character, any building, etc., which is not limited in this embodiment.
[0121] In practical applications, the screening and input of target objects may be the responsibility of artists and / or technicians. When the artists and / or technicians are familiar with the target objects, the screening efficiency is higher.
[0122] However, considering that the sample game has many screens, artists and / or technicians may not be familiar with each target object. In addition, the video data used as candidates is relatively long. It will take a long time for artists and / or technicians to browse and familiarize themselves with the sample game and screen the target objects. Therefore, a small number of target objects can be used as samples to pre-train the target detection network, such as R-CNN, YOLO, etc. Since the detection of target objects allows a certain amount of noise, a small number of samples can also meet the needs of using the target detection network to detect the target object.
[0123] Exemplarily, the target object includes an anime avatar. In this case, a target detection network may be loaded, and the second sample image data may be input into the target detection network to detect whether the anime avatar exists in the second sample image data.
[0124] This example automatically detects the target object and further provides it to artists and / or technicians for verification, which can greatly improve the efficiency of detecting the target object.
[0125] Step 1043: For the second sample image data in which the target object is detected, retain the second sample image data.
[0126] The second sample image data in which the target object is detected is beneficial to training the game image reconstruction network, and the second sample image data can be retained, that is, the second sample image data in which the target object is detected can all participate in the training of the game image reconstruction network.
[0127] Step 1044: For the second sample image data in which the target object is not detected, filter out part of the second sample image data.
[0128] Generally, the number of second sample image data in which the target object is not detected is relatively large. In order to balance the ratio between the second sample image data in which the target object is detected and the second sample image data in which the target object is not detected, part of the second sample image data can be filtered out. The filtered part of the second sample image data does not participate in the training of the game image reconstruction network, and the remaining part of the second sample image data participates in the training of the game image reconstruction network.
[0129] In an example of filtering, if the target object is not detected in a certain frame of the second sample image data, the second sample image data is divided into a preset set.
[0130] Determine a sampling rate, such as 1 / 2, 1 / 3, etc., randomly extract second sample image data from the set according to the sampling rate, and filter out the second sample image data that is not extracted from the set.
[0131] Step 105 , training the generative adversarial network into a game image reconstruction network using the first sample image data as a source of content and the second sample image data as a source of game style.
[0132] In this embodiment, a Generative Adversarial Network (GAN) may be pre-constructed.
[0133] Generally, a generative adversarial network includes a generator and a discriminator. The generator is responsible for generating content based on a random vector. In this embodiment, the content is image data, especially image data with a game style. The discriminator is responsible for determining whether the received content is real. The discriminator usually gives a probability representing the authenticity of the content.
[0134] The generator and the discriminator can use different structures. For the function of processing image data, these structures are not limited to artificially designed neural networks, such as convolutional layers, fully connected layers, etc., but can also be neural networks optimized by model quantization methods, neural networks searched for characteristics of game styles by NAS (Neural Architecture Search) methods, and so on. This embodiment does not impose any restrictions on this.
[0135] According to the generators and discriminators of different structures, the generative adversarial network can be divided into the following types:
[0136] DCGAN (Deep Convolutional Generative Adversarial Network), CGAN (Conditional Generative Adversarial Network), CycleGAN (Cycle Generative Adversarial Network), CoGAN (Coupled Generative Adversarial Network), ProGAN (Progressive Growing of Generative Adversarial Networks), WGAN (Wasserstein Generative Adversarial Network), SAGAN (Self-Attention Generative Adversarial Network), BigGAN (Big Generative Adversarial Network), StyleGAN (Style-based Generative Adversarial Network).
[0137] There is a confrontation between the generator and the discriminator. The so-called confrontation can refer to the process of alternating training in the generative adversarial network. Taking the generation of image data with a game style as an example, the generator is asked to generate some fake image data and real image data, and give them to the discriminator for discrimination. It is asked to learn to distinguish between the two, and give high scores to real image data (i.e., image data with a game style) and low scores to fake image data (i.e., image data without a game style). When the discriminator can skillfully judge the existing image data, the generator is asked to obtain high scores from the discriminator, and continuously generate better fake image data until it can deceive the discriminator. Repeat this process until the discriminator's predicted probability for any image data is close to 0.5, that is, it is impossible to distinguish the true or false image data, and then the training can be stopped.
[0138] In this embodiment, the first sample image data used to record the real world and the second sample image face data with a game style are samples for training a generative adversarial network. The first sample image data is the source of content, and the second sample image data is the source of the game style. The generative adversarial network is trained in this way, and the trained generative adversarial network is recorded as a game image reconstruction network, so that the game image reconstruction network can be used to reconstruct image data containing a game style.
[0139] Furthermore, the samples for training the generative adversarial network can be selected as paired data, which can improve the performance of the generative adversarial network. However, this requires collecting real-world image data corresponding to the second sample image data. However, in fact, most of the second sample image data do not have corresponding real-world image data. Therefore, the generative adversarial network in this embodiment supports training using unpaired data, such as CycleGAN, StyleGAN, and the like.
[0140] Take Learning to Cartoonize Using White-box Cartoon Representations as an example. The network contains three modules that can separate the original image and style map into three representations:
[0141] 1. Surface characterization
[0142] Extract surface representation to represent smooth surfaces of image data. Given image data, weighted low-frequency components can be extracted, where color components and surface textures are preserved, while edges, textures, and details are ignored, which can be used to achieve flexible and learnable feature representation of smooth surfaces.
[0143] 2. Structure characterization
[0144] Structural representation can effectively capture the global structural information and sparse color blocks in the celluloid cartoon style. Segmented regions are extracted from the input image data, and an adaptive colorization algorithm is applied to each segmented region to generate structural representation. Structural representation can imitate the celluloid cartoon style, which is characterized by clear boundaries and sparse color blocks.
[0145] 3. Texture representation
[0146] The texture representation contains the details and edges of the drawing. The input image data is converted into a single-channel intensity map, where color and brightness are removed and relative pixel intensity is preserved. The texture representation guides the network to learn high-frequency texture details independently, excluding color and brightness patterns.
[0147] The style of image data output is controlled by balancing the weights of surface representation, structure representation, and texture representation.
[0148] In this embodiment, first sample image data of the real world is collected and recorded; a sample game is determined; sample video data of the sample game recorded when the user controls the sample game is collected; second sample image data is extracted from the sample video data; and the generative adversarial network is trained as a game image reconstruction network with the first sample image data as the source of the content and the second sample image data as the source of the game style. The sample video data recorded when the user controls the sample game is used to collect the screen of the sample game, thereby improving the efficiency of making samples, and the generative adversarial network is trained in a non-paired data manner, so that the game image reconstruction network can reconstruct the image data into the game style. Reconstructing the game style belongs to post-processing, which can maintain the threshold for making video data and the time consumption for making video data, thereby greatly improving the efficiency of making video data with a game style.
[0149] Embodiment 2
[0150] Figure 2 This is a flowchart of an image reconstruction method provided in the second embodiment of the present invention. This embodiment is applicable to the case where image data is reconstructed into a game style based on a game image reconstruction network. The method can be executed by an image reconstruction device. The image reconstruction device can be implemented in the form of hardware and / or software. The image reconstruction device can be configured in an electronic device. Figure 2 As shown, the method includes:
[0151] Step 201, loading the game image reconstruction network.
[0152] In a specific implementation, a game image reconstruction network may be trained in advance according to the method described in Embodiment 1 of the present invention, wherein the game image reconstruction network may be used to reconstruct image data containing a game style.
[0153] When applying the game image reconstruction network, the game image reconstruction network and its parameters are loaded into the memory for running.
[0154] Step 202: Obtain original image data to be reconstructed.
[0155] Generally speaking, the structure of the game image reconstruction network is relatively large, occupies more resources, and is usually deployed on the server side. The server side can encapsulate the game image reconstruction network into interfaces, plug-ins, etc., and provide game style reconstruction services to users on the local area network or the public network. Users can call the interface, plug-in, etc. through the client or browser to transmit the image data of the game style to be reconstructed to the server side. For the sake of distinction, the image data of the game style to be reconstructed is recorded as the original image data.
[0156] Of course, if electronic devices such as personal computers and laptops have more local resources to meet the operation of the game image reconstruction network, the game image reconstruction network can be loaded and run locally on the electronic device. At this time, the original image data of the game style to be reconstructed can be input through the command line or other means. This embodiment does not limit this.
[0157] Among them, game style can refer to the style reflected by the sample game as a whole.
[0158] Step 203: input the original image data into the game image reconstruction network to reconstruct it into candidate image data containing the game style.
[0159] In this embodiment, the original image data is input into the game image reconstruction network, which processes the original image data according to its structure, and reconstructs the original image data into new image data containing the game style while maintaining the content of the original image data, which is recorded as candidate image data.
[0160] Step 204: Add stroke information to the candidate image data to obtain target image data.
[0161] In this embodiment, the candidate image data can be processed, and stroke information can be added on the basis of the candidate image data to obtain target image data. The stroke information refers to adding lines to the edges of each element, which has the effect of strengthening the edges of each element with a game style.
[0162] In a specific implementation, RVM (Relevance Vector Machine) can be used to cut out the portrait in the original image data, so as to extract the portrait image data containing the portrait information as a mask.
[0163] Use algorithms such as Anime2sketch to convert raw image data into sketch-style sketch image data.
[0164] The portrait image data (matrix) is multiplied with the sketch image data (matrix) to obtain a portrait sketch, which belongs to stroke information. At this time, the portrait sketch can be superimposed on the candidate image data to obtain the target image data.
[0165] In one example, Figure 3A The original image data shown is input into the game image reconstruction network, and the reconstruction is as follows Figure 3B The candidate image data shown is Figure 3B The candidate image data shown is compared to Figure 3A The original image data shown in the figure shows a significant change in the style of characters and floor tiles, which is more inclined to the sample game, highlighting the style of sketching (especially the stroke). Figure 3A The original image data shown is extracted as Figure 3C The portrait image data shown in FIG. Figure 3A The original image data shown is multiplied by the converted sketch image data to obtain Figure 3D The portrait sketch shown in Figure 3B The candidate image data shown is superimposed on the Figure 3D The portrait sketch shown in Figure 3E The target image data is shown.
[0166] In this embodiment, a game image reconstruction network is loaded; the original image data to be reconstructed is obtained; the original image data is input into the game image reconstruction network to be reconstructed into candidate image data containing a game style; and stroke information is added to the candidate image data to obtain target image data. When training the game image reconstruction network, the sample video data recorded when the user controls the sample game is used to collect the screen of the sample game, thereby improving the efficiency of making samples, and training the generative adversarial network in a non-paired data manner, so that the game image reconstruction network can reconstruct the image data into the game style. Reconstructing the game style belongs to post-processing, which can maintain the threshold for making video data and the time consumption for making video data, thereby greatly improving the efficiency of making video data with a game style.
[0167] Embodiment 3
[0168] Figure 4This is a flowchart of a video reconstruction method provided in the third embodiment of the present invention. This embodiment is applicable to the case where video data is reconstructed into a game style based on a game image reconstruction network. The method can be executed by a video reconstruction device. The video reconstruction device can be implemented in the form of hardware and / or software. The video reconstruction device can be configured in an electronic device. Figure 4 As shown, the method includes:
[0169] Step 401: Load the game image reconstruction network.
[0170] In a specific implementation, a game image reconstruction network may be trained in advance according to the method described in Embodiment 1 of the present invention, wherein the game image reconstruction network may be used to reconstruct image data containing a game style.
[0171] When applying the game image reconstruction network, the game image reconstruction network and its parameters are loaded into the memory for running.
[0172] Step 402: Acquire original video data that introduces the target game.
[0173] In this embodiment, artists can create video data for the target game to be promoted, that is, the content of the video data is used to introduce the game. For easy distinction, the game is recorded as the target game and the video data is recorded as the original video data.
[0174] The types of the target games may include MOBA, RPG, SLG, etc., which are not limited in this embodiment.
[0175] In a specific implementation, the content of the original video data can be divided into two main forms: game content and real plot. The plot can be further divided into the following categories:
[0176] 1. Sharing fake food
[0177] The original video data contains some food-related materials, which can attract the user's attention. Secondly, the gameplay of making money to eat food is implanted, and at the same time, it provides users with a clear goal of playing the game.
[0178] 2. Topics that are close to users’ daily lives
[0179] The original video data is close to the current life of the user, and the selling points of the game are implanted into all aspects of life. The user can use the game to make money and pay by purchasing props of the target game, eating, buying snacks, etc. The production of this type of material is also relatively simple, with a single scene and low shooting difficulty. The first half of the material is mainly a dialogue between two people, and the second half is an implanted clip of the game.
[0180] 3. Sitcom
[0181] The original video data contains sitcom material, in some cases celebrities endorse the game wearing clothing, and some of the plots are exaggerated to attract users' attention.
[0182] Generally speaking, the structure of the game image reconstruction network is relatively large, occupies more resources, and is usually deployed on the server side. The server side can encapsulate the game image reconstruction network into interfaces, plug-ins, etc., and provide game style reconstruction services to users on the local area network or the public network. Users can call the interface, plug-in, etc. through the client or browser to transmit the original video data of the game style to be reconstructed to the server side.
[0183] Of course, if electronic devices such as personal computers and laptops have more local resources to meet the operation of the game image reconstruction network, the game image reconstruction network can be loaded and run locally on the electronic device. At this time, the original video data of the game style to be reconstructed can be input through the command line or other methods. This embodiment does not limit this.
[0184] Among them, game style can refer to the style reflected by the sample game as a whole.
[0185] Step 403: input the original image data into the game image reconstruction network to reconstruct it into candidate image data containing the game style.
[0186] In a specific implementation, the original video data contains multiple frames of image data, which are recorded as original image data. For the original video data, each frame of original image data can be input into a game image reconstruction network. The game image reconstruction network processes the original image data according to its structure, and reconstructs the original image data into new image data containing a game style while maintaining the content of the original image data, which is recorded as candidate image data.
[0187] Step 404: Add stroke information to the candidate image data to obtain target image data.
[0188] In this embodiment, the candidate image data can be processed, and stroke information can be added on the basis of the candidate image data to obtain target image data. The stroke information refers to adding lines to the edges of each element, which has the effect of strengthening the edges of each element with a game style.
[0189] In a specific implementation, RVM (Relevance Vector Machine) can be used to cut out the portrait in the original image data, so as to extract the portrait image data containing the portrait information as a mask.
[0190] Use algorithms such as Anime2sketch to convert raw image data into sketch-style sketch image data.
[0191] The portrait image data (matrix) is multiplied with the sketch image data (matrix) to obtain a portrait sketch, which belongs to stroke information. At this time, the portrait sketch can be superimposed on the candidate image data to obtain the target image data.
[0192] Step 405: Replace the original image data with the target image data in the original video data to obtain the target video data.
[0193] In the original video data, the target image data may replace the corresponding original image data to obtain the target video data.
[0194] Thereafter, advertising element data related to the target game may be added to the target video data to obtain advertising video data, wherein the advertising element data includes the LOGO (icon) of the platform used to distribute the target game, Banner (banner ad), EC (ending clip, generally containing information of the target game (such as name, platform for distributing the target game, etc.)), and the like.
[0195] Advertising video data is published on designated channels (such as news information, short videos, novel reading, sports and health, etc.) so that when the client accesses the channel, the advertising video data is pushed to the client for playback. When the user is interested in the target game, he downloads the target game from the game distribution platform.
[0196] In this embodiment, a game image reconstruction network is loaded; the original video data of the target game is obtained, and the original video data has multiple frames of original image data; the original image data is input into the game image reconstruction network to be reconstructed into candidate image data containing the game style; the stroke information is added to the candidate image data to obtain the target image data; the target image data is replaced by the original image data in the original video data to obtain the target video data. When training the game image reconstruction network, the sample video data recorded when the user controls the sample game is used to collect the screen of the sample game, so as to improve the efficiency of making samples, and the generative adversarial network is trained in a non-paired data manner, so that the game image reconstruction network can reconstruct the image data into the game style. The reconstruction of the game style belongs to the post-processing, which can maintain the threshold of making video data and the time consumption of making video data, and greatly improve the efficiency of making video data with game style.
[0197] Embodiment 4
[0198] Figure 5 This is a schematic diagram of the structure of a training device for a game image reconstruction network provided by the fourth embodiment of the present invention. Figure 5 As shown, the device comprises:
[0199] The content sample collection module 501 is used to collect and record first sample image data of the real world;
[0200] A sample game determination module 502, used to determine a sample game;
[0201] The video data acquisition module 503 is used to acquire sample video data of the sample game recorded when the user controls the sample game;
[0202] A style sample extraction module 504, configured to extract second sample image data from the sample video data;
[0203] The generative adversarial network training module 505 is used to train the generative adversarial network into a game image reconstruction network using the first sample image data as a source of content and the second sample image data as a source of game style.
[0204] In one embodiment of the present invention, the video data acquisition module 503 is further used for:
[0205] Collecting candidate video data recorded when a user controls the sample game;
[0206] identifying in the candidate video data a first time point at which the character in the sample game is located and a second time point at which the user appears;
[0207] In the candidate video data, at least part of the first time points are connected to form a continuous first time range, and at least part of the second time points are connected to form a continuous second time range;
[0208] Deleting a portion of the area intersecting with the second time range within the first time range to obtain a third time range;
[0209] The candidate video data within the third time range is extracted as sample video data whose content is the sample game.
[0210] In one embodiment of the present invention, the video data acquisition module 503 is further used for:
[0211] determining the name of the sample game;
[0212] The name and the key words representing the manipulation are used to retrieve candidate video data recorded when the user manipulates the sample game in a preset video library.
[0213] In one embodiment of the present invention, the video data acquisition module 503 is further used for:
[0214] Querying the title of the candidate video data;
[0215] If the title contains a keyword indicating a collection, the candidate video data is divided into a plurality of video segments according to the scene;
[0216] Detecting the scene switching interval at the head or tail of each video segment;
[0217] Merging the video segments between the segmentation points into game segments, wherein the interval durations corresponding to the segmentation points exceed a preset first threshold;
[0218] Perform optical character recognition on the header of each of the game clips to obtain text information;
[0219] If the text information includes the name of the sample game, retaining the game clip in the candidate video data;
[0220] If the text information does not include the name of the sample game, the game clip is filtered out from the candidate video data.
[0221] In one embodiment of the present invention, the video data acquisition module 503 is further used for:
[0222] Searching a database of the sample game for a character that the user can control in the sample game, the character having portrait data;
[0223] Performing face detection in the candidate video data to obtain face data;
[0224] If the facial data matches the portrait data successfully, marking the time point of the facial data as the first time point;
[0225] If the face data fails to match the portrait data, calling a preset face identifier to identify whether the face data is a real human face;
[0226] If the facial data is a real human face, the time point at which the facial data is located is marked as the second time point.
[0227] In one embodiment of the present invention, the video data acquisition module 503 is further used for:
[0228] Counting the number of the third time range;
[0229] Counting the total length of all the third time ranges;
[0230] Linearly fusing the quantity and the total length into the confidence of the third time range, wherein the confidence is positively correlated with the total length and negatively correlated with the quantity;
[0231] If the confidence level is greater than or equal to a preset second threshold, the candidate video data within the third time range is extracted as sample video data whose content is the sample game.
[0232] In one embodiment of the present invention, the style sample extraction module 504 is further used for:
[0233] Extracting a frame of second sample image data from the sample video data at every preset time interval;
[0234] detecting a target object in the second sample image data;
[0235] For the second sample image data in which the target object is detected, retain the second sample image data;
[0236] For the second sample image data in which the target object is not detected, a portion of the second sample image data is filtered out.
[0237] In one embodiment of the present invention, the target object includes an anime avatar; the style sample extraction module 504 is further used to:
[0238] Load the object detection network;
[0239] The second sample image data is input into the target detection network to detect whether there is an animated avatar in the second sample image data.
[0240] In one embodiment of the present invention, the style sample extraction module 504 is further used for:
[0241] If the target object is not detected in the second sample image data of a certain frame, dividing the second sample image data into a preset set;
[0242] Determine the sampling rate;
[0243] Randomly extract the second sample image data from the set according to the sampling rate;
[0244] The second sample image data not extracted from the set is filtered out.
[0245] The training device for the game image reconstruction network provided in the embodiment of the present invention can execute the training method for the game image reconstruction network provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the training method for the game image reconstruction network.
[0246] Embodiment 5
[0247] Figure 6 FIG. 5 is a schematic diagram of the structure of an image reconstruction device provided in Embodiment 5 of the present invention. Figure 6As shown, the device comprises:
[0248] A reconstruction network loading module 601, used to load a game image reconstruction network trained according to the method of any embodiment of the present invention;
[0249] The original image data acquisition module 602 is used to acquire the original image data to be reconstructed;
[0250] A candidate image data generating module 603 is used to input the original image data into the game image reconstruction network to reconstruct it into candidate image data containing game styles;
[0251] The target image data generating module 604 is used to add stroke information to the candidate image data to obtain the target image data.
[0252] In one embodiment of the present invention, the target image data generating module 604 is further used for:
[0253] Extracting portrait image data containing portrait information from the original image data;
[0254] Converting the original image data into sketch image data in a sketch style;
[0255] Multiplying the portrait image data with the sketch image data to obtain a portrait sketch;
[0256] The portrait sketch is superimposed on the candidate image data to obtain target image data.
[0257] The image reconstruction device provided in the embodiment of the present invention can execute the image reconstruction method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the image reconstruction method.
[0258] Embodiment 6
[0259] Figure 7 This is a schematic diagram of the structure of a video reconstruction device provided by Embodiment 6 of the present invention. Figure 7 As shown, the device comprises:
[0260] A reconstruction network loading module 701, used to load a game image reconstruction network trained according to the method of any embodiment of the present invention;
[0261] The video data acquisition module 702 is used to acquire original video data whose content is to introduce the target game, wherein the original video data has multiple frames of original image data;
[0262] A candidate image data generating module 703 is used to input the original image data into the game image reconstruction network to reconstruct it into candidate image data containing game styles;
[0263] A target image data generating module 704 is used to add stroke information to the candidate image data to obtain target image data;
[0264] The target video data generating module 705 is used to replace the original image data with the target image data in the original video data to obtain the target video data.
[0265] In one embodiment of the present invention, it also includes:
[0266] An advertisement video data generating module, configured to add advertisement element data related to the target game to the target video data as advertisement video data;
[0267] The advertising video data publishing module is used to publish the advertising video data on a designated channel, so that when the client accesses the channel, the advertising video data is pushed to the client for playing.
[0268] The video reconstruction device provided in the embodiment of the present invention can execute the video reconstruction method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the video reconstruction method.
[0269] Embodiment 7
[0270] Figure 8 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0271] like Figure 8As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0272] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0273] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The processor 11 executes the various methods and processes described above, such as a training method for a game image reconstruction network, an image reconstruction method, or a video reconstruction method.
[0274] In some embodiments, the training method of the game image reconstruction network or the image reconstruction method or the video reconstruction method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the training method of the game image reconstruction network or the image reconstruction method or the video reconstruction method described above may be executed. Alternatively, in other embodiments, the processor 11 may be configured to execute the training method of the game image reconstruction network or the image reconstruction method or the video reconstruction method by any other appropriate means (e.g., by means of firmware).
[0275] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0276] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0277] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0278] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0279] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0280] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.
[0281] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.
[0282] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A training method for a game image reconstruction network, It is characterized in that include: Collect and record first sample image data of the real world; Identify sample games; Collecting sample video data recorded when the user controls the sample game; The content of the sample video data includes the sample game, which means that the screen of the sample video data is the screen of the sample game; extracting second sample image data from the sample video data; Taking the first sample image data as a source of content and the second sample image data as a source of game style, training a generative adversarial network into a game image reconstruction network; The collecting of sample video data recorded when the user controls the sample game includes: Collecting candidate video data recorded when a user controls the sample game; identifying in the candidate video data a first time point at which the character in the sample game is located and a second time point at which the user appears; In the candidate video data, respectively connecting at least part of the first time points into a continuous first time range and connecting at least part of the second time points into a continuous second time range, comprising: traversing each of the first time points in the candidate video data, and for two adjacent first time points, calculating a first time interval between the two first time points, and if the first time interval is less than or equal to a first interval threshold, connecting the two first time points to obtain a first candidate range; traversing each of the second time points in the candidate video data, and for two adjacent second time points, calculating a second time interval between the two second time points, and if the second time interval is less than or equal to a second interval threshold, connecting the two second time points to obtain a second candidate range; Deleting a portion of the area intersecting with the second time range within the first time range to obtain a third time range; The candidate video data within the third time range is extracted as sample video data whose content is the sample game.
2. The method according to claim 1, It is characterized in that The collecting of candidate video data recorded when the user manipulates the sample game includes: determining the name of the sample game; The name and the key words representing the manipulation are used to retrieve candidate video data recorded when the user manipulates the sample game in a preset video library.
3. The method according to claim 2, It is characterized in that The collecting of candidate video data recorded when the user manipulates the sample game also includes: Querying the title of the candidate video data; If the title contains a keyword indicating a collection, the candidate video data is divided into a plurality of video segments according to the scene; Detecting the scene switching interval at the head or tail of each video segment; Merging the video segments between the segmentation points into game segments, wherein the interval durations corresponding to the segmentation points exceed a preset first threshold; Perform optical character recognition on the header of each of the game clips to obtain text information; If the text information includes the name of the sample game, retaining the game clip in the candidate video data; If the text information does not include the name of the sample game, the game clip is filtered out from the candidate video data.
4. The method according to claim 1, It is characterized in that The step of identifying in the candidate video data a first time point at which the character in the sample game is located and a second time point at which the user appears, comprises: Searching a database of the sample game for a character that the user can control in the sample game, the character having portrait data; Performing face detection in the candidate video data to obtain face data; If the facial data matches the portrait data successfully, marking the time point of the facial data as the first time point; If the face data fails to match the portrait data, calling a preset face identifier to identify whether the face data is a real human face; If the facial data is a real human face, the time point at which the facial data is located is marked as the second time point.
5. The method according to claim 1, It is characterized in that The step of extracting the candidate video data within the third time range as sample video data containing the sample game includes: Counting the number of the third time range; Counting the total length of all the third time ranges; Linearly fusing the quantity and the total length into the confidence of the third time range, wherein the confidence is positively correlated with the total length and negatively correlated with the quantity; If the confidence level is greater than or equal to a preset second threshold, the candidate video data within the third time range is extracted as sample video data whose content is the sample game.
6. The method according to any one of claims 1 to 5, It is characterized in that The extracting second sample image data from the sample video data comprises: Extracting a frame of second sample image data from the sample video data at every preset time interval; detecting a target object in the second sample image data; For the second sample image data in which the target object is detected, retain the second sample image data; For the second sample image data in which the target object is not detected, a portion of the second sample image data is filtered out.
7. The method according to claim 6, It is characterized in that The target object includes an animated avatar; and detecting the target object in the second sample image data includes: Load the object detection network; The second sample image data is input into the target detection network to detect whether there is an animated avatar in the second sample image data.
8. The method according to claim 6, It is characterized in that The filtering out a portion of the second sample image data in which the target object is not detected comprises: If the target object is not detected in the second sample image data of a certain frame, dividing the second sample image data into a preset set; Determine the sampling rate; Randomly extract the second sample image data from the set according to the sampling rate; The second sample image data not extracted from the set is filtered out.
9. An image reconstruction method, It is characterized in that include: Loading a game image reconstruction network trained according to the method described in any one of claims 1 to 8; Acquire original image data to be reconstructed; Inputting the original image data into the game image reconstruction network to reconstruct it into candidate image data containing a game style; Add stroke information to the candidate image data to obtain target image data.
10. The method according to claim 9, It is characterized in that Adding the stroke information to the candidate image data to obtain the target image data includes: Extracting portrait image data containing portrait information from the original image data; Converting the original image data into sketch image data in a sketch style; Multiplying the portrait image data with the sketch image data to obtain a portrait sketch; The portrait sketch is superimposed on the candidate image data to obtain target image data.
11. A video reconstruction method, It is characterized in that include: Loading a game image reconstruction network trained according to the method described in any one of claims 1 to 8; The acquired content is original video data introducing a target game, wherein the original video data contains multiple frames of original image data; Inputting the original image data into the game image reconstruction network to reconstruct it into candidate image data containing a game style; Adding stroke information to the candidate image data to obtain target image data; The target image data is used to replace the original image data in the original video data to obtain the target video data.
12. The method according to claim 11, It is characterized in that Also includes: Adding advertisement element data related to the target game to the target video data as advertisement video data; The advertisement video data is published on a designated channel, so that when the client accesses the channel, the advertisement video data is pushed to the client for playing.
13. An electronic device, It is characterized in that The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the training method of the game image reconstruction network described in any one of claims 1-8, or the image reconstruction method described in any one of claims 9-10, or the video reconstruction method described in any one of claims 11-12.
14. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, which is used to enable a processor to implement the training method of the game image reconstruction network described in any one of claims 1-8, the image reconstruction method described in any one of claims 9-10, or the video reconstruction method described in any one of claims 11-12 when executed.