Game graphic design automatic scene matching method and system based on deep learning

Through deep learning-based methods, image segmentation, feature extraction and semantic label generation, the semantic understanding and element matching problems of game scene design in the prior art are solved, and efficient and personalized game scene design is achieved.

CN120107383AInactive Publication Date: 2025-06-06SHENZHEN AOKU TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510111155.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing game scene design technologies face key technical bottlenecks such as inaccurate image semantic understanding, unintelligent scene element matching, and lack of personalized design. It is difficult to achieve accurate segmentation and semantic understanding of complex scene backgrounds, and it is impossible to effectively evaluate the semantic correlation and degree of matching between planar elements and scene areas.

Method used

Using a deep learning-based method, we obtain the background map and plane element image library of the game scene, perform image segmentation and feature extraction, establish a semantic label generation model, generate scene semantic labels, and calculate the matching scores of plane elements and the area to be matched through a deep convolutional neural network to generate a complete game scene design diagram.

Benefits of technology

It realizes precise segmentation and semantic understanding of complex scene backgrounds, improves the intelligence of scene element matching and personalized design capabilities, optimizes the selection and layout of game materials, and improves the consistency of design efficiency and effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107383A_ABST
    Figure CN120107383A_ABST
Patent Text Reader

Abstract

The invention discloses a game graphic design automatic scene matching method and system based on deep learning, and relates to the technical field of image matching, and the method comprises the steps: obtaining a background base map and a plurality of plane element image libraries of a game scene; performing image segmentation on the background base image, dividing the background base image into a plurality of to-be-matched areas, and extracting an image feature vector of each to-be-matched area; establishing a semantic tag generation model, and generating a scene semantic tag of each to-be-matched area; according to the scene semantic tag, screening a plane element matched with each to-be-matched area from a plane element image library, and calculating a matching degree score of the plane element and the to-be-matched area; and placing the plane element with the highest matching degree score in the corresponding to-be-matched area to generate a complete game scene design drawing. The complete and coordinated game scene design drawing can be automatically generated, selection and layout of game materials are optimized, and design efficiency and effect consistency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image matching technology, and in particular to a method and system for automatic scene matching in game plane design based on deep learning. Background Art

[0002] With the rapid development of the digital entertainment industry, game art design has become a key factor in promoting the competitiveness of game products. Traditional game scene design often relies on the manual creation of artists, which not only consumes a lot of manpower and time costs, but is also limited by the personal experience and creativity of designers. In recent years, artificial intelligence technology, especially deep learning, has made breakthrough progress in the fields of image processing and computer vision, providing a new technical path and solution for the intelligent design of game scenes. Many game development companies and research institutions at home and abroad have begun to try to use machine learning algorithms to explore the possibility of automatic generation and intelligent matching of game scenes.

[0003] Existing game scene design technologies are mainly faced with key technical bottlenecks such as inaccurate image semantic understanding, unintelligent scene element matching, and lack of personalized design. Traditional methods usually rely on preset rules and manual experience, making it difficult to achieve accurate segmentation and semantic understanding of complex scene backgrounds, and are unable to effectively evaluate the semantic relevance and matching degree between planar elements and scene areas. These limitations often make it difficult to ensure the accuracy and beauty of automatic scene design, and cannot meet the urgent needs of game art design for high-quality, personalized scenes. Especially in the design of diversified and highly customized game scenes, existing technologies still have significant deficiencies in semantic understanding, element matching, and creative generation. Summary of the invention

[0004] In view of the problems existing in the above-mentioned automated FOTA testing method for lithium battery systems, the present invention is proposed.

[0005] Therefore, the present invention provides an automatic scene matching method for game plane design based on deep learning, which can solve the problems mentioned in the background technology.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0007] In a first aspect, an embodiment of the present invention provides a method for automatic scene matching of game plane design based on deep learning, which comprises:

[0008] Get the background map of the game scene and multiple graphic element image libraries;

[0009] Performing image segmentation on the background image to divide the background image into a plurality of to-be-matched regions, and extracting an image feature vector of each to-be-matched region;

[0010] Establishing a semantic label generation model, inputting the image feature vector into the semantic label generation model, and generating a scene semantic label for each of the to-be-matched areas;

[0011] According to the scene semantic label, plane elements matching each of the to-be-matched regions are screened from the plane element image library, and matching scores between the plane elements and the to-be-matched regions are calculated using a deep convolutional neural network;

[0012] The plane element with the highest matching score is placed in the corresponding area to be matched to generate a complete game scene design diagram.

[0013] As a preferred solution of the method for automatic scene matching of game plane design based on deep learning of the present invention, wherein: the plane element image library includes a character material library, a prop material library and a building material library;

[0014] The character material library includes human characters, biological characters and mechanical characters, and each character material is equipped with images from multiple perspectives;

[0015] The prop material library includes weapon props, consumable props and decorative props, and the prop material is in transparent background PNG format;

[0016] The building material library includes building bodies, building decorations and building ancillary facilities, and the building materials are marked with scale information;

[0017] Each material in the plane element image library contains preset tag information, and the preset tag information includes material category, usage scenario, style attribute and size information.

[0018] As a preferred solution of the method for automatic scene matching of game plane design based on deep learning of the present invention, extracting the image feature vector of each area to be matched includes:

[0019] Performing preliminary segmentation on the background map to obtain coarse-grained segmentation areas;

[0020] Refining the coarse-grained segmented area to divide the background map into a plurality of to-be-matched areas of different sizes;

[0021] Extracting a feature vector for each of the to-be-matched regions; the image feature vector includes color features, texture features, edge features, and spatial position features;

[0022] The extracted feature vector is normalized to generate a standardized image feature vector.

[0023] As a preferred solution of the method for automatic scene matching of game plane design based on deep learning of the present invention, the coarse-grained segmentation area is refined by using a superpixel segmentation algorithm; the distance metric formula of the superpixel segmentation algorithm is as follows:

[0024]

[0025] Among them, D is the comprehensive distance measurement value, d c is the color distance, d s is the spatial distance, d t is the texture distance, α, β, γ are adaptive weight coefficients, and the calculation method is:

[0026]

[0027] in, are the color gradient, position gradient and texture gradient of the image respectively, 1 , 2 , 3 is the balance factor, σ 1 is the scale parameter of the color gradient, σ 2 is the scale parameter of the position gradient, σ 3 is the scale parameter of the texture gradient.

[0028] As a preferred solution of the method for automatic scene matching of game plane design based on deep learning described in the present invention, wherein: the semantic label generation model includes a feature enhancement branch and a semantic understanding branch;

[0029] The feature enhancement branch adopts an improved Transformer encoder structure and adds a spatial attention module on the basis of the standard Transformer; the spatial attention module includes a position-aware self-attention layer and a feature enhancement layer;

[0030] The semantic understanding branch adopts a multi-scale feature extraction structure, which includes multiple feature extraction blocks of different scales, and each feature extraction block is composed of multiple residual convolution units connected in series;

[0031] The feature aggregation function of the semantic understanding branch is expressed as follows:

[0032]

[0033] Among them, α l is the weight coefficient of the first layer, CrossAttn is the cross attention module, F global and F local are global and local features respectively, and F semantic is the semantic feature output, L is the number of feature layers, Conv(F l) is the convolutional feature of the lth layer.

[0034] As a preferred solution of the automatic scene matching method for game plane design based on deep learning of the present invention, wherein: the plane elements matching each of the to-be-matched areas are screened from the plane element image library using a multi-level semantic association structure, including three levels of main category mapping, subcategory mapping and attribute mapping;

[0035] The main category mapping determines the basic material type; the subcategory mapping refines the specific material type; the attribute mapping matches style, size, and tone;

[0036] The matching scores include a semantic relevance score, a visual coordination score, a spatial adaptability score, and a scene coherence score.

[0037] As a preferred solution of the automatic scene matching method for game plane design based on deep learning of the present invention, wherein: the semantic relevance score quantifies the hierarchical matching degree of semantic labels by calculating the cosine similarity of the semantic vectors of the to-be-matched area and the candidate plane elements to obtain the compatibility score of the semantic attributes;

[0038] The visual harmony score quantifies the similarity of texture features by calculating the histogram matching of color distribution;

[0039] The spatial fitness score evaluates the matching degree between the size of the plane element and the area of ​​the area to be matched, analyzes the rationality of the spatial position by calculating the degree of fit of the shape features, and evaluates the naturalness of the transition with the adjacent area;

[0040] The scene coherence score evaluates the coordination of the overall scene atmosphere by calculating the style consistency with the surrounding matched elements.

[0041] The semantic relevance score, the visual coordination score, the spatial adaptability score and the scene coherence score are weightedly fused to obtain a matching score between the plane element and the area to be matched.

[0042] In the second aspect, in order to further solve the security problem existing in the automatic scene matching method of game plane design based on deep learning, the embodiment of the present invention provides an automatic scene matching system of game plane design based on deep learning, which includes:

[0043] The acquisition module is used to obtain the background map of the game scene and multiple plane element image libraries;

[0044] A segmentation module, used for performing image segmentation on the background base map, and dividing the background base map into a plurality of areas to be matched;

[0045] A feature vector extraction module, used to extract the image feature vector of each of the to-be-matched areas;

[0046] A model building module, used to establish a semantic label generation model, input the image feature vector into the semantic label generation model, and generate a scene semantic label for each of the to-be-matched areas;

[0047] A matching module, used to screen plane elements matching each of the to-be-matched regions from the plane element image library according to the scene semantic labels, and calculate a matching score between the plane elements and the to-be-matched regions through a deep convolutional neural network;

[0048] The design drawing generation module is used to place the plane element with the highest matching score into the corresponding area to be matched to generate a complete game scene design drawing.

[0049] In a third aspect, an embodiment of the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the method for automatic scene matching of game plane design based on deep learning as described in the first aspect of the present invention is implemented.

[0050] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the method for automatic scene matching of game plane design based on deep learning as described in the first aspect of the present invention.

[0051] The beneficial effect of the present invention is to provide an efficient method for automatic design of game scenes by using an improved image segmentation algorithm, a deep convolutional neural network, a semantic label mapping and a multi-dimensional matching scoring mechanism. First, an image segmentation technology based on an improved U-Net network and an adaptive weight adjustment are used to accurately identify and divide complex areas in the background image, especially those with rich textures. Subsequently, the matching features of the plane elements and the areas to be matched are extracted through the mapping of multi-level semantic labels and the feature fusion of the deep learning model, and the appropriate material elements are accurately screened in combination with the semantic, visual, spatial and scene coherence scores. Finally, the system can automatically generate a complete and coordinated game scene design drawing, optimize the selection and layout of game materials, and improve the consistency of design efficiency and effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them:

[0053] Figure 1 Flowchart of the automatic scene matching method for deep learning-based game graphics. DETAILED DESCRIPTION

[0054] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.

[0055] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0056] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0057] Example 1

[0058] Reference Figure 1 , which is the first embodiment of the present invention, and provides a method for automatic scene matching of game plane design based on deep learning, comprising the following steps:

[0059] S1: Obtain a background map of a game scene and a plurality of plane element image libraries, wherein the plane element image libraries contain game materials such as characters, props, and buildings;

[0060] A background map of a game scene and a plurality of plane element image libraries are obtained, wherein the resolution of the background map is not less than 1920×1080 pixels, and the background map adopts an RGB three-channel format; the plane element image library is divided into a character material library, a prop material library and a building material library according to a preset classification rule; the character material library includes human characters, biological characters and mechanical characters, and each character material is equipped with images from multiple perspectives; the prop material library includes weapon props, consumable props and decorative props, and the prop material adopts a transparent background PNG format; the building material library includes building main bodies, building decorations and building ancillary facilities, and the building materials are marked with scale information; each material in the plane element image library includes preset label information, and the preset label information includes material category, usage scenario, style attribute and size information.

[0061] S2: performing image segmentation on the background image, dividing the background image into a plurality of to-be-matched regions, and extracting an image feature vector of each to-be-matched region;

[0062] First, the background map is preliminarily segmented using a semantic segmentation algorithm based on an improved U-Net convolutional neural network. The improved U-Net network introduces an attention mechanism, and its attention weight $w_{att}$ is calculated as follows:

[0063]

[0064] Among them, W q , W k is the weight matrix, Q, K, V are the query, key, and value matrices, d k is the feature dimension, and σ is the softmax function.

[0065] Subsequently, the coarse-grained segmented area is refined based on the improved SLIC superpixel segmentation algorithm, and the background base map is divided into a plurality of to-be-matched areas of different sizes, wherein the distance measurement formula of the superpixel segmentation algorithm is as follows:

[0066]

[0067] Among them, D is the comprehensive distance metric, which indicates the similarity between two superpixel regions, d c is the color distance, d s is the spatial distance, d t is the texture distance, α, β, γ are adaptive weight coefficients, and the calculation method is:

[0068]

[0069] in, are the color gradient, position gradient and texture gradient of the image respectively, 1 , 2 , 3 is the balance factor, σ 1 is the scale parameter of the color gradient, σ 2 is the scale parameter of the position gradient, σ 3 is the scale parameter of the texture gradient.

[0070] Better, through practice it is found that the traditional SLIC algorithm only considers color distance and spatial distance, ignoring texture information, resulting in unsatisfactory segmentation effect in texture-rich game scenes, and it uses fixed weights for distance calculation, which cannot adapt to the differences in image features in different regions, and is prone to over-segmentation or under-segmentation, so it is improved; by increasing the adaptive weight coefficient as mentioned above and dynamically adjusting the importance of each distance term, the boundary recall rate is increased by 15.2%, the over-segmentation rate is reduced by 12.8%, and the segmentation accuracy in strong texture areas is increased by 18.5% compared with the traditional SLIC algorithm, but only 8.5% of the computational overhead is increased, so that complex texture areas in game scenes, such as grass, rocks and other natural elements, can be accurately identified, and the outline and detail features of buildings can be accurately maintained.

[0071] For each of the regions to be matched, feature vectors are extracted through an improved ResNet-50 deep convolutional network. The improved part is that by introducing a multi-scale feature fusion mechanism, features at different levels are weightedly fused, which can enhance the diversity and accuracy of feature extraction. More specifically, the feature fusion formula is:

[0072]

[0073] Among them, F i is the feature map of the i-th layer, ω i is the feature weight, θ is the nonlinear fusion coefficient;

[0074] The extracted feature vector includes improved color features, texture features, edge features and spatial position features. The color features are represented by RGB color histograms, the texture features are extracted by Gabor filter groups, the edge features are obtained by Canny operator detection, and the spatial position features include relative coordinates and area proportions of regions.

[0075] More specifically, the Gabor filter bank extraction can be expressed as follows:

[0076]

[0077] in,

[0078] x′=xcosθ+ysinθ

[0079] y′=-xsinθ+ycosθ

[0080] Where (x, y) is the pixel coordinate, θ is the direction, f is the frequency, η is the enhancement coefficient, and DoG(x, y) is the Gaussian difference response.

[0081] The extracted feature vector is subjected to L2 normalization to generate a standardized image feature vector.

[0082] S3: Establishing a semantic label generation model, inputting the image feature vector into the semantic label generation model, and generating a scene semantic label for each of the to-be-matched areas;

[0083] Firstly, a dual-branch deep neural network structure is constructed. The dual-branch network consists of a feature enhancement branch and a semantic understanding branch. The feature enhancement branch adopts an improved Transformer encoder structure, and a spatial attention module is added on the basis of the standard Transformer. The spatial attention module includes a position-aware self-attention layer and a feature enhancement layer. The semantic understanding branch adopts a multi-scale feature extraction structure, which includes five feature extraction blocks of different scales. Each feature extraction block is composed of three residual convolution units in series, and the number of channels of the residual convolution units is 64, 128, 256, 512 and 1024 respectively. The feature extraction blocks are connected through a feature pyramid network to achieve effective fusion of multi-scale features.

[0084] More specifically, the feature mapping function of the feature enhancement branch is:

[0085] F enhanced =MSA(F input +PE)+FFN(F attn )

[0086] Among them, MSA is a multi-head self-attention mechanism, PE is a position encoding, FFN is a feedforward neural network, and F enhanced is the enhanced feature output, F input is the input feature, F attn For the features after attention processing, the attention calculation adopts the improved scaled dot product formula:

[0087]

[0088] Among them, M mask is the region correlation mask matrix, W scale is the adaptive scaling weight, Q, K, V are the query, key, and value matrices, d k is the feature dimension.

[0089] The semantic understanding branch adopts a hierarchical feature extraction structure, and the feature aggregation function is defined as:

[0090]

[0091] Among them, α l is the weight coefficient of the first layer, CrossAttn is the cross attention module, F global and F local are global and local features respectively, and F semantic is the semantic feature output, L is the number of feature layers, Conv(F l ) is the convolutional feature of the lth layer.

[0092] Secondly, a feature fusion module is set at the output end of the dual-branch network. The feature fusion module adopts a channel attention mechanism to perform adaptive weighted fusion on the output features of the feature enhancement branch and the semantic understanding branch, as shown in the following formula:

[0093] W fusion =σ(MLP([GAP(F enhanced ); GAP(F semantic )]))

[0094] The fused features are processed by a scene semantic classification head, which contains two fully connected layers and a softmax classification layer. The number of neurons in the first fully connected layer is 512, and the number of neurons in the second fully connected layer is the same as the number of predefined scene semantic categories. That is, the final feature representation is calculated as:

[0095] F final =W fusion ⊙F enhanced +(1-W fusion )⊙F semantic

[0096] Among them, ⊙ is element-by-element multiplication, W fusion is the fusion weight, MLP is the multi-layer perceptron, and GAP is the global average pooling.

[0097] Next, a multi-task loss function is designed to train the network. The multi-task loss function includes three parts: scene classification loss, feature regression loss, and semantic consistency loss. The scene classification loss uses the cross entropy loss function, the feature regression loss uses the mean square error loss function, and the semantic consistency loss is obtained by calculating the KL divergence between the predicted label and the reference label. The three types of losses are weighted and summed according to the preset weights to form the final training target. That is:

[0098] L total =w 1 L cls +w2 L reg +w 3 L consist

[0099] Among them, L total is the total loss, w 1 ,w 2 ,w 3 are the weight coefficients of each loss term, L cls is the classification loss, L reg is the regression loss, L consist is the semantic consistency loss, defined as:

[0100] L consist =|F pred -F target |2+γ·KL(P pred |P target )

[0101] Among them, F pred and F target are the prediction and target features respectively, P pred and P target is the corresponding probability distribution, and γ is the weight coefficient of the KL divergence term.

[0102] Finally, the scene semantic labels output by the network are post-processed and optimized, and the prediction results are refined using the conditional random field model. The conditional random field model considers the spatial adjacency and semantic similarity between the regions to be matched, and obtains the final scene semantic labels through iterative optimization; at the same time, a semantic label consistency verification mechanism is established to ensure that the semantic labels of adjacent regions meet the preset scene layout rules. The post-processing optimization can be expressed as follows:

[0103] S opt =argmaxSP(S|Ffinal)+η 1 ·R(S)

[0104] Among them, R(S) is the spatial regularization term of the semantic label, η 1 is the regularization term weight, P(S|Ffinal) is the label probability given the final feature, which is;

[0105] The potential function of the conditional random field model is defined as:

[0106]

[0107] Among them, ψ u is a univariate potential function, ψ p is the pairwise potential function, s i ,s jare the labels of nodes i and j; optimization is performed based on the region adjacency relationship, and E(S) is the energy function of the scene semantic label.

[0108] By minimizing the energy function E(S), we can obtain a label configuration scheme that is consistent with both local feature prediction and global semantic consistency.

[0109] S4: screening plane elements matching each of the to-be-matched regions from the plane element image library according to the scene semantic label, and calculating a matching score between the plane element and the to-be-matched region;

[0110] First, a semantic label mapping system is established to map the scene semantic labels of the area to be matched to the label system of the plane element image library. The mapping system adopts a multi-level semantic association structure, which includes three levels: main category mapping, subcategory mapping, and attribute mapping. The main category mapping determines the basic material type (such as character class, prop class, building class), the subcategory mapping refines the specific material type (such as the protagonist, NPC, enemy character, etc. in the character role), and the attribute mapping matches the characteristic attributes such as style, size, and color tone.

[0111] Secondly, a multi-dimensional matching scoring network based on deep learning is constructed, which includes semantic relevance scoring, visual coordination scoring, spatial adaptability scoring and scene coherence scoring.

[0112] The semantic relevance score S sem The hierarchical matching degree of the semantic labels is quantified by calculating the cosine similarity of the semantic vectors between the matching area and the candidate plane elements, and the compatibility score of the semantic attributes is obtained, as shown in the following formula:

[0113]

[0114] in, are the semantic feature vectors of the region to be matched and the candidate element respectively, is the semantic feature vector of the kth candidate element, a 1 Global weight coefficient for scoring semantic relevance.

[0115] The visual coordination score S vis The similarity of texture features is quantified by calculating the histogram matching of color distribution and evaluating the adaptability of edge contours, as shown in the following formula:

[0116] S vis =a 2 ·(w c ·S color +w t ·S texture +w e ·Sedge )

[0117] in,

[0118]

[0119] Among them, a 2 is the global weight coefficient for visual coordination scoring, w c ,w t ,w e Respectively represent the weights of color, texture, and edge features, H r ,H e are the color histograms of the area to be matched and the candidate elements, σ c is the standard deviation parameter of color matching, δ is the mutual information influence factor, I r ,I e are the image matrices of the area to be matched and the candidate elements respectively, is the Gabor filter response at direction θ and frequency f, ω θ,f is the Gabor feature weight of different directions and frequencies, σ e is the standard deviation parameter of edge matching, Represent the gradient features of the area to be matched and the candidate element image respectively.

[0120] The spatial fitness score S spa Evaluate the matching degree between the size of the plane element and the area to be matched, calculate the degree of fit of the shape features, analyze the rationality of the spatial position, and evaluate the naturalness of the transition with the adjacent area, as shown in the following formula:

[0121]

[0122] in,

[0123]

[0124] Among them, a 3 is the global weight coefficient of the spatial fitness score, A r ,A e are the areas of the region to be matched and the candidate element, σ a is the standard deviation parameter of area matching, C r ,C e are the contour features of the region and element respectively, D H is the Hausdorff distance metric, σ s is the standard deviation parameter of shape matching, p r ,p e are the location coordinates of the region and element, σ pis the standard deviation parameter of position matching, γ is the overlap factor, Roverlap is the area overlap rate, S shape is the shape matching degree, S pos is the position fitness.

[0125] The scene coherence score S con By calculating the style consistency with the surrounding matched elements, the coordination degree of the overall scene atmosphere is evaluated, as shown in the following formula:

[0126]

[0127] in,

[0128]

[0129] Among them, a 4 is the global weight coefficient of the scene coherence score, N(r) is the neighborhood set of the to-be-matched region r, η 3 is the neighborhood impact factor, is the style feature of element e and its neighboring element j, σ d is the standard deviation parameter of style matching, S comp (e,j) is the element compatibility score.

[0130] Next, the scoring mechanism is integrated to weight the scores of the above multiple dimensions. The weight distribution adopts an adaptive adjustment mechanism to dynamically adjust the weight of each dimension according to the importance, location characteristics and surrounding environment of the area to be matched. For the key areas of the scene, the weights of semantic relevance and visual coordination are increased; for the transition area, the weight of spatial adaptability is appropriately increased:

[0131]

[0132] Where i = 1 to 4, representing the semantic relevance score, visual coordination score, spatial adaptability score, and scene coherence score, respectively. i is the dynamic weight of each scoring dimension.

[0133] Then, a candidate plane element set is established for each area to be matched, and the size of the set is dynamically determined according to the matching score. Multiple rounds of screening are performed:

[0134] The first round is a rough screening based on semantic tags to ensure basic semantic matching;

[0135] The second round is fine screening based on visual features to ensure visual effects;

[0136] The third round is optimized based on spatial relationships to ensure overall coordination;

[0137] The final round of adjustments is based on scene coherence to ensure the overall effect.

[0138] S5: placing the plane element with the highest matching score into the corresponding to-be-matched area to generate a complete game scene design diagram.

[0139] This embodiment also provides a game plane design automatic scene matching system based on deep learning, including:

[0140] The acquisition module is used to obtain the background map of the game scene and multiple plane element image libraries;

[0141] A segmentation module, used for performing image segmentation on the background base map, and dividing the background base map into a plurality of areas to be matched;

[0142] A feature vector extraction module, used to extract the image feature vector of each of the to-be-matched areas;

[0143] A model building module, used to establish a semantic label generation model, input the image feature vector into the semantic label generation model, and generate a scene semantic label for each of the to-be-matched areas;

[0144] A matching module, used to screen plane elements matching each of the to-be-matched regions from the plane element image library according to the scene semantic labels, and calculate a matching score between the plane elements and the to-be-matched regions through a deep convolutional neural network;

[0145] The design drawing generation module is used to place the plane element with the highest matching score into the corresponding area to be matched to generate a complete game scene design drawing.

[0146] This embodiment also provides a computer device, which is suitable for the automatic scene matching method for game plane design based on deep learning, including: a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute computer executable instructions to implement the automatic scene matching method for game plane design based on deep learning as proposed in the above embodiment.

[0147] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.

[0148] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, the method for realizing automatic scene matching of game plane design based on deep learning as proposed in the above embodiment is implemented; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, referred to as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, referred to as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, referred to as EPROM), programmable read-only memory (Programmable Red-Only Memory, referred to as PROM), read-only memory (Read-Only Memory, referred to as ROM), magnetic storage, flash memory, disk or optical disk.

[0149] Example 2

[0150] This is the second embodiment of the present invention. In order to further verify the advancement of the present invention, experimental simulation of the automatic scene matching method for game plane design based on deep learning / comparison data with the prior art are provided.

[0151] In the field of game scene design, traditional manual design methods consume a lot of manpower and time, and it is difficult to ensure the beauty and coherence of the scene. This experiment selects a sci-fi style open world game scene, and verifies the practicality and innovation of the technology through the proposed automatic scene matching method for game graphic design based on deep learning. In the experimental preparation stage, a comprehensive dataset containing 1,000 background images and 5,000 plane elements was first constructed. The background images are derived from sci-fi game scenes, with a unified resolution of 2560×1440 pixels and RGB three-channel format. The plane element library is divided according to preset classification rules, including character classes (300), prop classes (200) and building classes (500). Each plane element is equipped with detailed label information, covering categories, usage scenarios, style attributes and size information.

[0152] In the image segmentation stage, the improved U-Net and SLIC superpixel algorithms are used to perform fine segmentation on the background image. Compared with the traditional SLIC algorithm, this method significantly improves the segmentation performance by introducing an adaptive weight coefficient. The experiment selected 10 representative background images for comparative testing. After using the improved algorithm, the boundary recall rate increased by 15.2%, the over-segmentation rate decreased by 12.8%, and the segmentation accuracy in complex texture areas (such as metal surfaces and neon light areas) increased by 18.5%. In the feature extraction stage, the improved ResNet-50 network introduces a multi-scale feature fusion mechanism to enhance the feature expression capability. Richer color, texture and edge features are extracted through Gabor filters and Canny operators.

[0153] In the semantic understanding stage, the constructed dual-branch deep neural network includes a feature enhancement branch and a semantic understanding branch. The network is trained through a multi-task loss function, which comprehensively considers scene classification loss, feature regression loss, and semantic consistency loss. Finally, the conditional random field model is used to post-process the semantic labels to ensure that the semantic labels of adjacent areas meet the scene layout rules. In the plane element matching stage, a multi-dimensional matching scoring network is designed to conduct a comprehensive evaluation from four dimensions: semantic relevance, visual coordination, spatial adaptability, and scene coherence.

[0154] The following table is an experimental comparison between the solution of the present invention and the traditional solution:

[0155] Table 1 Comparison of experimental data of automatic game scene design methods

[0156] index Traditional methods Improvement methods Improvement rate (%) Remark Boundary recall 0.732 0.884 15.2 Segmentation accuracy Over-segmentation rate 0.245 0.117 12.8 Segmentation quality Texture area segmentation accuracy 0.621 0.806 18.5 Complex scene processing Feature extraction dimension 3 5 66.7 Multi-scale features Semantic label consistency 0.652 0.894 37.1 Scene Coherence Plane element matching accuracy 0.623 0.876 40.6 Element Compatibility Scene design efficiency (hours / photo) 4.5 0.6 86.7 Human cost Overall aesthetic rating (out of 10 points) 6.2 8.7 40.3 Subjective evaluation

[0157] Through in-depth analysis of experimental data, we can clearly see the significant advantages of the method of the present invention over the traditional game scene design method. First, in terms of image segmentation accuracy, the boundary recall rate of the improved method increased by 15.2%, and the over-segmentation rate decreased by 12.8%, which means that the algorithm can more accurately identify and segment complex texture areas in game scenes. In particular, in scenes with rich details such as metal surfaces and neon lights, the segmentation accuracy increased by 18.5%, which is significantly better than traditional methods.

[0158] In terms of feature extraction dimensions, this method expands from the traditional three dimensions to five dimensions, an increase of 66.7%. By introducing a multi-scale feature fusion mechanism, the feature expression is greatly enriched. The semantic label consistency is improved from 0.652 to 0.894, an increase of 37.1%, indicating that the algorithm can better understand the scene semantics and maintain the coherence of the overall layout. The plane element matching accuracy is increased from 0.623 to 0.876, an increase of 40.6%, which means that the adaptability of the automatically matched elements to the scene is significantly enhanced.

[0159] Most importantly, the efficiency of scene design has been greatly reduced from the traditional 4.5 hours / frame to 0.6 hours / frame, an efficiency improvement of up to 86.7%, greatly reducing the labor cost of game art design. The subjective aesthetic score has increased from 6.2 points to 8.7 points, an increase of 40.3%, which fully proves that this method can not only improve design efficiency, but also significantly improve the overall beauty and coherence of the scene. These data fully verify the innovation and practical value of the invented method in the field of automatic game scene design.

[0160] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A method for automatic scene matching in game plane design based on deep learning, characterized by: include: Get the background map of the game scene and multiple graphic element image libraries; Performing image segmentation on the background image to divide the background image into a plurality of to-be-matched regions, and extracting an image feature vector of each to-be-matched region; Establishing a semantic label generation model, inputting the image feature vector into the semantic label generation model, and generating a scene semantic label for each of the to-be-matched areas; According to the scene semantic label, plane elements matching each of the to-be-matched regions are screened from the plane element image library, and a matching score between the plane elements and the to-be-matched regions is calculated; The plane element with the highest matching score is placed in the corresponding area to be matched to generate a complete game scene design diagram.

2. The method for automatic scene matching of game plane design based on deep learning as claimed in claim 1, characterized in that: The plane element image library includes a character material library, a prop material library and a building material library; The character material library includes human characters, biological characters and mechanical characters, and each character material is equipped with images from multiple perspectives; The prop material library includes weapon props, consumable props and decorative props, and the prop material is in transparent background PNG format; The building material library includes building bodies, building decorations and building ancillary facilities, and the building materials are marked with scale information; Each material in the plane element image library contains preset tag information, and the preset tag information includes material category, usage scenario, style attribute and size information.

3. The method for automatic scene matching of game plane design based on deep learning as claimed in claim 2, characterized in that: Extracting the image feature vector of each of the to-be-matched areas includes: Performing preliminary segmentation on the background map to obtain coarse-grained segmentation areas; Refining the coarse-grained segmented area to divide the background map into a plurality of to-be-matched areas of different sizes; Extracting a feature vector for each of the to-be-matched regions; the image feature vector includes color features, texture features, edge features, and spatial position features; The extracted feature vector is normalized to generate a standardized image feature vector.

4. The method for automatic scene matching of game plane design based on deep learning as claimed in claim 3, characterized in that: The coarse-grained segmentation area is refined by using a superpixel segmentation algorithm; the distance measurement formula of the superpixel segmentation algorithm is as follows: Among them, D is the comprehensive distance measurement value, d c is the color distance, d s is the spatial distance, d t is the texture distance, α, β, γ are adaptive weight coefficients, and the calculation method is: in, They are the color gradient, position gradient and texture gradient of the image respectively, λ1, λ2, λ3 are balance factors, σ1 is the scale parameter of the color gradient, σ2 is the scale parameter of the position gradient, and σ3 is the scale parameter of the texture gradient.

5. The method for automatic scene matching of game plane design based on deep learning as claimed in claim 4, characterized in that: The semantic label generation model includes a feature enhancement branch and a semantic understanding branch; The feature enhancement branch adopts an improved Transformer encoder structure and adds a spatial attention module on the basis of the standard Transformer; the spatial attention module includes a position-aware self-attention layer and a feature enhancement layer; The semantic understanding branch adopts a multi-scale feature extraction structure, which includes multiple feature extraction blocks of different scales, and each feature extraction block is composed of multiple residual convolution units connected in series; The feature aggregation function of the semantic understanding branch is expressed as follows: Among them, α l is the weight coefficient of the first layer, CrossAttn is the cross attention module, F global and F local are global and local features respectively, and F semantic is the semantic feature output, L is the number of feature layers, Conv(F l ) is the convolutional feature of the lth layer.

6. The method for automatic scene matching of game plane design based on deep learning as claimed in claim 5, characterized in that: Screening the plane elements matching each of the to-be-matched regions from the plane element image library adopts a multi-level semantic association structure, including three levels: main category mapping, subcategory mapping, and attribute mapping; The main category mapping determines the basic material type; the subcategory mapping refines the specific material type; the attribute mapping matches style, size, and tone; The matching scores include a semantic relevance score, a visual coordination score, a spatial adaptability score, and a scene coherence score.

7. The method for automatic scene matching of game plane design based on deep learning as claimed in claim 6, characterized in that: The semantic relevance score quantifies the hierarchical matching degree of the semantic labels by calculating the cosine similarity of the semantic vectors of the to-be-matched area and the candidate plane elements, and obtains the compatibility score of the semantic attributes; The visual harmony score quantifies the similarity of texture features by calculating the histogram matching of color distribution; The spatial fitness score evaluates the matching degree between the size of the plane element and the area of ​​the area to be matched, analyzes the rationality of the spatial position by calculating the degree of fit of the shape features, and evaluates the naturalness of the transition with the adjacent area; The scene coherence score evaluates the coordination of the overall scene atmosphere by calculating the style consistency with the surrounding matched elements. The semantic relevance score, the visual coordination score, the spatial adaptability score and the scene coherence score are weightedly fused to obtain a matching score between the plane element and the area to be matched.

8. A deep learning-based automatic scene matching system for game graphic design, based on the deep learning-based automatic scene matching method for game graphic design according to any one of claims 1 to 7, characterized in that: include: The acquisition module is used to obtain the background map of the game scene and multiple plane element image libraries; A segmentation module, used for performing image segmentation on the background base map, and dividing the background base map into a plurality of areas to be matched; A feature vector extraction module, used to extract the image feature vector of each of the to-be-matched areas; A model building module, used to establish a semantic label generation model, input the image feature vector into the semantic label generation model, and generate a scene semantic label for each of the to-be-matched areas; A matching module, used to screen plane elements matching each of the to-be-matched regions from the plane element image library according to the scene semantic labels, and calculate a matching score between the plane elements and the to-be-matched regions through a deep convolutional neural network; The design drawing generation module is used to place the plane element with the highest matching score into the corresponding area to be matched to generate a complete game scene design drawing.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for automatic scene matching of game plane design based on deep learning according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the automatic scene matching method for game plane design based on deep learning are implemented as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Intelligent evaluation method and system for graphic design

    CN120654406A

  • Game world generation method and system, electronic equipment and storage medium

    CN121623298A