Game resource classification method and device, electronic equipment and storage medium

By preprocessing and encoding the multimodal data of game resources, a structured embedding vector is constructed, which solves the problems of low efficiency and insufficient accuracy in the existing technology of game resource classification, and realizes efficient and diversified game resource classification.

CN120849999APending Publication Date: 2025-10-28NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510912108.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

In existing technologies, the classification of game resources is inefficient and lacks accuracy and richness. In particular, when faced with a massive amount of UGC game resources to be classified, manual classification is inefficient, while text analysis suffers from insufficient accuracy and richness due to the single data format and inconsistent data quality.

Method used

By preprocessing multimodal data of game resources and constructing structured embedding vectors, and encoding them through a game resource classification model, the semantic information of spatial coordinates and resource attributes is fused to achieve efficient classification of game resources.

Benefits of technology

It improved the accuracy and efficiency of game resource classification, reduced labor costs, and enhanced the diversity and accuracy of classification through the fusion of multimodal data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849999A_ABST
    Figure CN120849999A_ABST
Patent Text Reader

Abstract

The invention discloses a game resource classification method and device, electronic equipment and a storage medium, and relates to the technical field of games.The method comprises the steps that firstly, multi-modal data of game resources is preprocessed to obtain a preprocessed data set, then the preprocessed data set is input into a game resource classification model for encoding processing, and a feature vector set is obtained; finally, the game resources are classified according to the feature vector set, the classification result of the game resources is determined, the multi-modal data at least comprise analysis files of the game resources, and the feature vector set comprises structured embedded vectors obtained after encoding preprocessing is conducted on the analysis files. And the structured embedded vector represents mapping semantic information of space coordinates and resource attributes in the game resources. Therefore, the game resource classification efficiency, accuracy and richness can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of game technology, and specifically to a method, apparatus, electronic device, and storage medium for classifying game resources. Background Art

[0002] Currently, all types of games encourage user-generated content. During the screening stage, user-uploaded UGC game resources need to be categorized, mainly through manual categorization and automatic categorization after analyzing the text of the game resources.

[0003] However, when faced with a massive amount of game resources to be categorized, manual categorization is inefficient, while text analysis suffers from insufficient categorization accuracy and richness due to the limited data format and inconsistent data quality. Summary of the Invention

[0004] This application provides a method, apparatus, electronic device, and storage medium for classifying game resources.

[0005] In a first aspect, embodiments of this application provide a method for classifying game resources, the method comprising:

[0006] The multimodal data of game resources is preprocessed to obtain a preprocessed dataset;

[0007] The preprocessed dataset is input into the game resource classification model for encoding processing to obtain a set of feature vectors.

[0008] The game resources are classified according to the set of feature vectors to determine the classification result of the game resources;

[0009] The multimodal data includes at least the parsed file of the game resource, and the feature vector set includes structured embedding vectors after encoding and preprocessing the parsed file. The structured embedding vectors represent the semantic information of the mapping between spatial coordinates and resource attributes in the game resource.

[0010] Secondly, embodiments of this application provide a game resource classification device, the device comprising:

[0011] The data processing module is used to preprocess the multimodal data of game resources to obtain a preprocessed dataset;

[0012] The data encoding module is used to input the preprocessed dataset into the game resource classification model for encoding processing to obtain a set of feature vectors;

[0013] The resource classification module is used to classify the game resources according to the feature vector set and determine the classification result of the game resources;

[0014] The multimodal data includes at least the parsed file of the game resource, and the feature vector set includes structured embedding vectors after encoding and preprocessing the parsed file. The structured embedding vectors represent the semantic information of the mapping between spatial coordinates and resource attributes in the game resource.

[0015] Thirdly, embodiments of this application also provide an electronic device, including a memory storing multiple instructions; a processor loading instructions from the memory to execute the steps of any of the game resource classification methods provided in embodiments of this application.

[0016] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute the steps of any of the game resource classification methods provided in embodiments of this application.

[0017] Fifthly, embodiments of this application also provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement the steps in any of the game resource classification methods provided in embodiments of this application.

[0018] The scheme adopted in this application embodiment can first preprocess the multimodal data of game resources to obtain a preprocessed dataset, then input the preprocessed dataset into a game resource classification model for encoding processing to obtain a set of feature vectors, and finally classify the game resources according to the set of feature vectors to determine the classification result. The multimodal data includes at least the parsed files of the game resources, and the set of feature vectors includes structured embedding vectors after encoding and preprocessing the parsed files. The structured embedding vectors represent the semantic information of the mapping between spatial coordinates and resource attributes in the game resources. Therefore, this application can extract the semantic information of the mapping between spatial coordinates and resource attributes from the parsed files of game resources, and integrate this information with other types of modal data to classify the game resources. This allows for the improvement of the accuracy, efficiency, and diversity of game resource classification based on the rich feature information, while reducing classification time and manual costs. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is an application scenario diagram of the game resource classification method provided in the embodiments of this application;

[0021] Figure 2 This is a flowchart illustrating a method for classifying game resources provided in an embodiment of this application;

[0022] Figure 3 This is another flowchart illustrating the method for classifying game resources provided in the embodiments of this application;

[0023] Figure 4 This is a schematic diagram illustrating the classification results of multiple game resources provided in an embodiment of this application;

[0024] Figure 5 This is a schematic diagram of the structure of the game resource classification device provided in the embodiments of this application;

[0025] Figure 6 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0026] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0027] Before providing a detailed explanation of the embodiments of this application, some terms involved in the embodiments of this application will be explained.

[0028] In the description of the embodiments of this application, the terms "first," "second," etc., may be used herein to describe various concepts, but unless specifically stated otherwise, these concepts are not limited by these terms. These terms are used only to distinguish one concept from another. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices.

[0029] It should be noted that the information (including but not limited to user input information, such as information entered into input boxes), data (including but not limited to data used for analysis, stored data, and displayed data, such as context code, all code of the current project, service pressure corresponding to operations performed on all code of the current project, and code development status of the current project), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with relevant laws, regulations, and standards. For example, the context code, operations performed on all code of the current project, the corresponding service pressure, and code development status involved in this application were all obtained with full authorization.

[0030] This application provides a method, apparatus, electronic device, and computer-readable storage medium for distributing game resources. Specifically, the game resource distribution method of this application can be executed by an electronic device, which can be a terminal or a server. The terminal can be a smartphone, tablet, laptop, touch screen, game console, personal computer (PC), personal digital assistant (PDA), or other terminal device. The terminal can also include a client, which can be a game application client, a browser client carrying a game program, or an instant messaging client. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0031] For example, such as Figure 1 As shown, the electronic device is illustrated using terminal 10 as an example. Terminal 10 can first preprocess the multimodal data of game resources to obtain a preprocessed dataset, then input the preprocessed dataset into the game resource classification model for encoding processing to obtain a set of feature vectors, and finally classify the game resources according to the set of feature vectors to determine the classification result of the game resources. The multimodal data includes at least the parsing file of the game resources, and the set of feature vectors includes the structured embedding vectors after encoding and preprocessing the parsing file. The structured embedding vectors represent the semantic information of the mapping between spatial coordinates and resource attributes in the game resources.

[0032] The following is a detailed description in conjunction with the accompanying drawings. It should be noted that the order of description of the following embodiments is not intended to limit the preferred order of the embodiments. Although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be performed in a different order than that shown in the drawings.

[0033] Please refer to Figure 2 The specific process for classifying game resources can be summarized in steps 201 to 203, where:

[0034] Step 201: Preprocess the multimodal data of the game resources to obtain a preprocessed dataset;

[0035] In the context of this application, game resources can be UGC (User-Generated Content) content such as game maps, game scene models, and game 3D animations.

[0036] Multimodal data can include at least parsed files of game resources. These parsed files are the core files of game resources, used to record information such as various components, objects, and their attributes. Parsed files are structured data, typically stored in JSON, XML, or platform-specific formats.

[0037] In practice, the parsed file can include the spatial coordinates of game resources, resource attributes, and logical connections. For example, spatial coordinates can be the three-dimensional coordinates (x, y, z) of each resource component (such as platforms, mechanisms, portals, etc.), and can also include spatial transformation parameters such as rotation angles and scaling ratios. Resource attributes can include the component's category (such as "bouncing mechanism," "portal," "obstacle," etc.), the component's behavioral parameters (such as triggering method, cooldown time, response conditions), and decorative or logical markers (such as interactive, dynamic movement, etc.). Logical connections can include the connection or triggering relationship between different resource components (such as "switch A controls door B"), level completion paths, checkpoint settings, etc.

[0038] In some embodiments, such as Figure 3 As shown, preprocessing can be performed on the multimodal data involved in game resources. Specifically, multimodal data from various modalities, such as initial classification, descriptive text, parsed files, background music, and screenshots, contained in UGC game resources can be collected and parsed. A unified data cleaning and filtering operation can then be performed on this multimodal data to remove samples with abnormal formats, semantic redundancy, or severe missing data. This results in a structurally complete and quality-controlled basic dataset, providing reliable input for subsequent model training and label generation.

[0039] In the data preprocessing stage, to improve the uniformity and scalability of the generated classification results, a multi-source semantic network can be further introduced, including synonyms, near-synonyms, and hyponyms / hypernyms from a general lexicon, as well as knowledge graph information related to the game domain. Based on this, multi-strategy semantic matching rules are constructed, comprehensively employing semantic normalization methods such as precise keyword matching, fuzzy mapping of synonyms and near-synonyms, and hierarchical reasoning of hyponyms / hypernyms to aggregate the vocabulary of historical game resource classification results and user-defined classification results. Through the above semantic normalization and mapping process, the problems of inconsistent label expressions and granularity can be effectively solved, ultimately constructing a standardized and structured basic label lexicon to support the generation of diverse classification results.

[0040] Step 202: Input the preprocessed dataset into the game resource classification model for encoding processing to obtain a set of feature vectors.

[0041] The game resource classification model is used to analyze the input preprocessed dataset and output the classification result of the game resource. For example, the classification result can be the category label of the gameplay corresponding to the game resource, such as "adventure", "story", "racing" etc.

[0042] Specifically, the game resource classification model can first perform structured embedding encoding on the modal data in the preprocessed dataset, transforming the data of each modality into a corresponding vector. It can be understood that different vectors carry information about different modalities, making the processed vectors usable for subsequent model learning and prediction.

[0043] The feature vector set includes structured embedding vectors after encoding and preprocessing the parsed files. The structured embedding vectors represent the semantic information of the mapping between spatial coordinates and resource attributes in game resources, thereby effectively expressing the spatial layout features and functional logic of the map structure in game resources, and providing key structural semantic support for the generation of multimodal and diversified classification results.

[0044] In some embodiments, step 202 may include:

[0045] Extract structured features from the parsed game resource files;

[0046] The structured features are encoded to obtain an intermediate vector;

[0047] The intermediate vector is fused with attention information from each resource component to obtain a structured embedding vector.

[0048] The structured features can include the resource attributes of multiple resource components and the spatial coordinates of each resource component in the game map. For example, an exemplary structured feature could be that the resource component corresponding to the spatial coordinates (1, 1, 1) is a tree, and the resource attributes are information such as the tree's color, height, width, and touch events.

[0049] Specifically, firstly, the structured feature data x can be analyzed. struct Encoding processing is performed using the structural encoding layer or sub-model in the game resource classification model, such as using BERT to extract the intermediate representation vector h1 = BERT(x struct Subsequently, an attention interaction mechanism between various map resource components can be introduced based on the intermediate vector to fuse their spatial and semantic association information, ultimately generating a structured embedding vector h. struct High-level semantic features used to characterize map structure information in game resources.

[0050] In some embodiments, the structured embedding vector can be obtained through the following steps:

[0051] The structured features are subjected to rotational position encoding to obtain a position embedding matrix that represents the relative positional relationships between the resource components in the game resources;

[0052] Based on the location embedding matrix, determine the attention index parameters that characterize the attention information between each resource component;

[0053] The attention index parameter is fused with the value vectors of other types of modal data to obtain fused information, and the intermediate vector is updated based on the fused information to obtain the structured embedding vector;

[0054] Other types of modal data are at least one type of modal data other than parsed files, such as text, images, and audio of game resources.

[0055] In specific implementation, RoPE-3D (Rotary Position Embedding-3D) can first be used to structure the feature x. struct The spatial coordinate information (x, y, z) of each resource component is encoded using 3D rotation to generate a position embedding matrix R(x, y, z). This position embedding matrix encodes the spatial orientation and relative positional relationships of the components in a block-diagonal manner to enhance the model's spatial awareness.

[0056] Specifically, the position embedding matrix can be represented as:

[0057]

[0058] Among them, R x R y Rz These are two-dimensional rotation matrices along the x, y, and z directions, used to transform the component's position coordinates in these directions into rotational transformations in vector space. The three are combined in a block-diagonal form to introduce rotational position representations when calculating attention scores.

[0059] Next, based on the generated positional encoding, the structural modality can be introduced as the query, and the embedding vectors of other types of modal data (such as text, images, or audio) can be used as the key. The attention metric parameters between the structural modality and other modalities can then be calculated based on the following formula:

[0060]

[0061] Among them, score m,n It is an attention metric parameter that characterizes the semantic information of the relative spatial location between resource components m and n, q m It is the query vector of the m-th component in the structural modality, k n It is the nth eigenvector in other modalities, and d is the dimension of the eigenvector.

[0062] Finally, the above attention metric parameter score can be used to... m,n Multiplying the value vector V of the corresponding modality achieves attention-weighted fusion, yielding cross-modal context fusion information:

[0063] h att =∑ n score m,n ·V n ;

[0064] Then the fused information is fed back to the structural modal intermediate vector h1 = BERT(x struct This process is repeated to obtain the final structured embedding vector h. struct .

[0065] Through the above methods, this application can achieve spatial-semantic fusion between structured data and other modalities, effectively model the relative positional relationships and multimodal semantic complementarity between map components, and further improve the accuracy of label generation and content awareness.

[0066] In some embodiments, the multimodal data may further include at least one of text, images, and audio from game resources, and step 202 may further include:

[0067] The text of game resources is segmented and sorted to obtain a text sequence, which is then encoded through the bidirectional semantic coding layer of the game resource classification model to obtain a text embedding vector.

[0068] The game resource images are segmented and normalized to obtain multiple image patches, and the multiple image patches are encoded through the visual transformation layer of the game resource classification model to obtain image embedding vectors;

[0069] The audio of the game resources is preprocessed to obtain Mel spectrum information, and then encoded through the speech pre-training layer of the game resource classification model to obtain audio embedding vectors. The feature vector set includes text embedding vectors, image embedding vectors, and audio embedding vectors.

[0070] In practice, feature extraction processing can be performed on multimodal data in UGC game resources to construct a set of feature vectors in a unified semantic space, providing basic support for subsequent multimodal semantic fusion and tag generation. The multimodal data includes, but is not limited to, text descriptions, map screenshots, and background audio, corresponding to text modality, image modality, and audio modality, respectively.

[0071] Specifically, text information in game resources (such as map names and descriptions) can be segmented into words or sub-word units, and special markers (such as [CLS], [SEP]) can be added to convert it into an input sequence that meets the requirements of a pre-trained language model. Subsequently, this sequence is processed using a bidirectional semantic coding structure (such as the BERT model), capturing the bidirectional contextual relationships between words in the text through multiple Transformer layers, and finally outputting a fixed-dimensional text embedding vector, denoted as:

[0072]

[0073] Specifically, map screenshots can be uniformly adjusted to 224×224 pixels and divided into 16×16 image patches, resulting in 196 image patches. After normalization, these image patches are input into a visual encoding model (e.g., the ViT-base model). Each image patch is linearly mapped to a 768-dimensional embedding, and this is combined with positional encoding to construct the input sequence. The model learns the spatial relationships between image patches through a self-attention mechanism and outputs an image embedding vector, denoted as:

[0074]

[0075] Specifically, the original audio data can be pre-emphasized and segmented into frames (25ms window) to extract the Mel spectrogram after short-time Fourier transform, and the energy of a 128-dimensional Mel filter bank within the 0–8kHz frequency band can be selected as input features. The spectrogram is then input into a pre-trained speech model (e.g., wav2vec2.0). The model first extracts local acoustic features through convolutional layers, then models long-term contextual relationships through a multi-layer Transformer structure, and finally obtains a fixed-dimensional audio embedding vector through average pooling, denoted as:

[0076]

[0077] In summary, a unified set of feature vectors containing text, image, and audio modalities can ultimately be obtained:

[0078] H = {h} text ,h image ,h audio};

[0079] As can be seen from the above, the feature vector set of this application will be used as the input of the subsequent cross-modal semantic alignment and fusion module to construct a multimodal semantic representation space and realize fine-grained label generation and type recognition of game resources.

[0080] Step 203: Classify the game resources according to the feature vector set and determine the classification result of the game resources.

[0081] The multimodal data includes at least the parsed files of game resources, and the feature vector set includes structured embedding vectors after encoding and preprocessing the parsed files. The structured embedding vectors represent the semantic information of the mapping between spatial coordinates and resource attributes in the game resources.

[0082] In practical implementation, target game resources can be classified and identified based on a multimodal feature vector set to determine their resource category. The multimodal feature vector set can include, but is not limited to, embedding vectors extracted from multiple sources such as text, images, audio, and map parsing files. Among these, the parsing file, as a key data source for the structural modality, contains the spatial coordinates and semantic attribute information of multiple map components. The system performs encoding preprocessing on the parsing file, first extracting the structured features of each component and constructing a spatial relationship representation using the Rotational Position Encoding (RoPE-3D) method. Then, a structural encoding network is used to semantically encode these features, resulting in structured embedding vectors representing the map's structural layout and functional logic.

[0083] This structured embedding vector, used as one of the inputs in conjunction with other modal features, is used for game resource tag generation and content classification tasks. It can effectively reflect the structural semantics and spatial features in map design, and improve the system's understanding of UGC game map resources and the accuracy of tag recommendation.

[0084] In some embodiments, step 203 may include:

[0085] The alignment layer of the game resource classification model performs semantic alignment on the feature vector set to obtain the aligned feature vector group.

[0086] The latent vector is obtained by jointly encoding each feature vector in the feature vector group.

[0087] The game resources are classified based on the latent vectors to obtain the classification results.

[0088] In practice, the alignment layer in the game resource classification model can be used to semantically align the features of each modality in the multimodal feature vector set, thus constructing an alignment feature vector group in a unified semantic space. For example, an exemplary alignment feature vector group can be represented as "{game resource ID is 1, text embedding vector, image embedding vector, audio embedding vector, structured embedding vector}".

[0089] Next, the modal features in this feature vector group can be jointly encoded, fusing their semantic and structural information to generate a unified latent vector representation. Finally, the target game resource can be classified based on this latent vector, outputting its classification result, such as the type of game resource, gameplay style, or functional tag.

[0090] In some embodiments, the training process for the alignment layer may include:

[0091] Multiple first feature vectors from the same game resource are combined to obtain a positive sample feature group;

[0092] Multiple second feature vectors from different game resources are combined to obtain negative sample feature groups;

[0093] Construct a loss function that represents the ability of the alignment layer to distinguish between positive and negative sample feature groups;

[0094] The initial alignment layer is trained iteratively until the loss function converges, resulting in the trained alignment layer.

[0095] In practice, this approach can optimize the alignment layer's ability to distinguish between semantically similar and dissimilar content by constructing positive and negative sample feature groups, thereby projecting multimodal data into a unified semantic space.

[0096] Multimodal feature vectors extracted from the same game resource (such as the same UGC map), such as the structure vector h. struct Text vector h text Image vector h image Audio vector h audio These can be combined into a homologous feature group, which can be used as a positive sample pair and should have a high similarity in the semantic space.

[0097] Modal feature vectors can also be sampled from different game resources and combined to form heterogeneous feature sets, which serve as negative sample pairs. These features come from semantically unrelated map resources and should maintain relatively low semantic similarity after alignment.

[0098] Specifically, the objective function of the alignment layer can be defined as a contrastive loss function, used to maximize the similarity between positive sample pairs while minimizing the similarity between positive and negative samples. Specifically, a symmetric InfoNCE loss based on cosine similarity is used, as shown in the following formula:

[0099]

[0100] Among them, h m ,h n represents the feature vectors of different modalities within the same game resource; N represents the set of negative samples, i.e., the modal vectors in different game resources; cos(h m ,h n ) represents the cosine similarity between two vectors; τ∈(0,1] is an adjustable temperature parameter that controls the discrimination of the sample distribution; M is the number of positive samples involved in the calculation.

[0101] Finally, positive and negative sample groups can be input into the alignment network, trained using the aforementioned loss function, and the parameters of the alignment layer can be iteratively updated using the gradient descent algorithm until the loss function converges, resulting in a converged multimodal alignment layer model. This model can map inputs from different modalities to a unified semantic space, making subsequent multimodal fusion and label generation more accurate.

[0102] In some embodiments, the latent vector can be obtained through the following steps:

[0103] The feature vector group is input into the joint encoding layer of the game resource classification model for processing to obtain the first fusion vector of the structure and image of the game resource, the second fusion vector of the structure and text, and the third fusion vector of the structure and audio.

[0104] The first fusion vector, the second fusion vector, and the third fusion vector are concatenated to obtain the latent vector.

[0105] To achieve multimodal semantic comprehensive modeling of UGC game resources, the aforementioned aligned multimodal feature vector group can be input into the joint coding layer for further processing. The joint coding layer models the semantic relationship between structural modalities and other modalities based on the cross-attention mechanism, and constructs cross-modal fusion representations respectively.

[0106] Specifically, the structural modality embedding vector can be used as the query vector, and cross-attention fusion can be performed with the feature vectors of the image modality, text modality, and audio modality respectively to generate:

[0107] The first fusion vector of the structure and image modalities is denoted as h. (struct,image) ;

[0108] The second fusion vector of structure and text modality is denoted as h. (struct,text) ;

[0109] The third fusion vector of structure and audio modality, denoted as h (struct,wave) .

[0110] The three fusion vectors mentioned above semantically capture the complementary features between structural information and other modalities, and can effectively express the potential relationship between map component layout and visual appearance, language description, and sound effects atmosphere.

[0111] To unify multimodal information, the system concatenates the three fusion vectors to form the final multimodal fusion representation latent vector h, calculated as follows:

[0112] h = Concat(h) (struct,image) ,h (struct,wave) ,h (struct,text) );

[0113] The hidden vector It integrates multimodal information from structure, image, text, and audio, providing a unified and rich semantic representation for subsequent label prediction and resource classification.

[0114] In some embodiments, the fusion vectors can be obtained through the following steps:

[0115] Using the structured embedding vector as the query vector and the image embedding vector, text embedding vector, and audio embedding vector as key-value pairs, calculate the first attention weight between structure and image, the second attention weight between structure and text, and the third attention weight between structure and audio.

[0116] According to the cross-attention mechanism, the first attention weight and the value vector of the image modality are processed to obtain the first cross-attention result, the second attention weight and the value vector of the text modality are processed to obtain the second cross-attention result, and the first attention weight and the value vector of the audio modality are processed to obtain the third cross-attention result.

[0117] A multi-head attention mechanism is used to fuse the features of the first cross-attention result, the second cross-attention result, and the third cross-attention result, respectively, to obtain the first fusion vector, the second fusion vector, and the third fusion vector.

[0118] Specifically, to achieve deep semantic fusion between structural modalities and other types of modalities (including images, text, and audio), the system adopts a cross-attention mechanism driven by structural embedding vectors and combines it with a multi-head attention structure to complete feature fusion.

[0119] Specifically, the system first uses structured embedding vectors As the query vector, and respectively with the image embedding vector K image Vimage Text embedding vector K text V text and audio embedding vector K wave V wave As key-value pairs, attention channels are constructed corresponding to the three modalities.

[0120] Attention weights between the structural modality and the image, text, and audio modalities can be calculated separately, specifically using a scaled dot product attention mechanism, the calculation method of which is as follows:

[0121]

[0122] Subsequently, the attention weights are applied to the value vectors of the corresponding modalities, and cross-attention calculation is performed to obtain the cross-attention output results for the three modalities:

[0123] CrossAttn(Q,K,V)=Attention(Q,K)·V;

[0124] To enhance the model's expressive power, a multi-head attention mechanism is further employed for feature fusion in each modality. Specifically, multiple sets of projection parameters are used to linearly map the query, key, and value vectors, the outputs of multiple attention heads are calculated, concatenated, and multiplied by the output weight matrix to obtain the final fused representation.

[0125] h (struct,modality) =Concat(head1,…,head) h W d ;

[0126] The calculation form for each attention head is as follows:

[0127]

[0128] Following the above method, the system calculates the following sequentially:

[0129] The first fusion vector h of structure and image modalities (struct,image) ;

[0130] The second fusion vector h of structure and text modality (struct,text) ;

[0131] The third fusion vector h of structure and audio modality (struct,wave) .

[0132] The three fused vectors can be further concatenated to construct a unified multimodal latent vector representation for use by downstream tag generation and resource classification modules.

[0133] In some embodiments, step 203 may further include:

[0134] The latent vectors are linearly transformed by the fully connected layer of the game resource classification model to obtain the original vectors corresponding to each category.

[0135] Activation functions are used to process multiple original vectors to obtain the probability distribution of game resources belonging to multiple defined categories, which serves as the classification result of the game resources.

[0136] In practice, to enable label prediction and game resource classification based on the fused multimodal features, the system adopts a classification prediction mechanism based on fully connected layers and normalized activation functions.

[0137] Specifically, the latent space feature vector output by the multimodal fusion module can be... The input is fed into the fully connected layer of the game resource classification model, where it undergoes a linear transformation to obtain the original predicted value (logits vector) corresponding to each preset label category. This process is achieved through matrix multiplication and bias addition, and the calculation formula is as follows:

[0138] z = W T h+b;

[0139] in, This is the weight matrix of the fully connected layer. As a bias term, C represents the total number of label categories. This is the original output vector, where each component z i This represents the response strength of the latent vector in the i-th label dimension.

[0140] To transform the raw output into an interpretable probability distribution, the system introduces a temperature-adjustable softma x activation function for normalization, generating the predicted probabilities of game resources under each category label. The calculation method is as follows:

[0141]

[0142] Here, τ∈(0.5,1.5) is a temperature parameter used to adjust the smoothness of the output distribution. Smaller temperature values ​​make the output more "sharp," emphasizing high-confidence labels; larger temperature values ​​make the output more uniform, suitable for multi-label recommendation scenarios.

[0143] Ultimately, the probability distribution vector that can be generated is P(y) = {P(y1), P(y2), ..., P(y...}}. C As a classification prediction result for game resources, the system can output one or more tags based on the highest probability tag or a set threshold strategy for tag recommendation, content classification, or map classification tasks.

[0144] In some embodiments, the method of this application may further include:

[0145] The candidate categories from the game resource categorization results are sent to the user, and the user's feedback on the candidate categories is obtained; the feedback action includes acceptance, replacement, or discard.

[0146] The candidate categories with feedback operations of "adoption" are used as the target categories and synonym expansion is performed to obtain the extended categories. The candidate categories with feedback operations of "discard" and "replace" are used as error categories and error analysis is performed to obtain the analysis results.

[0147] The game resource classification model is optimized based on the results of extended classification and error analysis, resulting in an optimized game resource classification model.

[0148] like Figure 4 As shown, after categorizing multiple game resources, the categorization results can be displayed. For example... Figure 4 The document provides game results corresponding to N game resources. For example, the classification results can include adventure, puzzle, parkour, and so on.

[0149] In practice, to further improve the prediction accuracy and system adaptability of the game resource classification model, the system designs a model optimization mechanism based on user feedback. By collecting user feedback on the label results, the model can achieve closed-loop updates and incremental learning.

[0150] Specifically, the system pushes the candidate classification results generated by the game resource classification model—for example, the top four labels with the highest probability values ​​in the probability distribution—to the user interface, highlighting them in the map editing area. It also provides operation components to support user feedback on each candidate classification. Feedback operations can include acceptance, replacement, and discard. Acceptance means confirming the classification as a valid label; replacement indicates the classification is inaccurate and a substitute label is selected; discarding explicitly marks the classification as invalid or irrelevant. The system can record these user feedback behaviors and use them to conduct subsequent model optimization processes.

[0151] For the adopted candidate classifications, the system treats them as positive samples and further performs semantic expansion operations, including synonym replacement and hyponym summarization, to generate expanded labels to enrich the training sample space. These expanded labels will serve as reinforcement learning signals to support the improvement of the label prediction model's generalization ability.

[0152] For candidate categories that are replaced or discarded, the system treats them as error samples, analyzes their potential semantic bias or feature interference during model prediction, and performs feature attribution and dynamic adjustment based on the correlation between the error category and the original input. Based on this analysis, the system uses strategies such as local gradient fine-tuning, sample weighting, or adversarial training to incrementally optimize the model parameters, thereby obtaining an updated optimized model.

[0153] Through the aforementioned feedback recording and adaptive update mechanism, the system can continuously learn and improve the performance of the game resource classification model, effectively enhancing its tag recommendation quality and user acceptance in practical applications.

[0154] This embodiment also provides a game resource classification device, which can be integrated into a terminal device. For example, such as... Figure 5 As shown, the game resource classification device may include:

[0155] Data processing module 301 is used to preprocess the multimodal data of game resources to obtain a preprocessed dataset;

[0156] Data encoding module 302 is used to input the preprocessed dataset into the game resource classification model for encoding processing to obtain a set of feature vectors;

[0157] The resource classification module 303 is used to classify game resources based on the feature vector set and determine the classification result of the game resources.

[0158] The multimodal data includes at least the parsed files of game resources, and the feature vector set includes structured embedding vectors after encoding and preprocessing the parsed files. The structured embedding vectors represent the semantic information of the mapping between spatial coordinates and resource attributes in the game resources.

[0159] In some embodiments, the data encoding module 302 can also be used for:

[0160] Based on the parsed files of the game resources, structured features are extracted; the structured features include the resource attributes of multiple resource components and the spatial coordinates of each resource component in the game map;

[0161] The structured features are encoded to obtain an intermediate vector;

[0162] The intermediate vector is fused with attention information from each resource component to obtain a structured embedding vector.

[0163] In some embodiments, the data encoding module 302 can also be used for:

[0164] The structured features are subjected to rotational position encoding to obtain a position embedding matrix that represents the relative positional relationships between the resource components in the game resources;

[0165] Based on the location embedding matrix, determine the attention index parameters that characterize the attention information between each resource component;

[0166] The attention index parameter is fused with the value vectors of other types of modal data to obtain fused information, and the intermediate vector is updated based on the fused information to obtain the structured embedding vector;

[0167] Other types of modal data refer to at least one type of modal data other than the parsed file.

[0168] In some embodiments, the multimodal data further includes at least one of text, images, and audio from game resources; the data encoding module 302 can also be used for:

[0169] The text of game resources is segmented and sorted to obtain a text sequence, which is then encoded through the bidirectional semantic coding layer of the game resource classification model to obtain a text embedding vector.

[0170] The game resource images are segmented and normalized to obtain multiple image patches, and the multiple image patches are encoded through the visual transformation layer of the game resource classification model to obtain image embedding vectors;

[0171] The audio of the game resources is preprocessed to obtain Mel spectrum information, and then encoded through the speech pre-training layer of the game resource classification model to obtain audio embedding vectors. The feature vector set includes text embedding vectors, image embedding vectors, and audio embedding vectors.

[0172] In some embodiments, the resource classification module 303 can also be used for:

[0173] The alignment layer of the game resource classification model performs semantic alignment on the feature vector set to obtain the aligned feature vector group.

[0174] The latent vector is obtained by jointly encoding each feature vector in the feature vector group.

[0175] The game resources are classified based on the latent vectors to obtain the classification results.

[0176] In some embodiments, the training process of the alignment layer includes:

[0177] Multiple first feature vectors from the same game resource are combined to obtain a positive sample feature group;

[0178] Multiple second feature vectors from different game resources are combined to obtain negative sample feature groups;

[0179] Construct a loss function that represents the ability of the alignment layer to distinguish between positive and negative sample feature groups;

[0180] The initial alignment layer is trained iteratively until the loss function converges, resulting in the trained alignment layer.

[0181] In some embodiments, the resource classification module 303 can also be used for:

[0182] The feature vector group is input into the joint encoding layer of the game resource classification model for processing to obtain the first fusion vector of the structure and image of the game resource, the second fusion vector of the structure and text, and the third fusion vector of the structure and audio.

[0183] The first fusion vector, the second fusion vector, and the third fusion vector are concatenated to obtain the latent vector.

[0184] In some embodiments, the resource classification module 303 can also be used for:

[0185] Using the structured embedding vector as the query vector and the image embedding vector, text embedding vector, and audio embedding vector as key-value pairs, calculate the first attention weight between structure and image, the second attention weight between structure and text, and the third attention weight between structure and audio.

[0186] According to the cross-attention mechanism, the first attention weight and the value vector of the image modality are processed to obtain the first cross-attention result, the second attention weight and the value vector of the text modality are processed to obtain the second cross-attention result, and the first attention weight and the value vector of the audio modality are processed to obtain the third cross-attention result.

[0187] A multi-head attention mechanism is used to fuse the features of the first cross-attention result, the second cross-attention result, and the third cross-attention result, respectively, to obtain the first fusion vector, the second fusion vector, and the third fusion vector.

[0188] In some embodiments, the resource classification module 303 can also be used for:

[0189] The latent vectors are linearly transformed by the fully connected layer of the game resource classification model to obtain the original vectors corresponding to each category.

[0190] Activation functions are used to process multiple original vectors to obtain the probability distribution of game resources belonging to multiple defined categories, which serves as the classification result of the game resources.

[0191] In some embodiments, the game resource classification device of this application can also be used for:

[0192] The candidate categories from the game resource categorization results are sent to the user, and the user's feedback on the candidate categories is obtained; the feedback action includes acceptance, replacement, or discard.

[0193] The candidate categories with feedback operations of "adoption" are used as the target categories and synonym expansion is performed to obtain the extended categories. The candidate categories with feedback operations of "discard" and "replace" are used as error categories and error analysis is performed to obtain the analysis results.

[0194] The game resource classification model is optimized based on the results of extended classification and error analysis, resulting in an optimized game resource classification model.

[0195] Accordingly, this application also provides an electronic device, which can be a terminal, such as a smartphone, tablet computer, laptop computer, touch screen, game console, personal computer (PC), personal digital assistant (PDA), or other terminal device. Alternatively, the electronic device can be a server.

[0196] like Figure 6 As shown, Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 400 includes a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, and a computer program stored in the memory 402 and executable on the processor. The processor 401 and the memory 402 are electrically connected. Those skilled in the art will understand that the electronic device structure shown in the figure does not constitute a limitation on the electronic device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0197] The processor 401 is the control center of the electronic device 400. It connects various parts of the electronic device 400 via various interfaces and lines. By running or loading software programs and / or units stored in the memory 402, and by calling data stored in the memory 402, it executes various functions and processes data of the electronic device 400, thereby providing overall monitoring of the electronic device 400. The processor 401 can be a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), etc., and can implement or execute the methods, steps, and logic diagrams disclosed in the embodiments of this application.

[0198] In this embodiment, the processor 401 in the electronic device 400 loads the instructions corresponding to the processes of one or more applications into the memory 402 according to the following steps, and the processor 401 runs the applications stored in the memory 402 to realize various functions. For example, it first preprocesses the multimodal data of game resources to obtain a preprocessed dataset, then inputs the preprocessed dataset into the game resource classification model for encoding processing to obtain a feature vector set, and finally classifies the game resources according to the feature vector set to determine the classification result of the game resources. The multimodal data includes at least the parsing file of the game resources, and the feature vector set includes the structured embedding vector after encoding and preprocessing the parsing file. The structured embedding vector represents the semantic information of the mapping between spatial coordinates and resource attributes in the game resources.

[0199] Therefore, this application can extract the semantic information of spatial coordinates and resource attributes from the parsed files of game resources, and integrate this information with other types of modal data to classify game resources. In this way, it can improve the accuracy, efficiency and diversity of game resource classification based on the above rich feature information, and reduce the time and manual cost of classification.

[0200] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0201] Optional, such as Figure 6 As shown, the electronic device 400 also includes: a touch display screen 403, a radio frequency circuit 404, an audio circuit 405, an input unit 406, and a power supply 407. The processor 401 is electrically connected to the touch display screen 403, the radio frequency circuit 404, the audio circuit 405, the input unit 406, and the power supply 407. Those skilled in the art will understand that... Figure 6 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0202] The touch display screen 403 can be used to display a graphical user interface (GUI) and receive operation commands generated by the user interacting with the GUI. The touch display screen 403 may include a display panel and a touch panel. The display panel can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the electronic device. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. Optionally, the display panel can be configured using a liquid crystal display (LCD), organic light-emitting diode (OLED), or other similar technologies. The touch panel can be used to collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel), generate corresponding operation commands, and execute the corresponding program according to the operation commands. Optionally, the touch panel may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch location and the signal generated by the touch operation, transmitting the signal to the touch controller. The touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 401. It can also receive and execute commands from the processor 401. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it transmits the information to the processor 401 to determine the type of touch event. Subsequently, the processor 401 provides corresponding visual output on the display panel based on the type of touch event. In this embodiment, the touch panel and the display panel can be integrated into the touch display screen 403 to achieve input and output functions. However, in some embodiments, the touch panel and the touch display screen 403 can be implemented as two independent components to achieve input and output functions. That is, the touch display screen 403 can also be used as part of the input unit 406 to achieve input functions.

[0203] The radio frequency circuit 404 can be used to transmit and receive radio frequency signals to establish wireless communication with network devices or other electronic devices, and to transmit and receive signals with network devices or other electronic devices.

[0204] Audio circuit 405 can be used to provide an audio interface between a user and an electronic device via a speaker and a microphone. Audio circuit 405 can convert received audio data into electrical signals and transmit them to the speaker, where the speaker converts them into sound signals for output. Conversely, the microphone converts collected sound signals into electrical signals, which are then received by audio circuit 405, converted back into audio data, and then processed by processor 401 before being transmitted via radio frequency circuit 404 to, for example, another electronic device, or output to memory 402 for further processing. Audio circuit 405 may also include an earphone jack to provide communication between peripheral headphones and electronic devices.

[0205] The input unit 406 can be used to receive input numbers, characters, or user characteristic information (such as fingerprints, iris, facial information, etc.), and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control.

[0206] Power supply 407 is used to supply power to various components of electronic device 400. Optionally, power supply 407 can be logically connected to processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. Power supply 407 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0207] although Figure 6 As not shown in the diagram, the electronic device 400 may also include a camera, sensor, wireless fidelity module, Bluetooth module, etc., which will not be described in detail here.

[0208] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0209] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0210] To this end, embodiments of this application provide a computer-readable storage medium storing multiple computer programs that can be loaded by a processor to execute any of the game resource classification methods provided in this application. The computer program can execute the following steps of the game resource classification method: first, preprocessing the multimodal data of the game resources to obtain a preprocessed dataset; then, inputting the preprocessed dataset into a game resource classification model for encoding processing to obtain a feature vector set; finally, classifying the game resources based on the feature vector set to determine the classification result. The multimodal data includes at least a parsing file of the game resources, and the feature vector set includes structured embedding vectors after encoding and preprocessing the parsing file. The structured embedding vectors represent the semantic information of the mapping between spatial coordinates and resource attributes in the game resources.

[0211] Therefore, this application can extract the semantic information mapping spatial coordinates and resource attributes from the parsed files of game resources, and integrate this information with other types of modal data to classify game resources. This allows for improved accuracy, efficiency, and diversity in game resource classification based on the rich feature information, while reducing classification time and labor costs. Specific implementation details of each of the above operations can be found in the preceding embodiments and will not be repeated here.

[0212] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0213] Since the computer program stored in the computer-readable storage medium can execute any of the game resource classification methods provided in the embodiments of this application, the beneficial effects that any of the game resource classification methods provided in the embodiments of this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.

[0214] According to one aspect of this application, a computer program product or computer program is also provided, comprising computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the methods provided in the various optional implementations of the above embodiments.

[0215] In the above embodiments of the game resource classification device, computer-readable storage medium, electronic device, and computer program product, the descriptions of each embodiment have different focuses. Parts not described in detail in a particular embodiment can be referred to in the relevant descriptions of other embodiments. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process and beneficial effects of the game resource classification device, computer-readable storage medium, computer program product, electronic device, and their corresponding units described above can be referred to the description of the game resource classification method in the above embodiments, and will not be repeated here.

[0216] The foregoing has provided a detailed description of a method, apparatus, electronic device, computer-readable storage medium, and computer program product for classifying game resources according to embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for classifying game resources, characterized in that, include: The multimodal data of game resources is preprocessed to obtain a preprocessed dataset; The preprocessed dataset is input into the game resource classification model for encoding processing to obtain a set of feature vectors. The game resources are classified according to the set of feature vectors to determine the classification result of the game resources; The multimodal data includes at least the parsed file of the game resource, and the feature vector set includes structured embedding vectors after encoding and preprocessing the parsed file. The structured embedding vectors represent the semantic information of the mapping between spatial coordinates and resource attributes in the game resource.

2. The method according to claim 1, characterized in that, The preprocessed dataset is input into a game resource classification model for encoding to obtain a set of feature vectors, including: Based on the parsed files of the game resources, structured features are extracted; the structured features include the resource attributes of multiple resource components and the spatial coordinates of each resource component in the game map; The structured features are encoded to obtain an intermediate vector; The intermediate vector is fused with attention information from each resource component to obtain the structured embedding vector.

3. The method according to claim 2, characterized in that, The intermediate vector is fused with attention information from each resource component to obtain the structured embedding vector, including: The structured features are subjected to rotational position encoding to obtain a position embedding matrix that represents the relative positional relationships between the resource components in the game resource; Based on the location embedding matrix, determine the attention index parameters that characterize the attention information between each of the resource components; The attention index parameter is fused with the value vectors of other types of modal data to obtain fused information, and the intermediate vector is updated based on the fused information to obtain the structured embedding vector; The other types of modal data refer to at least one type of modal data other than the parsed file.

4. The method according to claim 2, characterized in that, The multimodal data further includes at least one of the text, images, and audio of the game resources; the preprocessed dataset is input into a game resource classification model for encoding processing to obtain a feature vector set, which also includes: The text of the game resources is segmented and sorted to obtain a text sequence, which is then encoded through the bidirectional semantic coding layer of the game resource classification model to obtain a text embedding vector. The game resource images are segmented and normalized to obtain multiple image blocks, and the multiple image blocks are encoded through the visual transformation layer of the game resource classification model to obtain image embedding vectors; The audio of the game resource is preprocessed to obtain Mel spectrum information, and then encoded through the speech pre-training layer of the game resource classification model to obtain an audio embedding vector, wherein the feature vector set includes the text embedding vector, the image embedding vector, and the audio embedding vector.

5. The method according to claim 4, characterized in that, The step of classifying the game resources based on the feature vector set to determine the classification result of the game resources includes: The alignment layer of the game resource classification model performs semantic alignment on the feature vector set to obtain an aligned feature vector group. The latent vector is obtained by jointly encoding each feature vector in the feature vector group. The game resources are classified based on the latent vectors to obtain the classification results of the game resources.

6. The method according to claim 5, characterized in that, The training process for the alignment layer includes: Multiple first feature vectors from the same game resource are combined to obtain a positive sample feature group; Multiple second feature vectors from different game resources are combined to obtain negative sample feature groups; Construct a loss function that characterizes the ability of the alignment layer to distinguish between the positive sample feature group and the negative sample feature group; The initial alignment layer is iteratively trained until the loss function converges, resulting in the trained alignment layer.

7. The method according to claim 5, characterized in that, The latent vectors are obtained by jointly encoding each feature vector in the feature vector group, including: The feature vector group is input into the joint encoding layer of the game resource classification model for processing to obtain the first fusion vector of the structure and image of the game resource, the second fusion vector of the structure and text, and the third fusion vector of the structure and audio. The first fusion vector, the second fusion vector, and the third fusion vector are concatenated to obtain the latent vector.

8. The method according to claim 7, characterized in that, The feature vector set is input into the joint encoding layer of the game resource classification model for processing to obtain a first fusion vector of the game resource structure and image, a second fusion vector of the structure and text, and a third fusion vector of the structure and audio, including: Using the structured embedding vector as the query vector, and the image embedding vector, the text embedding vector, and the audio embedding vector as key-value pairs, calculate the first attention weight between structure and image, the second attention weight between structure and text, and the third attention weight between structure and audio. According to the cross-attention mechanism, the first attention weight and the value vector of the image modality are processed to obtain the first cross-attention result, the second attention weight and the value vector of the text modality are processed to obtain the second cross-attention result, and the first attention weight and the value vector of the audio modality are processed to obtain the third cross-attention result. A multi-head attention mechanism is used to fuse the features of the first cross-attention result, the second cross-attention result, and the third cross-attention result to obtain the first fusion vector, the second fusion vector, and the third fusion vector.

9. The method according to claim 7, characterized in that, The step of classifying the game resources based on the latent vectors to obtain the classification results of the game resources includes: The latent vectors are linearly transformed by the fully connected layer of the game resource classification model to obtain the original vectors corresponding to each classification. An activation function is used to process multiple original vectors to obtain the probability distribution of the game resource belonging to multiple defined categories, which serves as the classification result of the game resource.

10. The method according to claim 1, characterized in that, The method further includes: The candidate categories from the classification results of the game resources are sent to the user, and the user's feedback on the candidate categories is obtained; the feedback operation includes acceptance, replacement, or discard. The candidate categories with feedback operations of "adoption" are used as the target categories and synonym expansion is performed to obtain the extended categories. The candidate categories with feedback operations of "discard" and "replace" are used as error categories and error analysis is performed to obtain the analysis results. The game resource classification model is optimized based on the extended classification and the error analysis results to obtain the optimized game resource classification model.

11. A device for classifying game resources, characterized in that, The device includes: The data processing module is used to preprocess the multimodal data of game resources to obtain a preprocessed dataset; The data encoding module is used to input the preprocessed dataset into the game resource classification model for encoding processing to obtain a set of feature vectors; The resource classification module is used to classify the game resources according to the feature vector set and determine the classification result of the game resources; The multimodal data includes at least the parsed file of the game resource, and the feature vector set includes structured embedding vectors after encoding and preprocessing the parsed file. The structured embedding vectors represent the semantic information of the mapping between spatial coordinates and resource attributes in the game resource.

12. An electronic device, characterized in that, The system includes a processor and a memory, the memory storing multiple instructions; the processor loads instructions from the memory to perform the steps of the game resource classification method as described in any one of claims 1 to 9.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the steps of the game resource classification method as described in any one of claims 1 to 9.