Text-driven 3D model generation method and system for game development
By obtaining text descriptions of game design scenarios, extracting and mapping key information, and generating 3D models that meet game requirements, the problems of low generation efficiency and poor quality in existing technologies are solved, and efficient and automated 3D model generation is achieved.
Patent Information
- Application Number
- CN202511013511.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-07-23
AI Technical Summary
Existing technologies find it difficult to quickly and efficiently generate 3D models that meet game design requirements, especially when dealing with complex scene element descriptions in the form of multiple consecutive paragraphs, and the generated models lack details and accuracy.
By obtaining the target text description set of the game design scene, key information is extracted and processed, and the correlation mapping relationship between the scene features and the 3D model elements is established. The initial 3D model is generated using the preset model generation framework, and the structure is adjusted according to the game development constraints to generate the final model that meets the requirements.
It realizes the automatic generation from text to 3D models, improves generation efficiency and quality, reduces development costs, and provides a convenient and efficient 3D model generation solution.
Smart Images

Figure CN120526060B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of generative artificial intelligence technology, and in particular to a text-driven 3D model generation method and system for game development. Background Art
[0002] In the field of game development, the creation of 3D models is a key step in building visual elements such as game scenes and characters. Traditional 3D model creation methods mainly rely on professional modeling software and manual operations of designers. This method is not only time-consuming and labor-intensive, but also requires high professional skills of designers, resulting in high development costs and low efficiency. With the continuous growth of game development needs and the intensification of market competition, how to quickly and efficiently generate 3D models that meet game design requirements has become an urgent problem to be solved. In the related art, although some technologies for generating 3D models based on text descriptions have emerged, most of these technologies can only handle simple, isolated text descriptions, and it is difficult to handle complex scene element descriptions in the form of multiple consecutive paragraphs. In addition, the generated 3D models often lack details and accuracy, and cannot meet the high-quality requirements of game development. Summary of the Invention
[0003] In view of the above-mentioned problems, in combination with the first aspect of the present invention, an embodiment of the present invention provides a text-driven 3D model generation method for game development, the method comprising:
[0004] Obtaining a target text description set of a game design scene, wherein the target text description set includes scene element description contents in the form of multiple consecutive paragraphs;
[0005] Performing key information extraction processing on the target text description set to obtain a scene feature set including object type features, morphological feature descriptions, and spatial relationship features;
[0006] Establishing an association mapping relationship between each feature in the scene feature set and a 3D model element, wherein the 3D model element includes a geometric structure, surface material properties, and a topological connection relationship;
[0007] Based on the association mapping relationship, a preset model generation framework is called to generate initial 3D model data corresponding to the scene feature set;
[0008] The initial 3D model data is structurally adjusted according to the model adaptation constraint conditions of the game development to obtain a final 3D model output result that meets the development requirements.
[0009] On the other hand, an embodiment of the present invention also provides a text-driven 3D model generation system for game development, including a processor and a machine-readable storage medium, wherein the machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.
[0010] Based on the above aspects, the embodiment of the present invention obtains a target text description set of the game design scene and extracts key information from the set to obtain a scene feature set containing object type features, morphological feature descriptions, and spatial relationship features. Then, an association mapping relationship is established between each feature in the scene feature set and the 3D model elements, thereby achieving a precise conversion from text descriptions to 3D model elements. Based on the association mapping relationship, a preset model generation framework is called to generate initial 3D model data corresponding to the scene feature set, thereby achieving automated generation from text to 3D models. Finally, the initial 3D model data is structurally adjusted according to the model adaptation constraints of game development to obtain a final 3D model output result that meets the development requirements, ensuring that the generated 3D model meets the high-quality requirements of game development. This significantly improves the generation efficiency and quality of 3D models in game development, reduces development costs, and provides game developers with a more convenient and efficient 3D model generation solution. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 The figure is a schematic diagram of the execution flow of a text-driven 3D model generation method for game development provided by an embodiment of the present invention.
[0012] Figure 2 Schematic diagram of exemplary hardware and software components of a text-driven 3D model generation system for game development provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0013] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1 This is a flow chart of a text-driven 3D model generation method for game development provided by an embodiment of the present invention. The text-driven 3D model generation method for game development is introduced in detail below.
[0014] Step S110: Obtain a target text description set of the game design scene, wherein the target text description set includes scene element description contents in the form of multiple consecutive paragraphs.
[0015] In this embodiment, taking a medieval magic role-playing game as an example, game developers usually write detailed scene texts to describe each game scene. For a magic town scene, the target text description set may consist of multiple consecutive paragraphs.
[0016] For example, the first paragraph describes: "At the entrance to the magic town stands a tall archway, built of massive stones, its surface engraved with mysterious runes and emitting a faint glow." This paragraph provides preliminary information about the arch as a scene element for subsequent 3D model generation. The second paragraph reads: "The town's streets are lined with colorful magic shops, their windows gleaming with strange lights, as if displaying the mysterious goods within." This introduces the magic shop scene element. The third paragraph explains: "In the center of the town stands a massive fountain, its water gushing from the hands of a sculpture, surrounded by exquisite flower beds, the flowers exuding an enchanting fragrance." This further enriches the scene elements, including the fountain and flower beds.
[0017] These continuous paragraphs of scene element descriptions cover a wide range of scene elements, from buildings to natural landscapes, from static objects to dynamic effects. The target text description set obtained through the above method contains detailed descriptions of various elements in the game scene.
[0018] Step S120: performing key information extraction processing on the target text description set to obtain a scene feature set including object type features, morphological feature descriptions and spatial relationship features.
[0019] In order to transform the target text description set into information that can actually guide 3D model generation, key information extraction is required. For example, for the target text description set of the medieval magic town mentioned above, key information extraction is divided into multiple sub-steps.
[0020] Step S121: performing semantic segmentation processing on the target text description set, dividing the description content in the form of continuous paragraphs into multiple independent semantic units.
[0021] For example, each paragraph in the target text description set needs to be divided into independent semantic units based on semantic logic and grammatical structure. For example, the paragraph "At the entrance to the magic town, stands a tall arch. The arch is made of huge stones, engraved with mysterious runes on its surface, and emits a faint light" can be divided into the following independent semantic units: "At the entrance to the magic town" clarifies the location; "Stands a tall arch" points out the main object; "The arch is made of huge stones" describes the arch's material; "Mysterious runes are engraved on its surface" explains the arch's surface features; and "Emits a faint light" describes the arch's dynamic effect.
[0022] The paragraph "The streets of the town are lined with colorful magic shops, and the shop windows are shining with strange lights, as if displaying the mysterious goods in the stores" can be divided into independent semantic units such as "the streets of the town" (location information), "lined with colorful magic shops" (main objects), "the shop windows are shining with strange lights" (characteristics of the shop windows), and "as if displaying the mysterious goods in the stores" (further explanation of the light effect of the shop windows).
[0023] Similarly, the paragraph “There is a huge fountain in the center of the town, the water of the fountain gushes out from the hands of the sculpture, and it is surrounded by exquisite flower beds with flowers emitting charming fragrance” can be divided into independent semantic units such as “town center” (location information), “there is a huge fountain” (main object), “the water of the fountain gushes out from the hands of the sculpture” (dynamic characteristics of the fountain), “surrounded by exquisite flower beds” (spatial relationship between the fountain and the flower beds), and “the flowers emit charming fragrance” (characteristics of the flowers in the flower beds).
[0024] Therefore, through the above semantic segmentation process, the continuous paragraph information is split into independent semantic units.
[0025] Step S122: performing entity recognition processing on each of the independent semantic units to extract object type features representing specific game elements, wherein the object type features include character models, scene props, and environmental decorations.
[0026] After semantic segmentation, entity recognition is performed on each independent semantic unit to extract the object type features representing specific game elements. In the example of a medieval magic town, for the independent semantic unit "a tall arch stands," entity recognition can determine that "arch" belongs to the object type feature of a scene prop. This is because the arch is an integral part of the town scene, serving as decoration and spatial division.
[0027] For the independent semantic unit "colorfully lined magic shops," "magic shops" also belong to the scene props. Magic shops are places in the town where players interact and trade, and are an important part of the scene. Meanwhile, the "flowers" in "flowers exuding an enchanting fragrance" are environmental decorations, adding a natural and vibrant atmosphere to the town scene.
[0028] When performing entity recognition, some complex situations may arise. For example, in the independent semantic unit "water from the fountain gushes from the hands of the sculpture," both "fountain" and "sculpture" are scene props. For those skilled in the art, this requires careful analysis of the role and nature of each element in the game scene to accurately determine its object type characteristics.
[0029] Step S1221: Perform vocabulary filtering on the independent semantic unit to remove non-entity words that represent degree modification or emotional description.
[0030] During entity recognition, to more accurately extract object type features, lexical filtering is required for independent semantic units. For example, in the sentence "There stands a tall archway," "tall" is a degree modifier, primarily describing the archway's characteristics but not inherently representing a specific game element. Lexical filtering removes the non-entity word "tall," retaining only the key entity word "archway."
[0031] For "the flowers emit a charming fragrance", "charming" is an emotional descriptive word used to express the feeling given by the fragrance of flowers. It is not a specific game element entity, so it will also be removed in the vocabulary filtering process, and only "flowers" will be retained.
[0032] Similarly, in the sentence "The shop window shimmered with a strange light," "strange" is a degree modifier and is filtered out, leaving "shop window" as the key entity information. This word filtering process can reduce interference information and make subsequent entity recognition more accurate.
[0033] Step S1222: Perform domain dictionary matching processing on the remaining words, and classify the matched words into object type features corresponding to role models, scene props or environmental decorations.
[0034] After completing vocabulary filtering, the remaining words need to be matched against the domain dictionary. The domain dictionary is a pre-built dictionary that contains various common game elements and their object type characteristics. For example, in the domain dictionary, "arch" is clearly classified as a scene prop. By matching the filtered "arch" with the domain dictionary, we can determine that its object type characteristics are scene props.
[0035] The domain dictionary also has a corresponding record for "magic shop," categorizing it as a scene prop. This is because a magic shop is a functional building in the game and falls under the category of scene props. "Flowers" are categorized as environmental decorations in the domain dictionary, and matching accurately determines their object type characteristics.
[0036] If you encounter some new words, there may be no direct matches in the domain dictionary. At this time, the next step of context inference processing is required.
[0037] Step S1223: Perform context inference processing on the words that are not matched to the domain dictionary, and determine the corresponding object type features based on the object type feature association relationship of the preceding and following semantic units.
[0038] In actual processing, some words may not be directly matched in the domain dictionary. For example, a new word, "Mysterious Rune Crystal," appears in the target text description set, but the domain dictionary doesn't have a record for it. In this case, context inference is necessary based on the object type feature associations between the preceding and following semantic units.
[0039] Suppose the independent semantic unit for "Mysterious Rune Crystal" is "On the display stand in the magic shop, a mysterious rune crystal was placed, emitting a mysterious glow." In the preceding and following semantic units, "magic shop" is identified as a scene prop, and "display stand" is also a scene prop. The placement of the "Mysterious Rune Crystal" on the display stand in the magic shop indicates its connection to the scene props. The context suggests that the "Mysterious Rune Crystal" serves as a decorative item in the game scene, adding to the mysterious atmosphere. Therefore, it can be classified as a scene prop.
[0040] Through the above context inference processing, the problem of determining the object type features of words not included in the domain dictionary can be solved, thereby improving the accuracy of entity recognition.
[0041] Step S1224: performing duplicate item merging processing on all determined object type features, and retaining the first appearance position information of each object type feature.
[0042] After entity recognition, object type features may be duplicated. For example, the target text description contains multiple mentions of "flowers." To avoid information redundancy, all identified object type features need to be merged for duplicate entries.
[0043] For multiple "flower" object type features, only the location information of the first occurrence is retained. For example, if "flower" first appears in the independent semantic unit "There is a huge fountain in the center of the town, surrounded by exquisite flower beds, and the flowers exude an enchanting fragrance," then after the merging process, only the location information of "flower" in this independent semantic unit is recorded, and the location information of subsequent recurring "flower" is ignored.
[0044] By merging duplicate items, the storage and management of object type features can be optimized, unnecessary information redundancy can be reduced, and the efficiency of subsequent processing can be improved.
[0045] Step S1225: Generate an object type feature set including the object type name and the appearance position information as the object type feature.
[0046] After completing the above series of processes, all determined object type features and their first appearance locations are integrated to generate an object type feature set. Taking the medieval magic town as an example, the object type feature set might include: "Arch (at the entrance of the magic town)", "Magic Shop (on both sides of the town's streets)", "Flowers (a huge fountain in the center of the town)", "Mysterious Rune Crystal (on the display stand in the magic shop)", etc.
[0047] The object type feature set includes the object type names of various specific game elements in the game scene and their appearance position information in the target text description set.
[0048] Step S123: performing attribute description extraction processing on each of the independent semantic units to extract a morphological feature description describing the appearance of the object, wherein the morphological feature description includes a contour shape description, a surface texture description, and a scale size description.
[0049] After extracting object type features, we need to perform attribute description extraction on each independent semantic unit to obtain morphological features describing the object's appearance. For example, "There stands a tall arch, built of huge stones, with mysterious runes engraved on its surface." "Tall" can be used to describe the arch's size, indicating its large height; "built of huge stones" can be seen as a surface texture description, suggesting that the arch's surface has the rough texture of stacked stones; and "engraved with mysterious runes" is a surface texture description, further explaining the decorative features of the arch's surface.
[0050] For “The magic shop’s window is shimmering with strange light,” the outline shape of “the window” can be inferred from the context to be a rectangle (a common shape for general shop windows), and “shimmering with strange light” can be seen as a special effect describing the surface texture, reflecting the dynamic appearance characteristics of the window.
[0051] Step S1231: performing adjective phrase recognition processing on the independent semantic unit to extract adjective phrases used to describe shape, texture and size.
[0052] When extracting attribute descriptions, we first need to identify adjective phrases for independent semantic units. For example, in the sentence "There stands a tall archway," "tall" is an adjective phrase describing the arch's size, falling under the category of proportional size description. In the sentence "The flowers exude an enchanting fragrance," while "enchanting" isn't an adjective phrase describing shape, texture, or size, in the sentence "Exquisite flower bed," "exquisite" can be seen as a modification of the flower bed's overall appearance, potentially relating to its outline, shape, surface texture, and other aspects.
[0053] In the example "The shop window shimmered with a strange light," "strange" is an adjective phrase describing the characteristics of the window's light and is part of the surface texture description. Through the above adjective phrase recognition process, we can extract key descriptive information related to the object's morphological characteristics within the independent semantic units.
[0054] Step S1232: matching the adjective phrase with a preset morphological description dictionary to determine the corresponding contour shape type, surface texture type, and proportional size type.
[0055] After extracting adjective phrases, they need to be matched against a pre-set morphological description dictionary. This dictionary contains common adjective phrases and their corresponding contour shape types, surface texture types, and scale types. For example, in the morphological description dictionary, "tall" corresponds to "large" in the scale type. This matching process confirms that the arch described by "tall" is large in scale.
[0056] For example, the morphological description dictionary might match the word "exquisite" to the surface texture category "delicate texture," describing the delicate appearance of the flower bed. Similarly, the word "strange" might match the surface texture category "special luminous texture," reflecting the unique glow of the window. Through this matching process, we can accurately determine the morphological feature type corresponding to the adjective phrase.
[0057] Step S1233: performing semantic conversion processing on the comparative description contained in the adjective phrase, and converting the comparative word into a specific morphological reference object.
[0058] Adjective phrases may contain comparative descriptions. For example, if the target text description contains the phrase "The fountain is much higher than the surrounding flower beds," "much higher than..." is a comparative description. During semantic conversion, the comparative term needs to be converted into a specific morphological reference object. Here, "the surrounding flower beds" can be used as the morphological reference object, converting "much higher than the surrounding flower beds" into "The fountain has a significant height difference relative to the surrounding flower beds."
[0059] For the sentence "The magic shop's window is brighter than other shop windows," we use "other shop windows" as the morphological reference object and transform the comparative description into "The magic shop's window is brighter than other shop windows." This semantic transformation process more clearly expresses the morphological relationship between objects.
[0060] Step S1234: performing association and binding processing on the outline shape type, surface texture type, scale size type and morphological reference object to generate a morphological feature description corresponding to the object type feature in the independent semantic unit.
[0061] After completing the above steps, you need to associate and bind the outline shape type, surface texture type, scale type, and morphological reference object. For example, taking an arch as an example, the previous processing determines that its scale type is "large," and its surface texture types are "rough stone stacking texture" and "rune decoration texture." This information is associated and bound with the arch's object type features to generate a morphological feature description corresponding to the arch.
[0062] For the magic shop window, the outline shape type is a rectangle, and the surface texture type is a "special luminous texture." This information is associated and bound with the object type characteristics of the magic shop window to obtain its morphological feature description. Through this association and binding process, the various morphological feature information of the object can be integrated to form a complete morphological feature description.
[0063] Step S1235: perform ambiguity elimination processing on all generated morphological feature descriptions, and correct the contradictory morphological feature descriptions through consistency check of the morphological descriptions of the context semantic unit.
[0064] After generating morphological feature descriptions, some ambiguities or contradictions may exist. For example, the morphological feature descriptions of the same object may differ in different independent semantic units. For example, if an arch is described as "tall" in one independent semantic unit and "short" in another, this would be a contradiction.
[0065] This is where ambiguity resolution is necessary. This involves checking the consistency of the morphological descriptions of the contextual semantic units to correct the discrepancy. The more accurate description can be determined based on the overall logic of the context and more detailed information. For example, if the context surrounding the arch emphasizes the grandeur of the town and the importance of the arch, then the description "tall" is more consistent with the overall logic, so the description "short" can be corrected to "tall."
[0066] Even ambiguous descriptions can be clarified through context. For example, "The magic shop's window is a little bright." While the description "a little bright" is ambiguous, by examining the context, we can find that the window shimmers with a strange light. Therefore, "a little bright" can be corrected to "bright with a strange light." This ambiguity-removing process ensures the accuracy and consistency of morphological feature descriptions.
[0067] Step S124: performing relationship analysis on each of the independent semantic units to extract spatial relationship features describing the positional connections between objects, wherein the spatial relationship features include adjacent relationship descriptions, inclusion relationship descriptions, and occlusion relationship descriptions.
[0068] In addition to describing object type and morphological features, we also need to perform relational analysis on each independent semantic unit to extract spatial relationship features that describe the positional connections between objects. For example, consider the sentence "At the entrance to the magic town stands a tall arch, and the town's streets extend from behind the arch." There is a spatial relationship between "arch" and "town streets." "The streets extend from behind the arch" can be described as an adjacency relationship, indicating that the street and arch are spatially adjacent and have a sequential order.
[0069] For "There is a huge fountain in the center of the town, and the fountain is surrounded by exquisite flower beds", there is an inclusion relationship between "fountain" and "flower beds". The flower beds surround the fountain, which means that the flower beds surround the fountain in space.
[0070] Step S1241: performing prepositional phrase recognition processing on the independent semantic units to extract prepositional phrases representing positional connections.
[0071] When performing relationship analysis, we first need to identify prepositional phrases for independent semantic units. For example, in the sentence "At the entrance to the magic town stands a tall archway," "at the entrance" is a prepositional phrase that indicates the location of the archway and embodies the positional connection between the archway and the entrance to the magic town.
[0072] In the sentence "The fountain is surrounded by beautiful flower beds," while "around" isn't a typical preposition, in this context it can be considered a similar prepositional phrase that indicates a locational connection, explaining the spatial surrounding relationship between the flower beds and the fountain. Prepositional phrase recognition allows us to extract key locational connection information from independent semantic units.
[0073] Step S1242: matching the prepositional phrase with a preset spatial relationship dictionary to determine the corresponding adjacency relationship, inclusion relationship, or occlusion relationship type to obtain the spatial relationship type.
[0074] After extracting prepositional phrases, they need to be matched against a pre-set spatial relationship dictionary. This dictionary contains various common prepositional phrases and their corresponding spatial relationship types. For example, "at the entrance of..." corresponds to the adjacency relationship in the spatial relationship dictionary, indicating that the archway and the entrance to the magic town are spatially adjacent.
[0075] The spatial relationship dictionary maps "around" to a containment relationship (surrounding can be considered a special form of containment), indicating a containment relationship between the flower bed and the fountain. Through the above matching process, the spatial relationship type corresponding to the prepositional phrase can be accurately determined.
[0076] Step S1243: performing location processing on the two object type features involved in the prepositional phrase to determine the specific objects they refer to in the object type feature set.
[0077] After determining the spatial relationship type through prepositional phrase recognition and matching with the spatial relationship dictionary, the two object type features involved in the prepositional phrase need to be located. Taking the sentence "At the entrance to the magic town, stands a tall archway" as an example, the prepositional phrase "at the entrance..." involves two object type features: "magic town entrance" and "archway." In the previously generated object type feature set, "archway" has been identified as a scene prop type object. As for "magic town entrance," although it lacks a distinct entity like "archway," it can be semantically considered a special scene identifier and can also be classified as a scene prop. By searching and locating the two specific objects referred to in the object type feature set, the "archway" and "magic town entrance" are clearly identified, allowing for the subsequent accurate construction of spatial relationships.
[0078] For example, in the sentence "The fountain is surrounded by exquisite flower beds," the prepositional phrase "surrounding" refers to two object type features: "fountain" and "flower bed." Within the object type feature set, "fountain" and "flower bed" are categorized as scene props and environmental decorations, respectively. Positioning processing determines the specific locations and features of these two objects within the set, laying the foundation for constructing their spatial relationship.
[0079] Step S1244: performing triple construction processing on the spatial relationship type and the specific referenced object to generate a spatial relationship triple containing the first object, the spatial relationship type and the second object.
[0080] After determining the specific referents of the two object type features in the prepositional phrase, we need to construct a triplet of the spatial relationship type and the specific referent. For example, in the sentence "At the entrance to the magic town, there stands a tall archway," the first object is "archway," the spatial relationship type is "adjacent (the entrance can be understood as an adjacent relationship)," and the second object is "magic town entrance." The constructed spatial relationship triplet is (archway, adjacent, magic town entrance).
[0081] For the sentence "The fountain is surrounded by beautiful flower beds," the first object is "flower bed," the spatial relationship type is "contains (surrounding can be considered a containment relationship)," and the second object is "fountain." The resulting spatial relationship triple is (flower bed, contain, fountain). By constructing this spatial relationship triple, we can clearly express the spatial relationship between objects.
[0082] Step S1245: Performing logic conflict detection processing on all generated spatial relationship triples, by checking whether there are conflicting spatial relationship types between the same pair of objects, and correcting the conflicting spatial relationship features.
[0083] After generating all spatial relationship triples, logical conflicts may occur. For example, in the target text description set, contradictory descriptions such as "The arch is to the left of the entrance to the magic town" and "The arch is to the right of the entrance to the magic town" may appear simultaneously. When performing logical conflict detection, it is necessary to check whether there are conflicting spatial relationship types between the same pair of objects.
[0084] For the above example, we need to re-examine the context and determine which description is more logical. If, when describing the town layout, more information indicates that "the arch is to the left of the magic town entrance" is more logically related to the other scene elements, then the spatial relationship feature "the arch is to the right of the magic town entrance" can be revised to "the arch is to the left of the magic town entrance."
[0085] For example, if the descriptions "the fountain is surrounded by flower beds" and "the flower beds are inside the fountain" are contradictory, checking the context reveals that the previous description emphasizes the decorative role of the flower beds surrounding the fountain. Therefore, "the flower beds are inside the fountain" is corrected to "the fountain is surrounded by flower beds." This logical conflict detection ensures the accuracy and consistency of spatial relationship features.
[0086] Step S125: performing feature integration processing on the object type features, morphological feature descriptions and spatial relationship features according to the original order of semantic units to generate a scene feature set with context association.
[0087] After extracting object type features, generating morphological descriptions, and constructing spatial relationship features, these features need to be integrated according to the original order of semantic units. For example, in the mysterious forest scene described earlier, following the original order of the paragraphs in the target text description set, the description of "tall trees" that appears first is integrated with its object type features (trees, scene props), morphological descriptions (large trees with thick trunks and lush branches, rough texture of piled stones, and no obvious morphological reference objects), and spatial relationship features (adjacent to the forest floor).
[0088] Next is the description of “fallen leaves”, integrating its object type characteristics (fallen leaves, environmental decorations), morphological characteristics description (thin flakes, small objects with diverse colors, paper texture, no obvious morphological reference objects) and spatial relationship characteristics (contact with the forest floor).
[0089] Then comes the description of “mushroom”, integrating its object type features (mushroom, environmental decoration), morphological feature description (umbrella-shaped, small object with spots on the surface, fungal texture, no obvious morphological reference object) and spatial relationship features (growing on fallen leaves).
[0090] Finally, the description of “stream” integrates its object type features (stream, scene props), morphological feature description (winding flow, clear and transparent ribbon-like object, water texture, no obvious morphological reference objects) and spatial relationship features (through the forest, adjacent to trees, fallen leaves, mushrooms, etc.).
[0091] By integrating the features in the original order of semantic units, a contextual scene feature set is generated. This scene feature set not only contains the type, form, and spatial relationship information of each object, but also preserves their order and context in the original text.
[0092] Step S130: establishing an association mapping relationship between each feature in the scene feature set and a 3D model element, wherein the 3D model element includes a geometric structure, surface material properties, and a topological connection relationship.
[0093] In order to convert the information in the scene feature set into a 3D model, it is necessary to establish an association mapping relationship between each feature and the 3D model elements.
[0094] Step S131: perform element classification processing on the object type features in the scene feature set, correspond the character model to the main mesh element in the geometric structure, correspond the scene props to the additional component elements in the geometric structure, and correspond the environmental decorations to the background construction elements in the geometric structure to obtain the element classification results.
[0095] In the scene feature set, object type features include character models, scene props, and environmental decorations. For example, in the medieval magic town scene, if the target text description set mentions "a brave knight patrolling the town," then "knight," as the character model, will be mapped to the main mesh element in the geometric structure. The main mesh element is the most important structural component of the 3D model, used to construct the basic shape and outline of the character.
[0096] For scene props like "arches" and "fountains," map them to additional component elements within the geometry structure. Additional component elements are added to the main mesh elements to enrich the scene's detail and functionality. For example, arches and fountains can be important components of a town scene, and their construction in the 3D model can be achieved through additional component elements.
[0097] For environmental decorations, such as flowers and fallen leaves, we mapped them to background building elements within the geometry structure. Background building elements are used to create the overall atmosphere and environment of the scene. Flowers and fallen leaves can add a sense of nature and vitality to the town scene, and are presented in the 3D model through background building elements.
[0098] Step S132: Attribute mapping is performed on the morphological feature description in the scene feature set, the contour shape description is mapped to the vertex coordinate parameters of the geometric structure, the surface texture description is mapped to the texture map type of the surface material attribute, and the scale size description is mapped to the scaling coefficient parameter of the geometric structure to obtain the attribute mapping result.
[0099] When performing attribute mapping, the contour shape description is processed first. Taking the "tall arch" as an example, its contour shape description is "tall and has a certain curvature." This contour shape description is mapped to the vertex coordinate parameters of the geometric structure. The vertex coordinate parameters determine the position of each vertex in the 3D model. By analyzing "tall and has a certain curvature," the distribution of the arch's vertices in 3D space can be determined. For example, a tall feature may mean that the vertex coordinate value in the vertical direction is larger, while a certain curvature requires adjusting the vertex coordinate value in the horizontal direction to form a corresponding curved shape.
[0100] For surface texture descriptions such as "the arch is made of huge stone blocks with mysterious runes carved on the surface", the surface texture descriptions of "stone blocks" and "runes" are mapped to the texture map types of the surface material properties. The texture maps matching "stone blocks" and "runes" can be found in the preset texture map library, such as "stone block rough texture map" and "rune decoration texture map", and these texture map types are assigned to the surface material properties of the arch.
[0101] For scale size descriptions such as "the tall arch", which indicates that it is larger in scale size, this scale size description is mapped to the scaling factor parameter of the geometric structure. The scaling factor parameter is used to adjust the size of the 3D model, and for the tall arch, a larger scaling factor can be set to make the arch appear larger in size in the 3D model.
[0102] Step S1321: Geometric feature extraction processing is performed on the contour shape description to identify the shape constituent elements represented by straight lines, curves or surfaces.
[0103] When mapping the contour shape description, geometric feature extraction processing is first performed. Taking "the tall arch" as an example, its contour shape description is "tall and has a certain degree of curvature", and "a certain degree of curvature" indicates the presence of a curve shape constituent element. For the identification of curve shape, the keywords such as "curvature" and "bend" in the description can be analyzed.
[0104] For example, "the rectangular magic shop window", where "rectangular" indicates the presence of a straight line shape constituent element. By extracting the keyword "rectangular", the shape feature of being composed of four straight line edges is identified.
[0105] For "the spherical magic crystal", "spherical" indicates the presence of a curved surface shape constituent element. By analyzing the description of "spherical", the shape feature of having a continuous curved surface is identified.
[0106] Step S1322: The shape constituent elements are converted into the variation rule parameters of the vertex coordinates in the geometric structure to generate the initial setting scheme of the vertex coordinate parameters.
[0107] After identifying the shape constituent elements, they need to be converted into the variation rule parameters of the vertex coordinates in the geometric structure. Taking the curve shape of the arch as an example, for the curve shape constituent element, the shape feature of the curve can be converted into the variation rule of the vertex coordinates through mathematical methods. For example, using the piecewise linear approximation method, the curve is approximated as a plurality of small line segments, and the end points of each small line segment correspond to the vertices in the geometric structure. According to the characteristics of the curvature and length of the curve, the coordinate values of each vertex in the 3D space are determined, thereby generating the initial setting scheme of the vertex coordinate parameters.
[0108] For the rectangular magic shop window, its rectilinear shape means the vertex coordinates change relatively simply. Based on the rectangle's length and width, we can determine the coordinates of the four vertices in 3D space. For example, if one vertex has coordinates (x1, y1, z1), the coordinates of adjacent vertices increase or decrease along the corresponding coordinate axes based on the rectangle's side lengths. This generates an initial set of vertex coordinate parameters.
[0109] For spherical magic crystals, the curved shape requires more complex processing. A spherical coordinate system can be used to determine the coordinate values of each vertex on the sphere based on parameters such as the sphere's radius and angle, generating an initial setting for the vertex coordinate parameters.
[0110] Step S1323: performing texture feature extraction processing on the surface texture description to identify texture representation elements representing smoothness, roughness or patterning.
[0111] When processing surface texture descriptions, texture feature extraction is required. For example, for example, "The arch is built of massive stones, with mysterious runes carved into its surface," the "stones" represent a rough texture element, while the "runes" represent a patterned texture element. By analyzing keywords in the description, such as "rough," "smooth," and "patterned," texture elements can be identified.
[0112] In the example "The magic shop window shimmers with strange lights," the "shimmering light" can be considered a unique texture element. It's neither smooth nor patterned, but rather a dynamic, luminous texture. This unique texture element was identified by extracting keywords like "shimmer" and "light."
[0113] Step S1324: matching the texture representation element with a preset texture mapping library to determine the texture mapping type of the corresponding surface material attribute.
[0114] After identifying the texture element, it needs to be matched against a preset texture library. The preset texture library is a collection of various common texture types. For example, for the texture element "stone stacking," the texture library is searched for a texture type that matches "stone stacking," such as "stone stacking rough texture map."
[0115] For the "rune" texture element, a matching "rune decoration texture map" is searched. For a special texture element like "shimmering light," a texture map type with a glowing effect, such as a "special glowing texture map," may need to be searched in the texture library. Through matching, the texture map type corresponding to the surface material properties is determined.
[0116] Step S1325: relative relationship extraction processing is performed on the scale size description to identify the size scale elements representing the whole and the part.
[0117] When processing the scale size description, relative relationship extraction processing is needed. Take "a tall archway" as an example. Although "tall" does not have a specific size value, its scale size can be understood through its relative relationship with the surrounding environment or other objects. If it is mentioned in the description that "the archway is taller than the trees around it", "taller than the trees around it" is a relative relationship, which embodies the size scale element of the archway in the overall scene.
[0118] For "the window of the magic store is wider than the windows of other stores", "wider than the windows of other stores" embodies the size scale element of the magic store window when compared with other windows. Through the extraction of these relative relationships, the size scale elements representing the whole and the part are identified.
[0119] Step S1326: convert the size scale elements into scaling factor parameters of the vertex coordinates of each component in the geometric structure, and generate an initial setting scheme of the scaling factor parameters.
[0120] After identifying the size scale elements, they need to be converted into scaling factor parameters of the vertex coordinates of each component in the geometric structure. Take the archway as an example. According to the size scale element "taller than the trees around it", a scaling factor can be determined. If the height of the surrounding trees corresponds to a certain range of vertex coordinates, and the height of the archway is relatively high, a scaling factor greater than 1 can be set. By scaling the coordinates of the vertex of each component in the geometric structure of the archway, an initial setting scheme of the scaling factor parameters is generated.
[0121] For the window of the magic store, according to the size scale element "wider than the windows of other stores", a scaling factor is determined. If the size of the windows of other stores corresponds to a certain range of vertex coordinates, and the window of the magic store is wider, a scaling factor greater than 1 can be set to scale the coordinates of the vertex of each component in the geometric structure of the window, and an initial setting scheme of the scaling factor parameters is generated.
[0122] Step S1327: combine the initial setting scheme of the vertex coordinate parameters, the texture mapping type and the initial setting scheme of the scaling factor parameters to generate the result of the attribute mapping processing.
[0123] After completing the initial vertex coordinate parameter settings, determining the texture map type, and initial scaling factor parameter settings, this information is combined. For example, for an arch, the initial vertex coordinate parameter settings, the texture map types for "Rough Stone Stacking Texture Map" and "Rune Decoration Texture Map," and the corresponding initial scaling factor parameter settings are combined. This combination creates a complete attribute mapping result, containing detailed information about the arch's geometric structure and surface material properties.
[0124] Step S133: Perform connection mapping processing on the spatial relationship features in the scene feature set, map the adjacent relationship description to the adjacent surface sharing parameters in the topological connection relationship, map the inclusion relationship description to the nested hierarchy parameters in the topological connection relationship, and map the occlusion relationship description to the visibility mask parameters in the topological connection relationship to obtain the connection mapping result.
[0125] For the adjacency description in the spatial relationship feature, for example, "the arch is adjacent to the magic town entrance," this adjacency description is mapped to the adjacency face sharing parameter in the topological connection relationship. The adjacency face sharing parameter describes the connection between two adjacent objects in a 3D model. In this case, the arch and the magic town entrance may share some faces. By setting the adjacency face sharing parameter, we can ensure that their connection in the 3D model is natural and reasonable.
[0126] For containment descriptions, such as "a fountain is surrounded by beautiful flower beds," this containment description is mapped to the nesting level parameters in the topological connection relationship. Nesting level parameters are used to determine the hierarchical relationship of one object within or around another in a 3D model. For example, a flower bed surrounding a fountain can be understood as a nested relationship outside the fountain. By setting the nesting level parameters, the hierarchical structure of these objects in 3D space can be accurately represented.
[0127] For occlusion relationship descriptions, such as "In a forest, trees sometimes obscure part of a stream," this occlusion relationship description is mapped to visibility mask parameters in the topological connection relationship. Visibility mask parameters are used to control the visibility of objects in the 3D model. In this case, the part of the stream that is obscured by trees can be made invisible in the 3D model by setting the visibility mask parameters.
[0128] Step S134: performing consistency check processing on the element classification results, attribute mapping results and connection mapping results to ensure that there is no conflict in the geometric structure, surface material attributes and topological connection relationship mapping results of the same object type feature.
[0129] After completing feature classification, attribute mapping, and connection mapping, consistency verification is required. Taking an arch as an example, in feature classification, the arch is classified as a scene prop, corresponding to an additional component element in the geometric structure. In attribute mapping, the arch's vertex coordinate parameters, texture mapping type, and scaling factor parameters are determined. In connection mapping, shared parameters are set for the adjacent faces between the arch and the magic town entrance.
[0130] The consistency check checks for conflicts between these results. For example, it checks the compatibility of vertex coordinate parameters and adjacent face sharing parameters. If the vertex coordinate settings result in an inability to properly share adjacent faces between the arch and the entrance to the magic town, the vertex coordinate parameters or adjacent face sharing parameters need to be adjusted.
[0131] Similarly, check that the texture type and scaling factor parameters are consistent with the arch's classification as a scene prop. If the texture type or scaling factor parameters are set improperly, the arch's performance in the 3D model may not be as expected, requiring correction. Through consistency checking, we ensure that the mapping results of the geometric structure, surface material properties, and topological connection relationships of the same object type are coordinated and conflict-free.
[0132] Step S135: Generate an association mapping relationship set including a feature classification correspondence table, an attribute mapping correspondence table, and a connection mapping correspondence table as the association mapping relationship.
[0133] After completing the consistency check, the feature classification results, attribute mapping results, and connection mapping results are organized into a correspondence table. The feature classification correspondence table records the correspondence between object type characteristics and geometric structure elements, such as character models corresponding to main mesh elements, scene props corresponding to additional component elements, and environmental decorations corresponding to background construction elements.
[0134] The attribute mapping table records the correspondence between morphological feature descriptions and geometric structure and surface material attribute parameters, such as the correspondence between contour shape descriptions and vertex coordinate parameters, surface texture descriptions and texture map types, and scale factor parameters. The connection mapping table records the correspondence between spatial relationship features and topological connection relationship parameters, such as the correspondence between adjacency relationships and adjacent face sharing parameters, containment relationships and nesting level parameters, and occlusion relationships and visibility mask parameters.
[0135] These three correspondence tables are integrated to generate an association mapping relationship set consisting of a feature classification correspondence table, an attribute mapping correspondence table, and a connection mapping correspondence table. This association mapping relationship set, as a whole, ensures that when generating a 3D model, the corresponding 3D model elements can be accurately constructed based on the various features in the scene feature set.
[0136] Step S140: calling a preset model generation framework based on the association mapping relationship to generate initial 3D model data corresponding to the scene feature set.
[0137] After obtaining the set of association mapping relationships, the preset model generation framework can be called based on this to generate the initial 3D model data. The preset model generation framework is a pre-built system that includes multiple modules and processing units. It can generate the corresponding 3D model based on the input association mapping relationship and scene feature set.
[0138] Step S141: extracting a feature classification correspondence table from the association mapping relationship, and determining the grid generation module, component generation module, and background generation module that need to be called in the model generation framework according to the feature classification correspondence table.
[0139] First, we extract a feature classification table from the association mapping relationship set. For example, in the medieval magic town scene, the feature classification table records the correspondence between character models and main mesh elements, scene props and additional component elements, and environmental decorations and background elements. Based on this feature classification table, we determine the modules that need to be called in the model generation framework.
[0140] For character models, such as the "brave knight," the mesh generation module is called upon. This module is responsible for generating the basic mesh structure of the 3D model. For characters like the knight, this module constructs the knight's main mesh based on information such as vertex coordinate parameters determined in the associated mapping relationship.
[0141] For scene props, such as "arches" and "fountains," the component generation module needs to be called. The component generation module will generate 3D model components of these scene props based on the attribute mapping results of the scene props, including vertex coordinate parameters, texture mapping type, and scaling coefficient parameters.
[0142] For environmental decorations, such as flowers and fallen leaves, the background generation module is called. The background generation module generates background construction data for creating the scene atmosphere based on the attributes and spatial relationships of the environmental decorations.
[0143] Step S142: extracting an attribute mapping correspondence table from the association mapping relationship, and inputting the vertex coordinate parameters, texture mapping type and scaling coefficient parameters in the attribute mapping correspondence table into a corresponding module as generation parameters.
[0144] After extracting the attribute mapping table from the associated mapping relationship set, the vertex coordinate parameters, texture map type, and scaling factor parameters are input into the corresponding modules. For the mesh generation module, the vertex coordinate parameters corresponding to the character model are input. These parameters determine the shape and size of the character's main mesh. For example, for a knight model, the vertex coordinate parameters determine the position of each part of the knight's body in 3D space, thereby constructing the knight's basic appearance.
[0145] The component generation module takes the scene prop's vertex coordinate parameters, texture mapping type, and scaling factor as input. For example, for an arch, the vertex coordinate parameters determine its shape, the texture mapping type determines its surface appearance, and the scaling factor adjusts its size. By inputting these parameters into the component generation module, a 3D arch model component that matches the scene's characteristics can be generated.
[0146] The background generation module inputs the attribute parameters of the environmental decorations. For example, for a flower, the vertex coordinate parameters determine the flower's position in 3D space, the texture mapping type determines the flower's appearance, and the scaling factor parameter adjusts the flower's size, thereby generating realistic flower background data.
[0147] Step S143: extracting a connection mapping table from the association mapping relationship, and inputting the adjacent surface sharing parameters, nesting level parameters and visibility mask parameters in the connection mapping table into the connection control unit between modules.
[0148] After extracting the connection mapping table from the association mapping relationship set, the adjacent face sharing parameters, nesting level parameters, and visibility mask parameters are input into the inter-module connection control unit. The connection control unit is responsible for coordinating the connection and hierarchical relationship between the data generated by each module.
[0149] For adjacent face sharing parameters, such as the adjacent face sharing parameters of the arch and the magic town entrance, the connection control unit will adjust the 3D model data of the arch and the magic town entrance according to the parameters to ensure that the adjacent faces between them can be reasonably shared, so that the connection between the two looks natural in the 3D model.
[0150] For nested hierarchical parameters, such as the nested hierarchical parameters of fountains and flower beds, the connection control unit will determine the hierarchical relationship of the flower beds around the fountains based on the nested hierarchical parameters, and adjust the spatial position of the 3D model data of the flower beds and fountains to accurately present their hierarchical structure.
[0151] For visibility mask parameters, such as the visibility mask parameters of trees blocking the stream, the connection control unit will control the visibility of the trees and the stream in the 3D model according to the visibility mask parameters, so that the stream is not visible in the 3D model in the part blocked by the trees.
[0152] Step S144: generating main body mesh data, additional component data and background construction data respectively through the mesh generation module, component generation module and background generation module.
[0153] After inputting the corresponding parameters into each module, the mesh generation module begins generating the main mesh data. For example, for a knight, the mesh generation module constructs the main mesh of the knight based on the input vertex coordinate parameters. It begins with vertex coordinates for various parts of the knight, such as the head, torso, and limbs, and gradually connects these vertices to form the knight's basic mesh structure. During this process, vertex coordinates may be dynamically adjusted based on information such as the knight's posture and movements to generate the main mesh data for the knight in different poses.
[0154] The component generation module generates additional component data based on the input parameters. For example, for an arch, it can determine the arch's shape based on vertex coordinate parameters, apply a texture corresponding to the texture mapping type to the arch's surface, and adjust the arch's size based on the scaling factor parameter, ultimately generating the complete arch's 3D model component data.
[0155] The background generation module generates background construction data based on input parameters. For environmental decorations such as flowers and fallen leaves, it determines their position in 3D space based on vertex coordinate parameters, applies textures corresponding to the texture map type to their surfaces, adjusts their size based on the scaling factor parameter, and then integrates this data to generate the background construction data used to create the scene atmosphere.
[0156] Step S145: performing spatial position assembly processing on the main body mesh data, additional component data and background construction data according to the adjacent surface sharing parameters, nesting level parameters and visibility mask parameters through the connection control unit.
[0157] The connection control unit will perform spatial position assembly processing on the main mesh data, additional component data and background construction data based on the adjacent surface sharing parameters, nesting level parameters and visibility mask parameters.
[0158] Step S1451: extracting a set of faces that need to share vertices from the main body mesh data and the additional component data according to the adjacent face sharing parameters.
[0159] Taking the knight and his weapon (an additional component) as an example, based on the adjacent face sharing parameters, the connection control unit searches for face sets that need to share vertices in the main mesh data (the knight) and the additional component data (the weapon). For example, the portion of the weapon held in the knight's hand may have some faces that need to share vertices with faces in the knight's hand to ensure a natural and tight connection. By analyzing the adjacent face sharing parameters, these face sets that need to share vertices are accurately extracted.
[0160] Step S1452: performing vertex coordinate alignment processing on the face set so that the coordinate values of shared vertices remain consistent in the main mesh data and the additional component data.
[0161] After extracting the face sets that need to share vertices, these face sets are then aligned. For the knight and weapon example above, the connection control unit compares the coordinate values of the shared vertices in the main mesh data (knight) and the attached component data (weapon). If there are discrepancies, the coordinate values are adjusted to ensure that the shared vertex coordinates are consistent between the two. This ensures that the connection between the knight and the weapon in the 3D model is seamless and seamless, resulting in a natural connection.
[0162] Step S1453: Determine the nesting depth level of the additional component data in the main mesh data and the corresponding spatial offset according to the nesting level parameter.
[0163] For example, regarding the nesting level parameters for a fountain and a flower bed, the connection control unit determines the nesting depth of the flower bed around the fountain based on the nesting level parameters. If the nesting level parameters indicate that the flower bed is nested outside the fountain, its nesting depth level is determined. This nesting relationship also determines the corresponding spatial offset—the positional offset of the flower bed relative to the fountain in 3D space. By analyzing the nesting level parameters, the nesting depth level and spatial offset of the additional component data (the flower bed) within the main mesh data (the fountain) are accurately calculated.
[0164] Step S1454: performing spatial translation transformation processing on the additional component data, and adjusting its position in the main mesh data according to the nesting depth level and the spatial offset.
[0165] After determining the nesting depth level and spatial offset of the additional component data, the connection control unit performs a spatial translation transformation on the additional component data. For the flower bed, the 3D model data is translated in 3D space based on the calculated nesting depth level and spatial offset. This ensures that the flower bed is precisely positioned around the fountain, conforming to the pre-set nesting hierarchy and presenting a reasonable spatial layout in the 3D model.
[0166] Step S1455: extracting the area range of the background construction data that needs to block the main body mesh data or additional component data according to the visibility mask parameters.
[0167] For example, if a tree in a forest obscures a stream, the connection control unit extracts the area of the background build data (the tree) that needs to be obscured by the main mesh data or additional component data (the stream) based on the visibility mask parameters. By analyzing the visibility mask parameters, the specific location and extent of the tree obscuring the stream in 3D space are determined.
[0168] Step S1456: Perform visibility marking processing on the background construction data, mark the area range as an occlusion area and set corresponding transparency parameters.
[0169] After extracting the range of the occluded area, the connection control unit will perform visibility marking on the background construction data. For the above example of trees blocking the stream, the area of the trees corresponding to the occluded stream is marked as the occluded area. At the same time, the corresponding transparency parameters are set. When observing the 3D model from a specific perspective, the part marked as the occluded area will be displayed according to the transparency parameters, reflecting the occlusion effect. For example, setting a certain transparency parameter makes the part of the stream blocked by the trees appear translucent or completely invisible, enhancing the realism of the 3D model.
[0170] Step S1457: spatially superimpose the processed main body mesh data, additional component data, and background construction data to generate assembled model data.
[0171] After processing the main mesh data, additional component data, and background construction data, the connection control unit spatially overlays these processed data sets. The knight and weapon data, which have undergone vertex coordinate alignment, the fountain and flower bed data, which have undergone spatial translation transformations, and the tree and stream data, which have undergone visibility markings, are overlaid according to their position and hierarchical relationships in 3D space. This spatial overlay generates a complete, assembled model with interconnected components and backgrounds and a well-organized layout.
[0172] Step S146: The assembled model data is formatted in a unified manner to generate initial 3D model data having a unified coordinate system and data storage format.
[0173] After obtaining the assembled model data, it needs to be formatted uniformly. The data generated by different modules may use different coordinate systems and data storage formats. In order to facilitate subsequent use and processing, this data needs to be unified.
[0174] First, coordinate systems are unified. For example, the mesh generation module, component generation module, and background generation module may each use their own local coordinate systems when generating data. The connection control unit converts these local coordinate systems into a unified global coordinate system. By transforming the vertex coordinates of each data set, all data is represented in the same coordinate system, facilitating subsequent spatial analysis and rendering.
[0175] The data storage format is then unified. Data generated by different modules may be stored in different formats, such as binary or text. The connection control unit converts this data into a unified data storage format, such as a common 3D model file format. This way, the generated initial 3D model data has a unified coordinate system and data storage format, making it easier to use in subsequent game development.
[0176] Step S150: performing structural adjustment processing on the initial 3D model data according to the model adaptation constraint conditions of game development to obtain a final 3D model output result that meets the development requirements.
[0177] After generating the initial 3D model data, it is necessary to adjust its structure according to the model adaptation constraints of the game development to meet the specific needs of the game development.
[0178] For example, step S151: obtaining model adaptation constraints for game development, wherein the model adaptation constraints include polygon number restrictions, material resolution restrictions, and collision detection accuracy requirements.
[0179] Game development involves a series of model adaptation constraints, determined by factors such as game performance and platform requirements. Regarding polygon count limits, different gaming platforms and devices have varying tolerances for the number of polygons in 3D models. For example, mobile devices may have stricter polygon limits, as excessive polygon counts can lead to performance degradation and lag during game runtime.
[0180] Texture resolution limits are also a significant constraint. If the texture resolution is too high, it will take up a significant amount of memory and storage space, impacting game loading speed and performance. Therefore, it's necessary to limit the texture resolution to ensure smooth game performance.
[0181] Collision detection accuracy requirements are designed to ensure accurate and reliable collision detection between objects in the game. In some action games, high-precision collision detection can provide a more realistic gaming experience. If collision detection accuracy is insufficient, unreasonable phenomena such as object penetration may occur.
[0182] Step S152: performing polygon number statistical processing on the initial 3D model data to obtain polygon number distribution of the main body mesh data, additional component data and background construction data.
[0183] When polygon count processing is performed on the initial 3D model data, statistics are performed separately for the main mesh data, additional component data, and background construction data. For example, taking a knight, his equipment, and the surrounding environment, the connection control unit traverses each face in the main mesh data (the knight) and counts its polygon count. Similarly, polygon counts are performed for additional component data (such as weapons and armor) and background construction data (such as buildings, plants, and other objects in the scene).
[0184] Through the above statistical processing, we can obtain the polygon count distribution of each part. For example, we may find that the main mesh data of the knight has a large number of polygons, while some background construction data (such as flowers and plants in the distance) has a relatively small number of polygons.
[0185] Step S153: simplifying the component data with excessive polygon count according to the polygon count limit, and reducing the polygon count by vertex merging or face collapsing operations.
[0186] Based on polygon count statistics and polygon limit, component data with excessive polygon counts is simplified. For example, if the knight's main mesh exceeds the limit, the connection control unit will use vertex merging or face collapse operations.
[0187] Vertex merging combines adjacent, closely spaced vertices into a single vertex. In the knight's main mesh data, there may be some tiny details where the distances between adjacent vertices are very small, making little difference to the overall visual quality. By merging these adjacent vertices, the number of polygons can be reduced.
[0188] The face collapse operation merges or deletes smaller faces. The connection control unit collapses or deletes smaller faces in the knight's main mesh data that don't significantly impact the overall shape, reducing the polygon count. These simplification operations bring the knight's main mesh data within the polygon limit.
[0189] Step S154: performing resolution detection processing on the material map of the initial 3D model data to obtain actual resolution parameters of each texture map in the surface material attributes.
[0190] When performing resolution detection on the texture maps of the initial 3D model data, the connection control unit checks the actual resolution parameters of each texture map in the surface material properties. For texture maps such as the knight's armor and weapon, as well as the building textures in the scene, the actual resolution information is obtained.
[0191] By analyzing the pixel size and other information of the texture map, its actual resolution parameters are determined. For example, it is found that the texture map resolution of the knight's armor is too high, exceeding the material resolution limit.
[0192] Step S155: down-sampling the texture map with a resolution exceeding the limit according to the material resolution limit to generate a simplified texture map that meets the resolution requirement.
[0193] For textures with excessive resolution, the connection control unit downsamples them. For example, downsampling reduces the number of pixels in a texture to reduce its resolution, such as a high-resolution texture of a knight's armor.
[0194] You can use average sampling to merge multiple adjacent pixels in a texture map into a single pixel, and then calculate the color value of the merged pixel based on their color values. This method gradually reduces the number of pixels in the texture map and lowers its resolution. After downsampling, a simplified texture map is generated that meets the material's resolution limit, preserving the texture's essential characteristics while reducing memory and storage space usage.
[0195] Step S156: performing collision detection area identification processing on the initial 3D model data to extract the main mesh data area that needs to be subjected to collision detection.
[0196] In the game, not all areas of the 3D model require collision detection. The connection control unit identifies collision detection areas based on the initial 3D model data. For example, in a battle scene between a knight and an enemy, the areas requiring collision detection are primarily the knight's weapon, key body parts, and the corresponding enemy parts.
[0197] By analyzing the game rules and scenarios, the connection control unit extracts the main mesh data areas that require collision detection. For example, for a knight's sword, the main body of the sword will be used as the collision detection area; for the knight's body, key areas such as the chest and head will be used as collision detection areas.
[0198] Step S157: performing vertex encryption processing on the region according to the collision detection accuracy requirement, and increasing its vertex distribution density to enhance detection accuracy.
[0199] According to the collision detection accuracy requirements, the extracted collision detection area is subjected to vertex encryption processing. Taking the knight's sword as an example, if the collision detection accuracy requirement is high, the connection control unit will increase the distribution density of the vertices in the collision detection area of the main body mesh data of the sword.
[0200] The number of vertices can be increased by inserting new vertices between the original vertices, or subdividing the original faces, etc. After increasing the vertex distribution density, the collision between the sword and other objects can be more accurately judged during collision detection, enhancing the accuracy of collision detection. Similarly, the collision detection area of the key parts of the knight's body is also subjected to similar vertex encryption processing to meet the game's requirements for collision detection accuracy.
[0201] Step S158: The simplified part data, simplified texture map, and encrypted collision detection area are subjected to data integration processing to generate the final 3D model output result that meets all the model adaptation constraints.
[0202] After completing the polygon number simplification processing, material map downsampling processing, and collision detection area vertex encryption processing, the data sets after these processing need to be subjected to data integration processing.
[0203] The connection control unit integrates the part data simplified by vertex merging or face collapse operation, the simplified texture map subjected to downsampling processing, and the collision detection area data subjected to vertex encryption processing according to their positions and hierarchical relationships in the 3D space. The simplified knight main body mesh data, simplified armor and weapon texture maps, and encrypted collision detection area data of the sword and key parts of the body are integrated together. During the integration process, the data of each part is associated in a unified coordinate system, so that the position and posture of the simplified parts in the 3D space match the simplified texture map and the encrypted collision detection area.
[0204] For the simplified part data, its vertex coordinates and topological structure information are associated with other data. Taking the knight's weapon as an example, the simplified weapon part data contains processed vertex coordinates and adjusted face information. During integration, these data are matched with the simplified texture map of the weapon, so that the texture can be correctly mapped to the surface of the weapon. At the same time, the encrypted weapon collision detection area data is associated with the weapon part data to ensure accurate detection of the corresponding parts of the weapon during collision detection.
[0205] Simplified texture maps must be consistent with the corresponding component data in terms of space and texture coordinates. When a simplified texture map is applied to a component surface, the texture is correctly unwrapped and rendered based on the component's vertex coordinates and texture coordinate information. For example, for a knight's armor, the simplified texture map is accurately applied to the surface based on the texture coordinate information in the armor component data, ensuring the correct position and orientation of the texture.
[0206] The encrypted collision detection area data must be properly integrated with the overall 3D model data. During this integration process, the encrypted vertex coordinates and related detection logic are adjusted to work in conjunction with the rest of the data. For example, for key areas of a knight's body, the encrypted collision detection area data is merged with the knight's main mesh data. The collision detection algorithm can accurately identify and process these encrypted areas, improving collision detection accuracy.
[0207] Through this data integration process, the simplified component data, simplified texture maps, and encrypted collision detection area data are organically combined to generate a final 3D model output that meets all model adaptation constraints. This final 3D model not only meets the performance requirements of game development in terms of polygon count and material resolution, but also meets the interactive needs of the game in terms of collision detection accuracy. This provides high-quality 3D model resources for game development, which can be directly used in game scene construction and rendering, giving players a smoother and more realistic gaming experience.
[0208] Figure 2 A schematic diagram illustrates exemplary hardware and software components of a text-driven 3D model generation system 100 for game development that can implement the concepts of the present application, as provided in some embodiments of the present application. For example, the processor 120 can be used in the text-driven 3D model generation system 100 for game development and used to perform the functions of the present application.
[0209] The text-driven 3D model generation system 100 for game development can be a general-purpose server or a special-purpose server, both of which can be used to implement the text-driven 3D model generation method for game development of the present application. Although only one server is shown in this application, for convenience, the functions described in this application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.
[0210] For example, the text-driven 3D model generation system 100 for game development may include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as a disk, ROM, or RAM, or any combination thereof. Exemplarily, the text-driven 3D model generation system 100 for game development may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present application can be implemented according to these program instructions. The text-driven 3D model generation system 100 for game development also includes an I / O interface 150 between the computer and other input and output devices.
[0211] For ease of explanation, only one processor is described in the text-driven 3D model generation system 100 for game development. However, it should be noted that the text-driven 3D model generation system 100 for game development in the present application may also include multiple processors, so the steps performed by one processor described in the present application may also be performed jointly or individually by multiple processors. For example, if the processor of the text-driven 3D model generation system 100 for game development executes step A and step B, it should be understood that step A and step B may also be performed jointly by two different processors or individually in one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor execute steps A and B together.
[0212] In addition, an embodiment of the present invention further provides a readable storage medium, in which computer-executable instructions are preset. When a processor executes the computer-executable instructions, the above-mentioned text-driven 3D model generation method for game development is implemented.
[0213] It should be noted that in order to simplify the description of the present invention and thus help understand one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, multiple features are sometimes combined into one embodiment, figure or description thereof.
Claims
1. A text-driven 3D model generation method for game development, characterized in that: The method comprises: Obtaining a target text description set of a game design scene, wherein the target text description set includes scene element description contents in the form of multiple consecutive paragraphs; Performing key information extraction processing on the target text description set to obtain a scene feature set including object type features, morphological feature descriptions, and spatial relationship features; Establishing an association mapping relationship between each feature in the scene feature set and a 3D model element, wherein the 3D model element includes a geometric structure, surface material properties, and a topological connection relationship; Based on the association mapping relationship, a preset model generation framework is called to generate initial 3D model data corresponding to the scene feature set; Performing structural adjustment processing on the initial 3D model data according to the model adaptation constraint conditions of game development to obtain a final 3D model output result that meets the development requirements; The establishing of an association mapping relationship between each feature in the scene feature set and a 3D model element, wherein the 3D model element includes a geometric structure, surface material properties, and a topological connection relationship, includes: Performing element classification processing on the object type features in the scene feature set, mapping the character model to the main mesh element in the geometric structure, mapping the scene props to the additional component elements in the geometric structure, and mapping the environmental decorations to the background construction elements in the geometric structure, to obtain an element classification result; Performing attribute mapping processing on the morphological feature descriptions in the scene feature set, mapping the contour shape description to vertex coordinate parameters of the geometric structure, mapping the surface texture description to the texture map type of the surface material attribute, and mapping the scale size description to the scaling factor parameter of the geometric structure, to obtain an attribute mapping result; Performing connection mapping processing on the spatial relationship features in the scene feature set, mapping the adjacency relationship description to the adjacent surface sharing parameter in the topological connection relationship, mapping the inclusion relationship description to the nesting level parameter in the topological connection relationship, and mapping the occlusion relationship description to the visibility mask parameter in the topological connection relationship, to obtain a connection mapping result; Performing consistency check on the feature classification results, attribute mapping results, and connection mapping results to ensure that there is no conflict in the geometric structure, surface material attributes, and topological connection relationship mapping results of features of the same object type; Generate an association mapping relationship set including a feature classification corresponding table, an attribute mapping corresponding table, and a connection mapping corresponding table as the association mapping relationship; The calling of a preset model generation framework based on the association mapping relationship to generate initial 3D model data corresponding to the scene feature set includes: Extracting a feature classification correspondence table from the association mapping relationship, and determining the grid generation module, component generation module, and background generation module that need to be called in the model generation framework according to the feature classification correspondence table; Extracting an attribute mapping table from the association mapping relationship, and inputting vertex coordinate parameters, texture mapping type and scaling coefficient parameters in the attribute mapping table into a corresponding module as generation parameters; Extracting a connection mapping table from the association mapping relationship, and inputting the adjacent surface sharing parameters, nesting level parameters and visibility mask parameters in the connection mapping table into the connection control unit between modules; Generate main body grid data, additional component data and background construction data respectively through the grid generation module, component generation module and background generation module; Performing spatial position assembly processing on the main body mesh data, additional component data and background construction data by the connection control unit according to the adjacent surface sharing parameters, nesting level parameters and visibility mask parameters; The assembled model data is formatted in a unified manner to generate initial 3D model data with a unified coordinate system and data storage format.
2. The text-driven 3D model generation method for game development according to claim 1, characterized in that: The key information extraction process is performed on the target text description set to obtain a scene feature set including object type features, morphological feature descriptions and spatial relationship features, including: Performing semantic segmentation on the target text description set to divide the description content in the form of continuous paragraphs into multiple independent semantic units; Performing entity recognition processing on each of the independent semantic units to extract object type features representing specific game elements, wherein the object type features include character models, scene props, and environmental decorations; Performing attribute description extraction processing on each of the independent semantic units to extract a morphological feature description describing the appearance of the object, wherein the morphological feature description includes a contour shape description, a surface texture description, and a scale size description; Performing relationship analysis on each of the independent semantic units to extract spatial relationship features describing positional connections between objects, wherein the spatial relationship features include adjacency relationship description, inclusion relationship description, and occlusion relationship description; The object type features, morphological feature descriptions and spatial relationship features are subjected to feature integration processing according to the original order of semantic units to generate a scene feature set with context association.
3. The text-driven 3D model generation method for game development according to claim 2, characterized in that: The performing entity recognition processing on each of the independent semantic units to extract object type features representing specific game elements therein includes: Performing vocabulary filtering on the independent semantic units to remove non-entity words representing degree modification or emotional description; Perform domain dictionary matching on the remaining words and classify the matched words into object type features corresponding to character models, scene props or environmental decorations; Context inference is performed on words that are not matched to the domain dictionary, and the corresponding object type features are determined based on the object type feature association relationship of the preceding and following semantic units. Merge duplicates of all identified object type features and retain the first occurrence position information of each object type feature; An object type feature set including an object type name and occurrence position information is generated as the object type feature.
4. The text-driven 3D model generation method for game development according to claim 2, characterized in that: The performing of attribute description extraction processing on each of the independent semantic units to extract the morphological feature description describing the appearance of the object includes: performing adjective phrase recognition processing on the independent semantic units to extract adjective phrases used to describe shape, texture, and size; Matching the adjective phrase with a preset morphological description dictionary to determine the corresponding contour shape type, surface texture type, and proportional size type; Performing semantic conversion processing on the comparative description contained in the adjective phrase to convert the comparative word into a specific morphological reference object; Performing association and binding processing on the outline shape type, surface texture type, scale size type and morphological reference object to generate a morphological feature description corresponding to the object type feature in the independent semantic unit; All generated morphological feature descriptions are subjected to ambiguity elimination processing, and the contradictory morphological feature descriptions are corrected through the consistency check of the morphological descriptions of the context semantic units.
5. The text-driven 3D model generation method for game development according to claim 3, characterized in that: The performing of relationship analysis on each of the independent semantic units to extract spatial relationship features describing positional connections between objects includes: performing prepositional phrase recognition processing on the independent semantic units to extract prepositional phrases representing positional connections; Matching the prepositional phrase with a preset spatial relationship dictionary to determine the corresponding adjacency relationship, inclusion relationship, or occlusion relationship type to obtain the spatial relationship type; Performing location processing on the two object type features involved in the prepositional phrase to determine the specific objects referred to by the two object type features in the object type feature set; Performing triple construction processing on the spatial relationship type and the specific referred object to generate a spatial relationship triple including the first object, the spatial relationship type and the second object; All generated spatial relationship triples are subjected to logical conflict detection. By checking whether there are conflicting spatial relationship types between the same pair of objects, the conflicting spatial relationship features are corrected.
6. The text-driven 3D model generation method for game development according to claim 1, characterized in that: The attribute mapping process is performed on the morphological feature description in the scene feature set, mapping the contour shape description to the vertex coordinate parameters of the geometric structure, mapping the surface texture description to the texture map type of the surface material attribute, and mapping the scale size description to the scaling factor parameter of the geometric structure, including: Performing geometric feature extraction processing on the contour shape description to identify shape components representing straight lines, curves or curved surfaces; Converting the shape constituent elements into parameters of the change rules of vertex coordinates in a geometric structure to generate an initial setting scheme for the vertex coordinate parameters; Performing texture feature extraction on the surface texture description to identify texture elements representing smoothness, roughness or patterning; Matching the texture representation elements with a preset texture mapping library to determine the texture mapping type of the corresponding surface material attribute; Performing relative relationship extraction processing on the proportional size description to identify the size ratio elements representing the whole and the parts; Converting the size ratio elements into scaling coefficient parameters of vertex coordinates of each component in the geometric structure, and generating an initial setting scheme for the scaling coefficient parameters; The initial setting scheme of the vertex coordinate parameters, the texture mapping type and the initial setting scheme of the scaling coefficient parameters are combined to generate the result of the attribute mapping processing.
7. The text-driven 3D model generation method for game development according to claim 1, characterized in that: The process of performing spatial position assembly processing on the main body mesh data, the additional component data and the background construction data by the connection control unit according to the adjacent surface sharing parameters, the nesting level parameters and the visibility mask parameters includes: Extracting a face set that needs to share vertices from the main body mesh data and the additional component data according to the adjacent face sharing parameter; Performing vertex coordinate alignment processing on the face set so that coordinate values of shared vertices remain consistent in the main mesh data and the additional component data; Determining the nesting depth level of the additional component data in the main body mesh data and the corresponding spatial offset according to the nesting level parameter; Performing spatial translation transformation processing on the additional component data, and adjusting its position in the main grid data according to the nesting depth level and the spatial offset; Extracting the area range of the background construction data that needs to block the main body mesh data or additional component data according to the visibility mask parameters; Performing visibility marking processing on the background construction data, marking the area range as an occlusion area and setting corresponding transparency parameters; The processed main body mesh data, additional component data and background construction data are spatially superimposed to generate assembled model data.
8. A text-driven 3D model generation system for game development, characterized in that: It includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the text-driven 3D model generation method for game development as described in any one of claims 1 to 7 above.
Citation Information
Patent Citations
Three-dimensional model intelligent generation method and system based on natural language
CN120147551A