Intelligent content generation method and device based on large language model and medium

Through the cross-modal semantic analysis and multi-modal network linkage of large language models, the semantic deviation problem in the generation of multi-subject interactive images is solved. The generated images are in line with scientific laws at both the microscopic and macroscopic levels, ensuring that the image is consistent with the text description, and improving the logical rationality and detailed expression of the generated images.

CN120374776APending Publication Date: 2025-07-25浪潮智能终端有限公司
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510508995.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In complex image scenarios where multiple interactive subjects are generated, the prior art has semantic deviations in multi-subject interactions and mismatch of prompt word expressions and image understanding, resulting in the generated images being unable to accurately match text descriptions and unable to meet user needs.

Method used

Cross-modal semantic analysis is performed through large language models, multi-dimensional semantic embedding vectors are generated, including biological features, environmental elements and physical parameters, dynamically complement the missing interactive parameters, use the multi-modal control network to generate the initial image content, and link it with the multi-modal network through hierarchical encoding to ensure that the image content complies with scientific laws.

Benefits of technology

It significantly reduces generation distortion, ensures that the image is highly consistent with the text description, improves the overall consistency and logical rationality of the image, and realizes the full expression of detailed information and the multi-subject interaction relationship in accordance with design expectations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374776A_ABST
    Figure CN120374776A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an intelligent content generation method and device based on a large language model and a medium, and relates to the technical field of large language models.The method comprises the steps that text description information input by a user is obtained, cross-modal semantic analysis is conducted on the text description information through the large language model, and a multi-dimensional semantic embedding vector is generated; the multi-dimensional semantic embedding vector comprises a biological feature semantic embedding vector, an environment element semantic embedding vector and a physical parameter semantic embedding vector; on the basis of a multi-dimensional semantic embedding vector, carrying out dynamic reasoning complementation on missing interaction parameters, and generating optimized structured description data; and according to the structured description data and a preset multi-mode control network, generating initial image content, verifying the initial image content, and determining target image content. Potential information of multi-subject interaction is automatically complemented and inferred, it is ensured that the generated image and text description are kept highly consistent, and the problem of generation distortion caused by semantic errors is remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the technical field of large language models, and in particular, to an intelligent content generation method, device, and medium based on large language models. Background Art

[0002] In the context of the rapid development of generative artificial intelligence technology, text-to-image and image-to-image technologies have made significant progress in single-subject image generation tasks. Currently, text-to-image technology mainly relies on diffusion models, generative adversarial networks, etc., to generate corresponding images from text descriptions through deep learning models, and generate high-quality images through natural language input.

[0003] Traditional diffusion models establish text-image mapping through attention mechanisms. However, when faced with multiple interacting subjects, it is difficult for the model to analyze implicit spatial logical relationships and action coordination. When directly generating complex scenes, such as generating real-scene images with complex biological features and special material accessories, it is difficult to ensure the consistency between the image content and the text description, and there may be problems of semantic deviation. Especially when involving multiple subjects, complex environments, etc., the generated images may not accurately match the text description, and the generated images do not fully meet the user's expectations. For example, when inputting "A cute orange kitten is playing with a wool ball in a luxurious white castle", the model may generate problems such as the wrong color of the kitten, the wrong "castle" environment, and the lack of the element "wool ball". In addition, due to the ambiguity of natural language, small changes in the prompt words may lead to huge differences in the generation results. For example, inputting "A dog wearing sunglasses" may generate the correct image, but if it is changed to "A dog wearing fashionable sunglasses", the model may wrongly adjust the style, size, or position of the sunglasses, or even change the breed of the dog.

[0004] Therefore, in the scenario of generating complex images with multiple interacting subjects, there are problems of semantic deviation in multi-subject interaction and mismatch between prompt word expression and image understanding, resulting in the generated images not being able to accurately match the text description and not meeting the user's needs. Summary of the Invention

[0005] One or more embodiments of this specification provide an intelligent content generation method, device, and medium based on large language models to solve the following technical problems: In the scenario of generating complex images with multiple interacting subjects, there are problems of semantic deviation in multi-subject interaction and mismatch between prompt word expression and image understanding, resulting in the generated images not being able to accurately match the text description and not meeting the user's needs.

[0006] One or more embodiments of this specification adopt the following technical solutions:

[0007] One or more embodiments of this specification provide an intelligent content generation method based on a large language model. The method includes: obtaining text description information input by a user, performing cross-modal semantic parsing on the text description information through the large language model to generate a multi-dimensional semantic embedding vector, where the multi-dimensional semantic embedding vector includes a biometric semantic embedding vector, an environmental factor semantic embedding vector, and a physical parameter semantic embedding vector; based on the multi-dimensional semantic embedding vector, dynamically inferring and completing missing interaction parameters to generate optimized structured description data; and generating initial image content according to the structured description data and a preset multi-modal control network, and verifying the initial image content to determine target image content.

[0008] One or more embodiments of this specification provide an intelligent content generation device based on a large language model, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above method.

[0009] A non-volatile computer storage medium provided by one or more embodiments of this specification stores computer-executable instructions, and the computer-executable instructions are set to: execute the above method.

[0010] The above at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects: Through the embodiments of this specification, by means of the cross-modal attention mechanism of the large language model, the natural language description is decoupled into three semantic embedding vectors of biometric features, environmental elements, and physical parameters, eliminating semantic deviation problems such as subject loss and attribute confusion caused by text-image mapping dimension collapse in traditional methods. It not only analyzes the anatomical features of biological subjects, but also quantifies environmental optical parameters and accessory mechanical properties, making the generated images conform to scientific laws at both the microscopic and macroscopic levels, ensuring a high degree of consistency between the generated images and the text descriptions, and significantly reducing the generation distortion problems caused by semantic errors in traditional methods; accurately locates the missing biomechanical parameters, material physical properties, and environmental interaction rules in the semantic embedding vectors, dynamically complements the implicit constraints that are difficult to explicitly define in traditional methods, making the generated content conform to scientific laws in dimensions such as anatomical rationality, material durability, and environmental dynamic response, automatically maps unquantified descriptions to computable ecological parameters, and avoids the generation randomness caused by language ambiguity in traditional methods; through dynamic prompt optimization and text intelligent supplementation, automatically completes and infers potential information of multi-subject interactions, realizes the full expression of detailed information, ensures that the interaction relationships and visual display effects between the subjects in the image can meet the design expectations, and improves the overall consistency and logical rationality of the image; through hierarchical coding and multi-modal network linkage, converts the physical laws implicit in natural language, such as biomechanics, material deformation, and environmental interaction, into quantifiable and verifiable mathematical constraints, breaking through the limitations of traditional generation models that rely on data fitting. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:

[0012] Figure 1 is a schematic flowchart of an intelligent content generation method based on a large language model provided by the embodiments of this specification;

[0013] Figure 2 is an example diagram of the generated target image content provided by the embodiments of this specification;

[0014] Figure 3 is a schematic structural diagram of an intelligent content generation device based on a large language model provided by the embodiments of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of them. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.

[0016] The embodiments of this specification provide an intelligent content generation method based on a large language model. It should be noted that the execution subject in the embodiments of this specification can be a server or any device with data processing capabilities. Figure 1 As shown in the flowchart of an intelligent content generation method based on a large language model provided by the embodiments of this specification, Figure 1 it mainly includes the following steps:

[0017] Step S101, obtain the text description information input by the user, and perform cross-modal semantic parsing on the text description information through the large language model to generate a multi-dimensional semantic embedding vector.

[0018] Among them, the multi-dimensional semantic embedding vector includes a biometric semantic embedding vector, an environmental factor semantic embedding vector, and a physical parameter semantic embedding vector;

[0019] In an embodiment of this specification, the text description information input by the user is obtained. For example, the user inputs "A white rabbit wears sunglasses and stops in a flower field" or "A cute orange kitten is playing with a wool ball in a luxurious white castle". In the embodiments of this specification, taking "A white rabbit wears sunglasses and stops in a flower field" as an example, the intelligent content generation method based on a large language model will be described.

[0020] Perform cross-modal semantic parsing on the text description information through the large language model to generate a multi-dimensional semantic embedding vector, which specifically includes: parsing the text description information through the cross-modal attention mechanism of the large language model to determine multi-dimensional parsing features, where the multi-dimensional parsing features include anatomical features, kinematic parameters, and environmental optical characteristics; mapping the multi-dimensional parsing features to determine a multi-dimensional semantic embedding vector including bone key point constraints, material elastic modulus, and ecological distribution density.

[0021] In one embodiment of this specification, through the cross-modal attention mechanism of the large language model, the input text "A white rabbit wearing sunglasses pauses in a sea of flowers" is mapped into a mechatronic description containing biological parameters. Here, the large language model can adopt the DeepSeek-v3 model. DeepSeek has the advantage of large-scale pre-training, can accurately understand the key content in the text, including information such as the main object, background, lighting, style, etc., and generate corresponding semantic embedding vectors.

[0022] First, identify the anatomical features of the white rabbit (body length 38 cm ± 2%, ear length ratio 1:1.25), and establish a four-legged standing model in combination with kinematic parameters to emphasize the balanced force on the front and rear limbs and the stability of static balance. The center of gravity stability coefficient is 1.05 when standing. For the unconventional element of "animals wearing sunglasses", start the biomechanical adaptation protocol, calculate the contact stress distribution between the rabbit ear cartilage and the frame clamping surface, with a peak pressure of 18 kPa, and generate wearing stability parameters that meet the ISO 12312-1 standard. The environmental analysis module constructs a purple flower sea model based on the botanical database, with a distribution density of 45 plants / m 2 , and the diameter of a single inflorescence is about 1.5 - 2 cm. The optical properties of the petals can set the subsurface scattering coefficient to 0.35 according to the light scattering model of the spherical inflorescence, and the light transmittance in the wavelength range of 410 - 430 nm is about 30 - 35%. It should be noted that the anatomical features refer to the biological entities in the input text. In the example of "A white rabbit wearing sunglasses pauses in a sea of flowers", the biological entity corresponds to the white rabbit. The kinematic parameters are the parameters of the actions corresponding to the biological entity and the mechanical characteristics between the biological entity and the material, which are used to represent the motion state of the biological entity and the interaction between the biological entity and the material. In this example, they are standing and wearing sunglasses. The environmental optical characteristics correspond to the characteristics of the environmental elements in the input text. In this example, they are flowers. In addition, in the multi-dimensional semantic embedding vector obtained, which includes skeletal key point constraints, material elastic modulus, and ecological distribution density, the skeletal key point constraints correspond to the biological entity, the material elastic modulus corresponds to the material, such as non-biological entities like accessories and toys, and the environmental optical characteristics correspond to the environmental elements, such as flowers and grass.

[0023] Through the above technical solution, by means of the cross-modal attention mechanism, the natural language description is accurately mapped to the parameterized space of the biological-physical-environment linkage, solving the semantic collapse problem in traditional text-to-image technologies. For example, when the input is "A white rabbit wearing sunglasses stops in a flower field", the model not only analyzes the anatomical features of the biological subject, but also quantifies the environmental optical parameters and the mechanical properties of the accessories, making the generated image conform to scientific laws at both the microscopic and macroscopic levels; the generated semantic embedding vector contains the bone key point constraint matrix and the material elastic modulus tensor, providing computable physical rules for multi-subject interaction in complex scenes. When generating the scene of "A child playing frisbee with a pet dog", the throwing trajectory of the frisbee (initial velocity 18m / s ± 5%) is dynamically constrained by the kinematic parameters, the limb extension angle of the dog (knee joint ≤ 120°), and the center of gravity offset of the child (projection area ≥ foot contact surface 85%), eliminating common problems such as limb penetration and motion trajectory mismatch in traditional methods, and effectively improving the rationality of multi-subject interaction.

[0024] Step S102, based on the multi-dimensional semantic embedding vector, dynamically infer and complete the missing interaction parameters to generate optimized structured description data.

[0025] Based on the multi-dimensional semantic embedding vector, the missing interaction parameters are dynamically inferred and complemented to generate optimized structured description data, which specifically includes: performing fuzzy information recognition and element deconstruction processing on the multi-dimensional semantic embedding vector to determine the corresponding target complement entity, where the target complement entity includes any one or more of biological entities, environmental entities, and material entities; through a knowledge inference engine, performing interaction parameter complementation on the target complement entity to determine the biomechanical adaptation parameters that are not clearly specified, so as to generate an inference-complemented semantics; adopting a hierarchical coding strategy to perform hierarchical coding on the inference-complemented semantics to determine multiple hierarchical coding information, so as to determine the structured description data, where the hierarchical coding information includes a biokinetics layer, a material optics layer, and an ecological interaction layer. Performing fuzzy information recognition and element deconstruction processing on the multi-dimensional semantic embedding vector to determine the corresponding target complement entity specifically includes: using dependency syntactic analysis to parse the multi-dimensional semantic embedding vector to extract multiple target entities, where the target entities include biological subjects, environmental elements, and material characteristics; analyzing the description of each target entity to determine whether there are missing descriptions and unquantified descriptions for the target entity, and if so, determining it as a target complement entity. Through a knowledge inference engine, performing interaction parameter complementation on the target complement entity to determine the biomechanical adaptation parameters that are not clearly specified specifically includes: based on the bone key point constraints in the multi-dimensional semantic embedding vector, determining the biokinematic model corresponding to the biological characteristics, and through the biokinematic model, determining the biodynamic adaptation parameters to determine the biokinetics layer description data; according to the material elastic modulus in the multi-dimensional semantic embedding vector, constructing the finite element stress distribution data corresponding to the material contact surface, and based on the finite element stress distribution data, determining the material optics constraint conditions to determine the material optics layer description data; through the ecological distribution density in the multi-dimensional semantic embedding vector, generating the recursive growth structure corresponding to the environmental elements through a fractal algorithm, and establishing the contact dynamics model between the biological subject and the environmental elements to determine the ecological interaction layer description data.

[0026] In one embodiment of this specification, through the dynamic inference and hierarchical coding of the multi-dimensional semantic embedding vector, the problem of missing interaction parameters in complex scene generation is solved, and the physical rationality and logical consistency of the generated content are improved. Its core process is divided into three stages: fuzzy information recognition and element deconstruction, knowledge-driven parameter complementation, and hierarchical structured coding, forming a complete semantic optimization link.

[0027] First, deeply analyze the multi-dimensional semantic embedding vectors to identify the fuzzy or undefined semantic units therein. Through the cross-modal attention mechanism, the input information is deconstructed into three types of target entities: biological entities, environmental entities, and material entities. Analyze the anatomical structure features of organisms (such as bone topology and muscle group distribution) and kinematic parameters (such as joint range of motion and center of gravity offset threshold), extract the spatial distribution patterns of environmental entities (such as vegetation density gradient and terrain undulation coefficient) and physical environmental attributes (such as light attenuation rate and medium refractive index); quantify the physical properties of materials (such as elastic modulus and yield strength) and surface optical parameters (such as roughness and polarization characteristics). Conduct a completeness check on the description of each target entity. In the biological dimension, check whether quantitative parameters such as bone key point coordinates and joint range of motion are missing. In the environmental dimension, verify whether spatial features such as vegetation distribution density and terrain undulation coefficient are clearly defined. In the material dimension, determine whether physical properties such as elastic modulus and yield strength are complete. Based on a preset completeness threshold, such as a parameter missing rate ≥ of 20% or a confidence level < 85%, mark the entity nodes to be completed and generate a completion priority queue, which includes multiple target completion entities to be completed.

[0028] For the identified entities to be completed, activate the multi-domain knowledge inference engine to implement hierarchical parameter optimization. Based on the biological taxonomy database and the kinematic biomechanics model, deduce the unspecified dynamic interaction parameters. Construct an inverse kinematic chain based on the bone key point constraints to calculate the joint angle range and torque distribution scheme. For example, for the "standing still" action of quadruped animals, automatically complete the core parameters such as the torque distribution scheme of the four limbs joints and the contact pressure distribution curve of the foot pads. Through the material science knowledge graph and the industrial design standard library, deduce the missing physical properties. Construct a three-dimensional contact model based on the material elastic modulus to calculate the stress-strain distribution cloud map, and convert the stress data into bidirectional reflectance distribution function (BRDF) parameters through a physical rendering model. Define the coupling equation of material deformation and optical properties, such as the coating interference color shift caused by the bending of the frame. Based on the ecological distribution density parameters, use the L-system algorithm to generate the recursive growth topology of vegetation and establish the deformation interaction model between the biological foot and the vegetation stem, quantifying the bending stiffness and recovery rate. Combine the geographical ecological database and the physical simulation model to construct the dynamic response rules of environmental elements. For example, complete the deformation recovery rate of vegetation under the action of wind and the refraction and scattering coupling equation of water body and light. During the reasoning process, use the causal reasoning algorithm to ensure that the completed parameters meet the physical constraint conditions among multiple entities and eliminate the common parameter conflict problems in traditional methods.

[0029] The completed semantic information is transformed into a machine-executable structured description through a three-level coding strategy. The biokinetics layer encodes the kinematic chain model of the skeletal-muscular system, defines core parameters such as the joint degree-of-freedom constraint matrix and the inertia tensor distribution. The biokinetics layer contains kinematic and dynamic constraint conditions such as the skeletal key-point coordinate set, the joint degree-of-freedom constraint matrix, and the muscle group dynamics parameters, ensuring that the generated poses conform to anatomical laws; The material optics layer constructs a physically based rendering (PBR) material description system, including physical rendering rules such as the bidirectional reflectance distribution function (BRDF) parameter set, the definition of the material elastic modulus tensor, the surface bidirectional reflectance distribution function (BRDF) parameter set, and the stress-optical coupling equation; The ecological interaction layer describes the spatial interaction relationships among multiple agents, establishes a position repulsion field based on rigid body dynamics, an environmental growth rule library based on the L-system, and an energy conservation constraint equation. The ecological interaction layer includes spatial interaction logics such as encoding the environmental fractal structure topology, the contact dynamics model parameters, and the energy transfer constraint equation. The hierarchical coding realizes cross-level parameter coupling through tensor fusion technology. For example, the biological joint angle constraint is dynamically associated with the calculation of the ground reaction force, enabling the foot contact deformation of the virtual character to simultaneously satisfy anatomical rationality and material mechanics characteristics.

[0030] Transform the physical laws implicit in natural language into quantifiable and verifiable mathematical constraints, breaking through the limitations of empirical parameter settings in traditional generation models. Through hierarchical coding, cross-scale parameter synchronization optimization from the microscale (material molecular structure) to the mesoscale (biological motion unit) and then to the macroscale (ecosystem) is achieved. The knowledge inference engine supports online updating of the domain knowledge base, enabling the system to continuously adapt to the generation requirements of new materials and new species without retraining the basic model.

[0031] Next, take "A white rabbit wears sunglasses and stops in a sea of flowers" as an example to illustrate the above steps. Based on text semantic analysis, through the knowledge inference engine of the DeepSeek large model, the input text is dynamically optimized and expanded. For the possible ambiguity or incompleteness in the user input, such as the unspecified sunglasses style, the distribution of the sea of flowers, or the details of the stopping pose, first retrieve the biomechanics database and the industrial design standard library to automatically complete the key parameters: If the type of sunglasses is not specified, a suitable clip-on frame is matched based on the anatomical structure of the rabbit's ears. At the same time, use the causal reasoning mechanism to eliminate ambiguity. If the prompt word has multiple meanings, such as "purple sea of flowers" which can refer to various flowers, then combine the geographical knowledge base and seasonal lighting data to finally lock in the "pink-purple spherical flower clusters" as the most fitting visual feature. This strategy enables the generated images to reach engineering-level accuracy in dimensions such as subject interaction and environmental physical rationality.

[0032] The structured description generation stage adopts a hierarchical coding strategy to decompose semantic elements into three subsystems: biodynamics, material optics, and ecological interaction. The biodynamics layer contains 12 sets of standing posture parameters constrained by skeletal key points, for example, the hip joint fine-tuning angle is 85°±2° to ensure the balance of the limbs and overall static stability. The material optics layer defines the polarization filtering characteristics of the sunglasses lens, with an extinction ratio of 100:1, and the frame uses a finite element model of TR90 memory plastic with a Young's modulus of 2.1GPa. The lens color can be combined with the purple-pink gradient effect in the picture to set the interference color wavelength to 410-450nm and the thickness gradient to 150-200nm to simulate real optical coating. The material optics layer generates a flower cluster fractal structure through the L-system algorithm with an iteration depth of 6, and establishes a contact dynamics model between the rabbit's foot and the flower stem, with a stem bending stiffness of 1.8GPa and a deformation recovery rate of 92%. DeepSeek automatically constructs structured image description information based on semantic embedding. It not only identifies the key components of the image, but also extracts information such as style, color, and lighting to provide guidance for subsequent image generation.

[0033] Through the above technical solution, through dependency syntax analysis and fuzzy information recognition, the missing biomechanical parameters, material physical properties and environmental interaction rules in the semantic embedding vector are accurately located. Combined with multi-domain knowledge bases, such as biokinematic models, material science maps, and ecological fractal algorithms, implicit constraints that are difficult to explicitly define by traditional methods are dynamically completed, such as joint torque thresholds, frame contact stress peaks, and vegetation deformation recovery rates, so that the generated content conforms to scientific laws in terms of anatomical rationality, material durability, and environmental dynamic response. The hierarchical coding strategy constructs a mathematical description of the linkage between biology, materials, and the environment. The biodynamic layer transforms the key point constraints of the skeleton into the inverse kinematics solution space, and ensures that the posture generation meets the physiological limit through the rigid body dynamics model; the material optics layer defines the coupling relationship between the interference color of the lens coating and the deformation of the frame based on the finite element stress distribution data, and eliminates the material distortion problem; the ecological interaction layer quantifies the deformation interaction effect between the biological foot and the vegetation stem through the fractal growth model and the contact dynamics equation, so as to quantify the interaction effect between biological characteristics and environmental factors; real-time collaborative optimization of cross-level parameters is realized, avoiding the global mismatch problem caused by local optimization in traditional methods.

[0034] Step S103, generating initial image content according to the structured description data and the preset multimodal control network, so as to verify the initial image content and determine the target image content.

[0035] Generate initial image content based on the structured description data and a preset multimodal control network, specifically including: decomposing the structured description data into multiple hierarchical encoding information, where the hierarchical encoding information includes a biokinetics layer, a material optics layer, and an ecological interaction layer; determining the pose control network, material control network, and ecological control network in the multimodal control network; and generating the initial image content by constraining the biological motion pose through the pose control network according to the hierarchical encoding information, injecting physical optics parameters through the material control network, and constructing an environmental interaction model through the ecological control network.

[0036] In one embodiment of this specification, the input structured description data is decoupled into three types of machine-executable hierarchical encoding information. The biokinetics layer contains kinematic and dynamic constraint conditions such as a set of bone key point coordinates, a joint degree-of-freedom constraint matrix, and muscle group dynamics parameters. The material optics layer includes physical rendering rules such as defining a material elastic modulus tensor, a surface bidirectional reflectance distribution function (BRDF) parameter set, and a stress-optical coupling equation. The ecological interaction layer includes spatial action logics such as encoding the topological structure of the environmental fractal, contact dynamics model parameters, and energy transfer constraint equations. Establish the mapping relationship between the hierarchical encoding and the multimodal control network. The pose control network is dedicated to parsing the data of the biokinetics layer and generating poses through inverse kinematics algorithms and rigid body dynamics models. The material control network is used to parse the parameters of the material optics layer and drive the physically based rendering pipeline to achieve material authenticity. The ecological control network is used to process the information of the ecological interaction layer and construct a virtual physical environment with multi-agent dynamic interactions.

[0037] Load the set of bone key point coordinates, construct an inverse kinematics solution space, exclude abnormal poses that violate the joint angle threshold, and calculate the reaction force distribution of the end effector (such as the foot) of the limbs through the Newton-Euler equation to ensure that the generated pose meets the static equilibrium conditions. Convert the material elastic modulus into a stress distribution map of the finite element mesh and map it to the generated surface through normal mapping to simulate the real deformation effect. Based on the bidirectional reflectance distribution function (BRDF) parameters, calculate the gradient curve of the specular reflection highlight intensity changing with the viewing angle, and couple the subsurface scattering model to simulate the light transmission characteristics of the material. Establish a real-time correlation between the material stress and optical parameters (such as the increase in surface roughness caused by metal fatigue deformation) to achieve dynamic effects such as wear and aging. Recursively generate the vegetation fractal structure based on the L-system algorithm, control the ecological density through the iteration depth, define the deformation interaction rules between the biological foot and the ground / vegetation, and calculate the contact surface pressure distribution and deformation recovery rate, such as the grass recovering 95% of its original state 10 seconds after being pressed. Construct a rigid body collision volume and a repulsive force field to prevent object penetration (such as branches and hair interpenetration) or spatial layout mismatch during the generation process.

[0038] The output gradients of the three control networks are weighted and fused to balance the optimization objectives among the rationality of biological postures, the optical authenticity of materials, and the stability of environmental interactions. The iterative correction strategy adopts the Progressive Refinement algorithm to synchronously optimize the global scene layout and local detail features in the latent space. A lightweight physical engine is embedded in the rendering pipeline to verify in real time whether the generated content meets the constraint conditions defined by each layer of encoding, such as whether the joint torque exceeds the limit and whether the material reflectivity deviates from the BRDF parameters. The local latent space vector of the detected abnormal area is corrected to avoid wasting computing resources caused by global regeneration.

[0039] Next, take the example of "a white rabbit wearing sunglasses stops in a flower field" to illustrate the above steps. During the image generation process, an improved Stable Diffusion v2.1 model architecture is adopted to achieve the precise fusion of biological features and accessory components through a multi-modal control network. After receiving the structured description data, the model first generates a basic latent space representation (dimension 768×768), and then implements physically accurate rendering through a triple control mechanism. The anatomical model of the white rabbit (including 14 skeletal key points and 42 groups of muscle group dynamics parameters) is loaded through the posture control network to constrain the spine to remain basically upright (angle 2°±1°) in the quadruped standing posture, while ensuring balanced force on the four limbs (center of gravity balance coefficient 1.00±0.05). Then, the optical parameters of the sunglasses are injected through the material control network, including the gradient coating layer of the lens (thickness gradient 150 - 200nm, interference color wavelength 420 - 450nm) and the elastic modulus of the frame (TR90 plastic 1.8GPa), and the ear contact pressure distribution map (peak pressure ≤ 20kPa) is generated through finite element analysis; the spherical flower field model is loaded using the ecological control network, and a fractal structure (L-system iteration depth 6) is constructed based on the biological characteristics of the pink-purple spherical flowers to precisely control the scattering and reflectance ratio of the petals (about 0.35 - 0.40) and the bending stiffness of the flower stems (Young's modulus 1.2GPa). In the rendering pipeline, an anisotropic hair rendering engine is integrated. Since the main body is in a static standing state, motion blur processing is cancelled, and emphasis is placed on detail textures and static light and shadow effects. The background area uses instanced rendering technology to generate a batch of flowers (number of polygons per single plant ≥ 2500), while keeping the video memory occupancy ≤ 12GB. The chromaticity detection of the final output image shows that the contrast parameter ΔE between the white rabbit and the flower cluster is 32.7 (CIEDE2000 standard), achieving an organic unity of artistic expression and scientific authenticity.

[0040] Through the above technical solutions, by linking hierarchical coding with a multi-modal network, physical laws implicit in natural language, such as biomechanics, material deformation, and environmental interaction, are transformed into quantifiable and verifiable mathematical constraints, breaking through the limitations of traditional generative models that rely on data fitting; the hierarchical decoupling strategy effectively improves the anatomical rationality, material optical authenticity, and environmental interaction stability of the generated content, and parallelizes the processing of data at each level, effectively improving the generation speed.

[0041] Verify the initial image content to determine the target image content, specifically including: verifying the biomechanical adaptability and detecting the consistency of physical parameters of the initial image content to determine the contact adaptability verification result and the physical parameter consistency detection result; according to the contact adaptability verification result and the physical parameter consistency detection result, perform gradient correction on the latent space representation, and iteratively adjust the material parameters of the abnormal area through a physical renderer to output the target image content that meets the bio-physical joint verification.

[0042] Verify the biomechanical adaptability and detect the consistency of physical parameters of the initial image content to determine the contact adaptability verification result and the physical parameter consistency detection result, specifically including: loading a preset biological anatomy model, extracting the geometric features of the contact surface, performing finite element stress analysis based on the geometric features, and generating a contact pressure distribution map; comparing the pressure distribution map with the material yield strength threshold, marking the over-limit areas to determine the biomechanical adaptability verification result; extracting the material optical parameters from the hierarchical coding information of the structured description data, generating a theoretical reflection spectrum through ray tracing simulation; calculating the chromaticity difference between the actual reflection data corresponding to the initial image content and the theoretical reflection spectrum to verify whether the material reflection characteristics conform to physical constraints.

[0043] In one embodiment of this specification, by constructing a biological-physical joint verification system and a closed-loop optimization mechanism, double guarantees for the physical rationality and biomechanical authenticity of the generated content in complex scenarios are achieved. First, based on the biodynamics layer data (skeletal key point coordinates, joint degree-of-freedom constraints) in hierarchical encoding, a three-dimensional anatomical model is reconstructed, and the geometric features of the contact surface (such as the curvature radius of the ear cartilage, the contact area of the foot pad) are defined. According to the biokinematic model, the rigid body collision volumes of each part during the dynamic interaction process (such as the contact area between the frame clamping surface and the ear cartilage) are calculated. High-precision tetrahedral meshing is performed on the contact area, and boundary conditions such as material elastic modulus and load distribution are applied. Through an implicit finite element solver, the stress-strain distribution nephogram under static / dynamic loads is calculated, and the maximum contact pressure value (such as the peak value of the frame clamping pressure) is extracted. The calculated contact pressure distribution is compared with the material yield strength threshold (such as 25 MPa for TR90 plastic), and the over-limit areas are marked (such as ear pressure > 20 kPa). Combining with the biomechanics database, it is verified whether parameters such as joint torque and muscle stretch rate exceed the physiological limits of the species (such as knee joint torque > 18 Nm).

[0044] Obtain physical properties such as bidirectional reflectance distribution function (BRDF) parameters and coating interference color wavelength gradient from the material optical layer encoding. Based on the Monte Carlo path tracing algorithm, the theoretical reflection spectrum of the material under the preset lighting conditions (including specular highlight, subsurface scattering, and transmission effect) is simulated. The initial generated image is rendered from multiple angles, and the actual reflection intensity, chromaticity coordinates (CIE-Lab space), and polarization characteristics of each pixel are extracted. The change in the reflection characteristics of the material under the deformed state is detected (such as the shift of the coating interference color caused by the bending of the frame). Using the CIEDE2000 color difference formula, the global chromaticity difference (ΔE) between the actual reflection data and the theoretical spectrum is calculated, and the abnormal pixel area with ΔE > 5 is located. Through Stokes parameter analysis, it is verified whether the polarization filtering characteristics of the generated material are consistent with the physical constraints (such as the deviation angle of the lens polarization axis > 3°).

[0045] For the biomechanics over-limit areas (such as excessive contact pressure), calculate the Jacobian matrix of the potential space vector, and adjust the potential vector along the negative gradient direction to reduce the abnormal value. For the chromaticity difference area, generate a material optical compensation vector based on the ray tracing simulation results, and correct the potential space representation through weighted fusion. Perform Monte Carlo optimization in the abnormal area (such as the frame clamping surface), and screen the optimal combination from the feasible solution space of the material parameters (elastic modulus range, coating thickness gradient). Inject the optimized parameters into the rendering pipeline, and iteratively update the local material characteristics through Progressive Refinement while maintaining the global scene consistency. Perform the biomechanical adaptability verification and physical parameter consistency detection on the corrected image again until the double verification criteria (contact pressure≤ Threshold and ΔE ≤ 2). Record the full-process correction log (such as the number of iterations and records of key parameter adjustments), and output the target image content including the bio-physical verification report.

[0046] In the process of implementing "a white rabbit wearing sunglasses stops in a sea of flowers", a bio-physical joint verification mechanism is deployed in the intelligent optimization stage. The anatomical adaptability (contact area error < 2%) between the sunglasses frame and the auricular cartilage is detected through X-ray diffraction simulation, and the influence of local airflow on the static distribution of hair during stopping is verified using computational fluid dynamics. The dynamic regeneration technology performs potential space gradient correction (iteration step size 0.03) on the detected abnormal areas (such as petals penetrating the model), and at the same time applies a physics-based material optimizer to improve the authenticity of lens reflection (highlight intensity calibrated to 0.68 ± 0.02). Figure 2 An example diagram of the generated target image content provided by the embodiments of this specification is as Figure 2 shown. The final generated image is shown by chromaticity detection that the contrast between the white rabbit and the pink-purple flower cluster reaches ΔE = 32.7 (CIEDE2000 standard). The overall static performance and detail authenticity both meet the expected design requirements, achieving an organic unity of artistic expression and scientific truth. Through the fine-grained text semantic analysis of the large model, the key content in the input text, such as style, structure, details, etc., can be accurately extracted to ensure that the generated image is highly consistent with the text description, significantly reducing the generation distortion problem caused by semantic errors in traditional methods and improving the accuracy and controllability of image generation; it can handle more complex and diverse text descriptions, and through dynamic prompt optimization and text intelligent supplementation, the full expression of detailed information is realized, improving the user experience in intelligent content creation. Whether the user faces a simple description or a complex situation, they can obtain high-quality image output that meets expectations; through the refined text parsing and dynamic prompt optimization mechanism, the potential information of multi-agent interaction is automatically completed and inferred to ensure that the interaction relationship and visual display effect between the agents in the image both meet the design expectations, enhancing the overall consistency and logical rationality of the image. Due to its semantic parsing, semantic supplementation, semantic reasoning, intelligent optimization and other capabilities, it shows broad application prospects. It can achieve customized creative advertising design and brand image generation in the advertising and marketing field; assist in the design of animated characters and scene design in film and television and animation production; realize the automatic generation of game scenes and characters and virtual environment modeling in game design and virtual reality fields to enhance the game experience; in the education field, it can generate intelligent learning materials and match illustrations for teaching aids. Through cross-field applications, not only the work efficiency is improved and the cost is reduced, but also more personalized and customized innovative application scenarios are created for various industries.

[0047] Through the embodiments of this specification, through the cross-modal attention mechanism of the large language model, the natural language description is decoupled into three semantic embedding vectors of biological characteristics, environmental elements, and physical parameters, eliminating semantic deviation problems such as subject loss and attribute confusion caused by text-image mapping dimension collapse in traditional methods. It not only analyzes the anatomical characteristics of biological subjects, but also quantifies environmental optical parameters and accessory mechanical properties, making the generated images conform to scientific laws at both the microscopic and macroscopic levels, ensuring a high degree of consistency between the generated images and the text descriptions, and significantly reducing the generation distortion problems caused by semantic errors in traditional methods; accurately locates the missing biomechanical parameters, material physical properties, and environmental interaction rules in the semantic embedding vectors, dynamically complements the implicit constraints that are difficult to explicitly define in traditional methods, making the generated content conform to scientific laws in dimensions such as anatomical rationality, material durability, and environmental dynamic response, automatically maps unquantified descriptions to computable ecological parameters, and avoids the generation randomness caused by language ambiguity in traditional methods; through dynamic prompt optimization and text intelligent supplementation, automatically completes and infers the potential information of multi-subject interactions, realizes the full expression of detailed information, ensures that the interaction relationships and visual display effects between the subjects in the image can meet the design expectations, and improves the overall consistency and logical rationality of the image; through hierarchical coding and multi-modal network linkage, transforms the physical laws implicit in natural language, such as biomechanics, material deformation, and environmental interaction, into quantifiable and verifiable mathematical constraints, breaking through the limitations of traditional generation models that rely on data fitting.

[0048] The embodiments of this specification also provide an intelligent content generation device based on a large language model, as Figure 3 shown. The device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above method.

[0049] The embodiments of this specification also provide a non-volatile computer storage medium storing computer-executable instructions, and the computer-executable instructions are set to: execute the above method.

[0050] The various embodiments in this specification are all described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the embodiments of the device, equipment, and non-volatile computer storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0051] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0052] The devices and media provided by the embodiments of this specification correspond one-to-one with the methods. Therefore, the devices and media also have beneficial technical effects similar to those of their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be elaborated here.

[0053] Those skilled in the art should understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0054] This specification is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0055] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0056] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 steps of the functions specified in one block or multiple blocks.

[0057] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0058] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.

[0059] Computer-readable media includes permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0060] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, commodity or device comprising the element.

[0061] The above description is only for one or more embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various changes and modifications can be made to one or more embodiments of this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the scope of the claims of this specification.

Claims

1. An intelligent content generation method based on a large language model, characterized in that, The method includes: Obtain the text description information input by the user, and perform cross-modal semantic parsing on the text description information through a large language model to generate a multi-dimensional semantic embedding vector, where the multi-dimensional semantic embedding vector includes a biometric semantic embedding vector, an environmental factor semantic embedding vector, and a physical parameter semantic embedding vector; Based on the multi-dimensional semantic embedding vector, dynamically infer and complete the missing interaction parameters to generate optimized structured description data; According to the structured description data and a preset multi-modal control network, generate initial image content, and verify the initial image content to determine the target image content.

2. The intelligent content generation method based on a large language model according to claim 1, wherein Performing cross-modal semantic parsing on the text description information through a large language model to generate a multi-dimensional semantic embedding vector specifically includes: Parse the text description information through the cross-modal attention mechanism of the large language model to determine multi-dimensional parsing features, where the multi-dimensional parsing features include anatomical features, kinematic parameters, and environmental optical characteristics; Map the multi-dimensional parsing features to determine a multi-dimensional semantic embedding vector including bone key point constraints, material elastic modulus, and ecological distribution density.

3. The intelligent content generation method based on a large language model according to claim 1, characterized in that, Based on the multi-dimensional semantic embedding vector, dynamically infer and complete the missing interaction parameters to generate optimized structured description data, specifically including: Perform fuzzy information recognition and element deconstruction processing on the multi-dimensional semantic embedding vector to determine the corresponding target completion entity, where the target completion entity includes any one or more of a biological entity, an environmental entity, and a material entity; Through a knowledge inference engine, perform interaction parameter completion on the target completion entity to determine biomechanical adaptation parameters that are not clearly specified to generate an inference completion semantics; Adopt a hierarchical coding strategy to hierarchically code the inference completion semantics to determine multiple hierarchical coding information to determine the structured description data, where the hierarchical coding information includes a biokinetics layer, a material optics layer, and an ecological interaction layer.

4. An intelligent content generation method based on a large language model according to claim 3, characterized in that, Performing fuzzy information recognition and element deconstruction processing on the multi-dimensional semantic embedding vector to determine the corresponding target completion entity specifically includes: Use dependency syntactic analysis to parse the multi-dimensional semantic embedding vector to extract multiple target entities, where the target entities include biological subjects, environmental factors, and material characteristics; Analyze the description of each target entity to determine whether the target entity has missing descriptions and unquantified descriptions. If so, it is determined as a target completion entity.

5. An intelligent content generation method based on a large language model according to claim 3, characterized in that, Through a knowledge inference engine, perform interaction parameter completion on the target completion entity to determine biomechanical adaptation parameters that are not clearly specified, specifically including: Based on the bone key point constraints in the multi-dimensional semantic embedding vector, determine a biokinematic model corresponding to the biometric characteristics, and through the biokinematic model, determine biodynamic adaptation parameters to determine biokinetics layer description data; Construct the finite element stress distribution data corresponding to the material contact surface according to the material elastic modulus in the multi-dimensional semantic embedding vector, and determine the material optical constraint conditions based on the finite element stress distribution data to determine the material optical layer description data; Through the ecological distribution density in the multi-dimensional semantic embedding vector, generate the recursive growth structure corresponding to the environmental elements through the fractal algorithm, and establish the contact dynamics model between the biological entity and the environmental elements to determine the ecological interaction layer description data.

6. The intelligent content generation method based on a large language model according to claim 1, characterized in that, Generate the initial image content according to the structured description data and the preset multi-modal control network, specifically including: Decompose the structured description data into multiple hierarchical coding information, where the hierarchical coding information includes the biokinetics layer, the material optical layer, and the ecological interaction layer; Determine the pose control network, the material control network, and the ecological control network in the multi-modal control network; According to the hierarchical coding information, constrain the biological motion pose through the pose control network, inject physical optical parameters through the material control network, and construct an environmental interaction model through the ecological control network to generate the initial image content.

7. An intelligent content generation method based on a large language model according to claim 1, characterized in that, Verify the initial image content to determine the target image content, specifically including: Perform biodynamic adaptability verification and physical parameter consistency detection on the initial image content to determine the contact adaptability verification result and the physical parameter consistency detection result; According to the contact adaptability verification result and the physical parameter consistency detection result, perform gradient correction on the latent space representation, and iteratively adjust the material parameters of the abnormal area through the physical renderer optimizer to output the target image content that meets the bio-physical joint verification.

8. An intelligent content generation method based on a large language model according to claim 7, characterized in that, Perform biodynamic adaptability verification and physical parameter consistency detection on the initial image content to determine the contact adaptability verification result and the physical parameter consistency detection result, specifically including: Load the preset biological anatomy model, extract the geometric features of the contact surface, perform finite element stress analysis based on the geometric features, and generate a contact pressure distribution map; Compare the pressure distribution map with the material yield strength threshold, mark the over-limit area to determine the biodynamic adaptability verification result; Extract the material optical parameters from the hierarchical coding information of the structured description data, and generate a theoretical reflection spectrum through ray tracing simulation; Calculate the chromaticity difference between the actual reflection data corresponding to the initial image content and the theoretical reflection spectrum to verify whether the material reflection characteristics meet the physical constraints.

9. An intelligent content generation device based on a large language model, characterized in that, The device includes: At least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-8.

10. A non-volatile computer storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions are set to: execute the method according to any one of claims 1-8.

Citation Information

Cited By

  • Motion event analysis method, electronic equipment and medium

    CN122045712A

  • Motion event analysis method, electronic device, medium

    CN122045712B