Adaptive rendering method and system for AI digital human virtual and real scene fusion of Chinese characters and images

Through the neural network ordinary differential equation transformer and the luminous insect swarm optimization algorithm, precise alignment and dynamic rendering of Chinese text and images are achieved, which solves the problem of insufficient Chinese semantic understanding in existing technologies and improves the naturalness and interactivity of the AI ​​digital human image and text fusion.

CN120355827BActive Publication Date: 2025-09-19BAIGE ONLINE (XIAMEN) DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510846703.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-19
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

Existing AI digital human image and text fusion methods lack a deep understanding of Chinese semantics, making it difficult to achieve accurate alignment and dynamic rendering of Chinese text and images, resulting in a lack of semantic consistency and visual coordination in the fusion effect. The model structure also lacks an adaptive adjustment mechanism and cannot adapt to the dynamic behavior of digital humans and changes in scene lighting.

Method used

The neural network ordinary differential equation transformer and the luminous insect swarm optimization algorithm are used to realize the joint modeling of Chinese text and image features through the multi-head attention module and the variable time step ODE calculation module. The luminous insect swarm optimization algorithm is combined to optimize the model structure parameters, generate graphic rendering instructions, and perform posture and lighting adjustments to achieve high-quality virtual-reality fusion rendering.

Benefits of technology

It achieves dynamic, precise arrangement and high-quality rendering of Chinese text in images, improves the naturalness and intelligence of the fusion of images and text, adapts to a variety of virtual interactive scenarios, and enhances the immersion and interactivity of AI digital people's image and text expression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355827B_ABST
    Figure CN120355827B_ABST
Patent Text Reader

Abstract

The present invention discloses an adaptive rendering method and system for AI digital human virtual-real scene fusion of Chinese characters and images, including the following steps: S1, collecting Chinese text and image data, extracting semantics and regional features; S2, constructing a neural ordinary differential equation transformer model, inputting text and image features to generate a fusion representation; S3, extracting the attention weight matrix to determine the spatial arrangement of Chinese text in the image; S4, optimizing the structural parameter vector using a luminescent insect swarm optimization algorithm to obtain a fusion model; S5, generating a graphic rendering instruction set based on the preliminary arrangement and optimization model; S6, combining the digital human's motion state with background image information to complete posture mapping and lighting adjustment to generate a rendering frame; S7, inputting the rendered frame into a rendering engine to output video content fused with Chinese characters and images. The present invention achieves the precise fusion of Chinese semantics and images, significantly improving the naturalness and interactivity of AI digital human graphic presentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of virtual human rendering technology, and in particular to an adaptive rendering method and system for AI digital humans that integrates Chinese characters and images into virtual and real scenes. Background Art

[0002] Against the backdrop of the rapid development of artificial intelligence and digital content generation, AI digital humans, a combination of computer vision, speech generation, natural language processing, and other technologies, have been widely used in a variety of scenarios, including virtual live broadcasts, educational broadcasts, and cultural communication. As application demands continue to expand, AI digital humans must not only possess speech and motion generation capabilities but also be able to deeply integrate graphic content with real or virtual scenes. Especially in the Chinese context, achieving the accurate and natural integration of Chinese characters and background images has become a key challenge.

[0003] Most existing image-text fusion methods rely on static templates or rule-driven methods to arrange text in images. They lack a deep understanding of text semantics and are unable to adapt to the structure of Chinese expressions in complex contexts. At the same time, some systems use simple matching strategies based on image segmentation to locate text in image areas, but fail to combine the dynamic correspondence between text semantics and image content, resulting in a lack of semantic consistency and visual coordination in the fusion effect. In addition, during the image-text fusion process, existing methods often ignore the impact of AI digital humans in different movements, postures, and different scene lighting conditions on the image-text fusion display results, making it impossible to achieve dynamic, coherent, and high-quality rendering output.

[0004] In terms of model construction, traditional image and text processing methods primarily rely on convolutional neural networks or Transformer structures, which have inherent limitations in processing the temporal continuity and multi-scale dynamic characteristics of Chinese semantics and image information. This can easily lead to context loss or blurred semantic fusion, especially in dynamically changing scenarios. In terms of optimization, most current model structures rely on fixed parameter settings and lack adaptive adjustment mechanisms for structural elements such as model depth, attention mechanisms, and time control. This makes it impossible to dynamically optimize fusion effects based on different data distributions.

[0005] To address these issues, there is an urgent need for a text-and-image fusion rendering method and system that can accurately align Chinese semantic information with image region content, combine attention mechanisms with continuous-time modeling capabilities, and possess adjustable structural optimization capabilities. The system should also be able to adapt to the dynamic behavior of digital humans and changes in scene lighting. This system should support the joint modeling of Chinese text and images, possess semantic perception and visual linkage capabilities, and be able to automatically determine text layout positions, generate text-and-image rendering instructions, and synchronize posture and lighting adjustments. It should also present the fusion effect in the form of high-quality rendered frame output, thereby enhancing the AI ​​digital human's intelligent text-and-image expression capabilities and the naturalness of interaction in multiple scenarios.

[0006] Therefore, how to provide an adaptive rendering method and system for AI digital humans that integrates Chinese characters and images in virtual and real scenes is an urgent problem that needs to be solved by those skilled in the art. Summary of the Invention

[0007] One purpose of the present invention is to propose an AI digital human Chinese character and image fusion rendering method based on a neural ordinary differential equation transformer and a luminous insect swarm optimization algorithm. The present invention makes full use of natural language processing, multimodal feature modeling and continuous-time neural network technology, and describes in detail the joint modeling, layout optimization and dynamic rendering process of Chinese text semantics and image region features. It has the advantages of precise semantic alignment, natural visual fusion and strong ability to adapt to dynamic scenes.

[0008] The adaptive rendering method for AI digital human virtual-real scene fusion Chinese characters and images according to an embodiment of the present invention includes the following steps:

[0009] S1. Collect Chinese text and corresponding image data, perform word segmentation and semantic encoding on the Chinese text, and generate text feature vectors; perform region segmentation and texture extraction on the image data, and generate image region feature vectors;

[0010] S2. Construct a neural ordinary differential equation transformer model, including a multi-head attention module and a variable time step ODE calculation module. The structural parameter vector of the neural ordinary differential equation transformer model includes network depth, number of attention heads, and time step control parameters. Input the text feature vector and image region feature vector into the neural ordinary differential equation transformer model to generate a fused feature representation.

[0011] S3. Extract the attention weight matrix generated by the multi-head attention module from the fused feature representation, determine the spatial arrangement of the Chinese text in the image based on the attention weight matrix, and form a preliminary image-text fusion layout;

[0012] S4. Use the glowworm swarm optimization algorithm to jointly optimize the network depth, number of attention heads, and time step control parameters to obtain a fusion model with optimized structure;

[0013] S5. Based on the preliminary graphic and text fusion layout, generate a graphic and text rendering instruction set based on the structurally optimized fusion model;

[0014] S6. Combine the current action state of the digital human and the background image information to perform posture mapping and lighting adjustment on the graphic rendering instruction set to generate a virtual-reality fusion rendering frame;

[0015] S7. Input the virtual-reality fusion rendering frame into the digital human rendering engine to complete the output of the video content fusion of Chinese text and image.

[0016] Optionally, the S1 specifically includes:

[0017] S11. Collect Chinese text and corresponding image data, where the Chinese text is a short text with a specific expression intent, and the image data is an image containing a scene or object semantically related to the text;

[0018] S12. Perform Chinese language word segmentation processing on each Chinese text, using a word segmentation method that combines rule-based and data-driven methods to divide the original text into word sequences with word boundaries;

[0019] S13. After completing the language segmentation, encode the word sequence, extract the semantic features including the context dependency, and form a text feature vector of fixed dimension;

[0020] S14, performing region segmentation processing on each image data, using a region segmentation algorithm based on texture and boundary features to divide the image data into a number of image regions with consistent texture patterns;

[0021] S15. Perform texture feature operations on the divided image data regions to extract color distribution, edge direction, and local structural features of each region in the image to form an image region feature vector.

[0022] Optionally, the S2 specifically includes:

[0023] S21, taking the text feature vector as the language input sequence and the image region feature vector as the visual input sequence, and performing a unified encoding operation on the two types of input sequences;

[0024] S22. Input the uniformly encoded input sequence into a multi-head attention module. The multi-head attention module includes several parallel attention sub-heads. Each attention sub-head calculates the correlation between the text and the image and generates several attention representation matrices.

[0025] S23. Perform concatenation and linear mapping operations on each attention representation matrix to construct a fused attention output sequence as the output of the multi-head attention module;

[0026] S24. The fused attention output sequence is used as an input feature and input into the variable time step ODE calculation module. The input feature state changes over time based on the continuous time dynamic function. The continuous time dynamic function is composed of the following structure: a first linear transformation unit performs a joint transformation on the input feature and the time variable and outputs an intermediate representation; a time-aware gating control unit is set between the first linear transformation unit and the second linear transformation unit, and an input gate and a reset gate are generated based on the current time variable and the state of the input feature to dynamically adjust the channel information of the intermediate representation; the second linear transformation unit maps the gated feature representation and inputs it into the activation function unit. The activation function unit adopts a time-adjusted activation mechanism to control the response strength of the nonlinear output according to the adjustable coefficient generated by the time variable; the variable time step ODE calculation module uses the time step control parameter as the integral granularity control input, integrates the continuous time dynamic function over the time step, and outputs a multi-layer dynamic representation that evolves over time;

[0027] S25. The overall depth of the neural ordinary differential equation converter model is determined by the number of consecutive stacked layers of the included variable time step ODE calculation modules. The number of stacked layers serves as a network depth parameter to control the step size of each integration layer in the variable time step ODE calculation module. The time step control parameter is used to achieve a balance between feature calculation accuracy and calculation efficiency.

[0028] S26. Use the multi-layer dynamic representation output by the variable time step ODE calculation module as the fusion feature representation.

[0029] Optionally, the S3 specifically includes:

[0030] S31, extracting an attention weight matrix generated by the multi-head attention module based on the fused feature representation, wherein the attention weight matrix represents the matching strength between each word in the text feature vector and each region in the image region feature vector;

[0031] S32, determining the image area corresponding to each Chinese text based on the maximum weight value corresponding to each text in the attention weight matrix, and establishing a one-to-one mapping relationship between the Chinese text and the image area;

[0032] S33, extracting the spatial coordinate range of each corresponding image area based on the one-to-one mapping relationship, determining the arrangement position of the Chinese text on the image, and generating a preliminary image-text fusion layout including the display positions of all Chinese texts;

[0033] S34. Conflict detection and position adjustment processing are performed on the preliminary graphic and text fusion layout, and a minimum spacing threshold is set between any two adjacent Chinese words to ensure readability and non-overlapping areas of the text layout.

[0034] Optionally, the S4 specifically includes:

[0035] S41, initializing a population of glowworms, where each glowworm individual corresponds to a set of structural parameter vectors, wherein the structural parameter vectors are composed of three parts: network depth, number of attention heads, and time step control parameters in the neural ordinary differential equation transformer model;

[0036] S42. Setting a brightness function for each glowworm individual, wherein the brightness function uses the comprehensive performance index of the neural ordinary differential equation transformer model on the validation set as the objective function value, the comprehensive performance index being a weighted combination of three indicators: image and text arrangement accuracy, attention distribution sparsity, and fusion feature stability, where the sum of the weighted coefficients of each indicator is one;

[0037] S43. In each iteration, based on the relative difference between the individual brightness value and the brightness values ​​of other individuals in the neighborhood, and based on the difference between the current individual structure parameter vector and the structure parameter vector of the brighter individual in the neighborhood, a moving function is constructed through the brightness difference weight and the spatial distance factor to update the network depth, number of attention heads, and time step control parameters of the current individual;

[0038] S44. Setting a step-size attenuation mechanism and parameter boundary control conditions during the movement process to prevent the structural parameter vector from crossing the boundary and maintain the adaptability and convergence of the search area;

[0039] S45. When the preset maximum number of iterations or the brightness function convergence condition is reached, the lightworm individual with the optimal brightness value is selected, and the corresponding structural parameter vector is the optimal combination of network depth, number of attention heads, and time step control parameters;

[0040] S46. Based on the optimal combination, the neural ordinary differential equation converter model is reconstructed as the fusion model after structural optimization.

[0041] Optionally, the S5 graphic rendering instruction set includes: determining the font, color, transparency and display coordinates of each Chinese text in the image based on the structurally optimized fusion model, constructing graphic rendering instructions containing text style parameters and spatial position parameters, and combining them to generate a graphic rendering instruction set.

[0042] Optionally, S6 specifically includes: obtaining the skeletal posture parameters of the digital human in the current frame and the lighting distribution characteristics of the background image, performing coordinate transformation and brightness adjustment based on the posture angle and lighting direction according to the graphic rendering instruction set, and outputting the processed image frame as a virtual-reality fusion rendering frame.

[0043] The adaptive rendering system for AI digital human virtual-real scene fusion of Chinese characters and images according to an embodiment of the present invention includes the following modules:

[0044] The text feature extraction module is used to perform word segmentation and semantic encoding on Chinese text and generate text feature vectors;

[0045] Image feature extraction module, used to perform region segmentation and texture extraction on image data and generate image region feature vectors;

[0046] A fusion modeling module is used to input the text feature vector and the image region feature vector into the neural ordinary differential equation transformer model to generate a fusion feature representation;

[0047] The attention arrangement module is used to extract the attention weight matrix from the fused feature representation and generate a preliminary image-text fusion layout;

[0048] Parameter optimization module, used to optimize network depth, number of attention heads, and time step control parameters based on the glowworm swarm optimization algorithm;

[0049] An instruction generation module is used to generate an image and text rendering instruction set based on the preliminary image and text fusion layout;

[0050] The posture lighting mapping module is used to combine the digital human's action status with the background image information to generate a virtual-reality fusion rendering frame;

[0051] The graphic rendering engine module is used to output the virtual-reality fusion rendering frame into video content that is a fusion of Chinese characters and images.

[0052] The beneficial effects of the present invention are:

[0053] First, the present invention effectively integrates Chinese semantic features and image region features by constructing a neural network ordinary differential equation transformer model, solving the problems of inaccurate semantic alignment of images and texts and template-dependent arrangement in existing methods, and realizing dynamic reasoning of the position of Chinese text in the image and semantically consistent spatial arrangement, thereby improving the naturalness and intelligence of image-text fusion.

[0054] Secondly, the present invention introduces the luminous insect swarm optimization algorithm to jointly optimize the network depth, number of attention heads and time step control parameters of the neural ordinary differential equation transformer model, significantly enhancing the adaptability of the model structure to the complexity of different graphic samples, and improving the comprehensive performance of the fusion model in arrangement accuracy, attention sparsity and feature stability, effectively overcoming the defects of structural rigidity and poor generalization ability of existing methods.

[0055] In addition, the present invention combines the dynamic behavior of AI digital humans with scene background information, performs posture mapping and lighting adjustments on the graphic rendering instruction set, and generates dynamic rendering frames that integrate virtual and real elements. This achieves a high degree of coordination between Chinese content, image background, and digital human behavior, and is suitable for various application scenarios such as virtual live broadcast and intelligent interaction, enhancing the immersion and interactivity of AI digital humans' graphic expression. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0057] Figure 1 This is the overall flow chart of the adaptive rendering method for integrating Chinese characters and images in virtual and real scenes of AI digital humans proposed by the present invention;

[0058] Figure 2 This is a flowchart of the structural parameter vector optimization based on the luminous insect swarm optimization algorithm for the adaptive rendering method of Chinese characters and images in the virtual and real scene fusion of AI digital humans proposed in the present invention. DETAILED DESCRIPTION

[0059] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.

[0060] refer to Figure 1 and Figure 2 The adaptive rendering method for AI digital human virtual and real scene fusion Chinese characters and images includes the following steps:

[0061] S1. Collect Chinese text and corresponding image data, perform word segmentation and semantic encoding on the Chinese text, and generate text feature vectors; perform region segmentation and texture extraction on the image data, and generate image region feature vectors;

[0062] S2. Construct a neural ordinary differential equation transformer model, including a multi-head attention module and a variable time step ODE calculation module. The structural parameter vector of the neural ordinary differential equation transformer model includes network depth, number of attention heads, and time step control parameters. Input the text feature vector and image region feature vector into the neural ordinary differential equation transformer model to generate a fused feature representation.

[0063] S3. Extract the attention weight matrix generated by the multi-head attention module from the fused feature representation, determine the spatial arrangement of the Chinese text in the image based on the attention weight matrix, and form a preliminary image-text fusion layout;

[0064] S4. Use the glowworm swarm optimization algorithm to jointly optimize the network depth, number of attention heads, and time step control parameters to obtain a fusion model with optimized structure;

[0065] S5. Based on the preliminary graphic and text fusion layout, generate a graphic and text rendering instruction set based on the structurally optimized fusion model;

[0066] S6. Combine the current action state of the digital human and the background image information to perform posture mapping and lighting adjustment on the graphic rendering instruction set to generate a virtual-reality fusion rendering frame;

[0067] S7. Input the virtual-reality fusion rendering frame into the digital human rendering engine to complete the output of the video content that fuses Chinese text and images. Specifically, the virtual-reality fusion rendering frame after posture mapping and lighting adjustment is injected into the digital human rendering engine as input. The rendering engine generates a video sequence containing the fusion effect of Chinese text and image frame by frame based on the frame content, realizing the synchronous presentation of text content and image visual information in the digital human broadcast scene, and finally outputting fused video content with semantic consistency and visual coherence.

[0068] By introducing the neural ordinary differential equation transformer model and the luminous insect swarm optimization algorithm, this invention constructs a Chinese character and image fusion rendering method for AI digital humans. It can achieve accurate fusion of text and image areas in dynamic semantic scenes, overcoming the problems of inaccurate semantic alignment and static and rigid rendering in existing text and image arrangement methods, greatly improving the naturalness, matching and rendering coherence of the fused content, and adapting to a variety of virtual interactive scenes.

[0069] In this embodiment, S1 specifically includes:

[0070] S11. Collect Chinese text and corresponding image data, where the Chinese text is a short text with a specific expression intent, and the image data is an image containing a scene or object semantically related to the text;

[0071] S12. Perform Chinese language word segmentation processing on each Chinese text, divide the continuous text into meaningful word units through a lexical analyzer, and use a word segmentation method that combines rule-based and data-driven methods to divide the original text into word sequences with word boundaries;

[0072] S13. After completing the language segmentation, encode the word sequence, encode the segmented word sequence, convert each word into a corresponding numerical vector representation using a pre-trained word vector model or a context embedding model, retain semantic information and context dependency, extract semantic features containing context dependency, and form a fixed-dimensional text feature vector;

[0073] S14, performing region segmentation processing on each image data, using a region segmentation algorithm based on texture and boundary features to divide the image data into a number of image regions with consistent texture patterns;

[0074] S15. Perform texture feature operations on the divided image areas, use a texture coding network or a statistical feature algorithm to extract the texture pattern, edge structure, and local details of each area, generate features that reflect the visual characteristics of the area, extract the color distribution, edge direction, and local structure features of each area in the image, and form an image area feature vector.

[0075] This invention refines and standardizes the text segmentation, semantic coding and image region feature extraction processes for the preprocessing of Chinese text and image data, ensuring the unity of input data dimensions and semantic feature structure, providing a stable and compatible feature basis for subsequent model input, and solving the problem of model fusion difficulties caused by feature dimension mismatch in existing methods.

[0076] In this embodiment, S2 specifically includes:

[0077] S21. Use text feature vectors as language input sequences and image region feature vectors as visual input sequences, perform unified encoding operations on the two types of input sequences, and use an embedding transformation method with the same dimension to map the two types of inputs into a shared representation space, maintaining consistency between semantic and visual information.

[0078] S22. Input the uniformly encoded input sequence into a multi-head attention module. The multi-head attention module includes several parallel attention sub-heads. Each attention sub-head calculates the relevance between the text and the image. Specifically, based on the attention mechanism, a weighted evaluation is performed on the matching strength between each text position and the image region to capture the semantic dependency and correspondence between the two, and generate several attention representation matrices.

[0079] S23. Concatenate and linearly map each attention representation matrix. Specifically, the attention matrices output by each attention sub-head are concatenated along the feature dimension to form a unified multi-head attention representation matrix. Then, a linear transformation is performed through a fully connected layer to integrate the attention information of each sub-head and improve the representation ability. Finally, a fused attention output sequence is constructed as the output of the multi-head attention module.

[0080] S24. The fused attention output sequence is used as an input feature and input into the variable time step ODE calculation module. The input feature state changes over time based on the continuous time dynamic function. The continuous time dynamic function is composed of the following structure: a first linear transformation unit performs a joint transformation on the input feature and the time variable and outputs an intermediate representation; a time-aware gating control unit is set between the first linear transformation unit and the second linear transformation unit, and an input gate and a reset gate are generated based on the current time variable and the state of the input feature to dynamically adjust the channel information of the intermediate representation; the second linear transformation unit maps the gated feature representation and inputs it into the activation function unit. The activation function unit adopts a time-adjusted activation mechanism to control the response strength of the nonlinear output according to the adjustable coefficient generated by the time variable; the variable time step ODE calculation module uses the time step control parameter as the integral granularity control input, integrates the continuous time dynamic function over the time step, and outputs a multi-layer dynamic representation that evolves over time;

[0081] S25. The overall depth of the neural ordinary differential equation converter model is determined by the number of consecutive stacked layers of the included variable time step ODE calculation modules. The number of stacked layers serves as a network depth parameter to control the step size of each integration layer in the variable time step ODE calculation module. The time step control parameter is used to achieve a balance between feature calculation accuracy and calculation efficiency.

[0082] S26. Use the multi-layer dynamic representation output by the variable time step ODE calculation module as the fusion feature representation.

[0083] The present invention realizes the deep fusion of cross-modal Chinese text and image features in the continuous time domain by constructing a neural ordinary differential equation transformer model including a multi-head attention module and a variable time-step ODE calculation module. It further enhances the modeling ability of nonlinear semantic associations by improving the continuous-time dynamic function in the structure, and significantly improves the expression accuracy and temporal consistency of the multimodal fusion model.

[0084] In this embodiment, S3 specifically includes:

[0085] S31, extracting an attention weight matrix generated by the multi-head attention module based on the fused feature representation, wherein the attention weight matrix represents the matching strength between each word in the text feature vector and each region in the image region feature vector;

[0086] S32, determining the image area corresponding to each Chinese text based on the maximum weight value corresponding to each text in the attention weight matrix, and establishing a one-to-one mapping relationship between the Chinese text and the image area;

[0087] S33. Extract the spatial coordinate range of each corresponding image region based on the one-to-one mapping relationship and determine the layout position of the Chinese text on the image. Specifically, extract the spatial coordinate range of each focused image region, including the normalized position coordinates of its upper left corner and lower right corner in the image, based on the one-to-one mapping relationship between the Chinese text and the image region constructed by the attention weight matrix generated by the multi-head attention module. Then, positionally match the Chinese text according to the geometric center or high attention density point of the region, determine the layout position of each text segment on the image, and generate a preliminary image-text fusion layout including the display positions of all Chinese texts.

[0088] S34. Conflict detection and position adjustment processing are performed on the preliminary graphic and text fusion layout, and a minimum spacing threshold is set between any two adjacent Chinese words to ensure readability and non-overlapping areas of the text layout.

[0089] The present invention extracts attention weight information from the multi-head attention representation matrix and constructs a spatial arrangement strategy for text in the image, thereby achieving dynamic positioning matching between Chinese semantics and image areas. It effectively solves the problems of traditional arrangement methods being unable to understand language context and causing serious misalignment of images and text, provides a structured spatial basis for subsequent rendering instruction generation, and improves the intelligent level of image and text arrangement.

[0090] In this embodiment, the S4 specifically includes:

[0091] S41, initializing a population of glowworms, where each glowworm individual corresponds to a set of structural parameter vectors, wherein the structural parameter vectors are composed of three parts: network depth, number of attention heads, and time step control parameters in the neural ordinary differential equation transformer model;

[0092] S42. Setting a brightness function for each glowworm individual, wherein the brightness function uses the comprehensive performance index of the neural ordinary differential equation transformer model on the validation set as the objective function value, the comprehensive performance index being a weighted combination of three indicators: image and text arrangement accuracy, attention distribution sparsity, and fusion feature stability, where the sum of the weighted coefficients of each indicator is one;

[0093] S43. In each iteration, based on the relative difference between the individual brightness value and the brightness values ​​of other individuals in the neighborhood, and based on the difference between the current individual structure parameter vector and the structure parameter vector of the brighter individual in the neighborhood, a moving function is constructed through the brightness difference weight and the spatial distance factor to update the network depth, number of attention heads, and time step control parameters of the current individual;

[0094] S44. Setting a step-size attenuation mechanism and parameter boundary control conditions during the movement process, specifically: introducing a step-size attenuation mechanism and parameter boundary control conditions during the movement of individual lightworms. The step-size attenuation mechanism dynamically reduces the individual movement distance according to the number of iterations, realizing the transition from coarse-grained search to fine-grained optimization, and improving convergence stability. The parameter boundary control conditions are used to constrain the parameter values ​​of the individual after each update to remain within the preset legal range of the structural parameter vector, preventing the model structure from having invalid or abnormal configurations beyond the design range, ensuring that the optimization process is effective and controllable, preventing the structural parameter vector from crossing the boundary, and maintaining the adaptability and convergence of the search area.

[0095] S45. When the preset maximum number of iterations or the brightness function convergence condition is reached, the lightworm individual with the optimal brightness value is selected, and the corresponding structural parameter vector is the optimal combination of network depth, number of attention heads, and time step control parameters;

[0096] S46. Based on the optimal combination, the neural ordinary differential equation converter model is reconstructed as the fusion model after structural optimization.

[0097] The present invention adopts the luminous insect swarm optimization algorithm to dynamically update the structural parameter vector of the neural ordinary differential equation converter model. By guiding the structural evolution through the brightness function and combining the information transfer mechanism between adjacent individuals, the adaptive adjustment of the fusion model structure within the search space is achieved, which effectively overcomes the problems of strong structural rigidity and weak generalization ability of traditional models and improves the global optimal performance of the fusion model.

[0098] In this embodiment, the S5 graphic rendering instruction set includes: determining the font, color, transparency and display coordinates of each Chinese text in the image based on the structurally optimized fusion model; based on the structurally optimized fusion model, combining the attention weight distribution and the semantic style information carried in the fusion feature representation, respectively determining the visual attributes and spatial attributes of each Chinese text: the font is mapped by the semantic emphasis intensity, the color and transparency are set according to the emotional tendency of the text and the contrast of the background image, and the display coordinates are calculated by the attention center position of the corresponding image area; constructing a graphic rendering instruction containing text style parameters and spatial position parameters, and combining them to generate a graphic rendering instruction set; the graphic rendering instruction is generated by the fusion model and is used to drive the final graphic presentation, containing the text style parameters of each Chinese text, such as font type, font size, color, transparency, and spatial position parameters, including display coordinates and alignment on the image; the instruction set is used to guide the rendering engine to accurately superimpose the text content on the image area in a manner that adapts to the visual context, thereby achieving a natural fusion display of text and image.

[0099] In this embodiment, S6 specifically includes: obtaining the skeletal posture parameters of the digital human and the lighting distribution characteristics of the background image in the current frame, and executing coordinate transformation and brightness adjustment based on the posture angle and lighting direction according to the graphic rendering instruction set. The coordinate transformation is based on the current three-dimensional posture angle information of the digital human, and the text display coordinates defined in the graphic rendering instruction are affine transformed or perspective transformed so that the text position is visually adjusted synchronously with the digital human's movements to ensure that it fits the dynamic perspective. The brightness adjustment is based on the direction and intensity of the main light source of the current background image, and the brightness and darkness of the text in the target area that may be affected by the lighting is calculated, and the text color or transparency is increased or decreased to make it clear and visible in the rendered frame without being overly abrupt, thereby improving the overall visual fusion and naturalness, and outputting the processed image frame as a virtual-reality fusion rendering frame.

[0100] The adaptive rendering system for AI digital human virtual-real scene fusion of Chinese characters and images according to an embodiment of the present invention includes the following modules:

[0101] The text feature extraction module is used to perform word segmentation and semantic encoding on Chinese text and generate text feature vectors;

[0102] Image feature extraction module, used to perform region segmentation and texture extraction on image data and generate image region feature vectors;

[0103] A fusion modeling module is used to input the text feature vector and the image region feature vector into the neural ordinary differential equation transformer model to generate a fusion feature representation;

[0104] The attention arrangement module is used to extract the attention weight matrix from the fused feature representation and generate a preliminary image-text fusion layout;

[0105] Parameter optimization module, used to optimize network depth, number of attention heads, and time step control parameters based on the glowworm swarm optimization algorithm;

[0106] An instruction generation module is used to generate an image and text rendering instruction set based on the preliminary image and text fusion layout;

[0107] The posture lighting mapping module is used to combine the digital human's action status with the background image information to generate a virtual-reality fusion rendering frame;

[0108] The graphic rendering engine module is used to output the virtual-reality fusion rendering frame into video content that is a fusion of Chinese characters and images.

[0109] Example 1:

[0110] In order to verify the feasibility of the present invention in practice, the present invention was applied to a certain AI digital human virtual explanation system for the immersive broadcasting task of displaying digital cultural and museum content to the public. In this task, the system needs to fuse high-resolution exhibit images from the cultural relics image database with explanatory Chinese sentences to achieve the effect of AI digital human synchronous broadcasting and synchronous display of text and images. Traditional systems mainly use static templates or manual configuration to set the display position of text on the image, which is difficult to adapt to changes in text length, differences in image area complexity, and dynamic changes in broadcast rhythm. Ultimately, it leads to serious misalignment of text and images, rigid layout, and poor user experience.

[0111] The implementation of the present invention is based on the aforementioned virtual explanation scenario, taking exhibit description text and exhibit images as input. A semantic coding model is first used to segment and encode Chinese sentences, generating a text feature vector with a dimension of 512. Regional convolution processing and texture coding are used to extract local features of eight key regions in the image, forming an image region feature tensor of size 8256. Subsequently, a neural ordinary differential equation transformer model is constructed, with an initial network depth of 4 layers, a number of attention heads of 6, and a time step control parameter of 0.1. The aforementioned text and image features are input into the model to generate a fused feature representation tensor that captures the cross-modal attention association between text and image.

[0112] By extracting the attention weight matrix generated by the multi-head attention module, the model automatically infers the correspondence between each phrase in the text and the image region. It then calculates the text's layout within the image based on the position of the attention center and the relative weights. After generating the initial layout, the structural parameters were further optimized using the Luminous Swarm Optimization algorithm. The model ultimately converged to a depth of 6 layers and a number of attention heads of 8, significantly improving fusion accuracy. Using this optimized model, the system generates rendering instructions for each text line based on the fused layout, including font size (Fangsong, Microsoft Yahei, or Kaiti), font size (16–24), transparency (0.4–1.0), and coordinates (normalized based on the image's width and height). These instructions are synchronized with the digital human's presentation, combining skeletal motion with real-time lighting conditions to generate the rendered frame.

[0113] During this implementation, the traditional template layout method and the method of the present invention were compared in terms of image and text alignment, readability, and rendering consistency. The table is shown below:

[0114] Table 1: Comparative experimental results of the image-text fusion effect between the method of the present invention and the traditional method

[0115]

[0116] The above-mentioned "Comparative Experimental Table of the Image-Text Fusion Effects of the Method of the Present Invention and Traditional Methods" shows the evaluation data of 10 groups of image-text samples in actual applications, which fully reflects the performance advantages of the method of the present invention in image-text fusion. First, from the perspective of image-text alignment indicators, the alignment of sample numbers 01 to 10 all exceeds 90%, with the highest being 95.0% for sample 07, and the average alignment reaching 92.66%. This result shows that the present invention has a high degree of accuracy in processing the spatial correspondence between Chinese text and image areas, which is significantly better than the traditional template arrangement method.

[0117] In terms of text readability, all samples scored above 4.4, with the highest score reaching 4.8 and an average readability score of 4.53. This demonstrates that the text styles generated by this invention have high visual clarity and legibility in terms of font, size, and transparency control, conforming to user reading habits and enhancing the interactive experience.

[0118] The rendering coherence score reflects the stability and dynamic consistency of the system's graphic and text presentation during continuous broadcasting. Sample scores generally remained above 89%, with an average score of 91.26%, demonstrating strong fusion coherence and content consistency. This improvement in this metric is primarily attributed to the present invention's dynamic adaptation mechanism during the pose mapping and lighting adjustment stages, which effectively addresses the issues of abrupt and fragmented graphic presentation in traditional systems.

[0119] In addition, the model structure parameters also show good optimization stability. The structural depth is mostly 6 layers, the number of attention heads is mostly stable at 8, and some samples are slightly adjusted to adapt to changes in image complexity, reflecting that the glowworm swarm optimization algorithm has good adaptive adjustment capabilities while maintaining high model performance.

[0120] This table verifies the advantages of this invention in Chinese text and image fusion rendering from multiple perspectives. It not only surpasses traditional methods in terms of alignment, readability and coherence, but also reflects the effectiveness of the model structure design and optimization strategy, providing a reliable and efficient technical path for AI digital human text and image display.

[0121] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. An adaptive rendering method for AI digital human virtual and real scene fusion of Chinese characters and images, characterized by: The steps include: S1. Collect Chinese text and corresponding image data, perform word segmentation and semantic encoding on the Chinese text, and generate text feature vectors; perform region segmentation and texture extraction on the image data, and generate image region feature vectors; S2. Construct a neural ODE transformer model, including a multi-head attention module and a variable time-step ODE calculation module. The structural parameters of the neural ODE transformer model include network depth, number of attention heads, and time-step control parameters. Input the text feature vector and image region feature vector into the neural ODE transformer model to generate a fused feature representation. S3. Extract the attention weight matrix generated by the multi-head attention module from the fused feature representation, determine the spatial arrangement of the Chinese text in the image based on the attention weight matrix, and form a preliminary image-text fusion layout; S4. Using the glowworm swarm optimization algorithm, the network depth, number of attention heads, and time step control parameters are jointly optimized to obtain the neural ordinary differential equation converter model after structural optimization; S5. Based on the preliminary graphic and text fusion layout, generate a graphic and text rendering instruction set based on the structure-optimized neural ordinary differential equation transformer model; S6. Combine the current action state of the digital human and the background image information to perform posture mapping and lighting adjustment on the graphic rendering instruction set to generate a virtual-reality fusion rendering frame; S7, inputting the virtual-reality fusion rendering frame into the digital human rendering engine to complete the output of the video content fusion of Chinese text and image; The S2 specifically includes: S21, taking the text feature vector as the language input sequence and the image region feature vector as the visual input sequence, and performing a unified encoding operation on the two types of input sequences; S22. Input the uniformly encoded input sequence into a multi-head attention module. The multi-head attention module includes several parallel attention sub-heads. Each attention sub-head calculates the correlation between the text and the image and generates several attention representation matrices. S23. Perform concatenation and linear mapping operations on each attention representation matrix to construct a fused attention output sequence as the output of the multi-head attention module; S24. The fused attention output sequence is used as the input feature and input into the variable time step ODE calculation module, and the change of the input feature state over time is modeled based on the continuous time dynamic function; S25. The overall depth of the neural ordinary differential equation converter model is determined by the number of consecutive stacked layers of the included variable time step ODE calculation modules. The number of stacked layers serves as a network depth parameter to control the step size of each integration layer in the variable time step ODE calculation module. S26. Use the multi-layer dynamic representation output by the variable time step ODE calculation module as the fusion feature representation.

2. The adaptive rendering method for AI digital human virtual-real scene fusion Chinese characters and images according to claim 1 is characterized in that: Said S1 specifically includes: S11. Collect Chinese text and corresponding image data, where the Chinese text is a short text with a specific expression intent, and the image data is an image containing a scene or object semantically related to the text; S12. Perform Chinese language word segmentation processing on each Chinese text, using a word segmentation method that combines rule-based and data-driven methods to divide the original text into word sequences with word boundaries; S13. After completing the language segmentation, encode the word sequence, extract the semantic features including the context dependency, and form a text feature vector of fixed dimension; S14, performing region segmentation processing on each image data, using a region segmentation algorithm based on texture and boundary features to divide the image data into a number of image regions with consistent texture patterns; S15. Perform texture feature operations on the divided image data regions to extract color distribution, edge direction, and local structural features of each region in the image to form an image region feature vector.

3. The adaptive rendering method for AI digital human virtual-real scene fusion Chinese characters and images according to claim 2 is characterized in that: The S3 specifically includes: S31, extracting an attention weight matrix generated by the multi-head attention module based on the fused feature representation, wherein the attention weight matrix represents the matching strength between each word in the text feature vector and each region in the image region feature vector; S32, determining the image area corresponding to each Chinese text based on the maximum weight value corresponding to each text in the attention weight matrix, and establishing a one-to-one mapping relationship between the Chinese text and the image area; S33, extracting the spatial coordinate range of each corresponding image area based on the one-to-one mapping relationship, determining the arrangement position of the Chinese text on the image, and generating a preliminary image-text fusion layout including the display positions of all Chinese texts; S34. Conflict detection and position adjustment processing are performed on the preliminary graphic and text fusion layout, and a minimum spacing threshold is set between any two adjacent Chinese words to ensure readability and non-overlapping areas of the text layout.

4. The adaptive rendering method for AI digital human virtual-real scene fusion Chinese characters and images according to claim 3 is characterized in that: The S4 specifically includes: S41, initializing a population of glowworms, where each glowworm individual corresponds to a set of structural parameter vectors, wherein the structural parameter vectors are composed of three parts: network depth, number of attention heads, and time step control parameters in the neural ordinary differential equation transformer model; S42. Setting a brightness function for each glowworm individual, wherein the brightness function uses the comprehensive performance index of the neural ordinary differential equation transformer model on the validation set as the objective function value, the comprehensive performance index being a weighted combination of three indicators: image and text arrangement accuracy, attention distribution sparsity, and fusion feature stability, where the sum of the weighted coefficients of each indicator is one; S43. In each iteration, based on the relative difference between the individual brightness value and the brightness values ​​of other individuals in the neighborhood, and based on the difference between the current individual structure parameter vector and the structure parameter vector of the brighter individual in the neighborhood, a moving function is constructed through the brightness difference weight and the spatial distance factor to update the network depth, number of attention heads, and time step control parameters of the current individual; S44. Setting a step-size attenuation mechanism and parameter boundary control conditions during the movement process to prevent the structural parameter vector from crossing the boundary and maintain the adaptability and convergence of the search area; S45. When the preset maximum number of iterations or the brightness function convergence condition is reached, the lightworm individual with the optimal brightness value is selected, and the corresponding structural parameter vector is the optimal combination of network depth, number of attention heads, and time step control parameters; S46. Based on the optimal combination, reconstruct the neural ordinary differential equation converter model as the neural ordinary differential equation converter model after structural optimization.

5. The adaptive rendering method for AI digital human virtual-real scene fusion Chinese characters and images according to claim 4 is characterized in that: The S5 graphic rendering instruction set includes: determining the font, color, transparency and display coordinates of each Chinese text in the image based on the structurally optimized neural ordinary differential equation transformer model, constructing graphic rendering instructions containing text style parameters and spatial position parameters, and combining them to generate a graphic rendering instruction set.

6. The adaptive rendering method for AI digital human virtual-real scene fusion Chinese characters and images according to claim 5 is characterized in that: The S6 specifically includes: obtaining the skeleton posture parameters of the digital human in the current frame and the lighting distribution characteristics of the background image, performing coordinate transformation and brightness adjustment based on the posture angle and lighting direction according to the graphic rendering instruction set, and outputting the processed image frame as a virtual-reality fusion rendering frame.

7. An adaptive rendering system for AI digital human virtual-real scene fusion of Chinese characters and images, applied to the adaptive rendering method for AI digital human virtual-real scene fusion of Chinese characters and images according to any one of claims 1 to 6, characterized in that: Includes the following modules: The text feature extraction module is used to perform word segmentation and semantic encoding on Chinese text and generate text feature vectors; Image feature extraction module, used to perform region segmentation and texture extraction on image data and generate image region feature vectors; A fusion modeling module is used to input the text feature vector and the image region feature vector into the neural ordinary differential equation transformer model to generate a fusion feature representation; The attention arrangement module is used to extract the attention weight matrix from the fused feature representation and generate a preliminary image-text fusion layout; Parameter optimization module, used to optimize network depth, number of attention heads, and time step control parameters based on the glowworm swarm optimization algorithm; An instruction generation module is used to generate an image and text rendering instruction set based on the preliminary image and text fusion layout; The posture lighting mapping module is used to combine the digital human's action status with the background image information to generate a virtual-reality fusion rendering frame; The graphic rendering engine module is used to output the virtual-reality fusion rendering frame into video content that is a fusion of Chinese characters and images.

Citation Information

Patent Citations

  • Multi-site water quality prediction method based on spatio-temporal feature fusion

    CN119168176A

  • Dynamic monitoring and management platform based on vision and sensing fusion

    CN119918015A