A method for generating works based on the fusion of virtual and real technologies using AI and MR.
By capturing the spatiotemporal trajectory data of user sketches and performing semantic segmentation, combined with style transfer protocols and feature compensation, graphic works consistent with the user's creative intent are generated. This solves the problems of insufficient sketch feature parsing and shallow cross-modal alignment in existing technologies, and achieves efficient graphic generation and MR virtual-real fusion.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies cannot effectively analyze the spatiotemporal dynamic features of user sketches in the generation of graphic works, resulting in the loss of creative intent during the input stage. Furthermore, the cross-modal alignment of graphics and text is superficial, failing to achieve a deep mapping between the visual features of the sketch and the abstract semantics of the text.
By capturing the spatiotemporal trajectory data of user sketches, a raw dataset containing brush pressure and coordinate sequences is generated. Semantic segmentation processing is then performed, and combined with style transfer protocols and feature compensation sub-processes, the dynamic features of the sketches are mapped to high-level semantics. Adaptive text is then generated, and virtual-real fusion is achieved using MR technology.
It achieves end-to-end mapping from low-level brush stroke features of sketches to high-level semantics, generating graphic works that are highly consistent with the user's creative intent, breaking through the bottleneck of semantic fragmentation, and achieving real-time optimization through interactive commands, filling the technical gap in physical interaction scenarios.
Smart Images

Figure CN120833455B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent digital content generation technology, and relates to a method for generating works based on the fusion of virtual and real technologies using AI and MR. Background Technology
[0002] In recent years, the integration of artificial intelligence and mixed reality technologies has brought about a new revolution in digital art creation. AI technology, with its deep learning capabilities, has demonstrated powerful advantages in image generation, style transfer, and semantic understanding, while MR technology, through its interactive mode of overlaying virtual and real elements, significantly enhances the immersive experience of users in the creative process.
[0003] In this wave of technological advancement, the generation of intelligent graphic and textual works has gradually become a research hotspot. This direction focuses on using AI to automatically generate text that highly matches visual content and leveraging MR technology to achieve an immersive display effect that blends the virtual and real worlds. However, facing users' ever-increasing demand for personalized, high-quality content, how to efficiently generate graphic and textual works that accurately adapt to user intent has become a key area where current technology urgently needs breakthroughs.
[0004] Existing technologies also include solutions related to the generation of graphic and textual works. For example, the graphic and textual content generation method, apparatus, device, and storage medium in Chinese Patent Publication No. CN117032869A extracts key information from the text to be accompanied by an image, calls a pre-configured text-image generation model, utilizes its ability to generate semantically matching images based on text, generates images that match the key information, and then merges the text and images to obtain graphic and textual content. This process does not require paper books, which can reduce generation costs and ensure the quality of graphic and textual works.
[0005] Another Chinese patent publication, CN118484511A, describes a creative support method and system for artistic sketching based on process awareness and narrative-driven techniques. It first acquires an initial artistic sketch based on line drawing, allows the user to copy and modify the sketch, constructs an artistic sketch library and narrative text, modifies the next timeline content of the sketch based on image change keywords and stores it, constructs a global view with sketches as nodes, timelines and text as edges, and provides a corresponding system to help users imitate and learn, providing creative support.
[0006] While the above solutions propose methods for generating graphic works, existing technologies still have the following limitations:
[0007] 1. Current mainstream image generation relies heavily on text descriptions or complete sketch inputs, which cannot effectively analyze the spatiotemporal dynamic features of user sketch strokes. Consequently, it cannot transform the high-level semantics such as compositional logic and brushstroke emotion contained in the sketch outline into generating constraints. This results in the loss of semantics of the user's creative intent at the input stage, and the generated result deviates significantly from the original concept.
[0008] 2. Existing technologies suffer from shallow alignment defects in cross-modal text and image. Text generation generally adopts keyword matching mechanisms, such as extracting object labels from images and then retrieving related words. However, it lacks deep alignment between the overall artistic conception of the image and the tone of the text content, and thus cannot achieve synesthetic mapping between the visual features of the sketch and the abstract semantics of the text. Summary of the Invention
[0009] In view of this, in order to solve the problems mentioned in the background technology, a method for generating works based on the fusion of virtual and real technologies of AI and MR is proposed.
[0010] The objective of this invention can be achieved through the following technical solution: This invention provides a method for generating works based on the fusion of AI and MR technologies, including: capturing the spatiotemporal trajectory data of a user's sketch through an input device, and generating an original dataset containing pen pressure and coordinate sequences.
[0011] The original dataset is subjected to semantic segmentation. If the segmentation result meets the matching criteria of any template in the predefined style template library, the style transfer protocol associated with that template is output. Otherwise, the feature compensation sub-process is triggered to provide the user with a brushstroke integrity prompt and wait for supplementary input.
[0012] The user sketch is stylized and rendered according to the style transfer protocol, and corresponding adaptive mood text is generated. The spatial position of the adaptive mood text in the user sketch is determined by combining the preset visual focus arrangement rules, thereby generating a text-image binding scheme.
[0013] Receive user interaction instructions to determine the object to be corrected in the image and text binding scheme and perform the correction. Continuously receive and process user interaction instructions through a feedback iteration mechanism until there is no need for correction, then terminate the iteration and output the final image and text binding scheme.
[0014] Spatial coordinate encoding is performed on the final image and text binding scheme to generate a physical medium that can trigger MR virtual-real fusion.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0016] (1) This invention captures the spatiotemporal trajectory data of user sketches, performs semantic segmentation based on pressure behavior, geometric topology and dynamic behavior features, transforms the dynamic features of the sketches into generation constraints of style transfer protocols and assists in stylized rendering of the sketches, realizes end-to-end mapping from low-level brush stroke features of sketches to high-level semantics, helps to restore the user's creative intent, avoids the loss of initial input information, and promotes the generation results to be highly consistent with the concept.
[0017] (2) This invention generates adapted text that matches the artistic conception of the stylized sketch by quantifying the entity coverage, style similarity and structural fit of each poem in the poetry library. This breaks through the semantic separation bottleneck of existing image and text generation. Through data-driven quantitative analysis, the artistic conception of the text and the visual style can achieve a dynamic balance.
[0018] (3) This invention not only automatically identifies missing pen stroke semantics and guides users to complete them through feature compensation sub-process, but also achieves real-time iterative optimization based on interactive instructions, solving the problem of rigid processes in existing technologies. At the same time, by combining spatial coordinate encoding of physical media with MR interaction, the virtual creation content is deeply bound to the physical carrier, filling the technical gap of existing solutions lacking physical interaction scenarios. Attached Figure Description
[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart illustrating the implementation steps of the method of the present invention.
[0021] Figure 2 This is a logical diagram illustrating the semantic segmentation process of the original dataset in this invention.
[0022] Figure 3 This is a schematic diagram of the adaptive contextual text generation logic of the present invention. Detailed Implementation
[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] Please see Figure 1 As shown, the present invention provides a method for generating works based on the fusion of virtual and real technologies of AI and MR, including: S1. Capturing the spatiotemporal trajectory data of the user's sketch through an input device to generate an original dataset containing pen pressure and coordinate sequence.
[0025] It should be noted that the above input devices are digital drawing tablets or styluses with pressure-sensitive characteristics.
[0026] In a preferred embodiment of the present invention, the step of capturing the spatiotemporal trajectory data of the user's sketch through the input device includes: using a single complete pen movement as the processing benchmark, defining the continuous trajectory formed by the user from putting down the pen to lifting the pen as a pen touch line unit.
[0027] Assign a unique identifier to each pen stroke unit during the user sketching process.
[0028] The pen pressure value and absolute coordinate value of each sampling point in the pen touch line unit are stored in chronological order. Pen touch continuity verification is performed synchronously to eliminate trajectory breaks, and an original dataset is generated with the pen touch line unit as the index.
[0029] It should be noted that the absolute coordinate values of each sampling point within the above-mentioned pen stroke unit are converted in real time from the continuous pen stroke captured in a single instance into relative coordinate values with the center of the physical medium as the origin through the input device's built-in absolute coordinate system.
[0030] It should also be noted that the above-mentioned pen stroke continuity verification process includes: i. When the distance between adjacent coordinate points exceeds the minimum displacement resolution of the input device, supplementary points are inserted in a linear interpolation manner.
[0031] ii. When the time interval between lifting the pen and putting it down again is less than or equal to the preset physiological duration threshold for continuous pen strokes, and the distance between the new landing point and the end point of the previous stroke is less than the preset multiple of the coordinate positioning error of the input device, the two trajectories are merged into the same pen stroke unit.
[0032] The minimum displacement resolution and coordinate positioning error of the input device can be directly extracted and used by referring to the standard values in the technical specification guide of the input device.
[0033] S2. Perform semantic segmentation on the original dataset. If the segmentation result meets the matching criteria of any template in the predefined style template library, output the style transfer protocol associated with that template. Otherwise, trigger the feature compensation sub-process, provide the user with a brushstroke integrity prompt, and wait for supplementary input.
[0034] Please see Figure 2 As shown, in a preferred embodiment of the present invention, semantic segmentation processing is performed on the original dataset, including: based on the pen pressure sequence, coordinate sequence and timestamp information of each pen touch line unit, parallel analysis of the pressure behavior features, geometric topology features and dynamic motion features of each pen touch line unit.
[0035] It should be noted that the above-mentioned pressure behavior characteristics include the pressure mean, pressure fluctuation amplitude, and pressure change trend factor, which are obtained by fitting the arithmetic mean, standard deviation, and curve slope of the pen pressure values of each unit sampling point within the pen stroke line unit.
[0036] Geometric topological features include closure, tilt deviation, and parallel spacing. Closure is the ratio of the Euclidean distance between the first and last points of a pen-line unit to the total length of the unit. Tilt deviation is the standard deviation of the tangent direction angle of each sampling point within the pen-line unit. Parallel spacing specifically refers to the average vertical distance between two pen-line units when there are adjacent parallel pen-line units within a preset spatial domain. When there are no adjacent parallel pen-line units, the parallel spacing value is automatically set to 0.
[0037] Dynamic motion features include the mean velocity, the percentage of zero-velocity points, and the extreme value of acceleration during the process of the user drawing pen stroke lines. The mean velocity is the arithmetic mean of the movement rates of all unit sampling points within the line. The percentage of zero-velocity points is the percentage of sampling points whose movement rate is lower than the minimum displacement resolution of the input device. The extreme value of acceleration is the maximum absolute value of the velocity change between adjacent sampling points.
[0038] The feature groups obtained from the analysis are compared with the standard conditions of each type of line in the preset stroke semantic library, where each type of line includes contour lines, structural lines, shadow lines, and highlight lines.
[0039] It should be noted that the specifications for the above-mentioned types of lines are based on hard parameter settings in the pressure behavior, geometric topology, and dynamic motion characteristics of the pen stroke line unit.
[0040] For example, the criteria for determining the contour include closure, average rate, and pressure fluctuation amplitude.
[0041] The criteria for judging high-brightness lines emphasize high average pressure, high acceleration extreme values that reflect smoothness, and parallel spacing that conforms to design specifications.
[0042] The criteria for determining structural lines focus on geometric and topological characteristics, requiring that the tilt angle deviation be below the critical value, the parallel spacing meet the standard, and the average rate be stable within a specific range.
[0043] The criteria for determining the shaded line emphasize low average pressure, regularly changing pressure fluctuations, and a zero-velocity point ratio higher than the threshold to reflect the effect of drawing pauses.
[0044] Highlights require the average pressure to be in the high intensity range and the extreme acceleration value to express the smoothness of fast drawing. At the same time, the parallel spacing must conform to the density design specifications of highlight lines.
[0045] If the pen contact line feature group fully satisfies all the specification conditions of a certain type of line, then the corresponding type identifier is directly assigned to the pen contact line unit.
[0046] If the specifications for multiple line types are met simultaneously, then a type identifier is assigned to the pen stroke unit according to the preset priority order.
[0047] It should be noted that the above-preset priority order is outline, structural lines, shadow lines, and highlights. This priority order is mainly based on the underlying rules of drawing logic and the scientific layering of visual information weights. Specifically, outlines define the existence and basic shape of an object and are the starting point of human visual cognition, thus having absolute priority. Structural lines explain the internal structure and spatial relationships of an object and are a secondary priority for ensuring the stability of the sketch object. Shadow lines convey the texture of the object and the direction of light, affecting the sense of three-dimensionality but not a core element of recognition. Highlights only modify reflective properties and visual focus; their absence does not affect the essential perception of the object, therefore they have the lowest priority.
[0048] If the specification conditions for any type of line cannot be met, the pen contact line unit is marked as an undefined type and the feature compensation subprocess is triggered.
[0049] In a preferred embodiment of the present invention, the segmentation result satisfies the matching criteria of a predefined style template, including the following verification: a) verifying whether the segmentation result contains the stroke type required by the predefined style template.
[0050] b) Extract the spatial topology rule set of the pen strokes of the predefined style template, which includes spatial position constraints and connection type constraints, and verify whether the relative positional relationship and connection combination between pen stroke units in the segmentation result conforms to the rule set and whether the proportion of conformity is greater than the preset compliance ratio threshold.
[0051] It should be noted that the aforementioned spatial position constraints and connection type constraints refer to the numerical restrictions on relative position parameters and the restrictions on connection type identifiers, respectively. The relative position parameters include at least distance, included angle, and overlap rate, and the connection type identifiers include at least endpoint contact type and intersection status.
[0052] The process of obtaining the proportion of relative positional relationships and connection combinations between pen-stroke units in the segmentation result that conform to the rule set includes: performing pairwise combination analysis on all pen-stroke units in the segmentation result; if multiple line connections are involved, the analysis is expanded to tuple analysis; extracting relative positional parameters and connection type identifiers for each combination; comparing them one by one with the pen-stroke spatial topology rule set of the predefined style template; counting the number of pen-stroke units in the combination that satisfies the rule set and using it as the numerator; and performing a ratio calculation with the total number of pen-stroke units as the denominator to obtain the proportion of relative positional relationships and connection combinations between pen-stroke units in the segmentation result that conform to the rule set.
[0053] c) Based on the density of each type of pen stroke in the segmentation result and its relative area coverage of the user sketch, verify whether the visual expression information of the user sketch meets the information requirements of the predefined style template.
[0054] It should be noted that the process of obtaining the density of each type of pen stroke and its relative area coverage of the user sketch is as follows: target pen stroke units are selected from the segmentation results, the cumulative drawing length of the target pen stroke units is calculated by spatial coordinate integration, and the ratio of the cumulative drawing length to the area of the sketch area is used as the density.
[0055] The sketch is processed into a unit grid. When the proportion of the area covered by the pen strokes in a unit grid exceeds the preset area proportion threshold, the unit grid is marked as a covered grid. The proportion of covered grids is counted and used as the area coverage rate.
[0056] It should also be noted that the information requirements of the above-mentioned predefined style templates include the specification requirements for the density of each type of stroke and its relative area coverage of the user sketch.
[0057] When all three verifications (a), (b), and (c) pass, the segmentation result is determined to meet the matching criteria of the predefined style template.
[0058] In a preferred embodiment of the present invention, the feature compensation subprocess supports incremental data acquisition, including: capturing sketch stroke data supplemented by the user.
[0059] The supplementary sketch stroke data is then fused with the original dataset using feature fusion.
[0060] A second semantic segmentation process is performed on the merged dataset.
[0061] S3. Stylize the user sketch according to the style transfer protocol and generate corresponding adapted mood text. Combine the preset visual focus arrangement rules to determine the spatial position of the adapted mood text in the user sketch, thereby generating a text-image binding scheme.
[0062] In a preferred embodiment of the present invention, the process of stylizing and rendering a user sketch includes: parsing the typed rendering control vector contained in the style transfer protocol, wherein the typed rendering control vector is a set of rendering parameters configured for different types of brush strokes, and the rendering parameters include at least brush granularity, texture density, and light and shadow gradient.
[0063] Based on the typified rendering control vector, directional rendering transformation is performed on each type of stroke line after semantic segmentation to obtain a stylized rendering sketch.
[0064] This invention captures the spatiotemporal trajectory data of user sketches, performs semantic segmentation based on pressure behavior, geometric topology, and dynamic behavior features, transforms the dynamic features of the sketches into generation constraints of style transfer protocols, and assists in stylized rendering of the sketches. This achieves end-to-end mapping from low-level brush stroke features to high-level semantics, helps to restore the user's creative intent, avoids the loss of initial input information, and promotes a high degree of consistency between the generated results and the concept.
[0065] Please see Figure 3 As shown, in a preferred embodiment of the present invention, the process of generating adapted mood text includes: performing multi-semantic recognition on the region formed by the closure of contour lines in the stylized rendering sketch, and constructing an element set containing each semantic recognition element and its confidence level.
[0066] Retrieve the poetry library stored in the web cloud and quantify the entity coverage of each poem in the poetry library relative to the set of elements.
[0067] It should be noted that the specific quantification process of the entity coverage of each poem in the above-mentioned poetry library relative to the set of elements includes: comparing the entity keywords involved in the content of each poem in the poetry library with the set of elements one by one, accumulating the confidence of each semantic recognition element covered by the content of the poem, and using the accumulated result as the entity coverage, thereby obtaining the entity coverage of each poem in the poetry library relative to the set of elements.
[0068] Based on the style tags corresponding to the predefined style template to which the user sketch belongs, the similarity between the pre-stored style tags of each poem in the poetry library and the style tags of the user sketch is quantified.
[0069] It should be noted that the above can convert the style tags corresponding to the predefined style template to which the user sketch belongs, as well as the pre-stored style tags of each poem in the poetry library, into style tag vectors of equal dimensions through one-hot encoding or word embedding technology. The similarity between the pre-stored style tags of each poem in the poetry library and the style tags of the user sketch is then quantified by the cosine similarity calculation formula.
[0070] The structural fit between the brushstroke distribution entropy of the stylized rendering sketch and the tonal entropy of each poem in the poetry library is analyzed.
[0071] It should be noted that the above can be achieved by mapping the brushstroke distribution entropy of the stylized rendering sketch into a two-dimensional spatial entropy matrix, converting the tonal entropy of poetry into a time series entropy vector, using a dynamic time warping algorithm to calculate the minimum cumulative distance between the two entropy structures, and obtaining the structural fit index through normalization.
[0072] The process of analyzing the stroke distribution entropy of the stylized rendering sketch is as follows: statistically analyze the spatial distribution density and directional distribution dispersion of each type of stroke line unit in the stylized rendering sketch, construct a two-dimensional correlation matrix of stroke type-spatial position and a frequency histogram of stroke direction change rate, quantify the entropy index corresponding to the correlation matrix and the frequency histogram respectively using the Shannon entropy formula, and calculate the stroke distribution entropy of the stylized rendering sketch by arithmetic mean.
[0073] The process of parsing the tonal entropy of each poem in the poetry database is as follows: the poem text is mapped into a standardized sequence of tonal symbols, the global distribution entropy is calculated by statistically analyzing the symbol frequencies, the positional distribution entropy is obtained based on the positional stratification frequency, the transition entropy is solved using a binary transition probability matrix, and the arithmetic mean of the global distribution entropy, positional distribution entropy, and transition entropy is calculated to obtain the tonal entropy of each poem in the poetry database.
[0074] By linearly weighting and fusing entity coverage, style similarity, and structural fit, the matching degree of each poem in the poetry library relative to the stylized rendering sketch is obtained, and the poem with the highest matching degree is selected as the text that matches the artistic conception.
[0075] It should be added that the linear weighting of entity coverage, style similarity, and structural fit can be set based on industry experience or obtained based on data-driven methods. Specifically, in scenarios such as image generation and design evaluation, entity coverage, style similarity, and structural fit of different samples are collected, the correlation coefficients of the three factors on the final evaluation target are calculated, regression analysis or logistic regression models are used to quantify the contribution of each indicator to the evaluation results, and after normalization, the contribution is converted into linear weighting with a sum of 1.
[0076] It should be noted that the style similarity mentioned above refers to the similarity between the pre-stored style tags of poems in the poetry library and the style tags of user sketches.
[0077] This invention generates adapted text that matches the artistic conception of the stylized sketch by quantifying the entity coverage, style similarity, and structural fit of each poem in the poetry library. This breaks through the semantic fragmentation bottleneck of existing image and text generation, and achieves a dynamic balance between text artistic conception and visual style through data-driven quantitative analysis.
[0078] In a preferred embodiment of the present invention, determining the spatial position of the appropriate mood text in the user sketch based on a preset visual focus arrangement rule includes: performing adaptive grid division on the user sketch to obtain multiple grid sub-regions.
[0079] Obtain the distribution quantity and area ratio of semantic recognition elements in each grid sub-region, quantify the semantic clustering evaluation index of each grid sub-region, and arrange them in ascending order of clustering degree to form a localization retrieval sequence.
[0080] Adjacent blank area connectivity detection is performed sequentially on the grid sub-regions in the positioning retrieval sequence. When the spatial range of a certain continuous region meets the preset placement requirements of the contextual text and the region boundary maintains a preset safe distance from the highly visually salient grid sub-region, the geometric centroid coordinates of the continuous region are determined as the spatial position of the contextual text in the user sketch.
[0081] It should be noted that the aforementioned highly visually saliency grid sub-regions specifically refer to grid sub-regions whose semantic clustering evaluation index is greater than the preset threshold for visual focus.
[0082] S4. Receive user interaction instructions to determine the object to be corrected in the image and text binding scheme and perform the correction. Continuously receive and process user interaction instructions through a feedback iteration mechanism until there is no need for correction, then terminate the iteration and output the final image and text binding scheme.
[0083] In a preferred embodiment of the present invention, the user interaction instructions include at least one of the following: a position offset instruction or a content addition instruction for the adapted contextual text.
[0084] Instructions for replacing local brushstrokes or adjusting rendering parameters on the stylized rendering sketch.
[0085] In a preferred embodiment of the present invention, determining the object to be corrected for the image-text binding scheme and performing the correction includes: if a position offset instruction for the adapted contextual text is received, obtaining the drag vector of the text anchor point by the user, calculating the new anchor point position coordinates based on the drag vector, and using the new anchor point position coordinates as the corrected spatial position of the adapted contextual text.
[0086] If an instruction to add content to the adapted mood text is received, the latest set of keywords input by the user is parsed, and the semantic similarity between the content of each poem and the input set of keywords is calculated one by one according to the order of matching degree between the poems in the poetry library and the stylized rendering sketch from high to low. The first poem whose semantic similarity reaches the preset threshold is taken as the corrected adapted mood text.
[0087] If a local brushstroke replacement instruction is received from a stylized rendering sketch, the brushstroke area selected by the user is identified and the semantic segmentation type of the reconstructed brushstroke line unit is extracted. The targeted replacement is then carried out according to the corresponding type of brushstroke line configuration in the predefined style template.
[0088] If a rendering parameter adjustment instruction is received from a stylized rendering sketch, then the rendering parameter adjustment operation is responded to by performing an equivalent replacement of the rendering effect within the adjustment area.
[0089] S5. Encode the final image and text binding scheme using spatial coordinates to generate a physical medium that can trigger MR virtual-real fusion.
[0090] The embodiments of the present invention not only automatically identify missing pen stroke semantics and guide users to complete them through feature compensation sub-process, but also achieve real-time iterative optimization based on interactive instructions, solving the problem of rigid processes in existing technologies. At the same time, by combining spatial coordinate encoding of physical media with MR interaction, virtual creation content is deeply bound to physical carriers, filling the technical gap in existing solutions that lack physical interaction scenarios.
[0091] The above content is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined by the present invention, and all such modifications and additions should fall within the protection scope of the present invention.
Claims
1. A method for generating works based on the fusion of virtual and real technologies using AI and MR, characterized in that, include: The spatiotemporal trajectory data of the user's sketch is captured by the input device, generating a raw dataset containing pen pressure and coordinate sequences. The process of capturing the spatiotemporal trajectory data of the user's sketch via the input device includes: Based on a single complete pen stroke, the continuous trajectory formed by the user from putting down the pen to lifting it is defined as a pen stroke line unit. Assign a unique identifier to each pen stroke unit in the user sketching process; The pen pressure value and absolute coordinate value of each unit sampling point in the pen touch line unit are stored in chronological order, and the pen touch continuity verification is performed synchronously to eliminate trajectory breakage, generating an original dataset indexed by the pen touch line unit. The original dataset is subjected to semantic segmentation. If the segmentation result meets the matching criteria of any template in the predefined style template library, the style transfer protocol associated with that template is output. Otherwise, the feature compensation sub-process is triggered to provide the user with a brushstroke integrity prompt and wait for supplementary input. Semantic segmentation processing is performed on the original dataset, including: based on the pen pressure sequence, coordinate sequence and timestamp information of each pen touch line unit, parallel analysis of the pressure behavior features, geometric topology features and dynamic motion features of each pen touch line unit; The feature groups obtained by parsing are compared with the normative conditions of each type of line in the preset stroke semantic library, where each type of line includes contour lines, structural lines, shadow lines, and highlight lines. If the pen contact line feature group fully satisfies all the specification conditions of a certain type of line, then the corresponding type identifier is directly assigned to the pen contact line unit. If the specification conditions for multiple line types are met simultaneously, then a type identifier is assigned to the pen contact line unit according to the preset priority order. If the specification conditions for any type of line cannot be met, the pen contact line unit is marked as undefined type and the feature compensation subprocess is triggered; The user sketch is stylized and rendered according to the style transfer protocol, and corresponding adaptive mood text is generated. The spatial position of the adaptive mood text in the user sketch is determined by combining the preset visual focus arrangement rules, thereby generating a text-image binding scheme. Receive user interaction instructions to determine the object to be corrected in the image and text binding scheme and perform the correction. Continuously receive and process user interaction instructions through a feedback iteration mechanism until there is no need for correction, then terminate the iteration and output the final image and text binding scheme. Spatial coordinate encoding is performed on the final image and text binding scheme to generate a physical medium that can trigger MR virtual-real fusion.
2. The method for generating works based on the fusion of AI and MR technologies according to claim 1, characterized in that: The segmentation result satisfies the matching criteria of the predefined style template, including the following verification content: a) Verify whether the segmentation result contains the stroke type required by the predefined style template; b) Extract the spatial topology rule set of the pen strokes of the predefined style template, which includes spatial position constraints and connection type constraints, and verify whether the proportion of the relative positional relationship and connection combination between pen stroke units in the segmentation result that conforms to the rule set is greater than the preset compliance ratio threshold. c) Based on the density of each type of pen stroke in the segmentation results and its relative area coverage of the user sketch, verify whether the visual expression information of the user sketch meets the information requirements of the predefined style template. When all three verifications (a), (b), and (c) pass, the segmentation result is determined to meet the matching criteria of the predefined style template.
3. The method for generating works based on the fusion of AI and MR technologies according to claim 1, characterized in that: The feature compensation subprocess supports incremental data acquisition, including: Capture the brush stroke data of the sketches added by the user; The supplementary sketch stroke data is then fused with the original dataset using feature fusion. A second semantic segmentation process is performed on the merged dataset.
4. The method for generating works based on the fusion of AI and MR technologies according to claim 1, characterized in that: The process of stylizing and rendering user sketches includes: The style transfer protocol contains typed rendering control vectors, which are sets of rendering parameters configured for different types of brush strokes. The rendering parameters include at least brush granularity, texture density, and lighting gradient. Based on the typified rendering control vector, directional rendering transformation is performed on each type of stroke line after semantic segmentation to obtain a stylized rendering sketch.
5. The method for generating works based on the fusion of AI and MR technologies according to claim 4, characterized in that: The process of generating contextually appropriate text includes: Perform multi-semantic recognition on the regions formed by the closure of contour lines in the stylized rendering sketch, and construct an element set containing each semantic recognition element and its confidence level; Retrieve the poetry library stored in the WEB cloud and quantify the entity coverage of each poem in the poetry library relative to the set of elements. Based on the style tags corresponding to the predefined style template to which the user sketch belongs, quantify the similarity between the pre-stored style tags of each poem in the poetry library and the style tags of the user sketch. Analyze the structural fit between the brushstroke distribution entropy of the stylized rendering sketch and the tonal entropy of each poem in the poetry library; By linearly weighting and fusing entity coverage, style similarity, and structural fit, the matching degree of each poem in the poetry library relative to the stylized rendering sketch is obtained, and the poem with the highest matching degree is selected as the text that matches the artistic conception.
6. The method for generating works based on the fusion of AI and MR technologies according to claim 5, characterized in that: The step of determining the spatial position of the appropriate mood text in the user sketch based on preset visual focus arrangement rules includes: Adaptive mesh generation is performed on the user sketch to obtain multiple mesh sub-regions; Obtain the distribution quantity and area ratio of semantic recognition elements in each grid sub-region, quantify the semantic clustering evaluation index of each grid sub-region, and arrange them in ascending order of clustering degree to form a localization retrieval sequence; Adjacent blank area connectivity detection is performed sequentially on the grid sub-regions in the positioning retrieval sequence. When the spatial range of a certain continuous region meets the preset placement requirements of the contextual text and the region boundary maintains a preset safe distance from the highly visually salient grid sub-region, the geometric centroid coordinates of the continuous region are determined as the spatial position of the contextual text in the user sketch.
7. The method for generating works based on the fusion of AI and MR technologies according to claim 5, characterized in that: The user interaction instructions include at least one of the following: a position offset instruction or a content addition instruction for the adapted mood text; Instructions for replacing local brushstrokes or adjusting rendering parameters on the stylized rendering sketch.
8. The method for generating works based on the fusion of AI and MR technologies according to claim 7, characterized in that: The step of determining the object to be corrected for the image-text binding scheme and performing the correction includes: if a position offset instruction for the adapted mood text is received, then the drag vector of the user on the text anchor point is obtained, the new anchor point position coordinates are calculated based on the drag vector, and the new anchor point position coordinates are used as the corrected adapted mood text spatial position. If an instruction to add content to the adapted mood text is received, the latest set of keywords input by the user is parsed, and the semantic similarity between the content of each poem and the set of input keywords is calculated one by one according to the order of matching degree between the poems in the poetry library and the stylized rendering sketch from high to low. The first poem whose semantic similarity reaches the preset threshold is taken as the corrected adapted mood text. If a local brushstroke replacement instruction is received from a stylized rendering sketch, the brushstroke area selected by the user is identified and the semantic segmentation type of the reconstructed brushstroke line unit is extracted. The targeted replacement is then carried out according to the corresponding type of brushstroke line configuration in the predefined style template. If a rendering parameter adjustment instruction is received from a stylized rendering sketch, then the rendering parameter adjustment operation is responded to by performing an equivalent replacement of the rendering effect within the adjustment area.
Citation Information
Patent Citations
Image-text content generation method and device, equipment and storage medium
CN117032869A
Art sketch drawing creative support method and system based on process awareness and narrative driving
CN118484511A
Handwriting generation method and device, equipment and storage medium
CN113743302A
Digital work display method based on combination of artificial intelligence and mixed reality technology
CN119357479A