Character shape generation method and device based on time sequence consistency, equipment and medium
By dynamically updating styles and adjusting for temporal sequence, combined with style memory and discriminator evaluation, the problem of insufficient temporal consistency in traditional text generation methods is solved, achieving temporal consistency and style continuity in text shape generation, and improving the authenticity and standardization of the generated results.
Patent Information
- Application Number
- CN202510722341.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-16
AI Technical Summary
Traditional text generation methods struggle to maintain temporal consistency in text shapes in areas such as dynamic font synthesis, calligraphy animation, and stroke order reconstruction. This results in a lack of realism and fluidity in the generated results, impacting doctor-patient communication, the readability of medical documents, and the security and credibility of financial transactions.
By acquiring the initial text sequence for dynamic style updates, combining historical information from the style memory bank, performing time-aware adjustments and weighted summation of loss values, and using a preset discriminator to evaluate temporal consistency and authenticity, the generated text shapes are optimized.
It achieves temporal consistency and style continuity in the text shape generation process, improves the authenticity and standardization of the generated results, and enhances the credibility of medical documents and the security of financial documents.
Smart Images

Figure CN120655770A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, device, equipment and medium for generating text shapes based on temporal consistency. Background Art
[0002] The text generation process not only involves restoring static glyphs and expressing their style, but also has important applications in dynamic generation and temporal consistency, particularly in areas such as dynamic font synthesis, calligraphy animation, and stroke order reconstruction. Traditional text generation methods often focus solely on the outline and structure of individual characters, lacking the ability to model temporal information during the writing process (such as the order of strokes, dynamic changes in brush strokes, and changes in writing speed and force). As a result, while the generated results resemble the shape, they fail to truly reproduce the flow and rhythm of the writing process. Therefore, the lack of temporal consistency has become a major bottleneck restricting the development of traditional text generation technology towards higher quality and more realistic expression.
[0003] In the field of healthcare, traditional text generation methods are unable to accurately restore the personalized writing style of doctors' handwritten medical records and prescriptions. As a result, in applications such as medical record digitization and intelligent assisted recording, it is difficult to balance the standardization and authenticity of writing, affecting doctor-patient communication and the readability and authority of medical documents.
[0004] In the field of financial technology, traditional methods have difficulty maintaining consistency in writing style and detail in scenarios such as bill approval and handwritten signature generation. In particular, in seal simulation and anti-counterfeiting verification, they cannot meet the high-security and high-fidelity business needs, which restricts the effectiveness and credibility of intelligent and automated processing.
[0005] In summary, traditional text generation methods mainly rely on glyph outline interpolation or rule-based structural deformation. Although these methods can generate basic glyphs, they find it difficult to capture complex writing styles (such as the flying white effect of calligraphy strokes) and maintain style consistency during sequence generation.
[0006] Therefore, in the current technology, there is a problem that it is difficult to maintain consistency in the temporal sequence during the generation of text shapes. Summary of the Invention
[0007] The present invention provides a method, device, equipment and medium for generating character shapes based on temporal consistency, the main purpose of which is to solve the problem that temporal consistency is difficult to maintain during the generation process of character shapes.
[0008] In a first aspect, to achieve the above-mentioned purpose, the present invention provides a method for generating character shapes based on temporal consistency, comprising:
[0009] Obtaining an initial text sequence, and dynamically updating the style of the initial text sequence to obtain a text style code;
[0010] Performing temporal perception adjustment on the text style code to obtain a text image sequence;
[0011] Analyzing the temporal consistency and the authenticity of individual words of the text image sequence using a preset discriminator to obtain a discrimination result;
[0012] Calculating an adversarial loss value, a temporal loss value, and an authenticity loss value of the text image sequence according to the discrimination result, and performing a weighted summation of the adversarial loss value, the temporal loss value, and the authenticity loss value to obtain a total loss value;
[0013] The total loss value is used to perform style optimization on the text image sequence to obtain a target text shape.
[0014] In a second aspect, the present invention further provides a device for generating character shapes based on temporal consistency, comprising:
[0015] A style updating module is used to obtain an initial text sequence, dynamically update the style of the initial text sequence, and obtain a text style code;
[0016] A timing adjustment module, configured to perform timing-aware adjustment on the text style code to obtain a text image sequence;
[0017] An image analysis module is used to analyze the temporal consistency and the authenticity of individual words of the text image sequence using a preset discriminator to obtain a discrimination result;
[0018] a loss calculation module, configured to calculate an adversarial loss value, a temporal loss value, and an authenticity loss value of the text image sequence according to the discrimination result, and perform a weighted sum of the adversarial loss value, the temporal loss value, and the authenticity loss value to obtain a total loss value;
[0019] A style optimization module is used to perform style optimization on the text image sequence using the total loss value to obtain a target text shape.
[0020] In a third aspect, the present invention further provides an electronic device, comprising:
[0021] at least one processor; and,
[0022] a memory communicatively connected to the at least one processor; wherein,
[0023] The memory stores a computer program that can be executed by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can execute the above-mentioned method for generating a character shape based on temporal consistency.
[0024] In a fourth aspect, the present invention also provides a computer-readable storage medium, in which at least one computer program is stored. The at least one computer program is executed by a processor in an electronic device to implement the above-mentioned method for generating text shapes based on temporal consistency.
[0025] The present invention obtains an initial text sequence, performs dynamic style update on the initial text sequence, and obtains a text style code. While retaining the basic structural features of the text, the present invention combines the historical style information accumulated in the style memory library to achieve gradual optimization and personalized adjustment of the style features, ensuring that the generated text style code has both overall style consistency and fine-grained style differences. The text style code is adjusted temporally to obtain a text image sequence, which can achieve smooth transition and dynamic change of style over time, ensuring that the generated text image sequence has both consistency and continuity in style. The preset discriminator is used to analyze the temporal consistency and single-word authenticity of the text image sequence to obtain a discrimination conclusion. The results can comprehensively evaluate the generation effect from both the global and local levels, ensure the consistency of the image sequence in style continuity and time coherence, and at the same time ensure the authenticity and standardization of each single word in details such as structure and strokes. The adversarial loss value, temporal loss value and authenticity loss value of the text image sequence are calculated according to the discrimination results, and the adversarial loss value, the temporal loss value and the authenticity loss value are weighted and summed to obtain the total loss value. Each loss item is flexibly adjusted according to actual needs to avoid distortion or style drift caused by single indicator optimization. The total loss value is used to perform style optimization on the text image sequence to obtain the target text shape, which can keep the temporal sequence consistent during the generation process of the text shape. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0027] Figure 1 Schematic diagram of an application environment of a method for generating character shapes based on temporal consistency according to an embodiment of the present invention;
[0028] Figure 2 A flowchart of a method for generating character shapes based on temporal consistency according to an embodiment of the present invention is provided;
[0029] Figure 3A schematic diagram of a process for obtaining a target character shape in a character shape generation method based on temporal consistency provided by an embodiment of the present invention;
[0030] Figure 4 A schematic diagram of a module of a device for generating character shapes based on temporal consistency provided by one embodiment of the present invention;
[0031] Figure 5 A schematic structural diagram of an electronic device for implementing a method for generating character shapes based on temporal consistency according to an embodiment of the present invention;
[0032] Figure 6 Another structural diagram of an electronic device for implementing a method for generating character shapes based on temporal consistency provided by an embodiment of the present invention.
[0033] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0034] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, and to fully understand and implement how the present disclosure applies technical means to solve technical problems and achieve the corresponding technical effects, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. The embodiments of the present disclosure and the various features in the embodiments can be combined with each other without conflict, and the technical solutions formed are all within the scope of protection of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work should fall within the scope of protection of the present disclosure.
[0035] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, apparatus, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0036] The embodiment of the present application provides a method for generating text shapes based on temporal consistency, and the execution subject of the method for generating text shapes based on temporal consistency includes but is not limited to at least one of the electronic devices such as a server and a terminal that can be configured to execute the device provided by the embodiment of the present application. In other words, the method for generating text shapes based on temporal consistency can be executed by software or hardware installed on a terminal device or a server device. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0037] The embodiment of the present invention provides a method for generating character shapes based on temporal consistency, which can be applied in the following situations: Figure 1In the application environment. Among them, the client communicates with the server through the network. The server can obtain the initial text sequence through the client, dynamically update the style of the initial text sequence, and obtain the text style code. While retaining the basic structural features of the text, it can combine the historical style information accumulated in the style memory library to achieve gradual optimization and personalized adjustment of the style features, ensure that the generated text style code has both overall style consistency and fine-grained style differences, perform temporal perception adjustment on the text style code, obtain a text image sequence, and achieve smooth transition and dynamic change of style over time, ensure that the generated text image sequence is both consistent and continuous in style, use a preset discriminator to analyze the temporal consistency and single-word authenticity of the text image sequence, and obtain a discrimination result, which can be globally The generation effect is comprehensively evaluated at two levels, namely, the style continuity and temporal coherence of the image sequence are guaranteed, while the authenticity and standardization of each single word in terms of structure, strokes and other details are guaranteed. The adversarial loss value, the temporal loss value and the authenticity loss value of the text image sequence are calculated according to the discrimination result, and the adversarial loss value, the temporal loss value and the authenticity loss value are weighted and summed to obtain the total loss value. Each loss item is flexibly adjusted according to actual needs to avoid distortion or style drift caused by single indicator optimization. The total loss value is used to optimize the style of the text image sequence to obtain the target text shape, so that the temporal sequence in the generation process of the text shape can be kept consistent. Finally, the target text shape output is fed back to the user client. Among them, the client can be but is not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server can be implemented with an independent server or a server cluster composed of multiple servers. The present invention is described in detail below through specific embodiments.
[0038] The following is an explanation of the specification of the present invention. The present invention uses a preset discriminator to analyze the temporal consistency and single-word authenticity of the text image sequence to obtain a discrimination result. It can comprehensively evaluate the generation effect from both global and local levels, ensure the consistency of the image sequence in style continuity and temporal coherence, and at the same time ensure the authenticity and standardization of each single word in details such as structure and strokes. Each loss item is flexibly adjusted according to actual needs, avoiding distortion or style drift caused by single indicator optimization, so that the temporal sequence in the generation process of the text shape remains consistent.
[0039] Reference Figure 2 FIG. 1 is a flow chart of a method for generating character shapes based on temporal consistency according to an embodiment of the present invention. In this embodiment, the method for generating character shapes based on temporal consistency includes:
[0040] S1. Obtain an initial text sequence, perform dynamic style update on the initial text sequence, and obtain a text style code.
[0041] In an embodiment of the present invention, an initial text sequence is input into a mapping network to generate a corresponding basic style code, thereby realizing the conversion from character code to a generative style vector. The basic style code is input into a dynamic style memory mechanism constructed based on a gated recurrent unit (GRU). The style memory state is dynamically updated by combining the historical style state and the current step information in the memory bank. The style adjustment amount is calculated and fused with the basic style code. The influence of the basic and historical styles is balanced through an adjustable fusion coefficient, and finally a text style code with enhanced temporal coherence is obtained.
[0042] In specific healthcare scenarios, this technology is being applied to systems such as TCM AI consultations and personalized health report generation to improve text quality and user experience. For example, patient information and ancient text content are encoded and input into a mapping network to obtain a basic text style vector. This dynamic style memory mechanism combines the previously generated text style (such as the expression style in the previous chapter) with the current case description to dynamically adjust the style memory state. This style adjustment allows for the fusion of the ancient text style and the patient's personalized language, resulting in a coherent and natural language style for the generated medical record, while also providing personalized expression.
[0043] In specific FinTech scenarios, this approach is being applied in intelligent financial report interpretation, investment research report generation, and intelligent financial advisory, improving content consistency and industry adaptability. For example, the mapping network encodes company type, industry characteristics, and analysis theme to generate a basic financial style vector. A dynamic style memory mechanism combines the style of previously generated paragraphs (such as terminology, tone, and expression) with the current paragraph content to dynamically update the style memory state. By adjusting the style, the inherent expression habits of the company / industry are integrated with the current content requirements of the report, achieving consistent and accurate analytical reports.
[0044] In an embodiment of the present invention, the step of obtaining an initial text sequence and dynamically updating the style of the initial text sequence to obtain a text style code includes:
[0045] Mapping the initial character sequence to a preset network, and encoding the mapped initial character sequence to obtain an initial character code;
[0046] Generate a corresponding basic style code according to the initial text code;
[0047] Acquire a preset gated recurrent unit and historical style features of historical texts, and construct a historical text style memory library based on the preset gated recurrent unit and the historical style features;
[0048] Acquiring the historical style state of the historical text style memory library and the initial style state of the basic style code;
[0049] Updating the initial style state using the historical style state to obtain an updated style state;
[0050] Obtaining an adjustment weight matrix and an adjustment bias vector, and calculating a style adjustment amount according to the updated style state, the adjustment weight matrix, and the adjustment bias vector;
[0051] The style adjustment amount and the basic style code are weightedly fused to obtain a text style code.
[0052] In detail, the initial text sequence is input into a preset feature extraction network, and the structural information and semantic features of the text are extracted and represented layer by layer through modules such as convolution, embedding, and position encoding. Finally, the corresponding initial text code is output at the encoding layer to represent the basic font information and stroke structure features of the sequence, providing a semantic basis for subsequent style generation and adjustment.
[0053] Through multi-layer perceptron (MLP), self-attention mechanism or convolutional neural network, global feature information related to font style is extracted, content information is filtered out, and finally a basic style code that can represent the basic font style attributes is generated for subsequent style adjustment and fusion.
[0054] By utilizing the gated recurrent unit (GRU)'s ability to remember and update sequence information, a style memory bank is constructed. By recursively inputting text style codes, historical style features are gradually accumulated and updated, enabling the style memory bank to effectively store and express multi-stage and diverse text style information.
[0055] The historical style state at the current moment is extracted from the historical text style memory library to represent the stored text style context information. At the same time, the initial style state initialized by the basic style code is obtained as the starting reference for the current style generation. The two together serve as the input basis for subsequent style adjustments and dynamic updates, realizing the temporal association and personalized expression of style features.
[0056] Using the historical style state and the initial style state as input, the update gate and reset gate mechanism of the gated recurrent unit (GRU) are used to fuse the historical style memory with the current basic style information, dynamically adjust the feature weights, suppress redundant information, and highlight key style features. Finally, the updated style state representing the style characteristics after temporal evolution is output. The state update calculation formula is as follows:
[0057] h t =GRI(w t-1 ,h t-1 )
[0058] Among them, w t-1 represents the initial style state of the basic style encoding at the t-1th time step, h t-1 represents the historical style state at time step t-1, h t represents the updated style state at time step t.
[0059] The updated style state is used as input, and a weighted calculation is performed on it through a linear transformation. The bias term is combined to perform translation, and the feature information related to the target style difference is extracted and enhanced. Finally, the style adjustment Δw used to correct the basic style encoding is obtained. t , the calculation formula is as follows:
[0060] Δω t =W Δ ·h t +b Δ
[0061] Among them, W Δ Represents the adjustment weight matrix, b Δ Indicates the adjustment bias vector, h t represents the updated style state at time step t.
[0062] In an embodiment of the present invention, performing weighted fusion on the style adjustment amount and the basic style code to obtain the text style code includes:
[0063] Obtaining a memory influence strength parameter, and using the memory influence strength parameter to constrain the style adjustment amount to obtain a target adjustment amount;
[0064] Linearly superimposing the target adjustment amount and the basic style code to obtain a preliminary fusion style code;
[0065] The preliminary fusion style code is subjected to nonlinear transformation to obtain a text style code.
[0066] Specifically, the memory influence strength parameter reflects the degree of influence of memory on style adjustment. The memory influence is incorporated into the adjustment constraint process of the style adjustment amount to ensure that the style adjustment does not deviate from the predetermined range or target. After the constraint processing, the target adjustment amount is obtained. The target adjustment amount is introduced into the basic style encoding to refine and enhance the text style features, ultimately obtaining a text style encoding that combines the basic style and personalized adjustment features. The calculation formula is as follows:
[0067]
[0068] in, represents the basic style encoding at the t-th time step, α represents the memory influence strength parameter, tanh function represents the hyperbolic tangent function, Δwt Represents the amount of style adjustment at time step t. The target adjustment amount is limited to the interval -1, 1. When α = 0, the system is the standard StyleGAN2. Increasing α can enhance style coherence.
[0069] Dynamic style updates of the initial text sequence can preserve the basic structural features of the text while combining the historical style information accumulated in the style memory library to achieve gradual optimization and personalized adjustment of the style features, ensuring that the generated text style encoding has both overall style consistency and fine-grained style differences, thereby improving the naturalness, coherence and diversity of text style generation, and meeting the needs of font design and style migration in complex scenarios.
[0070] S2. Performing temporal perception adjustment on the text style code to obtain a text image sequence.
[0071] In an embodiment of the present invention, the text style encoding is input into an improved normalization layer, and a temporal perception offset from a dynamic style memory mechanism is introduced on the basis of traditional normalization, so that the text style encoding normalization process can perceive historical style changes and enhance long-range consistency. After normalization modulation, it is input into a synthesis network to generate a text image sequence with consistent style.
[0072] In specific medical and health scenarios, in the digitization project of ancient Chinese medical books, the system can automatically generate electronic medical records and health education graphics and texts with a unified style and ancient calligraphy effects based on patient information, so that the content conforms to modern medical standards while retaining the cultural charm of traditional Chinese medicine, thereby enhancing patients' trust and acceptance of Chinese medicine diagnosis and treatment.
[0073] In the specific scenario of financial technology, in the intelligent report generation of financial enterprises, the system can automatically output industry analysis reports and financial briefings that are consistent with the company's brand image style, ensuring the professionalism and visual unity of graphic expression, helping enterprises to efficiently complete external display and customer service, and enhance brand influence.
[0074] In an embodiment of the present invention, performing temporal perception adjustment on the text style code to obtain a text image sequence includes:
[0075] Normalizing the text style code to obtain a normalized style code;
[0076] generating a style feature map according to the normalized style code;
[0077] Determining a mean of the style feature map, and calculating a standard deviation of the style feature map based on the mean;
[0078] resizing the style feature map according to the mean and the standard deviation to obtain an adjusted feature map;
[0079] generating a style scaling parameter and a style offset parameter according to the normalized style code;
[0080] Normalizing the adjusted feature map according to the style scaling parameter and the style shift parameter to obtain a standard feature map;
[0081] Performing a linear projection on the updated style state to obtain a temporal perception offset;
[0082] The timing-aware offset is used to perform timing adjustment on the standard feature map to obtain a text image sequence.
[0083] Specifically, through methods such as standardization or minimum-maximum normalization, the numerical range of style codes is adjusted to a uniform scale (for example, between 0 and 1 or -1 and 1) to eliminate the numerical differences between different style codes, ensure that in the subsequent processing, the comparison and fusion between different codes are more stable and effective, and avoid calculation deviations caused by numerical differences.
[0084] A convolutional neural network (CNN) or other feature extraction network is used to generate a style feature map. The style code is spatially mapped through multi-layer convolution operations, converting the one-dimensional normalized style code into a two-dimensional feature map with spatial structure information. This can more intuitively express the distribution and changes of text style in the spatial domain, providing richer style features for subsequent stylized generation or migration tasks.
[0085] The mean and standard deviation of the style feature graph are calculated to measure its overall distribution and volatility. The calculation formula is as follows:
[0086]
[0087] Among them, μ(x i ) represents the mean, N represents the total number of pixels in the style feature map, x i Represents the value of the pixel in the i-th style feature map.
[0088]
[0089] Among them, μ(x i ) represents the mean, N represents the total number of pixels in the style feature map, x i represents the value of the pixel of the i-th style feature map, σ(x i ) represents the standard deviation.
[0090] Based on the calculated mean and standard deviation, the style feature map is resized (e.g., scaled or cropped) to conform to the specified size while maintaining the validity and consistency of the style information. This resizing process ensures that the style feature map is compatible with other modules, preventing size mismatches from affecting subsequent style generation or transfer.
[0091] Style scaling parameters and style shift parameters are generated through neural network or linear mapping calculations. The style scaling parameter adjusts the strength of style features, controlling overall amplification or reduction of the style; the style shift parameter shifts or offsets style features, adjusting local variations. The combination of these two parameters enables flexible and fine-grained adjustments to text style, achieving precise control and personalized expression of style.
[0092] By applying the style scaling parameter to the feature map, and using the style offset parameter to translate the feature map, it is made to meet specific standardization requirements in terms of numerical distribution. A standard feature map is obtained, which ensures the consistency of style features while being able to adapt to subsequent style fusion and generation tasks, ensuring that the output style image is highly consistent with the target style. The calculation formula is as follows:
[0093]
[0094] Among them, μ(x i ) represents the mean, σ(x i ) represents the standard deviation, γ i (w) represents the style scaling parameter, βi ( w) represents the style offset parameter, x i Represents the value of the pixel of the i-th style feature map, AdaIN(x i ,w) represents the standard feature map.
[0095] The updated style state is linearly projected, mapping it to a new space by applying a linear transformation layer (such as a fully connected layer or matrix multiplication) to extract an offset related to temporal variations. This projection process learns the temporal dependencies of the style state to generate a temporally aware offset that represents the dynamic changes and evolution of the style over time. This offset helps adjust the style features, making the generated style more consistent with temporal variations.
[0096] In an embodiment of the present invention, the step of using the timing-aware offset to perform timing adjustment on the standard feature map to obtain a text image sequence includes:
[0097] Obtaining the standard style features of each time step in the standard feature map;
[0098] Dynamically adjusting the standard style features using the temporal-aware offset to obtain a style feature map for each time step;
[0099] Deconvolution is performed on the style feature map to obtain a text image sequence.
[0100] Specifically, the time-aware offset weights or shifts the standard style features at each time step based on the temporal evolution of the style, capturing the details and trends of the style over time. This process yields a style feature map for each time step, ensuring that the style features evolve smoothly over time.
[0101] The deconvolutional neural network maps the style feature map at each time step back into high-dimensional space, restoring the corresponding text image. The deconvolution process gradually maps low-dimensional features back to the original image space, generating a high-resolution text image sequence through upsampling and convolution operations. Each frame accurately reflects the style features of the corresponding time step, ultimately resulting in a text image sequence that conforms to dynamic style changes. The calculation formula is as follows:
[0102]
[0103] Among them, x i represents the value of the pixel of the i-th style feature map, h t represents the updated style state at time step t, δ i (h t ) represents the updated style state h from GRU t The time-aware offset obtained by linear projection, Represents a sequence of text images.
[0104] By performing temporal-aware adjustments to the text style encoding, smooth transitions and dynamic changes in style over time can be achieved, ensuring that the generated text image sequence maintains both consistency and continuity in style. Normalizing the style encoding, generating a style feature map, and performing size and standardization adjustments can effectively optimize image detail and quality; the introduction of style scaling and offset parameters enhances the accuracy of style adjustment. Dynamic adjustment of the temporal-aware offset further enhances the temporal consistency of the style, making the resulting text image sequence more natural and smooth in style, and accurately reflecting the evolution of style over time. This approach provides greater flexibility and expressiveness for complex tasks such as font design and style transfer.
[0105] S3. Analyze the temporal consistency and single-word authenticity of the text image sequence using a preset discriminator to obtain a discrimination result.
[0106] In an embodiment of the present invention, the generated text image sequence is input into the discriminator, the temporal consistency of the text image sequence is evaluated based on structural features and style features through the sequence discrimination branch, and the authenticity of individual text images in the text image sequence is evaluated through the standard discrimination branch. Finally, the discrimination result reflecting the sequence coherence and the authenticity of individual words is output as the basis for subsequent optimization and generation quality control.
[0107] In specific medical and health scenarios, the text image sequence evaluation technology based on the sequence discriminator can be applied to the automatic review and repair of handwritten documents such as electronic medical records and medical prescriptions. By combining the structural and style features of the sequence discrimination branch, the temporal coherence of the medical record content and the consistency of the professional terminology are guaranteed. At the same time, the standard discrimination branch is used to ensure the clarity and authenticity of individual text images, effectively reducing the medical risks caused by illegible handwriting or forgery and tampering, and improving the digital credibility of medical documents.
[0108] In specific financial technology scenarios, it can be applied to aspects such as bill recognition, contract review and handwritten signature verification. The discriminator can evaluate the temporal consistency of the text image sequence in handwritten bills or contracts, detect traces of tampering or forged signatures, and combine the authenticity judgment of single words to ensure the reliability of each keyword information, thereby enhancing the automated risk control capabilities of financial documents and ensuring the security and compliance of financial business.
[0109] In an embodiment of the present invention, the use of a preset discriminator to analyze the temporal consistency and the authenticity of individual words of the text image sequence to obtain a discrimination result includes:
[0110] Obtain the sequence discrimination branch and standard discrimination branch of the preset discriminator;
[0111] Using the sequence discrimination branch to perform convolution feature extraction on the text image sequence to obtain a temporal consistency feature;
[0112] Performing global average pooling on the temporal consistency features to obtain global image features;
[0113] Performing full connection on the global image features to obtain a judgment result of temporal consistency of the text image sequence;
[0114] The standard discrimination branch is used to perform authenticity analysis on each character in the character image sequence to obtain a discrimination result of the authenticity of a single character in the character image sequence.
[0115] Specifically, multi-layer convolution operations are used to extract temporal features from image sequences, focusing specifically on the consistency of image changes over time. Through the convolutional layers, the model is able to capture subtle features of text image sequences that change over time, ultimately generating temporal consistency features that measure the continuity and stability of image sequences over time.
[0116] Global average pooling extracts global information of the image by calculating the mean of the entire feature map, while reducing the dimension of the feature map, allowing the model to focus more on global temporal consistency rather than local details. The obtained global image features can summarize the overall temporal characteristics of the text image sequence.
[0117] The global image features are fully connected and fed into a fully connected layer, which maps the features to a lower-dimensional output space through a weighted summation. The fully connected layer integrates the temporal consistency information in the global image features and performs appropriate nonlinear transformations to ultimately generate a score representing the temporal consistency of the text image sequence. This result can be used to assess the temporal coherence and stability of the image sequence, providing feedback for style generation or adjustment.
[0118] The standard discrimination branch is used to analyze the authenticity of each character in the text image sequence, extracting and distinguishing features from each character image through modules such as convolutional and fully connected layers. The standard discrimination branch focuses on the authenticity of each individual character's structure, strokes, and font, assessing whether it conforms to the morphological characteristics of real text. By analyzing each character image, an authenticity score is obtained for the character in the image sequence, providing accurate feedback for subsequent style optimization and image generation, ensuring that each character visually conforms to the standard text representation.
[0119] By using a preset discriminator to analyze the temporal consistency and individual word authenticity of the text image sequence, the generation effect can be comprehensively evaluated from both global and local levels, ensuring the consistency of the image sequence in terms of style continuity and temporal coherence, while ensuring the authenticity and standardization of each individual word in details such as structure and strokes. The sequence discrimination branch strengthens the temporal perception of overall style changes to avoid sudden changes or incoherence in style; the standard discrimination branch refines the quality control of individual word forms to improve the visual authenticity and usability of text images. Through the dual discrimination mechanism, the generated text image sequence reaches a high level in both style expression and content quality, providing a reliable guarantee for subsequent applications such as style transfer and font generation.
[0120] S4. Calculate the adversarial loss value, temporal loss value, and authenticity loss value of the text image sequence according to the discrimination result, and perform weighted summation on the adversarial loss value, the temporal loss value, and the authenticity loss value to obtain a total loss value.
[0121] In an embodiment of the present invention, an adversarial loss value, a temporal loss value, and an authenticity loss value are calculated based on the discrimination results, wherein: the adversarial loss value is based on the output of the standard discrimination branch, and adopts an R1 regularized non-saturated logical loss to improve the authenticity of the generated image; the temporal loss value is based on the deep features of adjacent characters in the generated sequence, and calculates the Euclidean distance difference of the stroke topological structure and the L1 norm difference of the ink and white texture, respectively, to constrain structural consistency and texture consistency; the authenticity loss value measures the similarity between the generated characters and the real characters in the high-level feature space to ensure structural authenticity; the above three types of losses are combined in a weighted manner to obtain a total loss value.
[0122] In specific medical and health scenarios, it can be applied to the high-fidelity generation and restoration of medical handwritten documents. By adversarial loss, the authenticity of the generated text images of medical records and prescriptions is enhanced. Temporal loss constrains the coherence of medical terms in stroke structure and texture details to avoid information ambiguity or error propagation. Authenticity loss ensures the consistency of high-level features between the generated characters and the real handwriting style, thereby achieving intelligent restoration of damaged or blurred handwritten documents and improving the digital quality and credibility of medical archives.
[0123] In specific financial technology scenarios, it can be applied to the detection and intelligent repair of handwritten bill forgeries, and combat loss to enhance the realism of generated bill characters. Timing loss ensures the structural and texture consistency of continuously written bill key fields (such as amount and payee). Authenticity loss is used to measure the similarity between generated bill characters and real samples at the semantic layer, thereby enhancing the reliability of bill information. Overall, through multi-loss weighted optimization, it effectively improves the automated verification and anti-counterfeiting capabilities of financial documents and reduces business risks.
[0124] In an embodiment of the present invention, the adversarial loss value, temporal loss value, and authenticity loss value of the text image sequence are calculated based on the discrimination results, wherein the adversarial loss value adopts the non-saturated logistic loss with R1 regularization, and the calculation formula is as follows:
[0125]
[0126] Among them, G represents the length of the text image sequence, y g represents the g-th image discrimination result in the text image sequence in the discriminator, and BCE represents the binary cross entropy loss function.
[0127] The temporal loss value is composed of structural consistency loss and texture consistency loss, where: structural consistency loss is based on VGGconv32 layer features Constrain the topological relationship of the strokes. The calculation formula is as follows:
[0128]
[0129] in, represents the structural consistency feature in the g-th image discrimination result in the text image sequence, represents the structural consistency feature in the discrimination result of the g+1th image in the text image sequence, represents the squared Euclidean distance.
[0130] Texture consistency loss uses conv42 layer features To keep the ink color and flying white effect stable, the calculation formula is as follows:
[0131]
[0132] in, represents the texture consistency feature in the g-th image discrimination result in the text image sequence, represents the texture consistency feature in the discrimination result of the g+1th image in the text image sequence, and |·|1 represents the L1 norm.
[0133] The formula for calculating the timing loss value is as follows:
[0134] L temp =γ·L struct +(1-γ)·L texture
[0135] Among them, γ∈[0,1] is the adjustable parameter controlling the strength of the structural constraint, L struct Represents the structural consistency loss value, L texture Indicates the texture consistency loss value.
[0136] The authenticity loss value measures the similarity between the generated characters and the real characters in the high-level feature space. The calculation formula is as follows:
[0137]
[0138] Among them, G represents the length of the text image sequence, I g represents the g-th image discrimination result in the text image sequence in the generator, represents the real image in the g-th image discrimination result, and ||·||1 represents the L1 norm.
[0139] The adversarial loss value, the timing loss value, and the authenticity loss value are weighted and summed to obtain the total loss value. The calculation formula is as follows:
[0140] L G =L GAN +λ1·L temp +λ2·L perc
[0141] Among them, λ1 and λ2 are balance hyperparameters, L GAN Represents the adversarial loss value, L temp Represents the timing loss value, L perc Represents the authenticity loss value.
[0142] By calculating the adversarial loss, temporal consistency loss, and authenticity loss of the text image sequence based on the discrimination results, and taking the weighted sum of the three to obtain the total loss value, it is possible to achieve multi-dimensional constraints and balanced optimization of the generation effect. The adversarial loss enhances the overall realism of the generated image, the temporal consistency loss ensures the coherence of the style and structure in the sequence, and the authenticity loss refines the morphological quality of individual characters. Through weighted fusion, each loss item can be flexibly adjusted according to actual needs, avoiding distortion or style drift caused by the optimization of a single indicator, thereby effectively improving the comprehensive performance of the generated text image sequence in terms of "authenticity, coherence, and aesthetics."
[0143] S5. Utilize the total loss value to perform style optimization on the text image sequence to obtain a target text shape.
[0144] In an embodiment of the present invention, based on the total loss value, the learnable parameters in the generator network, the discriminator network and the dynamic style memory mechanism are jointly optimized through back propagation, and cyclic iterative training is performed for the deviations of the text image sequence in style coherence, temporal consistency and structural authenticity. The stroke topology, ink texture and overall writing style of the characters in the generated sequence are continuously corrected, so that the generated text sequence gradually approaches the shape characteristics of the target font, and finally achieves the optimization effect of the text image sequence with a natural and smooth style, rigorous structure and visual authenticity.
[0145] In specific medical and health scenarios, it can be applied to the high-quality digital reconstruction of electronic medical records, medical image annotation, and historical handwritten archives. By jointly optimizing the generator, discriminator, and style memory mechanism, it iteratively corrects and restores the style inconsistencies and structural distortions in handwritten medical documents caused by writing differences, aging blurring, or damage. Ultimately, it generates high-fidelity text and image sequences with unified style, rigorous structure, and compliance with medical professional expression standards, thereby improving the readability and digital quality of medical documents.
[0146] In specific financial technology scenarios, it can be applied to tasks such as automatic bill generation, electronic contract archiving, and signature consistency verification. By cyclically optimizing the stroke structure, ink texture, and overall style of text image sequences, it ensures the writing continuity and information authenticity of important documents such as bills and contracts during digital transformation, effectively improving the usability and anti-counterfeiting capabilities of financial documents, and assisting in intelligent risk control and business automation.
[0147] Figure 3 A schematic diagram of a flow chart of obtaining a target character shape in a character shape generation method based on temporal consistency provided by an embodiment of the present invention.
[0148] In an embodiment of the present invention, the step of performing style optimization on the text image sequence using the total loss value to obtain a target text shape includes:
[0149] Performing back propagation on the text image sequence according to the total loss value, and calculating the parameter gradient value of the text image sequence during the back propagation process;
[0150] performing parameter update on the text image sequence using the parameter gradient value to obtain an updated image sequence;
[0151] A target text shape is generated according to the updated image sequence.
[0152] In detail, the chain rule is used to backpropagate the error information of the total loss to the parameters of each layer layer by layer, and the gradient of each parameter to the total loss is automatically calculated. For example, the gradient of the generator parameters affects the quality and style expression of the generated image, the gradient of the GRU unit parameters affects the temporal coherence of the style memory, and the gradient of the discriminator parameters affects the discriminator's ability to judge authenticity and consistency. The direction and magnitude of their impact on the generated results are quantified to provide a basis for subsequent optimization.
[0153] Based on the parameter gradients calculated through backpropagation, an optimization algorithm (such as Adam) is used to update the learnable parameters in the generator, dynamic style memory mechanism (GRU unit), and discriminator. Parameter values are adjusted according to the gradient direction, optimizing towards minimizing the total loss. The magnitude of each update is controlled by the learning rate. Generator parameter updates enhance the style quality and authenticity of the image, GRU unit parameters optimize the temporal style expression, and discriminator parameters enhance discrimination accuracy. These three factors evolve in tandem to continuously improve the generation of text image sequences.
[0154] Utilizing the updated generator parameters and dynamic style memory mechanism, the initial text sequence is forward propagated again to generate a new text image sequence. The updated parameters more accurately represent the target style characteristics and temporal coherence, further optimizing the generated results in terms of structural restoration, detail representation, and style consistency. Through multiple rounds of iteration, the gap between the generated image and the target style is gradually narrowed, ultimately achieving a target text shape that meets the desired shape and style requirements, achieving high-quality text shape generation.
[0155] By using the total loss value to perform style optimization on the text image sequence, the quality and style consistency of the generated images can be effectively improved. By calculating the parameter gradient value through backpropagation, it is ensured that the parameters of the generator, dynamic style memory mechanism and discriminator can be updated in the optimization direction, enhancing the structural authenticity and temporal coherence of the generated image. By updating the parameters, the generation results are further refined, making the text image sequence more accurate in details and more unified in style. The updated image sequence can better display the target text shape, which not only improves the generation quality, but also ensures the stability and consistency of the style, thereby achieving high-quality target text shape generation.
[0156] It should be understood that the order of execution of the steps in the above embodiments does not necessarily mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0157] like Figure 4 , which is a functional module diagram of a device for generating character shapes based on temporal consistency provided by one embodiment of the present invention.
[0158] In the embodiment of the present disclosure, a device for generating character shapes based on temporal consistency is provided. The device for generating character shapes based on temporal consistency corresponds to the method for generating character shapes based on temporal consistency in the above embodiment. Figure 4 As shown, the device 100 for generating text shapes based on temporal consistency can be installed in an electronic device. According to the functions to be implemented, the device 100 includes a style updating module 101, a temporal adjustment module 102, an image analysis module 103, a loss calculation module 104, and a style optimization module 105. The functional modules are described in detail as follows:
[0159] The style updating module 101 is used to obtain an initial text sequence, perform dynamic style updating on the initial text sequence, and obtain a text style code;
[0160] A timing adjustment module 102 is configured to perform timing-aware adjustment on the text style code to obtain a text image sequence;
[0161] The image analysis module 103 is used to analyze the temporal consistency and the authenticity of the single words of the text image sequence using a preset discriminator to obtain a discrimination result;
[0162] a loss calculation module 104 for calculating an adversarial loss value, a temporal loss value, and an authenticity loss value of the text image sequence according to the discrimination result, and performing a weighted summation of the adversarial loss value, the temporal loss value, and the authenticity loss value to obtain a total loss value;
[0163] The style optimization module 105 is configured to perform style optimization on the text image sequence using the total loss value to obtain a target text shape.
[0164] In one embodiment, when the style updating module 101 obtains an initial text sequence, performs dynamic style updating on the initial text sequence, and obtains a text style code, it is configured to:
[0165] Mapping the initial character sequence to a preset network, and encoding the mapped initial character sequence to obtain an initial character code;
[0166] Generate a corresponding basic style code according to the initial text code;
[0167] Acquire a preset gated recurrent unit and historical style features of historical texts, and construct a historical text style memory library based on the preset gated recurrent unit and the historical style features;
[0168] Acquiring the historical style state of the historical text style memory library and the initial style state of the basic style code;
[0169] Updating the initial style state using the historical style state to obtain an updated style state;
[0170] Obtaining an adjustment weight matrix and an adjustment bias vector, and calculating a style adjustment amount according to the updated style state, the adjustment weight matrix, and the adjustment bias vector;
[0171] The style adjustment amount and the basic style code are weightedly fused to obtain a text style code.
[0172] In one embodiment, when the style updating module 101 obtains an initial text sequence, performs dynamic style updating on the initial text sequence, and obtains a text style code, it is configured to:
[0173] Obtaining a memory influence strength parameter, and using the memory influence strength parameter to constrain the style adjustment amount to obtain a target adjustment amount;
[0174] Linearly superimposing the target adjustment amount and the basic style code to obtain a preliminary fusion style code;
[0175] The preliminary fusion style code is subjected to nonlinear transformation to obtain a text style code.
[0176] In one embodiment, when performing the timing-aware adjustment on the text style code to obtain the text image sequence, the timing adjustment module 102 is configured to:
[0177] Normalizing the text style code to obtain a normalized style code;
[0178] generating a style feature map according to the normalized style code;
[0179] Determining a mean of the style feature map, and calculating a standard deviation of the style feature map based on the mean;
[0180] resizing the style feature map according to the mean and the standard deviation to obtain an adjusted feature map;
[0181] generating a style scaling parameter and a style offset parameter according to the normalized style code;
[0182] Normalizing the adjusted feature map according to the style scaling parameter and the style shift parameter to obtain a standard feature map;
[0183] Performing a linear projection on the updated style state to obtain a temporal perception offset;
[0184] The timing-aware offset is used to perform timing adjustment on the standard feature map to obtain a text image sequence.
[0185] In one embodiment, when performing the timing-aware adjustment on the text style code to obtain the text image sequence, the timing adjustment module 102 is configured to:
[0186] Obtaining the standard style features of each time step in the standard feature map;
[0187] Dynamically adjusting the standard style features using the temporal-aware offset to obtain a style feature map for each time step;
[0188] Deconvolution is performed on the style feature map to obtain a text image sequence.
[0189] In one embodiment, when the image analysis module 103 uses a preset discriminator to analyze the temporal consistency and the authenticity of the single word of the text image sequence and obtains the discrimination result, it is used to:
[0190] Obtain the sequence discrimination branch and standard discrimination branch of the preset discriminator;
[0191] Using the sequence discrimination branch to perform convolution feature extraction on the text image sequence to obtain a temporal consistency feature;
[0192] Performing global average pooling on the temporal consistency features to obtain global image features;
[0193] Performing full connection on the global image features to obtain a judgment result of temporal consistency of the text image sequence;
[0194] The standard discrimination branch is used to perform authenticity analysis on each character in the character image sequence to obtain a discrimination result of the authenticity of a single character in the character image sequence.
[0195] In one embodiment, when performing style optimization on the text image sequence using the total loss value to obtain the target text shape, the style optimization module 105 is configured to:
[0196] Performing back propagation on the text image sequence according to the total loss value, and calculating the parameter gradient value of the text image sequence during the back propagation process;
[0197] performing parameter update on the text image sequence using the parameter gradient value to obtain an updated image sequence;
[0198] A target text shape is generated according to the updated image sequence.
[0199] In the present invention, a device for generating text shapes based on temporal consistency is provided. First, the present invention obtains an initial text sequence, performs dynamic style update on the initial text sequence, and obtains a text style code. While retaining the basic structural features of the text, the present invention combines the historical style information accumulated in the style memory library to achieve gradual optimization and personalized adjustment of the style features, ensuring that the generated text style code has both overall style consistency and fine-grained style differences. The text style code is temporally adjusted to obtain a text image sequence, which can achieve smooth transition and dynamic change of style over time, ensuring that the generated text image sequence is both consistent and continuous in style. The temporal consistency and single-character style of the text image sequence are determined by a preset discriminator. The authenticity of the characters is analyzed to obtain a discrimination result, which can comprehensively evaluate the generation effect from both global and local levels, ensure the consistency of the image sequence in style continuity and time coherence, and at the same time ensure the authenticity and standardization of each single character in details such as structure and strokes. Then, the adversarial loss value, temporal loss value and authenticity loss value of the character image sequence are calculated according to the discrimination result, and the adversarial loss value, the temporal loss value and the authenticity loss value are weighted and summed to obtain the total loss value. Each loss item is flexibly adjusted according to actual needs to avoid distortion or style drift caused by single indicator optimization. The total loss value is used to optimize the style of the character image sequence to obtain the target character shape, so that the temporal consistency in the generation process of the character shape can be maintained. Regarding the specific definition of a character shape generation device based on temporal consistency, please refer to the definition of a character shape generation method based on temporal consistency above, which will not be repeated here. The various modules in the above-mentioned character shape generation device based on temporal consistency can be fully or partially implemented by software, hardware and their combination. The above modules may be embedded in or independent of the processor in the computer device in the form of hardware, or may be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0200] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 5As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a text shape generation method based on temporal consistency.
[0201] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the client side of a text shape generation method based on temporal consistency.
[0202] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0203] Obtaining an initial text sequence, and dynamically updating the style of the initial text sequence to obtain a text style code;
[0204] Performing temporal perception adjustment on the text style code to obtain a text image sequence;
[0205] Analyzing the temporal consistency and the authenticity of individual words of the text image sequence using a preset discriminator to obtain a discrimination result;
[0206] Calculating an adversarial loss value, a temporal loss value, and an authenticity loss value of the text image sequence according to the discrimination result, and performing a weighted summation of the adversarial loss value, the temporal loss value, and the authenticity loss value to obtain a total loss value;
[0207] The total loss value is used to perform style optimization on the text image sequence to obtain a target text shape.
[0208] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and apparatuses can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the module division is merely a logical function division, and actual implementation may employ other division methods.
[0209] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or hardware plus software functional modules.
[0210] Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the claims are intended to be embraced therein. Any reference to a figure in a claim should not be construed as limiting the claim to which it relates.
[0211] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0212] In some implementations of this embodiment, a computer-readable storage medium is provided, on which a computer program is stored, characterized in that when the computer program is executed by a processor, the steps of the method described in the above embodiment are implemented.
[0213] The readable storage medium of the present invention stores a computer program, which, when executed by a processor of an electronic device, can implement:
[0214] Obtaining an initial text sequence, and dynamically updating the style of the initial text sequence to obtain a text style code;
[0215] Performing temporal perception adjustment on the text style code to obtain a text image sequence;
[0216] Analyzing the temporal consistency and the authenticity of individual words of the text image sequence using a preset discriminator to obtain a discrimination result;
[0217] Calculating an adversarial loss value, a temporal loss value, and an authenticity loss value of the text image sequence according to the discrimination result, and performing a weighted summation of the adversarial loss value, the temporal loss value, and the authenticity loss value to obtain a total loss value;
[0218] The total loss value is used to perform style optimization on the text image sequence to obtain a target text shape.
[0219] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0220] The computer-readable storage medium may also store at least one computer-executable program / instruction, such as a computer-readable instruction. Computer-readable storage media include, but are not limited to, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Computer-readable storage media may include, for example, read-only memory (ROM), a hard disk, a flash memory, etc. For example, a non-transitory computer-readable storage medium may be connected to a computing device such as a computer, and then, when the computing device executes the computer-readable instructions stored on the computer-readable storage medium, the various methods described above may be performed.
[0221] In addition, the computer device may also include (but is not limited to) a data bus, an input / output (I / O) bus, a display, and input / output devices (eg, keyboard, mouse, speaker, etc.).
[0222] The processor can communicate with external devices via an I / O bus via a wired or wireless network.
[0223] In one embodiment, the at least one computer executable instruction may also be compiled into or constitute a software product / computer program product, wherein one or more computer executable instructions are executed by a processor to perform the various functions and / or method steps in the embodiments described in the present technology.
[0224] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0225] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0226] In the embodiments provided in the present disclosure, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a portion of code, and the above-mentioned module, program segment or a portion of code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.
[0227] It should be noted that, in this disclosure, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element limited by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0228] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
[0229] It should be noted that if software tools or components other than those of our company appear in the embodiments of this application, they are only used for illustration and do not represent actual use.
Claims
1. A method for generating character shapes based on temporal consistency, characterized in that: The method comprises: Obtaining an initial text sequence, and dynamically updating the style of the initial text sequence to obtain a text style code; Performing temporal perception adjustment on the text style code to obtain a text image sequence; Analyzing the temporal consistency and the authenticity of individual words of the text image sequence using a preset discriminator to obtain a discrimination result; Calculating an adversarial loss value, a temporal loss value, and an authenticity loss value of the text image sequence according to the discrimination result, and performing a weighted summation of the adversarial loss value, the temporal loss value, and the authenticity loss value to obtain a total loss value; The total loss value is used to perform style optimization on the text image sequence to obtain a target text shape.
2. The method for generating character shapes based on temporal consistency according to claim 1, wherein: The step of obtaining an initial text sequence and dynamically updating the style of the initial text sequence to obtain a text style code includes: Mapping the initial character sequence to a preset network, and encoding the mapped initial character sequence to obtain an initial character code; Generate a corresponding basic style code according to the initial text code; Acquire a preset gated recurrent unit and historical style features of historical texts, and construct a historical text style memory library based on the preset gated recurrent unit and the historical style features; Acquiring the historical style state of the historical text style memory library and the initial style state of the basic style code; Updating the initial style state using the historical style state to obtain an updated style state; Obtaining an adjustment weight matrix and an adjustment bias vector, and calculating a style adjustment amount according to the updated style state, the adjustment weight matrix, and the adjustment bias vector; The style adjustment amount and the basic style code are weightedly fused to obtain a text style code.
3. The method for generating character shapes based on temporal consistency according to claim 2, wherein: The step of performing weighted fusion of the style adjustment amount and the basic style code to obtain the text style code includes: Obtaining a memory influence strength parameter, and using the memory influence strength parameter to constrain the style adjustment amount to obtain a target adjustment amount; Linearly superimposing the target adjustment amount and the basic style code to obtain a preliminary fusion style code; The preliminary fusion style code is subjected to nonlinear transformation to obtain a text style code.
4. The method for generating character shapes based on temporal consistency according to claim 2, wherein: The step of performing temporal perception adjustment on the text style code to obtain a text image sequence includes: Normalizing the text style code to obtain a normalized style code; generating a style feature map according to the normalized style code; Determining a mean of the style feature map, and calculating a standard deviation of the style feature map based on the mean; resizing the style feature map according to the mean and the standard deviation to obtain an adjusted feature map; generating a style scaling parameter and a style offset parameter according to the normalized style code; Normalizing the adjusted feature map according to the style scaling parameter and the style shift parameter to obtain a standard feature map; Performing a linear projection on the updated style state to obtain a temporal perception offset; The timing-aware offset is used to perform timing adjustment on the standard feature map to obtain a text image sequence.
5. The method for generating character shapes based on temporal consistency according to claim 4, wherein: The step of adjusting the timing of the standard feature map by using the timing-aware offset to obtain a text image sequence includes: Obtaining the standard style features of each time step in the standard feature map; Dynamically adjusting the standard style features using the temporal-aware offset to obtain a style feature map for each time step; Deconvolution is performed on the style feature map to obtain a text image sequence.
6. The method for generating character shapes based on temporal consistency according to claim 1, wherein: The method of analyzing the temporal consistency and the authenticity of individual words of the text image sequence using a preset discriminator to obtain a discrimination result includes: Obtain the sequence discrimination branch and standard discrimination branch of the preset discriminator; Using the sequence discrimination branch to perform convolution feature extraction on the text image sequence to obtain a temporal consistency feature; Performing global average pooling on the temporal consistency features to obtain global image features; Performing full connection on the global image features to obtain a judgment result of temporal consistency of the text image sequence; The standard discrimination branch is used to perform authenticity analysis on each character in the character image sequence to obtain a discrimination result of the authenticity of a single character in the character image sequence.
7. The method for generating character shapes based on temporal consistency according to claim 1, wherein: The performing style optimization on the text image sequence using the total loss value to obtain a target text shape includes: Performing back propagation on the text image sequence according to the total loss value, and calculating the parameter gradient value of the text image sequence during the back propagation process; performing parameter update on the text image sequence using the parameter gradient value to obtain an updated image sequence; A target text shape is generated according to the updated image sequence.
8. A device for generating character shapes based on temporal consistency, characterized in that: The device comprises: A style updating module is used to obtain an initial text sequence, dynamically update the style of the initial text sequence, and obtain a text style code; A timing adjustment module, configured to perform timing-aware adjustment on the text style code to obtain a text image sequence; An image analysis module is used to analyze the temporal consistency and the authenticity of individual words of the text image sequence using a preset discriminator to obtain a discrimination result; a loss calculation module, configured to calculate an adversarial loss value, a temporal loss value, and an authenticity loss value of the text image sequence according to the discrimination result, and perform a weighted sum of the adversarial loss value, the temporal loss value, and the authenticity loss value to obtain a total loss value; A style optimization module is used to perform style optimization on the text image sequence using the total loss value to obtain a target text shape.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for generating text shapes based on temporal consistency as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for generating character shapes based on temporal consistency as claimed in any one of claims 1 to 7 is implemented.