Drama text image generation method and device, equipment, storage medium and computer program product
By acquiring multi-source heterogeneous data from dramas, performing category identification and cross-language mapping, and establishing a hierarchically organized multimodal content resource library, the problem of unified organization of language data and performance data in drama text-image generation is solved, improving generation efficiency and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG MECHANICAL & ELECTRICAL COLLEGE
- Filing Date
- 2026-06-16
- Publication Date
- 2026-07-31
AI Technical Summary
In the current technology for generating dramatic text images, the correspondence between language data and performance data lacks a unified organization, making it difficult to fully utilize multimodal association features and resulting in poor generation efficiency and accuracy.
Acquire multi-source heterogeneous data related to drama, extract multi-dimensional feature information, perform category identification and cross-language mapping processing, establish a hierarchically organized multimodal content resource library, and generate drama text images based on user input content analysis of demand features.
It improves the efficiency and accuracy of generating theatrical text images, enables the orderly retrieval and retrieval of multimodal content, and enhances the correspondence between the generated results and user needs.
Smart Images

Figure CN122492885A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image generation technology, and in particular to a method, apparatus, device, storage medium, and computer program product for generating dramatic text images. Background Technology
[0002] Theatrical cultural content typically encompasses multiple information forms, including language expression, performance actions, character portrayal, costumes and props, set design, and audio-visual presentation, making it a typical multimodal information collection. In the processes of theatrical cultural dissemination, digital preservation, cross-linguistic organization, and content generation applications, related technologies usually collect and digitize text, images, audio, and video content. However, existing solutions mostly focus on the organization, storage, or display of single-type data, lacking a unified organization of the correspondence between linguistic and performance data, making it difficult to form a structured content system that can be uniformly retrieved and linked. Furthermore, in cross-linguistic scenarios, existing technologies usually only process text content separately, failing to integrate it with non-linguistic content such as performance and scene information for collaborative organization. This results in the multimodal correlation features in theatrical content being difficult to fully utilize, leading to poor efficiency and accuracy in the generation of theatrical text-images. Therefore, improving the efficiency and accuracy of theatrical text-image generation has become an urgent technical problem to be solved. Summary of the Invention
[0003] The main objective of this application is to provide a method, apparatus, device, storage medium, and computer program product for generating theatrical text images, aiming to solve the technical problem of how to improve the efficiency and accuracy of theatrical text image generation.
[0004] To achieve the above objectives, this application provides a method for generating dramatic text images, applied to a computer device, the method comprising the following steps: Acquire multi-source heterogeneous data related to drama, and extract corresponding multi-dimensional feature information based on the multi-source heterogeneous data, wherein the multi-source heterogeneous data includes at least language data and performance data; The multidimensional feature information is subjected to category identification processing and cross-language mapping processing to obtain corresponding feature identification results and language comparison results, and the multidimensional feature information is hierarchically organized based on the feature identification results; Based on the correspondence between the language data and the performance data, an association relationship is established between each level, and a multimodal content resource library for image generation is constructed based on the hierarchical organization results, the language comparison results, and the association relationship. Obtain user input content, parse the user input content, and obtain the demand features, feature relationships, and output format; Based on the required characteristics and the multimodal content resource library, target content units are determined. The target content units are associated and organized based on the feature relationships to obtain a target content set. Image semantic control parameters are generated based on the target content set. The image semantic control parameters are input into a preset image generation model for image feature generation processing. The generated image features are rendered and encapsulated based on the image output format to obtain the drama text image generation result corresponding to the user input content.
[0005] In one embodiment, the step of acquiring multi-source heterogeneous data related to drama and extracting corresponding multi-dimensional feature information based on the multi-source heterogeneous data includes: According to the preset acquisition rules, text data, image data, audio data, video data and spatial reconstruction data corresponding to the drama content are acquired, and the source identification and object correspondence processing of the acquired data are performed to obtain the original data set corresponding to the drama object. Data cleaning, format standardization, and content alignment are performed on various types of data in the original dataset to form basic representation data corresponding to language expression content, performance action content, scene presentation content, and carrier form content, respectively. Based on the aforementioned basic representation data, feature information of representation semantic content, expression form, scene attributes, and associated objects is extracted, and the extracted feature information is integrated to obtain multidimensional feature information corresponding to the multi-source heterogeneous data.
[0006] In one embodiment, the step of performing category identification processing and cross-language mapping processing on the multidimensional feature information to obtain corresponding feature identification results and language comparison results, and hierarchically organizing the multidimensional feature information based on the feature identification results, includes: Based on the content attributes, performance types and source categories corresponding to each feature item in the multidimensional feature information, the multidimensional feature information is classified, identified and labeled to determine the category and identification content corresponding to each feature item, and the feature identification result corresponding to the multidimensional feature information is obtained. Based on the feature identification results, the identification content related to language expression in each feature item is extracted, and the identification content related to language expression is converted and organized across languages according to the preset language correspondence rules to obtain the language comparison results corresponding to the feature identification results; Based on the feature identification results, the hierarchical attribution rules corresponding to each feature item in the multidimensional feature information are determined, and the multidimensional feature information is divided into the corresponding information levels according to the hierarchical attribution rules, so as to obtain the result of hierarchical organization of the multidimensional feature information based on the feature identification results.
[0007] In one embodiment, the step of establishing the association relationship between each level based on the correspondence between the language data and the performance data, and constructing a multimodal content resource library for image generation based on the hierarchical organization result, the language comparison result, and the association relationship, includes: Based on the correspondence between the language data and the performance data in terms of content theme, expression object, performance segment and time sequence position, the feature items of different information levels are associated and identified to determine the cross-level association clues between each information level, and the hierarchical association results between the language data and the performance data are obtained. Based on the hierarchical association results, the feature information of each level in the hierarchical organization results is associated and mapped and arranged, and the language mapping relationship corresponding to each level feature information is established in combination with the language comparison results to obtain the hierarchical association organization results corresponding to the hierarchical organization results. Based on the hierarchical association organization results, the feature information of each level and the corresponding language mapping content in the hierarchical organization results are collected, stored and structured to obtain a multimodal content resource library for image generation, including hierarchical structure, language correspondence relationship and cross-level association relationship.
[0008] In one embodiment, the step of acquiring user input content and parsing the user input content to obtain demand features, feature relationships, and output format includes: The system acquires user input search requests, generates instructions or interactive description information, and performs normalization processing on the user input content to determine the corresponding target expression unit and content pointing range, thereby obtaining the input parsing data corresponding to the user input content. Based on the input parsing data, the topic objects, modal orientations, content constraints and combination requirements involved in the user input content are identified, the feature elements corresponding to the user input content and the coordination relationship between each feature element are determined, and the demand features and feature relationships corresponding to the user input content are obtained. Based on the input parsing data, the demand features, and the feature relationships, the result presentation requirements corresponding to the user input content are identified, and the corresponding content organization form and output expression type are determined to obtain the output format after parsing the user input content.
[0009] In one embodiment, the steps of determining target content units based on the demand features and the multimodal content resource library, associating and organizing the target content units based on the feature relationships to obtain a target content set, generating image semantic control parameters based on the target content set, inputting the image semantic control parameters into a preset image generation model for image feature generation processing, and rendering and encapsulating the generated image features based on the image output format to obtain a drama text image generation result corresponding to the user input content include: Based on the demand characteristics, the content items corresponding to the demand characteristics are retrieved and matched in the multimodal content resource library, and target content units that are adapted to the demand characteristics are filtered based on the matching results. Based on the characteristic relationship, the content pointing relationship, combination order relationship and cooperation constraint relationship between the target content units are associated and arranged to obtain the target content set corresponding to the demand characteristics. Based on the semantic content, performance content, scene content, and object relationships corresponding to each target content unit in the target content set, semantic control information for characterizing image generation requirements is extracted, and the semantic control information is parameterized and transformed to generate image semantic control parameters corresponding to the target content set. The image semantic control parameters are then input into a preset image generation model to control the preset image generation model to perform image feature generation processing according to the content semantics corresponding to the target content set, thereby obtaining the corresponding image feature generation result. Based on the image output format, the image feature generation result is processed by image composition adjustment, presentation adaptation, and result rendering encapsulation to obtain the drama text image generation result corresponding to the user input content.
[0010] Furthermore, to achieve the above objectives, this application also proposes a theatrical text image generation apparatus, which includes: The data acquisition module is used to acquire multi-source heterogeneous data related to drama, and extract corresponding multi-dimensional feature information based on the multi-source heterogeneous data, wherein the multi-source heterogeneous data includes at least language data and performance data; The hierarchical organization module is used to perform category identification processing and cross-language mapping processing on the multidimensional feature information to obtain corresponding feature identification results and language comparison results, and to hierarchically organize the multidimensional feature information based on the feature identification results; The resource library construction module is used to establish the association between each level based on the correspondence between the language data and the performance data, and to construct a multimodal content resource library for image generation based on the hierarchical organization results, the language comparison results and the association. The input parsing module is used to acquire user input content and parse the user input content to obtain demand features, feature relationships and output format; The result generation module is used to determine target content units based on the required features and the multimodal content resource library, associate and organize the target content units based on the feature relationships to obtain a target content set, generate image semantic control parameters based on the target content set, input the image semantic control parameters into a preset image generation model for image feature generation processing, and render and encapsulate the generated image features based on the image output format to obtain the drama text image generation result corresponding to the user input content.
[0011] Furthermore, to achieve the above objectives, this application also proposes a theatrical text image generation device, the device comprising: a memory, a processor, and a theatrical text image generation program stored in the memory and executable on the processor, the theatrical text image generation program being configured to implement the steps of the theatrical text image generation method as described above.
[0012] In addition, to achieve the above objectives, this application also proposes a storage medium storing a theatrical text image generation program, which, when executed by a processor, implements the steps of the theatrical text image generation method described above.
[0013] In addition, to achieve the above objectives, this application also proposes a computer program product comprising a computer program that, when executed by a processor, implements the steps of the dramatic text image generation method described above.
[0014] This application acquires multi-source heterogeneous data related to drama and extracts corresponding multi-dimensional feature information based on the multi-source heterogeneous data, wherein the multi-source heterogeneous data includes at least language data and performance data; performs category identification processing and cross-language mapping processing on the multi-dimensional feature information to obtain corresponding feature identification results and language comparison results, and organizes the multi-dimensional feature information hierarchically based on the feature identification results; establishes the association relationship between each level based on the correspondence between language data and performance data, and constructs a multimodal content resource library for image generation based on the hierarchical organization results, language comparison results, and association relationships; acquires user input content and parses the user input content to obtain demand features, feature relationships, and output format; determines target content units based on demand features and the multimodal content resource library, organizes the target content units based on feature relationships to obtain a target content set, generates image semantic control parameters based on the target content set, inputs the image semantic control parameters into a preset image generation model for image feature generation processing, and renders and encapsulates the generated image features based on the image output format to obtain the drama text image generation result corresponding to the user input content. This application acquires multi-source heterogeneous data related to drama and extracts multi-dimensional feature information. It then combines category identification, cross-language mapping, and hierarchical organization to form a structured multimodal content resource library. Furthermore, it establishes associations based on the correspondence between language-related data and performance-related data, enabling the orderly retrieval and retrieval of drama content. On this basis, it parses user input to obtain demand features, feature relationships, and output formats. Based on this, it identifies target content units in the multimodal content resource library, organizes target content sets, and completes format conversion. This reduces disordered filtering and irrelevant retrieval during the generation process, improving the efficiency of drama text image generation. Simultaneously, it enhances the correspondence between the generated results and user needs, improving the accuracy of drama text image generation. Attached Figure Description
[0015] Figure 1 This is a flowchart illustrating the first embodiment of the dramatic text image generation method of this application; Figure 2 This is a schematic diagram of a sub-process in the second embodiment of the dramatic text image generation method of this application; Figure 3 This is a schematic diagram of a sub-process in the third embodiment of the dramatic text image generation method of this application; Figure 4 This is a schematic diagram of the module structure of the theatrical text image generation device according to an embodiment of this application; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the dramatic text image generation method in this application embodiment.
[0016] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0017] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of this application.
[0018] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0019] It should be noted that theatrical cultural content typically encompasses multiple information forms simultaneously, including language expression, performance actions, character portrayal, costumes and props, set design, and audio-visual presentation, making it a typical multimodal information collection. In the processes of theatrical cultural dissemination, digital preservation, cross-linguistic organization, and content generation applications, relevant technologies usually collect and digitize text, images, audio, and video content. However, existing solutions mostly focus on the organization, storage, or display of single-type data, lacking a unified organization of the correspondence between linguistic and performance data, making it difficult to form a structured content system that can be uniformly retrieved and linked. Furthermore, in cross-linguistic scenarios, existing technologies typically process text content separately, failing to integrate it with non-linguistic content such as performance and scene information for collaborative organization. This results in the multimodal correlation features in theatrical content being difficult to fully utilize, leading to poor efficiency and accuracy in the generation of theatrical text-images. Therefore, improving the efficiency and accuracy of theatrical text-image generation has become an urgent technical problem to be solved.
[0020] The main solution of this application is as follows: Acquire multi-source heterogeneous data related to drama, and extract corresponding multi-dimensional feature information based on the multi-source heterogeneous data, wherein the multi-source heterogeneous data includes at least language data and performance data; perform category labeling and cross-language mapping processing on the multi-dimensional feature information to obtain corresponding feature labeling results and language comparison results, and organize the multi-dimensional feature information hierarchically based on the feature labeling results; establish the association relationship between each level based on the correspondence between language data and performance data, and construct a multimodal content resource library for image generation based on the hierarchical organization results, language comparison results, and association relationships; acquire user input content, parse the user input content to obtain demand features, feature relationships, and output format; determine target content units based on demand features and the multimodal content resource library, organize the target content units based on feature relationships to obtain a target content set, generate image semantic control parameters based on the target content set, input the image semantic control parameters into a preset image generation model for image feature generation processing, and render and encapsulate the generated image features based on the image output format to obtain the drama text image generation result corresponding to the user input content.
[0021] This application acquires multi-source heterogeneous data related to drama and extracts multi-dimensional feature information. It then combines category identification, cross-language mapping, and hierarchical organization to form a structured multimodal content resource library. Furthermore, it establishes associations based on the correspondence between language-related data and performance-related data, enabling the orderly retrieval and retrieval of drama content. On this basis, it parses user input to obtain demand features, feature relationships, and output formats. Based on this, it identifies target content units in the multimodal content resource library, organizes target content sets, and completes format conversion. This reduces disordered filtering and irrelevant retrieval during the generation process, improving the efficiency of drama text image generation. Simultaneously, it enhances the correspondence between the generated results and user needs, improving the accuracy of drama text image generation.
[0022] It should be noted that the executing entity of the method in this embodiment can be a computing service device with data processing, network communication, and program execution functions, or it can be the aforementioned drama text image generation device with the same or similar functions. This embodiment and the following embodiments will be described using a drama text image generation device as an example.
[0023] Based on this, a first embodiment of the dramatic text image generation method of this application is proposed. Please refer to... Figure 1 , Figure 1 This is a schematic flowchart of the first embodiment of the dramatic text image generation method of this application.
[0024] In this embodiment, the method includes the following steps: S1: Acquire multi-source heterogeneous data related to drama, and extract corresponding multi-dimensional feature information based on the multi-source heterogeneous data, wherein the multi-source heterogeneous data includes at least language data and performance data; S2: Perform category identification processing and cross-language mapping processing on the multidimensional feature information to obtain corresponding feature identification results and language comparison results, and organize the multidimensional feature information hierarchically based on the feature identification results; It should be noted that multi-source heterogeneous data refers to a data set with different sources and different forms of expression, but which are used to represent the content of a drama. This includes at least language data used to represent lines, lyrics, dialogue, and explanatory text, as well as performance data used to represent actions, gestures, expressions, costumes, scenes, and character presentation. Multi-dimensional feature information refers to various types of information extracted from the multi-source heterogeneous data that reflect the attributes and performance characteristics of the drama content. Category identification processing refers to the process of classifying, assigning, and labeling each feature information based on its corresponding content attributes, performance forms, source types, or functional categories. Cross-language mapping processing refers to the process of establishing correspondences between different languages for the language-related content in the multi-dimensional feature information. Feature identification results refer to the results formed after category identification processing, used to represent the category affiliation and attribute labels of each feature information. Language comparison results refer to the correspondences between different language expressions formed after cross-language mapping processing. Hierarchical organization refers to the processing method of dividing multi-dimensional feature information into different hierarchical structures according to preset hierarchical rules based on the feature identification results.
[0025] Specifically, firstly, multi-source heterogeneous data related to drama is acquired, and corresponding multi-dimensional feature information is extracted based on this data. Specifically, language data and performance data related to the drama content can be collected. Language data is used to represent lines, lyrics, dialogue, plot descriptions, and character expressions in the drama, while performance data is used to represent action performances, stage presentation, character images, costumes, props, and scene information. After acquiring the aforementioned multi-source heterogeneous data, content analysis and feature extraction are performed on different types of data. Feature information related to semantic expression, subject matter, character orientation, and text content is extracted from the language data, while feature information related to action performance, visual presentation, scene attributes, character states, and performance cues is extracted from the performance data, thereby forming multi-dimensional feature information corresponding to the multi-source heterogeneous data.
[0026] Furthermore, after obtaining the multidimensional feature information, category identification and cross-language mapping processing are performed on the multidimensional feature information. Specifically, the multidimensional feature information can be classified, identified, and labeled according to the content attributes, expression forms, source categories, or expression objects corresponding to each feature information, determining the category affiliation and attribute labels of each feature information, and obtaining the corresponding feature identification results. Then, cross-language correspondence processing is performed on the feature content related to language expression to establish content correspondence relationships between different languages, obtaining the corresponding language comparison results. Furthermore, based on the feature identification results, the multidimensional feature information is classified, organized, and hierarchically allocated according to preset hierarchical division rules, so that feature information of different categories and different attributes enters the corresponding hierarchical structure, thereby completing the hierarchical organization of the multidimensional feature information.
[0027] By first acquiring multi-source heterogeneous data related to drama and extracting corresponding multi-dimensional feature information, the originally scattered and diverse drama content can be transformed into uniformly processed feature information. Further, through category identification and cross-language mapping, the multi-dimensional feature information can have clear category affiliations and language correspondences. Then, based on the feature identification results, the multi-dimensional feature information is hierarchically organized, allowing different types of drama feature content to form an orderly hierarchical structure. This provides a standardized foundation for establishing the correlation between language-based and performance-based data, and also provides structured support for the construction of a multimodal content resource library and matching user needs. Therefore, it helps improve the efficiency of content retrieval during the drama text-image generation process and enhances the accuracy of the correspondence between the generated results and the drama content.
[0028] S3: Based on the correspondence between the language data and the performance data, establish the association between each level, and construct a multimodal content resource library for image generation based on the hierarchical organization results, the language comparison results, and the association. S4: Obtain user input content, parse the user input content, and obtain demand features, feature relationships, and output format; It should be noted that the correspondence refers to the mutual correspondence between language data and performance data in terms of theme content, character objects, performance segments, temporal position, or expressive direction; the association between different levels refers to the content association, directional association, or cooperative association established between feature information at different levels after hierarchical organization; the hierarchical organization result refers to the hierarchical organization result formed after multidimensional feature information is hierarchically divided; the multimodal content resource library refers to the structured content collection used to store and call up various types of drama-related content, constructed based on the hierarchical organization result, language comparison result, and association relationship; user input content refers to the search requests, generation instructions, or interactive description information entered by the user during the generation process; demand features refer to the feature information obtained from the user input content that represents the user's target content demand; feature relationship refers to the combination relationship, cooperative relationship, or limiting relationship between the demand features; and output format refers to the result presentation form and content organization method corresponding to the user input content.
[0029] Specifically, based on the correspondence between the language data and the performance data, a multimodal content resource library is constructed, establishing relationships between different levels. This library is built upon the hierarchical organization results, the language comparison results, and the relationships. Specifically, based on the correspondence between language data and performance data in terms of theme, character, performance segment, temporal position, or expressive direction, feature information belonging to different levels is identified and associated to determine content correspondence and association clues between different levels, thus establishing relationships between each level. On this basis, the feature information of each level in the hierarchical organization results is matched with the language comparison results, and the feature information between different levels is organized and arranged in conjunction with the relationships, ensuring a unified association between language expression information, performance information, and corresponding content between different languages within the same dramatic content.
[0030] Furthermore, after completing the above processing, based on the hierarchical organization results, the language comparison results, and the association relationships, the relevant content in each level can be collected, stored, and structured to form a multimodal content resource library that can be called later. Then, user input content is acquired and parsed. Specifically, user input retrieval requests, generation instructions, or interactive description information can be received, and the user input content can be standardized and identified to determine the relevant topic objects, content orientation, modal requirements, content constraints, and result presentation requirements. Based on this, requirement features related to the user's target content are extracted, the combination, constraint, or coordination relationships between these requirement features are identified, and the corresponding output expression type and content organization form are further determined, thereby obtaining the requirement features, feature relationships, and output format corresponding to the user input content.
[0031] By establishing relationships between different levels based on the correspondence between language and performance data, and constructing a multimodal content resource library by combining hierarchical organization results and language comparison results, the previously scattered theatrical language content, performance content, and different linguistic expressions can be organized into a unified, structured system, thus providing a clear foundation for subsequent content retrieval. Furthermore, by acquiring user input and parsing it to obtain demand features, feature relationships, and output formats, the subsequent generation process no longer relies solely on simple keyword matching but can instead perform targeted retrieval based on the content orientation, combination requirements, and presentation methods in the user's needs. Therefore, this step enhances the organization and retrieval of theatrical content resources on the one hand, and improves the clarity of user demand identification on the other, providing prerequisite support for the accurate matching, association, and result generation of subsequent target content units, thereby improving the efficiency and accuracy of theatrical text image generation.
[0032] S5: Determine target content units based on the demand characteristics and the multimodal content resource library, associate and organize the target content units based on the feature relationships to obtain a target content set, generate image semantic control parameters based on the target content set, input the image semantic control parameters into a preset image generation model for image feature generation processing, and render and encapsulate the generated image features based on the image output format to obtain the drama text image generation result corresponding to the user input content.
[0033] It should be noted that the matching result refers to the matching analysis result obtained after inputting the demand features into the multimodal content resource library, based on the correspondence between the demand features and the content items in the resource library; the target content unit refers to the content item corresponding to the user's demand that is selected and determined based on the matching result; the target content set refers to the content combination result formed after associating and organizing multiple target content units based on the feature relationship; and the target generation result refers to the final output content obtained after converting the target content set according to the output format.
[0034] Specifically, the demand characteristics are input into the multimodal content resource library for matching. Specifically, based on the topic object, content orientation, modal requirements, and limiting conditions included in the demand characteristics, corresponding content items can be retrieved from the multimodal content resource library, and the correspondence between each content item and the demand characteristics can be matched and analyzed to determine candidate content suitable for the demand characteristics. Further, the candidate content can be screened and confirmed by combining the coordination of different feature elements in the demand characteristics, thereby determining the target content unit corresponding to the current user's needs. This process ensures that the target content unit is not selected randomly, but is based on demand-oriented matching, thus providing a clear input basis for subsequent content organization.
[0035] After determining the target content units, they are organized and associated based on the characteristic relationships. Specifically, according to the combination, limitation, and content coordination relationships reflected in the characteristic relationships, the order, correspondence, referencing, and combination methods of multiple target content units are organized and arranged to form a target content set corresponding to user needs. Then, based on the output format, the content in the target content set is format-adapted and the results are converted, so that the target content set is presented according to the corresponding content organization form and output expression type, ultimately obtaining the target generation result.
[0036] By inputting the aforementioned demand features into the multimodal content resource library for matching, and determining target content units based on the matching results, the subsequent generation process can be built upon content corresponding to user needs, reducing the involvement of irrelevant content in the generation process. Furthermore, by associating and organizing the target content units based on the aforementioned feature relationships, multiple target content units can form an orderly and interconnected set of target content according to the combination requirements and content relationships in the user's needs. Finally, by converting the target content set based on the specified output format, the final output content can be adapted to the presentation format required by the user. Therefore, this step transforms the dramatic text image generation process from simple content retrieval to a demand-oriented matching, organization, and conversion process, thereby improving the targeting and processing efficiency of content retrieval and enhancing the accuracy of the correspondence between the target generation results and user needs.
[0037] This embodiment acquires multi-source heterogeneous data related to drama and extracts corresponding multi-dimensional feature information based on the multi-source heterogeneous data. The multi-source heterogeneous data includes at least language data and performance data. Category identification and cross-language mapping processing are performed on the multi-dimensional feature information to obtain corresponding feature identification results and language comparison results. The multi-dimensional feature information is then hierarchically organized based on the feature identification results. Based on the correspondence between language data and performance data, association relationships are established between each level. A multi-modal content resource library for image generation is constructed based on the hierarchical organization results, language comparison results, and association relationships. User input content is acquired and parsed to obtain demand features, feature relationships, and output formats. Target content units are determined based on the demand features and the multi-modal content resource library. These target content units are then associated and organized based on feature relationships to obtain a target content set. Image semantic control parameters are generated based on the target content set. These parameters are input into a preset image generation model for image feature generation processing. The generated image features are then rendered and encapsulated based on the image output format to obtain the drama text image generation result corresponding to the user input content. This embodiment acquires multi-source heterogeneous data related to drama and extracts multi-dimensional feature information. It combines category identification, cross-language mapping, and hierarchical organization to form a structured multimodal content resource library. Then, it establishes associations based on the correspondence between language data and performance data, enabling drama content to be retrieved and called in an orderly manner. On this basis, it obtains demand features, feature relationships, and output formats by parsing user input content. Based on this, it determines target content units in the multimodal content resource library, organizes target content sets, and completes format conversion. This reduces disordered filtering and irrelevant calls in the generation process, improves the efficiency of drama text image generation, enhances the correspondence between the generated results and user needs, and improves the accuracy of drama text image generation.
[0038] Based on the first embodiment described above, a second embodiment of the dramatic text image generation method of this application is proposed. Please refer to... Figure 2 , Figure 2 This is a schematic diagram of a sub-process in the second embodiment of the dramatic text image generation method of this application.
[0039] like Figure 2 As shown, in this embodiment, step S1 includes: S11: According to the preset acquisition rules, acquire text data, image data, audio data, video data and spatial reconstruction data corresponding to the drama content, and perform source identification and object correspondence processing on the acquired data to obtain the original data set corresponding to the drama object; S12: Perform data cleaning, format normalization and content alignment on various types of data in the original data set to form basic representation data corresponding to language expression content, performance action content, scene presentation content and carrier form content respectively; S13: Based on the basic representation data, extract the feature information of the representation semantic content, expression form, scene attributes and associated objects, and integrate the extracted feature information to obtain the multi-dimensional feature information corresponding to the multi-source heterogeneous data.
[0040] It should be noted that the preset data acquisition rules refer to the pre-defined scope, objects, methods, and processing requirements for acquiring data related to theatrical content; spatial reconstruction data refers to data used to characterize the spatial form of theatrical scenes, costumes, props, or related objects; source identifiers refer to information tags used to indicate the source channel, acquisition method, or carrier of various types of data; object correspondence processing refers to the process of associating the acquired data with corresponding theatrical characters, scenes, performance segments, or content objects; the original data set refers to the initial data set corresponding to theatrical objects formed after source identifier and object correspondence processing; and basic representation data refers to the data obtained from theatrical content. The data in the original dataset, after being cleaned, organized, and aligned, represents the basic attributes of different theatrical content. Language expression content refers to content related to lines, lyrics, dialogues, and explanatory statements in the drama. Performance action content refers to content related to actions, postures, expressions, and performance styles in the drama. Scene presentation content refers to content related to the stage environment, set design, visual presentation, or scene composition. Carrier form content refers to content related to the appearance and structural state of costumes, props, cultural relics, or other carrier objects. Associated objects refer to the character objects, scene objects, performance objects, or content objects corresponding to the characteristic information.
[0041] Specifically, firstly, text data, image data, audio data, video data, and spatial reconstruction data corresponding to the theatrical content are acquired according to preset acquisition rules. Specifically, based on the expressive needs of the theatrical content, language materials, performance materials, visual materials, and spatial materials in the drama are collected in a categorized manner. Text data can be used to reflect lines, lyrics, dialogue, and explanatory information; image, video, and audio data can be used to reflect character images, performance actions, stage scenes, and sound performances; and spatial reconstruction data can be used to reflect the spatial structure and morphological information of scenes, costumes, props, or related objects. After acquiring the above types of data, source identification and object mapping are performed on each type of data, establishing a correspondence between data from different sources and corresponding theatrical characters, scenes, performance segments, or content objects, thereby forming an original data set corresponding to the theatrical objects.
[0042] Furthermore, after obtaining the original dataset, data cleaning, format standardization, and content alignment are performed on various types of data in the original dataset to form basic representation data. Specifically, redundant, irrelevant, or non-standard content in the original data can be cleaned up, and the storage, expression, and organization of different types of data can be standardized. Furthermore, the corresponding content between different data can be aligned in conjunction with the dramatic object, content theme, performance segment, or scene orientation, so that the original data respectively form basic representation data corresponding to language expression content, performance action content, scene presentation content, and carrier form content. Then, based on the basic representation data, feature information of semantic content, expression form, scene attributes, and associated objects is extracted, and the extracted feature information is integrated to form a unified feature result of different types of feature content, ultimately obtaining the multi-dimensional feature information corresponding to the multi-source heterogeneous data.
[0043] By acquiring multiple types of data corresponding to theatrical content according to preset collection rules, and processing these data types by source identification and object correspondence, a unified initial correspondence foundation can be established for data from different sources and in different forms around the same theatrical object. Further, by cleaning, formatting, and aligning the original dataset, the originally scattered, disorganized, and differently expressed data can be transformed into consistent and comparable basic representation data. Based on this, the feature information of semantic content, performance style, scene attributes, and associated objects is extracted and integrated, transforming multiple types of theatrical content into unified multidimensional feature information. Therefore, this step provides a standardized and structured data foundation for subsequent category identification processing, cross-language mapping processing, hierarchical organization, and the construction of a multimodal content resource library. This is beneficial for improving the content processing efficiency in the subsequent theatrical text image generation process and enhancing the accuracy of the generated results in representing the theatrical content.
[0044] Based on the first embodiment described above, in this embodiment, step S2 includes: S21: Based on the content attributes, performance types and source categories corresponding to each feature item in the multidimensional feature information, classify, identify and label the multidimensional feature information, determine the category belonging and identification content corresponding to each feature item, and obtain the feature identification result corresponding to the multidimensional feature information; S22: Based on the feature identification results, extract the identification content related to language expression from each feature item, and perform cross-language conversion and corresponding organization on the identification content related to language expression according to the preset language correspondence rules to obtain the language comparison results corresponding to the feature identification results; S23: Based on the feature identification results, determine the hierarchical attribution rules corresponding to each feature item in the multidimensional feature information, and divide the multidimensional feature information into the corresponding information levels according to the hierarchical attribution rules, so as to obtain the result of hierarchical organization of the multidimensional feature information based on the feature identification results.
[0045] It should be noted that: Feature item refers to each specific information unit constituting the multidimensional feature information; content attribute refers to the content nature or content orientation reflected by each feature item; presentation type refers to the presentation category corresponding to each feature item; source category refers to the data source type corresponding to each feature item; classification identification and labeling processing refers to the process of determining the category of each feature item and assigning corresponding labels based on its content attribute, presentation type, and source category; label content refers to the label content used to characterize the category affiliation, attribute characteristics, or source type of each feature item; preset language correspondence rules refer to the pre-set conversion and correspondence rules between different language expressions; cross-language conversion refers to the process of converting label content related to language expression into another language expression form; correspondence organization refers to the process of establishing correspondences and organizing related content under different language expressions; hierarchical attribution rules refer to the rules used to determine which information level each feature item should be classified to; information level refers to the information structure of different levels formed according to a predetermined classification logic; hierarchical organization result refers to the hierarchical organization result formed after classifying the multidimensional feature information to the corresponding information level according to the hierarchical attribution rules.
[0046] Specifically, firstly, based on the content attributes, performance types, and source categories corresponding to each feature item in the multidimensional feature information, the multidimensional feature information is classified, identified, and labeled. Specifically, the content nature, performance form, and source type corresponding to each feature item can be differentiated to determine whether each feature item belongs to the language expression category, performance presentation category, scene description category, carrier form category, or other related categories. Based on this, each feature item is assigned a corresponding category identifier and attribute label, thereby forming a feature identification result reflecting the category affiliation and identified content of each feature item. Then, based on the feature identification result, the identified content related to language expression is extracted, and the identified content related to language expression is cross-language converted and organized according to preset language correspondence rules, so that the same language expression content forms a corresponding organizational relationship in different language forms, thereby obtaining a language comparison result corresponding to the feature identification result.
[0047] Furthermore, after completing the above processing, the hierarchical classification rules corresponding to each feature item in the multidimensional feature information are determined based on the feature identification results. Specifically, the target information level to which each feature item should enter can be determined according to its category classification, attribute label, and source type. The multidimensional feature information is then divided into corresponding information levels according to the hierarchical classification rules, allowing feature items of different categories and attributes to enter their corresponding hierarchical structures. Through the above processing, the multidimensional feature information originally in the same feature set is further organized into a hierarchical information structure with category identification, language correspondence, and hierarchical classification relationships.
[0048] By classifying, identifying, and labeling multidimensional feature information based on the content attributes, expression types, and source categories corresponding to each feature item, the originally mixed multi-type feature content can be given clear category affiliation and identified content. Furthermore, by extracting the identified content related to language expression and performing cross-language conversion and corresponding organization according to preset language correspondence rules, a stable correspondence relationship can be established between related content in different language forms. Then, based on the feature identification results, the hierarchical affiliation rules corresponding to each feature item are determined, and the multidimensional feature information is divided into corresponding information levels, forming an orderly hierarchical organizational structure. Therefore, this step can further organize the multidimensional feature information extracted from multi-source heterogeneous data into structured information results with category distinctions, language correspondences, and hierarchical relationships. This provides a clearer organizational foundation for subsequently establishing the relationships between different levels and constructing a multimodal content resource library, thereby improving the efficiency of content retrieval and retrieval in the subsequent drama text image generation process and enhancing the accuracy of the correspondence between the generated results and the drama content.
[0049] This embodiment acquires multi-source heterogeneous data related to drama and extracts corresponding multi-dimensional feature information based on the multi-source heterogeneous data. The multi-source heterogeneous data includes at least language data and performance data. Category identification and cross-language mapping processing are performed on the multi-dimensional feature information to obtain corresponding feature identification results and language comparison results. The multi-dimensional feature information is then hierarchically organized based on the feature identification results. Based on the correspondence between language data and performance data, association relationships are established between each level. A multi-modal content resource library for image generation is constructed based on the hierarchical organization results, language comparison results, and association relationships. User input content is acquired and parsed to obtain demand features, feature relationships, and output formats. Target content units are determined based on the demand features and the multi-modal content resource library. These target content units are then associated and organized based on feature relationships to obtain a target content set. Image semantic control parameters are generated based on the target content set. These parameters are input into a preset image generation model for image feature generation processing. The generated image features are then rendered and encapsulated based on the image output format to obtain the drama text image generation result corresponding to the user input content. This embodiment acquires multi-source heterogeneous data related to drama and extracts multi-dimensional feature information. It combines category identification, cross-language mapping, and hierarchical organization to form a structured multimodal content resource library. Then, it establishes associations based on the correspondence between language data and performance data, enabling drama content to be retrieved and called in an orderly manner. On this basis, it obtains demand features, feature relationships, and output formats by parsing user input content. Based on this, it determines target content units in the multimodal content resource library, organizes target content sets, and completes format conversion. This reduces disordered filtering and irrelevant calls in the generation process, improves the efficiency of drama text image generation, enhances the correspondence between the generated results and user needs, and improves the accuracy of drama text image generation.
[0050] Based on the second embodiment described above, a third embodiment of the dramatic text image generation method of this application is proposed. Please refer to... Figure 3 , Figure 3 This is a schematic diagram of a sub-process in the third embodiment of the dramatic text image generation method of this application.
[0051] In this embodiment, step S3 includes: S31: Based on the correspondence between the language data and the performance data in terms of content theme, expression object, performance segment and time sequence position, the feature items of different information levels are associated and identified to determine the cross-level association clues between each information level, and the hierarchical association result between the language data and the performance data is obtained. S32: Based on the hierarchical association results, perform association mapping and relationship arrangement on the feature information of each level in the hierarchical organization results, and establish a language mapping relationship corresponding to the feature information of each level in combination with the language comparison results, so as to obtain the hierarchical association organization results corresponding to the hierarchical organization results; S33: Based on the hierarchical association organization results, the feature information of each level and the corresponding language mapping content in the hierarchical organization results are collected, stored and structured to obtain a multimodal content resource library including hierarchical structure, language correspondence relationship and cross-level association relationship.
[0052] It should be noted that: feature information refers to the specific content representation information contained within each information level in the hierarchical organization result; association identification refers to the process of determining the correspondence and identifying the relationship between related feature items in different information levels based on the correspondence between different data; cross-level association clues refer to the association information used to represent the content correspondence, pointing relationship, or connecting relationship between different information levels; hierarchical association result refers to the association result between each information level determined based on the correspondence between language data and performance data; association mapping refers to the process of establishing and mapping the correspondence between related feature information in different information levels; and relationship arrangement. Hierarchical association refers to the process of organizing and structuring the relationships between different levels; language mapping relationship refers to the correspondence established between the feature information of each level and the corresponding language expression content; hierarchical association organization result refers to the organizational result formed after mapping and arranging the relationship of feature information at each level and establishing language mapping relationship in combination with language comparison result; collection and storage refers to the process of summarizing and storing relevant content according to a preset method; structured organization refers to the processing method of organizing and structuring content according to a predetermined data structure and association logic; cross-level association relationship refers to the content association relationship or pointing relationship that exists between different information levels.
[0053] Specifically, based on the correspondence between the language data and the performance data in terms of content theme, expression object, performance segment, and temporal position, feature items at different information levels are associated and identified. Specifically, focusing on the consistency of thematic content, character object consistency, performance segment correspondence, and temporal connection within the dramatic content, relevant feature items distributed across different information levels are compared and analyzed to identify cross-level connections between language data and performance data. Through this process, cross-level connection clues corresponding to each information level can be determined, forming a hierarchical association result between the language data and the performance data. Then, based on the hierarchical association result, the feature information at each level in the hierarchical organization result is associated, mapped, and arranged so that relevant feature information at different information levels can be linked according to content correspondence; simultaneously, combined with the language comparison result, corresponding language mapping relationships are configured for the feature information at each level, so that the feature information at each level not only has cross-level connections but also corresponding language content, thereby obtaining a hierarchical association organization result corresponding to the hierarchical organization result.
[0054] Furthermore, after obtaining the hierarchical association organization result, the feature information of each level and the corresponding language mapping content in the hierarchical organization result are further collected, stored, and structured based on the hierarchical association organization result. Specifically, the feature information in each level, the association content between each level, and the corresponding language mapping content can be uniformly organized and stored according to a preset data organization method, so that all content forms a structured content organization result with clear hierarchical distribution, clear language correspondence, and traceable cross-level associations within the same resource system. Through the above processing, a multimodal content resource library including hierarchical structure, language correspondence relationship, and cross-level association relationship can be finally formed, so that the language expression content, performance presentation content, and the correspondence between different levels in drama are uniformly incorporated into a callable content resource system.
[0055] By identifying the correspondence between language and performance data in terms of content theme, expressive objects, performance segments, and temporal positions, feature items at different information levels can be correlated and identified. This establishes clear cross-level connections between language and performance content that were originally scattered across different levels. Further, based on the hierarchical correlation results, feature information at each level is mapped and arranged, and language mapping relationships are established in conjunction with language comparison results. This ensures that content at different levels not only forms an orderly connection but also possesses a foundation of language correspondence. Then, by collecting, storing, and structuring the feature information at each level and its corresponding language mapping content, a multimodal content resource library is formed, including hierarchical structure, language comparison relationships, and cross-level correlation relationships. This transforms the scattered storage of drama-related content into a unified, associative, and callable resource system. Therefore, this step provides a more complete resource and correlation foundation for subsequent content matching, target content unit determination, and result generation based on user needs. It is beneficial to improve the efficiency of content retrieval during the drama text image generation process and enhance the accuracy of the correspondence between the generated results and the context of the drama content.
[0056] Based on the second embodiment described above, in this embodiment, step S4 includes: S41: Obtain the user's input search request, generate instructions or interactive description information, and standardize the user input content to determine the corresponding target expression unit and content pointing range, and obtain the input parsing data corresponding to the user input content; S42: Based on the input parsing data, identify the topic objects, modal orientations, content constraints and combination requirements involved in the user input content, determine the feature elements corresponding to the user input content and the coordination relationship between each feature element, and obtain the demand features and feature relationships corresponding to the user input content; S43: Based on the input parsing data, the demand features, and the feature relationships, identify the result presentation requirements corresponding to the user input content, determine the corresponding content organization form and output expression type, and obtain the output format after parsing the user input content.
[0057] It should be noted that: a retrieval request refers to a query-type request information entered by a user to obtain specific drama-related content; a generation instruction refers to the generation control information entered by a user to obtain target text or image content; interactive description information refers to descriptive information entered by the user during the interaction process to explain the target content requirements or presentation requirements; normalization processing refers to the process of uniformly organizing, formatting, and standardizing the expression of user input content; target expression unit refers to the basic expression content unit determined from the user input content to represent the user's intention; content scope refers to the scope of target content covered by the user input content; input parsing data refers to the parsing result data obtained after normalizing the user input content; and the subject object is... The following are considered in the context of user input: 1. **Drama Character, Scene, Scene, Theme, or Related Content Object:** 2. **Modal Indication:** 3. **Content Constraints:** 4. **Scope Restrictions, Attribute Restrictions, or Content Constraints:** 5. **Combination Requirements:** ...
[0058] Specifically, firstly, the system acquires the user's input search request, generating instructions, or interactive description information, and then standardizes the user input. Specifically, text descriptions, instructions, or search statements in the user input are uniformly organized to eliminate differences caused by different expression methods, transforming them into a standardized content format that is easy to identify and process later. Based on this, the standardized user input is preliminarily parsed to determine the target expression units and content scope contained within, thereby obtaining the input parsing data corresponding to the user input. Subsequently, based on the input parsing data, the system identifies the subject matter, modal orientation, content constraints, and combination requirements involved in the user input, extracts various feature elements that can characterize user needs, and analyzes the correspondence, coordination, or constraint relationships between these feature elements, thereby obtaining the demand characteristics and feature relationships corresponding to the user input.
[0059] Furthermore, after obtaining the required features and the feature relationships, the result presentation requirements corresponding to the user input content are further identified based on the input parsing data, the required features, and the feature relationships. Specifically, the content organization form and output expression type of the final output content can be determined by combining the requirements of the user input content regarding the expression method, content combination method, and presentation type of the output content, so that the output result is consistent with the user's requirements in terms of content structure and presentation. Through the above processing, a continuous parsing process can be completed from user input content acquisition, input parsing data formation, required features and feature relationships identification, to result presentation requirements identification and output format determination, thereby obtaining the output format after parsing the user input content.
[0060] By acquiring user input search requests, generation instructions, or interactive descriptions, and standardizing the user input, user input from different modes of expression can be uniformly converted into analyzable input parsing data. Further, based on the input parsing data, the subject matter, modal orientation, content constraints, and combination requirements are identified, and demand characteristics and feature relationships are determined. This transforms vague descriptions of user needs into structured demand information with clear content orientation and organizational relationships. Then, based on the input parsing data and the demand characteristics and feature relationships, the presentation requirements are identified, and the corresponding content organization form and output expression type are determined. This provides clear output constraints and presentation directions for subsequent content matching and generation processes. Therefore, this step provides a clear demand foundation for the accurate determination of subsequent target content units, the orderly organization of target content sets, and the format conversion of target generation results. This improves the processing efficiency in the theatrical text image generation process and enhances the accuracy of the correspondence between the generated results and user input requirements.
[0061] In this embodiment, step S5 includes: S51: Based on the demand characteristics, the content items corresponding to the demand characteristics are retrieved and matched in the multimodal content resource library, and target content units that are adapted to the demand characteristics are filtered based on the matching results. Based on the characteristic relationship, the content pointing relationship, combination order relationship and cooperation constraint relationship between the target content units are associated and arranged to obtain the target content set corresponding to the demand characteristics. S52: Based on the semantic content, performance content, scene content and object relationship corresponding to each target content unit in the target content set, extract semantic control information to characterize the image generation requirements, and perform parameterized transformation processing on the semantic control information to generate image semantic control parameters corresponding to the target content set. Input the image semantic control parameters into a preset image generation model to control the preset image generation model to perform image feature generation processing according to the content semantics corresponding to the target content set, and obtain the corresponding image feature generation result. S53: Based on the image output format, perform image composition adjustment, presentation adaptation and result rendering encapsulation processing on the image feature generation result to obtain the drama text image generation result corresponding to the user input content.
[0062] It should be noted that: "Content item" refers to the basic content object in the multimodal content resource library that is relevant to user needs and can be retrieved and called; "content pointing relationship" refers to the correspondence between different target content units in terms of theme expression, object description, or semantic pointing; "combination order relationship" refers to the sequential arrangement of multiple target content units during organization; "cooperation constraint relationship" refers to the cooperation rules or restrictions that multiple target content units should satisfy when jointly expressing the same generation requirement; "semantic control information" refers to the descriptive information extracted from the target content set used to constrain the semantics of the generated image content; "parameterization conversion processing" refers to the process of converting semantic control information into a parameter expression form that can be called by the image generation model; "image semantic control parameters" refers to the parameter information formed after parameterization conversion, used to control the direction and content composition of image generation; and "preset" refers to the parameters used to control the direction and content composition of image generation. An image generation model refers to a pre-configured model used to generate image features based on input control parameters; image feature generation processing refers to the process of generating image features corresponding to a target content set using a preset image generation model; image feature generation result refers to the image feature result obtained after image feature generation processing; image output format refers to the presentation format requirements corresponding to the output of the image generation result; image composition adjustment refers to the process of adjusting the layout, hierarchy, or presentation mode of each component in the generated image; presentation format adaptation refers to the process of making the generated image conform to the predetermined expression format requirements; result rendering and encapsulation processing refers to the process of organizing, rendering, and outputting the generated image features; the dramatic text image generation result refers to the final image result obtained based on the image feature generation result and combined with image output format processing.
[0063] Specifically, firstly, based on the stated requirement characteristics, content items corresponding to the requirement characteristics are retrieved and matched in the multimodal content resource library, and target content units adapted to the requirement characteristics are selected based on the matching results. Specifically, content items corresponding to the theatrical themes, characters, performance content, scene information, and expression requirements involved in the user's requirements can be searched in the multimodal content resource library, and filtered based on the compatibility between the content items and the requirement characteristics, thereby determining target content units suitable for the current generation task. Then, based on the stated feature relationships, the content orientation relationships, combination order relationships, and coordination constraints among the target content units are associated and arranged, so that multiple target content units are connected in an orderly manner according to the content logic and expression requirements in the user's requirements, thus obtaining a set of target content corresponding to the requirement characteristics. Therefore, the subsequent image generation process can be based on the filtered and organized target content.
[0064] Furthermore, after obtaining the target content set, semantic control information for characterizing image generation requirements is extracted based on the semantic content, performance content, scene content, and object relationships corresponding to each target content unit in the target content set. This semantic control information is then parameterized to generate image semantic control parameters corresponding to the target content set. Subsequently, these image semantic control parameters are input into a preset image generation model to control the model to perform image feature generation processing according to the content semantics corresponding to the target content set, resulting in the corresponding image feature generation result. Finally, based on the image output format, the image feature generation result undergoes image composition adjustment, presentation adaptation, and result rendering encapsulation processing to ensure that the generated image meets the requirements corresponding to the user input content in terms of composition, content presentation, and output format, thereby obtaining a theatrical text image generation result corresponding to the user input content.
[0065] By retrieving, matching, and filtering target content units from a multimodal content resource library based on demand characteristics, the content upon which image generation is based can be established on a resource foundation adapted to user needs. Furthermore, by associating and arranging target content units based on feature relationships to form a target content set, multiple content units can be organized in an orderly manner according to the logical relationships and coordination requirements in user needs. Then, semantic control information is extracted from the target content set to generate image semantic control parameters, enabling the preset image generation model to generate image features around the semantic content represented by the target content set. Finally, the generated image features are adjusted in terms of image composition, adapted in terms of presentation, and rendered and encapsulated in conjunction with the image output format, ensuring that the output image maintains consistency with user input requirements in terms of content composition and presentation. Therefore, this step can effectively transform the theatrical content in the multimodal content resource library into a control basis for the image generation model and complete the image result output corresponding to user needs, thereby improving the processing efficiency of theatrical text image generation and enhancing the accuracy of the correspondence between the generated results and user needs and the semantics of the theatrical content.
[0066] This embodiment acquires multi-source heterogeneous data related to drama and extracts corresponding multi-dimensional feature information based on the multi-source heterogeneous data. The multi-source heterogeneous data includes at least language data and performance data. Category identification and cross-language mapping processing are performed on the multi-dimensional feature information to obtain corresponding feature identification results and language comparison results. The multi-dimensional feature information is then hierarchically organized based on the feature identification results. Based on the correspondence between language data and performance data, association relationships are established between each level. A multi-modal content resource library for image generation is constructed based on the hierarchical organization results, language comparison results, and association relationships. User input content is acquired and parsed to obtain demand features, feature relationships, and output formats. Target content units are determined based on the demand features and the multi-modal content resource library. These target content units are then associated and organized based on feature relationships to obtain a target content set. Image semantic control parameters are generated based on the target content set. These parameters are input into a preset image generation model for image feature generation processing. The generated image features are then rendered and encapsulated based on the image output format to obtain the drama text image generation result corresponding to the user input content. This embodiment acquires multi-source heterogeneous data related to drama and extracts multi-dimensional feature information. It combines category identification, cross-language mapping, and hierarchical organization to form a structured multimodal content resource library. Then, it establishes associations based on the correspondence between language data and performance data, enabling drama content to be retrieved and called in an orderly manner. On this basis, it obtains demand features, feature relationships, and output formats by parsing user input content. Based on this, it determines target content units in the multimodal content resource library, organizes target content sets, and completes format conversion. This reduces disordered filtering and irrelevant calls in the generation process, improves the efficiency of drama text image generation, enhances the correspondence between the generated results and user needs, and improves the accuracy of drama text image generation.
[0067] In one embodiment, after extracting multi-dimensional feature information from multi-source heterogeneous data related to drama and completing category identification processing, corresponding credibility markers and conflict status markers can be generated for each feature item. The credibility marker characterizes the data completeness, source consistency, and content stability of the corresponding feature item, while the conflict status marker characterizes whether there are inconsistencies in content, time, object, or meaning between the corresponding feature item and other feature items. Specifically, for language and performance data originating from different carriers, different acquisition stages, or different forms of expression, consistency comparisons can be performed on their corresponding feature items. When multiple feature items match in terms of theme, character, performance segment, or temporal position, the credibility marker of the corresponding feature item can be increased. When there are content conflicts or deviations in direction among multiple feature items, a conflict status marker can be added to the corresponding feature item, and the corresponding conflict type and source can be recorded.
[0068] When constructing a multimodal content resource library, the credibility markers and conflict status markers can be written as extended attributes into the storage structure corresponding to the feature information of each level. This allows each content unit in the resource library to retain not only the hierarchical structure, language correspondence, and cross-level association, but also credibility and conflict information that can be used for content filtering. Furthermore, when inputting demand features into the multimodal content resource library for matching, target content units whose credibility markers meet preset conditions and whose conflict status markers meet preset constraints can be prioritized. For candidate content items with conflict status markers, conflicting content can be selectively discarded, retained in parallel, or organized in an annotation manner by combining the topic object, modality orientation, content limitation conditions, or output format in the user input content. For example, when the same drama segment corresponds to different action descriptions or scene descriptions in different sources, content items with a high degree of correspondence to the current demand features and a high credibility marker can be prioritized, or multiple content items with differences can be organized in parallel into the target content set for subsequent output conversion.
[0069] In some implementations, a consistency review can be performed on the target content set before obtaining the target generation result. Specifically, based on the credibility markers, conflict status markers, and interrelationships of each target content unit in the target content set, it can be determined whether there are content breaks, object mismatches, or temporal jumps within the target content set. If any of these situations are found, the target content units in the target content set can be reordered, replaced, partially deleted, or supplemented to form a more consistent output content organization result. Then, based on the output format, the adjusted target content set undergoes format adaptation, presentation conversion, and result encapsulation processing to obtain the corresponding target generation result.
[0070] This application also provides a theatrical text image generation apparatus, please refer to... Figure 4 , Figure 4 This is a schematic diagram of the module structure of the theatrical text image generation device according to an embodiment of this application. The theatrical text image generation device includes: The data acquisition module 401 is used to acquire multi-source heterogeneous data related to drama, and extract corresponding multi-dimensional feature information based on the multi-source heterogeneous data, wherein the multi-source heterogeneous data includes at least language data and performance data; The hierarchical organization module 402 is used to perform category identification processing and cross-language mapping processing on the multidimensional feature information to obtain corresponding feature identification results and language comparison results, and to hierarchically organize the multidimensional feature information based on the feature identification results; The resource library construction module 403 is used to establish the association relationship between each level based on the correspondence between the language data and the performance data, and to construct a multimodal content resource library for image generation based on the hierarchical organization results, the language comparison results and the association relationship; The input parsing module 404 is used to acquire user input content and parse the user input content to obtain demand features, feature relationships and output format; The result generation module 405 is used to determine target content units based on the demand features and the multimodal content resource library, associate and organize the target content units based on the feature relationships to obtain a target content set, generate image semantic control parameters based on the target content set, input the image semantic control parameters into a preset image generation model for image feature generation processing, and render and encapsulate the generated image features based on the image output format to obtain a drama text image generation result corresponding to the user input content.
[0071] The theatrical text image generation apparatus provided in this application, employing the theatrical text image generation method described in the above embodiments, can solve the technical problem of how to improve the efficiency and accuracy of theatrical text image generation. Compared with the prior art, the beneficial effects of the theatrical text image generation apparatus provided in this application are the same as those of the theatrical text image generation method described in the above embodiments, and other technical features in the theatrical text image generation apparatus are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0072] This application provides a theatrical text image generation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the theatrical text image generation method in the above embodiments.
[0073] The following is for reference. Figure 5 , Figure 5 This is a schematic diagram of the hardware operating environment involved in the dramatic text image generation method in the embodiments of this application, showing a schematic diagram of the structure of a dramatic text image generation device suitable for implementing the embodiments of this application. Figure 5 The dramatic text image generation device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0074] like Figure 5As shown, the theatrical text image generation device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the theatrical text image generation device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the theatrical text image generation device to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows theatrical text image generation devices with various systems, it should be understood that it is not required to implement or possess all of the systems shown. More or fewer systems may be implemented alternatively.
[0075] In particular, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. When the computer program is executed by the processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0076] The theatrical text image generation device provided in this application, employing the theatrical text image generation method in the above embodiments, can solve the technical problem of how to improve the efficiency and accuracy of theatrical text image generation. Compared with the prior art, the beneficial effects of the theatrical text image generation device provided in this application are the same as those of the theatrical text image generation method provided in the above embodiments, and other technical features in this theatrical text image generation device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0077] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0078] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0079] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the dramatic text image generation method in the above embodiments.
[0080] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by the drama text-image generation device, the drama text-image generation device performs the following actions: acquires multi-source heterogeneous data related to drama, and extracts corresponding multi-dimensional feature information based on the multi-source heterogeneous data, wherein the multi-source heterogeneous data includes at least language data and performance data; performs category identification processing and cross-language mapping processing on the multi-dimensional feature information to obtain corresponding feature identification results and language comparison results, and organizes the multi-dimensional feature information hierarchically based on the feature identification results; establishes the association relationship between each level based on the correspondence between language data and performance data, and constructs a multimodal content resource library based on the hierarchical organization results, language comparison results, and association relationships; acquires user input content, parses the user input content, and obtains demand features, feature relationships, and output format; determines target content units based on demand features and the multimodal content resource library, organizes the target content units based on feature relationships to obtain a target content set, and performs format conversion on the target content set based on the output format to obtain the target generation result. Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0081] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0082] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0083] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described dramatic text image generation method, thereby solving the technical problem of how to improve the efficiency and accuracy of dramatic text image generation. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the dramatic text image generation method provided in the above embodiments, and will not be repeated here.
[0084] This application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the dramatic text image generation method described above.
[0085] The computer program product provided in this application can solve the technical problem of how to improve the efficiency and accuracy of generating theatrical text images. Compared with the prior art, the beneficial effects of the computer program product provided in the embodiments of this application are the same as the beneficial effects of the theatrical text image generation method provided in the above embodiments, and will not be repeated here.
[0086] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent scope of this application.
Claims
1. A method for generating theatrical text images, applied to computer equipment, characterized in that, The method includes: Acquire multi-source heterogeneous data related to drama, and extract corresponding multi-dimensional feature information based on the multi-source heterogeneous data, wherein the multi-source heterogeneous data includes at least language data and performance data; The multidimensional feature information is subjected to category identification processing and cross-language mapping processing to obtain corresponding feature identification results and language comparison results, and the multidimensional feature information is hierarchically organized based on the feature identification results; Based on the correspondence between the language data and the performance data, an association relationship is established between each level, and a multimodal content resource library for image generation is constructed based on the hierarchical organization results, the language comparison results, and the association relationship. Obtain user input content, parse the user input content, and obtain the demand features, feature relationships, and output format; Based on the required characteristics and the multimodal content resource library, target content units are determined. The target content units are associated and organized based on the feature relationships to obtain a target content set. Image semantic control parameters are generated based on the target content set. The image semantic control parameters are input into a preset image generation model for image feature generation processing. The generated image features are rendered and encapsulated based on the image output format to obtain the drama text image generation result corresponding to the user input content.
2. The method as described in claim 1, characterized in that, The step of acquiring multi-source heterogeneous data related to drama and extracting corresponding multi-dimensional feature information based on the multi-source heterogeneous data includes: According to the preset acquisition rules, text data, image data, audio data, video data and spatial reconstruction data corresponding to the drama content are acquired, and the source identification and object correspondence processing of the acquired data are performed to obtain the original data set corresponding to the drama object. Data cleaning, format standardization, and content alignment are performed on various types of data in the original dataset to form basic representation data corresponding to language expression content, performance action content, scene presentation content, and carrier form content, respectively. Based on the aforementioned basic representation data, feature information of representation semantic content, expression form, scene attributes, and associated objects is extracted, and the extracted feature information is integrated to obtain multidimensional feature information corresponding to the multi-source heterogeneous data.
3. The method as described in claim 1, characterized in that, The steps of performing category identification processing and cross-language mapping processing on the multidimensional feature information to obtain corresponding feature identification results and language correspondence results, and hierarchically organizing the multidimensional feature information based on the feature identification results, include: Based on the content attributes, performance types and source categories corresponding to each feature item in the multidimensional feature information, the multidimensional feature information is classified, identified and labeled to determine the category and identification content corresponding to each feature item, and the feature identification result corresponding to the multidimensional feature information is obtained. Based on the feature identification results, the identification content related to language expression in each feature item is extracted, and the identification content related to language expression is converted and organized across languages according to the preset language correspondence rules to obtain the language comparison results corresponding to the feature identification results; Based on the feature identification results, the hierarchical attribution rules corresponding to each feature item in the multidimensional feature information are determined, and the multidimensional feature information is divided into the corresponding information levels according to the hierarchical attribution rules, so as to obtain the result of hierarchical organization of the multidimensional feature information based on the feature identification results.
4. The method as described in claim 1, characterized in that, The step of establishing relationships between different levels based on the correspondence between the language data and the performance data, and constructing a multimodal content resource library for image generation based on the hierarchical organization results, the language comparison results, and the relationships, includes: Based on the correspondence between the language data and the performance data in terms of content theme, expression object, performance segment and time sequence position, the feature items of different information levels are associated and identified to determine the cross-level association clues between each information level, and the hierarchical association results between the language data and the performance data are obtained. Based on the hierarchical association results, the feature information of each level in the hierarchical organization results is associated and mapped and arranged, and the language mapping relationship corresponding to each level feature information is established in combination with the language comparison results to obtain the hierarchical association organization results corresponding to the hierarchical organization results. Based on the hierarchical association organization results, the feature information of each level and the corresponding language mapping content in the hierarchical organization results are collected, stored and structured to obtain a multimodal content resource library for image generation, including hierarchical structure, language correspondence relationship and cross-level association relationship.
5. The method as described in claim 1, characterized in that, The steps of obtaining user input and parsing the user input to obtain demand features, feature relationships, and output format include: The system acquires user input search requests, generates instructions or interactive description information, and performs normalization processing on the user input content to determine the corresponding target expression unit and content pointing range, thereby obtaining the input parsing data corresponding to the user input content. Based on the input parsing data, the topic objects, modal orientations, content constraints and combination requirements involved in the user input content are identified, the feature elements corresponding to the user input content and the coordination relationship between each feature element are determined, and the demand features and feature relationships corresponding to the user input content are obtained. Based on the input parsing data, the demand features, and the feature relationships, the result presentation requirements corresponding to the user input content are identified, and the corresponding content organization form and output expression type are determined to obtain the output format after parsing the user input content.
6. The method as described in claim 1, characterized in that, The steps of determining target content units based on the demand characteristics and the multimodal content resource library, associating and organizing the target content units based on the feature relationships to obtain a target content set, generating image semantic control parameters based on the target content set, inputting the image semantic control parameters into a preset image generation model for image feature generation processing, and rendering and encapsulating the generated image features based on the image output format to obtain the drama text image generation result corresponding to the user input content include: Based on the demand characteristics, the content items corresponding to the demand characteristics are retrieved and matched in the multimodal content resource library, and target content units that are adapted to the demand characteristics are filtered based on the matching results. Based on the characteristic relationship, the content pointing relationship, combination order relationship and cooperation constraint relationship between the target content units are associated and arranged to obtain the target content set corresponding to the demand characteristics. Based on the semantic content, performance content, scene content, and object relationships corresponding to each target content unit in the target content set, semantic control information for characterizing image generation requirements is extracted, and the semantic control information is parameterized and transformed to generate image semantic control parameters corresponding to the target content set. The image semantic control parameters are then input into a preset image generation model to control the preset image generation model to perform image feature generation processing according to the content semantics corresponding to the target content set, thereby obtaining the corresponding image feature generation result. Based on the image output format, the image feature generation result is processed by image composition adjustment, presentation adaptation, and result rendering encapsulation to obtain the drama text image generation result corresponding to the user input content.
7. A theatrical text image generation device, applied to computer equipment, characterized in that, The device includes: The data acquisition module is used to acquire multi-source heterogeneous data related to drama, and extract corresponding multi-dimensional feature information based on the multi-source heterogeneous data, wherein the multi-source heterogeneous data includes at least language data and performance data; The hierarchical organization module is used to perform category identification processing and cross-language mapping processing on the multidimensional feature information to obtain corresponding feature identification results and language comparison results, and to hierarchically organize the multidimensional feature information based on the feature identification results; The resource library construction module is used to establish the association between each level based on the correspondence between the language data and the performance data, and to construct a multimodal content resource library for image generation based on the hierarchical organization results, the language comparison results and the association. The input parsing module is used to acquire user input content and parse the user input content to obtain demand features, feature relationships and output format; The result generation module is used to determine target content units based on the required features and the multimodal content resource library, associate and organize the target content units based on the feature relationships to obtain a target content set, generate image semantic control parameters based on the target content set, input the image semantic control parameters into a preset image generation model for image feature generation processing, and render and encapsulate the generated image features based on the image output format to obtain the drama text image generation result corresponding to the user input content.
8. A computer device, characterized in that, The device includes: a memory, a processor, and a theatrical text image generation program stored in the memory and executable on the processor, the theatrical text image generation program being configured to implement the steps of the theatrical text image generation method as claimed in any one of claims 1 to 6.
9. A storage medium, characterized in that, The storage medium stores a theatrical text image generation program, which, when executed by a processor, implements the steps of the theatrical text image generation method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the dramatic text image generation method as described in any one of claims 1 to 6.