Controllable Text Generation Method for Generative Large Language Models
Through the generative large model, the attribute processing and interactive determination of controllable information on image or video files is solved, and the problem that the existing technology cannot meet the user's controllable and personalized needs is achieved, and high-quality and multi-level controllable text generation is achieved.
Patent Information
- Application Number
- CN202510335337.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-20
AI Technical Summary
The existing image and video text generation technologies cannot meet the user's controllable and personalized needs, and lack in-depth understanding and precise control of user specific needs.
Through a generative large model, the image or video file is processed at attributes, the image or video file is analyzed, the controllable template is generated, the controllable information is determined interactively, the file is divided and the controllable sub-information is extracted, and the content is transformed, and multi-level controllable text is generated.
It realizes accurate response to user needs. The generated text is not only accurate in content, but also in the structure and length of the text meet user needs, improving the quality and practicality of text generation.
Smart Images

Figure CN119850898B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to data processing technologies, and in particular to a method for generating controllable text of a generative large model. Background Art
[0002] Currently, text generation technologies for images and videos mainly rely on traditional text generation models. In the process of processing, these models often adopt relatively general algorithms, lacking accurate control over the content of images and videos and in-depth understanding of specific user needs. That is, although existing generative models can output relatively smooth descriptive text based on the input image or video content, most of them stay in the "overall description" or "automatic summary" stage, lacking fine-grained support for user controllable needs. They cannot meet the controllable and personalized text output of users.
[0003] Therefore, how to perform controllable and personalized text output in combination with user needs has become a technical problem to be urgently solved. Summary of the Invention
[0004] An embodiment of the present invention provides a method for generating controllable text of a generative large model, which can perform controllable and personalized text output in combination with user needs.
[0005] In a first aspect of an embodiment of the present invention, a method for generating controllable text of a generative large model is provided, including:
[0006] The generative large model processes the attributes of the first modal file to be analyzed to generate a corresponding parsing controllable template;
[0007] Interacts with the parsing controllable template to determine the controllable information of the controllable text of the first modal file;
[0008] Based on the controllable information, the first modal file is segmented according to corresponding attributes to obtain a plurality of segmented fragments and the controllable sub-information of each segmented fragment;
[0009] Content extraction and conversion processing are performed on each segmented fragment to generate a corresponding initial text, and the initial text is processed based on the controllable sub-information to obtain the multi-level controllable text of the first modal file.
[0010] Optionally, the generative large model processes the attributes of the first modal file to be analyzed to generate a corresponding parsing controllable template, including:
[0011] A1. If the attributes of the first modal file include image attributes, an image interaction preview box corresponding to the image attributes is generated;
[0012] A2. Adaptive processing is performed on the interaction preview box based on the specifications in the attributes of the first modal file so that the interaction preview box fits the first modal file correspondingly;
[0013] A3. After extracting the corresponding specifications of the interactive preview box to generate the corresponding coordinate selection lines, generate the corresponding parseable and controllable template.
[0014] Optionally, adaptively processing the interactive preview box according to the specifications in the first modal file attributes to make the interactive preview box fit the first modal file correspondingly, including:
[0015] Obtain the number of horizontal pixel points and the number of vertical pixel points of the first modal file;
[0016] Adaptively process the interactive preview box according to the specifications of the number of horizontal pixel points and the number of vertical pixel points, so that the boundary of the interactive preview box fits the first modal file.
[0017] Optionally, after extracting the corresponding specifications of the interactive preview box to generate the corresponding coordinate selection lines, generating the corresponding parseable and controllable template includes:
[0018] Determine the target horizontal frame and the target vertical frame in the interactive preview box, and take the intersection point of the target horizontal frame and the target vertical frame as the starting point;
[0019] Taking the starting point as the initial position, establish the corresponding coordinate lines according to the target horizontal frame and the target vertical frame respectively, and the end points of the target horizontal frame and the target vertical frame are the cut-off positions;
[0020] Calculate the coordinates of the unit number of pixel points based on the number of pixel points between the initial position and the cut-off position, select the pixel points corresponding to the preset coordinate values to add coordinate selection lines, and obtain the parseable and controllable template.
[0021] Optionally, interacting with the parseable and controllable template to determine the controllable information of the controllable text of the first modal file includes:
[0022] After determining that the user selects any one of the coordinate selection lines, receive the first mark added by the user to the coordinate selection line, and generate the corresponding first controllable slot and the first confirmation slot at the parseable and controllable template;
[0023] After determining that the first confirmation slot is triggered, count the first closed area formed by all the coordinate selection lines with the same first mark;
[0024] If it is determined that the user selects a first closed area, use it as the pre-segmentation area, and receive the controllable sub-information about the pre-segmentation area based on the first controllable slot. The controllable information includes at least the text requirements of the controllable text.
[0025] Optionally, splitting and processing the first modal file according to the corresponding attributes based on the controllable information to obtain multiple split segments and the controllable sub-information of each split segment, including:
[0026] The first modal file is segmented according to the pre-segmented area to obtain multiple segmented fragments;
[0027] Extract the controllable sub-information of each segmented fragment, where the controllable sub-information includes at least one of the text threshold quantity and the text ratio.
[0028] Optionally, content extraction and conversion processing are performed on each segmented fragment to generate corresponding initial text, and the initial text is processed based on the controllable sub-information to obtain the multi-level controllable text of the first modal file, including:
[0029] Content extraction is performed on each segmented fragment to obtain the first text content and / or image content, and the image content is converted into the preset second text content. The initial text includes the first text content and the second text content;
[0030] Based on the text threshold quantity and the text ratio corresponding to the controllable sub-information, the initial text is fused to obtain the level and controllable sub-text of each segmented fragment;
[0031] The levels and controllable sub-texts of the segmented fragments are counted and filled into the preset area of the output template to obtain the multi-level controllable text of the first modal file.
[0032] Optionally, the initial text is fused based on the text threshold quantity and the text ratio corresponding to the controllable sub-information to obtain the level and controllable sub-text of each segmented fragment, including:
[0033] If it is determined that the text threshold quantity of the controllable sub-information is not set and the text ratio is not set, and the number of controllable sub-texts is less than or equal to the first text threshold quantity, then save;
[0034] If the number of controllable sub-texts is greater than the first threshold quantity, the controllable sub-text is input into a text simplification model for text reduction processing until it is less than or equal to the first text threshold quantity and then saved;
[0035] Generate corresponding levels according to the number of pixel points in each segmented fragment, and store the levels and controllable sub-texts of each segmented fragment correspondingly.
[0036] Optionally, the initial text is fused based on the text threshold quantity and the text ratio corresponding to the controllable sub-information to obtain the level and controllable sub-text of each segmented fragment, including:
[0037] If it is determined that the text threshold quantity of the controllable sub-information is not set and the text ratio is set, calculate the pixel point ratio of all segmented fragments according to the number of pixel points in each segmented fragment;
[0038] Extract the highest ratio value in the pixel point ratio and the corresponding number of controllable sub - texts as the highest number, and determine the second threshold number of the controllable sub - information of other segmentation segments based on the highest number and the pixel point ratio;
[0039] Based on the second threshold number and the number of pixel points in each segmentation segment, obtain the level of each segmentation segment and store the controllable sub - texts correspondingly.
[0040] Optionally, the obtaining the level of each segmentation segment and storing the controllable sub - texts correspondingly based on the second threshold number and the number of pixel points in each segmentation segment includes:
[0041] If the number of controllable sub - texts in the segmentation segment is less than or equal to the second threshold number of the text, save it;
[0042] If the number of controllable sub - texts is greater than the second threshold number, input the controllable sub - texts into a text simplification model for text reduction processing until it is less than or equal to the second threshold number of the text and then save it;
[0043] Generate corresponding levels according to the number of pixel points in each segmentation segment, and store the levels and controllable sub - texts of each segmentation segment correspondingly.
[0044] Optionally, the generative large - model processes the attributes of the first modal file to be analyzed and generates a corresponding parsing controllable template, including:
[0045] If the attributes of the first modal file include video attributes, generate a video interaction preview box corresponding to the video attributes, and the video interaction preview box includes a video interaction axis;
[0046] Based on the specifications in the attributes of the first modal file, perform adaptive processing on the interaction preview box so that the interaction preview box fits and corresponds to the first modal file;
[0047] After extracting the video interaction axis to generate a corresponding time - point selection line, generate a corresponding parsing controllable template.
[0048] Optionally, the extracting the video interaction axis to generate a corresponding time - point selection line and then generating a corresponding parsing controllable template includes:
[0049] Starting from the 0 - point of the video interaction axis, determine a time point at intervals of a preset time period and generate a corresponding time - point selection line.
[0050] Optionally, the interacting with the parsing controllable template to determine the controllable information of the controllable text of the first modal file includes:
[0051] After determining that the user has selected any time - point selection line, receive the second mark added by the user to the time - point selection line, and generate a second controllable slot and a second confirmation slot corresponding to the second mark at the parsed controllable template;
[0052] After determining that the second confirmation slot is triggered, count the minimum time and the maximum time with the same second mark to obtain a time period;
[0053] Receive controllable sub - information about the corresponding time period based on the first controllable slot, and the controllable information at least includes the text requirements of the controllable text.
[0054] Optionally, it further includes:
[0055] If it is determined that the user has selected a frame image at any moment in the video interaction axis and input the image attributes, execute steps A1 to A3.
[0056] In this solution, an image interaction preview box is generated for the first - mode file with image attributes, and adaptive processing is performed based on the image specifications to make it fit the image precisely. At the same time, the preview - box specifications are extracted to generate a coordinate selection line, thereby obtaining a parsed controllable template. When processing the video - attribute file, a video interaction preview box containing a video interaction axis is generated, and a time - point selection line is generated for the video interaction axis according to a preset time period to form a parsed controllable template. This method provides an intuitive and precise operation interface for users. Users can conveniently specify the region of interest or time period by clicking on the coordinate selection line or the time - point selection line, greatly improving the operation accuracy and interactivity, and providing a solid foundation for subsequent determination of controllable information.
[0057] In this solution, the user interacts with the parsed controllable template, determines the controllable information of the controllable text through operations such as adding marks and generating controllable slots. The system processes the first - mode file according to the attributes based on this information, obtains multiple segmented fragments, and extracts the controllable sub - information of each fragment. In the image field, the text requirements can be determined according to the user - marked area, such as the text threshold quantity, text ratio, etc.; in the video field, the text - generation requirements can be clarified according to the time period marked by the user. This mechanism enables text generation to be customized for different regions or time periods, realizes the refined operation of multimedia files, and improves the matching degree between the generated text and the user's needs.
[0058] This solution extracts and converts the content of each segmented fragment to generate the initial text, and then processes the initial text based on the controllable sub-information. According to the number of text thresholds and text ratios, etc., technologies such as the text simplification model are used to obtain the hierarchy and controllable sub-text of each segmented fragment. Finally, statistics are made and filled into the output template to obtain the multi-level controllable text. This process fully considers the diverse needs of users. The generated text not only accurately reflects the actual situation of each part of the image or video in terms of content, but also has good hierarchy and rationality in terms of structure and length. Taking images as an example, the segmented fragments of important regions can generate detailed and higher-level text descriptions; in the case of videos, the text length and hierarchy of key plot time periods can be reasonably set according to their importance. This greatly improves the quality and practicality of text generation, provides users with text output results that better meet expectations, and meets the complex requirements for text generation in different scenarios. Brief Description of the Drawings
[0059] Figure 1 is a schematic flowchart of a method for generating controllable text of a generative large model provided by an embodiment of the present invention;
[0060] Figure 2 is a schematic diagram of an image interaction preview box provided by an embodiment of the present invention. Detailed Description of the Embodiments
[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0062] See Figure 1 , which is a schematic flowchart of a method for generating controllable text of a generative large model provided by an embodiment of the present invention. The method includes:
[0063] S1. The generative large model processes the attributes of the first modal file to be analyzed and generates a corresponding parsing controllable template.
[0064] In the method for generating controllable text of the generative large model of the present invention, processing the attributes of the first modal file and generating a parsing controllable template is a key starting step. For the first modal file with image attributes, a parsing controllable template is generated by generating an image interaction preview box, performing adaptive processing, and extracting coordinate selection lines, providing a basis and operation interface for subsequent determination of controllable information, segmentation of the file, and generation of controllable text, enabling the entire text generation process to be carried out under the precise control of the user.
[0065] The generative large model processes the attributes of the first modal file to be analyzed and generates a corresponding parseable and controllable template, including:
[0066] A1. If the attributes of the first modal file include image attributes, an image interaction preview box corresponding to the image attributes is generated.
[0067] See Figure 2 , when the attributes of the first modal file to be analyzed include image attributes, the generative large model will generate an image interaction preview box corresponding to the image attributes. This image interaction preview box is a visual interface element, and its main function is to provide an intuitive operation platform for users to facilitate subsequent interactive operations on the image.
[0068] A2. Adaptive processing of the interaction preview box based on the specifications in the attributes of the first modal file to make the interaction preview box fit the first modal file.
[0069] The image interaction preview box has the same shape as the image, but is larger than the image to wrap the image. For example, for a common landscape image, the generative large model will identify its format and generate an image interaction preview box that wraps the landscape image according to the system display settings and user interface layout requirements.
[0070] Among them, the adaptive processing of the interaction preview box based on the specifications in the attributes of the first modal file to make the interaction preview box fit the first modal file includes:
[0071] Obtain the number of horizontal pixel points and the number of vertical pixel points of the first modal file;
[0072] Adaptive processing of the interaction preview box based on the specifications of the number of horizontal pixel points and the number of vertical pixel points to make the interaction preview box fit the boundary of the first modal file.
[0073] To ensure that the image interaction preview box can accurately reflect the actual range and content of the image, it is necessary to perform adaptive processing on the interaction preview box based on the specifications in the attributes of the first modal file to make the interaction preview box fit the first modal file. The specific implementation method is as follows:
[0074] First, the system will obtain the number of horizontal pixel points and the number of vertical pixel points of the first modal file. These two parameters are key information describing the image size and can be accurately obtained by reading the metadata of the image file or using image analysis algorithms.
[0075] Then, based on the specifications of the obtained number of horizontal pixel points and vertical pixel points, the interactive preview frame is adaptively adjusted. The adjustment process mainly dynamically modifies the size, proportion, etc. of the interactive preview frame so that its boundary fits the boundary of the first modal file.
[0076] For example, if the number of horizontal pixel points of an image is 1920 and the number of vertical pixel points is 1080, the system will adjust the width and height of the image interactive preview frame according to this specification to ensure that the size of the preview frame is consistent with the shape of the image, but it needs to be slightly larger. For example, 10 pixel points larger in each of the top, bottom, left, and right directions to wrap the image, so that the image can be displayed completely and without distortion within the preview frame. The advantage of doing this is that users can more accurately select and process each part of the image in subsequent operations.
[0077] A3. After extracting the corresponding specifications of the interactive preview frame to generate the corresponding coordinate selection lines, generate the corresponding parseable and controllable template.
[0078] After completing the adaptive processing of the interactive preview frame, it is necessary to extract the corresponding specifications of the interactive preview frame to generate the corresponding coordinate selection lines, and then generate the corresponding parseable and controllable template.
[0079] In some embodiments, the extracting the corresponding specifications of the interactive preview frame to generate the corresponding coordinate selection lines and then generating the corresponding parseable and controllable template includes:
[0080] Determine the target horizontal frame and target vertical frame in the interactive preview frame, and use the intersection point of the target horizontal frame and target vertical frame as the starting point. It can be understood that the target horizontal frame and target vertical frame can be the lower horizontal line and left vertical line of the interactive preview frame, and the starting point is the lower left corner point of the interactive preview frame.
[0081] Taking the starting point as the initial position, establish the corresponding coordinate lines according to the target horizontal frame and target vertical frame respectively, and use the end points of the target horizontal frame and target vertical frame as the end positions. In this solution, the end points (the endpoints of the line) of the target horizontal frame and target vertical frame are used as the end positions. In this way, a two-dimensional coordinate system is formed to locate each position in the image.
[0082] Calculate the coordinates of the unit number of pixel points based on the number of pixel points between the initial position and the end position, select the pixel points corresponding to the preset coordinate values to add coordinate selection lines, and obtain the parseable and controllable template.
[0083] Exemplarily, if the number of horizontal pixels between the initial point and the cut-off point is 1000 and its length is 10 cm, then one pixel corresponds to 0.01 cm. If the preset coordinate value is 0.5 cm, for example, then the corresponding number of pixels is 50. At this time, a coordinate selection line can be added every 50 pixels. By adding the coordinate selection line, the user can more conveniently select a specific area in the image.
[0084] After adding the coordinate selection line, an analytically controllable template containing coordinate information and selection functions is obtained. This template provides a basis for subsequent interaction with the user and determination of controllable information. The user can specify the image area to be processed by clicking on the coordinate selection line.
[0085] S2. Interact with the analytically controllable template to determine the controllable information of the controllable text of the first modal file.
[0086] In the method for generating controllable text of the generative large model of the present invention, interacting with the analytically controllable template and determining the controllable information is a crucial link connecting the previous and the next. The analytically controllable template generated in the previous steps provides an operation interface for the user, and this step, through the interaction between the user and the template, clarifies the processing requirements for specific areas in the first modal file, that is, determines the controllable information of the controllable text. These controllable information will guide the subsequent segmentation of the first modal file and the text generation process to ensure that the finally generated text meets the specific needs of the user.
[0087] In some embodiments, the interacting with the analytically controllable template to determine the controllable information of the controllable text of the first modal file includes:
[0088] S21. After determining that the user has selected any one of the coordinate selection lines, receive the first mark added by the user to the coordinate selection line, and generate a first controllable slot and a first confirmation slot corresponding to the first mark at the analytically controllable template.
[0089] After determining that the user has selected any one of the coordinate selection lines, the system will enter the information interaction process. The user's operation of selecting the coordinate selection line indicates that the user starts to specify the part of the first modal file that they are interested in. At this time, the system receives the first mark added by the user to the coordinate selection line. The first mark is a user-defined identifier used to distinguish different areas or contents. For example, the user may add marks such as "1", "2", etc.
[0090] Next, the system generates a first controllable slot and a first confirmation slot corresponding to the first marker at the parsed controllable template. The first controllable slot is an input area where the user can enter specific requirements regarding the marked area, such as the number of words to be generated for the text. The first confirmation slot is used for the user to confirm whether the input information is correct and complete. From a technical implementation perspective, the system dynamically creates these two slots on the interface of the parsed controllable template and associates them with the first marker added by the user.
[0091] S22. After determining that the first confirmation slot is triggered, count the first enclosed area formed by all coordinate selection lines with the same first marker.
[0092] After the user completes the input of information in the first controllable slot, the first confirmation slot is triggered. After receiving the confirmation signal, the system counts the first enclosed area formed by all coordinate selection lines with the same first marker. This counting process involves geometric analysis and logical judgment of the coordinate selection lines. The system determines which coordinate selection lines form a closed figure based on the positions and connection relationships of the coordinate selection lines.
[0093] For example, in an image interaction preview box, if the user adds the marker "1" to multiple coordinate selection lines, the system analyzes the positions of these coordinate selection lines through an algorithm to determine whether they can enclose a closed area. If so, this closed area is recorded as the first enclosed area. This enclosed area will serve as a basic unit for subsequent processing.
[0094] It can be understood that several coordinate selection lines marked with "2" can also form the first enclosed area, that is, there can be multiple first enclosed areas.
[0095] S23. If it is determined that the user selects a first enclosed area, use it as the pre-segmentation area, and receive controllable sub-information regarding the pre-segmentation area based on the first controllable slot. The controllable information at least includes the text requirements of the controllable text.
[0096] The pre-segmentation area is the basis for subsequent segmentation processing of the first modal file, which clarifies the specific range where specific text generation is required. Then, receive controllable sub-information regarding the pre-segmentation area based on the first controllable slot. The controllable information at least includes the text requirements of the controllable text, such as the word count range of the text, the content focus (such as describing the appearance, function, etc. of an object). The system associates these controllable sub-information with the pre-segmentation area to process according to these requirements during the subsequent text generation process.
[0097] Through the above interaction process with the parseable controllable template, the processing requirements of the user for specific regions in the first-modal file can be accurately determined, that is, the controllable information of the controllable text. This interaction method gives the user control, and the user can flexibly specify the processing region and text requirements according to their own needs. At the same time, the system can accurately identify and record this information, providing clear guidance for subsequent file segmentation and text generation, greatly improving the pertinence and accuracy of the generated text, and meeting the diverse needs of different users in different scenarios.
[0098] S3. Based on the controllable information, the first-modal file is segmented and processed according to the corresponding attributes to obtain multiple segmentation segments and the controllable sub-information of each segmentation segment.
[0099] In the method for generating controllable text of the generative large model of the present invention, the segmentation and processing of the first-modal file based on the controllable information is a key step that links the preceding and the following. The preceding steps have determined the controllable information and the pre-segmentation region, and this step actually segments the first-modal file according to this information to obtain multiple segmentation segments, and extracts the controllable sub-information of each segment. This lays a foundation for subsequent content extraction and conversion processing of each segmentation segment, and then generating the corresponding controllable text, enabling the text generation to process each specific region more precisely according to the user's needs.
[0100] In some embodiments, the segmentation and processing of the first-modal file based on the controllable information according to the corresponding attributes to obtain multiple segmentation segments and the controllable sub-information of each segmentation segment includes:
[0101] S31. Segment and process the first-modal file according to the pre-segmentation region to obtain multiple segmentation segments.
[0102] After determining the pre-segmentation region, the system will segment and process the first-modal file according to these pre-segmentation regions, thereby obtaining multiple segmentation segments. The system will cut out the image from the whole according to the coordinate range corresponding to the pre-segmentation region. For example, if the pre-segmentation region is a rectangular region determined by the coordinate selection line, the system will extract the image pixel information within the rectangular region and use it as an independent segmentation segment. This process involves an image cropping algorithm, and the system will accurately extract the corresponding image part according to the boundary coordinates of the pre-segmentation region.
[0103] S32. Extract the controllable sub-information of each segmentation segment, where the controllable sub-information includes at least one of the text threshold quantity and the text ratio.
[0104] After obtaining multiple segmentation segments, the system will extract the controllable sub-information of each segmentation segment. The controllable sub-information includes at least one of the text threshold quantity and the text ratio.
[0105] Among them, the text threshold quantity refers to the maximum or minimum word count limit allowed when generating text for the segmentation fragment. For example, for a segmentation fragment of a person's face in an image, the user may set the text threshold quantity to 200 words, requiring that the generated descriptive text does not exceed this word count or is exactly 200 words. When the system processes this segmentation fragment subsequently, it will control the length of the generated text according to this threshold.
[0106] The text ratio is used to determine the length of the generated text according to the ratio of the number of pixel points between segmentation fragments. For example, some segmentation fragments do not set the text threshold quantity, while some do. Then, this solution can determine the word count of the text for the segmentation fragments without setting according to the set text threshold quantity and text ratio.
[0107] The system accurately extracts the controllable sub-information corresponding to each segmentation fragment by associating with the previously recorded controllable information. For each segmentation fragment, these controllable sub-information will serve as important constraint conditions for subsequent content extraction and text generation.
[0108] By segmenting the first-modal file according to the pre-segmented regions and extracting the controllable sub-information of each segmentation fragment, this step realizes the refined processing of the first-modal file. Splitting the file into multiple fragments enables subsequent text generation to be targeted at different local regions, improving the pertinence of text generation. At the same time, the extraction of controllable sub-information provides clear constraints and guidance for the text generation process, ensuring that the generated text meets the user's expectations in terms of word count, length, etc., and better satisfying the diverse needs of users.
[0109] S4. Perform content extraction and conversion processing on each segmentation fragment to generate the corresponding initial text, and obtain the multi-level controllable text of the first-modal file after processing the initial text based on the controllable sub-information.
[0110] In the complete process of the method for generating controllable text of the generative large model in the present invention, this step is in a crucial final stage, undertaking the important task of converting each segmentation fragment obtained from the previous segmentation processing into controllable text. Through content extraction, conversion of the segmentation fragment, and processing based on the controllable sub-information, multi-level controllable text that meets the specific needs of users is finally generated, ensuring that the generated text not only accurately reflects the content of the first-modal file but also meets the diverse requirements of users in terms of structure and length.
[0111] In some embodiments, the performing content extraction and conversion processing on each segmentation fragment to generate the corresponding initial text, and obtaining the multi-level controllable text of the first-modal file after processing the initial text based on the controllable sub-information includes:
[0112] S41. For each segmented fragment, content extraction is performed to obtain the first text content and / or image content. The image content is converted into a preset second text content. The initial text includes the first text content and the second text content.
[0113] For each segmented fragment, the system will perform a comprehensive content extraction operation. In this process, according to the actual situation of the segmented fragment, the first text content and / or image content will be obtained. If the segmented fragment itself contains text information, such as captions in pictures, paragraphs in documents, etc., this content will be directly extracted as the first text content. When there are image elements in the segmented fragment, the system will use advanced image recognition and understanding technologies to convert the image content into a preset second text content.
[0114] For example, when processing an image segmented fragment containing animals, the image recognition algorithm will identify features such as the species, posture, and color of the animals, and convert this information into a text description to form the second text content. Finally, the first text content and the second text content are combined to form the initial text of each segmented fragment.
[0115] S42. Based on the text threshold quantity and text ratio corresponding to the controllable sub-information, the initial text is fused and processed to obtain the level and controllable sub-text of each segmented fragment.
[0116] In this step, according to different setting situations of the controllable sub-information, different strategies are adopted to fuse and process the initial text to obtain the controllable sub-text of each segmented fragment.
[0117] Among them, the step of fusing and processing the initial text based on the text threshold quantity and text ratio corresponding to the controllable sub-information to obtain the level and controllable sub-text of each segmented fragment includes:
[0118] S421. If it is judged that the text threshold quantity of the controllable sub-information is not set and the text ratio is not set, and the number of controllable sub-texts is less than or equal to the first text threshold quantity, then save it.
[0119] When it is judged that the text threshold quantity of the controllable sub-information is set but the text ratio is not set, the system will first check the relationship between the number of controllable sub-texts and the first text threshold quantity. If the number of controllable sub-texts is less than or equal to the first text threshold quantity, it means that the current generated text length meets the requirements set by the user. At this time, the system will directly save the controllable sub-text.
[0120] S422. If the number of controllable sub-texts is greater than the first threshold quantity, then the controllable sub-text is input into a text simplification model for text reduction processing until it is less than or equal to the first text threshold quantity and then saved.
[0121] If the number of controllable sub - texts is greater than the first threshold number, in order to meet the text length limit set by the user, the system will input the controllable sub - texts into a text simplification model. The text simplification model will use natural language processing technology to streamline and refine the text, remove redundant information, simplify the expression, and reduce the number of texts to be less than or equal to the first threshold number before saving. For example, the initial text is: "This remarkable new mobile phone has an extremely fashionable and unique appearance design. Its body lines are smooth and elegant, like lively notes jumping in space." The simplified text is "This new mobile phone has a fashionable and unique appearance, and its body lines are smooth." Among them, redundant modifiers are removed, such as "remarkable", "extremely", and "like lively notes jumping in space" and other expressions with strong modification but little impact on the core information are removed, making the text more concise and clear. Of course, it can also be processed by manual removal. This is prior art and will not be elaborated here.
[0122] S423. Generate corresponding levels according to the number of pixel points in each segmentation fragment, and store the levels of each segmentation fragment and the controllable sub - texts correspondingly.
[0123] After completing the processing of the number of texts, the system will generate corresponding levels according to the number of pixel points in each segmentation fragment. Generally speaking, the segmentation fragment with a larger number of pixel points may represent more important and complex content, so it will be assigned a higher level. Finally, store the levels of each segmentation fragment and the controllable sub - texts correspondingly for subsequent integration and output.
[0124] Based on the above - mentioned embodiments, the initial text is fusion - processed according to the text threshold number and text ratio corresponding to the controllable sub - information, and the levels of each segmentation fragment and the controllable sub - texts are obtained, including:
[0125] S424. If it is judged that the text threshold number of the controllable sub - information is not set and the text ratio is set, calculate the pixel - point ratio of all segmentation fragments according to the number of pixel points in each segmentation fragment.
[0126] When the text threshold number of the controllable sub - information is not set but the text ratio is set, the system will calculate according to the number of pixel points in each segmentation fragment to obtain the pixel - point ratio of all segmentation fragments. This ratio reflects the proportion of each segmentation fragment in the entire first - mode file. The larger the ratio, the more words the finally generated text will have.
[0127] S425. Extract the highest ratio value in the pixel - point ratio and the corresponding number of controllable sub - texts as the highest number, and determine the second threshold number of the controllable sub - information of other segmentation fragments based on the highest number and the pixel - point ratio.
[0128] S426. Based on the second threshold quantity and the number of pixel points in each segmented fragment, obtain the level of each segmented fragment and store the corresponding controllable sub-text.
[0129] The system extracts the highest proportion value in the pixel point proportion and the corresponding number of controllable sub-texts as the highest quantity. Then, based on this highest quantity and the pixel point proportions of each segmented fragment, calculate the second threshold quantity of the controllable sub-information of other segmented fragments. This can reasonably allocate the text generation space according to the relative importance and scale of the segmented fragments. Suppose a large event poster image containing multiple elements is processed and segmented into 5 fragments A, B, C, D, and E. The number of pixel points in each fragment is counted, and the pixel point proportion is calculated. For example, fragment A has 20,000 pixel points, accounting for 0.2. After comparison, the pixel point proportion of fragment C is the highest at 0.3, and the number of its controllable sub-texts, 300 words, is set as the highest quantity. According to the formula "Second threshold quantity = Highest quantity × Pixel point proportion of this segmented fragment ÷ Highest proportion value", calculate that the second threshold quantity of fragment A is 200 words, fragment B is 150 words, fragment D is 100 words, and fragment E is 250 words, so as to reasonably allocate the text space according to the relative importance and scale of each fragment.
[0130] Among them, the obtaining the level of each segmented fragment and storing the corresponding controllable sub-text based on the second threshold quantity and the number of pixel points in each segmented fragment includes:
[0131] S4261. If the number of controllable sub-texts of the segmented fragment is less than or equal to the text second threshold quantity, save it.
[0132] If the number of controllable sub-texts of the segmented fragment is less than or equal to the text second threshold quantity, it means that the text space meets the requirements calculated according to the proportion, and the system will directly save this controllable sub-text.
[0133] S4262. If the number of controllable sub-texts is greater than the second threshold quantity, input the controllable sub-text into the text simplification model for text reduction processing until it is less than or equal to the text second threshold quantity and then save it.
[0134] If the number of controllable sub-texts is greater than the second threshold quantity, the system will input the controllable sub-text into the text simplification model for text reduction processing until its quantity is less than or equal to the text second threshold quantity and then save it.
[0135] S4263. Generate the corresponding level according to the number of pixel points in each segmented fragment, and store the level and controllable sub-text of each segmented fragment correspondingly.
[0136] Similarly, corresponding levels are generated according to the number of pixel points in each segmentation segment, and the levels and controllable sub-texts are stored correspondingly.
[0137] S43. Statistically analyze the levels and controllable sub-texts of the segmentation segments and fill them into the preset area of the output template to obtain the multi-level controllable text of the first modal file.
[0138] After completing the processing and storage of the levels and controllable sub-texts of each segmentation segment, the system will statistically analyze the levels and controllable sub-texts of all segmentation segments. Then, these information will be filled into the preset area of the output template according to certain rules. The output template is pre-designed to standardize the presentation format of the multi-level controllable text. By filling the information of each segmentation segment into the template in an orderly manner, the multi-level controllable text of the first modal file is finally obtained.
[0139] Through the above steps, the present invention realizes the in-depth content mining and accurate text generation of the first modal file. From content extraction and conversion to fine processing based on controllable sub-information, and then to the final text integration and output, the whole process fully considers the diverse needs of users and can flexibly adjust the text generation strategy according to different controllable sub-information settings. The generated multi-level controllable text not only accurately reflects each part of the first modal file in terms of content, but also has good hierarchy and rationality in terms of structure and length, greatly improving the quality and practicality of text generation and providing users with a text output result that better meets their expectations.
[0140] It can be understood that the above embodiments are for processing the first modal file with image attributes. In some other embodiments, the present solution can also process video attributes. The generative large model processes the attributes of the first modal file to be analyzed and generates a corresponding parsing controllable template, including:
[0141] If the attributes of the first modal file include video attributes, a video interaction preview box corresponding to the video attributes is generated, and the video interaction preview box includes a video interaction axis.
[0142] This embodiment expands the application scope of the method for generating controllable text of the generative large model of the present invention from processing the first modal file with image attributes to processing the video attribute file. By generating a specific parsing controllable template for video attributes and realizing interaction with users to determine controllable information, this method can more comprehensively process different types of multimedia data, laying a foundation for subsequent video content segmentation and controllable text generation, and further enhancing the generality and practicality of the invention.
[0143] In this embodiment, when the attributes of the first modal file to be analyzed include video attributes, the generative large model generates a video interaction preview box corresponding to the video attributes. This preview box contains a video interaction axis, which can be located below the video interaction preview box. The video interaction axis visually presents the time dimension of the video, providing a user interface basis for operating and selecting video time points.
[0144] Based on the specifications in the attributes of the first modal file, the interaction preview box is adaptively processed so that the interaction preview box fits and corresponds to the first modal file.
[0145] Based on the specifications in the attributes of the first modal file, such as the video window size, the interaction preview box is adaptively processed so that the interaction preview box wraps the video window. This process is similar to the processing of the image interaction preview box, aiming to make the interaction preview box fit and correspond to the video content, ensuring that the video display effect seen by the user in the preview box is accurate and complete, and providing a good visual basis for subsequent time point selection operations.
[0146] After extracting the video interaction axis to generate the corresponding time point selection line, a corresponding parsing controllable template is generated.
[0147] Among them, after extracting the video interaction axis to generate the corresponding time point selection line, generating the corresponding parsing controllable template includes:
[0148] Starting from the 0 point on the video interaction axis, a time point is determined at intervals of a preset time period and a corresponding time point selection line is generated.
[0149] Extract the video interaction axis, starting from the 0 point as the starting point, determine the time points according to the preset time period and generate the corresponding time point selection lines. For example, if the preset time period is 1 minute, then a time point will be determined every 1 minute on the video interaction axis and a time point selection line will be generated. These time point selection lines constitute the time segmentation markers of the video. The user can click on these selection lines to specify the video time period of interest, and finally form a parsing controllable template.
[0150] Among them, interacting with the parsing controllable template to determine the controllable information of the controllable text of the first modal file includes:
[0151] After determining that the user has selected any time point selection line, receive the second mark added by the user to the time point selection line, and generate a second controllable slot and a second confirmation slot corresponding to the second mark at the parsing controllable template.
[0152] When the system determines that the user has selected any one of the time - point selection lines, it will receive the second mark added by the user to this time - point selection line. The second mark is used to distinguish different video time periods. For example, the user may add marks such as "1", "2", etc. At the same time, at the parsing controllable template, a second controllable slot and a second confirmation slot corresponding to the second mark are generated. The second controllable slot is for the user to input specific requirements for this time period, such as the number of text characters, key description content, etc.; the second confirmation slot is for the user to confirm the input information.
[0153] After determining that the second confirmation slot is triggered, the minimum moment and the maximum moment with the same second mark are counted to obtain the time period.
[0154] When the user triggers the second confirmation slot, the system will count the minimum moment and the maximum moment corresponding to all the time - point selection lines with the same second mark, so as to determine a complete time period. For example, if the user adds the mark "1" to 2 time - point selection lines with an interval, the system will find the earliest moment and the latest moment among these marked time points to determine a complete time period.
[0155] Receive controllable sub - information about the corresponding time period based on the second controllable slot. The controllable information at least includes the text requirements of the controllable text.
[0156] Receive controllable sub - information about the corresponding time period based on the second controllable slot. The controllable information at least includes the text requirements of the controllable text, such as the word - count range of the text, content focus, etc. These information will guide the subsequent text generation of the video content in this time period.
[0157] It is worth emphasizing that after obtaining the controllable information, text data can be generated in combination with the video content of the corresponding time period. Taking a 10 - second time period as an example, if the video content in this period shows a dog playing football, then the generated text data will revolve around this scene. It should be noted that for the technology of identifying video content such as "a dog playing football", this solution adopts existing mature technologies and does not make specific limitations and improvements. The core highlight of this solution is that after completing video segmentation processing, it can accurately meet diverse customized text generation requirements according to the characteristics of different time periods.
[0158] In the above - mentioned embodiment, it further includes:
[0159] If it is determined that the user selects a frame image at any moment in the video interaction axis and inputs the image attributes, then steps A1 to A3 are executed.
[0160] If it is determined that the user selects a frame image at any moment in the video interaction axis and inputs image attributes, the system will execute steps A1 to A3 for image attribute processing before. That is, an image interaction preview box corresponding to the frame image is generated, the preview box is adaptively processed to fit the image, then the preview box specifications are extracted to generate a coordinate selection line, and finally an analysis and controllable template for the frame image is generated, organically combining video processing and image processing to achieve more flexible content processing.
[0161] The present invention also provides a storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it is used to implement the methods provided by the above various embodiments.
[0162] Among them, the storage medium can be a computer storage medium or a communication medium. The communication medium includes any medium convenient for transmitting a computer program from one place to another. The computer storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer. For example, the storage medium is coupled to the processor, so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an application specific integrated circuit (ASIC). In addition, the ASIC can be located in the user equipment. Of course, the processor and the storage medium can also exist as discrete components in the communication device. The storage medium can be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0163] The present invention also provides a program product, which includes execution instructions stored in the storage medium. At least one processor of the device can read the execution instructions from the storage medium, and the execution of the execution instructions by at least one processor enables the device to implement the methods provided by the above various embodiments.
[0164] In the above embodiments of the terminal or the server, it should be understood that the processor may be a central processing unit (CPU for short), or other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in conjunction with the present invention may be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0165] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A generative large model controllable text generation method, characterized in that: include: The generative large model processes the attributes of the first modal file to be analyzed and generates a corresponding analytical controllable template, including: A1. If the attributes of the first modal file include image attributes, generating an image interactive preview box corresponding to the image attributes; A2. Adaptively process the interactive preview frame based on the specifications in the first modal file attributes, so that the interactive preview frame fits the first modal file; A3. After extracting the corresponding specifications of the interactive preview frame and generating the corresponding coordinate selection lines, a corresponding analytical controllable template is generated, including: Determine the target horizontal frame and the target vertical frame in the interactive preview frame, and take the intersection of the target horizontal frame and the target vertical frame as the starting point; Taking the starting point as the initial point, corresponding coordinate lines are established according to the target horizontal frame and the target vertical frame, and the end points of the target horizontal frame and the target vertical frame are the end points; Calculate the coordinates of the unit number of pixels based on the number of pixels between the initial point and the cutoff point, select the pixel points corresponding to the preset coordinate values and add the coordinate selection lines to obtain the analytically controllable template; interacting with the parsed controllable template to determine controllable information of the controllable text of the first modal file; Based on the controllable information, the first modal file is segmented according to corresponding attributes to obtain a plurality of segmented segments and controllable sub-information of each segmented segment; Content extraction and conversion processing are performed on each segmented segment to generate a corresponding initial text, and the initial text is processed based on the controllable sub-information to obtain a multi-level controllable text of the first modal file.
2. The method according to claim 1, characterized in that The adaptive processing of the interactive preview frame based on the specification in the first modal file attribute so that the interactive preview frame is aligned with the first modal file includes: Obtain the number of horizontal pixels and the number of vertical pixels of the first modal file; The interactive preview frame is adaptively processed based on the specifications of the number of horizontal pixels and the number of vertical pixels to make the interactive preview frame fit the boundary of the first modal file.
3. The method according to claim 1, characterized in that The interacting with the parsed controllable template to determine the controllable information of the controllable text of the first modality file includes: After determining that the user has selected any one of the coordinate selection lines, receiving a first mark added by the user to the coordinate selection line, and generating a first controllable slot and a first confirmation slot corresponding to the first mark at the parsed controllable template; After determining that the first confirmation slot is triggered, counting the first closed area formed by all coordinate selection lines having the same first mark; If it is determined that the user has selected a first closed area, it is used as a pre-divided area, and controllable sub-information about the pre-divided area is received based on the first controllable slot, and the controllable information at least includes text requirements of the controllable text.
4. The method according to claim 3, characterized in that The first modal file is segmented and processed according to corresponding attributes based on the controllable information to obtain a plurality of segmented segments and controllable sub-information of each segmented segment, including: Segmenting the first modal file according to the pre-segmented regions to obtain a plurality of segmented segments; Controllable sub-information of each segmented segment is extracted, wherein the controllable sub-information includes at least one of a text threshold quantity and a text ratio.
5. The method according to claim 1, characterized in that The extracting and converting of content from each segmented segment generates a corresponding initial text, and the processing of the initial text based on the controllable sub-information obtains a multi-level controllable text of the first modal file, including: Extracting the content of each segmented segment to obtain first text content and / or image content, converting the image content into a preset second text content, wherein the initial text includes the first text content and the second text content; Based on the text threshold number and text ratio corresponding to the controllable sub-information, the initial text is fused to obtain the hierarchy and controllable sub-text of each segmented segment; The levels and controllable sub-texts of the segmented segments are counted and filled into a preset area of the output template to obtain multi-level controllable texts of the first modal file.
6. The method according to claim 5, characterized in that The initial text is fused based on the text threshold quantity and text ratio corresponding to the controllable sub-information to obtain the level and controllable sub-text of each segmented segment, including: If it is determined that the text threshold number of controllable sub-information is set without setting the text ratio, and the number of controllable sub-texts is less than or equal to the first text threshold number, then save; If the number of controllable subtexts is greater than the first threshold number, the controllable subtexts are input into the text simplification model to be reduced to a number less than or equal to the first threshold number of texts and then saved; A corresponding level is generated according to the number of pixels in each segmented segment, and the level and controllable sub-text of each segmented segment are stored correspondingly.
7. The method according to claim 5, characterized in that The initial text is fused based on the text threshold quantity and text ratio corresponding to the controllable sub-information to obtain the level and controllable sub-text of each segmented segment, including: If it is determined that the controllable sub-information does not set the text threshold number and the text ratio, the pixel ratio of all the segmented segments is calculated according to the number of pixels in each segmented segment; Extracting the highest ratio value in the pixel point ratio and the corresponding controllable sub-text quantity as the highest quantity, and determining the second threshold quantity of controllable sub-information of other segmented segments based on the highest quantity and the pixel point ratio; Based on the second threshold number and the number of pixels in each segmented segment, the level of each segmented segment and the corresponding storage of the controllable sub-text are obtained.
8. The method according to claim 7, characterized in that The step of obtaining the level of each segmented segment and the corresponding storage of the controllable subtext based on the second threshold number and the number of pixels in each segmented segment includes: If the number of controllable sub-texts of the segmented fragment is less than or equal to the second threshold number of texts, save; If the number of controllable subtexts is greater than the second threshold number, the controllable subtexts are input into the text simplification model to be reduced to a number less than or equal to the second threshold number of texts and then saved; A corresponding level is generated according to the number of pixels in each segmented segment, and the level and controllable sub-text of each segmented segment are stored correspondingly.
9. The method according to claim 1, characterized in that: The generative large model processes the attributes of the first modal file to be analyzed and generates a corresponding analytical controllable template, including: If the attributes of the first modal file include video attributes, generating a video interaction preview frame corresponding to the video attributes, wherein the video interaction preview frame includes a video interaction axis; Adaptively process the interactive preview frame based on the specifications in the first modal file attributes, so that the interactive preview frame fits and corresponds to the first modal file; After extracting the video interaction axis to generate the corresponding time point selection line, a corresponding analytically controllable template is generated.
10. The method according to claim 9, characterized in that After extracting the video interaction axis to generate the corresponding time point selection line, a corresponding analytical controllable template is generated, including: The video interaction axis takes point 0 as the starting point, determines a time point at a preset time interval, and generates a corresponding time point selection line.
11. The method according to claim 9, characterized in that The interacting with the parsed controllable template to determine the controllable information of the controllable text of the first modality file includes: After determining that the user selects any time point selection line, receiving a second mark added by the user to the time point selection line, and generating a second controllable slot and a second confirmation slot corresponding to the second mark at the parsed controllable template; After determining that the second confirmation slot is triggered, the minimum time and the maximum time with the same second mark are counted to obtain a time period; The controllable sub-information about the corresponding time period is received based on the first controllable slot, and the controllable information at least includes text requirements of the controllable text.
12. The method according to claim 11, characterized in that Also includes: If it is determined that the user selects a frame image at any time in the video interaction axis and inputs image attributes, steps A1 to A3 are executed.
Citation Information
Patent Citations
Image description text determination method and related equipment thereof
CN114021646A